Skip to content
decosa

Apple Silicon · self-host

Run it on a Mac Studio.

32 of 86 tools run entirely on an Apple Silicon Mac, natively on MLX: the same open models, no NVIDIA GPU, and no cloud call. 25 of them passed their own end-to-end smoke check on a Mac Studio, and the live scribe keeps up with speech. One command sets it up.

Measured, not estimated

Each row is the tool’s own sample run end to end against a local decosa-api, with every model call signed by the Mac’s key. A Blackwell card is faster, by 1.4x on the long pipelines to several times on short checks. A Mac Studio is a desk, not a data centre: one user, one job at a time.

  • Visit copilotfirst captionMac1.5 sRTX PRO 60001.4 s
  • Visit copilotdiarize a 123 s visitMac20 sRTX PRO 60009 s
  • Visit copilotstop to final note, with pass 2Mac81-88 sRTX PRO 600044-75 s
  • Grounding check5 judged sentences, signed reportMac8.6 sRTX PRO 60000.9-3.0 s (gateway, quiet)
  • Deposition and hearing digestfictional pair, 44 calls, 28k tokensMac105 sRTX PRO 600034-76 s (gateway)
  • Filing pre-flightdemo brief, 15 calls, 47k tokens, with public case lookupsMac145 sRTX PRO 600028-95 s (gateway)
  • Typed-judgment APIsupport-ticket sample, method single, 7 callsMac12.5 sRTX PRO 600035 s for the 35-call samples variant (gateway, loaded)
  • Privilege review and privilege log12-email sample, 10 callsMac23 sRTX PRO 600017 s self-hosted (58 calls)
  • Decosa StudioMusic3, 30 s clip, 30 steps, with loadMac102 sRTX PRO 600043-73 s (45 s clip)
  • Decosa Studioimage, Z-Image-Turbo 1024x1024, 9 stepsMac32 sRTX PRO 600099 s for Qwen-Image, 50 steps
  • Decosa StudioLTX-2.5 distilled, 5 s 1280x704 with audio, with loadMac184 sRTX PRO 600063-65 s

Quiet GPU for the numbers shown. A first pass with another MLX job and a graphics app on the same GPU was 1.5-2x slower. Blackwell figures come from each use case's stack.json and several were on the shared gateway route.

Which Mac

  1. Any Apple Silicon MacAny Mac

    Auditing an endpoint needs no model at all.

    1 tool starts here

  2. 32 GB or more32 GB

    Text tools: Qwen3.8-27B in 4-bit is 16 GB, and the server sits near 18 GB. One job at a time.

    27 tools start here

  3. 48 GB or more48 GB

    Live voice: the language model, Voxtral and the diarizer together measured 27 GB at rest.

    7 tools start here

  4. 64 GB or more64 GB and up

    Studio music and images next to the text model, or long filings with headroom.

    2 tools start here

Only the M3 Ultra was measured; the smaller tiers come from the measured memory footprints. Decode speed follows memory bandwidth, so a Max chip is slower than an Ultra and a Pro slower again. With 192 GB or more, the best tier’s DeepSeek-V4-Flash (about 156 GB in MXFP4 for MLX) fits beside the rest; the setup script does not wire it in yet.

One command

Docker on macOS cannot reach the GPU, so the Mac kit skips containers for the models. They run on the host through MLX, and decosa-api runs from the checkout with uv.

  • Checks the chip and memory, and installs uv with Homebrew if needed.
  • Downloads the MLX weights: Qwen3.8-27B 4-bit (16 GB); with the live profile, Voxtral Mini 4B Realtime and MOSS-Transcribe-Diarize (5 GB).
  • Starts the model servers and the API on 127.0.0.1 and mints a local API key.
  • Names the exact MLX weights in every receipt, with a hash of the downloaded files.

Each tool’s Self-host tab also has a Mac prompt for your coding agent, for example grounding-mac.md.

git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> && cd decosa-api
scripts/mac/setup.sh                   # text tools
scripts/mac/setup.sh --profile live    # + live speech and diarization

Faster single answers: add --engine omlx for oMLX with multi-token prediction (61 tok/s against 31.8). It does not help the pipelines that make many calls at once, and it returns no log-probabilities.

6 live-voice tools (Visit copilot, Field reports, Live translation, Tamper-evident record, Clinical AI assurance monitor, Structured oral assessment) need --profile live.

Every tool

Each part of a tool’s standard stack maps to a Mac equivalent. “Runs on a Mac Studio” means every part runs on the Mac; the build checks it against each stack. “Runs, measured” means that part was run on the M3 Ultra in at least one tool; rows marked “Measured on the M3 Ultra” have numbers in the table above.

GPU part, Mac part

  • Hy-MT2-7B (language pack)Untested on a Mac
    NVIDIA GPU
    BF16 on vLLM 0.29 (Transformers backend, eager)
    Mac
    Not run on a Mac; the vendor publishes a GGUF build (tencent/Hy-MT2-7B-GGUF) for llama.cpp, untested here
  • decosa-note-detail-modernbert-large (M17 detail checker)Runs, not measured
    NVIDIA GPU
    PyTorch on CPU (8 threads)
    Mac
    The same service with PyTorch on CPU (about 2.5 GB of RAM); expected to run, not timed on a Mac
  • decosa-api moduleRuns, measured
    NVIDIA GPU
    Python on CPU
    Mac
    The same Python module, run with uv
  • Qwen3.8-27BRuns, measured
    NVIDIA GPU
    NVFP4 on vLLM 0.29 (Blackwell)
    Mac
    MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option

    31.8 tok/s decode on an M3 Ultra with mlx_lm.server, 61 tok/s with oMLX and MTP; about 132 tok/s on one RTX PRO 6000.

  • Qwen3.5-4B-BaseRuns, not measured
    NVIDIA GPU
    BF16 on vLLM
    Mac
    MLX (mlx_lm.server)

    Only needed to record a new auditor reference. A reference recorded on MLX describes the MLX stack, not vLLM.

  • Voxtral Mini 4B RealtimeRuns, measured
    NVIDIA GPU
    vLLM realtime WebSocket
    Mac
    MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py

    Captions keep up with speech: first caption 1.5 s, median gap 0.4 s.

  • MOSS-Transcribe-Diarize 0.9BRuns, measured
    NVIDIA GPU
    transformers on CUDA
    Mac
    MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py

    A 123 s visit diarized in 20 s (9 s on the GPU).

  • c2pa-pythonRuns, not measured
    NVIDIA GPU
    CPU
    Mac
    Native macOS arm64 wheel

    CPU only; PyPI ships a macOS arm64 build. Not run on the Mac here.

  • TrustMarkRuns, not measured
    NVIDIA GPU
    PyTorch on CPU
    Mac
    PyTorch on CPU

    Runs on CPU on the GPU stack too. Not run on the Mac here.

  • Playwright + ChromiumRuns, measured
    NVIDIA GPU
    CPU
    Mac
    Playwright's macOS arm64 Chromium

    The flight recorder and test runs smoke checks passed with Playwright's macOS Chromium.

  • GROBID 0.8.2Untested on a Mac
    NVIDIA GPU
    Java service on CPU (Docker)
    Mac
    Docker Desktop, amd64 image under emulation

    The published image is amd64 only, so an Apple Silicon Mac runs it emulated. Not tried.

  • Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (evidence retrieval block)Runs, measured
    NVIDIA GPU
    transformers on CUDA (services/retrieval), about 12 GB
    Mac
    The same service with PyTorch on MPS (bf16): same scores within rounding, about 12 GB, rerank of 40 passages 5.1 s (p50) vs 1.8 s on the GPU
  • all-MiniLM-L6-v2 (ONNX)Runs, not measured
    NVIDIA GPU
    ONNX Runtime on CPU
    Mac
    ONNX Runtime on CPU (macOS arm64 wheels)
  • LAION CLAP + librosaRuns, not measured
    NVIDIA GPU
    PyTorch on CPU
    Mac
    PyTorch on CPU

    CPU only on the GPU stack too. Not run on the Mac here.

  • MiniMax-Music3Runs, measured
    NVIDIA GPU
    ComfyUI on CUDA
    Mac
    Community MLX port (PocketAiHub/MiniMax-Music3-MLX, INT8)

    A 30 s instrumental at 30 steps took 102 s including load, 50 GB peak. Community port, not the ComfyUI path the hosted studio uses.

  • ACE-Step 1.5Untested on a Mac
    NVIDIA GPU
    PyTorch on CUDA
    Mac
    PyTorch on MPS
  • Qwen-Image-2512Untested on a Mac
    NVIDIA GPU
    diffusers on CUDA
    Mac
    mflux 0.20 (mlx-community/Qwen-Image-2512-4bit, 26 GB)

    mflux supports it; not run here for lack of disk. Z-Image-Turbo, the studio's lite image model, took 32 s for 1024x1024 in 9 steps.

  • Kokoro-82MRuns, not measured
    NVIDIA GPU
    kokoro package on CPU
    Mac
    kokoro package on CPU

    The GPU stack already runs it on CPU. Not run on the Mac here.

  • IndexTTS-2.5Untested on a Mac
    NVIDIA GPU
    PyTorch on CUDA
    Mac
    PyTorch on MPS
  • LTX-2.5 22B distilled (video with audio)Runs, measured
    NVIDIA GPU
    ComfyUI on CUDA, int8-convrot transformer and Gemma encoder
    Mac
    MLX int8 (group 64) on ltx-2-mlx 0.15.9, pack converted from Lightricks/LTX-2.5 with mlx-forge (42 GB)

    A 5 s 1280x704 clip with audio (121 frames, 8 + 3 distilled steps) took 184-186 s including load (3 runs, 2026-09-26, a graphics app open on the same GPU), against 63-65 s on an RTX PRO 6000: about 2.9x slower, 37 s per second of video. A same-seed re-render was bit-identical. Self-host only under the LTX-2 Community License, as on the GPU stack.

  • Wan 2.1 / 2.2 videoNeeds a CUDA GPU
    NVIDIA GPU
    ComfyUI on CUDA
    Mac
    None practical

    A 5 s Wan 14B clip takes about 35 minutes on a Blackwell card; ComfyUI on MPS would be far slower.

  • MiniMax H3 Max via falHosted API
    NVIDIA GPU
    fal's hosted API
    Mac
    The same hosted API, called from the Mac

    The render happens on fal, not on your machine.

What is different on a Mac

  • The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route runs. The quality numbers on each Stack tab were measured on NVFP4; rerun a tool’s eval if you rely on them.
  • Receipts are signed by the Mac’s own key and name the MLX weights. There is no gateway countersignature on a self-hosted Mac: it is your attestation, not a proof.
  • Video renders (Wan, LTX, MiniMax-H3) still need a CUDA GPU or a hosted API. So do the Blackwell-only engines.
  • The live speech server is a small shim over mlx-audio that speaks the same WebSocket dialect as the GPU stack, so decosa-api runs unchanged.

Self-host on an NVIDIA GPU Tools that run on a Mac