Skip to content
decosa
LiveHostedSelf-hostMac

Run an open model behind the OpenAI API

Qwen3.8-27B behind a standard OpenAI-style /v1 endpoint: try it hosted, with a signed receipt per call, or run it on your own GPU so nothing leaves the box. Hosted replies are capped at 2,048 tokens and don't take tool calls yet.

Measured478 / 500Coding benchmark, 5 core tasks, self-hosted stack (out of 500)
On production4.3 smedian on production (2026-09-30); slower when the service is busy
List price~$0.046 per 100 requestsmeasured, at list price

Try it live

Live

Qwen3.8-27B on a first-party GPU. Prompts are not retained. Every answer carries a receipt. Do not paste secrets or proprietary code into a public demo.

Try one of these:

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the private code assistant API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
code

Use the hosted API

# Decosa Private code assistant: use the hosted API

You are wiring Decosa's code assistant into this project or editor. It is an OpenAI-compatible chat completions API
serving Qwen3.8-27B through the Decosa API. Every response carries a signed receipt.
Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Model: `qwen3.8-27b`
- Health check: `GET https://api.decosa.ai/healthz` returns `{"ok": true, "asr": bool, "llm": bool, "live_sessions": n, "queue": n}`.

## Auth: API key (or a demo session)
1. Preferred: an API key. Create one on the tool page with "Get an API key"; it looks like `dk_…` and is shown
   only once. Keep it in an environment variable, never in code: `DECOSA_API_KEY=dk_…`. Send
   `Authorization: Bearer $DECOSA_API_KEY` on calls that need auth; WebSockets take `?token=$DECOSA_API_KEY` in the URL.
2. Without a key, use a short demo session: `POST https://api.decosa.ai/demo/session` with JSON `{"vertical": "code"}` returns
   `{"token": "<opaque>", "expires_at": <unix seconds>, "budget": {"seconds_audio": 300, "llm_tokens": 20000}}`.
3. Send `Authorization: Bearer <token>` on calls that need it (session-bound calls such as live audio, replay, chat and
   render jobs). WebSockets take `?token=<token>` in the URL instead. These need no token: `GET /healthz`,
   `GET /demo/scripts`, `GET /demo/recordings`, `GET /demo/recordings/{id}/events`, `GET /studio/gallery`,
   `GET /studio/jobs/{id}`, `GET /receipts/{id}`.
4. Demo-session limits: a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz), and a global cap on concurrent live audio sessions. Over a limit the API answers
   HTTP 429 with a `Retry-After` header (seconds): wait that long, then retry. Reuse a token until `expires_at`.
5. The API keeps no PII; session transcripts live in memory and are deleted when the session ends.

Your API key (or a demo token) is the bearer key for chat completions. Token budget per session: 20,000 LLM tokens.

## Chat completions
`POST https://api.decosa.ai/v1/chat/completions` (OpenAI-compatible). Send `"stream": true` to get SSE.
- `max_tokens` defaults to, and is capped at, 2,048 on the hosted API. A reply that stops there has
  `finish_reason: "length"`: ask it to continue, or self-host for longer outputs.
- Prompts up to 48,000 characters (413 above that). `model` must be `qwen3.8-27b` (404 otherwise).
- Tool calling works on the hosted route: send OpenAI `tools`, the answer comes back as `tool_calls` (finish_reason
  `tool_calls`) with a receipt, and you send the tool result back as a `tool` message. Each call stops at 2,048 generated tokens.
- Bad input gets a 4xx with `{"error": "..."}`; 401 means a missing or revoked key.
- The hosted route computes the whole answer, then streams it, so the text arrives in one burst.
- Each response has the header `x-decosa-receipt: <completion id>`.
- The final SSE chunk includes `"receipt": {"id", "model", "gateway_sig", "provider"}`.

## Receipts (public, read-only)
`GET https://api.decosa.ai/receipts/{id}` returns
`{"id", "model", "weights_root", "request_hash", "output_hash", "provider": {"miner_id", "pubkey", "sig"}, "gateway": {"pubkey", "sig"}, "proof": {"format", "verified"}, "checks": [{"name", "ok", "detail"}]}`.

## Example: TypeScript (`npm i openai`)
```ts
import OpenAI from "openai";

const BASE = "https://api.decosa.ai";
// Prefer your API key (dk_…); fall back to a short demo session.
const token: string =
  process.env.DECOSA_API_KEY ??
  (await fetch(`${BASE}/demo/session`, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ vertical: "code" }),
}).then((r) => r.json())).token;

const client = new OpenAI({ baseURL: `${BASE}/v1`, apiKey: token });
const { data: stream, response } = await client.chat.completions
  .create({ model: "qwen3.8-27b", stream: true, messages: [{ role: "user", content: "Write a retry helper in TypeScript." }] })
  .withResponse();
console.log("receipt:", `${BASE}/receipts/${response.headers.get("x-decosa-receipt")}`);
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
```

## Example: Python (`pip install openai requests`)
```python
import os
import requests
from openai import OpenAI

BASE = "https://api.decosa.ai"
token = os.environ.get("DECOSA_API_KEY") or requests.post(f"{BASE}/demo/session", json={"vertical": "code"}).json()["token"]  # dk_… API key preferred
client = OpenAI(base_url=f"{BASE}/v1", api_key=token)

raw = client.chat.completions.with_raw_response.create(
    model="qwen3.8-27b",
    stream=True,
    messages=[{"role": "user", "content": "Write a retry helper in Python."}],
)
print("receipt:", f"{BASE}/receipts/{raw.headers.get('x-decosa-receipt')}")
for chunk in raw.parse():
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
```

## What to build
1. Configure this project's OpenAI-compatible client (or editor extension) with the base URL `https://api.decosa.ai/v1`,
   model `qwen3.8-27b`, and `DECOSA_API_KEY` (your `dk_…` key) as the API key; without a key, a demo token refreshed
   from `/demo/session` when it expires.
2. On HTTP 429, wait for `Retry-After` seconds and retry.
3. Log or show the receipt link for each completion.

The hosted API is a demo with small budgets. For private repositories, self-host so source never leaves your machines.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Private code assistant: run it yourself (containers)

You are setting up Decosa Private code assistant to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/code.zip (1 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py code` (the api image carries the same bundle under /app/rehearsal/code/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py code --bundle code.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the answer defines merge_intervals", "the answer carries doctests", "the answer finished within the token limit"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"code"}'`
   should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/code-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"code"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py code

Download the mock-data bundle (1 KB, 9 checks)expected.json

One short coding request (a merge_intervals function with a docstring and three doctests) sent to the OpenAI-compatible chat route. The answer must hold the function and its doctests, and the call must carry a receipt that is public at /receipts/{id} and whose checks all pass. There is no verify route for a single completion, so there is no tamper step; the receipt's own checks are the proof.

What the rehearsal checks
  • the answer defines merge_intervals
  • the answer carries doctests
  • the answer finished within the token limit
  • the route reports token usage
  • the answer carries a signed (hosted) or attested (self-hosted) receipt
  • the receipt is public at /receipts/{id}
  • the public receipt lists its checks
  • every check on the receipt passes
  • every model call has a signed receipt

Licence: Synthetic: a coding prompt written for Decosa (the nightly smoke check's prompt). No real code base or people. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Private code assistant: run it yourself (containers)

You are setting up Decosa Private code assistant to run entirely on this machine's NVIDIA GPU(s). Nothing is sent to Decosa's
hosted API and there are no Decosa charges. The local service speaks the same API as the hosted one, so apps built
against the hosted API only need a new base URL.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Hardware: 1x RTX PRO 6000 (96 GB), or 2x RTX 5090 (32 GB each). Linux x86_64 with a recent NVIDIA driver.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/code.zip (1 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py code` (the api image carries the same bundle under /app/rehearsal/code/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py code --bundle code.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the answer defines merge_intervals", "the answer carries doctests", "the answer finished within the token limit"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Check the GPU and driver: `nvidia-smi`. If it fails, stop and tell me; do not install drivers without asking.
   Check free disk: the first start downloads model weights (tens of GB).
2. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Add me to the `docker` group only if I agree.
3. NVIDIA Container Toolkit: if `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the toolkit using
   NVIDIA's official instructions (docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html),
   then run `sudo nvidia-ctk runtime configure --runtime=docker` and `sudo systemctl restart docker`. Re-run the check.
4. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. If it references an `.env` file or variables, ask me for any values. Never print secrets.
5. Pull and start: `docker compose pull && docker compose up -d`.
6. Wait for health. Find the host port that compose.yaml publishes for the API (`docker compose ps`), then poll
   `curl -fsS http://localhost:<PORT>/healthz` every 10 s until it returns `"ok": true` with `"asr": true` and
   `"llm": true`. The first start can take a while as weights download. Show me `docker compose logs --tail=50` if it
   has not come up after 20 minutes.
7. Smoke test: `curl -fsS -X POST http://localhost:<PORT>/demo/session -H 'Content-Type: application/json' -d '{"vertical":"code"}'`
   should return a token.
8. Report back: GPU model(s) and memory, Docker and toolkit versions, the `/healthz` output, and the local base URL.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/code-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsPrivate code assistant on GeForce RTX 5090: use the Standard · one 96 GB card (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 24-32 GB Blackwell card (fits): Not measured here. The NVFP4 weights are about 20 GB, so use a shorter max-model-len and fewer sequences. A third-party report gives about 50 tok/s single-stream with NVFP4 + MTP on a 24 GB Blackwell card.

Standard · one 96 GB card (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Coding model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Private code assistant, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Private code assistant on my hardware

Fetch https://decosa.ai/prompts/code-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=code)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one 96 GB card (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Coding model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/code-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Private code assistant: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Private code assistant on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/code.zip (1 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py code` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the answer defines merge_intervals", "the answer carries doctests", "the answer finished within the token limit"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Coding model (chat, edit, agent tool calls) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py code`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Single-stream chat is where oMLX with MTP helps most (61 tok/s).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 4.3 s · ~<$0.001 per run · 1 receipt

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers

Measured cost to run: about $0.046 per 100 requests (hosted, 30 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: the fallback llm image builds, the compose file validates, the api starts and the sample passes end to end (attested receipts, p50 1.5 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Tool calling was not re-verified: the local server ran without the tool-parser flags.

Known limits (5)
  • Hosted answers stop at 2,048 generated tokens (finish_reason "length"); self-host for longer outputs.
  • Tool calling works on the hosted route (checked on production 2026-09-30: a request with `tools` returns `tool_calls` and a receipt, and the follow-up turn with the tool result is answered). Each call is still one request of at most 2,048 generated tokens.
  • The hosted route computes the whole answer before streaming it, so text arrives in one burst.
  • Hosted p50 is for an answer of about 340 tokens (7 calls on 2026-09-30, fastest 2.5 s, slowest 5.1 s); under load on 2026-09-25 the same call took up to 37 s.
  • Token counts on hosted receipts are the gateway's metering, which on 2026-09-25 overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the private code assistant API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB), or 2× RTX 5090.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

An OpenAI-compatible coding model on your own GPU, so your source code never leaves the machine.

One open-weight coding model, Qwen3.8-27B, runs on vLLM with NVFP4 weights and its built-in MTP draft head, behind a standard /v1 endpoint. Point OpenCode, Continue, Codex-style CLIs or any OpenAI-SDK tool at it. Self-hosted, prompts and code stay on the box. On the hosted demo, every completion comes back with a gateway-signed receipt that anyone can re-check. It suits teams that want an agentic coding assistant without sending their repository to a third-party API.

Deployment
Hosted or self-host
Regulatory
Self-hosted, source code and prompts stay on your machine. The hosted demo sends prompts to a shared GPU and our gateway, so don't paste proprietary code there. Model licences: Apache-2.0 (Qwen3.8-27B), MIT (optional DeepSeek-V4-Flash tier).
Architecture
Text description

A coding agent or IDE on the developer's machine (OpenCode, Continue, a Codex-style CLI) sends OpenAI-style requests either to decosa-api on port 8445, which adds sessions, budgets and receipts, or straight to vLLM on port 8000 (no sessions, budgets or receipts). vLLM 0.29.0 serves Qwen3.8-27B (Apache-2.0) with NVFP4 weights, an MTP draft head proposing 3 tokens, and an FP8 KV cache on one RTX PRO 6000. Everything in this box stays local when self-hosted. An optional larger tier runs DeepSeek-V4-Flash (MIT) across two GPUs. On the receipt path, decosa-api's gateway route sends the request to our gateway's metered route; the completion comes back with a signed receipt (commitment hash, metering proofs, the serving machine's attestation), the gateway countersigns it with its Ed25519 key, the id is returned in the x-decosa-receipt header and at GET /receipts/{id}.

Architecture

At a glance

Data retention
Hosted: prompts and answers are not stored; only the receipt (hashes, token counts, cost) is kept. Self-hosted: nothing leaves the machine unless you turn on the network profile.
What leaves the box (self-host)
Nothing. vLLM is bound to 127.0.0.1 and the optional api service signs receipts with a local key.
Inputs
OpenAI-compatible chat messages, up to 48,000 characters per request on the hosted API; model qwen3.8-27b.
Typical hosted cost
A fraction of a cent per short answer at the gateway's list price (measured). Each run shows its own measured cost.
Hosted limits
Demo session: 20,000 generated tokens. API key: 200,000 generated tokens a day, 60 requests a minute.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    runs on one 32 GB Blackwell card

    Same Qwen3.8-27B NVFP4 weights and engine as Standard. You give up context length and concurrent sessions, not model quality.

    Models
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX 5090 32 GB (Blackwell, NVFP4)
    Quality evidence
    • Coding benchmark, 5 core tasks, direct API (/500)478 (same weights; measured on a 96 GB card, single run)coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16
    • On a 32 GB card specificallynot measured yetnot measured yet
    Latency
    estimate: similar single-stream speed to Standard (decode is memory-bound on the same weights), but with a much smaller KV cache, so plan for a shorter max-model-len and one or two sessions.
    Verification
    Proof: strongSelf-host onlyQwen3.8-27B is the hosted model qwen3.8-27b. Self-hosted with the direct route there are no receipts.
  • In the hosted demo

    Standard

    one 96 GB card (hosted demo)

    Qwen3.8-27B NVFP4 + MTP3 on the pinned vLLM 0.29.0 with a full 262k context and room for about 32 concurrent streams. This is the hosted demo.

    Models
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Coding benchmark, 5 core tasks, direct API (/500)478 (vLLM NVFP4 + MTP, single run)coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16
    • Same weights on other engines (/500)vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (engine bugs, not the model)coding-agent-bench results/RESULTS.md
    • Code smoke on the exact pinned stack12/12Decosa model benchmarks (Sep 2026)
    Latency
    measured on our server: time to first token in milliseconds and fast single-stream decoding locally. The hosted route's end-to-end time is the measured figure on this page.
    Verification
    Proof: strongEvery hosted completion carries a gateway-signed receipt (commitment hash, Ed25519 countersignature, metering proofs).
  • Best

    two 96 GB cards

    DeepSeek-V4-Flash (284B MoE, MIT) across both GPUs for longer agentic runs. It takes the whole box. The 5-task suite is near its ceiling, so it doesn't clearly beat Standard one-shot; its edge shows with an agent harness.

    Models
    • DeepSeek-V4-Flash (NVFP4)
    Hardware
    2x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • Coding benchmark, 5 core tasks, agent harness (/500)495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses); single runscoding-agent-bench results/RESULTS.md, official_prime_results.json and ds4_*_results.json, 2026-08-15/16
    • Same suite, direct API (/500)459 uncapped (single run)coding-agent-bench results/RESULTS.md, ds4_api_uncapped_results.json
    • Caveatnot measured yetcoding-agent-bench results/RESULTS.md
    Latency
    measured by the dsv4-flash-nvfp4-sm120 kit: 150.6 tok/s single-stream with MTP, about 328 tok/s aggregate at 4 streams.
    Verification
    No proof yetSelf-host onlyNone today: DeepSeek-V4-Flash is not a hosted model, so this tier produces no receipts. It becomes strong (text LLM coverage) once it is hosted with a pinned engine.
  • Needs more compute

    Wanted: the best setup

    the two most-used open coding models

    GLM-5.3-Flash and DeepSeek-V4.1-Flash, served by network providers. Neither fits one workstation card: GLM needs two 96 GB cards (its sm_120 runtime hangs on our server today) and V4.1-Flash a datacentre node.

    Models
    • GLM-5.3-Flash
    • DeepSeek-V4.1-Flash
    Hardware
    Network providers: 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash; an 8x H200-class node for DeepSeek-V4.1-Flash. Estimate.
    Quality evidence
    • Coding benchmark, 4 tiebreaker tasks, Prime Agent harness (/400)372.3 three-run mean (389, 359, 369), community TR3 4bpw build, 8k thinking budgetcoding-agent-bench results/RESULTS.md, h2h/glm53_tr3_t8k_tb_r1-3.json, 2026-08-28
    • DeepSeek-V4.1-Flash on the same benchmarknot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot hosted yet, so no receipts today.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Coding model (chat, edit, agent tool calls)Qwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strong
Optional larger tier (two GPUs)DeepSeek-V4-Flash (NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active)No proof yet
Network-hosted flash model (wanted)GLM-5.3-Flashnvidia/GLM-5.3-Flash-NVFP4 on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yet
Coding model, datacentre classDeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab)
552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • Terminal coding agent; add the endpoint as an @ai-sdk/openai-compatible provider in opencode.json.

  • IDE extension (VS Code, JetBrains); add a model with provider: openai and apiBase pointing at the endpoint.

  • Terminal coding agent; add a custom model_providers entry in ~/.codex/config.toml pointing at vLLM's /v1 (it serves /v1/responses and /v1/chat/completions).

Services

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM 0.29.0 serving Qwen3.8-27B NVFP4 + MTP3 as qwen3.8-27b. /v1/chat/completions, /v1/responses, /v1/models. Publish on 127.0.0.1 only; with --enable-auto-tool-choice --tool-call-parser qwen3_coder it serves agent tool calls.

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    OpenAI-compatible /v1/chat/completions proxy with demo-session auth, budgets and receipts (x-decosa-receipt header, GET /receipts/{id}). Tool calling works on this route (OpenAI `tools`; the answer comes back as `tool_calls` with a receipt); max_tokens capped at 2048, prompts at 48,000 characters.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: weights 19.9 GiB, KV space for 1.76M tokens at 262k max length, 132 tok/s single-stream, about 1,170-1,520 tok/s at 32 streams. This is the pinned stack.

  • 24-32 GB Blackwell card Fits

    Not measured here. The NVFP4 weights are about 20 GB, so use a shorter max-model-len and fewer sequences. A third-party report gives about 50 tok/s single-stream with NVFP4 + MTP on a 24 GB Blackwell card.

  • Hopper or Ada 48-80 GB (no NVFP4) Fits

    Use the official FP8 checkpoint Qwen/Qwen3.8-27B-FP8 (Apache-2.0, 29 GB). Measured on our server at 49 tok/s single-stream without MTP, 90-94 with MTP3; not measured on Hopper or Ada.

  • 2x RTX PRO 6000 Blackwell 96 GB Fits

    Optional DeepSeek-V4-Flash tier, TP2 on both cards with the patched B12X build from the dsv4-flash-nvfp4-sm120 kit. It replaces the Qwen tier on that box.

  • Cards under 24 GB Does not fit

    The NVFP4 weights alone are about 20 GB, leaving no useful KV cache.

Latency per lane

  • time to first token, local vLLM, 1 stream, ~539-token prompt92 ms

    Measuredmeasured on our server 2026-09-23 (NVFP4 + MTP3 + FP8 KV, benchmark page 17)

  • time per output token, local vLLM, 1 stream7.4 ms

    Measuredmeasured on our server 2026-09-23 (benchmark page 17, ~132 tok/s)

  • time to first token, local vLLM, 32 streams, ~2k-token prompts504 ms

    Measuredmeasured on our server 2026-09-23 (benchmark page 17, median; p99 6,510 ms)

  • hosted /v1/chat/completions, ~350 output tokens, gateway route with signed receipt4.3 s

    Measuredmeasured on production 2026-09-30 (7 calls through api.decosa.ai, median; 2.5-5.1 s end to end for about 340 output tokens; hosted streaming arrives all at once, so TTFT equals this)

  • DeepSeek-V4-Flash tier, 1 stream decode6.6 ms

    Measuredderived from the dsv4-flash-nvfp4-sm120 kit's measurement on 2x RTX PRO 6000 (150.6 tok/s single-stream with MTP)

Notes

  • Qwen3.8-Flash-Next is excluded: its Qwen Community License 1.0 requires a separate licence from Qwen for any Model-as-a-Service business, which covers serving it through the network. Its uncensored and abliterated derivatives carry the same restriction.
  • Tool calling for agents uses vLLM's qwen3_coder parser from the NVIDIA model card recipe; that flag set has not yet been tested on the pinned stack.
  • The hosted route re-emits a finished completion as SSE, so tokens don't arrive progressively there. Self-hosted vLLM streams normally.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

code/assemble-prompt.md174 lines
You are setting up a **private code assistant** on this Linux machine: Qwen3.8-27B (Apache-2.0) on vLLM 0.29.0, with NVFP4 weights, its built-in MTP draft head (3 tokens) and an FP8 KV cache, behind an OpenAI-compatible `/v1` endpoint. Coding agents on this machine will use it, and source code must never leave the box. Work step by step, show me each command before running anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/code.zip (1 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py code` (the api image carries the same bundle under /app/rehearsal/code/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py code --bundle code.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the answer defines merge_intervals", "the answer carries doctests", "the answer finished within the token limit"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 1. Check the GPU, driver and Docker
- Run `nvidia-smi`. The pinned stack needs a Blackwell GPU (RTX PRO 6000, B200, RTX 50-series) for NVFP4, and a driver of 580 or newer (the images were built against it). A 96 GB card is the measured target. 24–32 GB Blackwell cards can work with a shorter context, but that hasn't been measured.
- On Hopper or Ada (no NVFP4), use `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main` below. That is the official FP8 checkpoint, Apache-2.0, 29 GB.
- If `docker` is missing, install Docker Engine from the official docs (https://docs.docker.com/engine/install/). If `docker run --rm --gpus all ubuntu nvidia-smi` fails, install the NVIDIA Container Toolkit (https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html), then run `sudo nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker` and repeat the check.
- Make sure there are about 60 GB free for images and weights.

## 2. Images and weights (pinned)
- `${DECOSA_REGISTRY}/decosa-llm:0.1.0` is vLLM 0.29.0 serving the model. `${DECOSA_REGISTRY}/decosa-api:0.1.0` is an optional proxy that adds sessions, budgets and receipt ids. **Both are publishing soon.** Try `docker pull` first.
- Fallback for `decosa-llm`: it's a thin layer over the pinned vLLM image, so build it locally.
  ```bash
  mkdir -p ~/decosa-code/llm && cat > ~/decosa-code/llm/Dockerfile <<'EOF'
  FROM vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
  ENTRYPOINT ["vllm", "serve"]
  EOF
  ```
- Fallback for `decosa-api`: build it from source (`docker build -f docker/api/Dockerfile .`) once its repo is published. Until then, skip the `api` service. Coding agents only need vLLM.
- Weights: `nvidia/Qwen3.8-27B-NVFP4` at revision `482ca0f3832238542f8f5295dde86b5f22711d80` (about 21 GB). They download into the `hf-cache` volume on first start.

## 3. Write `~/decosa-code/docker-compose.yml`
```yaml
name: decosa-code
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:0.1.0
    build: ./llm                      # used only if the image can't be pulled
    ipc: host
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices: [{ driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] }]
    environment: { HF_TOKEN: "${HF_TOKEN:-}" }
    volumes: [hf-cache:/root/.cache/huggingface]
    ports: ["127.0.0.1:8000:8000"]    # localhost only: nothing on the network can reach the model
    command:
      - ${LLM_MODEL:-nvidia/Qwen3.8-27B-NVFP4}
      - --revision
      - ${LLM_REVISION:-482ca0f3832238542f8f5295dde86b5f22711d80}
      - --served-model-name
      - qwen3.8-27b
      - --language-model-only
      - --max-model-len
      - "${LLM_MAX_LEN:-262144}"
      - --gpu-memory-utilization
      - "${LLM_GPU_UTIL:-0.90}"
      - --max-num-seqs
      - "32"                          # MTP helps up to ~32 streams and hurts above (measured)
      - --kv-cache-dtype
      - fp8_e4m3
      - --speculative-config
      - '{"method":"mtp","num_speculative_tokens":3}'
      - --enable-auto-tool-choice     # agent tool calls (flags from the NVIDIA model card)
      - --tool-call-parser
      - qwen3_coder
      - --reasoning-parser
      - qwen3
      - --seed
      - "0"
      - --host
      - 0.0.0.0
      - --port
      - "8000"
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"]
      interval: 15s
      retries: 5
      start_period: 900s
  api:                                # optional: sessions, budgets, receipt ids
    image: ${DECOSA_REGISTRY}/decosa-api:0.1.0
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct        # local model; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_BUDGET_LLM_TOKENS: "2000000"
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_CODE_MAX_TOKENS: "8192"
      DECOSA_CODE_MAX_INPUT_CHARS: "400000"
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
volumes: { hf-cache: {}, decosa-data: {} }
```
Start it with `cd ~/decosa-code && docker compose up -d`. The first start downloads the weights, and `llm` takes about 5–10 minutes to become healthy. On a smaller card, lower `LLM_MAX_LEN` (for example 32768) and `LLM_GPU_UTIL`.

## 4. Smoke test
```bash
curl -s localhost:8000/v1/models | jq -r '.data[].id'          # expect qwen3.8-27b (and its older alias)
curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' -d '{
  "model":"qwen3.8-27b","max_tokens":400,"chat_template_kwargs":{"enable_thinking":false},
  "messages":[{"role":"user","content":"Write a Python function merge_intervals(intervals). Code only."}]}' \
  | jq -r '.choices[0].message.content'                        # expect a Python function
curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' -d '{
  "model":"qwen3.8-27b","messages":[{"role":"user","content":"Read the file README.md"}],
  "tools":[{"type":"function","function":{"name":"read_file","parameters":{"type":"object",
  "properties":{"path":{"type":"string"}},"required":["path"]}}}]}' \
  | jq '.choices[0].message.tool_calls'                        # expect a read_file call
```
If you started `api`, check it as well:
```bash
curl -s localhost:8445/healthz        # "llm": true ("asr" is false here because this stack has no speech model)
TOKEN=$(curl -s localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"code"}' | jq -r .token)
curl -si localhost:8445/v1/chat/completions -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Say hi"}]}' | grep -i x-decosa-receipt
```
The response body carries `"receipt": {"status": "attested", ...}` in direct mode. The `api` route passes `tools` through: on the hosted service a request with `tools` returns `tool_calls` and a receipt (checked 2026-09-30). On your own box that needs the two tool-call flags above and has not been re-verified there, so test one call first, or point agents straight at vLLM.

## 5. Point your coding agents at the local endpoint
Base URL `http://127.0.0.1:8000/v1`, model `qwen3.8-27b`. vLLM started without `--api-key` accepts any key. If you add one, pass the same value below. Generic OpenAI-SDK tools: `export OPENAI_BASE_URL=http://127.0.0.1:8000/v1 OPENAI_API_KEY=local`.
- **OpenCode** (`opencode.json` in the project, or `~/.config/opencode/opencode.json`):
  ```json
  { "$schema": "https://opencode.ai/config.json",
    "provider": { "local": { "npm": "@ai-sdk/openai-compatible", "name": "Local Qwen",
      "options": { "baseURL": "http://127.0.0.1:8000/v1" },
      "models": { "qwen3.8-27b": { "name": "Qwen3.8-27B (local)" } } } },
    "model": "local/qwen3.8-27b" }
  ```
- **Continue** (`~/.continue/config.yaml`):
  ```yaml
  name: Local
  version: 1.0.0
  schema: v1
  models:
    - name: Qwen3.8-27B (local)
      provider: openai
      model: qwen3.8-27b
      apiBase: http://127.0.0.1:8000/v1
      apiKey: local
      roles: [chat, edit, apply]
      capabilities: [tool_use]
  ```
- **Codex-style CLIs** (`~/.codex/config.toml`; vLLM serves both `/v1/responses` and `/v1/chat/completions`, so choose the `wire_api` your version supports):
  ```toml
  model = "qwen3.8-27b"
  model_provider = "local"
  [model_providers.local]
  name = "Local vLLM"
  base_url = "http://127.0.0.1:8000/v1"
  env_key = "LOCAL_LLM_KEY"      # export LOCAL_LLM_KEY=local
  wire_api = "responses"
  ```
- vLLM 0.29.0 also serves an Anthropic-style `/v1/messages` route. We haven't tested agents against it.
- Tool calling via `qwen3_coder` follows the NVIDIA model card. It hasn't been soak-tested on this exact pin yet, so run your agent on a throwaway branch first.

## Notes
- Optional larger tier: DeepSeek-V4-Flash (MIT, 284B total / 13B active) on **two** RTX PRO 6000 cards, using the public MIT kit https://github.com/hikarioyama/dsv4-flash-nvfp4-sm120 (`serve_b12x_tp2.sh`, patched B12X vLLM build, MTP on). Stop `llm` first, because it takes both GPUs. It is served as `DeepSeek-V4-Flash` with the `deepseek_v4` tool parser. It isn't a hosted model, so it produces no receipts.
- Qwen3.8-Flash-Next is deliberately excluded. Its Qwen Community License requires a separate licence for any Model-as-a-Service business.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
An OpenAI-compatible coding model on your own GPU, so your source code never leaves the machine.
Who it's for
Teams in software and ai ops.
Where it runs
Hosted or self-host
Key numbers
  • 478 / 500 Coding benchmark, 5 core tasks, direct API (Qwen3.8-27B, vLLM NVFP4 + MTP) (test split, n = 5)
  • vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (/500) Same weights on other engines (test split, n = 5)
  • 12/12 Code smoke on the exact pinned stack (test split, n = 12)
  • 4.3 s Median end-to-end run, hosted (QA sweep 2026-09-30)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Hosted or self-host
Checks
Strong
Output
Notes, reports and drafts
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

Does the self-hosted coding assistant send my source code anywhere?

When the private code assistant is self-hosted, nothing leaves the machine: vLLM is bound to 127.0.0.1 and the optional api service signs receipts with a local key, unless you turn on the network profile. The hosted demo is different. It sends prompts to a shared GPU and our gateway, so don't paste proprietary code there. Hosted prompts and answers are not stored; only the receipt (hashes, token counts, cost) is kept.

Which coding tools can use this self-hosted model?

The private code assistant serves Qwen3.8-27B on vLLM behind a standard OpenAI-compatible /v1 endpoint, so you can point OpenCode, Continue, Codex-style CLIs or any OpenAI-SDK tool at it. Tool calling works on the hosted route too (OpenAI `tools`, with a receipt for each call). One limit matters for agents: a hosted answer stops at 2,048 generated tokens per call.

Can I run a coding LLM locally, and on what GPU?

The private code assistant has three self-host tiers. Lite runs the same Qwen3.8-27B NVFP4 weights on one 32 GB RTX 5090, giving up context length and concurrent sessions, not model quality; this tier was not measured specifically. Standard uses one 96 GB RTX PRO 6000 Blackwell with the full 262k context and room for about 32 concurrent streams. Best runs DeepSeek-V4-Flash across two 96 GB cards.

How well does Qwen3.8-27B score on coding tasks here?

On the owner's own coding benchmark, the private code assistant's Qwen3.8-27B scored 478 / 500 on vLLM with NVFP4 and MTP, and passed a 12/12 code smoke on the exact pinned stack. Caveats: the benchmark has only 5 tasks and is not an independent public benchmark, the results are single runs with run-to-run variation not measured, and a 32 GB card was not measured.

Why does the serving engine matter for a self-hosted coding model?

The same Qwen3.8-27B weights scored very differently depending on the engine in the private code assistant's tests: 478 / 500 on vLLM NVFP4 with MTP, 387 on vLLM FP8 eager, 193 on vLLM FP8 with MTP and 63 on llama.cpp CUDA Q8_0. These gaps were engine bugs, not the model. They are single runs on a 5-task benchmark written by the owner.

What does the hosted coding endpoint cost?

On the hosted API, the private code assistant costs about $0.0005 per 300-token answer at the gateway's list price of $0.30 per million prompt tokens and $1.50 per million generated, measured 2026-09-25. An API key allows 200,000 generated tokens a day and 60 requests a minute. On that date the gateway's metering overstated prompt tokens by about 25-80%, and a fix is in progress.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Private code assistant

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.