{"schema_version":"1","site":"https://decosa.ai","id":"code","num":"03","name":"Private code assistant","tool_name":"Run an open model behind the OpenAI API","short":"Code","blurb":"A coding agent on an open model that never sends your source to a third party.","status":"live","labels":{"industry":["software"],"job":["draft"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["text"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["software"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":[],"models":"Qwen3.8-27B","where":"Hosted or self-host","hardware":"1× RTX PRO 6000 (96 GB), or 2× RTX 5090","final_artifact":"Streamed answers, each with a receipt.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Strong","manual_qa":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":4268,"p95_ms":null,"runs":null,"receipts_per_run":1,"cost_per_run_usd":0.00046},"selfhost":{"date":"2026-09-25","result":"pass","method":"Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers","notes":"Verified on 2026-09-25: the fallback llm image builds, the compose file validates, the api starts and the sample passes end to end (attested receipts, p50 1.5 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Tool calling was not re-verified: the local server ran without the tool-parser flags."},"known_limits":["Hosted answers stop at 2,048 generated tokens (finish_reason \"length\"); self-host for longer outputs.","Tool calling works on the hosted route (checked on production 2026-09-30: a request with `tools` returns `tool_calls` and a receipt, and the follow-up turn with the tool result is answered). Each call is still one request of at most 2,048 generated tokens.","The hosted route computes the whole answer before streaming it, so text arrives in one burst.","Hosted p50 is for an answer of about 340 tokens (7 calls on 2026-09-30, fastest 2.5 s, slowest 5.1 s); under load on 2026-09-25 the same call took up to 37 s.","Token counts on hosted receipts are the gateway's metering, which on 2026-09-25 overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress)."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Coding benchmark, 5 core tasks, direct API (Qwen3.8-27B, vLLM NVFP4 + MTP)","value":"478 / 500","unit":null,"n":5,"split":"test","note":"single run"},{"name":"Same weights on other engines","value":"vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (/500)","unit":null,"n":5,"split":"test","note":"engine bugs, not the model"},{"name":"Code smoke on the exact pinned stack","value":"12/12","unit":null,"n":12,"split":"test","note":null},{"name":"Best tier (DeepSeek V4-Flash), agent harness","value":"495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses) / 500","unit":null,"n":5,"split":"test","note":"measured on the DSpark serving build, not the NVFP4 kit this tier uses"}],"dataset":"coding-agent-bench (the owner's own coding benchmark, 5 core tasks scored /500), results/RESULTS.md on our server, plus the 12-task code smoke from the model benchmark page.","held_out":false,"caveats":["The owner's own benchmark with only 5 tasks; not an independent public benchmark.","Single runs; run-to-run variation not measured.","A 32 GB card (lite tier) was not measured specifically.","The best-tier figures were measured on a different serving build than the one this tier ships."],"date":"2026-09-23","doc_url":null},"stack":{"summary":"One open-weight coding model, Qwen3.8-27B, runs on vLLM with NVFP4 weights and its built-in MTP draft head, behind a standard /v1 endpoint. Point OpenCode, Continue, Codex-style CLIs or any OpenAI-SDK tool at it. Self-hosted, prompts and code stay on the box. On the hosted demo, every completion comes back with a gateway-signed receipt that anyone can re-check. It suits teams that want an agentic coding assistant without sending their repository to a third-party API.","tagline":"An OpenAI-compatible coding model on your own GPU, so your source code never leaves the machine.","deployment":"hosted-or-self-host","regulatory_note":"Self-hosted, source code and prompts stay on your machine. The hosted demo sends prompts to a shared GPU and our gateway, so don't paste proprietary code there. Model licences: Apache-2.0 (Qwen3.8-27B), MIT (optional DeepSeek-V4-Flash tier).","components":[{"id":"llm","role":"Coding model (chat, edit, agent tool calls)","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (OpenAI-compatible server)","receipt_coverage":"strong","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"llm-large","role":"Optional larger tier (two GPUs)","name":"DeepSeek-V4-Flash (NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts (MTP draft experts MXFP4)","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM, B12X native sm_120 build, TP2, MTP on, with two mount-overlay patches from the dsv4-flash-nvfp4-sm120 serving kit (MIT)","receipt_coverage":"none","in_hosted_demo":null,"tiers":["best"],"alternative_to":null},{"id":"llm-wanted","role":"Network-hosted flash model (wanted)","name":"GLM-5.3-Flash","hf_repo":"nvidia/GLM-5.3-Flash-NVFP4","license":"MIT","params":"321B","quant":"NVFP4 (NVIDIA quant of the FP8 release)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang sm_120 build (vLLM on sm_120 is broken per the upstream issue); not running on our server today","receipt_coverage":"none","in_hosted_demo":null,"tiers":["wanted"],"alternative_to":null},{"id":"dsv41-wanted","role":"Coding model, datacentre class","name":"DeepSeek-V4.1-Flash","hf_repo":"deepseek-ai/DeepSeek-V4.1-Flash","license":"MIT","params":"552B backbone (763B incl. Engram tables)","quant":"Official FP8 (block 32x32) + FP4 experts, 476 GB on disk","vram_gb":null,"memory_gb_estimate":476,"engine":"vLLM >= 0.30 tagged image on 8x H200, GB200 NVL4 or 4x B200 (DeepSeek's recipe); no sm_120 path found","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · runs on one 32 GB Blackwell card","summary":"Same Qwen3.8-27B NVFP4 weights and engine as Standard. You give up context length and concurrent sessions, not model quality.","components":["llm"],"hardware":"1x RTX 5090 32 GB (Blackwell, NVFP4)","quality_evidence":[{"metric":"Coding benchmark, 5 core tasks, direct API (/500)","value":"478 (same weights; measured on a 96 GB card, single run)","source":"coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16"},{"metric":"On a 32 GB card specifically","value":"not measured yet","source":"not measured yet"}],"latency_note":"estimate: similar single-stream speed to Standard (decode is memory-bound on the same weights), but with a much smaller KV cache, so plan for a shorter max-model-len and one or two sessions.","in_hosted_demo":false,"receipt_coverage":"strong","receipt_note":"Qwen3.8-27B is the hosted model qwen3.8-27b. Self-hosted with the direct route there are no receipts.","hosting":null},{"id":"standard","label":"Standard · one 96 GB card (hosted demo)","summary":"Qwen3.8-27B NVFP4 + MTP3 on the pinned vLLM 0.29.0 with a full 262k context and room for about 32 concurrent streams. This is the hosted demo.","components":["llm"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Coding benchmark, 5 core tasks, direct API (/500)","value":"478 (vLLM NVFP4 + MTP, single run)","source":"coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16"},{"metric":"Same weights on other engines (/500)","value":"vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (engine bugs, not the model)","source":"coding-agent-bench results/RESULTS.md"},{"metric":"Code smoke on the exact pinned stack","value":"12/12","source":"Decosa model benchmarks (Sep 2026)"}],"latency_note":"measured on our server: time to first token in milliseconds and fast single-stream decoding locally. The hosted route's end-to-end time is the measured figure on this page.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every hosted completion carries a gateway-signed receipt (commitment hash, Ed25519 countersignature, metering proofs).","hosting":null},{"id":"best","label":"Best · two 96 GB cards","summary":"DeepSeek-V4-Flash (284B MoE, MIT) across both GPUs for longer agentic runs. It takes the whole box. The 5-task suite is near its ceiling, so it doesn't clearly beat Standard one-shot; its edge shows with an agent harness.","components":["llm-large"],"hardware":"2x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"Coding benchmark, 5 core tasks, agent harness (/500)","value":"495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses); single runs","source":"coding-agent-bench results/RESULTS.md, official_prime_results.json and ds4_*_results.json, 2026-08-15/16"},{"metric":"Same suite, direct API (/500)","value":"459 uncapped (single run)","source":"coding-agent-bench results/RESULTS.md, ds4_api_uncapped_results.json"},{"metric":"Caveat","value":"measured on the DSpark serving build, not on the NVFP4 kit this tier uses; NVFP4 kit not measured yet","source":"coding-agent-bench results/RESULTS.md"}],"latency_note":"measured by the dsv4-flash-nvfp4-sm120 kit: 150.6 tok/s single-stream with MTP, about 328 tok/s aggregate at 4 streams.","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"None today: DeepSeek-V4-Flash is not a hosted model, so this tier produces no receipts. It becomes strong (text LLM coverage) once it is hosted with a pinned engine.","hosting":null},{"id":"wanted","label":"Wanted · the two most-used open coding models","summary":"GLM-5.3-Flash and DeepSeek-V4.1-Flash, served by network providers. Neither fits one workstation card: GLM needs two 96 GB cards (its sm_120 runtime hangs on our server today) and V4.1-Flash a datacentre node.","components":["llm-wanted","dsv41-wanted"],"hardware":"Network providers: 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash; an 8x H200-class node for DeepSeek-V4.1-Flash. Estimate.","quality_evidence":[{"metric":"Coding benchmark, 4 tiebreaker tasks, Prime Agent harness (/400)","value":"372.3 three-run mean (389, 359, 369), community TR3 4bpw build, 8k thinking budget","source":"coding-agent-bench results/RESULTS.md, h2h/glm53_tr3_t8k_tb_r1-3.json, 2026-08-28"},{"metric":"DeepSeek-V4.1-Flash on the same benchmark","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"Not hosted yet, so no receipts today.","hosting":"network"}],"alternates":[],"services":[{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM 0.29.0 serving Qwen3.8-27B NVFP4 + MTP3 as qwen3.8-27b. /v1/chat/completions, /v1/responses, /v1/models. Publish on 127.0.0.1 only; with --enable-auto-tool-choice --tool-call-parser qwen3_coder it serves agent tool calls."},{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"OpenAI-compatible /v1/chat/completions proxy with demo-session auth, budgets and receipts (x-decosa-receipt header, GET /receipts/{id}). Tool calling works on this route (OpenAI `tools`; the answer comes back as `tool_calls` with a receipt); max_tokens capped at 2048, prompts at 48,000 characters."}],"tools":[{"name":"OpenCode","url":"https://opencode.ai","license":"MIT","purpose":"Terminal coding agent; add the endpoint as an @ai-sdk/openai-compatible provider in opencode.json."},{"name":"Continue","url":"https://continue.dev","license":"Apache-2.0","purpose":"IDE extension (VS Code, JetBrains); add a model with provider: openai and apiBase pointing at the endpoint."},{"name":"Codex CLI","url":"https://github.com/openai/codex","license":"Apache-2.0","purpose":"Terminal coding agent; add a custom model_providers entry in ~/.codex/config.toml pointing at vLLM's /v1 (it serves /v1/responses and /v1/chat/completions)."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: weights 19.9 GiB, KV space for 1.76M tokens at 262k max length, 132 tok/s single-stream, about 1,170-1,520 tok/s at 32 streams. This is the pinned stack."},{"tier":"24-32 GB Blackwell card","fits":true,"notes":"Not measured here. The NVFP4 weights are about 20 GB, so use a shorter max-model-len and fewer sequences. A third-party report gives about 50 tok/s single-stream with NVFP4 + MTP on a 24 GB Blackwell card."},{"tier":"Hopper or Ada 48-80 GB (no NVFP4)","fits":true,"notes":"Use the official FP8 checkpoint Qwen/Qwen3.8-27B-FP8 (Apache-2.0, 29 GB). Measured on our server at 49 tok/s single-stream without MTP, 90-94 with MTP3; not measured on Hopper or Ada."},{"tier":"2x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Optional DeepSeek-V4-Flash tier, TP2 on both cards with the patched B12X build from the dsv4-flash-nvfp4-sm120 kit. It replaces the Qwen tier on that box."},{"tier":"Cards under 24 GB","fits":false,"notes":"The NVFP4 weights alone are about 20 GB, leaving no useful KV cache."}],"latency":[{"lane":"time to first token, local vLLM, 1 stream, ~539-token prompt","typical_ms":92,"source":"measured on our server 2026-09-23 (NVFP4 + MTP3 + FP8 KV, benchmark page 17)"},{"lane":"time per output token, local vLLM, 1 stream","typical_ms":7.4,"source":"measured on our server 2026-09-23 (benchmark page 17, ~132 tok/s)"},{"lane":"time to first token, local vLLM, 32 streams, ~2k-token prompts","typical_ms":504,"source":"measured on our server 2026-09-23 (benchmark page 17, median; p99 6,510 ms)"},{"lane":"hosted /v1/chat/completions, ~350 output tokens, gateway route with signed receipt","typical_ms":4268,"source":"measured on production 2026-09-30 (7 calls through api.decosa.ai, median; 2.5-5.1 s end to end for about 340 output tokens; hosted streaming arrives all at once, so TTFT equals this)"},{"lane":"DeepSeek-V4-Flash tier, 1 stream decode","typical_ms":6.6,"source":"derived from the dsv4-flash-nvfp4-sm120 kit's measurement on 2x RTX PRO 6000 (150.6 tok/s single-stream with MTP)"}],"benchmark":null,"notes":["Qwen3.8-Flash-Next is excluded: its Qwen Community License 1.0 requires a separate licence from Qwen for any Model-as-a-Service business, which covers serving it through the network. Its uncensored and abliterated derivatives carry the same restriction.","Tool calling for agents uses vLLM's qwen3_coder parser from the NVIDIA model card recipe; that flag set has not yet been tested on the pinned stack.","The hosted route re-emits a finished completion as SSE, so tokens don't arrive progressively there. Self-hosted vLLM streams normally."]},"buyer_facts":[{"label":"Data retention","value":"Hosted: prompts and answers are not stored; only the receipt (hashes, token counts, cost) is kept. Self-hosted: nothing leaves the machine unless you turn on the network profile."},{"label":"What leaves the box (self-host)","value":"Nothing. vLLM is bound to 127.0.0.1 and the optional api service signs receipts with a local key."},{"label":"Inputs","value":"OpenAI-compatible chat messages, up to 48,000 characters per request on the hosted API; model qwen3.8-27b."},{"label":"Typical hosted cost","value":"A fraction of a cent per short answer at the gateway's list price (measured). Each run shows its own measured cost."},{"label":"Hosted limits","value":"Demo session: 20,000 generated tokens. API key: 200,000 generated tokens a day, 60 requests a minute."}],"data_handling":{"page":"/data#code","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Hosted: prompts and answers are not stored; only the receipt (hashes, token counts, cost) is kept. Self-hosted: nothing leaves the machine unless you turn on the network profile.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/developer/code","input":"chat","lanes":[],"samples":[],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/code-hosted.md","selfhost":"/prompts/code-selfhost.md","assemble":"/prompts/code-assemble.md","mac":"/prompts/code-mac.md"},"rehearsal":{"bundle":"/samples/code.zip","bundle_url":"https://decosa.ai/samples/code.zip","folder":"/samples/code/","expected":"/samples/code/expected.json","files":["/samples/code/expected.json","/samples/code/inputs/prompt.txt"],"bytes":1347,"checks":["the answer defines merge_intervals","the answer carries doctests","the answer finished within the token limit","the route reports token usage","the answer carries a signed (hosted) or attested (self-hosted) receipt","the receipt is public at /receipts/{id}","the public receipt lists its checks","every check on the receipt passes","every model call has a signed receipt"],"licence":"Synthetic: a coding prompt written for Decosa (the nightly smoke check's prompt). No real code base or people. Part of decosa-api, AGPL-3.0-or-later.","about":"One short coding request (a merge_intervals function with a docstring and three doctests) sent to the OpenAI-compatible chat route. The answer must hold the function and its doctests, and the call must carry a receipt that is public at /receipts/{id} and whose checks all pass. There is no verify route for a single completion, so there is no tamper step; the receipt's own checks are the proof.","run":{"containers":"docker compose exec api python scripts/rehearse.py code","checkout":"python scripts/rehearse.py code --bundle code.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py code"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=code","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":192,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":1320,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/tools/developer/code","json":"/use-cases/code.json","metrics":"/metrics/code","console":"/tools/developer/code","console_sample":"/tools/developer/code?sample=1&autorun=0","stack":"/tools/developer/code#stack","try_live":"/tools/developer/code","watch":"/tools/developer/code","build":"/tools/developer/code#build","self_host":"/tools/developer/code#self-host","prompts":{"hosted":"/prompts/code-hosted.md","selfhost":"/prompts/code-selfhost.md","assemble":"/prompts/code-assemble.md","mac":"/prompts/code-mac.md"}}}