Skip to content
decosa

03 · Software and AI ops · live

Private code assistant

Open the toolJSON

Eval results

Not held outRun 23 Sep 2026

  • Coding benchmark, 5 core tasks, direct API (Qwen3.8-27B, vLLM NVFP4 + MTP)478 / 500test splitn = 5single run
  • Same weights on other enginesvLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (/500)test splitn = 5engine bugs, not the model
  • Code smoke on the exact pinned stack12/12test splitn = 12
  • Best tier (DeepSeek V4-Flash), agent harness495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses) / 500test splitn = 5measured on the DSpark serving build, not the NVFP4 kit this tier uses

Dataset

coding-agent-bench (the owner's own coding benchmark, 5 core tasks scored /500), results/RESULTS.md on our server, plus the 12-task code smoke from the model benchmark page.

Caveats

  • The owner's own benchmark with only 5 tasks; not an independent public benchmark.
  • Single runs; run-to-run variation not measured.
  • A 32 GB card (lite tier) was not measured specifically.
  • The best-tier figures were measured on a different serving build than the one this tier ships.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
30 Sep 2026
Latency, this run
n/a
p50 over passed runs
4.3 s
Receipts
1
Model calls
n/a
Tokens
n/a
Cost per run
<$0.001

Self-host verification

Verified on 25 Sep 2026: Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers

Verified on 2026-09-25: the fallback llm image builds, the compose file validates, the api starts and the sample passes end to end (attested receipts, p50 1.5 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Tool calling was not re-verified: the local server ran without the tool-parser flags.

Rehearsal bundle: code.zip (1 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted answers stop at 2,048 generated tokens (finish_reason "length"); self-host for longer outputs.
  • Tool calling works on the hosted route (checked on production 2026-09-30: a request with `tools` returns `tool_calls` and a receipt, and the follow-up turn with the tool result is answered). Each call is still one request of at most 2,048 generated tokens.
  • The hosted route computes the whole answer before streaming it, so text arrives in one burst.
  • Hosted p50 is for an answer of about 340 tokens (7 calls on 2026-09-30, fastest 2.5 s, slowest 5.1 s); under load on 2026-09-25 the same call took up to 37 s.
  • Token counts on hosted receipts are the gateway's metering, which on 2026-09-25 overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · runs on one 32 GB Blackwell card (2)
  • Coding benchmark, 5 core tasks, direct API (/500): 478 (same weights; measured on a 96 GB card, single run)coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16
  • On a 32 GB card specifically: not measured yetnot measured yet
Standard · one 96 GB card (hosted demo) (3)
  • Coding benchmark, 5 core tasks, direct API (/500): 478 (vLLM NVFP4 + MTP, single run)coding-agent-bench results/RESULTS.md, qwen38_nvfp4_results.json, 2026-08-16
  • Same weights on other engines (/500): vLLM FP8 eager 387; vLLM FP8 + MTP 193; llama.cpp CUDA Q8_0 63 (engine bugs, not the model)coding-agent-bench results/RESULTS.md
  • Code smoke on the exact pinned stack: 12/12Decosa model benchmarks (Sep 2026)
Best · two 96 GB cards (3)
  • Coding benchmark, 5 core tasks, agent harness (/500): 495 (Prime Agent), 485 (Claude Code harness), 481-483 (other harnesses); single runscoding-agent-bench results/RESULTS.md, official_prime_results.json and ds4_*_results.json, 2026-08-15/16
  • Same suite, direct API (/500): 459 uncapped (single run)coding-agent-bench results/RESULTS.md, ds4_api_uncapped_results.json
  • Caveat: measured on the DSpark serving build, not on the NVFP4 kit this tier uses; NVFP4 kit not measured yetcoding-agent-bench results/RESULTS.md
Wanted · the two most-used open coding models (2)
  • Coding benchmark, 4 tiebreaker tasks, Prime Agent harness (/400): 372.3 three-run mean (389, 359, 369), community TR3 4bpw build, 8k thinking budgetcoding-agent-bench results/RESULTS.md, h2h/glm53_tr3_t8k_tb_r1-3.json, 2026-08-28
  • DeepSeek-V4.1-Flash on the same benchmark: not measured yet

How we measure · All tools