Skip to content
decosa

The Decosa agent

A coding agent that can keep your code on your machine

A coding agent on an open model, Qwen3.8-27B, with the Decosa tools built in. Run it on your own GPU and nothing leaves your machine, or use it on Decosa's hosted service, where every model call comes back with a signed receipt.

decosa setup && decosa agent

Private beta, 29 Sep 2026. The install line works for beta accounts with repository access; the public package is waiting on release approval.

Where the model runs

On your own GPU

decosa agent --local http://127.0.0.1:8000/v1
  • The same agent against your own vLLM: your code never leaves your machine. No key or account needed.
  • Standard: one 96 GB RTX PRO 6000 (what Decosa runs, measured). Lite: one 32 GB RTX 5090 at a shorter context, not yet benchmarked by us.
  • Set it up with the code kit's prompt (decosa.ai/apps/code), then start vLLM with --enable-auto-tool-choice --tool-call-parser qwen3_coder.
  • Some of the kit's container images are still being published; the kit says how to build them.

Hosted on Decosa

decosa agent
  • Calls go only to Decosa's hosted service (the gateway fails closed rather than use anyone else's), never to a third-party AI lab.
  • Prompts and code are not stored; each call's signed receipt (hashes, token counts, cost) is kept, and decosa usage totals them.
  • The hosted route follows the code tool's rule: don't send proprietary code. Use your own GPU for that.
  • Needs a Decosa API key (dk_...).

Who it's for

  • Teams that can't send code to a third partyHealth, finance, defence suppliers, anyone whose policy says code stays in-house: run the agent on your own GPU and nothing leaves. The hosted beta is for code you're allowed to share with a vendor.
  • Decosa customers automating Decosa toolsRun Decosa tools from a script or your editor and keep the receipts: add the Decosa tools to the agent you already use (Claude Code, Cursor, Codex), or run decosa agent --loop "<job>", the small built-in loop the numbers below measure. The coding agent (decosa agent) can run them too, but it's slower and costs more for tool jobs.

What it's good at, and what it isn't

Decosa jobs

93%
of 46 Decosa tasks right with the Decosa tools (100% with thinking on)
46%
the same model as a generic agent with our docs and raw HTTP
$0.007
per Decosa task at list price, with the built-in loop (decosa agent --loop)

Our own eval of 46 tasks with the built-in loop, 28 Sep 2026; it's not a public benchmark. The coding agent (pi) takes more calls for tool jobs: a 3-product tariff batch took 46 calls, about $0.27, in a cold-user test.

General coding, measured against Claude Code

AgentTasks passedTime per task (median)Cost per taskWhere your code went
Decosa agent (Qwen3.8-27B)27 of 34 (79%); held-out: 19 of 2037 s; held-out: 50 s$0.013; held-out: $0.008Decosa's hosted service only; every call receipted
Claude Code (Opus 5.5)34 of 34 (100%); held-out: 20 of 2012 s; held-out: 10 s$0.066; held-out: $0.051Anthropic

First set: the 34 Python exercises of the Aider polyglot benchmark. Held-out set: 20 other Exercism Python exercises, picked at random before the run. Exercism content, MIT licence; 12 minutes per task, one run each, 28–29 Sep 2026. Times and costs are medians: the Decosa agent at our list price, Claude Code as it reported. On the first set, the Decosa agent's 7 misses were 5 runaway loops and 2 test runs that never finished, which pushed its mean cost to $0.25 a task. It now stops a stuck task (60 calls, $0.50, or the same step 8 times) and time-boxes test runs; the held-out set was run after that change.

So: Claude Code is the stronger coder. The Decosa agent is for code that can't leave your company or your machine: it passed most of the tasks, at a fifth of the typical cost, and every model call came with a signed receipt.

How it works

  • An open harness, unmodifiedThe agent is pi (MIT), with its read, bash, edit and write tools, sessions and a terminal UI. We don't fork it; decosa agent gives it Decosa's model and the Decosa skill, with pi's network features and telemetry off.
  • Receipts that cover the tool callsEach call's receipt is signed by Decosa's gateway: a hash of the tool definitions and the conversation, a hash of the answer with its tool calls, tokens, cost and the serving machine. The agent checks each answer against its receipt, and decosa verify checks a receipt offline. A receipt proves what the gateway signed; hardware-attested proof of where a call ran is what confidential access adds.
  • Your key stays out of the agentThe agent talks to a small passthrough on 127.0.0.1 that adds your key and forwards each request unchanged. The key isn't written to disk or shown to the model.

Questions

Can I send proprietary code to the hosted agent?
Not on the hosted route yet: it follows the code tool's rule (don't send proprietary code). Run the agent on your own GPU with --local for that; nothing leaves your machine.
Is my code used to train models?
No. Hosted prompts and answers are not stored; the receipt keeps hashes, token counts and cost. Self-hosted, nothing leaves your machine.
Which model is it?
Qwen3.8-27B (Apache-2.0), NVFP4 on Decosa's hosted service. The receipt names the model and its pinned weights.
Can I use it from Claude Code, Cursor or Codex instead?
Yes: the Decosa tools are an MCP server (decosa-mcp). Your agent keeps its own model; the tools run Decosa jobs and verify receipts. See the setup page.
What does it need?
Python 3.10+ for the decosa command, and Node 22.19+ for the coding agent (it fetches pi, pinned, the first time).
Install the Decosa agent →