Examiner prompt
waitingCriteria still without evidence, and neutral follow-ups that passed the guard.
A draft score for each rubric criterion that quotes the candidate's own words, which the examiner confirms or overrides with a reason on the record.
Built on: Live speech to text, Speaker diarization, Signed record
Examiner side only: the tool drafts, a person decides every score. Demo only: use the synthetic samples or your own consented test session, never a real student's or candidate's exam. Tell everyone the session is recorded.
Start recording or run a sample.
Criteria still without evidence, and neutral follow-ups that passed the guard.
Leading questions and talk balance.
Entries, head hash and signed checkpoints.
Written when you stop; words the recognisers disagree on are marked.
One level per criterion, the lines it rests on, and an evidence check.
Sealed when the session ends; the review chains to it.
Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.
Intro statistics viva, partly correct, with a leading question (synthetic). Recorded from a real run of the hosted pipeline on 2026-09-30 (synthetic voices, speech recognition, speaker pass, scoring and signing all real), played back at 2x. After it, the examiner's review as it was signed.
Press play.
Criteria still without evidence, and neutral follow-ups that passed the guard.
Leading questions and talk balance.
Entries, head hash and signed checkpoints.
Written when you stop; words the recognisers disagree on are marked.
One level per criterion, the lines it rests on, and an evidence check.
Sealed when the session ends; the review chains to it.
Each step is signed: which model ran, and a fingerprint of what went in and came out, so it can be checked later.
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)# Decosa structured oral assessment: use the hosted API
You are wiring Decosa's oral-assessment assistant into this project. For vivas and structured interviews it gives the
examiner a live prompt (rubric criteria still without evidence, neutral follow-ups, leading questions flagged) and,
afterwards, a draft score per rubric criterion that cites the candidate's transcript lines, with a grounding check of
that evidence. The examiner confirms or overrides every score with a reason, and both steps come back as signed,
chained records. Each model call has its own signed receipt. Use only what is listed below. If you need something else,
stop and ask me.
- Base URL: `https://api.decosa.ai` (WebSocket: `wss://api.decosa.ai`)
- Health check: `GET https://api.decosa.ai/healthz`.
- The hosted API is for the synthetic samples and for consented test sessions. Real students' exams are education
records (FERPA) and interviews are personal data: for real assessments use the self-host prompt instead.
- It drafts; an examiner decides. Show every score as a draft, never as a grade, and never skip the review step.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "oral-assessment"}` returns `{"token", "expires_at", "budget"}`.
a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); 300 s of audio and a token allowance per session (the `budget` in the session response). Over a limit: HTTP 429 with `Retry-After`.
## Endpoints
- Live audio: `wss://api.decosa.ai/ws/live?vertical=oral-assessment&token=<t>&rubric=intro-stats-viva` (or
`support-specialist-interview`). Send 16 kHz mono PCM16 frames of about 100 ms, then `{"type":"stop"}`. Lanes:
`prompt`, `guard`, `chain`, then `final_transcript`, `scores`, `record`. `done.summary.draft.scores` and
`done.summary.record` (the signed draft) feed the review.
- `POST /oral/score` (token): `{"rubric_id": "intro-stats-viva" | "rubric": {...}, "lines": [{"role": "examiner"|"candidate", "text"}] | "transcript": "Examiner: ...\nStudent: ..." | "sample_id", "names"?: [...], "minor"?: false, "stream"?: true}`.
Streams `ready` (blinding report, leading questions), a `receipt` per model call, one `score` per criterion
(`level`, `max_level`, `p`, `reason`, `evidence: [{n, text}]`, `grounding`, `flags`), then `draft` with the signed
`record`. With `"stream": false`: `{ready, draft, budget}`. Limits: 400 lines, 60,000 characters, 12 criteria.
- `POST /oral/review` (token): `{"draft": <draft.record>, "examiner": "<name or id>", "decisions": [{"criterion", "action": "confirm"|"override", "level"?, "reason"?}]}`.
One decision per criterion; an override needs a new level and a reason. Returns `{final_total, max_total, overrides, bundle, check}`.
- `POST /oral/verify` (no token) `{"bundle": ...}` → `{ok, checks, summary}`. `POST /record/verify` checks either record alone.
- `POST /oral/prompt` (token): the examiner prompt for a transcript so far. `GET /oral/info`, `GET /oral/samples`,
`GET /oral/rubrics/{id}`, `GET /attest/signing-key` (no token).
## Example: score a transcript, review it, keep the bundle (Python, `pip install httpx`)
```python
import httpx, json, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
s = httpx.get(f"{API}/oral/samples/oral-stats-partial").json() # your transcript has the same shape
d = httpx.post(f"{API}/oral/score", headers=H, timeout=300,
json={"rubric_id": s["rubric"], "lines": s["lines"], "stream": False}).json()["draft"]
for x in d["scores"]:
print(x["criterion"], x["level"], x["flags"], [e["n"] for e in x["evidence"]], x["reason"])
decisions = [{"criterion": x["criterion"], "action": "confirm"} for x in d["scores"]] # the examiner's real decisions go here
rv = httpx.post(f"{API}/oral/review", headers=H, timeout=60,
json={"draft": d["record"], "examiner": "Examiner 7", "decisions": decisions}).json()
pathlib.Path("viva-bundle.json").write_text(json.dumps(rv["bundle"]))
print(httpx.post(f"{API}/oral/verify", json={"bundle": rv["bundle"]}).json()["summary"])
```
## Honest limits
- Scores are drafts from an open model. On synthetic held-out answers they matched careful labels on about nine in ten
criteria and were never more than one level off; real answers and real examiners will differ. See the Stack tab.
- The flags do not reliably mark the scores that are wrong: the examiner must decide every criterion.
- Speech recognition makes more mistakes on some accents; words the two recognisers disagree on are marked, and a score
resting on them is flagged for a re-listen rather than lowered.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa structured oral assessment: run it yourself (containers)
You are setting up the Decosa oral-assessment assistant on this machine, so students' and candidates' recordings,
transcripts and scores never leave it. It gives the examiner a live prompt during a viva or structured interview,
drafts a cited score per rubric criterion afterwards, and seals the draft and the examiner's decisions in signed,
chained records. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/oral-assessment.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py oral-assessment` (the api image carries the same bundle under /app/rehearsal/oral-assessment/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py oral-assessment --bundle oral-assessment.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 6 rubric criteria get a draft score", "the wrong p-value definition scores 0 or 1", "agreeing with the misreading (the null is probably true) scores 0 or 1"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` (Qwen3.8-27B), `asr` (Voxtral Mini 4B Realtime) and `api` services; for transcripts only,
`llm` and `api` are enough. For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
`DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_SIGNER_NAME` to the institution that signs the records, and bind every port
to 127.0.0.1. Keep the `decosa-data` volume: it holds the signing key an appeal will need.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the health checks (the first start downloads
about 30 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/oral/info` lists the two built-in rubrics; `GET /attest/signing-key`
shows this box's public key. Show me the key: it is what an appeals panel pins to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"oral-assessment"}`, run
`POST /oral/score {"sample_id": "oral-stats-partial", "stream": false}`, then `POST /oral/review` with the returned
`draft.record`, an examiner name and one decision per criterion (override one with a reason), then `POST /oral/verify`
with the returned `bundle`: `ok` must be true, and false after editing the override's reason.
6. Report back: the public key and key id, the six draft levels with their cited lines, the verify results and how
long scoring took.
For students under 18 send `"minor": true` with `"consent_basis": "school_official"` or `"parental_consent"`: the signed
records then keep hashes only. Tell every candidate the session is recorded and scored with AI help; a person decides.
Off, and keep it off on a box that holds student or candidate data: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/oral-assessment-mac.md instead.
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Voxtral Mini 4B Realtime needs a GPU.
Needs about 40 GB of GPU memory at the smallest settings; 24 GB available.
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
The standard tier fits with changes: Voxtral Mini 4B Realtime: run it at its smallest setting (about 16 GB instead of 24 GB), with a shorter context and fewer parallel sessions.
The standard tier does not fit: Needs about 53.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
The standard tier fits (85.6 of 96 GB).
The standard tier fits (85.6 of 192 GB).
The standard tier fits (48 of 96 GB).
The standard tier fits (48 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa
curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yamlThe first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz
# {"ok": true, "asr": true, "llm": true, ...}
curl -fsS -X POST http://localhost:<PORT>/demo/session \
-H 'Content-Type: application/json' -d '{"vertical":"oral-assessment"}'expected.json. Every check must print PASS.docker compose exec api python scripts/rehearse.py oral-assessment
Download the mock-data bundle (4 KB, 11 checks)expected.json
A short synthetic intro-statistics viva in which the student gives partly correct answers and agrees with a leading question from the examiner (that a p-value is the probability the null is true). The scorer must score all six rubric criteria with cited lines, give low marks where the answers are wrong or vague, flag the leading question, and seal a signed draft; an examiner's review with one override must verify, and an edited override reason must be caught.
Licence: Synthetic: a viva script written for Decosa (no real student or examiner) and the built-in intro-statistics rubric. Part of decosa-api, AGPL-3.0-or-later.
# Decosa structured oral assessment: run it yourself (containers)
You are setting up the Decosa oral-assessment assistant on this machine, so students' and candidates' recordings,
transcripts and scores never leave it. It gives the examiner a live prompt during a viva or structured interview,
drafts a cited score per rubric criterion afterwards, and seals the draft and the examiner's decisions in signed,
chained records. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/oral-assessment.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py oral-assessment` (the api image carries the same bundle under /app/rehearsal/oral-assessment/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py oral-assessment --bundle oral-assessment.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 6 rubric criteria get a draft score", "the wrong p-value definition scores 0 or 1", "agreeing with the misreading (the null is probably true) scores 0 or 1"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Also install the NVIDIA container toolkit and
check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` (Qwen3.8-27B), `asr` (Voxtral Mini 4B Realtime) and `api` services; for transcripts only,
`llm` and `api` are enough. For the `api` service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`,
`DECOSA_LLM_MODEL=qwen3.8-27b`, `DECOSA_SIGNER_NAME` to the institution that signs the records, and bind every port
to 127.0.0.1. Keep the `decosa-data` volume: it holds the signing key an appeal will need.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the health checks (the first start downloads
about 30 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/oral/info` lists the two built-in rubrics; `GET /attest/signing-key`
shows this box's public key. Show me the key: it is what an appeals panel pins to verify my records.
5. Smoke test: get a token with `POST /demo/session {"vertical":"oral-assessment"}`, run
`POST /oral/score {"sample_id": "oral-stats-partial", "stream": false}`, then `POST /oral/review` with the returned
`draft.record`, an examiner name and one decision per criterion (override one with a reason), then `POST /oral/verify`
with the returned `bundle`: `ok` must be true, and false after editing the override's reason.
6. Report back: the public key and key id, the six draft levels with their cited lines, the verify results and how
long scoring took.
For students under 18 send `"minor": true` with `"consent_basis": "school_official"` or `"parental_consent"`: the signed
records then keep hashes only. Tell every candidate the session is recorded and scored with AI help; a person decides.
Off, and keep it off on a box that holds student or candidate data: joining serves other people's requests on this GPU.
If I ask for it later, on a separate machine, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 48 GB of unified memory or more): use https://decosa.ai/prompts/oral-assessment-mac.md instead.
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitStructured oral assessment on GeForce RTX 5090
Needs about 48 GB of GPU memory at the smallest settings; 32 GB available.
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
The self-host prompt for Structured oral assessment, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Structured oral assessment on my hardware Fetch https://decosa.ai/prompts/oral-assessment-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=oral-assessment) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · one 48 GB card, captions only (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Live captions: Voxtral Mini 4B Realtime (mistralai/Voxtral-Mini-4B-Realtime-2602), 24 GB - Lite tier: Qwen3.8-27B (official FP8) (Qwen/Qwen3.8-27B-FP8), 33.6 GB Warning: the fit check says this tier does not fit: Needs about 48 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/oral-assessment-assemble.md
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 48 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh --profile live
# Decosa Structured oral assessment: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Structured oral assessment on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 48 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/oral-assessment.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py oral-assessment` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 6 rubric criteria get a draft score", "the wrong p-value definition scores 0 or 1", "agreeing with the misreading (the null is probably true) scores 0 or 1"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Live captions (streaming, no speakers): the examiner prompt reads these | vLLM realtime WebSocket | MLX 4-bit (mlx-community/Voxtral-Mini-4B-Realtime-2602-4bit) on mlx-audio 0.5.6, behind scripts/mac/asr_server.py | Runs, measured | | After the session: examiner and candidate lines, each with a receipt over its audio | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured | | Examiner prompt, speaker roles, one score per criterion, grounding check of the evidence | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 48 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh --profile live`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model, plus about 5 GB for speech recognition and diarization), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`, `"asr": true` and `"diarize": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py oral-assessment`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --profile live --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 18 s · ~$0.007 per run · 42 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, api image built, the prompt's api service (named volume) against the running local model servers, then torn down
Measured cost to run: about $0.036 per assessment (hosted, 25 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The step 6 smoke passed as written: six scores with cited lines, the leading question at line 3 flagged, one override signed, the bundle verified, and after editing the override's reason verification failed at that decision. The audio replay also passed (speaker lines, signed draft of 87 entries). Model-server startup itself not re-verified (no new GPU load).
For departments bringing back oral exams, certification bodies, and hiring teams running structured interviews. During the session the examiner sees which rubric criteria still lack evidence and neutral follow-ups, and leading questions are flagged. Afterwards an open diarization model writes who said what, the words two recognisers disagree on are marked for a re-listen, names and personal details are removed, and the language model drafts a score for each criterion that cites the candidate's lines, checked by a grounding judge. The examiner confirms or overrides every score with a reason. Both steps are sealed in signed, chained records. Examiner side only; the tool never grades on its own.
Microphone audio streams to decosa-api. Voxtral writes live captions with signed speech receipts; Qwen3.8-27B keeps the examiner prompt, and a rule-based guard drops leading follow-ups and flags leading questions. On stop, MOSS-Transcribe-Diarize writes examiner and candidate lines; words the two recognisers disagree on are marked; names and personal details are removed; Qwen3.8-27B scores each criterion with cited lines, checked by a grounding judge. A signed draft record holds it all; the examiner's confirm-or-override decisions go into a review record chained to the draft, and anyone can verify the pair. Pasted transcripts enter at the scoring step. Everything inside the dashed box runs on one machine when self-hosted.
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
one 48 GB card, captions only
Live captions and the examiner prompt, then scores from the captions with roles guessed from question marks. No speaker labels and no uncertainty marks.
the hosted demo, two recognisers
Captions and the examiner prompt live; then speaker lines, uncertainty marks, blinding, cited scores with a grounding check, and the signed records. Every model call carries a gateway-signed receipt.
DeepSeek-V4-Flash scores and checks
Standard with a larger scorer and grounding judge on two more 96 GB cards.
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
| Live, while it happens | ||||
Live captions (streaming, no speakers): the examiner prompt reads theseVoxtral Mini 4B Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 on Hugging Face (opens in a new tab) 4.4B · 24 GBProof: partialIn the hosted demo | LiteStandardBest | 4.4B · 24 GB | Proof: partialIn the hosted demo | |
| ||||
| After the session | ||||
After the session: examiner and candidate lines, each with a receipt over its audioMOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab) 0.9BProof: partialIn the hosted demo | StandardBest | 0.9B | Proof: partialIn the hosted demo | |
| ||||
Examiner prompt, speaker roles, one score per criterion, grounding check of the evidenceQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | StandardBest | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Lite tier: the same scorer on the FP8 checkpoint; captions become the transcript (roles guessed)Qwen3.8-27B (official FP8)Qwen/Qwen3.8-27B-FP8 on Hugging Face (opens in a new tab) 27.8BProof: strongSelf-host only | Lite | 27.8B | Proof: strongSelf-host only | |
| ||||
Best tier: scorer and grounding judgeDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yetSelf-host only | Best | 284B (13B active) · 192 GB | No proof yetSelf-host only | |
| ||||
The answer format and parser for each criterion's level and stated confidence.
Checks that the cited lines back what the model says the candidate said.
Verify the draft record in the browser without calling any server.
${DECOSA_REGISTRY}/decosa-api:0.1.0Live lanes, scoring, review and verify routes (/ws/live, /oral/*). No GPU. Binds 127.0.0.1 by default.
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.
${DECOSA_REGISTRY}/decosa-asr:0.1.0vLLM realtime endpoint for Voxtral Mini 4B Realtime. Not needed for text-only scoring.
MOSS-Transcribe-Diarize 0.9B speaker pass (decosa-api services/diarize). No published image yet; without it the captions become lines with guessed roles and no uncertainty marks.
Hosted demo layout on our server: Voxtral and MOSS-TD on GPU0, Qwen3.8-27B on GPU1 (two cards). The one-card split is the record tool's compose, not measured for this one.
Only the Qwen3.8-27B server is needed; verified self-hosted on 2026-09-25 against the running local server.
Not measured. FP8 LLM plus Voxtral; no speaker pass.
Measuredmeasured on our server 2026-09-25: a decosa-api pre-release test instance, POST /demo/replay of the demo scripts, one session at a time, gateway route shared with other workloads (1.1-1.6 s over 3 replays)
Measuredmeasured on our server 2026-09-25: a decosa-api pre-release test instance, POST /demo/replay of the demo scripts, one session at a time, gateway route shared with other workloads (10.4-12.0 s; the prompt waits for about 20 words)
Measuredmeasured on our server 2026-09-25: a decosa-api pre-release test instance, POST /demo/replay of the demo scripts, one session at a time, gateway route shared with other workloads (16.9 s and 17.6 s for 85 s and 60 s sessions; 81.8 s once for the 117 s session while the shared gateway was loaded: scoring took 55 s of it)
Measuredmeasured on our server 2026-09-25: 40 runs, two requests in parallel on the shared gateway, p90 19.4 s; 2.8 s for one request on a quiet gateway
Measuredmeasured on our server 2026-09-25: fresh clone, api container against the local Qwen3.8-27B server, one run
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa structured oral assessment on this machine
You are setting up a self-hosted oral-assessment assistant on this Linux machine, for vivas and structured interviews: live captions and an examiner prompt while the session runs, then (when it stops) examiner and candidate lines, words the two recognisers disagree on marked for a re-listen, draft rubric scores that cite the candidate's lines, and a signed draft record. The examiner then confirms or overrides every score with a reason, which is sealed in a review record chained to the draft. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:** the tool drafts, a person decides every score. Tell every candidate that the session is recorded and scored with AI help, and get consent where the law requires it. Oral exam recordings and scores are education records (FERPA); for hiring in New York City, Local Law 144 needs a bias audit and notice first. Repeat these points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/oral-assessment.zip (4 KB, 11 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py oral-assessment` (the api image carries the same bundle under /app/rehearsal/oral-assessment/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py oral-assessment --bundle oral-assessment.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "all 6 rubric criteria get a draft score", "the wrong p-value definition scores 0 or 1", "agreeing with the misreading (the null is probably true) scores 0 or 1"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `asr` | `${DECOSA_REGISTRY}/decosa-asr:0.1.0` (vLLM 0.27.1 + `mistral-common[audio]`) | `mistralai/Voxtral-Mini-4B-Realtime-2602`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` (optional) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |
Audio, transcripts, scores and records stay on this machine. If I only want to score transcripts I already have (no audio), skip `asr` and `diarize`: the `/oral/*` routes need only `llm` and `api`. Without `diarize`, live sessions still work, but roles are guessed from question marks and nothing is marked uncertain.
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4).
- Hopper (H100/H200): set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`.
- 48 GB Ada/L40S: FP8 checkpoint as above, plus `LLM_MAX_LEN=16384`, `LLM_GPU_UTIL=0.70`, `ASR_GPU_UTIL=0.22`, `DECOSA_LIVE_CAP=2`.
- Under 48 GB: stop and tell me it will not fit.
Only the Blackwell defaults have been measured; the other rows are starting points.
2. Check `docker --version` and `docker compose version`. If Docker is missing, install Docker Engine from Docker's official apt/dnf repository for this distro.
3. Check `docker run --rm --gpus all ubuntu nvidia-smi`. If it fails, install the NVIDIA Container Toolkit (`nvidia-container-toolkit`) from NVIDIA's repository, run `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
4. Confirm about 80 GB of free disk for images and weights.
## 2. Get the images
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,asr,api}:0.1.0`. If a pull fails (not published yet, or no access), build from source once the `decosa-api` source is published: clone it, then `docker compose build llm asr api` in the repo, which builds the same tags from `docker/`. If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-oral/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.60
ASR_GPU_UTIL=0.25
DECOSA_LLM_ROUTE=direct # local model; receipts are signed by this box's own key ("attested")
DECOSA_SIGNER_NAME="<who signs these records, e.g. Riverbend College, Statistics department>"
DECOSA_LIVE_CAP=4
```
Create `~/decosa-oral/docker-compose.yml` with exactly these services:
```yaml
name: decosa-oral
x-gpu: &gpu
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
<<: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
asr:
image: ${DECOSA_REGISTRY}/decosa-asr:${DECOSA_TAG}
<<: *gpu
ipc: host
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } } # start after llm so the memory split is stable
volumes: [hf-cache:/root/.cache/huggingface]
command: ["--model", "mistralai/Voxtral-Mini-4B-Realtime-2602", "--tokenizer-mode", "mistral", "--config-format", "mistral",
"--load-format", "mistral", "--compilation-config", '{"cudagraph_mode":"PIECEWISE"}', "--max-model-len", "45000",
"--max-num-batched-tokens", "8192", "--max-num-seqs", "16", "--gpu-memory-utilization", "${ASR_GPU_UTIL}",
"--served-model-name", "voxtral-realtime", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, asr: { condition: service_healthy } }
environment:
DECOSA_ASR_WS: ws://asr:8000/v1/realtime
DECOSA_LLM_ROUTE: ${DECOSA_LLM_ROUTE}
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LIVE_CAP: ${DECOSA_LIVE_CAP}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_AUDIO_S: "3600"
DECOSA_BUDGET_LLM_TOKENS: "200000"
DECOSA_SESSION_TTL_S: "28800"
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_DIARIZE_URL: "" # step 4 sets this for speaker labels and uncertainty marks
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
For transcripts only, delete the `asr` service and the `asr` line under `api.depends_on`. Run `docker compose up -d`, then poll `docker compose ps` until the services are healthy (the LLM takes 5–10 minutes the first time) and `curl -s localhost:8445/healthz` shows `"llm": true` (and `"asr": true` with audio). If `llm` runs out of memory, lower `LLM_GPU_UTIL` or `LLM_MAX_LEN`; the GPU shares must add up to less than about 0.9.
## 4. Optional: speaker labels and uncertainty marks (diarize service)
The `diarize` service (MOSS-Transcribe-Diarize 0.9B, Apache-2.0) has no published image yet. To add it, run `services/diarize` from the decosa-api source on this box (install with the `uv` commands at the top of `services/diarize/requirements.txt`, then `DIARIZE_HOST=172.17.0.1 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0 .venv-diarize/bin/python services/diarize/server.py`; `172.17.0.1` is the docker bridge address from `ip -4 addr show docker0`). Add `extra_hosts: ["host.docker.internal:host-gateway"]` to `api` and set `DECOSA_DIARIZE_URL: http://host.docker.internal:8092`; `/healthz` then shows `"diarize": true`. Lower `LLM_GPU_UTIL` to 0.55 so it fits (an estimate, not measured). If the source isn't available, skip this step.
## 5. The signing key
1. `curl -s localhost:8445/attest/signing-key` returns `{"scheme":"ed25519","pubkey":"<64 hex>",...}`. Show me the `pubkey`.
2. Tell me to **back up the `decosa-data` volume**: it holds the private key (`/data/attest/ed25519.pem`, mode 0600). An appeal months later needs the same key to show the records as issued here.
3. Receipts from this box say `status: "attested"`: signed by our own key. They prove nothing was changed after signing and who signed, not that a score is right.
## 6. Smoke test: score, review, verify, tamper
```bash
H='content-type: application/json'
TOKEN=$(curl -s localhost:8445/demo/session -H "$H" -d '{"vertical":"oral-assessment"}' | jq -r .token)
curl -s localhost:8445/oral/score -H "authorization: Bearer $TOKEN" -H "$H" -d '{"sample_id":"oral-stats-partial","stream":false}' > /tmp/oral.json
jq '{summary: .draft.summary, check: .draft.record_check, leading: [.ready.leading[].n], scores: [.draft.scores[] | {criterion, level, evidence: [.evidence[].n], flags}]}' /tmp/oral.json
jq '{draft: .draft.record, examiner: "Examiner 7", decisions: [.draft.scores[] | if .criterion == "q1-definition"
then {criterion, action: "override", level: 0, reason: "The student then agreed with the misreading at line 4."}
else {criterion, action: "confirm"} end]}' /tmp/oral.json > /tmp/review-req.json
curl -s localhost:8445/oral/review -H "authorization: Bearer $TOKEN" -H "$H" --data-binary @/tmp/review-req.json > /tmp/review.json
jq '{final_total, max_total, overrides, check}' /tmp/review.json
jq '{bundle: .bundle}' /tmp/review.json | curl -s localhost:8445/oral/verify -H "$H" --data-binary @- | jq '{ok, summary}'
jq '{bundle: (.bundle | (.review.entries[] | select(.kind=="decision" and .action=="override") | .text) |= "edited later")}' /tmp/review.json \
| curl -s localhost:8445/oral/verify -H "$H" --data-binary @- | jq '{ok, summary}'
```
Pass if: six scores with evidence line numbers, `record_check.ok` true, line 3 in `leading`; the review prints `check.ok: true` with `overrides: 1`; the first verify prints `ok: true`; the second prints `ok: false` naming the edited decision. Levels can differ by one from the hosted demo's.
With `asr` running, also try the audio path: `curl -sN localhost:8445/demo/replay -H "authorization: Bearer $TOKEN" -H "$H" -d '{"vertical":"oral-assessment","script_id":"oral-stats-partial"}'` streams `transcript` events, the `prompt`, `guard` and `scores` lanes, and a `done` whose `summary.draft` and `summary.record` feed the same review (68 s of audio, paced in real time).
## 7. Point the app at the local API
- Base URL: `http://localhost:8445` (for the web app, `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add the app's origin to `DECOSA_CORS_ORIGINS` if it is not localhost.
- Live: `ws://localhost:8445/ws/live?vertical=oral-assessment&token=<t>&rubric=intro-stats-viva`, 16 kHz mono PCM16 frames of about 100 ms, then `{"type":"stop"}`.
- Your own rubric: `POST /oral/score` with `rubric` (the shape of `GET /oral/rubrics/intro-stats-viva`) and `lines` or `transcript`. Pass `names` to remove the candidate's and examiners' names; for students under 18 set `minor: true` with `consent_basis`, which keeps text out of the signed records.
- Keep the bundle from `/oral/review` with the grade. Anyone can check it: `POST /oral/verify`.
- Keep the API on `127.0.0.1`; put a TLS reverse proxy in front for the LAN and set `DECOSA_TRUSTED_PROXIES`.
Finish with a summary: what is running, the health output, the signing key's pubkey, both verify results, and the consent, FERPA, LL144 and human-decides reminders.Loading the watch status…
10 laws, rules and guidance pages cited; 10 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Last reviewed
On 120 held-out synthetic test criteria the draft level matched the label 90.8% of the time and was always within one level (quadratic weighted kappa 0.971). One AI author wrote the answers and the labels, so treat this as an upper bound on real vivas.
No. It drafts a level for each rubric criterion with the candidate's lines cited, and the examiner confirms or overrides every criterion, with a reason for an override. The final total comes only from the examiner's decisions.
As a drafting aid, with the teacher deciding. Here the model drafts a level for each criterion and cites the candidate's words; the examiner confirms or overrides every one, and only those decisions make the grade. Exam recordings and scores are FERPA education records, so run real exams self-hosted or under the school-official agreement. The EU AI Act lists evaluating learning outcomes as high-risk, with duties from 2 Dec 2027. Not legal advice.
On 120 held-out synthetic criteria the draft level matched the label 90.8% of the time and was always within one level (quadratic weighted kappa 0.971). The labels were written by an AI agent and the data is synthetic, so treat it as an upper bound on real vivas.
In the eval, strong answers in plain, non-native English scored the same as fluent ones (24 of 24 each). That was measured on written text, not real accented speech, which is untested.
A signed bundle: the draft record with the cited evidence and the examiner's decisions with reasons, in a hash chain. Changing one word of the transcript or a reason makes verification fail at that entry.
Exam recordings and scores are FERPA education records, so real exams should run self-hosted, where nothing leaves the box. The server keeps receipts (hashes), not text, and for students under 18 the records keep hashes only.
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…