{"schema_version":"1","site":"https://decosa.ai","id":"oral-assessment","num":"35","name":"Structured oral assessment","tool_name":"Score an oral exam against a rubric","short":"Oral assessment","blurb":"Vivas and structured interviews scored against a rubric. While it runs, the examiner sees which criteria still lack evidence and neutral follow-ups. Afterwards each criterion gets a draft score with the candidate's words cited and checked, the examiner confirms or overrides every score with a reason, and both steps are sealed in signed records that survive an appeal. Examiner side only.","status":"live","labels":{"industry":["education","hr-recruiting"],"job":["review","attest"],"input":["voice","text"],"deploy":["hosted","selfhost"],"status":"live","output":["record","data"],"data":["pii","minors"],"hardware":"gpu-96","licence":"permissive"},"industries":["education","hr-recruiting"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["live-asr","diarize","signed-record"],"models":"Voxtral 4B · MOSS-TD 0.9B · Qwen3.8-27B","where":"Self-host for real students and candidates; hosted for pilots on synthetic or consented data","hardware":"1× RTX PRO 6000 (96 GB) for Voxtral, the diarizer and Qwen3.8-27B; text-only scoring needs only the Qwen server","final_artifact":"Draft scores with cited evidence, then the examiner's decisions: two signed, chained records (the bundle) that an appeals panel can re-check.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Receipt per caption, line and model call; signed draft and review records, chained","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":17600,"p95_ms":null,"runs":null,"receipts_per_run":42,"cost_per_run_usd":0.007},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone, api image built, the prompt's api service (named volume) against the running local model servers, then torn down","notes":"The step 6 smoke passed as written: six scores with cited lines, the leading question at line 3 flagged, one override signed, the bundle verified, and after editing the override's reason verification failed at that decision. The audio replay also passed (speaker lines, signed draft of 87 entries). Model-server startup itself not re-verified (no new GPU load)."},"known_limits":["Scores are drafts. On 120 held-out synthetic criteria they matched labels written by an AI agent (Claude) 90.8% of the time and were never more than one level off; real answers and real examiners are not measured yet.","The review flags did not catch the model's disagreements (0 of 11), and on the hosted route the stated confidence is nearly always 0.95; the examiner has to decide every criterion.","Speech was measured on synthetic TTS voices only. Words the two recognisers disagree on are marked (81% of real errors found in the eval), but accented or overlapping human speech is untested.","Rubrics are JSON: two built in, custom ones through the API; no rubric editor in the console yet.","No LMS or ATS export yet; keep the bundle JSON with the grade."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Draft level equals the label (exact)","value":"90.8%","unit":null,"n":120,"split":"test","note":"Within one level: 100%. 120 criteria come from 60 distinct answer-criterion pairs."},{"name":"Quadratic weighted kappa vs labels","value":"0.971","unit":null,"n":120,"split":"test","note":null},{"name":"Misconception or mixed answers scored exactly","value":"70.8%","unit":null,"n":null,"split":"test","note":"Strong 100%, partial 91.7%, weak 91.7%."},{"name":"Fluency penalty: plain non-native English vs fluent strong answers","value":"none measured (24 of 24 each)","unit":null,"n":24,"split":"test","note":"Written text, not real accented speech."},{"name":"Test-retest, same level","value":"118 of 120 (98.3%)","unit":null,"n":120,"split":"test","note":null},{"name":"Real recognition errors found by the disagreement marks: precision / recall","value":"65% / 74%","unit":null,"n":36,"split":"synthetic","note":"27 errors on synthetic TTS voices, clean and in noise. Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M, American and British English only) and re-run; the first build (macOS voices, six English accents): 53% / 81%, 36 errors."}],"dataset":"20 synthetic transcripts (10 per rubric: intro-statistics viva and customer-support interview), 6 criteria each, 120 scored criteria per run; labels written with the answers before any model run. Prompts developed on the four demo scripts only; the test set was not used for tuning.","held_out":true,"caveats":["Labels are one AI author's (the building agent), who also wrote the answers to hit the levels: agreement here is an upper bound on real vivas.","Synthetic only; no examiner labels and no examiner-examiner agreement to compare with.","Review flags did not predict the model's disagreements (0 of 11 flagged); the examiner must decide every criterion.","Speech measured only on synthetic voices; real accented, fast or overlapping speech will have many more recognition errors.","Only the two built-in rubrics were tested."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/oral-assessment"},"stack":{"summary":"For departments bringing back oral exams, certification bodies, and hiring teams running structured interviews. During the session the examiner sees which rubric criteria still lack evidence and neutral follow-ups, and leading questions are flagged. Afterwards an open diarization model writes who said what, the words two recognisers disagree on are marked for a re-listen, names and personal details are removed, and the language model drafts a score for each criterion that cites the candidate's lines, checked by a grounding judge. The examiner confirms or overrides every score with a reason. Both steps are sealed in signed, chained records. Examiner side only; the tool never grades on its own.","tagline":"Vivas and structured interviews scored against a rubric, with the candidate's words cited, the examiner deciding every score, and records that survive an appeal.","deployment":"hosted-or-self-host","regulatory_note":"As of 25 Sep 2026, not legal advice. Education: exam recordings, transcripts and scores are education records under FERPA (34 CFR Part 99); a vendor gets them only as a school official under the institution's control, so run real exams self-hosted or under that agreement. Students under 18: set minor (hashes-only records) and self-host; COPPA applies to services directed at children under 13. EU AI Act Annex III lists evaluating learning outcomes (point 3) and evaluating candidates (point 4) as high-risk; those duties apply from 2 Dec 2027 after the AI Omnibus. Hiring in New York City: Local Law 144 requires a bias audit and candidate notice before an automated tool substantially assists the decision; this tool does not collect demographics, so it cannot run that audit itself. A person decides every score (GDPR Art. 22). Tell everyone the session is recorded and get consent where the law requires it.","components":[{"id":"asr-live","role":"Live captions (streaming, no speakers): the examiner prompt reads these","name":"Voxtral Mini 4B Realtime","hf_repo":"mistralai/Voxtral-Mini-4B-Realtime-2602","license":"Apache-2.0","params":"4.4B","quant":"BF16","vram_gb":24,"memory_gb_estimate":null,"engine":"vLLM realtime WebSocket (/v1/realtime)","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["all"],"alternative_to":null},{"id":"asr-pass2","role":"After the session: examiner and candidate lines, each with a receipt over its audio","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"transformers (trust_remote_code) + moss_transcribe_diarize package; hosted: decosa-diarize service, 127.0.0.1:8092, GPU0","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard","best"],"alternative_to":null},{"id":"llm","role":"Examiner prompt, speaker roles, one score per criterion, grounding check of the evidence","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard","best"],"alternative_to":null},{"id":"llm-lite","role":"Lite tier: the same scorer on the FP8 checkpoint; captions become the transcript (roles guessed)","name":"Qwen3.8-27B (official FP8)","hf_repo":"Qwen/Qwen3.8-27B-FP8","license":"Apache-2.0","params":"27.8B","quant":"FP8","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (same decosa-llm image)","receipt_coverage":"strong","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"llm-best","role":"Best tier: scorer and grounding judge","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts + FP8 (about 159–176 GB of weights)","vram_gb":192,"memory_gb_estimate":null,"engine":"vLLM B12X community build, TP2 on 2× RTX PRO 6000, MTP draft fixed by the kit's patches","receipt_coverage":"none","in_hosted_demo":false,"tiers":["best"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card, captions only","summary":"Live captions and the examiner prompt, then scores from the captions with roles guessed from question marks. No speaker labels and no uncertainty marks.","components":["asr-live","llm-lite"],"hardware":"1x L40S or RTX 6000 Ada 48 GB (not measured)","quality_evidence":[{"metric":"scoring agreement from captions with guessed roles","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Direct route: model calls and captions are attested by the box's own key; no gateway receipts.","hosting":null},{"id":"standard","label":"Standard · the hosted demo, two recognisers","summary":"Captions and the examiner prompt live; then speaker lines, uncertainty marks, blinding, cited scores with a grounding check, and the signed records. Every model call carries a gateway-signed receipt.","components":["asr-live","asr-pass2","llm"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB (measured on two cards on our server)","quality_evidence":[{"metric":"draft level vs labels, 120 held-out criteria: exact / within one / QWK","value":"90.8% / 100% / 0.971","source":"decosa-api docs/evals/oral-assessment.md, 2026-09-25; synthetic answers, labels by Claude (an AI agent), not examiners"},{"metric":"fluency penalty: strong answers in plain, non-native English","value":"none measured (24 of 24 same level)","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"evidence: levels above 0 citing the right answer; grounding supported","value":"100%; 99.2%","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"test-retest same level","value":"98.3%","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"10% of words mis-recognised: levels changed; lowered scores flagged when marked","value":"7.5%; 11 of 12","source":"decosa-api docs/evals/oral-assessment.md"},{"metric":"real recognition errors found by the disagreement marks","value":"74% recall, 65% precision (27 errors, synthetic voices)","source":"decosa-api docs/evals/oral-assessment.md"}],"latency_note":"measured on our server on the shared gateway: the signed draft seconds after the audio ends; the text route in seconds with two requests in parallel","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Language-model calls: gateway-signed receipts, embedded in the draft record. Speech: attested receipts from decosa-api's key.","hosting":null},{"id":"best","label":"Best · DeepSeek-V4-Flash scores and checks","summary":"Standard with a larger scorer and grounding judge on two more 96 GB cards.","components":["asr-live","asr-pass2","llm-best"],"hardware":"2x RTX PRO 6000 96 GB for the scorer plus the standard card","quality_evidence":[{"metric":"scoring agreement","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Not a hosted model: model calls are attested by the box's key only.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Live lanes, scoring, review and verify routes (/ws/live, /oral/*). No GPU. Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network."},{"name":"decosa-asr","port":8000,"image":"${DECOSA_REGISTRY}/decosa-asr:0.1.0","purpose":"vLLM realtime endpoint for Voxtral Mini 4B Realtime. Not needed for text-only scoring."},{"name":"decosa-diarize","port":8092,"image":null,"purpose":"MOSS-Transcribe-Diarize 0.9B speaker pass (decosa-api services/diarize). No published image yet; without it the captions become lines with guessed roles and no uncertainty marks."}],"tools":[{"name":"Typed-judgment engine (Decosa tool 24)","url":null,"license":"AGPL-3.0-or-later (decosa-api)","purpose":"The answer format and parser for each criterion's level and stated confidence."},{"name":"Grounding check (Decosa tool 17)","url":null,"license":"AGPL-3.0-or-later (decosa-api)","purpose":"Checks that the cited lines back what the model says the candidate said."},{"name":"WebCrypto Ed25519 and SHA-256","url":"https://developer.mozilla.org/en-US/docs/Web/API/SubtleCrypto/verify","license":"browser built-in","purpose":"Verify the draft record in the browser without calling any server."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"The layout on our server when this was measured: Voxtral and MOSS-TD on GPU0, Qwen3.8-27B on GPU1 (two cards). The one-card split is the record tool's compose, not measured for this one."},{"tier":"Text-only scoring (pasted transcripts)","fits":true,"notes":"Only the Qwen3.8-27B server is needed; verified self-hosted on 2026-09-25 against the running local server."},{"tier":"1x L40S / RTX 6000 Ada 48 GB, lite tier","fits":null,"notes":"Not measured. FP8 LLM plus Voxtral; no speaker pass."}],"latency":[{"lane":"first caption","typical_ms":1200,"source":"measured on our server 2026-09-25: a decosa-api pre-release test instance, POST /demo/replay of the demo scripts, one session at a time, gateway route shared with other workloads (1.1-1.6 s over 3 replays)"},{"lane":"first examiner prompt","typical_ms":11800,"source":"measured on our server 2026-09-25: a decosa-api pre-release test instance, POST /demo/replay of the demo scripts, one session at a time, gateway route shared with other workloads (10.4-12.0 s; the prompt waits for about 20 words)"},{"lane":"signed draft after the audio ends","typical_ms":17600,"source":"measured on our server 2026-09-25: a decosa-api pre-release test instance, POST /demo/replay of the demo scripts, one session at a time, gateway route shared with other workloads (16.9 s and 17.6 s for 85 s and 60 s sessions; 81.8 s once for the 117 s session while the shared gateway was loaded: scoring took 55 s of it)"},{"lane":"text route, 6 criteria (12 calls)","typical_ms":15400,"source":"measured on our server 2026-09-25: 40 runs, two requests in parallel on the shared gateway, p90 19.4 s; 2.8 s for one request on a quiet gateway"},{"lane":"text route, self-hosted (direct route with logprobs)","typical_ms":6900,"source":"measured on our server 2026-09-25: fresh clone, api container against the local Qwen3.8-27B server, one run"}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"Measured latency","value":"The signed draft arrives seconds after the audio ends (median of hosted replays, gateway route, one session at a time); longer when the shared gateway is loaded."},{"label":"Typical run cost","value":"A fraction of a cent per short viva on the hosted route (a couple of dozen model calls at the gateway list price), plus a speech receipt per caption and per transcript line signed by the server's key at no model cost. A pasted transcript costs less. Each run shows its own measured cost."},{"label":"What the examiner gets","value":"Per criterion: a draft level, the candidate's lines it rests on, a grounding check of that evidence and flags (re-listen, leading question, low confidence). Then a signed bundle: the draft record and the examiner's decisions with reasons."},{"label":"Fairness guards","value":"The scorer never sees names, pronouns or stated personal details (0 of 32 planted ones survived). It scores content, not fluency: strong answers in plain, non-native English scored the same as fluent ones (24 of 24)."},{"label":"Data retention","value":"Nothing stored: transcripts and records live in memory for the request or session and come back to you. The server keeps receipts (hashes), not text. For students under 18 the records keep hashes only."},{"label":"What leaves the box (self-host)","value":"Nothing, with DECOSA_LLM_ROUTE=direct. Anyone can re-check a bundle with POST /oral/verify or a single record in the browser."}],"hosted_now":{"needs":["diarize","live-asr","qwen3.8-27b"],"off":["diarize","live-asr"],"live_by_default":false,"live_status":"https://api.decosa.ai/status"},"data_handling":{"page":"/data#oral-assessment","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: transcripts and records live in memory for the request or session and come back to you. The server keeps receipts (hashes), not text. For students under 18 the records keep hashes only.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/operations/oral-assessment","input":"oral","lanes":[{"id":"prompt","title":"Examiner prompt","kind":"markdown"},{"id":"guard","title":"Consistency guard","kind":"markdown"},{"id":"chain","title":"Record chain","kind":"markdown"},{"id":"final_transcript","title":"Transcript (speakers)","kind":"markdown"},{"id":"scores","title":"Draft scores (cited)","kind":"markdown"},{"id":"record","title":"Signed draft record","kind":"markdown"}],"samples":[{"n":1,"id":"oral-stats-partial","title":"Statistics viva: partly correct, with a leading question (synthetic)","deep_link":"/tools/operations/oral-assessment?sample=1&autorun=0"},{"n":2,"id":"oral-stats-strong","title":"Statistics viva: strong answers, non-native speaker (synthetic)","deep_link":"/tools/operations/oral-assessment?sample=2&autorun=0"},{"n":3,"id":"oral-stats-weak","title":"Statistics viva: weak answers (synthetic)","deep_link":"/tools/operations/oral-assessment?sample=3&autorun=0"},{"n":4,"id":"oral-support-interview","title":"Support specialist interview: mixed answers (synthetic)","deep_link":"/tools/operations/oral-assessment?sample=4&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/oral-assessment-hosted.md","selfhost":"/prompts/oral-assessment-selfhost.md","assemble":"/prompts/oral-assessment-assemble.md","mac":"/prompts/oral-assessment-mac.md"},"rehearsal":{"bundle":"/samples/oral-assessment.zip","bundle_url":"https://decosa.ai/samples/oral-assessment.zip","folder":"/samples/oral-assessment/","expected":"/samples/oral-assessment/expected.json","files":["/samples/oral-assessment/expected.json","/samples/oral-assessment/inputs/rubric-intro-stats-viva.json","/samples/oral-assessment/inputs/viva-transcript.json"],"bytes":4137,"checks":["all 6 rubric criteria get a draft score","the wrong p-value definition scores 0 or 1","agreeing with the misreading (the null is probably true) scores 0 or 1","the vague 'bigger study' design scores 0 or 1","the interval-width score cites the student's answer (line 6)","the interval-width answer (larger sample, narrower interval) scores 2 or 3","the examiner's leading question at line 3 is flagged","the signed draft record passes the server's own check","the review bundle (draft + examiner decisions, one override) verifies","a bundle with the override reason edited no longer verifies","every model call has a signed receipt"],"licence":"Synthetic: a viva script written for Decosa (no real student or examiner) and the built-in intro-statistics rubric. Part of decosa-api, which will be released under AGPL-3.0-or-later; until then the source is on request.","about":"A short synthetic intro-statistics viva in which the student gives partly correct answers and agrees with a leading question from the examiner (that a p-value is the probability the null is true). The scorer must score all six rubric criteria with cited lines, give low marks where the answers are wrong or vague, flag the leading question, and seal a signed draft; an examiner's review with one override must verify, and an edited override reason must be caught.","run":{"containers":"docker compose exec api python scripts/rehearse.py oral-assessment","checkout":"python scripts/rehearse.py oral-assessment --bundle oral-assessment.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py oral-assessment"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=oral-assessment","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":85.6,"basis":"estimate","unknown":[]},{"id":"best","gpu_gb":220,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":48}},"links":{"page":"/tools/operations/oral-assessment","json":"/use-cases/oral-assessment.json","metrics":"/metrics/oral-assessment","console":"/tools/operations/oral-assessment","console_sample":"/tools/operations/oral-assessment?sample=1&autorun=0","stack":"/tools/operations/oral-assessment#stack","try_live":"/tools/operations/oral-assessment","watch":"/tools/operations/oral-assessment","build":"/tools/operations/oral-assessment#build","self_host":"/tools/operations/oral-assessment#self-host","prompts":{"hosted":"/prompts/oral-assessment-hosted.md","selfhost":"/prompts/oral-assessment-selfhost.md","assemble":"/prompts/oral-assessment-assemble.md","mac":"/prompts/oral-assessment-mac.md"}}}