35 · Education · HR and recruiting · live
Structured oral assessment
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Draft level equals the label (exact)90.8%test splitn = 120Within one level: 100%. 120 criteria come from 60 distinct answer-criterion pairs.
- Quadratic weighted kappa vs labels0.971test splitn = 120
- Misconception or mixed answers scored exactly70.8%test splitStrong 100%, partial 91.7%, weak 91.7%.
- Fluency penalty: plain non-native English vs fluent strong answersnone measured (24 of 24 each)test splitn = 24Written text, not real accented speech.
- Test-retest, same level118 of 120 (98.3%)test splitn = 120
- Real recognition errors found by the disagreement marks: precision / recall65% / 74%syntheticn = 3627 errors on synthetic TTS voices, clean and in noise. Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M, American and British English only) and re-run; the first build (macOS voices, six English accents): 53% / 81%, 36 errors.
Dataset
20 synthetic transcripts (10 per rubric: intro-statistics viva and customer-support interview), 6 criteria each, 120 scored criteria per run; labels written with the answers before any model run. Prompts developed on the four demo scripts only; the test set was not used for tuning.
Caveats
- Labels are one AI author's (the building agent), who also wrote the answers to hit the levels: agreement here is an upper bound on real vivas.
- Synthetic only; no examiner labels and no examiner-examiner agreement to compare with.
- Review flags did not predict the model's disagreements (0 of 11 flagged); the examiner must decide every criterion.
- Speech measured only on synthetic voices; real accented, fast or overlapping speech will have many more recognition errors.
- Only the two built-in rubrics were tested.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 18 s
- Receipts
- 42
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.007
Self-host verification
Verified on 25 Sep 2026: fresh clone, api image built, the prompt's api service (named volume) against the running local model servers, then torn down
The step 6 smoke passed as written: six scores with cited lines, the leading question at line 3 flagged, one override signed, the bundle verified, and after editing the override's reason verification failed at that decision. The audio replay also passed (speaker lines, signed draft of 87 entries). Model-server startup itself not re-verified (no new GPU load).
Rehearsal bundle: oral-assessment.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Scores are drafts. On 120 held-out synthetic criteria they matched labels written by an AI agent (Claude) 90.8% of the time and were never more than one level off; real answers and real examiners are not measured yet.
- The review flags did not catch the model's disagreements (0 of 11), and on the hosted route the stated confidence is nearly always 0.95; the examiner has to decide every criterion.
- Speech was measured on synthetic TTS voices only. Words the two recognisers disagree on are marked (81% of real errors found in the eval), but accented or overlapping human speech is untested.
- Rubrics are JSON: two built in, custom ones through the API; no rubric editor in the console yet.
- No LMS or ATS export yet; keep the bundle JSON with the grade.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Live captions (streaming, no speakers): the examiner prompt reads theseVoxtral Mini 4B RealtimeApache-2.0
- After the session: examiner and candidate lines, each with a receipt over its audioMOSS-Transcribe-Diarize 0.9BApache-2.0
- Examiner prompt, speaker roles, one score per criterion, grounding check of the evidenceQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card, captions only (1)
- scoring agreement from captions with guessed roles: not measured yet
Standard · the hosted demo, two recognisers (6)
- draft level vs labels, 120 held-out criteria: exact / within one / QWK: 90.8% / 100% / 0.971decosa-api docs/evals/oral-assessment.md, 2026-09-25; synthetic answers, labels by Claude (an AI agent), not examiners
- fluency penalty: strong answers in plain, non-native English: none measured (24 of 24 same level)decosa-api docs/evals/oral-assessment.md
- evidence: levels above 0 citing the right answer; grounding supported: 100%; 99.2%decosa-api docs/evals/oral-assessment.md
- test-retest same level: 98.3%decosa-api docs/evals/oral-assessment.md
- 10% of words mis-recognised: levels changed; lowered scores flagged when marked: 7.5%; 11 of 12decosa-api docs/evals/oral-assessment.md
- real recognition errors found by the disagreement marks: 74% recall, 65% precision (27 errors, synthetic voices)decosa-api docs/evals/oral-assessment.md
Best · DeepSeek-V4-Flash scores and checks (1)
- scoring agreement: not measured yet