Skip to content
decosa

35 · Education · HR and recruiting · live

Structured oral assessment

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Draft level equals the label (exact)90.8%test splitn = 120Within one level: 100%. 120 criteria come from 60 distinct answer-criterion pairs.
  • Quadratic weighted kappa vs labels0.971test splitn = 120
  • Misconception or mixed answers scored exactly70.8%test splitStrong 100%, partial 91.7%, weak 91.7%.
  • Fluency penalty: plain non-native English vs fluent strong answersnone measured (24 of 24 each)test splitn = 24Written text, not real accented speech.
  • Test-retest, same level118 of 120 (98.3%)test splitn = 120
  • Real recognition errors found by the disagreement marks: precision / recall65% / 74%syntheticn = 3627 errors on synthetic TTS voices, clean and in noise. Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M, American and British English only) and re-run; the first build (macOS voices, six English accents): 53% / 81%, 36 errors.

Dataset

20 synthetic transcripts (10 per rubric: intro-statistics viva and customer-support interview), 6 criteria each, 120 scored criteria per run; labels written with the answers before any model run. Prompts developed on the four demo scripts only; the test set was not used for tuning.

Caveats

  • Labels are one AI author's (the building agent), who also wrote the answers to hit the levels: agreement here is an upper bound on real vivas.
  • Synthetic only; no examiner labels and no examiner-examiner agreement to compare with.
  • Review flags did not predict the model's disagreements (0 of 11 flagged); the examiner must decide every criterion.
  • Speech measured only on synthetic voices; real accented, fast or overlapping speech will have many more recognition errors.
  • Only the two built-in rubrics were tested.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
18 s
Receipts
42
Model calls
n/a
Tokens
n/a
Cost per run
$0.007

Self-host verification

Verified on 25 Sep 2026: fresh clone, api image built, the prompt's api service (named volume) against the running local model servers, then torn down

The step 6 smoke passed as written: six scores with cited lines, the leading question at line 3 flagged, one override signed, the bundle verified, and after editing the override's reason verification failed at that decision. The audio replay also passed (speaker lines, signed draft of 87 entries). Model-server startup itself not re-verified (no new GPU load).

Rehearsal bundle: oral-assessment.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Scores are drafts. On 120 held-out synthetic criteria they matched labels written by an AI agent (Claude) 90.8% of the time and were never more than one level off; real answers and real examiners are not measured yet.
  • The review flags did not catch the model's disagreements (0 of 11), and on the hosted route the stated confidence is nearly always 0.95; the examiner has to decide every criterion.
  • Speech was measured on synthetic TTS voices only. Words the two recognisers disagree on are marked (81% of real errors found in the eval), but accented or overlapping human speech is untested.
  • Rubrics are JSON: two built in, custom ones through the API; no rubric editor in the console yet.
  • No LMS or ATS export yet; keep the bundle JSON with the grade.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card, captions only (1)
  • scoring agreement from captions with guessed roles: not measured yet
Standard · the hosted demo, two recognisers (6)
  • draft level vs labels, 120 held-out criteria: exact / within one / QWK: 90.8% / 100% / 0.971decosa-api docs/evals/oral-assessment.md, 2026-09-25; synthetic answers, labels by Claude (an AI agent), not examiners
  • fluency penalty: strong answers in plain, non-native English: none measured (24 of 24 same level)decosa-api docs/evals/oral-assessment.md
  • evidence: levels above 0 citing the right answer; grounding supported: 100%; 99.2%decosa-api docs/evals/oral-assessment.md
  • test-retest same level: 98.3%decosa-api docs/evals/oral-assessment.md
  • 10% of words mis-recognised: levels changed; lowered scores flagged when marked: 7.5%; 11 of 12decosa-api docs/evals/oral-assessment.md
  • real recognition errors found by the disagreement marks: 74% recall, 65% precision (27 errors, synthetic voices)decosa-api docs/evals/oral-assessment.md
Best · DeepSeek-V4-Flash scores and checks (1)
  • scoring agreement: not measured yet

How we measure · All tools