Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Decision records (use case 184): parity and tamper eval

Generated by scripts/eval_decision_record.py on 2026-09-29 (Node v24.18.0, Python 3.12.3). Every number below comes from that run and from docs/evals/decision-record.json; re-run the script to update both. Synthetic data only; no model calls and no network.

What was tested

The TypeScript kit @decosa/decision-record (clients/decision-record) and decosa-api's Python record format (hashchain.py, attest.py, decision_record.py) must make and check the same records. The kit records automated steps about people as hashes, the app's own ids and a keyed reference per person; it never holds personal data.

Parity on the shared fixture

  • Entries in the fixture: 10.
  • Python rebuilds the TypeScript record byte for byte: yes; same Ed25519 signature: yes.
  • TypeScript rebuilds its committed fixture: yes.

Tampering

12 kinds of tampering, each applied to a record made by TypeScript and to one made by Python (24 cases). Both verifiers rejected 24 of 24. They agreed on the verdict, the first bad entry and the failing checks in 24 of 24.

Tampering What the Python verifier says TypeScript agrees
a result changed Verification failed at entry 3, automated step (score): entry fields were changed: the entry hash does not recompute. yes
an entry deleted Verification failed at entry 5, automated step (send): sequence number is 6, expected 5. yes
two entries swapped Verification failed at entry 6, automated step (send): sequence number is 7, expected 6. yes
a link rewired Verification failed at entry 2, automated step (screen): chain link broken: prev does not equal the previous entry's hash. yes
the signed count changed Verification failed: count, signature. yes
the signature replaced Verification failed: signature. yes
a checkpoint forged Verification failed: checkpoints. yes
free text added Verification failed at entry 4, automated step (rank): text was changed: its sha256 no longer matches text_sha256. yes
a raw id instead of the keyed ref Verification failed at entry 3, automated step (score): entry fields were changed: the entry hash does not recompute. yes
a float in a result Verification failed at entry 3, automated step (score): $.result.score: use an integer (for example per mille), not a float. yes
an entry appended after sealing Verification failed at entry 10, automated step (other): sequence number is 9, expected 10. yes
not a record This is not a Decosa session record. yes

The personal-data guard

34 shared cases (16 should be refused: names, emails, URLs, phone numbers, free text, floats, deep nesting). Python matched the expected verdict in 34, TypeScript in 34, and the two agreed in 34.

Known limits: it is a guard against mistakes, not a proof. A short label can still be personal data if a caller puts it there. Date-times written with a space ("2026-09-29 10:00") are refused as possible phone numbers; use ISO 8601.

A larger synthetic log

600 steps about 60 synthetic people (screens, scores, rankings, follow-ups sent without approval, human decisions).

  • Entries: 601; signed checkpoints: 25.
  • Python and TypeScript build identical records: yes.
  • Python verifies the TypeScript record: yes; TypeScript verifies the Python record: yes.
  • Record size: 359,544 bytes, about 598 bytes per entry.
  • Synthetic ids and names searched for in the record: 0 found.

Time on this machine

  • Python: build 23 ms, verify 23 ms.
  • Node (includes starting Node): build 120 ms, verify 100 ms; starting Node and verifying the small fixture, median of 5: 64 ms.

What this does not show

  • That a model ran, or that its output was right. The signature is the operator's attestation over the recorded hashes and results.
  • That a hiring process is lawful. The records support notices, explanations and retention; what a law requires is for the employer and counsel.
  • Anything about an app's real candidates: this eval uses synthetic data only.