Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: signed lab notebook (vertical 44)

Run 25-26 Sep 2026 on our server (CPU for the eval; the end-to-end runs used the shared Qwen3.8-27B gateway). Script: scripts/notebook_eval.py (results in docs/evals/signed-lab-notebook/results.json and real-export.json). Nothing was tuned against these numbers: the verifier's rules were written first, the eval then run, and the only changes made after a run were to the attacker (making the insider smarter, below), never to the verifier to pass a case.

Data

Synthetic notebooks from decosa_api/verticals/notebook/testkit.py: members, entries with instrument files attached by SHA-256, amendments, AI analyses (with locally attested model receipts), witness, review and approval signatures on about 70% of entries, and RFC 3161 tokens from an offline test Time-Stamp Authority (real DER tokens signed with a throwaway EC key, so the verifier's full TSA path runs without the network). Plus one real export: the demo notebook recorded against the pre-release server (a gateway-signed receipt for the AI analysis, personal Ed25519 keys for the witness and the PI, and a real FreeTSA token). No personal or real lab data. Licences: all synthetic, generated here.

Genuine exports verify

300 / 300 synthetic exports (5 to 60 numbered entries, all three configurations) verify; median check 4.4 ms in Python. The real recorded export verifies in Python and in the browser verifier (Node 26 WebCrypto), with the FreeTSA token's issuer matching the pinned FreeTSA root.

Python and TypeScript agree on 219 / 219 exports (genuine and altered, all configurations): same pass or fail.

Tamper detection

Two attackers, 40 notebooks per configuration, 13 alterations in five groups:

  • edit: entry text, author, an attachment hash, the AI's output, an AI analysis relabelled as a human entry;
  • delete: an entry (the insider also deletes the signatures and amendments that depend on it);
  • reorder: two adjacent entries swapped (the insider keeps the times in order);
  • backdate: an entry moved a day earlier; a new entry slipped in among old ones with a matching old time;
  • forged signature: meaning changed (review to approval), random signature bytes, an approval added under a member's name with another key, a personal-key signature replaced by an unsigned "account" one.

The outsider edits the exported file. The insider holds the server's signing key: after the change it renumbers, repairs every reference, recomputes all hashes, re-signs the checkpoints, the export and any model receipt the server signed itself, and deletes timestamp entries whose tokens no longer match. It does not hold members' personal keys and cannot get a TSA token for an earlier time.

Configuration Attacker edit delete reorder backdate forged signature
personal keys + TSA outsider 200/200 40/40 40/40 80/80 160/160
personal keys + TSA insider 200/200 38/40 39/40 78/80 160/160
account signatures + TSA outsider 200/200 40/40 40/40 80/80 80/80
account signatures + TSA insider 29/200 2/40 14/40 40/80 40/80
account signatures, no TSA outsider 200/200 40/40 40/40 80/80 80/80
account signatures, no TSA insider 29/200 2/40 16/40 40/80 40/80

Checked against an earlier export of the same notebook as well (the verify page's optional second file), every insider alteration is caught: 1,400 / 1,400. (An appended forgery extends the earlier copy rather than rewriting it; those were caught by the signature checks.)

What this means:

  • Anyone editing an export without keys is caught every time (1,400 / 1,400), and the failure names the entry.
  • Against the operator, what helps is something the operator does not control. With personal e-signature keys and TSA, 515 of 520 insider changes were caught on the export alone; the 5 misses were near the end of a notebook, where no personal signature or surviving timestamp covers the changed entries. An earlier export kept by a witness or an inspector caught all of them.
  • A TSA token is only as good as its survival: an insider can delete tokens that stop matching, which is why the account-signature rows look the same with and without TSA. The verify page reports how much of the chain the last surviving token covers. A token kept outside the notebook (or a published copy of the head) would close this; nothing like that is offered yet.
  • On the hosted route the AI analysis carries the gateway's receipt, a second signer: on the real export the insider's edit of the AI output was caught ("gateway receipt covers a different output"), while in the synthetic notebooks (receipts signed by the same server key) the insider re-signed it.
  • Real export, insider: 10 of 12 alterations caught on the export alone (the misses: deleting the amendment together with its two signatures, and swapping two adjacent entries' contents); all 12 outsider alterations caught.

Performance at 10,000 entries

One notebook with 10,000 numbered entries (17,107 chain entries including 7,082 personal-key signatures, 3 member enrolments and 21 RFC 3161 tokens, one every 500 entries), 2,545 AI analyses:

  • build, including every write-time check and signature verification: 9.9 s (4.1 s in a quieter run), Python, one core;
  • export and sign: 0.15 s; export size 17.1 MB;
  • verify in Python: 2.6 s; verify in the browser code under Node 26: 4.7 s.

The verify page takes files up to 64 MB. POST /notebook/verify accepts up to DECOSA_RECORD_MAX_BYTES (8 MB by default), so a notebook this size is checked in the browser, or the server limit is raised on self-host.

End-to-end (real model, real TSA)

scripts/smoke/signed-lab-notebook.py against the pre-release server (gateway route): 12 chain entries, one AI analysis (1,567 tokens, $0.00105 at list price), signed receipt covering the stored output, FreeTSA timestamp, export verified, and the export with one data value changed in the AI's prompt rejected. Whole run 22.4 s, of which the model call took 22.2 s on the shared, busy gateway. scripts/notebook_record_demo.py recorded the site's replay the same way in 23.7 s.

On that run the model's analysis named the outlying well (F7, 1.6 mM, 61.23 against 92.88 and 94.10) and declined to report Km and Vmax without addressing it. The number check found 13 numbers in the output: 7 appear in the data, 4 are roundings of data values, and 2 (a calculated 8.5-fold ratio and a 0.998 threshold) do not, which is what the reviewer is pointed to. This is one run, not an accuracy measurement of the analysis itself.

Not measured

  • The quality of AI analyses in general (the product records and receipts them; it does not grade them).
  • Real-world clock drift against the TSA (the check allows 120 s).
  • Revocation status of TSA certificates (not checked; the issuer is pinned instead).