44 · Science and research · Compliance and trust · live
Signed lab notebook
Eval results
Not held outRun 26 Sep 2026Eval write-up (decosa-api, access required)
- Genuine synthetic exports that verify300 / 300syntheticn = 300Median check 4.4 ms in Python.
- Outsider alterations caught (editing the export without keys)1,400 / 1,400syntheticn = 1,400
- Insider alterations caught, personal keys + TSA, export alone515 of 520syntheticn = 520All caught when checked against an earlier export (1,400 / 1,400 across configurations).
- Insider edits caught, account signatures + TSA, export alone29/200syntheticn = 200Deletes 2/40, reorders 14/40, backdates 40/80, forged signatures 40/80.
- Real recorded export, insider alterations caught on the export alone10 of 12syntheticn = 12All 12 outsider alterations caught.
- Python and TypeScript verifiers agree219 / 219syntheticn = 219
Dataset
Synthetic notebooks from the vertical's test kit (entries, attachments by SHA-256, amendments, AI analyses with receipts, signatures, RFC 3161 tokens from an offline test TSA), attacked by an outsider and an insider holding the server key across 13 alterations in three configurations; plus one real recorded demo export with a FreeTSA token.
Caveats
- Synthetic notebooks only; no personal or real lab data.
- No dev/test split: the verifier's rules were written first and not tuned to pass cases; only the attacker was changed after a run.
- Against the operator, account signatures alone catch few changes: an insider can delete timestamp tokens that stop matching.
- The quality of the AI analyses is not measured; the product records and receipts them, it does not grade them.
- Real-world clock drift against the TSA and TSA certificate revocation are not measured.
Nightly smoke check
Loading the nightly status…
- Result
- partial
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 24 s
- Receipts
- 1
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.001
Self-host verification
Verified on 26 Sep 2026: Fresh clone of the branch into a clean directory on our server, image built from docker/api/Dockerfile, api started with compose (named volume, python healthcheck), then the assemble prompt's smoke steps 1-9; torn down afterwards.
Ran against the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114 instead of the compose llm service; model-server startup not re-verified. Analysis receipt attested, FreeTSA token in 159 ms, export verified, altered export rejected. Host networking and port 8438 because other services held the default ports.
Rehearsal bundle: signed-lab-notebook.zip (6 KB, 7 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted check ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route): console flow with personal keys, own entry with a file, AI analysis, self-witness refused, timestamp rate limit, export, audit CSV, tamper buttons, verify page, Watch replay, Build example as written, 390 px layout. Production runs it once merged and deployed.
- The server signs exports with its own key: an operator could rebuild a notebook that has only account signatures and no surviving TSA token. Personal keys and export copies kept by others catch that.
- Signer identity is the name the key-holder sends unless the signer enrolled a personal key; no SSO or two-component sign-in (21 CFR 11.200) yet.
- AI analyses are receipted, not graded; the number check only points at numbers not in the data.
- FreeTSA is a free service with no SLA; TSA certificate revocation is not checked (the issuer is pinned).
- POST /notebook/verify takes 8 MB by default; larger exports are checked in the browser (a 17 MB, 10,000-entry export took 4.7 s).
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Notebook: hash chain, members, amendments, e-signature rules, RFC 3161 client, export, audit CSV and verification (no model; CPU)decosa-api notebook module (decosa_api/verticals/notebook)AGPL-3.0-or-later
- Writes the AI analysis note from the selected entries and attached data filesQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · notebook, signatures and timestamps on CPU (self-host) (4)
- Genuine exports verified: 300 / 300docs/evals/signed-lab-notebook.md
- Alterations caught, file edited without keys (13 kinds in 5 groups, 3 configurations): 1,400 / 1,400docs/evals/signed-lab-notebook.md
- Alterations caught, insider with the server key, personal keys + TSA: 515 / 520 on the export alone; 520 / 520 with an earlier exportdocs/evals/signed-lab-notebook.md
- Verify a 10,000-entry notebook (17,107 chain entries): 2.6 s Python; 4.7 s browser code (Node)docs/evals/signed-lab-notebook.md
Standard · Qwen3.8-27B writes receipted AI analyses (hosted demo) (4)
- AI analyses with a signed receipt covering the stored output: 5 / 5 end-to-end runs (smoke, recorder, self-host, two console analyses)scripts/smoke/signed-lab-notebook.py; docs/evals/signed-lab-notebook.md
- Insider edit of an AI output with a gateway receipt (real export): caught: the gateway receipt covers a different outputdocs/evals/signed-lab-notebook.md
- Sample analysis: outlier named, numbers checked: named well F7 in the run recorded; 7 of 13 numbers in the data, 4 rounded, 2 flagged (one run, not an accuracy measure)docs/evals/signed-lab-notebook.md
- Tamper detection and 10k performance: as the lite tier (same code)docs/evals/signed-lab-notebook.md