A third-party AI field indexer

Rosicodex

A provenance instrument for AI-assisted science

When a probabilistic model reads instrument data and returns an interpretation, Rosicodex catalogues the moment: exactly what it saw, a confidence you can recompute, and a tamper-evident record a reviewer can verify without re-running the model.

/ Instrument/ Provenance/ 01–05

An MRI for AI

Making the invisible decisions of generative science visible, verifiable, reproducible

Before magnetic resonance imaging, a physician who needed to see inside soft tissue had two options: cut, or guess. The MRI made the interior legible without opening the body.

Generative AI has put science in the pre-MRI position. A model reads data and returns an interpretation, fluent and often excellent. What it saw, how confident it should be, and whether tomorrow's re-run would agree, stay invisible. Rosicodex does not open the model. It produces, at the moment of inference, a readable record of the decision.

Fig. 1 — the receipt pipeline
Content hash binds the exact prompt the model received
Source row IDs pin the exact records the model saw
Receipt — schema 1.3.0
receipt_id7f21…c93a
query█████ flagged samples
confidence0.87
sources3 files, 9 rows
modelnemotron-49b
provenancea94e…14b0
PASS
Deterministic confidence, recomputable, never the model's self-report
Provenance chain makes any later edit detectable
Plate I — a receipt, annotated
/ Problem/ Generative AI/ 02–05

The black box

What a fluent answer does not tell you

Today those questions get resolved by cutting or guessing: expensive re-run validation on a sliver of outputs, or trusting the model's own account of itself. Neither scales, and neither satisfies a reviewer, a partner, or a reproducibility standard.

A dashboard sits beside a workflow and reports on it after the fact. An instrument sits inside the workflow and produces the primary measurement. Rosicodex is built as the latter.

Illustrative rendering of an opaque compute system, styled as a diagnostic plate
Which candidate, and whyRanked by criteria no one downstream can see
Confidence, uncheckedA number with no basis a third party can recompute
No record on re-runNothing pins what changed if this runs again tomorrow
Plate II — inside the machine, unexamined
/ Mechanism/ The pipeline/ 03–05

The instrument

Five stages, one receipt

The receipt opens before the model is called and closes before the answer returns. Nothing is reconstructed afterward. A blocked or out-of-domain query never reaches the model at all, and the reason is recorded in its place.

Guards
Injection and answerability checked first
Retrieval
Exact rows pinned and hashed before the call
Confidence
Recomputable, never self-reported
Decision
Answer, caveat, or abstain to human review
Receipt
Sealed into a tamper-evident ledger
data coverage recency volume query specificity
/ Policy/ Executive Order 14303/ 04–05

The standard

Gold Standard Science, mechanized

Gold Standard Science asks that a result be reproducible, transparent, honest about error and uncertainty, skeptical of its own findings, and structured for falsifiability. Those have been properties a reviewer took on faith for an AI-assisted result. The receipt turns each into a field.

TenetReceipt mechanism
ReproducibleA content hash of the full prompt and the exact source records the model saw. A reviewer re-pulls precisely what it was given.
TransparentEvery source touched, the filter applied, and the rows returned, recorded in plain terms.
Communicative of errorConfidence is a deterministic function with every input exposed, recomputable independent of the model's own claim.
Skeptical of its findingsA grey-zone flag fires when an answer was permitted but sat close to the abstain line.
FalsifiableThe instrument refuses to answer when the data cannot ground a claim, rather than producing a confident, unfalsifiable paragraph.
/ Application/ Generative biology/ 05–05

Why now

Where the black box gets expensive

Consider the most demanding version of this problem: designing a therapeutic with generative AI. A platform can screen millions of AI-generated candidates a week, filtered by an internal scoring model before anything reaches a bench. Only a fraction are ever confirmed by direct structural evidence; the rest carry a confidence that is, today, the model's word alone.

The model retrains between cycles, so a candidate now in a screen was designed by a model that no longer exists in that form. And a January 2025 FDA draft framework for AI in drug and biological product submissions already asks sponsors for exactly this kind of record: model context, data provenance, versioned parameters, an audit trail a reviewer who wasn't in the room can still judge.

Rosicodex is built to generate that trail prospectively, at the moment of inference, instead of reconstructing it under deadline.

/ Field notes/ Status
Filing
NSF Project Pitch submitted, Scientific Instrumentation topic area — awaiting invitation decision
Build
Working, model-agnostic proof of concept, passing golden-case evaluation suite
Record
Replay-grade model record and hash-chained provenance ledger, prototyped
/ Correspondence/ Request access

Building toward an NSF pilot. Talking to early lab partners now.

If your team is putting AI-assisted results into a scientific record and needs a way to prove what the model actually saw, we want to hear about your workflow.