Working is the easy half
A clinical data pipeline that transforms EDC exports into submission-ready datasets is a solved engineering problem. Parse, map, validate, emit. The hard half arrives eighteen months later, when someone asks why subject 4021's visit date changed between two extracts, and the honest answer is that nobody knows.
Regulatory inspection does not test whether software is correct. It tests whether you can demonstrate what the software did to a specific record at a specific time, and show that the demonstration itself was not edited afterwards. Those are different properties, and the second one has to be designed in. It cannot be added under deadline.
Three properties that must be structural
1. Every value carries its provenance. A transformed value that does not know where it came from is an assertion, not evidence. The unit of storage is not the value: it is the value plus the source record, the rule that produced it, and the version of that rule.
interface DerivedValue<T> {
value: T;
sourceOID: string; // the ODM item this came from
ruleId: string; // which mapping produced it
ruleVersion: string; // pinned, never "latest"
derivedAt: string; // ISO 8601, UTC
inputHash: string; // hash of the inputs, not the output
}
Hashing the inputs rather than the output is the part people get backwards. An output hash proves the value has not been tampered with. An input hash proves the value is reproducible: re-run the pinned rule against the same inputs and you must get the same answer. Only the second one survives being asked "show me."
2. Controlled terminology is pinned, not fetched. NCI EVS publishes updates quarterly. A pipeline that resolves codelists at runtime produces different output on different days from identical input, and the difference is invisible until an inspector diffs two extracts. Pin the dictionary version per study, store it alongside the data, and treat a terminology upgrade as a deliberate migration with its own record, never as a background refresh.
3. The audit trail is append-only or it is decorative. If the process that writes audit records can also update or delete them, the trail proves nothing beyond the good intentions of whoever held the credentials. This is the single most common gap, and it is usually justified by a cleanup job that "only removes duplicates."
The failure that is hardest to catch
A pipeline that fails loudly is a good pipeline. The dangerous one succeeds while doing nothing.
Fallback behaviour is the usual culprit. A service that substitutes static content when a query fails is correct at runtime: a visitor should still get a page. Run that same code at build time and it bakes the fallback into an artifact that ships as though it were real, and every downstream check passes, because from the outside a complete dataset and a complete substitute dataset are indistinguishable.
The rule worth internalising: gate on build versus runtime, not on production versus non-production. A production build satisfies every "is this production?" check while having no business touching a database at all. Make the build fail on an unreachable data source, and make an empty-but-successful query a non-failure, because those two states are genuinely different and conflating them produces either false alarms or false confidence.
What to test
Conformance suites check that output matches a schema. They do not check that the output is this input's output. Add tests that:
- Re-run a pinned rule against a stored input hash and assert byte equality with the recorded result.
- Attempt to update an audit row and assert the write is rejected: not that it is "not done", that it cannot be done.
- Assert the build exits non-zero when the data source is unreachable. A green build is only evidence if a red one was possible.
That last one is the test people skip, because it feels like testing the framework. It is not. It is testing the one assumption every other test rests on.
The CRF-XL and clinical data mapper case studies walk through the AST and streaming-transform mechanics; Cadence Clinical covers the hexagonal boundary that keeps the audit surface small enough to reason about. You can drive the rule evaluator directly at /crf.