Declare the cases. Get the patients. See the evidence.
Test your healthcare AI against the patients your sandbox doesn’t have. Tell us the healthcare situations your software must handle. Supermerco constructs the test patients, verifies what is actually present, and gives you reproducible evidence.
Coverage
– met– short– not verifiable
- SCN-MEDSAFE-001Medication prescribed despite documented allergy to it110 / 50
- SCN-MEDSAFE-002Renally cleared medication with impaired kidney function85 / 60
- SCN-MEDSAFE-003Potassium-raising drug with existing hyperkalaemiaNot verifiable
- SCN-MEDSAFE-007No allergy record of any kind exists150 / 60
- SCN-MEDSAFE-008Medication present with no supporting indication anywhereNot verifiable
- SCN-DATAINT-001Observation references an encounter that does not exist20 / 20
Abridged from a real run on 27 Sep 2026, on a MacBook Pro (Apple M3, 8 GB). The values are real; the log layout is ours, not the tool’s exact output. The tick on each bar marks the target.
Example run RUN-745029d75136: 16 of 21 scenarios met, 0 short and 5 not verifiable, 3,000 patients in 7.6 seconds. Rebuilding it from its manifest matched both inputs and results.
- FHIR R4 output
- 21 declared scenarios
- A Data Passport per run
- Reproducible from a manifest
- No runtime network calls
- One runtime dependency
A large test dataset is not automatically a useful one.
- Healthcare software and AI teams need specific patient situations for development, integration testing, QA, model evaluation and regression testing.
- If a required situation is absent from the test population, that situation is not being tested, and nobody can tell from the population alone what it covers.
- Rare, interacting, temporal, contradictory or deliberately malformed cases can be hard to obtain in the exact quantities needed.
- Learning-based synthetic data tools reproduce the patterns of their source, so rare situations stay rare.
- When software fails, reproducing the exact patient state and keeping it as a regression test can be hard.
Construct, then verify.
Tools that learn from existing data reproduce that data’s patterns, so rare situations stay rare. You cannot filter for something a generator never produces.
We construct patients to a specification instead of sampling them, then verify the result with separate code that does not know how the patients were made. We report the result as coverage of what you asked for, not as a single quality score.
Generation is the engine. Scenario control, verification and evidence are the product.
Seven steps from test plan to evidence
Each step is a command that exists today. Pick one to see what it produces.
Step 1 of 7
Declare
The situations to test are written as a specification in YAML: each scenario, how many patients need it, and the checkable conditions that define it.
You getA versioned test plan
# One scenario entry, simplified
- id: SCN-MEDSAFE-002
situation: Renally cleared medication with
impaired kidney function
severity: critical
target: 60
# ...plus the checkable conditions
# that define itWhat you receive
Five things, for every run.
FHIR R4 data
Six resource types in a population bundle, checked against the pinned R4 specification. Not a full HL7 validator, and we say so.
- Patient
- Encounter
- Condition
- Observation
- MedicationRequest
- AllergyIntolerance
Coverage report
Every declared scenario: met, short or not verifiable, with built versus by-chance counts.
Data Passport
Eleven sections. The “Not tested” section is never empty and cannot be switched off.
Test fixtures
One bundle per patient, as files your test suite can load.
Reproducible manifest
Every input, with its hash. Rebuild from it and you get the same population.
population sha-256c7e006ad71994ff1…inputs MATCH · results MATCH
Why you can trust it
Each rule is enforced by an automated test, not by good intentions.
- 01
The checker cannot see how the data was made
Verification and evaluation code are forbidden from importing construction code.
- 02
No combined quality score, anywhere
One number would hide the trade we make on purpose: we change the population so it covers your cases, which makes it look less like any source.
- 03
Every measurement declares what it cannot detect
It is a required field with no default.
- 04
A missing clinical fact is “unknown”, never “none”
The engine refuses to count a scenario that depends on an unsigned clinical fact.
- 05
Clinical rules live in versioned files
Each has a named reviewer and is never hidden in code, so a clinician can correct it.
- 06
Nothing leaves your machine
No telemetry and no network calls at runtime.
- 07
Every run is reproducible
From its manifest. The run ID is derived from the inputs, not from a clock.
- 08
Our wording is tested
An automated check fails the build if our output makes any of the legal or safety claims our policy forbids.
By the numbers
Measured 27 Sep 2026 on a MacBook Pro (Apple M3, 8 GB), seed 7 unless stated.
100,000 patients, construct and verify only
18.0 s
Library level, 494 MB, no FHIR export.
The FHIR export builds the whole bundle in memory, so end-to-end runs are measured up to 30,000 patients on this 8 GB machine. 100,000 patients is a construct-and-verify figure only.
~0.15 ms per patient to construct, measured at 1k, 3k, 10k, 30k and 100k.
- declared scenarios in 7 families
- 21
- In a 3,000-patient run, 16 are met in the stated quantities and 5 are openly marked not verifiable while they wait for clinical review.
- give a byte-identical population
- Same inputs
- Any run can be rebuilt from its manifest, and the run ID is derived from the inputs.
- FHIR R4 validation issues in valid mode
- 0
- At 3,000, 10,000 and 30,000 patients, under the project’s stated validator scope.
- network calls at runtime
- 0
- Enforced by an automated test. The one exception is the command that sends fixtures to an address you give it.
- runtime dependency
- 1
- PyYAML.
- automated tests
- 1,296
- CI on Linux, Windows and macOS.
Where this is going: the chart is not the patient
In development Programmable healthcare test worlds with verifiable ground truth. Each patient will carry its clinical truth, the record your system sees, and what the workflow did, with declared errors between them.
What exists today, and what we are building
Area by area. Anything not marked Today is in development.
| Area | Today | In development |
|---|---|---|
| Data construction | Scenario-driven synthetic patient construction. | Coverage-gap construction against your existing test fixtures. |
| Verification | Independent scenario verification and structured measurements. | A framework for judging system behaviour (oracles). |
| Evidence | Data Passport and manifest. | Test Passport with failures and regression lineage. |
| System execution | A fixture sender exists, tested only against a stand-in. | Real FHIR endpoint and model integrations. |
| Counterfactuals | Controlled mutations exist at library level. | Counterfactual twins: one declared divergence, an exact diff and verified lineage. |
| Clinical reality | Not modelled separately yet. | Separate clinical truth, recorded state and workflow state, with multiple clocks. |
| Regression | Failure minimisation exists at library level. | Stored regression suites, rerun on each release. |
| AI model | The core product does not require an AI model. | Possibly a self-hosted model for assisted language tasks. |
| Interface | Command line. | Other interfaces only once customer use validates them. |
| Clinical content | Sourced but unsigned, so partly not verifiable. | Clinician-reviewed rules with provenance. |
| Benchmark | No quotable comparative result. | A fair external benchmark, published whichever way it falls. |
What we do not claim
For a technical healthcare buyer, the limits matter as much as the features.
- The Data Passport does not establish that the data is anonymous, that its use is lawful, or that any system tested against it is clinically safe.
- Nothing has been signed by a clinician yet. Scenarios that depend on an unsigned clinical fact are counted as not verifiable, and 8 scenarios have no named clinical reviewer.
- We do not publish a benchmark comparison. No valid result exists yet.
- Sending fixtures to a FHIR server has been tested only against a stand-in, never against a real FHIR server.
- End-to-end runs are measured up to 30,000 patients. Above that, the FHIR export does not fit in 8 GB of memory.
- We construct declared test trajectories. We never predict what will happen to a real patient.
Ways to work with us today
Prices are not decided yet, so none are listed.
- 01Scenario coverage pilotWe turn your list of situations your software must handle into a scenario specification, build the population, and deliver FHIR R4 data, per-patient test fixtures, the Data Passport and the reproducible manifest.
- 02Custom scenario authoringNew scenarios written for your domain.
- 03Data-quality and robustness test setsPopulations with declared defects: missing values by stated mechanism, duplicates, conflicting records, broken references, adversarial mode.
- 04Guided test executionBetaWe send the fixtures to your FHIR endpoint and report which scenarios your system handled.
Questions
No. Every patient is constructed from a specification and stated distributions. No real patient record is used to build populations.
We do not describe it that way. It is built without real patient records, and the Data Passport reports what we measured about re-identification and what we did not test. Whether data counts as anonymous in your jurisdiction is for you and your counsel to decide.
FHIR R4 JSON: a population bundle, and one bundle per patient as a test fixture. The internal format is JSON too.
Yes. Every run writes a manifest. Rebuilding from the manifest produces the same population, and the run ID is derived from the inputs.
No. It makes no network calls while running. The one exception is when you ask it to send fixtures to a system you name.
Today, medication safety review (18 scenarios) and data integrity (3 scenarios). We add scenarios for your use case as part of a pilot.
Bring the list of situations your software must handle.
We will turn it into a scenario specification and show you what a pilot would deliver. Or email us at hello@supermerco.com.