Declare the cases. Get the patients. See the evidence.

Test your healthcare AI against the patients your sandbox doesn’t have. Tell us the healthcare situations your software must handle. Supermerco constructs the test patients, verifies what is actually present, and gives you reproducible evidence.

Coverage

– met– short– not verifiable

  • SCN-MEDSAFE-001Medication prescribed despite documented allergy to it110 / 50
  • SCN-MEDSAFE-002Renally cleared medication with impaired kidney function85 / 60
  • SCN-MEDSAFE-003Potassium-raising drug with existing hyperkalaemiaNot verifiable
  • SCN-MEDSAFE-007No allergy record of any kind exists150 / 60
  • SCN-MEDSAFE-008Medication present with no supporting indication anywhereNot verifiable
  • SCN-DATAINT-001Observation references an encounter that does not exist20 / 20
All 21 scenarios

Abridged from a real run on 27 Sep 2026, on a MacBook Pro (Apple M3, 8 GB). The values are real; the log layout is ours, not the tool’s exact output. The tick on each bar marks the target.

Example run RUN-745029d75136: 16 of 21 scenarios met, 0 short and 5 not verifiable, 3,000 patients in 7.6 seconds. Rebuilding it from its manifest matched both inputs and results.

  • FHIR R4 output
  • 21 declared scenarios
  • A Data Passport per run
  • Reproducible from a manifest
  • No runtime network calls
  • One runtime dependency

A large test dataset is not automatically a useful one.

  • Healthcare software and AI teams need specific patient situations for development, integration testing, QA, model evaluation and regression testing.
  • If a required situation is absent from the test population, that situation is not being tested, and nobody can tell from the population alone what it covers.
  • Rare, interacting, temporal, contradictory or deliberately malformed cases can be hard to obtain in the exact quantities needed.
  • Learning-based synthetic data tools reproduce the patterns of their source, so rare situations stay rare.
  • When software fails, reproducing the exact patient state and keeping it as a regression test can be hard.

Construct, then verify.

Tools that learn from existing data reproduce that data’s patterns, so rare situations stay rare. You cannot filter for something a generator never produces.

We construct patients to a specification instead of sampling them, then verify the result with separate code that does not know how the patients were made. We report the result as coverage of what you asked for, not as a single quality score.

Generation is the engine. Scenario control, verification and evidence are the product.

Seven steps from test plan to evidence

Each step is a command that exists today. Pick one to see what it produces.

Step 1 of 7

Declare

The situations to test are written as a specification in YAML: each scenario, how many patients need it, and the checkable conditions that define it.

You getA versioned test plan

medication_safety.v1.yamlIllustration
# One scenario entry, simplified
- id: SCN-MEDSAFE-002
  situation: Renally cleared medication with
             impaired kidney function
  severity: critical
  target: 60
  # ...plus the checkable conditions
  #    that define it

Why you can trust it

Each rule is enforced by an automated test, not by good intentions.

  1. 01

    The checker cannot see how the data was made

    Verification and evaluation code are forbidden from importing construction code.

  2. 02

    No combined quality score, anywhere

    One number would hide the trade we make on purpose: we change the population so it covers your cases, which makes it look less like any source.

  3. 03

    Every measurement declares what it cannot detect

    It is a required field with no default.

  4. 04

    A missing clinical fact is “unknown”, never “none”

    The engine refuses to count a scenario that depends on an unsigned clinical fact.

  5. 05

    Clinical rules live in versioned files

    Each has a named reviewer and is never hidden in code, so a clinician can correct it.

  6. 06

    Nothing leaves your machine

    No telemetry and no network calls at runtime.

  7. 07

    Every run is reproducible

    From its manifest. The run ID is derived from the inputs, not from a clock.

  8. 08

    Our wording is tested

    An automated check fails the build if our output makes any of the legal or safety claims our policy forbids.

By the numbers

Measured 27 Sep 2026 on a MacBook Pro (Apple M3, 8 GB), seed 7 unless stated.

End to end, secondsConstruct, verify, FHIR bundle, validation and Passport
0s30s60s90s010k20k30k7.6s25.9s79.7s

100,000 patients, construct and verify only

18.0 s

Library level, 494 MB, no FHIR export.

The FHIR export builds the whole bundle in memory, so end-to-end runs are measured up to 30,000 patients on this 8 GB machine. 100,000 patients is a construct-and-verify figure only.

~0.15 ms per patient to construct, measured at 1k, 3k, 10k, 30k and 100k.

declared scenarios in 7 families
21
In a 3,000-patient run, 16 are met in the stated quantities and 5 are openly marked not verifiable while they wait for clinical review.
give a byte-identical population
Same inputs
Any run can be rebuilt from its manifest, and the run ID is derived from the inputs.
FHIR R4 validation issues in valid mode
0
At 3,000, 10,000 and 30,000 patients, under the project’s stated validator scope.
network calls at runtime
0
Enforced by an automated test. The one exception is the command that sends fixtures to an address you give it.
runtime dependency
1
PyYAML.
automated tests
1,296
CI on Linux, Windows and macOS.

Where this is going: the chart is not the patient

In development Programmable healthcare test worlds with verifiable ground truth. Each patient will carry its clinical truth, the record your system sees, and what the workflow did, with declared errors between them.

Same patient · same seed · identical history until 17:45

09:00blood collected10:00result produced10:03visible in record17:45review time18:10action timeLaterdeclared branch
Clinical truthPotassium 6.2Medicine stoppedStable
Recorded stateResult producedResult visibleChange recorded
WorkflowBlood collectedResult reviewedMedicine changedFollow-up arranged
In developmentIllustration of a counterfactual twin. Exactly one declared divergence; every later difference follows from declared rules, and the pair is verified with a machine-readable diff. A declared branch, not a prediction.

What exists today, and what we are building

Area by area. Anything not marked Today is in development.

AreaTodayIn development
Data constructionScenario-driven synthetic patient construction.Coverage-gap construction against your existing test fixtures.
VerificationIndependent scenario verification and structured measurements.A framework for judging system behaviour (oracles).
EvidenceData Passport and manifest.Test Passport with failures and regression lineage.
System executionA fixture sender exists, tested only against a stand-in.Real FHIR endpoint and model integrations.
CounterfactualsControlled mutations exist at library level.Counterfactual twins: one declared divergence, an exact diff and verified lineage.
Clinical realityNot modelled separately yet.Separate clinical truth, recorded state and workflow state, with multiple clocks.
RegressionFailure minimisation exists at library level.Stored regression suites, rerun on each release.
AI modelThe core product does not require an AI model.Possibly a self-hosted model for assisted language tasks.
InterfaceCommand line.Other interfaces only once customer use validates them.
Clinical contentSourced but unsigned, so partly not verifiable.Clinician-reviewed rules with provenance.
BenchmarkNo quotable comparative result.A fair external benchmark, published whichever way it falls.

What we do not claim

For a technical healthcare buyer, the limits matter as much as the features.

  • The Data Passport does not establish that the data is anonymous, that its use is lawful, or that any system tested against it is clinically safe.
  • Nothing has been signed by a clinician yet. Scenarios that depend on an unsigned clinical fact are counted as not verifiable, and 8 scenarios have no named clinical reviewer.
  • We do not publish a benchmark comparison. No valid result exists yet.
  • Sending fixtures to a FHIR server has been tested only against a stand-in, never against a real FHIR server.
  • End-to-end runs are measured up to 30,000 patients. Above that, the FHIR export does not fit in 8 GB of memory.
  • We construct declared test trajectories. We never predict what will happen to a real patient.

No. Every patient is constructed from a specification and stated distributions. No real patient record is used to build populations.

We do not describe it that way. It is built without real patient records, and the Data Passport reports what we measured about re-identification and what we did not test. Whether data counts as anonymous in your jurisdiction is for you and your counsel to decide.

FHIR R4 JSON: a population bundle, and one bundle per patient as a test fixture. The internal format is JSON too.

Yes. Every run writes a manifest. Rebuilding from the manifest produces the same population, and the run ID is derived from the inputs.

No. It makes no network calls while running. The one exception is when you ask it to send fixtures to a system you name.

Today, medication safety review (18 scenarios) and data integrity (3 scenarios). We add scenarios for your use case as part of a pilot.

Bring the list of situations your software must handle.

We will turn it into a scenario specification and show you what a pilot would deliver. Or email us at hello@supermerco.com.

Talk to us about a pilot