Synthetic healthcare data — eligibility, medical and pharmacy claims

Claims data that tells
a coherent clinical story.

No model runs when a corpus is generated. Every claim is assembled from a template that is coherent by construction, and every row traces back to the clinical event that caused it.

Healthcare Corpus generates synthetic members, eligibility, professional and institutional medical claims, and pharmacy claims — priced against public CMS pricing files, adjudicated against each plan's benefit design, and reproducible to the byte from a seed.

Request a Demo See How It Works ↓
12
Tables in every corpus,
plus a run manifest
None
Model calls when
a corpus is generated
0
Unresolved codes tolerated —
one aborts the run
100%
Synthetic: no real person,
provider or payer
All data this product produces is synthetic
It describes no real person, provider or payer, every fact row carries data_source = SYNTHETIC, and it must never be used for clinical, financial or coverage decisions.
Real people in the data
None
Model calls at runtime
None
Same inputs, same seed
Same bytes
The status quo

Test data that tests nothing.

The systems that most need testing are built to notice exactly what ordinary synthetic data gets wrong — or never contains.

Hand-built fixtures are too clean

They are small, and they pass the checks a pipeline was built to make because they contain nothing those checks exist to catch.

Random codes are clinical nonsense

Valid code by code, wrong as a whole: a diagnosis that does not justify the procedure, a procedure billed where it cannot happen, a drug for a condition the member does not have.

Production data cannot travel

Development, offshore and contractor work, vendor evaluations and sales demonstrations all need realistic claims — without access to production data.

No ground truth to measure against

Scoring a fraud rule or a risk model means knowing which members should be flagged. A corpus whose behaviours were authored knows; real data rarely says.

How it works

From persona to priced claim.

A persona describes an archetypal patient as rates and ratios. A cohort drawn from a mix of personas is enrolled in plans and played forward across a timeline.

01
Author the specifications
A cohort spec, one benefit spec per plan, clinical profiles and personas — reviewed, version-controlled YAML, so a clinician or benefits analyst can check the content without reading code.
02
Load and gate
Every reference table is loaded and verified, and every authored code must resolve against its version-pinned table. A missing table or an unresolved code stops the run before a row exists.
03
Play the cohort forward
Members are grouped into households and enrolled. Conditions produce episodes, episodes produce encounters, and each encounter renders into claims through its template.
04
Price, adjudicate, write
Allowed amounts from public CMS pricing files, cost sharing against plan accumulators, denials and payment lags as each plan authors them. Integrity checks run first: a corpus that fails is never written.
Healthcare Corpus — clinical_event trace Illustrative
Member M-104233 — Persona A
Commercial PPO · Episode: acute cardiac event · data_source = SYNTHETIC
Persona A Inpatient admission anchored
Claims from this encounter
Facility claim · MS-DRG 309837I / UB-04
Professional claim · POS 21 · Dx I48.0837P / CMS-1500
Why does this row exist?
One encounter, two claims — the admission yields a paired professional and facility claim that agree by construction.
allowed_basis = anchored — priced from a real public CMS table, not invented.
Generated from specifications — no model call
Where the model runs

The model writes specifications, never rows.

AI assistance is confined to authoring the specification files, which a person reviews before they are committed. Generation is deterministic Python over the specifications, the reference tables and a seed.

Authoring — AI-assisted
A persona narrative becomes a clinical profile of rates and ratios, drawing on reusable catalogues of episodes, encounter templates, condition code sets and drug regimens.
Review — recorded on each document
Every specification names the reviews it has received — clinical, billing, benefits — with reviewer and date. Changing reviewed content empties the stamp, so a name never sits over content it did not cover.
Execution — no model, ever
The generator never calls a model, and the model never emits a data row. Anything clinical the generator seems to need to decide belongs in a specification instead.
4
Specification artifacts, one concern eachCohort, benefit design, clinical profile and persona. A cross-file validator enforces the seams, and nothing in the generator branches on a persona.
12
Tables, plus a run manifestMembers, eligibility, professional, institutional and pharmacy claims, providers and the coverage dimensions — and clinical_event, which answers for any claim why the row exists.
3
Tiers of read-back validationA written corpus is read back off disk and checked for structure, for codes against the pinned versions, and for clinical plausibility — by a harness that shares no code with the generator.
1 seed
Byte-identical outputThe same specifications, reference versions and seed give the same bytes. Member n of a persona is the same in a run of 100 members or 100,000, on one worker or forty.
Built for your team

Known ground truth,
for every team.

Realistic, internally consistent claims data without protected health information — for the work where knowing exactly what is in the data is the point.

Test the pipeline
against a golden file.

Exercise ingestion, canonical-model mapping and warehouse loads with all three claim families, header and line grains, and eligibility segments with gaps and plan changes — at any volume, reproducibly.

3
Claim families:
professional, institutional, pharmacy
12
Tables
per run
Regression testingIdentical inputs give identical bytes, so a pipeline change that alters downstream results shows up in a diff.
Scale and performance testingSize is set in member-years, not row counts. A larger cohort is a one-line change, and running inside JetStore spreads it across nodes.
Data-quality rule developmentGuaranteed integrity — claims within coverage, dates ordered, totals equal to their lines — gives a known-clean baseline for tuning your own rules.
Eligibility with real churnHouseholds enrolled as coverage segments, with gaps, plan changes and terminations as the cohort authors them.

Measure detection
against a known answer.

Behavioural signatures are authored rather than hoped for, and every row carries the persona that produced it — so a rule can be scored for sensitivity and for false positives, with persona_id as the label.

Fraud, waste and abuseThe opioid-seeking demonstration persona produces multiple prescribers, early refills and overlapping fills at authored rates. Every other persona produces none.
Pharmacy and PBM analyticsFormulary tiers, specialty drugs, adherence against authored proportion-of-days-covered ranges and controlled-substance schedules, priced against public drug pricing files.
Care management and risk stratificationMulti-chronic, geriatric and paediatric rare-disease members with trajectories and mortality, for stratification logic and care-gap workflows.
Behavioural health and substance usePsychiatric admission, substance use disorder and controlled-substance dispensing, modelled accurately and without euphemism.

Hold the population constant.
Change the plan.

A clinical profile carries no money and no plan mechanics, so the same population can be enrolled in two plans: clinically identical, different only in benefit design.

Adjudication and pricing engine testingAllowed amounts anchored to CMS files; cost sharing through deductibles, coinsurance and out-of-pocket accumulators, individual and family; denials at authored rates.
Benefit design what-ifCompare cost sharing, deductible burn-down and member liability with the clinical content held exactly constant.
Medicare Advantage workflowsMedicare-specific eligibility columns and status vocabularies, alongside commercial members.
Real money told from invented moneyEvery allowed amount is graded anchored, positioned or declared, so a consumer can tell which amounts a real public table priced.

Realistic on day one.
Shareable with anyone.

Because the corpus contains no PHI, it can go where production data cannot: to a prospect, a competing vendor, a new hire or a contractor.

Product demonstrations and salesA realistic population from the first day, shareable with a prospect.
Vendor and platform evaluationGive competing vendors the same reproducible corpus.
Training and onboardingAnalysts and engineers learn claims data on records whose every row can be traced to the reason it exists.
PHI-free development environmentsDevelopment, offshore and contractor work proceeds without access to production data.
Trust by construction

Every row traceable.
Every licence respected.

A corpus states what it is: which persona and event produced each row, which money is real, which codes are real, and which version of every table answered.

Synthetic only
data_source = SYNTHETIC on every fact row
Pinned references
Version, date, SHA-256 and row count
Full traceability
Every fact row carries run_id, persona_id, event_id and data_source. The run manifest records the seed, the generator version, and the version and checksum of every specification and reference table used.
Honest about fabricated money
Each allowed amount is graded anchored (a real public pricing table priced it), positioned (fabricated, placed against a real table's priced population) or declared (fabricated with nothing public behind it).
Licensing-aware by design
Public code sets ship with the product. Licensed ones — CPT, NUBC, NCPDP, X12 and others — are supplied by your installation, which holds the licence; the product only ever reads them and never copies one.
A fake code always looks fake
Where a site lacks a licensed table, it uses a stand-in in which every code begins ZZ and every description begins SYNTHETIC — so a fabricated code can never pass for a real one.
A missing table is fatal
If a table the run needs is absent, the run aborts at startup, naming the table and the path searched. A corpus that silently lacked its CPT lines would look complete and not be.
No network connection at runtime
The installed command never opens a network connection. Reference data is retrieved and refreshed by separate, supervised scripts, never by a run.
Standalone, or inside JetStore
Run it from a command line on one machine, or as an operator in a JetStore pipeline, partitioned by household across worker nodes — with output measured byte-identical to the standalone command.
What it is not

Stated plainly, so expectations are set.

Not calibrated to a real population

It reproduces the rates its authors wrote, not true prevalence, cost distributions or utilisation curves. Calibration against published benchmarks is future work.

It will not surprise you

Everything in it was authored, and the long tail is only as long as it is written. It is for testing, demos and training — not for learning what is happening in a real population.

Deliberate gaps, recorded

Coordination of benefits is out of scope, and denial reason codes are not written: a denial drawn at a rate has no reason the corpus can honestly name.

On the roadmap, not yet produced

Authorisations, lab results, health-risk and social-determinants data, X12 or FHIR serialisation, referral networks, capitation and value-based payment.

See a corpus built for
your population.

In a deployment engagement, your populations, plans and utilisation patterns are authored as specifications, reviewed, and generated inside your own environment under your own code-set licences.

Request a Demo Ask for the Technical Overview
Contact

Have a question? Let's talk.

Email us for more information about Healthcare Corpus, deployment engagements, or to start a conversation with our team.

info@artisoft.io