No model runs when a corpus is generated. Every claim is assembled from a template that is coherent by construction, and every row traces back to the clinical event that caused it.
Healthcare Corpus generates synthetic members, eligibility, professional and institutional medical claims, and pharmacy claims — priced against public CMS pricing files, adjudicated against each plan's benefit design, and reproducible to the byte from a seed.
The systems that most need testing are built to notice exactly what ordinary synthetic data gets wrong — or never contains.
They are small, and they pass the checks a pipeline was built to make because they contain nothing those checks exist to catch.
Valid code by code, wrong as a whole: a diagnosis that does not justify the procedure, a procedure billed where it cannot happen, a drug for a condition the member does not have.
Development, offshore and contractor work, vendor evaluations and sales demonstrations all need realistic claims — without access to production data.
Scoring a fraud rule or a risk model means knowing which members should be flagged. A corpus whose behaviours were authored knows; real data rarely says.
A persona describes an archetypal patient as rates and ratios. A cohort drawn from a mix of personas is enrolled in plans and played forward across a timeline.
AI assistance is confined to authoring the specification files, which a person reviews before they are committed. Generation is deterministic Python over the specifications, the reference tables and a seed.
Realistic, internally consistent claims data without protected health information — for the work where knowing exactly what is in the data is the point.
Exercise ingestion, canonical-model mapping and warehouse loads with all three claim families, header and line grains, and eligibility segments with gaps and plan changes — at any volume, reproducibly.
Behavioural signatures are authored rather than hoped for, and every row carries the persona that produced it — so a rule can be scored for sensitivity and for false positives, with persona_id as the label.
A clinical profile carries no money and no plan mechanics, so the same population can be enrolled in two plans: clinically identical, different only in benefit design.
Because the corpus contains no PHI, it can go where production data cannot: to a prospect, a competing vendor, a new hire or a contractor.
A corpus states what it is: which persona and event produced each row, which money is real, which codes are real, and which version of every table answered.
It reproduces the rates its authors wrote, not true prevalence, cost distributions or utilisation curves. Calibration against published benchmarks is future work.
Everything in it was authored, and the long tail is only as long as it is written. It is for testing, demos and training — not for learning what is happening in a real population.
Coordination of benefits is out of scope, and denial reason codes are not written: a denial drawn at a rate has no reason the corpus can honestly name.
Authorisations, lab results, health-risk and social-determinants data, X12 or FHIR serialisation, referral networks, capitation and value-based payment.
In a deployment engagement, your populations, plans and utilisation patterns are authored as specifications, reviewed, and generated inside your own environment under your own code-set licences.
Email us for more information about Healthcare Corpus, deployment engagements, or to start a conversation with our team.
info@artisoft.io