Pasteur¶
Clinical AI stress testing
Pasteur
Find where clinical models become brittle before deployment.
Test model behavior under realistic data failures. Pasteur simulates missing measurements, measurement noise, and transitions between patient cohorts, then quantifies how predictions respond. All computation stays on your machine.
Remove a feature and its measurement companions to test missingness.
Add controlled measurement noise and measure prediction stability.
Interpolate between differently labeled patients and locate decision flips.
100% local |
Runs on your machine’s CPU. No server, daemon, or admin rights. |
No network, no telemetry |
No network calls at runtime. Data, models, and results never leave the machine. |
Standard tooling |
|
Note
Reviewing Pasteur for a hospital or health system? Start with
Security and data handling, or download the
one-page overview (PDF).
What Pasteur produces¶
One simulate run creates a reproducible bundle containing the clean cohort
and each stress-test variant. evaluate scores one ONNX model;
compare places several models on the same rows and can write a parquet of
per-row predictions.
Command |
Input |
Result |
|---|---|---|
|
Local parquet data and optional cohort labels |
Clean, blackout, jitter, and flipper parquets |
|
Simulation bundle, labels, and one ONNX model |
Baseline and stability metrics (Reading the results) |
|
Simulation bundle and multiple ONNX models |
Model comparison and optional row-level predictions |
|
Simulation bundle and provenance |
A Hugging Face-compatible dataset card |
Quick look¶
pasteur-cli simulate \
--input patients.parquet \
--id-col patient_id \
--feature glucose \
--output ./output
This writes clean/, blackout/, and jitter/ beneath ./output.
Add --labels groups.parquet to generate flipper pairs. For a complete
run from CSV to a model comparison, see Walkthrough: choosing between models.
Designed for evidence, not a pass/fail badge¶
Pasteur does not claim that a model is safe. It produces concrete evidence about model behavior under declared perturbations: what changed, where it changed, and how strongly. Those results belong alongside intended-use documentation, local validation, clinical review, and deployment monitoring. Pasteur is a research and evaluation tool, not a medical device.