Reading the results¶
evaluate and compare report a small set of numbers per model. This page
defines each one, says which direction is better, and lists the caveats that
matter when you use them to choose between models.
All metrics are computed on the rows in your simulation bundle, with labels
from --labels and --positive-group-id.
Metrics¶
Field |
Definition |
Reading it |
|---|---|---|
|
ROC AUC of the model on the clean cohort. |
Ordinary discrimination. Higher is better; 0.5 is chance. |
|
ROC AUC on the blackout variant ÷ ROC AUC on the clean cohort. |
1.0 means missing the feature cost nothing. 0.9 means the model kept 90% of its AUC. Higher is better. |
|
|
1.0 means predictions did not move under noise. Lower means individual patients’ scores move when the measurement is noisy. Higher is better. |
|
For each sampled pair of opposite-label patients, the position |
Describes where the decision boundary sits between real patients. It is a description, not a score. Compare models against each other, not against a target value. |
|
Not implemented. Always |
Ignore it. |
Caveats¶
One simulation type per run. evaluate blackout loads only
blackout/. Metrics for simulation types that were not loaded are reported
at their neutral value: resiliency and jitter_stability are 1.0,
and flipper_stability is null. Run once per simulation type and read
the matching metric from each run:
Run |
Metric to read |
|---|---|
|
|
|
|
|
|
ROC AUC needs both classes. If the clean cohort has no positive rows, or no negative rows, ROC AUC is reported as 0.5 rather than as an error. Check your label counts first.
Tied scores count half. ROC AUC counts a tied positive/negative pair as half a win, whatever order the rows are in. Before this release ties were broken by row order, which mattered for models that emit many identical scores (tree ensembles, models fed a constant blackout fill).
Non-integer IDs are labelled negative. A row is positive only when its ID
parses as an integer and appears in member_ids. Identifiers such as
A00123 are silently treated as negative. Map them to integers before you
run Pasteur.
Missing values reach the model as NaN. Unless the model declares an
absent_sentinel or you pass --null-fill, blacked-out values are sent
to the model as NaN. Most exported scikit-learn models cannot handle NaN and
Pasteur stops with an error instead of reporting bad numbers. The simplest fix
is to include the imputer in the exported pipeline; see
Walkthrough: choosing between models.
Per-patient predictions¶
compare --predictions-out writes one row per patient per variant:
Column |
Meaning |
|---|---|
|
The patient’s original ID (synthetic path ID for flipper rows) |
|
|
|
1 if the patient is in the positive group, else 0 |
one column per model |
Positive-class probability. Named after the model file, without
|
|
True when every model puts the patient on the same side of 0.5
(fixed; |
For multiclass and multilabel models, label becomes one
label__<class> column per class, and each model gets one
<model>__<class> column per class. all_agree compares the argmax class
(multiclass) or the set of labels at or over each model’s thresholds
(multilabel). Compared models must list the same output.classes.
Use it to find the specific patients whose score moves, and to review them with clinicians.
Multiclass and multilabel models¶
For a model with K outputs, every metric is computed per class (multiclass)
or per label (multilabel). The headline fields above hold the macro
average over classes, so they read on the same scale as a binary model’s.
The extra multi section of the JSON holds what the averages hide. It is
absent for binary models.
Metric |
How it generalizes |
|---|---|
ROC AUC |
Each class against the rest (multiclass), or each label on its own
(multilabel). |
Resiliency |
Blackout AUC ÷ clean AUC for each label, then averaged. Labels whose
clean AUC is 0.5 or below are skipped: a ratio against chance-level
performance means nothing. |
Jitter stability |
Multilabel: |
Flipper stability |
Multiclass: pairs are drawn from two different classes a and b, and
the flip is where |
Read ``flipper_never_flipped`` next to ``flipper_stability``. Pairs that
never cross count as 1.0 in the mean, so a model that rarely changes its mind
along these paths scores high for that reason alone.
multi.flipper_never_flipped reports the share.
Regenerate the flipper grid when the task or groups change. Multiclass
grids store class indices in --positive-group-id order, and multilabel
grids store the label index in pair_label. Evaluating a grid built for a
different task is an error. Evaluating one built with the same task but
different groups is not detected.
Regression models¶
A single-target regression model (--task regression) predicts a value in
the target’s units instead of a probability. The JSON keeps the same shape,
with these fields:
Field |
Definition |
|---|---|
|
Root-mean-square error, mean absolute error, and R² on the clean cohort. RMSE and MAE are in the target’s units; lower is better. R² is 1 for a perfect model, 0 for one no better than predicting the mean, and negative for one worse than that. |
|
Blackout R² ÷ clean R²: the share of the model’s skill over predicting
the cohort mean that survives blackout. 1.0 means blackout cost
nothing; 0 means the model fell back to no better than the mean;
negative means blackout made it worse than the mean. It can exceed
1.0 when blackout happens to help, which usually means the
blacked-out feature was hurting the model. A model whose clean R² is
0 or below has no skill to keep and is reported at the neutral 1.0;
read |
|
|
|
For each pair of one patient whose true target is below the
|
Why an R² ratio and not an RMSE ratio. An error ratio rewards a model
for having little accuracy to lose. When a model loses a feature entirely and
falls back to the mean, clean RMSE ÷ blackout RMSE is sqrt(1 − R²): about
0.58 for a model with R² 0.66, but 0.86 for one with R² 0.26. On a synthetic
HbA1c cohort with glucose blacked out for every patient, a 1-nearest-neighbour
model dropped from R² 0.26 to −0.09 (worse than predicting the mean) and still
had the best RMSE ratio, 0.82, of the models that used glucose. Its R² ratio
is −0.35, the worst.
The regression section reports the same behaviour in the target’s units:
target_sd, blackout_rmse, blackout_r2, jitter_prediction_sd (the mean
per-patient standard deviation of the prediction across jitter draws),
flip_threshold, and flipper_never_flipped. It is absent for
classifiers.
The flipper cutoff is a clinical choice. There is no decision boundary in
a regression model, so flipper needs one: the threshold your service acts on,
such as HbA1c 6.5%. Pass the same --flip-threshold to simulate and to
evaluate. A grid evaluated against a cutoff that some pair does not
straddle is rejected. Without a cutoff, simulate skips flipper.
Read ``flipper_never_flipped`` here too. Pairs are chosen by their true targets, but crossings are read from predictions. A model whose predictions stay on one side of the cutoff for patients just above it never crosses, and those pairs count as 1.0. Regression models shrink towards the mean, so when the cutoff sits in the tail of the target this is common: on a synthetic cohort with 8% of HbA1c values at or above 6.5%, 57–68% of pairs never crossed for well-fitted models, and those pairs dominate the mean.
Jitter depends on the cohort. Because jitter is scaled by the target’s variance in the clean cohort, compare jitter stability between models on the same cohort only.
What these numbers do not show¶
They describe behaviour under the perturbations you declared: one feature removed at one rate, noise at one scale, straight-line paths between real patients. They do not measure calibration, fairness across subgroups, performance on another site’s population, or clinical benefit.