Shortcut susceptibility

nIPD measures how a supervised probe changes when its training data becomes confounder-biased. Negative nIPD means net degradation as training-set confounding increases, near zero means little or no net change, and positive nIPD means net improvement. A near-zero value can hide offsetting changes, so it is not proof of stable performance at every point.

Interactive evidence browser

Select a cohort, then an encoder, to open the sampled trajectory behind its nIPD value. Pick a second encoder under Compare with to overlay its trajectory on the same scale. DINOv2-B is shown separately as a natural-image control where available; it is not ranked with the pathology encoders.

Rows are ranked by the normalized change at V = 1, the endpoint of the trajectory, not by nIPD. nIPD averages over the whole confounding range, so an early gain can pay for a late collapse and a curve ending at chance can outrank one that never moved; an endpoint cannot cancel with itself. Rows ending at or below -90% are marked ≈ chance, and panels whose scale reaches the -100% floor draw it as the chance level.

Loading the committed nIPD evidence…

The CRoMa and downstream susceptibility scatter below the trajectory pairs each encoder’s median CRoMa at m=5 with its nIPD. Hover or tab to any point to name it with its two values. The fitted trend and the Spearman coefficient use ranked pathology encoders only, so DINOv2-B is excluded from both. PCaBiop holds n=5 encoders, so its coefficient is descriptive and no trend is fitted.

How nIPD is computed

A logistic probe predicts the biological class from frozen embeddings. Its training composition moves from balanced to fully confounded while the test rows stay fixed, and at each Cramér’s-V value nIPD compares mean balanced accuracy with the balanced baseline. The change is divided by the baseline’s margin over chance, so nIPD measures the share of above-chance performance that is lost. ID is the primary mechanistic endpoint: its acquisition groups remain represented in training, isolating susceptibility to the shortcut. OOD also includes transfer effects because its acquisition groups were unseen during training. See the API definition for the equation and input contract.

Results

Per-cohort tables, ranked by the normalized change at V = 1, live on the cohort pages: Camelyon, TCGA-4×4, Tolkach-ESCA and PCaBiop. Download the exact float basis as JSON or the complete tabular view as CSV.

APD remains available as the PathoROB-faithful continuity reduction. It normalizes by raw baseline accuracy; nIPD is the primary result because it normalizes by the margin over chance and integrates over the observed Cramér’s-V coordinates.