PCaBiop (slide-level)¶
1,000 prostate biopsy slides sourced from PANDA, labelled benign or cancer, contributed by two data providers (Karolinska Institutet and Radboud UMC). Unlike the three tile cohorts, the evaluation unit is a whole slide: each of the 5 encoders is a whole-slide foundation model producing one embedding per slide, and every number below is computed over those slide embeddings.
This is a different panel. The slide roster shares no encoder ranking with the 26-model tile panel, so this cohort takes no part in the cross-cohort aggregate and its ranks — the table below orders these five encoders against each other, nothing more. There is also no natural-image control at slide level, so no row carries the † mark.
Don't see a slide encoder? Request an evaluation
Model |
bio bacc |
conf bacc |
|
|
|
F(0) |
LTM₁₀ |
support |
|---|---|---|---|---|---|---|---|---|
PRISM2 |
0.975 |
0.995 |
0.434 |
0.367 |
0.27 |
0.187 |
-0.24 |
20.9% |
PRISM |
0.968 |
0.992 |
0.281 |
0.195 |
0.26 |
0.288 |
-0.39 |
10.1% |
MOOZY |
0.971 |
1.000 |
0.236 |
0.181 |
-0.02 |
0.535 |
-0.41 |
24.3% |
TITAN |
0.915 |
1.000 |
0.015 |
0.001 |
-0.30 |
0.895 |
-0.60 |
49.1% |
Prov-GigaPath |
0.720 |
0.995 |
0.020 |
0.006 |
-0.41 |
0.990 |
-0.59 |
97.7% |
The confounder is unusually dominant here: every encoder’s confounder k-NN accuracy is
at or near 1.0 — the data provider is essentially perfectly decodable from every slide
embedding — and 3 of the five encoders hold a negative
median margin. Support spans 10% to
98%, the widest spread of any published cohort, so the
count-based indices rest on very different amounts of evidence per row: Prov-GigaPath’s
low RI is measured on nearly every slide, while PRISM’s rests on a tenth of them.
The two rankings, on this cohort alone:
Median CRoMa against tail severity LTM₁₀. Better is up and to the right; ringed
points are undominated on both axes. Hover or tab to any point to name it with its two
values. With five encoders the panel is sparse by construction — read it as a picture of
the table, not as a frontier claim of tile-panel strength.
The distribution explorer¶
The same explorer as the tile panel’s, restricted to the slide roster. Click a row to move the detail, drag across the detail curve to count the slides in any range, and pick a second encoder under Compare with to overlay its shape.
The shape is again where the medians stop telling the story. PRISM2 and PRISM sit
0.014 apart on median CRoMa —
indistinguishable on that column — while PRISM carries
1.5× the confounder-dominant mass and
1.7× the tail severity. Overlay
the two above to see it.
The operating point¶
This cohort is reported at k*, not the tile panel’s shared median-k: each encoder
is evaluated at its own kNN-optimal k (here spanning k =
3–13; the per-encoder values are recorded in
results/PROVENANCE.json). With only five encoders, a shared median is dominated by
panel composition — adding a single encoder to the panel moved the would-be shared k
from 3 to 9 — so pinning one k would make every number hostage to who else happens to
be on the roster. The trade is stated rather than hidden: RI and MaRI are
protocol-dependent, so those two columns are not measured at one shared operating point
here, while CRoMa, F(0) and LTM₁₀ are k-free and unaffected.
Shortcut susceptibility¶
Shortcut susceptibility for the five whole-slide encoders, in domain (ID) and out of
domain (OOD). Rows are ranked by Change at V = 1, the normalized change at maximum
confounding, because nIPD averages over the whole range: an early gain there can pay
for a late collapse, so a curve ending at chance can outrank one that never moved. Rows
ending at or below -0.900 are marked ≈ chance, where none of the above-chance margin
survives.
Bold marks the leading ranked encoder in each column where higher is better, so a
column that disagrees with the ranking shows it at a glance. Each caption reports
Spearman ρ, the rank correlation between the CRoMa and nIPD columns: how closely
the two order the encoders the same way.
Shortcut susceptibility defines the measure and holds the interactive explorer.
Model |
Median CRoMa (m=5) |
Change at |
|
Baseline balanced accuracy |
|---|---|---|---|---|
PRISM2 |
0.270 |
-0.216 |
-0.032 |
0.957 |
PRISM |
0.257 |
-0.293 |
-0.078 |
0.953 |
MOOZY |
-0.016 |
-0.313 |
-0.065 |
0.967 |
TITAN |
-0.295 |
-0.803 |
-0.189 |
0.912 |
Prov-GigaPath |
-0.413 |
-0.970 ≈ chance |
-0.245 |
0.803 |
Model |
Median CRoMa (m=5) |
Change at |
|
Baseline balanced accuracy |
|---|---|---|---|---|
PRISM2 |
0.270 |
-0.040 |
-0.011 |
0.945 |
PRISM |
0.257 |
-0.304 |
-0.083 |
0.917 |
MOOZY |
-0.016 |
-0.759 |
0.048 |
0.790 |
TITAN |
-0.295 |
-0.931 ≈ chance |
-0.126 |
0.855 |
Prov-GigaPath |
-0.413 |
-1.161 ≈ chance |
-0.107 |
0.548 |
Read MOOZY’s positive OOD value with care. Its pretraining corpus includes the PANDA
slides every probe trains on, so the apparent improvement under confounding plausibly
reflects familiarity with the training data rather than robustness on the held-out PAR
cohort, which no encoder has seen.