Results¶
26 encoders — 25 pathology foundation models and one natural-image control — scored on three tile cohorts from the PathoROB study, plus a separate 5-encoder slide-level cohort, PCaBiop, published on its own page. Every number on this page is read at build time from the committed CSVs, never transcribed. The method is described in Beyond counts: A distributional robustness margin for pathology foundation models.
Two rankings, one frontier¶
Model |
mean rank |
|
tail rank |
Camelyon |
TCGA-4×4 |
Tolkach-ESCA |
|---|---|---|---|---|---|---|
Mascaret (TCGA-exposed pretraining) |
1.7 |
2.0 |
1.3 |
0.29/-0.02 |
0.27/-0.11 |
0.51/0.01 |
RudolfV-2-S |
4.3 |
2.3 |
6.3 |
0.32/-0.02 |
0.19/-0.16 |
0.49/-0.00 |
RudolfV-2 |
4.5 |
4.3 |
4.7 |
0.24/-0.04 |
0.17/-0.12 |
0.41/-0.01 |
RudolfV-2-B |
5.5 |
4.7 |
6.3 |
0.24/-0.05 |
0.17/-0.14 |
0.41/-0.02 |
CONCHv1.5 |
7.5 |
7.3 |
7.7 |
0.19/-0.14 |
0.15/-0.13 |
0.39/-0.03 |
GenBio-PathFM (TCGA-exposed pretraining) |
8.5 |
7.0 |
10.0 |
0.19/-0.07 |
0.16/-0.19 |
0.39/-0.02 |
CONCH |
9.2 |
6.0 |
12.3 |
0.20/-0.20 |
0.15/-0.15 |
0.44/-0.04 |
Virchow2 |
9.2 |
8.3 |
10.0 |
0.20/-0.11 |
0.13/-0.17 |
0.35/-0.04 |
H-optimus-1 |
9.3 |
12.7 |
6.0 |
0.08/-0.14 |
0.09/-0.10 |
0.26/-0.04 |
Phaet (TCGA-exposed pretraining) |
10.0 |
12.7 |
7.3 |
0.11/-0.18 |
0.08/-0.11 |
0.25/-0.03 |
H0-mini (TCGA-exposed pretraining) |
12.0 |
9.3 |
14.7 |
0.17/-0.16 |
0.12/-0.19 |
0.38/-0.07 |
Virchow |
12.2 |
10.3 |
14.0 |
0.16/-0.18 |
0.09/-0.18 |
0.37/-0.05 |
Midnight-12k (TCGA-exposed pretraining) |
12.2 |
4.7 |
19.7 |
0.11/-0.35 |
0.40/-0.21 |
0.58/-0.08 |
H-optimus-0 |
12.3 |
15.3 |
9.3 |
0.05/-0.15 |
0.05/-0.11 |
0.23/-0.08 |
UNI2-h |
13.3 |
15.3 |
11.3 |
0.04/-0.21 |
0.07/-0.12 |
0.22/-0.06 |
mSTAR (TCGA-exposed pretraining) |
14.3 |
16.7 |
12.0 |
0.02/-0.18 |
0.06/-0.12 |
0.19/-0.09 |
MUSK (TCGA-exposed pretraining) |
14.5 |
15.0 |
14.0 |
0.04/-0.22 |
0.05/-0.14 |
0.29/-0.07 |
UNI |
16.0 |
19.3 |
12.7 |
-0.03/-0.22 |
0.05/-0.12 |
0.17/-0.08 |
Prov-GigaPath |
17.3 |
19.3 |
15.3 |
0.01/-0.19 |
0.05/-0.15 |
0.13/-0.16 |
GPFM (TCGA-exposed pretraining) |
18.3 |
19.0 |
17.7 |
-0.10/-0.36 |
0.04/-0.15 |
0.24/-0.10 |
Hibou-B |
20.5 |
21.3 |
19.7 |
-0.09/-0.36 |
0.04/-0.18 |
0.13/-0.17 |
Phikon (TCGA-exposed pretraining) |
21.7 |
21.0 |
22.3 |
-0.20/-0.48 |
0.04/-0.20 |
0.17/-0.19 |
Phikon-v2 (TCGA-exposed pretraining) |
22.5 |
23.7 |
21.3 |
-0.21/-0.50 |
0.03/-0.19 |
0.12/-0.17 |
Prost40M (TCGA-exposed pretraining) |
24.0 |
24.0 |
24.0 |
-0.32/-0.64 |
-0.01/-0.25 |
0.13/-0.24 |
Hibou-L |
24.2 |
23.3 |
25.0 |
-0.44/-0.66 |
0.04/-0.25 |
0.11/-0.29 |
DINOv2-B † |
— |
— |
— |
0.05/-0.18 |
0.01/-0.12 |
0.18/-0.07 |
Bold marks the Pareto frontier: the encoders no other pathology encoder beats on both rankings at once. † marks the unranked natural-image control (see The natural-image control). Orange rows mark encoders whose disclosed pretraining overlaps TCGA — one of the three cohorts behind these ranks (legend).
The CRoMa and tail ranks are each encoder’s mean rank across the three cohorts — by
median CRoMa and by tail severity LTM₁₀ respectively — and the mean rank averages the
two. It is a reading order, not a score: a model can hold a strong median margin and still
be brittle on a subgroup, so the two ranks stay in view beside their average.
Midnight-12k and H-optimus-1 show the two ways that plays out — the former ranks
far better by median margin than by tail severity, the latter the opposite.
The claim the table makes is the frontier — a set, not an order. Anything on it is a defensible choice; which one you want depends on whether you care more about the typical sample or the worst tenth of them.
The same two rankings as axes. Better is up and to the right; the ringed, named points are undominated, and the shaded region below-left of the staircase is dominated on both axes. Hover or tab to any point to name it with its two mean ranks.
The distribution explorer¶
The per-sample CRoMa distribution is the object every number above is read from, so it
is shown whole. The list is every encoder on the cohort, in the tables’ order; the
highlighted row is drawn in full beneath it. Click a row to move the detail, drag across
the large histogram to count the samples in any range, and pick a second encoder under
Compare with to overlay its shape on the same axes — the readout then counts both.
The histograms are 200 bins of the per-sample CRoMa at the headline radius m = 5,
read from the same committed export as the tables.
The cohorts¶
Each cohort has its own page: the full column set sorted by median CRoMa, the cohort’s
own median-versus-tail Pareto panel, and what is specific to reading it. Columns are
explained under Reading the columns; † marks the natural-image control
(The natural-image control). Every encoder is evaluated at the cohort’s shared operating point —
the cohort median of the per-model biological k* — with tau resolved per model
(see Choosing tau).
Camelyon is scored entirely outside TCGA and is the most discriminating of the three; TCGA-4×4 must be read with pretraining overlap in mind; Tolkach-ESCA is the mildest, and the one where the count-based indices stop separating models.
The slide-level cohort¶
PCaBiop evaluates five whole-slide encoders — one embedding per slide — on 1,000 PANDA prostate biopsies. It is a different roster on a different evaluation unit, so it has its own page, its own distribution explorer, and no part in the aggregate ranks above (see Scope).
Reading the columns¶
Column |
Meaning |
|---|---|
bio bacc / conf bacc |
Balanced accuracy of a k-NN classifier predicting the biological label and the center. Diagnostics, not scores: a high confounder accuracy marks a representation that encodes the center strongly, so its maximum is never bolded. |
|
The pooled count-based and distance-weighted indices, in |
Δ |
|
|
The median signed margin at the headline radius, in |
F(0) |
The fraction of samples with |
LTM₁₀ |
The lower-tail mean: the mean of the lowest decile of the per-sample |
support |
The support fraction: how many samples contribute to |
The natural-image control¶
DINOv2-B, the natural-image control, is pretrained on natural images and has never seen
a whole-slide image. Its measurements are shown with the panel, but it is excluded from the
pathology ranks and frontier because it is a calibration floor, not a competitor. Its
positive margin is not evidence that a natural-image model beats pathology encoders: it has
the lowest biological retrieval accuracy in the panel on every cohort, and CRoMa
compares two neighbour distances, so a representation with weak structure of either kind
can score positively simply by having no strong confounder structure either. That is
precisely what makes it useful — it calibrates what a positive margin is worth on a poor
representation.
Scope¶
The tile roster is fixed across the three tile cohorts, and the ranks above are computed on it alone. The slide-level cohort PCaBiop is a different roster, so it never shares a table, a rank, or an explorer dropdown with the tile panel. And ranks are within a panel: they say which of these encoders is more robust on these cohorts, not how any of them would behave on yours.
The operating point¶
The three tile cohorts are reported under the median-k protocol: one shared k per cohort,
the cohort median of the per-model biological k*. A single operating point is what makes
a rank across encoders meaningful — comparing a model evaluated at k = 5 against one at
k = 91 compares two different questions.
tau is never pinned. Each encoder gets the median typed-neighbour distance of its own
embedding at that k, which is the only setting under which MaRI is comparable across
models (see Choosing tau). CRoMa is reported at its headline averaging radius,
m = 5, and LTM₁₀ at α = 0.10.
The slide-level cohort is the exception: with five encoders a shared median k is
dominated by panel composition, so PCaBiop reports k* — each encoder
at its own kNN-optimal k — and says so on its page.