Results

26 encoders — 25 pathology foundation models and one natural-image control — scored on three tile cohorts from the PathoROB study, plus a separate 5-encoder slide-level cohort, PCaBiop, published on its own page. Every number on this page is read at build time from the committed CSVs, never transcribed. The method is described in Beyond counts: A distributional robustness margin for pathology foundation models.

Two rankings, one frontier

25 ranked pathology encoders plus control

Model

mean rank

CRoMa rank

tail rank

Camelyon

TCGA-4×4

Tolkach-ESCA

Mascaret

(TCGA-exposed pretraining)

1.7

2.0

1.3

0.29/-0.02

0.27/-0.11

0.51/0.01

RudolfV-2-S

4.3

2.3

6.3

0.32/-0.02

0.19/-0.16

0.49/-0.00

RudolfV-2

4.5

4.3

4.7

0.24/-0.04

0.17/-0.12

0.41/-0.01

RudolfV-2-B

5.5

4.7

6.3

0.24/-0.05

0.17/-0.14

0.41/-0.02

CONCHv1.5

7.5

7.3

7.7

0.19/-0.14

0.15/-0.13

0.39/-0.03

GenBio-PathFM

(TCGA-exposed pretraining)

8.5

7.0

10.0

0.19/-0.07

0.16/-0.19

0.39/-0.02

CONCH

9.2

6.0

12.3

0.20/-0.20

0.15/-0.15

0.44/-0.04

Virchow2

9.2

8.3

10.0

0.20/-0.11

0.13/-0.17

0.35/-0.04

H-optimus-1

9.3

12.7

6.0

0.08/-0.14

0.09/-0.10

0.26/-0.04

Phaet

(TCGA-exposed pretraining)

10.0

12.7

7.3

0.11/-0.18

0.08/-0.11

0.25/-0.03

H0-mini

(TCGA-exposed pretraining)

12.0

9.3

14.7

0.17/-0.16

0.12/-0.19

0.38/-0.07

Virchow

12.2

10.3

14.0

0.16/-0.18

0.09/-0.18

0.37/-0.05

Midnight-12k

(TCGA-exposed pretraining)

12.2

4.7

19.7

0.11/-0.35

0.40/-0.21

0.58/-0.08

H-optimus-0

12.3

15.3

9.3

0.05/-0.15

0.05/-0.11

0.23/-0.08

UNI2-h

13.3

15.3

11.3

0.04/-0.21

0.07/-0.12

0.22/-0.06

mSTAR

(TCGA-exposed pretraining)

14.3

16.7

12.0

0.02/-0.18

0.06/-0.12

0.19/-0.09

MUSK

(TCGA-exposed pretraining)

14.5

15.0

14.0

0.04/-0.22

0.05/-0.14

0.29/-0.07

UNI

16.0

19.3

12.7

-0.03/-0.22

0.05/-0.12

0.17/-0.08

Prov-GigaPath

17.3

19.3

15.3

0.01/-0.19

0.05/-0.15

0.13/-0.16

GPFM

(TCGA-exposed pretraining)

18.3

19.0

17.7

-0.10/-0.36

0.04/-0.15

0.24/-0.10

Hibou-B

20.5

21.3

19.7

-0.09/-0.36

0.04/-0.18

0.13/-0.17

Phikon

(TCGA-exposed pretraining)

21.7

21.0

22.3

-0.20/-0.48

0.04/-0.20

0.17/-0.19

Phikon-v2

(TCGA-exposed pretraining)

22.5

23.7

21.3

-0.21/-0.50

0.03/-0.19

0.12/-0.17

Prost40M

(TCGA-exposed pretraining)

24.0

24.0

24.0

-0.32/-0.64

-0.01/-0.25

0.13/-0.24

Hibou-L

24.2

23.3

25.0

-0.44/-0.66

0.04/-0.25

0.11/-0.29

DINOv2-B †

0.05/-0.18

0.01/-0.12

0.18/-0.07

Bold marks the Pareto frontier: the encoders no other pathology encoder beats on both rankings at once. † marks the unranked natural-image control (see The natural-image control). Orange rows mark encoders whose disclosed pretraining overlaps TCGA — one of the three cohorts behind these ranks (legend).

The CRoMa and tail ranks are each encoder’s mean rank across the three cohorts — by median CRoMa and by tail severity LTM₁₀ respectively — and the mean rank averages the two. It is a reading order, not a score: a model can hold a strong median margin and still be brittle on a subgroup, so the two ranks stay in view beside their average. Midnight-12k and H-optimus-1 show the two ways that plays out — the former ranks far better by median margin than by tail severity, the latter the opposite.

The claim the table makes is the frontier — a set, not an order. Anything on it is a defensible choice; which one you want depends on whether you care more about the typical sample or the worst tenth of them.

The same two rankings as axes. Better is up and to the right; the ringed, named points are undominated, and the shaded region below-left of the staircase is dominated on both axes. Hover or tab to any point to name it with its two mean ranks.

The distribution explorer

The per-sample CRoMa distribution is the object every number above is read from, so it is shown whole. The list is every encoder on the cohort, in the tables’ order; the highlighted row is drawn in full beneath it. Click a row to move the detail, drag across the large histogram to count the samples in any range, and pick a second encoder under Compare with to overlay its shape on the same axes — the readout then counts both.

The histograms are 200 bins of the per-sample CRoMa at the headline radius m = 5, read from the same committed export as the tables.

The cohorts

Each cohort has its own page: the full column set sorted by median CRoMa, the cohort’s own median-versus-tail Pareto panel, and what is specific to reading it. Columns are explained under Reading the columns; † marks the natural-image control (The natural-image control). Every encoder is evaluated at the cohort’s shared operating point — the cohort median of the per-model biological k* — with tau resolved per model (see Choosing tau).

Camelyon is scored entirely outside TCGA and is the most discriminating of the three; TCGA-4×4 must be read with pretraining overlap in mind; Tolkach-ESCA is the mildest, and the one where the count-based indices stop separating models.

The slide-level cohort

PCaBiop evaluates five whole-slide encoders — one embedding per slide — on 1,000 PANDA prostate biopsies. It is a different roster on a different evaluation unit, so it has its own page, its own distribution explorer, and no part in the aggregate ranks above (see Scope).

Reading the columns

Column

Meaning

bio bacc / conf bacc

Balanced accuracy of a k-NN classifier predicting the biological label and the center. Diagnostics, not scores: a high confounder accuracy marks a representation that encodes the center strongly, so its maximum is never bolded.

RI / MaRI

The pooled count-based and distance-weighted indices, in [0, 1], neutral at 0.5. See Metrics.

Δ

MaRI RI. Informative in its sign — whether weighting by distance helps or hurts — rather than ordered by its size, so it is never bolded either.

CRoMa

The median signed margin at the headline radius, in (-1, 1), neutral at 0.

F(0)

The fraction of samples with CRoMa <= 0: confounder-dominant neighbourhoods, over the samples on which CRoMa is defined. Lower is better, so its bold is the minimum. See The confounder-dominant fraction.

LTM₁₀

The lower-tail mean: the mean of the lowest decile of the per-sample CRoMa distribution. How bad the worst tenth actually is.

support

The support fraction: how many samples contribute to RI/MaRI at all. A high index over a thin support is not a strong result — see undefined neighbourhoods.

The natural-image control

DINOv2-B, the natural-image control, is pretrained on natural images and has never seen a whole-slide image. Its measurements are shown with the panel, but it is excluded from the pathology ranks and frontier because it is a calibration floor, not a competitor. Its positive margin is not evidence that a natural-image model beats pathology encoders: it has the lowest biological retrieval accuracy in the panel on every cohort, and CRoMa compares two neighbour distances, so a representation with weak structure of either kind can score positively simply by having no strong confounder structure either. That is precisely what makes it useful — it calibrates what a positive margin is worth on a poor representation.

Scope

The tile roster is fixed across the three tile cohorts, and the ranks above are computed on it alone. The slide-level cohort PCaBiop is a different roster, so it never shares a table, a rank, or an explorer dropdown with the tile panel. And ranks are within a panel: they say which of these encoders is more robust on these cohorts, not how any of them would behave on yours.

The operating point

The three tile cohorts are reported under the median-k protocol: one shared k per cohort, the cohort median of the per-model biological k*. A single operating point is what makes a rank across encoders meaningful — comparing a model evaluated at k = 5 against one at k = 91 compares two different questions.

tau is never pinned. Each encoder gets the median typed-neighbour distance of its own embedding at that k, which is the only setting under which MaRI is comparable across models (see Choosing tau). CRoMa is reported at its headline averaging radius, m = 5, and LTM₁₀ at α = 0.10.

The slide-level cohort is the exception: with five encoders a shared median k is dominated by panel composition, so PCaBiop reports k* — each encoder at its own kNN-optimal k — and says so on its page.