Camelyon

20,400 breast lymph-node tiles, labelled tumour or normal, contributed by two medical centers (RUMC and UMCU) and scored at k = 11. Scored entirely outside TCGA, so no encoder holds an in-distribution advantage from its pretraining corpus — the cleanest of the three cohorts to read, and the most discriminating: 7 pathology encoders score below zero, meaning their typical neighbourhood is closer to a different-biology tile from the same center than to a same-biology tile from another.

Camelyon, sorted by median CRoMa. Columns are explained under Reading the columns; † marks the natural-image control (The natural-image control).

Model

bio bacc

conf bacc

RI

MaRI

Δ

CRoMa

F(0)

LTM₁₀

support

RudolfV-2-S

0.984

0.823

0.932

0.940

+0.008

0.32

0.047

-0.02

58.8%

Mascaret

0.981

0.849

0.909

0.914

+0.004

0.29

0.041

-0.02

72.7%

RudolfV-2

0.989

0.915

0.895

0.903

+0.008

0.24

0.065

-0.04

38.3%

RudolfV-2-B

0.987

0.921

0.878

0.888

+0.010

0.24

0.071

-0.05

38.3%

Virchow2

0.988

0.958

0.806

0.823

+0.017

0.20

0.129

-0.11

31.4%

CONCH

0.971

0.956

0.662

0.626

-0.037

0.20

0.225

-0.20

35.9%

GenBio-PathFM

0.983

0.928

0.842

0.850

+0.008

0.19

0.092

-0.07

38.3%

CONCHv1.5

0.971

0.915

0.774

0.763

-0.012

0.19

0.174

-0.14

46.3%

H0-mini

0.969

0.927

0.741

0.718

-0.023

0.17

0.180

-0.16

38.7%

Virchow

0.980

0.946

0.751

0.708

-0.043

0.16

0.221

-0.18

26.4%

Phaet

0.967

0.943

0.708

0.686

-0.023

0.11

0.219

-0.18

48.4%

Midnight-12k

0.976

0.984

0.478

0.408

-0.070

0.11

0.354

-0.35

19.8%

H-optimus-1

0.986

0.978

0.664

0.677

+0.013

0.08

0.219

-0.14

17.2%

DINOv2-B †

0.919

0.912

0.561

0.507

-0.053

0.05

0.345

-0.18

68.0%

H-optimus-0

0.982

0.966

0.659

0.652

-0.007

0.05

0.315

-0.15

23.7%

UNI2-h

0.986

0.987

0.515

0.548

+0.033

0.04

0.370

-0.21

13.0%

MUSK

0.958

0.983

0.366

0.297

-0.069

0.04

0.403

-0.22

28.0%

mSTAR

0.979

0.984

0.460

0.434

-0.026

0.02

0.418

-0.18

18.0%

Prov-GigaPath

0.979

0.991

0.375

0.369

-0.007

0.01

0.470

-0.19

14.4%

UNI

0.982

0.999

0.108

0.092

-0.015

-0.03

0.651

-0.22

9.6%

Hibou-B

0.973

0.999

0.057

0.041

-0.017

-0.09

0.737

-0.36

13.4%

GPFM

0.955

0.999

0.034

0.017

-0.017

-0.10

0.753

-0.36

20.9%

Phikon

0.955

1.000

0.009

0.004

-0.005

-0.20

0.905

-0.48

16.6%

Phikon-v2

0.954

1.000

0.019

0.008

-0.011

-0.21

0.932

-0.50

16.9%

Prost40M

0.926

1.000

0.015

0.002

-0.012

-0.32

0.922

-0.64

27.2%

Hibou-L

0.971

1.000

0.013

0.001

-0.011

-0.44

0.993

-0.66

12.1%

Read the support column carefully here. Two biological classes across two centers is a sparse neighbourhood: no encoder’s support fraction clears 73%, and the floor is 10%. A high RI over that little evidence is not the same claim as one over TCGA-4×4’s near-total support — the same two indices, resting on very different amounts of evidence.

The two rankings, on this cohort alone:

Median CRoMa against tail severity LTM₁₀ on Camelyon. Better is up and to the right; ringed points are undominated on both axes and named, and the shaded region is dominated on both. Hover or tab to any point to name it with its two values. The natural-image control is excluded — the frontier is a pathology-only claim.

The shape says more than the median. Virchow2 and CONCH sit within 0.002 of each other on median CRoMa — indistinguishable on that column alone — while CONCH carries 1.7× the confounder-dominant mass and 1.9× the tail severity. Overlay the two in the explorer to see it.