TCGA-4×4

5,760 tiles spanning four cancer types — breast invasive carcinoma, colon adenocarcinoma, and lung adeno- and squamous cell carcinoma — contributed by four TCGA tissue source sites (Asterand, Christiana Healthcare, Roswell Park, University of Pittsburgh) and scored at k = 71.

Read this cohort with its pretraining overlap in mind. TCGA is the most widely used pretraining corpus in computational pathology, and many of these encoders have seen it. A strong score can reflect an in-distribution advantage rather than robustness, and this page cannot tell the two apart — the paper quantifies the overlap encoder by encoder. Midnight-12k, which tops the cohort by a wide margin, is pretrained on TCGA and on nothing else.

TCGA-4×4, sorted by median CRoMa. Columns are explained under Reading the columns; † marks the natural-image control (The natural-image control); row tint marks pretraining overlap (legend below).

Model

bio bacc

conf bacc

RI

MaRI

CRoMa

F(0)

LTM₁₀

support

Midnight-12k

(TCGA-exposed pretraining)

0.882

0.559

0.898

0.936

0.40

0.140

-0.21

99.4%

Mascaret

(TCGA-exposed pretraining)

0.876

0.434

0.908

0.932

0.27

0.119

-0.11

99.9%

RudolfV-2-S

0.849

0.542

0.821

0.844

0.19

0.189

-0.16

100.0%

RudolfV-2-B

0.866

0.565

0.822

0.842

0.17

0.188

-0.14

100.0%

RudolfV-2

0.876

0.562

0.833

0.849

0.17

0.177

-0.12

100.0%

GenBio-PathFM

(TCGA-exposed pretraining)

0.851

0.695

0.757

0.780

0.16

0.242

-0.19

99.9%

CONCHv1.5

0.811

0.492

0.828

0.851

0.15

0.193

-0.13

100.0%

CONCH

0.790

0.487

0.801

0.825

0.15

0.216

-0.15

100.0%

Virchow2

0.825

0.594

0.771

0.795

0.13

0.239

-0.17

100.0%

H0-mini

(TCGA-exposed pretraining)

0.820

0.617

0.737

0.754

0.12

0.257

-0.19

99.9%

Virchow

0.789

0.656

0.697

0.706

0.09

0.301

-0.18

100.0%

H-optimus-1

0.878

0.676

0.775

0.786

0.09

0.220

-0.10

99.8%

Phaet

(TCGA-exposed pretraining)

0.786

0.569

0.760

0.767

0.08

0.254

-0.11

100.0%

UNI2-h

0.847

0.737

0.711

0.725

0.07

0.265

-0.12

99.3%

mSTAR

(TCGA-exposed pretraining)

0.817

0.690

0.696

0.703

0.06

0.298

-0.12

99.9%

H-optimus-0

0.846

0.692

0.707

0.711

0.05

0.287

-0.11

99.9%

MUSK

(TCGA-exposed pretraining)

0.724

0.589

0.674

0.680

0.05

0.332

-0.14

100.0%

Prov-GigaPath

0.831

0.672

0.683

0.679

0.05

0.308

-0.15

99.9%

UNI

0.807

0.712

0.673

0.678

0.05

0.320

-0.12

100.0%

Hibou-L

0.759

0.737

0.577

0.541

0.04

0.405

-0.25

99.7%

GPFM

(TCGA-exposed pretraining)

0.759

0.725

0.612

0.604

0.04

0.387

-0.15

100.0%

Phikon

(TCGA-exposed pretraining)

0.792

0.788

0.577

0.577

0.04

0.396

-0.20

99.2%

Hibou-B

0.768

0.722

0.600

0.590

0.04

0.387

-0.18

99.9%

Phikon-v2

(TCGA-exposed pretraining)

0.790

0.782

0.550

0.544

0.03

0.409

-0.19

99.4%

DINOv2-B †

0.607

0.580

0.540

0.513

0.01

0.468

-0.12

100.0%

Prost40M

(TCGA-exposed pretraining)

0.635

0.718

0.464

0.459

-0.01

0.528

-0.25

100.0%

The row tint is that overlap, made visible — the same convention as the paper’s dagger:

  • Orange — TCGA-exposed. TCGA appears in the encoder’s disclosed pretraining corpus or institutional provenance. Discount an advantage here.

  • Untinted — no disclosed overlap. No TCGA in the encoder’s disclosed corpus. For the proprietary corpora this reflects the paper’s description, not an independent audit.

Only 1 encoder falls below zero, and support is near-total — every model sits at 99% or above, so RI and MaRI rest on essentially every tile. In the explorer below, every encoder here resolves into a single tight mode close to zero, strong and weak alike; on Camelyon the same panel spreads across most of the scale. Two cohorts, one roster, very different separability — the argument for reporting more than one.

The two rankings, on this cohort alone:

Median CRoMa against tail severity LTM₁₀ on TCGA-4×4. Better is up and to the right; ringed points are undominated on both axes and named, and the shaded region is dominated on both. Hover or tab to any point to name it with its two values. The natural-image control is excluded, and pretraining exposure is not marked — the caveat above applies to every point.

The distribution explorer

The same explorer as the aggregate page’s, pinned to TCGA-4×4. Click a row to move the detail, drag across the detail curve to count the samples in any range, and pick a second encoder under Compare with to overlay its shape.

Shortcut susceptibility

Shortcut susceptibility for every encoder on this cohort, in domain (ID) and out of domain (OOD); the natural-image control sits last. Rows are ranked by Change at V = 1, the normalized change at maximum confounding, because nIPD averages over the whole range: an early gain there can pay for a late collapse, so a curve ending at chance can outrank one that never moved. Rows ending at or below -0.900 are marked ≈ chance, where none of the above-chance margin survives.

Bold marks the leading ranked encoder in each column where higher is better, so a column that disagrees with the ranking shows it at a glance. Each caption reports Spearman ρ, the rank correlation between the CRoMa and nIPD columns: how closely the two order the encoders the same way. Shortcut susceptibility defines the measure and holds the interactive explorer.

TCGA-4×4 — ID; Spearman ρ = 0.91; n=25 ranked pathology encoders

Model

Median CRoMa (m=5)

Change at V = 1

nIPD

Baseline balanced accuracy

RudolfV-2

0.168

-0.007

-0.004

0.848

CONCHv1.5

0.153

-0.008

0.003

0.777

RudolfV-2-B

0.169

-0.011

-0.004

0.838

RudolfV-2-S

0.192

-0.011

-0.003

0.827

Mascaret

0.270

-0.016

-0.004

0.869

GenBio-PathFM

0.155

-0.022

-0.003

0.838

Virchow2

0.128

-0.033

-0.007

0.817

Midnight-12k

0.396

-0.037

-0.015

0.873

Phaet

0.081

-0.046

-0.014

0.772

CONCH

0.146

-0.051

-0.017

0.753

mSTAR

0.059

-0.100

-0.027

0.801

H-optimus-1

0.088

-0.103

-0.029

0.856

H0-mini

0.121

-0.104

-0.028

0.809

UNI

0.047

-0.109

-0.041

0.797

UNI2-h

0.075

-0.120

-0.037

0.836

Virchow

0.092

-0.123

-0.024

0.776

H-optimus-0

0.054

-0.134

-0.040

0.812

MUSK

0.051

-0.137

-0.028

0.696

GPFM

0.037

-0.207

-0.053

0.750

Prov-GigaPath

0.051

-0.209

-0.062

0.795

Hibou-B

0.036

-0.220

-0.048

0.757

Hibou-L

0.039

-0.290

-0.062

0.760

Phikon-v2

0.029

-0.300

-0.069

0.774

Phikon

0.036

-0.309

-0.082

0.788

Prost40M

-0.010

-0.442

-0.143

0.632

DINOv2-B †

0.006

-0.203

-0.029

0.617

TCGA-4×4 — OOD; Spearman ρ = 0.89; n=25 ranked pathology encoders

Model

Median CRoMa (m=5)

Change at V = 1

nIPD

Baseline balanced accuracy

Midnight-12k

0.396

0.063

0.038

0.755

Mascaret

0.270

0.028

0.010

0.820

GenBio-PathFM

0.155

0.018

0.021

0.838

CONCHv1.5

0.153

0.003

0.027

0.741

RudolfV-2-B

0.169

-0.004

0.011

0.815

RudolfV-2-S

0.192

-0.004

0.009

0.807

RudolfV-2

0.168

-0.006

0.009

0.823

Virchow2

0.128

-0.047

-0.003

0.783

H-optimus-1

0.088

-0.051

-0.011

0.835

CONCH

0.146

-0.060

-0.016

0.711

Phaet

0.081

-0.076

-0.029

0.751

H-optimus-0

0.054

-0.110

-0.023

0.790

H0-mini

0.121

-0.112

-0.006

0.753

Virchow

0.092

-0.118

-0.008

0.738

UNI2-h

0.075

-0.126

-0.033

0.844

UNI

0.047

-0.141

-0.046

0.771

Prov-GigaPath

0.051

-0.154

-0.042

0.775

mSTAR

0.059

-0.157

-0.048

0.784

Hibou-B

0.036

-0.164

-0.014

0.705

Hibou-L

0.039

-0.179

-0.038

0.732

MUSK

0.051

-0.201

-0.053

0.641

GPFM

0.037

-0.266

-0.085

0.709

Phikon-v2

0.029

-0.282

-0.064

0.712

Phikon

0.036

-0.367

-0.100

0.744

Prost40M

-0.010

-0.678

-0.228

0.554

DINOv2-B †

0.006

-0.385

-0.104

0.560