TCGA-4×4¶
5,760 tiles spanning four cancer types — breast invasive carcinoma, colon adenocarcinoma,
and lung adeno- and squamous cell carcinoma — contributed by four TCGA tissue source sites
(Asterand, Christiana Healthcare, Roswell Park, University of Pittsburgh) and scored at
k = 71.
Read this cohort with its pretraining overlap in mind. TCGA is the most widely used
pretraining corpus in computational pathology, and many of these encoders have seen it. A
strong score can reflect an in-distribution advantage rather than robustness, and this page
cannot tell the two apart — the paper quantifies the
overlap encoder by encoder. Midnight-12k, which tops the cohort by a wide margin, is
pretrained on TCGA and on nothing else.
Model |
bio bacc |
conf bacc |
|
|
Δ |
|
F(0) |
LTM₁₀ |
support |
|---|---|---|---|---|---|---|---|---|---|
Midnight-12k (TCGA-exposed pretraining) |
0.882 |
0.559 |
0.898 |
0.936 |
+0.037 |
0.40 |
0.140 |
-0.21 |
99.4% |
Mascaret (TCGA-exposed pretraining) |
0.876 |
0.434 |
0.908 |
0.932 |
+0.024 |
0.27 |
0.119 |
-0.11 |
99.9% |
RudolfV-2-S |
0.849 |
0.542 |
0.821 |
0.844 |
+0.023 |
0.19 |
0.189 |
-0.16 |
100.0% |
RudolfV-2-B |
0.866 |
0.565 |
0.822 |
0.842 |
+0.020 |
0.17 |
0.188 |
-0.14 |
100.0% |
RudolfV-2 |
0.876 |
0.562 |
0.833 |
0.849 |
+0.016 |
0.17 |
0.177 |
-0.12 |
100.0% |
GenBio-PathFM (TCGA-exposed pretraining) |
0.851 |
0.695 |
0.757 |
0.780 |
+0.023 |
0.16 |
0.242 |
-0.19 |
99.9% |
CONCHv1.5 |
0.811 |
0.492 |
0.828 |
0.851 |
+0.024 |
0.15 |
0.193 |
-0.13 |
100.0% |
CONCH |
0.790 |
0.487 |
0.801 |
0.825 |
+0.024 |
0.15 |
0.216 |
-0.15 |
100.0% |
Virchow2 |
0.825 |
0.594 |
0.771 |
0.795 |
+0.025 |
0.13 |
0.239 |
-0.17 |
100.0% |
H0-mini (TCGA-exposed pretraining) |
0.820 |
0.617 |
0.737 |
0.754 |
+0.017 |
0.12 |
0.257 |
-0.19 |
99.9% |
Virchow |
0.789 |
0.656 |
0.697 |
0.706 |
+0.009 |
0.09 |
0.301 |
-0.18 |
100.0% |
H-optimus-1 |
0.878 |
0.676 |
0.775 |
0.786 |
+0.012 |
0.09 |
0.220 |
-0.10 |
99.8% |
Phaet (TCGA-exposed pretraining) |
0.786 |
0.569 |
0.760 |
0.767 |
+0.007 |
0.08 |
0.254 |
-0.11 |
100.0% |
UNI2-h |
0.847 |
0.737 |
0.711 |
0.725 |
+0.013 |
0.07 |
0.265 |
-0.12 |
99.3% |
mSTAR (TCGA-exposed pretraining) |
0.817 |
0.690 |
0.696 |
0.703 |
+0.007 |
0.06 |
0.298 |
-0.12 |
99.9% |
H-optimus-0 |
0.846 |
0.692 |
0.707 |
0.711 |
+0.004 |
0.05 |
0.287 |
-0.11 |
99.9% |
MUSK (TCGA-exposed pretraining) |
0.724 |
0.589 |
0.674 |
0.680 |
+0.006 |
0.05 |
0.332 |
-0.14 |
100.0% |
Prov-GigaPath |
0.831 |
0.672 |
0.683 |
0.679 |
-0.005 |
0.05 |
0.308 |
-0.15 |
99.9% |
UNI |
0.807 |
0.712 |
0.673 |
0.678 |
+0.005 |
0.05 |
0.320 |
-0.12 |
100.0% |
Hibou-L |
0.759 |
0.737 |
0.577 |
0.541 |
-0.037 |
0.04 |
0.405 |
-0.25 |
99.7% |
GPFM (TCGA-exposed pretraining) |
0.759 |
0.725 |
0.612 |
0.604 |
-0.008 |
0.04 |
0.387 |
-0.15 |
100.0% |
Phikon (TCGA-exposed pretraining) |
0.792 |
0.788 |
0.577 |
0.577 |
-0.000 |
0.04 |
0.396 |
-0.20 |
99.2% |
Hibou-B |
0.768 |
0.722 |
0.600 |
0.590 |
-0.010 |
0.04 |
0.387 |
-0.18 |
99.9% |
Phikon-v2 (TCGA-exposed pretraining) |
0.790 |
0.782 |
0.550 |
0.544 |
-0.006 |
0.03 |
0.409 |
-0.19 |
99.4% |
DINOv2-B † |
0.607 |
0.580 |
0.540 |
0.513 |
-0.028 |
0.01 |
0.468 |
-0.12 |
100.0% |
Prost40M (TCGA-exposed pretraining) |
0.635 |
0.718 |
0.464 |
0.459 |
-0.005 |
-0.01 |
0.528 |
-0.25 |
100.0% |
The row tint is that overlap, made visible — the same convention as the paper’s dagger:
Orange — TCGA-exposed. TCGA appears in the encoder’s disclosed pretraining corpus or institutional provenance. Discount an advantage here.
Untinted — no disclosed overlap. No TCGA in the encoder’s disclosed corpus. For the proprietary corpora this reflects the paper’s description, not an independent audit.
Only 1 encoder falls below zero, and support is
near-total — every model sits at 99% or above, so RI
and MaRI rest on essentially every tile. In the explorer, every
encoder here resolves into a single tight mode close to zero, strong and weak alike; on
Camelyon the same panel spreads across most of the scale. Two cohorts,
one roster, very different separability — the argument for reporting more than one.
The two rankings, on this cohort alone:
Median CRoMa against tail severity LTM₁₀ on TCGA-4×4. Better is up and to the right;
ringed points are undominated on both axes and named, and the shaded region is dominated
on both. Hover or tab to any point to name it with its two values. The natural-image
control is excluded, and pretraining exposure is not marked — the caveat above applies to
every point.