Camelyon¶
20,400 breast lymph-node tiles, labelled tumour or normal, contributed by two medical
centers (RUMC and UMCU) and scored at k = 11. Scored entirely
outside TCGA, so no encoder holds an in-distribution advantage from its pretraining corpus
— the cleanest of the three cohorts to read, and the most discriminating:
7 pathology encoders score below zero, meaning their
typical neighbourhood is closer to a different-biology tile from the same center than to
a same-biology tile from another.
Model |
bio bacc |
conf bacc |
|
|
Δ |
|
F(0) |
LTM₁₀ |
support |
|---|---|---|---|---|---|---|---|---|---|
RudolfV-2-S |
0.984 |
0.823 |
0.932 |
0.940 |
+0.008 |
0.32 |
0.047 |
-0.02 |
58.8% |
Mascaret |
0.981 |
0.849 |
0.909 |
0.914 |
+0.004 |
0.29 |
0.041 |
-0.02 |
72.7% |
RudolfV-2 |
0.989 |
0.915 |
0.895 |
0.903 |
+0.008 |
0.24 |
0.065 |
-0.04 |
38.3% |
RudolfV-2-B |
0.987 |
0.921 |
0.878 |
0.888 |
+0.010 |
0.24 |
0.071 |
-0.05 |
38.3% |
Virchow2 |
0.988 |
0.958 |
0.806 |
0.823 |
+0.017 |
0.20 |
0.129 |
-0.11 |
31.4% |
CONCH |
0.971 |
0.956 |
0.662 |
0.626 |
-0.037 |
0.20 |
0.225 |
-0.20 |
35.9% |
GenBio-PathFM |
0.983 |
0.928 |
0.842 |
0.850 |
+0.008 |
0.19 |
0.092 |
-0.07 |
38.3% |
CONCHv1.5 |
0.971 |
0.915 |
0.774 |
0.763 |
-0.012 |
0.19 |
0.174 |
-0.14 |
46.3% |
H0-mini |
0.969 |
0.927 |
0.741 |
0.718 |
-0.023 |
0.17 |
0.180 |
-0.16 |
38.7% |
Virchow |
0.980 |
0.946 |
0.751 |
0.708 |
-0.043 |
0.16 |
0.221 |
-0.18 |
26.4% |
Phaet |
0.967 |
0.943 |
0.708 |
0.686 |
-0.023 |
0.11 |
0.219 |
-0.18 |
48.4% |
Midnight-12k |
0.976 |
0.984 |
0.478 |
0.408 |
-0.070 |
0.11 |
0.354 |
-0.35 |
19.8% |
H-optimus-1 |
0.986 |
0.978 |
0.664 |
0.677 |
+0.013 |
0.08 |
0.219 |
-0.14 |
17.2% |
DINOv2-B † |
0.919 |
0.912 |
0.561 |
0.507 |
-0.053 |
0.05 |
0.345 |
-0.18 |
68.0% |
H-optimus-0 |
0.982 |
0.966 |
0.659 |
0.652 |
-0.007 |
0.05 |
0.315 |
-0.15 |
23.7% |
UNI2-h |
0.986 |
0.987 |
0.515 |
0.548 |
+0.033 |
0.04 |
0.370 |
-0.21 |
13.0% |
MUSK |
0.958 |
0.983 |
0.366 |
0.297 |
-0.069 |
0.04 |
0.403 |
-0.22 |
28.0% |
mSTAR |
0.979 |
0.984 |
0.460 |
0.434 |
-0.026 |
0.02 |
0.418 |
-0.18 |
18.0% |
Prov-GigaPath |
0.979 |
0.991 |
0.375 |
0.369 |
-0.007 |
0.01 |
0.470 |
-0.19 |
14.4% |
UNI |
0.982 |
0.999 |
0.108 |
0.092 |
-0.015 |
-0.03 |
0.651 |
-0.22 |
9.6% |
Hibou-B |
0.973 |
0.999 |
0.057 |
0.041 |
-0.017 |
-0.09 |
0.737 |
-0.36 |
13.4% |
GPFM |
0.955 |
0.999 |
0.034 |
0.017 |
-0.017 |
-0.10 |
0.753 |
-0.36 |
20.9% |
Phikon |
0.955 |
1.000 |
0.009 |
0.004 |
-0.005 |
-0.20 |
0.905 |
-0.48 |
16.6% |
Phikon-v2 |
0.954 |
1.000 |
0.019 |
0.008 |
-0.011 |
-0.21 |
0.932 |
-0.50 |
16.9% |
Prost40M |
0.926 |
1.000 |
0.015 |
0.002 |
-0.012 |
-0.32 |
0.922 |
-0.64 |
27.2% |
Hibou-L |
0.971 |
1.000 |
0.013 |
0.001 |
-0.011 |
-0.44 |
0.993 |
-0.66 |
12.1% |
Read the support column carefully here. Two biological classes across two centers is a
sparse neighbourhood: no encoder’s support fraction clears
73%, and the floor is
10%. A high RI over that little evidence is not the
same claim as one over TCGA-4×4’s near-total support — the same two
indices, resting on very different amounts of evidence.
The two rankings, on this cohort alone:
Median CRoMa against tail severity LTM₁₀ on Camelyon. Better is up and to the
right; ringed points are undominated on both axes and named, and the shaded region is
dominated on both. Hover or tab to any point to name it with its two values. The
natural-image control is excluded — the frontier is a pathology-only claim.
The shape says more than the median. Virchow2 and CONCH sit within
0.002 of each other on median CRoMa —
indistinguishable on that column alone — while CONCH carries
1.7× the confounder-dominant mass
and 1.9× the tail severity.
Overlay the two in the explorer to see it.