Tolkach-ESCA

9,000 tiles of oesophageal tissue across six classes — tumour, regression, adventitia, muscularis propria, oesophageal and gastric mucosa — from three centers (UKK, WNS and CHA), scored at k = 61. The TCGA cohort of the original Tolkach dataset is held out, following PathoROB, so like Camelyon this is scored outside TCGA.

Tolkach-ESCA, sorted by median CRoMa. Columns are explained under Reading the columns; † marks the natural-image control (The natural-image control).

Model

bio bacc

conf bacc

RI

MaRI

Δ

CRoMa

F(0)

LTM₁₀

support

Midnight-12k

0.976

0.728

0.943

0.941

-0.002

0.58

0.051

-0.08

99.0%

Mascaret

0.977

0.445

0.972

0.973

+0.001

0.51

0.030

0.01

100.0%

RudolfV-2-S

0.984

0.599

0.967

0.975

+0.007

0.49

0.032

-0.00

99.8%

CONCH

0.973

0.654

0.951

0.957

+0.006

0.44

0.045

-0.04

99.8%

RudolfV-2

0.986

0.637

0.966

0.969

+0.003

0.41

0.032

-0.01

99.4%

RudolfV-2-B

0.984

0.664

0.960

0.968

+0.009

0.41

0.036

-0.02

99.0%

CONCHv1.5

0.973

0.633

0.952

0.964

+0.012

0.39

0.043

-0.03

99.9%

GenBio-PathFM

0.981

0.598

0.960

0.964

+0.004

0.39

0.038

-0.02

99.9%

H0-mini

0.967

0.642

0.935

0.946

+0.011

0.38

0.058

-0.07

99.7%

Virchow

0.970

0.703

0.935

0.943

+0.008

0.37

0.053

-0.05

99.6%

Virchow2

0.978

0.613

0.954

0.957

+0.002

0.35

0.040

-0.04

99.6%

MUSK

0.969

0.739

0.924

0.931

+0.008

0.29

0.063

-0.07

99.2%

H-optimus-1

0.977

0.680

0.940

0.949

+0.008

0.26

0.049

-0.04

99.1%

Phaet

0.967

0.619

0.935

0.946

+0.011

0.25

0.050

-0.03

100.0%

GPFM

0.969

0.841

0.883

0.902

+0.018

0.24

0.095

-0.10

97.6%

H-optimus-0

0.971

0.731

0.911

0.919

+0.007

0.23

0.076

-0.08

98.5%

UNI2-h

0.976

0.767

0.916

0.927

+0.012

0.22

0.064

-0.06

97.7%

mSTAR

0.970

0.818

0.881

0.898

+0.017

0.19

0.087

-0.09

97.7%

DINOv2-B †

0.905

0.535

0.876

0.867

-0.008

0.18

0.099

-0.07

100.0%

Phikon

0.962

0.898

0.772

0.782

+0.010

0.17

0.179

-0.19

81.5%

UNI

0.973

0.835

0.880

0.891

+0.011

0.17

0.086

-0.08

95.3%

Hibou-B

0.967

0.916

0.774

0.775

+0.000

0.13

0.188

-0.17

85.7%

Prov-GigaPath

0.962

0.929

0.707

0.746

+0.039

0.13

0.236

-0.16

79.1%

Prost40M

0.911

0.838

0.701

0.721

+0.019

0.13

0.277

-0.24

96.2%

Phikon-v2

0.956

0.896

0.741

0.746

+0.005

0.12

0.225

-0.17

83.3%

Hibou-L

0.955

0.960

0.624

0.586

-0.038

0.11

0.315

-0.29

67.0%

One provenance caveat: the RudolfV-2 family’s disclosed Charité/LMU institutional corpus creates a possible institutional/source-domain overlap with this cohort’s CHA center. Exact patient or slide overlap is unknown, so this does not establish leakage — but read the family’s scores here with it in mind.

The mildest of the three cohorts — 0 encoders fall below zero — and the one where the count-based indices run out of room: 16 of the 25 ranked encoders score above 0.90 on RI, so RI and MaRI have largely stopped separating models here while CRoMa still spreads the panel.

The two rankings, on this cohort alone:

Median CRoMa against tail severity LTM₁₀ on Tolkach-ESCA. Better is up and to the right; ringed points are undominated on both axes and named, and the shaded region is dominated on both. Hover or tab to any point to name it with its two values. The natural-image control is excluded — the frontier is a pathology-only claim.

Several encoders are also visibly bimodal in the explorer — one population of neighbourhoods comfortably biology-dominant, another close to the line. A median reports where the middle of that lands and says nothing about the split, which is the case tail reporting exists for.