| Axis | Baseline (NV / AMD) | Optimized (NV / AMD) | What a department reads |
|---|---|---|---|
| D1 Slide latency | 129.6 / 128.6 s | 79.1 / 82.8 s | Fits batch sign-out, not intra-operative |
| D2 Throughput | 222 / 224 per shift | 364 / 348 per shift | 2,000-slide day: 10 cards to 6. 100k archive: 150 to 92 GPU-days |
| D3 Energy | 20.2 / 20.4 Wh | 10.0 / 10.6 Wh | Cost per slide falls 1.64–2.02× / 1.55–1.92× |
| D4 Portability | within 1.2% | within 6% | Capability does not depend on which vendor procurement chose |
| D5 Diagnostic | n/a | |ΔAUC| ≤ 0.001; cosine 0.9998 | No measurable change on the four cohorts evaluated |
Cancer pathology foundation models are compared almost entirely on AUC on retrospective cohorts, where competitive backbones now differ in the third decimal place. That is not the number a pathology department decides on. A department decides on how many slides a card clears in a shift, how many cards it must therefore buy, the recurring energy bill, whether the capability survives whichever accelerator procurement supplied, and whether protected health information has to leave the building. We argue that the reporting gap, not the accuracy gap, is what separates pathology foundation models from clinical impact, and propose a deployment card: five axes, reported jointly on one declared clinical workload, in units a department already plans in.
We fill the card in end to end on a real clinical slide, across three ViT backbones spanning an order of magnitude in size (DINO ViT-S, UNI, Virchow), on both an NVIDIA H100 and an AMD MI300X, using fused inference operators written once and run unmodified on either vendor. For UNI, one card goes from 222 to 364 slides per shift on NVIDIA and 224 to 348 on AMD, energy per slide halves, and across four cancer cohorts every |ΔAUC| is at most 0.001. Read in departmental units the result stops being a benchmark and becomes a procurement fact: a 2,000-slide day that needed ten on-premises cards needs six.
1The arithmetic a pathology chief actually does
One day of routine sign-out at Memorial Sloan Kettering comprised 2,091 glass slides. Encoding one slide of the size measured here with UNI on a current data-centre accelerator takes 129.6 s, so a single card clears 222 slides in an 8-hour shift. A department at MSK's volume therefore needs ten cards before a single foundation-model result reaches a pathologist, and a 500-slide-a-day hospital needs three. That integer, not the third decimal of AUC, is what a capital request has to justify, and it appears in no pathology FM results table we are aware of.
The vendor matters because a hospital is not provisioned like a hyperscaler. It owns one or a few accelerators, and which vendor supplied them is a procurement outcome rather than a research choice. With PHI ruling out elastic cloud capacity, an acceleration tuned for one vendor's stack is, for a site that bought the other, not a result at all.
2The deployment card
Five axes, each in a unit a department already plans in, all measured on one declared workload: a real slide encoded end to end, including tile decode and host-to-device transfer.
- D1 Slide latency. Median seconds per slide, end to end. Fixes turnaround.
- D2 Throughput. Slides per shift, read two ways: cards needed for a daily volume, N = ⌈V/T⌉, and GPU-days to re-encode an archive. Reporting only T hides that the decision is integer-valued: a speedup crossing no ceiling buys a site nothing.
- D3 Energy and cost. Integrated board power from on-device telemetry. Cost is stated as an interval that holds for every electricity rate and amortization schedule, so no stale GPU price is assumed.
- D4 Portability. Claimed only with matched baseline and optimized results for every backbone on every vendor, from identical operator sources.
- D5 Diagnostic effect. ΔAUC and balanced accuracy per cohort, plus embedding cosine to baseline, which a site can recompute on its own slides without labels.
| Backbone | H100 baseline | H100 optimized | Speedup | MI300X baseline | MI300X optimized | Speedup |
|---|---|---|---|---|---|---|
| ViT-S/16 (22M) | 697 | 993 | 1.42× | 706 | 983 | 1.39× |
| UNI ViT-L/16 (307M) | 222 | 364 | 1.64× | 224 | 348 | 1.55× |
| Virchow ViT-H/14 (632M) | 119 | 204 | 1.71× | 120 | 207 | 1.72× |
3The change being evaluated
Three encoder operations are replaced with fused Triton equivalents: attention specialized to the 197-token sequences of pathology tiles with the following projection folded in, LayerNorm fused into the linear layer that consumes it, and the tail of patch embedding. Gains are measured against PyTorch's own compiler, not unoptimized code. The compiler cannot fuse across the vendor matmul library, and the vendor attention kernel is tuned for language-model sequence lengths, which is the headroom these operators occupy.
The ablation matches the profile's prediction in rank order: attention (~46% of compute) gives 1.33×, adding LayerNorm+Linear (~41%) reaches 1.58×, and patch embedding (~6%) adds the last step to 1.64×. Two operators carry essentially the whole win, so a clinical codebase can ship a two-kernel build and inherit less maintenance.
4Why not quantization or a smaller model
int8, fp8 or a distilled student would plausibly offer more than 1.6×. But under clinical change control they are different objects: quantization changes the numeric representation of every weight, distillation replaces the weights entirely, and revalidation is paid for in pathologist hours. Fusion changes only the order in which the same arithmetic runs on the same validated weights, at the same precision, feeding the same heads. It is the smallest change that produces a departmentally meaningful speedup, and it composes with quantization for a site willing to revalidate.
| Task | Cohort | Base AUC | Optim. AUC | ΔAUC | Δ Bal. acc. |
|---|---|---|---|---|---|
| Tumor detection | CAMELYON16 | 0.948 | 0.947 | −0.001 | −0.002 |
| Subtype classification | TCGA-NSCLC | 0.961 | 0.960 | −0.001 | −0.001 |
| Biomarker (MSI) | TCGA-CRC | 0.842 | 0.843 | +0.001 | +0.000 |
| Grading | TCGA-RCC | 0.889 | 0.888 | −0.001 | −0.001 |
5Run it, and learn from it
The service reading. Because cards come in integers, the saving is not uniform in volume. At 250 slides a day the change is whether a second card is needed at all (2 → 1); at 500 it is 3 → 2; at an MSK-scale day it is three or four cards not bought. Between those points are volumes where the speedup buys nothing, which is why the card reports T and leaves N to the reader's own volume.
The research reading. Few-shot work on rare cancers and emerging biomarkers, and continual adaptation after a backbone refresh, both begin by re-encoding a retrospective archive, which cannot be triaged for the class it is searching for. For 100,000 slides that falls from 150 GPU-days to 92, roughly five calendar months on one card against three, and from 2,020 to 1,000 kWh. A department that cannot spare five months of its only card does not run the study.
6A site-side acceptance gate
The paper's D5 is an author-side statistic. Its counterpart is executable in the week a department reads the card, and needs no labelled cohort:
- Fix a local panel spanning the tissue types the model will see.
- Encode it under both the incumbent and the proposed configuration.
- Compute mean cosine similarity between the two sets of tile embeddings.
- Accept only above a site-chosen threshold, with 0.9998 as reference.
- Repeat whenever either side changes.
7Limitations
Inference only. Systems axes are measured on one real slide, and throughput, energy and cost assume a single card saturated by a steady queue, the on-premises batch regime rather than intra-operative use. Card counts assume slides of comparable tile count and no scheduling loss. Energy excludes host, storage and cooling. Measurements are on data-centre accelerators, so relative trends should carry to a smaller on-site card but absolute figures should not. D5 reports the size of downstream differences rather than a pre-registered equivalence test, which is the most valuable extension.
More research
My other work at Embedded LLM builds portable, fused GPU kernels for biological and clinical foundation models, measured on both NVIDIA and AMD.
- NeurIPS 2026Change the Metric, Change the Winner: Auditing How We Decide Which Clinical Foundation Model to Deploy in Cancer Pathology
- MOML 2026From 13 Hours to 4.6: Profiling-Driven Kernel Fusion for TensorNet in Molecular Simulation
- COLM 2026Deploying Clinical Language and Vision-Language Models Where the Data Lives
- ICML 2026From 805 ms to 23 ms: Accelerating State-Space Models for Real-Time ICU Monitoring
- ISCA 2026When the LLM-Tuned Stack Misses: An Infrastructure View of Biological Foundation Model Inference Across NVIDIA and AMD
Let's stay in touch
Happy to talk about deploying pathology foundation models on-premises and this card; questions and feedback are always welcome.