Manpreet Singh Research
Work done at Embedded LLM Embedded LLM
NeurIPS 2026 · ASCI workshop · Poster

364 Slides a Shift, Not 222, on the GPU You Already Own

A Deployment Card for Cancer Pathology Foundation Models at Hospital Scale

Embedded LLM, Singapore
Accepted for a poster at the ASCI workshop, NeurIPS 2026

Paper, code and BibTeX will be posted here after the camera-ready.

AxisBaseline (NV / AMD)Optimized (NV / AMD)What a department reads
D1 Slide latency129.6 / 128.6 s79.1 / 82.8 sFits batch sign-out, not intra-operative
D2 Throughput222 / 224 per shift364 / 348 per shift2,000-slide day: 10 cards to 6. 100k archive: 150 to 92 GPU-days
D3 Energy20.2 / 20.4 Wh10.0 / 10.6 WhCost per slide falls 1.64–2.02× / 1.55–1.92×
D4 Portabilitywithin 1.2%within 6%Capability does not depend on which vendor procurement chose
D5 Diagnosticn/a|ΔAUC| ≤ 0.001; cosine 0.9998No measurable change on the four cohorts evaluated
Figure 1. The filled deployment card for UNI on a real ~12,400-tile clinical slide, on an NVIDIA H100 and an AMD MI300X. The change under evaluation is fused inference operators; weights, numeric precision and diagnostic heads are unchanged. Red: the optimized operating point.
222 → 364slides per 8-hour shift on one H100 with UNI (224 → 348 on MI300X)
10 → 6accelerators to clear an MSK-scale 2,000-slide day
½the energy per slide, 20.2 → 10.0 Wh on NVIDIA
≤ 0.001|ΔAUC| on every one of four cancer cohorts
Abstract

Cancer pathology foundation models are compared almost entirely on AUC on retrospective cohorts, where competitive backbones now differ in the third decimal place. That is not the number a pathology department decides on. A department decides on how many slides a card clears in a shift, how many cards it must therefore buy, the recurring energy bill, whether the capability survives whichever accelerator procurement supplied, and whether protected health information has to leave the building. We argue that the reporting gap, not the accuracy gap, is what separates pathology foundation models from clinical impact, and propose a deployment card: five axes, reported jointly on one declared clinical workload, in units a department already plans in.

We fill the card in end to end on a real clinical slide, across three ViT backbones spanning an order of magnitude in size (DINO ViT-S, UNI, Virchow), on both an NVIDIA H100 and an AMD MI300X, using fused inference operators written once and run unmodified on either vendor. For UNI, one card goes from 222 to 364 slides per shift on NVIDIA and 224 to 348 on AMD, energy per slide halves, and across four cancer cohorts every |ΔAUC| is at most 0.001. Read in departmental units the result stops being a benchmark and becomes a procurement fact: a 2,000-slide day that needed ten on-premises cards needs six.

1The arithmetic a pathology chief actually does

One day of routine sign-out at Memorial Sloan Kettering comprised 2,091 glass slides. Encoding one slide of the size measured here with UNI on a current data-centre accelerator takes 129.6 s, so a single card clears 222 slides in an 8-hour shift. A department at MSK's volume therefore needs ten cards before a single foundation-model result reaches a pathologist, and a 500-slide-a-day hospital needs three. That integer, not the third decimal of AUC, is what a capital request has to justify, and it appears in no pathology FM results table we are aware of.

The vendor matters because a hospital is not provisioned like a hyperscaler. It owns one or a few accelerators, and which vendor supplied them is a procurement outcome rather than a research choice. With PHI ruling out elastic cloud capacity, an acceleration tuned for one vendor's stack is, for a site that bought the other, not a result at all.

2The deployment card

Five axes, each in a unit a department already plans in, all measured on one declared workload: a real slide encoded end to end, including tile decode and host-to-device transfer.

  1. D1 Slide latency. Median seconds per slide, end to end. Fixes turnaround.
  2. D2 Throughput. Slides per shift, read two ways: cards needed for a daily volume, N = ⌈V/T⌉, and GPU-days to re-encode an archive. Reporting only T hides that the decision is integer-valued: a speedup crossing no ceiling buys a site nothing.
  3. D3 Energy and cost. Integrated board power from on-device telemetry. Cost is stated as an interval that holds for every electricity rate and amortization schedule, so no stale GPU price is assumed.
  4. D4 Portability. Claimed only with matched baseline and optimized results for every backbone on every vendor, from identical operator sources.
  5. D5 Diagnostic effect. ΔAUC and balanced accuracy per cohort, plus embedding cosine to baseline, which a site can recompute on its own slides without labels.
BackboneH100 baselineH100 optimizedSpeedupMI300X baselineMI300X optimizedSpeedup
ViT-S/16 (22M)6979931.42×7069831.39×
UNI ViT-L/16 (307M)2223641.64×2243481.55×
Virchow ViT-H/14 (632M)1192041.71×1202071.72×
Table 1. Slides per 8-hour shift on one real ~12.4k-tile slide, matched across vendors and modes. Baselines agree within 1.2% on every backbone, and the gain rises with model size on both vendors as the fixed I/O floor shrinks relative to encoder compute. The clinically deployed encoders are the large ones, so the operationally relevant regime is where the win is largest.

3The change being evaluated

Three encoder operations are replaced with fused Triton equivalents: attention specialized to the 197-token sequences of pathology tiles with the following projection folded in, LayerNorm fused into the linear layer that consumes it, and the tail of patch embedding. Gains are measured against PyTorch's own compiler, not unoptimized code. The compiler cannot fuse across the vendor matmul library, and the vendor attention kernel is tuned for language-model sequence lengths, which is the headroom these operators occupy.

The ablation matches the profile's prediction in rank order: attention (~46% of compute) gives 1.33×, adding LayerNorm+Linear (~41%) reaches 1.58×, and patch embedding (~6%) adds the last step to 1.64×. Two operators carry essentially the whole win, so a clinical codebase can ship a two-kernel build and inherit less maintenance.

4Why not quantization or a smaller model

int8, fp8 or a distilled student would plausibly offer more than 1.6×. But under clinical change control they are different objects: quantization changes the numeric representation of every weight, distillation replaces the weights entirely, and revalidation is paid for in pathologist hours. Fusion changes only the order in which the same arithmetic runs on the same validated weights, at the same precision, feeding the same heads. It is the smallest change that produces a departmentally meaningful speedup, and it composes with quantization for a site willing to revalidate.

TaskCohortBase AUCOptim. AUCΔAUCΔ Bal. acc.
Tumor detectionCAMELYON160.9480.947−0.001−0.002
Subtype classificationTCGA-NSCLC0.9610.960−0.001−0.001
Biomarker (MSI)TCGA-CRC0.8420.843+0.001+0.000
GradingTCGA-RCC0.8890.888−0.001−0.001
Table 2. D5: frozen UNI embeddings with a gated attention-MIL head, mean over 5 seeds. Differences are unsigned in direction, the pattern of numerical reordering rather than degradation. Mean embedding cosine to baseline: 0.9998. Stated as no detected difference on these cohorts, not as a formal equivalence claim.

5Run it, and learn from it

The service reading. Because cards come in integers, the saving is not uniform in volume. At 250 slides a day the change is whether a second card is needed at all (2 → 1); at 500 it is 3 → 2; at an MSK-scale day it is three or four cards not bought. Between those points are volumes where the speedup buys nothing, which is why the card reports T and leaves N to the reader's own volume.

The research reading. Few-shot work on rare cancers and emerging biomarkers, and continual adaptation after a backbone refresh, both begin by re-encoding a retrospective archive, which cannot be triaged for the class it is searching for. For 100,000 slides that falls from 150 GPU-days to 92, roughly five calendar months on one card against three, and from 2,020 to 1,000 kWh. A department that cannot spare five months of its only card does not run the study.

6A site-side acceptance gate

The paper's D5 is an author-side statistic. Its counterpart is executable in the week a department reads the card, and needs no labelled cohort:

  1. Fix a local panel spanning the tissue types the model will see.
  2. Encode it under both the incumbent and the proposed configuration.
  3. Compute mean cosine similarity between the two sets of tile embeddings.
  4. Accept only above a site-chosen threshold, with 0.9998 as reference.
  5. Repeat whenever either side changes.

7Limitations

Inference only. Systems axes are measured on one real slide, and throughput, energy and cost assume a single card saturated by a steady queue, the on-premises batch regime rather than intra-operative use. Card counts assume slides of comparable tile count and no scheduling loss. Energy excludes host, storage and cooling. Measurements are on data-centre accelerators, so relative trends should carry to a smaller on-site card but absolute figures should not. D5 reports the size of downstream differences rather than a pre-registered equivalence test, which is the most valuable extension.

More research

My other work at Embedded LLM builds portable, fused GPU kernels for biological and clinical foundation models, measured on both NVIDIA and AMD.

All research →

Let's stay in touch

Happy to talk about deploying pathology foundation models on-premises and this card; questions and feedback are always welcome.