PRISM2 doesn't just scan pathology slides — it argues with itself about them

PRISM2 doesn't just scan pathology slides — it argues with itself about them

Paige and Microsoft trained a 4-billion-parameter model to answer diagnostic questions instead of just labeling tissue, using dialogue pulled from nearly 700,000 real cancer cases.

Written by OutOfToken AI

August 10, 2026 · 4 min read · Synthesized from reporting by AI News · How this works

AI Verified · 9/10

Most pathology AI still works like a very expensive stamp: feed it a slide, get back a classification. PRISM2, built jointly by digital pathology company Paige and Microsoft, takes a different approach — it reads whole-slide images the way a resident might discuss a case with an attending, generating text answers to diagnostic questions rather than sorting pixels into buckets.

Turning tissue into text

A single whole-slide image can contain thousands of individual tile-level regions, each capturing a different patch of tissue at extreme magnification. PRISM2 uses a perceiver-based encoder to compress all of those tile embeddings into one unified slide-level representation, rather than analyzing each patch in isolation.

Dialogue as supervision

What sets PRISM2 apart is what it's trained to predict. Instead of just aligning a slide's representation to a static clinical report, the model is trained on question-and-answer dialogue generated from those reports using GPT-4o, giving it fine-grained diagnostic supervision rather than a single broad label.

"PRISM2 is trained on 2.3 million whole-slide images spanning nearly 700,000 specimens, with dialogue supervision drawn from 685,507 pathology reports from Memorial Sloan Kettering Cancer Center."

A three-part architecture

The model's pipeline runs through a slide encoder, a language encoder, and a large language model decoder built on a 4 billion parameter LLM. That decoder is what lets PRISM2 generate actual diagnostic text — answering questions about a slide — rather than outputting a fixed classification score, which its creators describe as the largest multi-modal slide-level pathology foundation model built to date.

Whether question-answering supervision translates into meaningfully better diagnostic accuracy in clinical practice — rather than just a more articulate model — is something only broader validation can settle. But PRISM2 signals where slide-level pathology AI is heading: away from single-label classifiers and toward models that can be interrogated, case by case, in something closer to clinical conversation.

Editorial Note

The research sources comprehensively corroborate all major factual claims in the article—architecture components, training dataset size, specimen counts, report sources, and technical approach. Sources 1 and 2 appear to be official/academic documentation. The article's sole skeptical note about validation in clinical practice remains appropriate and is not contradicted by the sources.

Claim Tracker

AI-assessed

VerifiedPRISM2 is built jointly by Paige and Microsoft

Source 1, 2, 4, and 6 all confirm Paige and Microsoft as builders of PRISM2.

VerifiedThe model is trained on 2.3 million whole-slide images spanning nearly 700,000 specimens

Source 1 states 'nearly 700,000 specimens, totaling 2.3 million WSIs.' Source 2 confirms '2.3 million WSIs.' Source 4 and 6 confirm the 2.3 million figure.

VerifiedDialogue supervision is drawn from 685,507 pathology reports from Memorial Sloan Kettering Cancer Center

Source 4 explicitly states '685,507 MSK pathology reports.' Source 6 corroborates this number and source attribution.

VerifiedThe decoder is built on a 4 billion parameter LLM

Source 2 states 'a 4 billion parameter large language model (LLM).' Source 1 confirms the three-component architecture including an LLM decoder.

VerifiedPRISM2 uses a perceiver-based encoder to compress tile embeddings into a unified slide-level representation

Source 3 and 6 both state PRISM2 'reads whole-slide images through a perceiver-based encoder trained jointly on tissue tiles' and 'aggregates thousands of tile embeddings per slide into one representation.'

Ask AI about this story

// discussion

sign in to join the discussion