07/04/2026
Stanford proved that GPT-5, Gemini, and Claude can appear to see your images when they are not actually looking at them.
The illusion of visual understanding. 🙌🏻
Researchers at Stanford removed the images from visual AI benchmarks and asked frontier models to answer questions about them anyway. No images. Nothing to look at. Blank.
The models described the images in detail. Gave confident diagnoses. Identified objects and abnormalities. In images that did not exist.
They did this over 60% of the time. Zero uncertainty. No "I don't see an image." With standard evaluation prompts, the rate went up to 90 to 100%.
Stanford calls this the "mirage effect." Not a hallucination. A hallucination is getting details wrong about a real input. A mirage is fabricating the entire input, then reasoning about it as if it exists.
They tested GPT-5.1, Gemini-3-Pro, Gemini-2.5-Pro, and Claude Opus 4.5 on six major benchmarks. Removed every image. The models still retained 70 to 80% of their original scores. On medical benchmarks, up to 99%.
Then Stanford did something that broke the entire field.
They took a 3-billion-parameter text-only model. Never seen a single image. Trained it on radiology questions with the images removed.
This blind model outperformed every frontier multimodal model on the held-out chest X-ray benchmark. It outperformed human radiologists by more than 10%.
A model that has never seen an image beat the world's best AI and human doctors at reading chest X-rays. Because the test was never actually testing vision. It was testing text.
When Stanford removed every question models could answer without images, 74 to 77% of each benchmark was eliminated.
The medical bias is the most dangerous part. When these models hallucinate scans, they do not hallucinate healthy results. They hallucinate heart attacks. Melanoma. Carcinoma. Brain nodules. Conditions that trigger emergency intervention.
This paper is co-authored by Fei-Fei Li, arguably the most important figure in the history of computer vision. The person who created ImageNet.
230 million people ask AI health questions every day. The models they are asking can answer confidently without ever looking at the images. And nobody can tell the difference from the output alone.
- The blind radiology model is the finding that should end careers.
Stanford took a 3-billion-parameter text model. Never trained on a single image. Fed it radiology questions with the images stripped out. It beat GPT-5.
It beat Gemini. It beat Claude. It beat human radiologists by over 10%.
Every vision benchmark score you have ever seen is now suspect.
- The medical bias is the part that should terrify you.
When these models hallucinate scans that do not exist, they do not hallucinate healthy results.
They hallucinate heart attacks. Melanoma. Carcinoma. Brain nodules. Conditions that trigger emergency intervention.
The model does not err on the side of caution. It errs on the side of the worst possible diagnosis.
- Stanford tested GPT-5.1 on MicroVQA, a microscopy imaging benchmark. With images: 61.5% accuracy.
After removing every question the model could answer without seeing the image: 15.4%.
Three quarters of its "visual understanding" was pattern matching in text. The actual vision ability is roughly one sixth of what the benchmarks report.
- The mirage effect is worse than hallucination and nobody is treating it that way.
A hallucination is getting details wrong about something real. A mirage is fabricating the entire input and then building a complete analysis on top of it.
With reasoning traces indistinguishable from real ones. With full confidence. With no acknowledgment that anything is missing.
You cannot detect it from the output alone.
- This paper is co-authored by Fei-Fei Li. She created ImageNet, the dataset that started the entire deep learning era.
She is arguably the single most important person in the history of computer vision.
She is now telling you that high scores on visual benchmarks do not mean these models actually understand what they see.
The person who taught machines to see is telling you the tests we use to prove it are broken.
-/arxiv.org/abs/2603.21687