Share this article and save a life!

The FDA just funded a jury of AI judges for radiology reports. 🧠

Not one AI. A panel of them. Cross-examining each other.

Last week, the FDA awarded Cognita Imaging a $1.29 million research contract to build what they call an “LLMs-as-a-jury” evaluation framework for autonomous, generative AI radiology reports. According to HIT Consultant, the 18-month contract was funded under the FDA’s Broad Agency Announcement program for advanced research and development in regulatory science.

Here is why this matters to anyone building in this space.

🔹 Traditional human reader studies evaluate only a few hundred cases.
🔹 They miss rare edge cases, subtle equipment discrepancies, and regional workflow variations.
🔹 Cognita’s framework will be stress-tested across approximately 1 million patient imaging exams from a diverse U.S. clinical cohort.
🔹 The system analyzes performance variance across demographics, care settings, imaging hardware manufacturers, and low-prevalence pathologies.

That gap between a few hundred cases and 1 million is not a footnote. It is the entire problem.

The architecture itself is worth understanding. Rather than trusting a single LLM to score diagnostic drafts, Cognita aggregates outputs across an ensemble of distinct models acting as a virtual panel of expert reviewers. The platform then isolates high-discrepancy outliers for expert human radiologist review. Researchers can see whether errors came from the generative drafting model, disagreement within the jury, or ambiguity in the original human radiologist’s ground truth.

This builds on Cognita’s open-source GREEN benchmark, Generative Radiologist Evaluation and Error Notation, which was designed to capture clinically meaningful discrepancies between ground-truth radiologist reports and AI-generated text.

The project is led by Dr. Akshay Chaudhari, Cognita co-founder and Associate Professor of Radiology and Biomedical Data Science at Stanford University, alongside co-investigator and CEO Dr. Louis Blankemeier.

⚡ What most people are missing here: this is the FDA funding the infrastructure to evaluate generative AI in radiology, not just approve individual products. That is a structural shift in how the agency approaches this category.

At Oatmeal Health, we are building AI for lung cancer screening. Every day I think about what it means to deploy AI at scale across diverse patient populations and imaging environments. The question of how you validate performance across millions of real-world cases, not just curated test sets, is one of the hardest problems in this field.

The FDA is starting to take it seriously. That should matter to all of us.

👉 Follow Jonathan Govette, CEO of Oatmeal Health, for daily healthcare insights on LinkedIn. Deeper dives in The Oatmeal Bite on Substack: https://news.oatmealhealth.com

Share this article and save a life!

Author:


Guest post on Oatmeal Health and reach millions of healthcare professionals. Tell us your story!

Recent Posts