Evidence explainer

Digital health and AI

Model Cards and Nutrition Labels for Health AI, Explained

A health AI model card is a one-page summary of what a model does, who it was built for, how it performs, and where it should not be used. Voluntary ones help buyers compare; FDA labeling binds.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Why the food-label metaphor fits
  3. What sits on the card
  4. Who checks the numbers
  5. How this differs from an FDA label
  6. Reading a card without being fooled

A health AI model card, sometimes called a nutrition label, is a short standardized summary that tells you what an artificial intelligence tool does, which patients it was built and tested on, how well it performed, and where it should not be used. Think of the panel on the side of a cereal box: the point is not to sell the product but to put the facts a careful buyer needs in one predictable place. The Coalition for Health AI (CHAI) has published such a template as a voluntary aid for hospitals comparing products before purchase. That is a very different thing from the label the U.S. Food and Drug Administration (FDA) clears when it authorizes a machine-learning-enabled medical device, which carries the force of regulation. Both push toward more honest disclosure, but they do not carry the same weight.

Key points#

Why the food-label metaphor fits#

Nutrition labels work because they are boring and predictable. Every box uses the same fields in the same order, so you can compare two cereals in seconds without reading the marketing. A model card borrows that discipline. Instead of chasing performance claims across sales decks, slide-ware, and scattered PDFs, a procurement team can line up several vendors and read the same fields, in the same order, for each one. CHAI describes its template as a starting point for the people evaluating AI tools during purchasing, and as an open, cost-free resource rather than a members-only benefit. Participation is voluntary, and a developer can publish a card regardless of coalition membership.

What sits on the card#

A well-built card collects the items a buyer would otherwise have to reconstruct from a sales conversation. Three groups do most of the work.

Who made it and what it is for#

The card names the developer, the intended uses, and the specific patient population the tool targets. Scope is not a footnote here; it is the headline. A sepsis-prediction model validated on adult inpatients is a different product from one aimed at a pediatric emergency department, and the card exists to make that boundary explicit rather than letting a buyer assume the tool travels everywhere.

What it runs on#

This is the ingredient list: the type of model, the kinds of data it consumes, and the security and compliance accreditations behind it. It answers a plain question, which is what has to be true about your data and your systems for this tool to work as described.

How it performs, and where it fails#

Here you find the key performance metrics, maintenance needs, known risks, out-of-scope uses, documented bias, and ethical considerations. These last few are the most valuable and the easiest to leave out of a brochure. A card that names its own blind spots is doing its job. A card that reads like an advertisement is not.

The payoff of all this structure is comparability. Read three cards side by side and the differences that matter, wrong population, thin evidence, missing limitations, tend to surface on their own.

Who checks the numbers#

A label is only as good as the testing behind it, so the harder question is who produced the metrics a card reports. CHAI has advanced a certification framework for independent assurance labs, the organizations that would evaluate a model and populate those fields. The framework draws on ISO 17025, the main international standard for testing and calibration laboratories, and was developed with the ANSI National Accreditation Board.

Two design choices stand out. First, the framework calls for mandatory disclosure of conflicts of interest between an assurance lab and the developer whose model it tests. That is the classic failure mode of any voluntary review scheme, and naming it up front is the right instinct. Second, the framework leans on FDA thinking about data quality. The ambition is a clean chain of custody for the numbers: an accredited lab runs the evaluation, discloses its independence, and the result lands on a card in a shared format. As of CHAI's 2024 announcement, both the certification framework and the model card were described as forthcoming and open to stakeholder review, so it is fair to treat them as serious proposals rather than settled, widely adopted standards.

How this differs from an FDA label#

The regulatory track is a separate mechanism with separate teeth. On June 13, 2024, the FDA, Health Canada, and the United Kingdom's MHRA jointly published "Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles," building on the ten good-machine-learning-practice principles the same three agencies set out in 2021. In that framework, transparency means the degree to which appropriate information about a device, including its intended use, development, and performance, is clearly communicated to the people who rely on it.

Three differences are structural.

Legal weight. A CHAI card is a voluntary disclosure. FDA-cleared labeling is a regulatory instrument. A device that drifts from its cleared label can face enforcement; a vendor whose voluntary card oversells faces mainly reputational and contractual fallout.

Reach. FDA oversight attaches only to software that meets the definition of a medical device. A large share of health AI, including administrative triage, ambient documentation, and operational tools, falls outside that boundary and gets no mandatory label at all. Voluntary frameworks are one of the few disclosure mechanisms that reach those products.

What is actually guaranteed. Guiding principles are principles, not a pass-or-fail checklist. They shape what regulators expect to see across a device's life cycle. A model card, by contrast, is a fixed document you can read today. The two complement each other: the regulatory principles set direction, and voluntary labels try to make disclosure concrete and comparable, often before and beyond the point where regulation applies.

Reading a card without being fooled#

A model card is a claim, not a verdict, and a few habits keep it honest. Start with the population the card describes and ask whether it matches the patients in front of you, because a strong metric earned on the wrong group is not reassuring. Next, look at who generated the numbers and whether any independence was disclosed, since a vendor grading its own homework is the weakest form of evidence. Finally, read the out-of-scope and known-bias sections first rather than last, because a card that lists no limitations is usually incomplete rather than flawless. Treat the label as the opening of due diligence, not the end of it.

Sources and further reading

  1. CHAI: Advances Assurance Lab Certification and Nutrition Label for Health AI
  2. FDA: Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles
  3. FDA: Good Machine Learning Practice for Medical Device Development: Guiding Principles

Questions and answers

Is a model card the same as FDA clearance?

No. A model card, including the CHAI nutrition label, is a voluntary disclosure a developer chooses to publish. FDA clearance is a regulatory authorization, and the label the FDA clears is legally binding. A tool can have a polished card and no FDA involvement, or FDA clearance and no public card.

Does every health AI tool need FDA authorization?

No. FDA oversight applies only to software that meets the definition of a medical device. Many tools used in care settings, such as scheduling, documentation, and operational systems, sit outside that definition, which is exactly where a voluntary card can add the transparency regulation does not require.

What is the single most useful field on a card?

The intended population, read together with the out-of-scope and known-bias sections. Together they tell you whether a tool was built for patients like yours, and where its makers already expect it to struggle.