The AI DownsideDocumenting AI's downsides

Safety

What an AI Model Card Actually Tells You — and What It Leaves Out

It arrives looking like a safety certificate, but it is closer to a brochure the manufacturer chose to print — and the most revealing part is often the section that isn’t there.

Editorial illustration for “What an AI Model Card Actually Tells You — and What It Leaves Out”.

A new model lands, and it no longer arrives alone. Alongside the launch post and the benchmark charts comes a sober-looking PDF — a “model card,” or on the larger releases a “system card” — dense with evaluation tables, red-team summaries and risk ratings. It has the typography of a regulatory filing and the heft of something official. It looks, in short, like a safety certificate.

It is not one. A model card is a document the company wrote about its own product. That is not a scandal, and it is not nothing: a decade ago you got a blog post and a vibe, so a structured account of what a model is for and how it was tested is real progress. But the register of the thing — the charts, the clinical prose, the appendices — invites you to read it as an independent audit, when it is closer to a brochure the manufacturer chose to print.

So the useful question is not “what does the card say.” It is “what does it leave out, and who checked it.” Answer those two and the card becomes genuinely useful. Take it at face value and it becomes a very polished way to feel informed without being so.

What a model card is, and where it came from

The idea is specific and traceable. In a 2019 paper titled Model Cards for Model Reporting, a group of researchers — among them Margaret Mitchell and Timnit Gebru, then at Google — proposed that every released machine-learning model should ship with a short, one-to-two-page record. The explicit analogy was the warning label on food and electronics: a standard place to state what the thing is for, what it is not for, how it was tested, and where it is likely to fail.

The proposed sections were sensible and still shape the format: model details, intended use (primary uses, primary users, and pointedly, out-of-scope uses), the factors and groups performance might vary across, the metrics, the evaluation and training data, ethical considerations, and caveats. A companion idea, “datasheets for datasets,” did the same job for the data underneath. The goal was to let someone decide whether a model suited their context before they wired it into something that mattered.

What is easy to miss, because almost nobody reads to the end of the paper, is that the authors built the escape hatch themselves. A model card, they wrote, is only as good as the honesty of whoever fills it in: its “usefulness and accuracy … relies on the integrity of the creator(s) of the card itself.” They went further, doubting it was even possible, “at least in the near term,” to standardise cards enough “to prevent misleading representations of model results.” The inventors shipped the disclaimer in the box. Everything below is a footnote to that sentence.

What a good card does tell you

Start with the steel-man, because the best cards are substantial and it would be reverse-hype to pretend otherwise. The frontier labs have turned the one-page idea into something much larger, and a modern system card can run to dozens of pages of genuinely checkable material. OpenAI’s GPT-5 system card is a fair example: it documents fairness and bias testing, a battery of safety evaluations, and the results of pointing the model at the things you least want it to be good at.

The red-teaming alone is not trivial. The GPT-5 card reports more than 5,000 hours of adversarial testing by over 400 external experts, probing for violent-attack planning, jailbreaks, prompt injection and help with bioweapons. It maps results onto a published risk policy with named capability thresholds. This is the opposite of “we take safety seriously”: it is specific, dated and, crucially, the sort of claim that could embarrass the company if it turned out to be wrong. That is exactly the test we set in what ‘AI safety’ actually means — prefer the claim that is falsifiable — and on capabilities, a good card passes it.

Worth separating two documents that get conflated here. A system card is about one release; it is not the same as the standing risk policy — the safety frameworks the labs promised, such as OpenAI’s Preparedness Framework or Anthropic’s Responsible Scaling Policy. The framework is the rulebook the company says it follows; the card is its homework for one model. You need both to judge either, and a card that cites its framework is doing the right thing.

The section that’s almost always blank

Now the gap, and it is remarkably consistent. The single most important thing about a model — what it was trained on — is the thing cards disclose least. This is not a hunch; it has been measured. Stanford’s Foundation Model Transparency Index scored ten major developers, including OpenAI, Google and Meta, against 100 transparency indicators in 2023. The mean score was 37 out of 100. The best anyone managed was 54.

The worst-scoring region was what the researchers call “upstream”: the data, the labour and the compute that built the model. Several developers scored zero across the entire upstream category. Most strikingly, no company scored a single point for disclosing who created its training data or the copyright and licence status of that data. That blank is not an oversight; it sits directly on top of the lawsuits over who owns the words that trained your AI. A detailed data section is a litigation exhibit waiting to be subpoenaed, so the section stays vague by design.

Some of that reticence is defensible — genuine trade secrets exist, and “we scraped the open web” is at least partly true and partly unknowable even to the lab. But the effect on you is the same whatever the motive: the card describes the engine in loving detail and stays quiet about the fuel.

Capabilities get a chapter; limitations get a sentence

There is a structural tilt in what cards emphasise, and once you see it you cannot unsee it. The Stanford index found that developers were reasonably forthcoming about what their models can do — capabilities, demonstrations, benchmark wins — and markedly worse at the adjacent, less flattering subdomains: limitations, risks, and the effectiveness of their own mitigations. Just two of the ten developers meaningfully demonstrated their models’ limitations. None offered externally reproducible or third-party assessments of whether their safety mitigations actually worked.

This is not only a frontier-lab habit; it runs through the whole ecosystem. A systematic analysis of 32,111 model cards on the Hugging Face hub, published in Nature Machine Intelligence, found that while 44% of models carried a card at all, the sections were wildly uneven. A training section was common. An evaluation section appeared in only 15% of cards, a limitations section in 17%, and an environmental-impact section in a vanishing 2%. The researchers noted a drift they politely called an “increasing reluctance to address the limitations of models.”

A model card is the brochure the manufacturer chose to print — and the section that isn’t there is data too.

The incentive is obvious and human. The card doubles as marketing, and nobody writes a glowing limitations section for their own launch. But a document that lists every strength and hurries past every weakness is not a safety artefact; it is a sales sheet wearing a lab coat. The shape of the omissions — loud on capability, quiet on failure — tells you what the document is really for.

Graded by the people who sat the exam

Even the numbers that are there deserve a particular kind of reading, because of who produced them. The evaluations in a system card are, with few exceptions, designed by the company, run by the company, and reported by the company. There is no fixed standard that says which tests a card must include, no requirement that the methodology be reproducible, and — unlike a scientific paper — no peer review standing between the claim and the reader.

That matters more than it sounds, because self-reported performance skews optimistic in entirely predictable ways. A 2025 study of 500 widely used models found that around 88% of authors overstated their model’s performance in its own card, and 96% gave no account of bias, risks or limitations at all. These were mostly smaller community models, not frontier releases — but the direction of the bias is the same one the incentives predict everywhere, and the frontier labs face the larger marketing pressure, not the smaller.

It is the same trap we flagged in why benchmarks mean less than you think: a number is only as trustworthy as the test behind it, and a test you cannot inspect or reproduce is a marketing figure with error bars drawn in pencil. A reported eval score is a reason to ask how the eval was run. It is not, on its own, a reason to believe the answer.

No two cards are the same shape

The 2019 authors worried that cards could not be standardised enough to prevent misleading representations, and the worry aged well. There is still no agreed schema. One lab’s “system card” is another’s “model card” is another’s “transparency report,” and the contents vary as much as the names. That makes the single most useful operation — comparing two models on the same basis — surprisingly hard, because the two cards rarely answer the same questions in the same units.

It also makes absence ambiguous in a way a standard would fix. When a card omits a robustness evaluation, you cannot tell whether the lab ran the test and buried a bad result, or never ran it, or considered it out of scope. A standard form would at least force a “not assessed” where today there is only silence. The freedom to choose your own sections is the freedom to make your weak spots disappear by not naming them.

The score that went up — and the asterisk

Concede the genuinely good news, because there is some. When Stanford re-ran its index six months later, in May 2024, the mean score had jumped from 37 to 58 out of 100. Every developer assessed in both rounds improved. Public pressure, it turns out, works: name the opacity and some of it recedes. That is the whole theory of an index, and on this evidence it holds.

But read the asterisk. The second-round process was different: instead of only searching for public information, the researchers asked companies to submit their own transparency reports, and the developers duly disclosed, on average, 16.6 indicators’ worth of information that had not been public before. So the leap partly measures what firms will say when a respected institution asks nicely and publishes the league table. Useful — but prompted disclosure, to a friendly auditor, is a softer thing than routine transparency. And the regions that stayed stubbornly opaque across both rounds were the familiar ones: copyright status, data access, data labour and downstream impact.

From convention to law

For most of their life, model cards were etiquette — a norm the research community adopted and the labs mostly honoured on their own terms. That is changing. Under the EU’s AI Act, providers of general-purpose AI models have had to meet real documentation duties since 2 August 2025. They must draw up and maintain technical documentation covering the model’s training, testing and evaluation, keep it available to regulators, and pass structured information to the downstream developers who build on the model. A voluntary Code of Practice even supplies a standard Model Documentation Form — the closest thing yet to the schema the field has lacked since 2019.

This is the most consequential shift in the whole story, and it is genuinely pro-consumer: documentation you could previously only hope for is becoming something a regulator can demand. But keep expectations calibrated. The Act obliges makers to publish a “summary” of training data, not the data itself, and the duties are threaded through with protections for trade secrets and confidential business information. Much of the detailed documentation goes to the regulator, not to you. The law turns the brochure into a filing — a real improvement — without obliging anyone to make it a confession.

How to actually read one

None of this means ignore the card. It means read it the way you would read a company’s own annual report: valuable, and written by an interested party. The framework the field keeps returning to — the US standards body’s AI Risk Management Framework — treats documentation as one input to trust, not the whole of it. A few habits turn the document from reassurance into evidence:

  • Read the gaps before the graphs. Note which standard sections are missing — limitations, training data, downstream impact — because on the numbers those are the ones most often left out, and the omission is the finding.
  • Start with intended use and out-of-scope. This is the one part written to protect you, and the part that tells you, in the maker’s own words, where they expect the thing to break.
  • Ask who ran the test. An evaluation by the vendor is a claim; one by a named third party, or with a reproducible method, is closer to evidence. Weight them differently.
  • Distrust the polish. A card heavy on capability charts and light on limitations, data and method is optimised for the launch, not for your risk assessment.
  • Cross-check the quiet risks. Known weaknesses the card underplays — that models still hallucinate, for instance — are your responsibility to remember, because the card has every incentive to let you forget.

The brochure and the inspection report

Model cards are one of the better ideas the field has had about its own accountability, and they are getting longer, more detailed and, thanks to the EU, more compulsory. A good system card is worth reading closely; the capability and red-team sections of the frontier ones are real, checkable work that did not exist a few years ago. Credit where due.

But the document was designed to behave like an inspection report while remaining, structurally, a brochure: written by the maker, graded by the maker, free to omit what flatters least, and blank exactly where the stakes are highest. The fix is not to discard it but to read it for what it is — the most polished account the company was willing to give of its own product, no more binding than that. The sections it skips are telling you something. The trick is to keep reading after the charts run out.

Frequently asked questions

What is the difference between a model card and a system card?

The model card is the original 2019 concept: a short document about one model’s intended use, evaluation and limits. “System card” is the label the big labs now use for a longer release document covering the whole deployed system — safety evaluations, red-teaming and risk classifications — around a model such as GPT-5 or Claude. In practice the terms are used loosely, and both are written and published by the maker rather than an independent party.

Does a model card tell you what the AI was trained on?

Almost never in any detail. Training data is the single biggest blind spot: Stanford’s transparency index found that no major developer disclosed the copyright or licence status of its training data, an omission tangled up with the ongoing copyright lawsuits. The EU AI Act now requires a public “summary” of training data for general-purpose models, but confidentiality carve-outs mean you get a description, not the actual list.

Can I trust the numbers in a system card?

Treat them as the maker’s own figures. The evaluations are usually designed, run and reported by the company itself, to no fixed standard, with no peer review and seldom any independent audit. That does not make them false — frontier cards contain real, useful data — but it means they are a starting point for scrutiny, not a verified certificate. A missing section is information too.

Are AI model cards legally required?

Increasingly. Under the EU AI Act, providers of general-purpose AI models have had to keep technical documentation and share it with regulators and downstream developers since 2 August 2025, with a voluntary Code of Practice offering a standard documentation form. In most other places it remains convention rather than law. The direction of travel is from etiquette to obligation — though the obligations still bend to trade-secret protection.

How should I actually read a model card?

Start with what is missing. Read the intended-use and out-of-scope sections first, then check whether limitations and training data are genuinely described or merely waved at, and note who ran the evaluations. A glossy card full of capability charts but thin on limitations, data and reproducible method is marketing until proven otherwise — useful, but not the independent safety report its design is borrowed from.

Sources

  1. Model Cards for Model Reporting — Mitchell, Wu, Zaldivar, Barnes, Vasserman, Hutchinson, Spitzer, Raji & Gebru: proposes the model card (a short, 1–2 page document of intended use, evaluation and caveats) and warns its “usefulness and accuracy … relies on the integrity of the creator(s)”arXiv — ACM FAccT 2019
  2. GPT-5 System Card — sections on fairness and bias, red teaming and the Preparedness Framework; reports more than 5,000 hours of red-teaming by over 400 external testers across violent-attack planning, jailbreaks, prompt injection and bioweaponisationOpenAI
  3. The Foundation Model Transparency Index v1.0 — 100 indicators across 10 major developers; mean score 37/100, top score 54; upstream data, labour and compute are the worst-disclosed areas and no company scores for the copyright/licence status of its training dataStanford HAI / CRFM (arXiv:2310.12941)
  4. The Foundation Model Transparency Index v1.1 (May 2024) — mean score rose to 58/100, but largely because developers submitted their own transparency reports, disclosing on average 16.6 previously non-public indicators; copyright, data access and downstream impact stayed opaqueStanford HAI / CRFM
  5. What’s documented in AI? Systematic Analysis of 32K AI Model Cards — Liang et al.: of 32,111 cards, 44.2% of models carry one; a limitations section appears in 17.4%, evaluation in 15.4% and environmental impact in 2.0%; unlike papers, cards face no peer reviewNature Machine Intelligence (2024)
  6. Model Hubs and Beyond: Analyzing Model Popularity, Performance, and Documentation — a study of 500 models finding roughly 80% lack detailed documentation, about 88% of authors overstate their model’s performance in the card, and 96% omit bias, risks and limitationsICWSM 2025
  7. Article 53: Obligations for providers of general-purpose AI models — GPAI providers must keep technical documentation (Annex XI), share information with downstream providers (Annex XII), put in place a copyright policy and publish a public summary of training data; obligations apply from 2 August 2025EU AI Act (Regulation (EU) 2024/1689)
  8. AI Risk Management Framework (AI RMF 1.0, NIST AI 100-1) — the US standards body’s framework, which treats transparency and documentation as core characteristics of trustworthy AI, independent of any single disclosure formatNIST

Related grievances

All articles →