<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ai-Technical-Documentation |</title><link>https://hwyler.github.io/tags/ai-technical-documentation/</link><atom:link href="https://hwyler.github.io/tags/ai-technical-documentation/index.xml" rel="self" type="application/rss+xml"/><description>Ai-Technical-Documentation</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Fri, 31 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://hwyler.github.io/media/icon_hu_cd51c91342a84ed6.png</url><title>Ai-Technical-Documentation</title><link>https://hwyler.github.io/tags/ai-technical-documentation/</link></image><item><title>Tips for Implementing and Assessing AI Model Cards and Bills of Materials</title><link>https://hwyler.github.io/blog/tips-for-implementing-and-assessing-ai-model-cards-and-bills-of-materials/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/tips-for-implementing-and-assessing-ai-model-cards-and-bills-of-materials/</guid><description>&lt;p&gt;Pull ten AI model cards from ten different vendors. Read the limitations section on each one.&lt;/p&gt;
&lt;p&gt;Most say close to nothing.&lt;/p&gt;
&lt;p&gt;A line about ongoing monitoring. A sentence about responsible use. No numbers, no subgroup breakdown, no named owner, no version tied to the model actually running in production right now.&lt;/p&gt;
&lt;p&gt;That gap is about to matter more than it ever has. High-risk AI systems in the EU now need technical documentation that survives a regulator&amp;rsquo;s questions, not a marketing page. Auditors are starting to ask for the AI components behind a model, the machine learning bill of materials that inventories what actually went into it, not just the card that summarizes it. Most organizations still treat both documents as something you generate once at launch and never open again.&lt;/p&gt;
&lt;p&gt;This piece covers both properly. Start with the bill of materials, the structural inventory a model card sits on top of. Then walk through what belongs in an actual model card, field by field. Then get to the ten tips that decide whether either document holds up when someone outside your team actually reads it.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/07/chatgpt-image-jul-31-2026-05_55_55-pm-edited.png" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="understanding-the-bill-of-material-the-framework-underneath-every-model-card"&gt;Understanding the Bill of Material, the Framework Underneath Every Model Card&lt;/h2&gt;
&lt;p&gt;A model card is a summary. The ML-BOM Machine Learning Bill of Materials is the inventory that summary is supposed to be honest about.&lt;/p&gt;
&lt;p&gt;Think of it as the AI equivalent of a software bill of materials, the practice that got standardized industry-wide once organizations realized nobody could answer &amp;ldquo;which of our systems use this vulnerable library&amp;rdquo; without one. The ML-BOM does the same job for a machine learning model. It answers a blunter question. What, exactly, is inside this thing, and where did each piece come from.&lt;/p&gt;
&lt;p&gt;Model identifiers pin the model to a specific, unambiguous reference, not a friendly nickname that could point to five different checkpoints.&lt;/p&gt;
&lt;p&gt;1. Model metadata covers the basics: name, version, license, developer, purpose, and the parameters that shape behavior. Model architecture documents the network design and how information moves through it.&lt;/p&gt;
&lt;p&gt;2. Datasets records what trained and tested the model and how that data was selected, arguably the hardest field to get right and the one most often left thin. Tokenizers and prompt templates capture how raw input gets converted into something the model actually processes, which matters more than most teams assume once a template changes without notice.&lt;/p&gt;
&lt;p&gt;3. Hardware, software, and frameworks lists every library, runtime, and dependency the model relies on, plus the protocols used when the model operates inside a larger agent or workflow.&lt;/p&gt;
&lt;p&gt;4. Training and testing details cover the computational environment, the hyperparameters, and the evaluation setup. Intended use and ethical considerations state what the model is for, its known limits, and the guardrails around it.&lt;/p&gt;
&lt;p&gt;5. Environmental impact records the resource cost, increasingly a real procurement question rather than a disclosure nobody reads.&lt;/p&gt;
&lt;p&gt;Smaller teams will not populate every one of these on day one, and that is fine. Start with identifiers and datasets, the two fields that carry the most risk if they are wrong, and build outward from there.&lt;/p&gt;
&lt;p&gt;When constructing a machine learning bill of materials, establish the exact model identifier before you document another word. A stable, unique identifier allows your risk systems to automatically match the asset against vulnerability feeds, license databases, and dependency trackers.&lt;/p&gt;
&lt;p&gt;A model referenced only by a friendly display name is a governance dead end. It cannot be mapped to anything systematically. I mandate that teams anchor this identifier first, even if the rest of the documentation remains thin. Once the identifier is locked, every subsequent control in the technical file has a verifiable center of gravity.&lt;/p&gt;
&lt;p&gt;The second failure pattern occurs in data documentation. Move past treating the dataset field as a casual description, you must treat it as evidentiary documentation I routinely reject model cards that summarize data provenance with a single line stating &amp;ldquo;proprietary internal data&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;That phrasing tells an auditor absolutely nothing. It obscures selection bias, masks consent violations, and hides whether your training set overlaps with your evaluation set. Require a precise accounting of the source, the collection methodology, and the known representation gaps. Most critically, demand a direct declaration confirming that your training and evaluation data are strictly disjoint.&lt;/p&gt;
&lt;p&gt;In my practice, this single data provenance field predicts more downstream regulatory exposure than any other metric in your entire technical file.&lt;/p&gt;
&lt;h2 id="what-belongs-in-a-model-card-field-by-field"&gt;What Belongs in a Model Card, Field by Field&lt;/h2&gt;
&lt;p&gt;A model card lives inside the ML-BOM as the description of the model component itself. It breaks into three groups: what the model is, how it performs, and what to watch out for.&lt;/p&gt;
&lt;h3 id="model-parameters"&gt;Model Parameters&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Approach: the general learning method behind the model. Common values include supervised, unsupervised, reinforcement learning, semi-supervised, and self-supervised. This one field tells a reviewer what kind of failure modes to expect before reading another line.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Task: the specific job the model does. Classification, regression, clustering, anomaly detection, generation, and recommendation are typical values. A card that skips this field is asking the reader to guess.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Architecture family: the broad category of network design, such as a transformer, a convolutional network, or a recurrent network. This tells a technical reviewer what kind of behavior to expect at a glance.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Model architecture: the specific implementation, named precisely enough that someone could locate the actual class or configuration behind it, not just a marketing label.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Datasets: what trained and evaluated the model, cross-referenced against the ML-BOM entry rather than restated loosely.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Inputs and outputs: the exact data types the model accepts and produces, described concretely enough to catch a mismatch before integration.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Configuration parameters and hyperparameters: the settings that shaped training and inference, recorded so a future reviewer can tell whether a performance change came from the model itself or from a config tweak.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="quantitative-analysis"&gt;Quantitative Analysis&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Benchmarks: the specific, named tests the model was measured against, not a vague reference to industry standards.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Metrics: the measurements actually reported, defined precisely enough that two different teams would calculate them the same way.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Performance metrics: the results themselves, broken out by the subgroups that matter for your deployment, not one blended number.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Graphics: visual evidence, distributions, and error curves that a single summary statistic cannot show on its own.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="considerations"&gt;Considerations&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Users and use cases: who the model is actually built for, and just as important, who it is not built for.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Technical limitations: the conditions under which the model is known to underperform, stated plainly rather than buried in a footnote.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Performance tradeoffs: what improves and what degrades depending on how the model gets tuned or deployed.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Fairness assessments: how the model performs across the groups relevant to your specific context, with an actual test result attached, not a claim.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ethical considerations: risks named specifically enough to act on, each paired with what mitigates it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Environmental impact: the energy and resource cost of training and running the model, increasingly a line item procurement teams ask for directly.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When I review a model card, I skip the technical specifications and go straight to the considerations section. I do this because it is almost always hollow. Your model parameters and quantitative analysis will usually look perfectly fine. That happens because those metrics are pulled straight out of the training pipeline by an automated script. They require zero additional effort. The considerations section is entirely different. It requires an actual human being to sit down, step back from the code, and critically think through how the system will behave in the real world.&lt;/p&gt;
&lt;p&gt;Because it requires actual judgment, it is exactly the section that gets abandoned the moment an engineering team feels deadline pressure. If your schedule only gives you enough time to review a single part of a model card, make it this one. It tells you instantly whether you are looking at a real risk assessment or just a box-checking exercise.&lt;/p&gt;
&lt;h2 id="field-by-field-assessment-guide-for-model-cards"&gt;Field-by-Field Assessment Guide for Model Cards&lt;/h2&gt;
&lt;p&gt;Model cards started as a fix for a specific problem: AI teams were shipping models with almost no record of what the model was trained on, how it performed across different groups of people, or where it was likely to fail. A model card is the answer to that gap, a structured document meant to travel with the model itself, so that anyone deciding whether to trust it, deploy it, or build a control around it has something concrete to work from instead of a marketing page.&lt;/p&gt;
&lt;p&gt;The approach below treats a model card the way an auditor treats a set of financial statements: every field is either present and adequate, present and thin, or missing entirely, and each of those three states tells you something different about the risk you&amp;rsquo;re inheriting by using the model. A field that&amp;rsquo;s simply absent isn&amp;rsquo;t neutral, it&amp;rsquo;s a signal that either nobody thought to document it or nobody wanted to. Reviewing a model card well means reading past the narrative language vendors tend to favor and asking, field by field, whether what&amp;rsquo;s written actually supports the decision you need to make: approve this model for the use case in front of you, reject it, or send it back with a list of what&amp;rsquo;s missing before a decision can be made responsibly.&lt;/p&gt;
&lt;p&gt;The fields below are ordered the way they typically appear across widely used model card structures, starting with basic identity and working through intended use, technical characteristics, data provenance, performance, fairness, safety, and finally the operational and compliance information that governs the model once it&amp;rsquo;s live. For each field, you&amp;rsquo;ll find the kinds of values you should expect to see, worked examples, how to actually review it, and the vulnerabilities and risks a thin or missing entry tends to expose.&lt;/p&gt;
&lt;h3 id="model-identity-and-basic-details"&gt;Model Identity and Basic Details&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A model name and version string (for example, &amp;ldquo;FraudScore-v3.2&amp;rdquo; or &amp;ldquo;Qwen-7B-Instruct&amp;rdquo;), the model family or architecture type (transformer, gradient-boosted tree, diffusion model), the developing organization, a named contact or team responsible for the model, a license type (&amp;ldquo;Apache 2.0,&amp;rdquo; &amp;ldquo;proprietary, internal use only,&amp;rdquo; &amp;ldquo;research use only, no commercial deployment&amp;rdquo;), and a release date alongside the date of the last update.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Confirm the model can be traced to exactly one accountable owner, not a generic team mailbox, and that the versioning is specific enough to distinguish this release from the last one. Check that the license terms actually match what you intend to do with the model; a &amp;ldquo;research use only&amp;rdquo; license attached to a model someone wants to put into a customer-facing product is an immediate stop, not a footnote.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; A model with no clear owner or inconsistent versioning is a governance failure waiting to surface at the worst possible time, usually during an incident, when nobody can say with confidence which version was actually running in production. This maps directly to the cybersecurity and model drift risk categories referenced in AI assurance frameworks such as the NIST AI Risk Management Framework, and it should be treated as a release blocker for anything classified as high-risk, not a documentation nicety to fix later.&lt;/p&gt;
&lt;h3 id="intended-purpose-and-use-cases"&gt;Intended Purpose and Use Cases&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A description of the model&amp;rsquo;s purpose (&amp;ldquo;triage chatbot for customer support inquiries,&amp;rdquo; &amp;ldquo;credit risk scoring for personal loan applications&amp;rdquo;), the intended task type (classification, generation, forecasting, decision support), the intended user roles (developers, clinicians, customer support agents, automated downstream systems), the intended deployment environment (cloud, on-device, specific geographic regions), and, critically, an explicit list of out-of-scope or prohibited uses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Compare the stated purpose against your actual planned deployment, not against a loose paraphrase of it. If the card lists out-of-scope uses, check every one of them against what your organization or its users might realistically attempt, deliberately or not. If out-of-scope uses aren&amp;rsquo;t listed at all, treat that absence as a documentation gap rather than an implicit &amp;ldquo;anything goes.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Misalignment between what a model was built for and what it actually gets used for is one of the most common root causes of AI-related harm on record, a research model repurposed into a safety-critical workflow, a general-purpose chatbot pressed into a role requiring domain expertise it was never evaluated on. Regulatory frameworks including the EU AI Act treat this misalignment as a primary driver of foreseeable risk to health, safety, and fundamental rights, which makes this field one of the highest-priority checks in the entire card, particularly for anything touching credit, employment, health, or law enforcement decisions.&lt;/p&gt;
&lt;h3 id="model-architecture-and-technical-characteristics"&gt;Model Architecture and Technical Characteristics&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A high-level architecture description (encoder-decoder transformer, convolutional network, ensemble of decision trees), parameter count or model size, input and output formats (text, image, tabular data, bounding boxes, class probabilities), preprocessing and postprocessing steps (tokenization, normalization, output thresholding), and dependencies on external components such as embeddings, retrieval systems, or feature stores.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Check that stated input and output formats actually match what your integration expects; a mismatch here produces silent failures rather than obvious errors, which is worse. Look specifically at any external dependency, a retrieval index, a third-party embedding service, because that dependency now sits inside your risk boundary whether or not it was your engineering decision.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Complex architectures with opaque internal logic raise interpretability risk, which matters most in regulated or high-stakes decisions where a person affected by the output has a right to understand roughly why the model reached its conclusion. Undocumented external dependencies are a supply-chain risk hiding in plain sight: if the retrieval index or embedding provider changes or degrades, your model&amp;rsquo;s behavior changes with it, and nothing in your own testing history would have caught it.&lt;/p&gt;
&lt;h3 id="training-data-description-and-provenance"&gt;Training Data Description and Provenance&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; Data sources (internal transaction logs, licensed third-party datasets, public web-scraped corpora, user-generated content), the time period the data covers, geographic and demographic coverage, collection methods (scraping, sensor data, manual annotation, purchased datasets), known gaps or exclusions, and governance notes covering consent and legal basis for use.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Ask specifically whether the data reflects the population you&amp;rsquo;ll actually be applying the model to. A fraud model trained predominantly on urban transaction patterns and deployed against a largely rural customer base has a documented representativeness gap the moment you check this field, regardless of how strong its aggregate accuracy numbers look. Flag vague provenance statements like &amp;ldquo;collected from the internet&amp;rdquo; as a finding in their own right, not as an acceptable summary.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; This is where the majority of bias and fairness failures originate, since a model can only be as representative as the data it learned from, and it&amp;rsquo;s also where privacy exposure tends to start, since personal or sensitive data folded into a training set without a documented legal basis becomes a downstream liability the moment the model memorizes and later reproduces it. Widely cited work on model documentation, including the original Model Cards for Model Reporting proposal by Mitchell and colleagues, and the EU AI Act&amp;rsquo;s technical documentation requirements under Annex IV, both treat training data provenance as one of the two or three fields that most determines whether the rest of the card can be trusted.&lt;/p&gt;
&lt;h3 id="evaluation-data-and-test-conditions"&gt;Evaluation Data and Test Conditions&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A description of the evaluation dataset&amp;rsquo;s source, size, and coverage, an explicit statement of whether it overlaps with training data, the test environment (offline benchmark, simulated environment, limited pilot deployment), and a stated rationale for why that particular evaluation set was chosen.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; The single most important check here is independence: confirm the evaluation data doesn&amp;rsquo;t overlap with the training data, because contamination between the two produces performance numbers that look excellent and mean almost nothing about real-world behavior. Then check whether the evaluation set actually reflects your deployment distribution, language, region, user population, rather than a convenient benchmark that happened to be available.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Undetected train-test contamination is a data leakage risk that inflates every downstream metric in the card, meaning a reviewer who trusts the accuracy numbers without checking this field is building a risk assessment on a number that was never real. Evaluation on a narrow or non-representative dataset produces a second, quieter failure: strong reported performance that simply doesn&amp;rsquo;t transfer to your actual users, a gap that typically isn&amp;rsquo;t discovered until the model is already live and something has gone wrong.&lt;/p&gt;
&lt;h3 id="performance-metrics-and-results"&gt;Performance Metrics and Results&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; Aggregate metrics appropriate to the task, accuracy, F1 score, area under the ROC curve, BLEU or ROUGE for generation tasks, mean absolute error for regression, along with task-specific figures like precision and recall for the classes that matter most, latency, and throughput. Stronger cards also report robustness under noise or adversarial conditions and confidence or uncertainty estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Match the reported metric to the actual cost of errors in your use case, a high overall accuracy figure can hide an unacceptable false-negative rate on the one category that matters most, a missed fraud case or a missed medical finding, so ask for the specific metric, not just the headline number. Treat a single aggregate figure reported without any breakdown or confidence interval as an incomplete answer rather than a final one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; A model with strong average performance but no reported robustness or calibration information carries hidden risk in exactly the conditions where a control failure would matter most, noisy inputs, distribution shift, adversarial manipulation. This is a well-established gap in AI assurance literature: metrics chosen and reported without transparent methodology or uncertainty bounds create false confidence, and that false confidence is precisely what leads organizations to under-resource the human oversight a model actually needs.&lt;/p&gt;
&lt;h3 id="disaggregated-performance-and-fairness-considerations"&gt;Disaggregated Performance and Fairness Considerations&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; Performance metrics broken out by relevant subgroup, demographic categories, language, geography, device type, alongside fairness metrics such as disparate impact ratio or differences in false positive and false negative rates across groups, and a narrative explanation of any observed disparities and what was attempted to address them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Look specifically for whether the subgroups tested match the population your deployment will actually affect, and check the sample size behind each subgroup figure; a fairness metric computed on a handful of examples from an underrepresented group carries far less statistical weight than the headline percentage suggests. A commonly cited screening threshold in employment and lending contexts, the four-fifths rule, treats a selection rate for any group below 80% of the highest-performing group&amp;rsquo;s rate as a signal warranting further review, a useful sanity check even outside those specific regulatory contexts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Aggregate metrics reported without disaggregation routinely conceal serious disparities that only become visible once you split the results by group, which is exactly why this field carries some of the highest regulatory weight in frameworks like the EU AI Act for any system affecting access to credit, employment, housing, or public services. A card that reports strong overall accuracy but skips this section entirely should be treated as materially incomplete for any use case touching individual people, not as a model that simply &amp;ldquo;didn&amp;rsquo;t need it.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="known-limitations-failure-modes-and-risk-statements"&gt;Known Limitations, Failure Modes, and Risk Statements&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; Documented weaknesses such as degraded performance on rare classes, unsupported languages, or out-of-domain inputs, specific behavioral failure modes for generative models, fabricated citations, sycophantic agreement with a user&amp;rsquo;s incorrect premise, repetitive output loops under certain decoding settings, and explicit statements about conditions likely to produce unreliable output.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Read this section for specificity rather than reassurance. A card stating &amp;ldquo;the model may occasionally produce inaccurate information&amp;rdquo; is not meaningfully different from saying nothing, whereas a card describing the specific conditions under which inaccuracy spikes, long documents beyond a certain token count, ambiguous multi-step reasoning, out-of-domain queries in an underrepresented language, gives you something you can actually build a control around.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; This section is the single richest source of information for building your own risk register entries, because it&amp;rsquo;s the vendor or development team telling you, in their own words, where the model is expected to break. A card with a suspiciously clean &amp;ldquo;no known major limitations&amp;rdquo; statement on a capable, general-purpose model should be treated with active suspicion rather than comfort; every capable model has documented failure modes in the broader research literature, so their absence here usually means nobody looked hard enough, not that none exist.&lt;/p&gt;
&lt;h3 id="safety-security-and-adversarial-considerations"&gt;Safety, Security, and Adversarial Considerations&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; Documented exposure to known AI-specific threats, prompt injection for language models, adversarial example evasion for classifiers, model inversion or membership inference against models handling sensitive training data, along with the specific defenses in place, input and output filtering, rate limiting, access controls, and a statement of residual risk that remains even after those defenses are applied.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Compare the threats the card discusses against your own deployment&amp;rsquo;s actual attack surface. A model exposed to untrusted public input carries a fundamentally different risk profile than the same model running behind an internal, authenticated interface, and the card should reflect that context, not a generic list copied across every deployment scenario. Where the card claims a mitigation is in place, ask what evidence supports that claim, a red-team test result, an adversarial benchmark score, rather than accepting the mitigation&amp;rsquo;s existence as self-evidently sufficient.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; For any model accepting input from users you don&amp;rsquo;t fully control, this is one of the two or three fields that most determines deployment risk, alongside training data provenance and intended use. A capable generative model with no adversarial testing or prompt injection discussion documented anywhere in its card should be assumed vulnerable by default rather than assumed safe by omission, a principle consistent with how established security assessment practice treats undocumented attack surfaces in conventional software.&lt;/p&gt;
&lt;h3 id="privacy-and-data-protection-considerations"&gt;Privacy and Data Protection Considerations&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A statement on whether personal or sensitive data was used in training, data minimization and anonymization practices applied, privacy risk assessments covering re-identification or unintended memorization, and compliance notes addressing data-subject rights where applicable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Even when the card states no personal data was directly stored, check whether the model could still expose privacy risk indirectly, through memorization of rare training examples or through inference of sensitive attributes from otherwise non-sensitive inputs. This distinction, between a model storing data and a model that can be made to reveal information about the data it learned from, is frequently missed in a quick read of this section.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Membership inference and model inversion are established, demonstrated attack classes against models trained on sensitive data, meaning the absence of any privacy discussion in a card for a model trained on personal information should trigger an internal privacy impact assessment before deployment proceeds, not after. This maps directly onto data protection impact assessment expectations found in privacy regulation generally and is treated as a required documentation element under the EU AI Act&amp;rsquo;s technical file requirements for high-risk systems.&lt;/p&gt;
&lt;h3 id="human-oversight-control-and-operational-use"&gt;Human Oversight, Control, and Operational Use&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A stated oversight model, fully automated decision-making, human-in-the-loop review of every output, or human-on-the-loop spot-checking, guidance for how a human reviewer should interpret model outputs, defined escalation thresholds, and any override or manual correction mechanism available to operators.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Check that the recommended oversight level actually matches the stakes of the decision the model informs; a model influencing credit or medical decisions with a card recommending only spot-check review, rather than review of every output, is a mismatch worth escalating regardless of how strong the model&amp;rsquo;s other metrics look. Confirm the guidance given to human reviewers is concrete enough to act on, not a generic instruction to &amp;ldquo;use judgment.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Ambiguous or missing oversight guidance is a leading contributor to automation bias, the tendency of a human reviewer to defer to a model&amp;rsquo;s output even when they have reason to question it, simply because no clear threshold was given for when to intervene. Regulatory frameworks increasingly treat documented, technically enforced human oversight as a non-negotiable requirement for high-risk AI systems rather than a best practice, which makes a thin entry here a strong candidate for a formal finding rather than a minor gap.&lt;/p&gt;
&lt;h3 id="monitoring-maintenance-and-lifecycle-management"&gt;Monitoring, Maintenance, and Lifecycle Management&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A stated monitoring plan covering which metrics are tracked and how often, defined triggers for retraining, a documented version history summarizing what changed between releases, and criteria for eventually retiring or replacing the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Confirm the monitoring plan tracks something meaningful, actual drift in input distribution or output accuracy, rather than only infrastructure uptime, which tells you the system is running but says nothing about whether it&amp;rsquo;s still behaving correctly. Check whether the documentation itself has a stated update cadence tied to the model&amp;rsquo;s own version history, since documentation that isn&amp;rsquo;t updated alongside the model quietly becomes inaccurate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; Every model degrades over time as the world it operates in shifts away from the distribution it was trained on, so the absence of a monitoring and retraining plan is itself an operational risk, not a placeholder to fill in later. This corresponds to the model drift risk category tracked across most AI assurance frameworks, and for any model influencing a recurring, high-volume decision, it deserves the same review rigor as the model&amp;rsquo;s original performance metrics.&lt;/p&gt;
&lt;h3 id="ethical-societal-and-impact-considerations"&gt;Ethical, Societal, and Impact Considerations&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A discussion of potential societal effects, labor displacement, misinformation risk, environmental cost, alongside a named framework of ethical principles the development team applied, fairness, transparency, accountability, and concrete recommendations for responsible use.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Assess whether the stated recommendations are specific enough to act on rather than generic statements of good intent, and consider whether the model could enable harmful uses even outside its stated intended purpose, a general-purpose generation model capable of producing convincing synthetic media, for instance, regardless of what its intended use case was.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; This section matters most for powerful, widely deployable models where the realistic misuse surface extends well beyond the documented intended use, and its absence in a capable model should be read as a gap worth raising with whoever is responsible for use-case approval, not dismissed as a soft or unquantifiable concern.&lt;/p&gt;
&lt;h3 id="environmental-considerations"&gt;Environmental Considerations&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; Estimated energy consumption at different lifecycle stages, training, fine-tuning, and inference, the energy source powering that consumption, and reported carbon dioxide equivalent figures alongside any claimed offsets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Where figures are reported, check whether they cover just training or the full lifecycle including ongoing inference, since a model queried millions of times a day can accumulate an inference-phase footprint that dwarfs its one-time training cost. Treat the complete absence of any environmental disclosure on a large-scale model as a documentation gap rather than an indication the cost doesn&amp;rsquo;t exist, since most providers currently under-disclose this figure rather than having genuinely measured zero impact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; This is a lower-severity field relative to safety, fairness, or privacy, but it is an increasingly explicit regulatory disclosure expectation for general-purpose AI models under emerging AI-specific regulation, and its absence is worth noting in any formal technical file review even where it doesn&amp;rsquo;t block a deployment decision on its own.&lt;/p&gt;
&lt;h3 id="compliance-and-regulatory-alignment-notes"&gt;Compliance and Regulatory Alignment Notes&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A statement of whether the model has been assessed against a specific regulatory classification, such as a high-risk categorization under applicable AI regulation, references to harmonized standards applied during development, and pointers to more detailed supporting technical documentation or risk assessments held elsewhere.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Treat a high-level compliance claim as a pointer, not a conclusion, always ask for the underlying documentation it references rather than accepting the summary sentence as sufficient evidence on its own. Verify that any cited standard or framework is actually applicable to your jurisdiction and use case rather than assumed to transfer automatically from wherever the model was originally developed and assessed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; A vague compliance statement unsupported by an underlying technical file is one of the more common findings in a rigorous model card review, and for any system likely to fall under a high-risk classification in your operating jurisdiction, this gap should be resolved before deployment, not tracked as an open item to close later.&lt;/p&gt;
&lt;h3 id="caveats-and-recommendations-for-deployers"&gt;Caveats and Recommendations for Deployers&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Typical values and examples:&lt;/strong&gt; A consolidated list of known caveats already discussed elsewhere in the card, paired here with concrete deployment guidance, recommended confidence thresholds, suggested human review checkpoints, monitoring configuration recommendations, and rate-limiting guidance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to review:&lt;/strong&gt; Cross-reference every caveat listed here against the corresponding evidence earlier in the card; a caveat mentioned in this closing section without a matching discussion in the performance or limitations fields is a sign the documentation was assembled inconsistently rather than derived from a single coherent evaluation. Check that the recommendations are specific and testable, &amp;ldquo;implement human review for low-confidence outputs&amp;rdquo; is actionable, &amp;ldquo;use responsibly&amp;rdquo; is not.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vulnerabilities, threats, and priority:&lt;/strong&gt; This section is where an incomplete card most often reveals itself, because vague or generic recommendations here usually indicate the underlying evaluation work was equally generic. Treat a strong, specific, evidence-backed recommendations section as one of the better proxies available for judging whether the rest of the card can be trusted, and a thin one as grounds to request the underlying technical assessment before relying on the model for any consequential decision.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/07/chatgpt-image-jul-31-2026-06_00_59-pm.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="field-reference-table-for-model-card-considerations"&gt;Field Reference Table for Model Card Considerations&lt;/h2&gt;
&lt;p&gt;The table below walks through every chapter and field found in the source considerations block, ordered within each chapter from the fields that appear most consistently across model cards to the more specialized, model-specific entries that show up less often. Use it as a companion to the review guidance above: this version focuses on exactly what values each field can take and what each one is documenting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to read this table in practice:&lt;/strong&gt; start at the top of each chapter and work down. The fields near the top of each section are the ones you should expect to find populated in nearly every reasonably complete model card, their absence is a meaningful gap. The fields toward the bottom of each section, the architecture-specific quirks, the granular fairness methodology, the per-lifecycle-stage energy breakdown, show up mostly in the more mature, detailed cards. Their presence is a positive signal about how seriously the model&amp;rsquo;s governance was handled; their absence isn&amp;rsquo;t automatically disqualifying, but it does mean you&amp;rsquo;re working with less information than you could have, and that gap should be logged, not silently assumed away.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Chapter&lt;/th&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;th&gt;Potential Values (with examples)&lt;/th&gt;
&lt;th&gt;Explanation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Users and Use Cases&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Users&lt;/td&gt;
&lt;td&gt;Intended User Roles&lt;/td&gt;
&lt;td&gt;Role labels such as &amp;ldquo;Academic Researcher&amp;rdquo; &amp;ldquo;Enterprise Security Analyst,&amp;rdquo; &amp;ldquo;Edge Device Engineer,&amp;rdquo; &amp;ldquo;Local AI Enthusiast / Privacy-First User&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Names the categories of people expected to interact with or deploy the model. This is the anchor field for the whole section, every use case listed should trace back to at least one of these roles.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Use Cases&lt;/td&gt;
&lt;td&gt;Concrete Use Case Descriptions&lt;/td&gt;
&lt;td&gt;Free-text scenarios, e.g., &amp;ldquo;real-time code completion within an IDE,&amp;rdquo; &amp;ldquo;translating business content while preserving tone and cultural nuance,&amp;rdquo; &amp;ldquo;low-latency triage chatbot escalating complex queries,&amp;rdquo; &amp;ldquo;summarizing long-form research using a 128K context window,&amp;rdquo; &amp;ldquo;on-device visual perception paired with natural-language navigation,&amp;rdquo; &amp;ldquo;analyzing internal security logs without data leaving the firewall&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Describes specific, real applications tied to the roles above. The more concrete the use case (naming a context window size, a deployment environment, a data-sensitivity constraint), the more useful the field is for matching the model against your actual deployment.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Technical Limitations&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hallucination and Inaccuracy&lt;/td&gt;
&lt;td&gt;Plausibility Over Accuracy&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;prioritizes plausible-sounding text over factual accuracy (sycophancy)&amp;rdquo;&lt;/td&gt;
&lt;td&gt;The most universally documented limitation across generative models. Flags that fluent output is not the same as correct output.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Context Window Constraints&lt;/td&gt;
&lt;td&gt;Memory Boundaries&lt;/td&gt;
&lt;td&gt;Token limits, e.g., &amp;ldquo;32,768 native tokens,&amp;rdquo; &amp;ldquo;128K via extended scaling&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Describes how much text the model can process or &amp;ldquo;remember&amp;rdquo; in a single interaction before earlier content is dropped or degraded.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Reasoning and Math Deficiencies&lt;/td&gt;
&lt;td&gt;Multi-Step Logic Gaps&lt;/td&gt;
&lt;td&gt;Descriptive text on struggles with complex, multi-step logic or arithmetic&lt;/td&gt;
&lt;td&gt;Common across LLM families regardless of size; signals where a model needs external tools (calculators, solvers) rather than being trusted to reason unaided.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Knowledge Cutoff&lt;/td&gt;
&lt;td&gt;Frozen-in-Time Knowledge&lt;/td&gt;
&lt;td&gt;A date or version marker, e.g., &amp;ldquo;training data through \[month/year\]&amp;rdquo;&lt;/td&gt;
&lt;td&gt;The model has no access to events or information after this point unless paired with retrieval or search tools.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Opacity (Lack of Traceable Reasoning)&lt;/td&gt;
&lt;td&gt;Black-Box Architecture&lt;/td&gt;
&lt;td&gt;Descriptive text on inability to trace how a specific output was generated&lt;/td&gt;
&lt;td&gt;Explains why standard explainability methods struggle with large, complex architectures, relevant to any interpretability requirement.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Probabilistic Output Inconsistency&lt;/td&gt;
&lt;td&gt;Non-Deterministic Output&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;same prompt yields different results across seeds or context carryover&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Notes that outputs aren&amp;rsquo;t guaranteed to repeat exactly, which matters for testing, auditing, and reproducibility expectations.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Bias Reinforcement&lt;/td&gt;
&lt;td&gt;Training-Data Bias Amplification&lt;/td&gt;
&lt;td&gt;Descriptive text, often flagging synthetic-data effects&lt;/td&gt;
&lt;td&gt;Explains how a model can replicate or amplify biases in its source data, a risk that has grown as synthetic training data use has increased.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(Model-specific examples)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Architecture-Specific Quirks&lt;/td&gt;
&lt;td&gt;E.g., &amp;ldquo;Greedy Decoding Degradation,&amp;rdquo; &amp;ldquo;Native Context Window Boundaries,&amp;rdquo; &amp;ldquo;Synthetic Data &amp;lsquo;Sanding&amp;rsquo; Effects&amp;rdquo; (model collapse on rare cases), &amp;ldquo;Thinking Mode History Overhead&amp;rdquo;&lt;/td&gt;
&lt;td&gt;These appear less consistently across cards because they&amp;rsquo;re specific to a model family&amp;rsquo;s architecture or training method rather than universal LLM limitations, still important, but narrower in applicability.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. Performance Tradeoffs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Accuracy vs. Interpretability&lt;/td&gt;
&lt;td&gt;Explainability Cost&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;complex models are black boxes; simpler models sacrifice performance for transparency&amp;rdquo;&lt;/td&gt;
&lt;td&gt;The most commonly cited tradeoff, relevant to any regulated or high-stakes use where explainability is a requirement, not a nice-to-have.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Accuracy vs. Speed/Latency&lt;/td&gt;
&lt;td&gt;Inference Time Cost&lt;/td&gt;
&lt;td&gt;Descriptive text, sometimes with numeric latency figures&lt;/td&gt;
&lt;td&gt;Highly accurate models often cost more compute per response; production systems frequently favor a faster, slightly less accurate model.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Bias vs. Variance (Generalization)&lt;/td&gt;
&lt;td&gt;Overfitting/Underfitting Balance&lt;/td&gt;
&lt;td&gt;Descriptive text on flexible (low-bias, high-variance) vs. simple (high-bias) models&lt;/td&gt;
&lt;td&gt;Explains why a model that performs well on training data may not generalize, or why an overly simple model misses real patterns.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Complexity vs. Resource Constraints (Cost)&lt;/td&gt;
&lt;td&gt;Compute/Budget Tradeoff&lt;/td&gt;
&lt;td&gt;Descriptive text, sometimes with hardware specs (GPU/CPU requirements)&lt;/td&gt;
&lt;td&gt;Larger models need more data, training time, and compute, a direct cost and deployment-feasibility constraint.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Precision vs. Recall&lt;/td&gt;
&lt;td&gt;False Positive/Negative Balance&lt;/td&gt;
&lt;td&gt;Descriptive text, sometimes with numeric thresholds&lt;/td&gt;
&lt;td&gt;For classification tasks, states whether the model is tuned to minimize false positives or false negatives, critical for fraud, medical, or safety contexts.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(Model-specific examples)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Family-Specific Tradeoffs&lt;/td&gt;
&lt;td&gt;E.g., &amp;ldquo;Intelligence Plateau in Domain-Specific Tasks,&amp;rdquo; &amp;ldquo;Enhanced Quantization Sensitivity,&amp;rdquo; &amp;ldquo;Context Window Consistency,&amp;rdquo; &amp;ldquo;Conciseness vs. Contextual Nuance,&amp;rdquo; &amp;ldquo;Agentic Capability Limitations,&amp;rdquo; &amp;ldquo;Hardware Efficiency vs. Throughput,&amp;rdquo; &amp;ldquo;Decoding Strategy Rigidity&amp;rdquo;&lt;/td&gt;
&lt;td&gt;These are narrower, model-size or architecture-specific tradeoffs. They appear in more detailed cards and matter most when comparing versions within the same model family (e.g., 7B vs. 32B parameter variants).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4. Ethical Considerations&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Name&lt;/td&gt;
&lt;td&gt;Consideration Name/Description&lt;/td&gt;
&lt;td&gt;Short label plus expanded description, e.g., &amp;ldquo;Algorithmic and Cultural Bias,&amp;rdquo; &amp;ldquo;Vulnerability to Adversarial Attacks (Jailbreaking),&amp;rdquo; &amp;ldquo;Misinformation or Hallucinations,&amp;rdquo; &amp;ldquo;Privacy/PII Content Leakage,&amp;rdquo; &amp;ldquo;Environmental Impact (Inference Energy),&amp;rdquo; &amp;ldquo;Instruction Misalignment&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Names a specific ethical risk tied to the model, since there&amp;rsquo;s no universal standard list, well-written cards use this field to add clarifying context beyond the label itself.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Mitigation Strategy&lt;/td&gt;
&lt;td&gt;Recommended Mitigation&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;use RLAIF and rule-based rewards,&amp;rdquo; &amp;ldquo;implement an input/output safety filter,&amp;rdquo; &amp;ldquo;use RAG to ground responses,&amp;rdquo; &amp;ldquo;deploy locally with PII scrubbing,&amp;rdquo; &amp;ldquo;apply 4-bit quantization to reduce power draw,&amp;rdquo; &amp;ldquo;standardize output formats with system prompts&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Pairs each named risk with a concrete, actionable step. A risk listed without a paired mitigation should be read as an incomplete entry.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;5. Fairness Assessments&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Group At Risk&lt;/td&gt;
&lt;td&gt;At-Risk Group Identification&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;people identified by race, gender, or disability status,&amp;rdquo; &amp;ldquo;non-English/non-Spanish speakers,&amp;rdquo; &amp;ldquo;speakers of regional dialects or specific geographic regions&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Identifies the specific population the assessment is evaluating for disparate treatment. This is the field that determines whether the rest of the assessment is even relevant to your deployment population.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Harms&lt;/td&gt;
&lt;td&gt;Documented Harm&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;discriminatory outcomes in task assignment,&amp;rdquo; &amp;ldquo;quality-of-service harm: oversimplified or hallucinated answers in non-primary languages&amp;rdquo;&lt;/td&gt;
&lt;td&gt;States the specific negative outcome observed during testing, ideally with a concrete example rather than a generic statement.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Mitigation Actions&lt;/td&gt;
&lt;td&gt;Fairness Mitigation&lt;/td&gt;
&lt;td&gt;Descriptive text, e.g., &amp;ldquo;RLAIF and rule-based rewards aligned to legal standards,&amp;rdquo; &amp;ldquo;multilingual supervised fine-tuning on reasoning tasks&amp;rdquo;&lt;/td&gt;
&lt;td&gt;The corrective action recommended or applied to reduce the documented harm.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(Underlying methodology, less commonly itemized directly)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Assessment Method&lt;/td&gt;
&lt;td&gt;Data Bias Auditing, Disaggregated Performance Metrics, Impact Assessments, Adversarial Testing, Algorithmic Fairness Interventions&lt;/td&gt;
&lt;td&gt;These describe how the fairness assessment was conducted across the model lifecycle. More rigorous cards name which of these methods were used; many cards only report the outcome (&lt;code&gt;groupAtRisk&lt;/code&gt;/&lt;code&gt;harms&lt;/code&gt;/&lt;code&gt;mitigationStrategy&lt;/code&gt;) without specifying methodology, which is itself worth flagging as a gap.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;6. Environmental Considerations, Energy Consumption&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Activity&lt;/td&gt;
&lt;td&gt;Lifecycle Stage&lt;/td&gt;
&lt;td&gt;One of: design, data-collection, data-preparation, training, fine-tuning, validation, deployment, inference, other&lt;/td&gt;
&lt;td&gt;Identifies which phase of the model lifecycle the reported energy figure applies to. Training is reported most often; inference (the ongoing, per-query cost) is reported far less often despite frequently being the larger cumulative cost.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Energy Sources&lt;/td&gt;
&lt;td&gt;Energy Source Type&lt;/td&gt;
&lt;td&gt;One of: coal, oil, natural-gas, nuclear, wind, solar, geothermal, hydropower, biofuel, unknown, other&lt;/td&gt;
&lt;td&gt;States what generated the electricity used for that activity, central to any claimed environmental benefit.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Energy Description&lt;/td&gt;
&lt;td&gt;Provider Identity&lt;/td&gt;
&lt;td&gt;Organization name, address, and description, e.g., a named data center and its location&lt;/td&gt;
&lt;td&gt;Documents who supplied the energy, supporting traceability and verification of the reported figures.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Activity Energy Cost&lt;/td&gt;
&lt;td&gt;Total Energy Cost&lt;/td&gt;
&lt;td&gt;Numeric value in kilowatt-hours (kWh)&lt;/td&gt;
&lt;td&gt;The raw energy consumption figure for the activity, the base number every other environmental figure derives from.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;CO2 Cost Equivalent&lt;/td&gt;
&lt;td&gt;Carbon Cost (Debit)&lt;/td&gt;
&lt;td&gt;Numeric value in tonnes of CO2 equivalent (tCO2eq)&lt;/td&gt;
&lt;td&gt;The greenhouse gas impact of the reported energy cost, standardized so it can be compared across energy sources and activities.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;CO2 Cost Offset&lt;/td&gt;
&lt;td&gt;Carbon Offset (Credit)&lt;/td&gt;
&lt;td&gt;Numeric value in tonnes of CO2 equivalent (tCO2eq)&lt;/td&gt;
&lt;td&gt;Any offset applied against the debit above. Reported least consistently of all environmental fields, and worth checking against the debit figure rather than accepting the net claim at face value.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="the-10-implementation-tips-that-actually-decide-card-quality"&gt;The 10 Implementation Tips That Actually Decide Card Quality&lt;/h2&gt;
&lt;p&gt;Everything above is structure. This is judgment, the part that decides whether a completed card actually protects you or just looks complete.&lt;/p&gt;
&lt;p&gt;I have sat in enough of these reviews to recognize the pattern by now. Someone asks for the fairness section. Someone says it is coming in the next revision. The next revision never quite arrives, and six months later the card still says exactly what it said at launch.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Treat a missing field as a finding, not a blank.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A card with no fairness section, no adversarial testing discussion, or a suspiciously clean &amp;ldquo;no known limitations&amp;rdquo; line rarely means the system is clean. Far more often it means nobody looked, or somebody looked and did not want to write down what they found. Every review should end with an explicit list of what is absent, not only an assessment of what is present. Silence is not neutral. Silence is a finding waiting to be named.&lt;/p&gt;
&lt;ol start="2"&gt;
&lt;li&gt;Map every field to the regulatory requirement it satisfies.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A field like intended purpose does more than tidy up documentation. Under the EU AI Act, it directly satisfies Article 13(3)(b)(i). A field like disaggregated performance satisfies a separate obligation in the same article. Reviewing or producing a card without this mapping means nobody can say with confidence whether it would survive a conformity assessment. Build the mapping once, per use case category, and reuse it. Do not rebuild it from scratch every time.&lt;/p&gt;
&lt;ol start="3"&gt;
&lt;li&gt;Never accept an aggregate metric without asking for the subgroup breakdown.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is the single highest-leverage check in the entire process. A strong overall accuracy number can hide a disparity that fails badly for one specific group, language, region, or device type, and that gap only becomes visible once someone insists on the breakdown. If the card reports one number and stops there, the review is incomplete. Not finished. Incomplete.&lt;/p&gt;
&lt;ol start="4"&gt;
&lt;li&gt;Require every named risk to carry a paired, evidenced mitigation.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Name plus mitigation is the right structure. A mitigation listed without supporting evidence, a test result, a red team score, an attack success rate, is a promise dressed up as a control. A mitigation only counts once it is paired with a measurable threshold that proves it actually works.&lt;/p&gt;
&lt;ol start="5"&gt;
&lt;li&gt;Check for train test contamination before trusting any performance number.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is the most commonly skipped verification step, and one of the most consequential. If evaluation data overlaps with training data, every metric downstream of that overlap is inflated. A card that does not explicitly state the two sets are disjoint should be treated as unverified, not assumed clean. This one check protects you from building risk decisions on numbers that were never real.&lt;/p&gt;
&lt;ol start="6"&gt;
&lt;li&gt;Version the model, the prompt, the retrieval source, and the evaluation together, and retest after any one of them changes.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A model card is not a one-time artifact. Swap a model version, adjust a prompt template, or update a retrieval index, and the prior evidence stops applying even when nothing else in the card changes. A card that does not tie its results to a specific, dated version combination is documenting a system that no longer exists by the time anyone reads it.&lt;/p&gt;
&lt;ol start="7"&gt;
&lt;li&gt;Assign a named owner and a review cadence to the card itself, not only to the model.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A model card that is not refreshed on a defined schedule becomes actively misleading. A reader has no way to tell stale information from current information just by looking at it. Attach an owner. Set a quarterly review at minimum, more often for anything that moves fast. That is the difference between a static PDF and a living control, and it is the difference that actually holds up under audit.&lt;/p&gt;
&lt;ol start="8"&gt;
&lt;li&gt;Match the human oversight level to the actual stakes of the decision, not to a generic default.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A card recommending a spot check for a model that influences credit, medical, or employment decisions is a mismatch worth escalating on its own, regardless of how good the model&amp;rsquo;s other metrics look. This is one of the fastest checks in a review because it needs no technical evaluation. It only needs a comparison between the stated oversight mechanism and the real consequence of the model being wrong.&lt;/p&gt;
&lt;ol start="9"&gt;
&lt;li&gt;Remember the model is not the system. Evaluate the integration, too.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A vendor&amp;rsquo;s safety testing on a base model says very little about what happens once that model is wired into your product, with your retrieval layer, your tool access, your identities and permissions attached. The most dangerous vulnerabilities usually live in that integration layer. A card review that stops at the vendor&amp;rsquo;s own documentation and never asks what your architecture adds to the attack surface has covered half the assessment at best.&lt;/p&gt;
&lt;ol start="10"&gt;
&lt;li&gt;Prefer quantitative security and robustness metrics over narrative safety claims.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&amp;ldquo;The model has been safety tested&amp;rdquo; is not a data point. An attack success rate against a defined adversarial benchmark, a prompt injection success rate, a membership inference score, these are data points, because they are measurable, comparable across versions, and provably false if they turn out to be wrong. Cards built around reassurance instead of numbers should go back for the underlying test results before anyone relies on them for anything that matters.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/07/chatgpt-image-jul-31-2026-06_17_52-pm.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="risk-and-control-practices-for-model-cards"&gt;Risk and Control Practices for Model Cards&lt;/h2&gt;
&lt;p&gt;Four habits apply across every stage above, from the ML-BOM through the card through the ten tips. None of them are about the documents themselves. They are about what keeps the documents honest once the initial review is over.&lt;/p&gt;
&lt;p&gt;I constantly see teams treat the model card and the ML-BOM as two completely isolated chores. Don&amp;rsquo;t do this. Wire them to each other. If I read a card that references a dataset completely missing from the BOM, or a BOM that contradicts the card’s own training specs, I know instantly that you lack a single source of truth. Pick one artifact to be your system of record. Force your tooling to generate the other from it.&lt;/p&gt;
&lt;p&gt;Here is a hard reality about engineering culture. The people who built the model are the absolute worst people to document its flaws. This is just the natural byproduct of deadline pressure mixing with builder&amp;rsquo;s optimism.&lt;/p&gt;
&lt;p&gt;Hand the limitations section to someone entirely outside the build team. A fresh, slightly cynical set of eyes on that one specific section catches more actual exposure than a second pass on the entire technical file. You also need to stop leaving fields blank. If you leave a box empty, the auditor reviewing it later cannot tell if you skipped it on purpose or simply forgot it existed. Writing &amp;ldquo;Not applicable; this model has no user-facing output&amp;rdquo; is a highly defensible control. A blank space is just an unquantified liability. Document your intentional exclusions so nobody has to hunt down the original engineer a year later to figure out what happened.&lt;/p&gt;
&lt;p&gt;Finally, look at how you actually store these things. A model card passed around as a PDF attachment or a slide deck is useless. The moment it hits someone’s downloads folder, it stops being a control and turns into a rumor about what the model used to be.&lt;/p&gt;
&lt;p&gt;Store both documents as versioned, machine-readable records anchored directly to your model registry. When you can run a diff across versions to see exactly what changed between releases, you have a surviving audit trail. Anything else is just paperwork.&lt;/p&gt;
&lt;h2 id="why-model-cards-and-ai-bills-of-materials-matter-for-governance-roles"&gt;Why Model Cards and AI Bills of Materials Matter for Governance Roles&lt;/h2&gt;
&lt;p&gt;When an organization adopts AI, the model card and the AI bill of materials are the foundational documents that make the system legible to anyone who wasn&amp;rsquo;t in the room when it was built. Without them, governance roles are flying blind. Here&amp;rsquo;s why each role specifically depends on them.&lt;/p&gt;
&lt;h2 id="auditors"&gt;Auditors&lt;/h2&gt;
&lt;p&gt;Auditors need an artifact to test against. A model card gives them the declared intended use, performance metrics, training data provenance, and known limitations.Tthese are the claims they verify. If the card says the model achieves 94% accuracy on a specific benchmark, the auditor re-runs that benchmark. If the card says training data was deduplicated and PII-filtered, the auditor checks the pipeline logs.&lt;/p&gt;
&lt;p&gt;The AI BOM goes deeper: it lists every component in the supply chain, such as pre-trained base models, third-party datasets, open-source libraries, APIs, and firmware versions. This is what makes a security audit or SOC 2 examination possible. An auditor cannot assess supply-chain risk (a poisoned dependency, a license violation, a deprecated vulnerable library) without a complete inventory. Under the EU AI Act,
explicitly requires documentation of &amp;ldquo;recourse to pre-trained systems or tools provided by third parties and how those were used, integrated or modified.&amp;rdquo; The BOM is that documentation.&lt;/p&gt;
&lt;p&gt;Without these documents, an audit becomes anecdotal, spot-checking what the auditor happens to think of, rather than systematic.&lt;/p&gt;
&lt;h2 id="compliance-officers"&gt;Compliance Officers&lt;/h2&gt;
&lt;p&gt;Compliance officers map organizational practice to legal obligations. The EU AI Act&amp;rsquo;s
requires that high-risk AI systems be accompanied by instructions for deployers covering provider identity, system capabilities and limitations, accuracy metrics, human oversight measures, and data specifications. The model card is the natural container for most of that information; the BOM covers the supply-chain transparency requirements.&lt;/p&gt;
&lt;p&gt;Compliance officers also need to demonstrate that the organization performed due diligence before deployment. If a regulator asks &amp;ldquo;did you know this model was trained on data scraped without consent?&amp;rdquo; or &amp;ldquo;did you know the base model had a known prompt-injection vulnerability?&amp;rdquo;. The answer needs to be &amp;ldquo;yes, we documented it in the model card and BOM, assessed the risk, and applied mitigations.&amp;rdquo; Ignorance is not a defensible position under the AI Act&amp;rsquo;s risk-based framework (
requires a documented risk management system).&lt;/p&gt;
&lt;p&gt;The model card also supports the conformity assessment process.
requires listing harmonised standards applied and attaching the EU declaration of conformity. These reference the technical documentation, which the model card and BOM feed into.&lt;/p&gt;
&lt;h2 id="risk-managers"&gt;Risk Managers&lt;/h2&gt;
&lt;p&gt;Risk managers quantify and prioritize. They need to know what can go wrong, how likely it is, and how severe the impact would be. The model card surfaces known failure modes, fairness disparities, hallucination rates, and adversarial vulnerabilities , these are the risk inputs. The risk review columns in the checklist you just received (evaluation methods, control checks, vulnerability coverage) are essentially a risk register in spreadsheet form.&lt;/p&gt;
&lt;p&gt;The BOM adds a dimension that traditional risk management hasn&amp;rsquo;t fully grappled with: software supply-chain risk in ML systems. A model can inherit vulnerabilities from its base model (e.g., a fine-tuned model that inherits a data-poisoning susceptibility), from its training data (e.g., a dataset containing copyrighted or consent-violating material), or from its inference infrastructure (e.g., a vulnerable inference server). The BOM makes these transitive risks visible and manageable.&lt;/p&gt;
&lt;p&gt;Risk managers also need the model card&amp;rsquo;s post-market monitoring plan (
) to set up ongoing risk surveillance, drift detection, incident response, performance degradation alerts.&lt;/p&gt;
&lt;h2 id="caios-chief-ai-officers"&gt;CAIOs Chief AI Officers&lt;/h2&gt;
&lt;p&gt;CAIOs sit at the intersection of strategy, accountability, and governance. They are typically the person who signs off on AI deployment decisions and who answers to the board, regulators, and customers. They need the model card and BOM for three reasons:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Strategic visibility&lt;/strong&gt;: The CAIO needs to know what AI systems exist in the organization, what they do, what data they depend on, and what risks they carry. The model card and BOM are the inventory that enables portfolio-level decisions: which models to invest in, which to retire, which to restrict.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Accountability&lt;/strong&gt;: Under the EU AI Act, the provider (and in many cases the deployer) bears legal responsibility. If something goes wrong, such as a discriminatory outcome, a data breach, a hallucination that caused harm, the CAIO is the person who will be asked &amp;ldquo;what did you know and when did you know it?&amp;rdquo; The model card is the record of what was known at deployment time.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cross-functional alignment&lt;/strong&gt;: The CAIO orchestrates auditors, compliance, risk, engineering, and legal teams. The model card and BOM are the shared artifact that all these functions reference. Without a common document, each team maintains its own partial picture, gaps go unnoticed, and accountability diffuses.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="related-reading"&gt;Related Reading&lt;/h2&gt;
&lt;p&gt;
writes regularly on AI governance, evidence, and audit-ready documentation. A few pieces that connect directly to the ground covered here:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;
, on what actually counts as evidence once an AI system is live, not just at launch.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;
, on why controls that look complete on paper collapse the moment someone asks for proof they operate.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;
, on the policy layer that sits above the documentation covered in this piece.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;
, on why evidence collection without quantified analysis behind it stops being useful to anyone outside compliance.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="key-references"&gt;Key References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act, Article 11 and Annex IV, technical documentation requirements for high-risk AI systems, enforceable from August 2, 2026.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act, Article 13, transparency and instructions for use, the article behind the field-to-requirement mapping in tip two.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework and its Generative AI Profile, for the broader risk categories a model card should reflect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001, the AI management system standard, for how card review fits into an ongoing governance program rather than a one-time exercise.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="where-this-actually-goes-wrong-and-what-it-looks-like-done-right"&gt;Where This Actually Goes Wrong, and What It Looks Like Done Right&lt;/h2&gt;
&lt;p&gt;Treated as a compliance artifact, a model card gets written once, right before a launch or an audit, by whoever drew the short straw that week. It gets filed, forgotten, and quietly contradicted by the model within a few months, because nothing forces it to update when the model does. The first time anyone reads it again is during an incident, a regulator&amp;rsquo;s request, or a board question nobody can answer cleanly, and by then it describes a system that no longer exists. That version of a model card protects nobody. It just proves, on paper, that a document once got created.&lt;/p&gt;
&lt;p&gt;Treated as an operational tool, the same card becomes something else entirely. It is versioned alongside the model it describes. It has a named owner who knows keeping it current is their job. It gets checked at every meaningful change, not once a year. It answers a procurement team&amp;rsquo;s questions before they ask them, an auditor&amp;rsquo;s questions before they escalate, and an incident responder&amp;rsquo;s questions before the incident gets worse. It gets read constantly, by people who trust it, because it has earned that trust field by field.&lt;/p&gt;
&lt;p&gt;A model card is either a record of what someone once claimed, or it is a record of what you can actually prove. Only one of those survives contact with a regulator.&lt;/p&gt;
&lt;p&gt;By Prof. Hernan Huwyler, CAIO MBA CPA&lt;br&gt;
&lt;br&gt;
&lt;br&gt;
&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and advisory work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative
predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and internationally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance, technical and business requirements.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item><item><title>AI Model Cards That Improves Transparency, Governance, and Real-World Use</title><link>https://hwyler.github.io/blog/ai-model-cards-that-improves-transparency-governance-and-real-world-use/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/ai-model-cards-that-improves-transparency-governance-and-real-world-use/</guid><description>&lt;h2 id="why-model-cards-matter-more-now-than-ever"&gt;Why Model Cards Matter More Now Than Ever&lt;/h2&gt;
&lt;p&gt;Most AI model cards fail for one reason.&lt;/p&gt;
&lt;p&gt;They are written after the fact, for compliance theater, by people who are too far from the model’s actual design and operation. The result is familiar. A neat summary of the model type, a few metrics, vague notes on limitations, and almost nothing that helps product teams, auditors, operators, or governance leads understand how the model should and should not be used. The document exists. The value does not.&lt;/p&gt;
&lt;p&gt;A strong AI model card is different. It is a working record of what the model is, what it was built to do, what data shaped it, where it performs well, where it struggles, what risks matter, and how it should be monitored in production. This post shows you how to build model cards that support transparency, accountability, and practical use, including when to use system cards for multiple interacting models.&lt;/p&gt;
&lt;p&gt;Suggested visual: A model card layout showing sections for basics, intended use, data and evaluation, risks, trustworthiness, and monitoring.&lt;/p&gt;
&lt;h2 id="understanding-the-core-framework-for-ai-model-cards"&gt;Understanding the Core Framework for AI Model Cards&lt;/h2&gt;
&lt;p&gt;An AI model card is a structured document that describes an AI model in a way that helps technical and non-technical stakeholders understand its purpose, training basis, performance, limitations, and operational requirements.&lt;/p&gt;
&lt;p&gt;That sounds simple. In practice, model cards often become either too technical to be useful or too shallow to be trustworthy.&lt;/p&gt;
&lt;p&gt;The framework I use has four core purposes. Transparency, decision support, accountability, and operational continuity. If your model card does not support these four things, it is probably just another document in a repository.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/futuristic-robotics-lab.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h3 id="1-transparency"&gt;1. Transparency&lt;/h3&gt;
&lt;p&gt;The model card should make the model understandable at the right level for the audience. It should clearly state what the model is intended to do, what it was trained on, what it was evaluated against, and where it may fail.&lt;/p&gt;
&lt;p&gt;Transparency matters because AI systems are often used by people who did not build them. Product managers, compliance teams, security reviewers, customer-facing operators, and auditors all need enough visibility to make sound decisions.&lt;/p&gt;
&lt;p&gt;Implementation tip: Write the model card so a smart non-specialist in your company can understand the model’s role, limits, and risks without reading code.&lt;/p&gt;
&lt;h3 id="2-decision-support"&gt;2. Decision support&lt;/h3&gt;
&lt;p&gt;A good model card helps people decide whether the model is suitable for a given use case, population, environment, or workflow. It should not only describe the model. It should support judgment.&lt;/p&gt;
&lt;p&gt;This means documenting intended use, out-of-scope use, performance tradeoffs, fairness patterns, and integration assumptions. These details help teams decide when the model is fit for purpose and when it is not.&lt;/p&gt;
&lt;p&gt;Implementation tip: Include a short section called “Use this model when…” and another called “Do not use this model when…” Those two fields improve practical judgment fast.&lt;/p&gt;
&lt;h3 id="3-accountability"&gt;3. Accountability&lt;/h3&gt;
&lt;p&gt;Model cards help create accountability by recording decisions, assumptions, versioning, evidence, and known limitations. They also support audits and governance reviews.&lt;/p&gt;
&lt;p&gt;Without that record, teams rely too heavily on memory and informal handoffs. That becomes risky when models are updated, integrated into broader systems, or reviewed months later by people who were not there at the start.&lt;/p&gt;
&lt;p&gt;Implementation tip: Treat the model card as evidence, not marketing. If the tone feels like product positioning, the document is probably too soft.&lt;/p&gt;
&lt;h3 id="4-operational-continuity"&gt;4. Operational continuity&lt;/h3&gt;
&lt;p&gt;Model cards are not only useful before launch. They help after deployment too. They give support teams, operators, and new team members a way to understand what the model is supposed to do, how it should be monitored, and what known weaknesses require attention.&lt;/p&gt;
&lt;p&gt;This is especially important when teams change, vendors are involved, or multiple models are connected in one workflow.&lt;/p&gt;
&lt;p&gt;Implementation tip: Keep the model card in the same operating ecosystem as version logs, deployment records, and monitoring references. Documentation that lives far from operations gets ignored.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/coding-in-the-dark.png?w=775" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="why-ai-model-cards-often-fall-short"&gt;Why AI Model Cards Often Fall Short&lt;/h2&gt;
&lt;p&gt;The most common issue is incompleteness.&lt;/p&gt;
&lt;p&gt;Teams document the model architecture and a few metrics, then skip training data limitations, demographic performance patterns, deployment assumptions, monitoring plans, and operator guidance. That creates a document that looks respectable and helps almost nobody.&lt;/p&gt;
&lt;p&gt;Another issue is staleness. A model card may describe version 1.2 while production is already on version 1.5 with new prompts, new tuning, new training data, or a new serving setup. Once that happens, trust in the document drops.&lt;/p&gt;
&lt;p&gt;There is also a structural issue. Some organizations write model cards for individual models but never produce a system card for the combined AI system. In real deployments, several models often work together. If only the parts are documented and not the whole, the most important interaction risks stay hidden.&lt;/p&gt;
&lt;p&gt;Implementation tip: If more than one model materially influences the output, produce both model cards and a system card. The interaction layer matters.&lt;/p&gt;
&lt;h2 id="stage-1-define-the-purpose-scope-and-audience-of-the-model-card"&gt;Stage 1: Define the Purpose, Scope, and Audience of the Model Card&lt;/h2&gt;
&lt;p&gt;Before filling in fields, decide what the model card is meant to support and who needs to use it.&lt;/p&gt;
&lt;p&gt;The responsible parties are the model owner, data scientists, AI engineers, product owner, and AI governance lead. Legal, privacy, compliance, and security should review where the use case is high impact or regulated.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the model card template, audience definition, governance requirements, and version control approach. These should shape the depth and style of the final document.&lt;/p&gt;
&lt;p&gt;What to implement: Define the model’s intended purpose, deployment context, and key stakeholders. Be explicit about whether the card is meant for internal developers only, for cross-functional governance, for customers, or for auditors. In most organizations, one detailed internal version and one simplified external-facing variant work better than trying to force one document to satisfy every audience.&lt;/p&gt;
&lt;p&gt;This is also where you should decide whether the model card covers one model or whether you also need a system card describing several models working together. If a ranking model, retrieval system, classifier, and large language model all interact in one user experience, the single-model view is incomplete.&lt;/p&gt;
&lt;p&gt;Implementation tip: Put the audience and intended use of the model card at the top of the document. That helps reviewers understand the level of detail and the purpose of the content.&lt;/p&gt;
&lt;h2 id="stage-2-document-the-model-basics-clearly"&gt;Stage 2: Document the Model Basics Clearly&lt;/h2&gt;
&lt;p&gt;This section sounds administrative. It is more important than teams expect.&lt;/p&gt;
&lt;p&gt;The responsible parties are the model owner, data scientists, engineering leads, and documentation owner. Product and governance should review for consistency with system and registry records.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the model registry entry, release notes, source references, license records, and technical glossary. These support traceability.&lt;/p&gt;
&lt;p&gt;What to implement: Include the model type using a standardized taxonomy and a brief plain-language description. Record licenses, citations, intellectual property references, date, version, and release history. Add a glossary for technical terms and a reference list covering tools, methods, and sources that shaped the model.&lt;/p&gt;
&lt;p&gt;This section should answer a few basic but important questions. What model is this. What kind of model is it. Where did it come from. What version is under discussion. What prior work or external assets shaped it.&lt;/p&gt;
&lt;p&gt;Versioning is critical. If the model card is not tied to a specific version and release date, it will become unreliable quickly.&lt;/p&gt;
&lt;p&gt;Implementation tip: Use the same model identifier across the model card, model registry, deployment records, and monitoring dashboard. Inconsistent naming creates support and audit problems.&lt;/p&gt;
&lt;h2 id="stage-3-define-intended-uses-and-legal-or-contextual-boundaries"&gt;Stage 3: Define Intended Uses and Legal or Contextual Boundaries&lt;/h2&gt;
&lt;p&gt;This is one of the most valuable parts of the model card because it helps prevent misuse.&lt;/p&gt;
&lt;p&gt;The responsible parties are the product owner, model owner, legal, compliance, and AI governance lead. Domain experts should review because intended use often depends on business context.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the use case definition, approved deployment scope, policy restrictions, and legal review notes.&lt;/p&gt;
&lt;p&gt;What to implement: Describe the purpose and scope of the model in practical language. Explain what the model is intended to do, who is expected to use it, in what context, and under what constraints. Then document the key legal, compliance, and contextual considerations relevant to deployment.&lt;/p&gt;
&lt;p&gt;This section should also identify unsupported or inappropriate uses. If a model works well for English-language support summarization but not for legal advice or multilingual risk scoring, say so directly. If human review is required, say that too.&lt;/p&gt;
&lt;p&gt;Teams often avoid writing strong boundaries because they fear limiting adoption. The opposite is usually true. Clear boundaries improve trust and reduce misuse.&lt;/p&gt;
&lt;p&gt;Implementation tip: Put explicit use restrictions in the same section as intended use. Splitting them into a hidden appendix makes them easier to ignore.&lt;/p&gt;
&lt;h2 id="stage-4-explain-training-data-evaluation-and-performance-honestly"&gt;Stage 4: Explain Training Data, Evaluation, and Performance Honestly&lt;/h2&gt;
&lt;p&gt;This is the section most people look for first, and it needs to be more than a metric dump.&lt;/p&gt;
&lt;p&gt;The responsible parties are data scientists, data engineers, AI engineers, and model owners. Governance and domain experts should review for clarity and practical usefulness.&lt;/p&gt;
&lt;p&gt;The critical artifacts are dataset documentation, preprocessing notes, evaluation reports, test data records, metric definitions, and performance summaries by relevant subgroup or scenario.&lt;/p&gt;
&lt;p&gt;What to implement: Describe the data sources used for training, validation, and testing. Include the type of data, source, time period, preprocessing steps, and important inclusion or exclusion choices. Then explain how the model was tested and validated, which metrics were used, and why those metrics were appropriate for the use case.&lt;/p&gt;
&lt;p&gt;This section also needs performance limitations. Under what conditions does the model perform less well. Are there failure patterns by language, geography, document type, user behavior, or demographic group. If there are tradeoffs between metrics, explain them clearly.&lt;/p&gt;
&lt;p&gt;Metric choice deserves justification too. Accuracy may matter in one case. Recall or false negative rate may matter more in another. The card should explain the reasoning, not simply list values.&lt;/p&gt;
&lt;p&gt;Implementation tip: Show performance in slices that matter for use, not only in overall averages. Overall performance often hides the conditions where users will struggle most.&lt;/p&gt;
&lt;h2 id="stage-5-document-risks-ethics-and-system-integration"&gt;Stage 5: Document Risks, Ethics, and System Integration&lt;/h2&gt;
&lt;p&gt;Strong model cards acknowledge that technical performance is only part of the story.&lt;/p&gt;
&lt;p&gt;The responsible parties are model owners, product, AI governance, legal, compliance, security, and domain experts. Risk or ethics review functions may also need to contribute.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the risk assessment, known failure scenarios, mitigation plan, system architecture, and integration documentation.&lt;/p&gt;
&lt;p&gt;What to implement: Describe potential risk scenarios associated with the model’s deployment and operation. Include likely misuse, harmful failure patterns, and mitigation strategies. Address ethical concerns that are relevant to the model’s use, especially where outputs can affect fairness, safety, privacy, dignity, or access to opportunity.&lt;/p&gt;
&lt;p&gt;Also describe how the model interfaces with other systems and the broader IT environment. This matters because many model risks emerge from integration, not just from the model itself. A classifier feeding a workflow engine, or a ranking model feeding a human review queue, can create downstream effects that matter operationally and ethically.&lt;/p&gt;
&lt;p&gt;Implementation tip: Write at least three realistic failure scenarios in plain language. Technical readers and non-technical readers both benefit from concrete examples.&lt;/p&gt;
&lt;h2 id="stage-6-cover-trustworthy-ai-factors-in-a-way-that-is-usable"&gt;Stage 6: Cover Trustworthy AI Factors in a Way That Is Usable&lt;/h2&gt;
&lt;p&gt;The “trustworthy” section should not become a generic paragraph about principles. It needs operational substance.&lt;/p&gt;
&lt;p&gt;The responsible parties are data scientists, AI engineers, governance, security, privacy, and product. Domain experts should review fairness and explainability claims to make sure they are meaningful in context.&lt;/p&gt;
&lt;p&gt;The critical artifacts are fairness analyses, explainability methods, robustness tests, and security review findings.&lt;/p&gt;
&lt;p&gt;What to implement: Describe fairness by showing how the model performs across relevant human groups or operational segments, and what mitigation steps were taken where bias or imbalance appeared. Describe explainability methods such as feature importance, confidence indicators, or decision logic support used to help users interpret outputs. Describe security and resilience by summarizing robustness against adversarial attacks, model misuse, or system vulnerabilities.&lt;/p&gt;
&lt;p&gt;This section should also stay realistic. If the model has limited explainability, say so. If fairness testing could not be done fully because protected-group data was unavailable, say what was done instead and what limitations remain.&lt;/p&gt;
&lt;p&gt;Implementation tip: Avoid claiming that a model is “fair” or “explainable” without context. Describe the actual tests, methods, and limits instead.&lt;/p&gt;
&lt;h2 id="stage-7-add-monitoring-updates-and-operator-guidance"&gt;Stage 7: Add Monitoring, Updates, and Operator Guidance&lt;/h2&gt;
&lt;p&gt;A model card should help after deployment, not only before it.&lt;/p&gt;
&lt;p&gt;The responsible parties are product, engineering, MLOps, operations, support teams, and governance. Training or enablement teams may also contribute if the system requires formal operator guidance.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the monitoring plan, update process, retraining criteria, user guidance, operator playbooks, and support materials.&lt;/p&gt;
&lt;p&gt;What to implement: Describe how the model will be updated in response to new data, vulnerabilities, or changing conditions. Explain how performance and fairness will be monitored after deployment. Provide guidance and training resources for users and operators so they understand what the model does, how to use it appropriately, and when to escalate issues.&lt;/p&gt;
&lt;p&gt;This section matters because a model card that ends at launch is only half useful. Teams need to know what to watch, what changes trigger reassessment, and how to handle model behavior in practice.&lt;/p&gt;
&lt;p&gt;Implementation tip: Include links to the live monitoring dashboard and incident workflow where possible. A model card should connect people to action, not just description.&lt;/p&gt;
&lt;h2 id="system-cards-when-one-model-card-is-not-enough"&gt;System Cards: When One Model Card Is Not Enough&lt;/h2&gt;
&lt;p&gt;Many AI systems use several models together. A recommender feeds a ranking model. A retrieval system supplies a language model. A moderation classifier filters outputs. A detection model triggers a workflow engine.&lt;/p&gt;
&lt;p&gt;In these cases, system cards are essential. A system card aggregates the relevant model information and explains how the components interact, where responsibility sits, and what combined risks matter.&lt;/p&gt;
&lt;p&gt;The responsible parties are the product owner, lead architect, AI governance, and owners of the underlying models. Security, legal, and operations should review where the integrated behavior creates new risk.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the architecture diagram, component model cards, system risk assessment, and workflow documentation.&lt;/p&gt;
&lt;p&gt;What to implement: Use system cards to describe the broader AI system, not just its component models. Explain the flow of data, decision points, human review steps, and combined behavior. Highlight risks that emerge only when the models interact.&lt;/p&gt;
&lt;p&gt;Implementation tip: If a user experiences one product but the documentation is split across five isolated model cards, you probably also need a system card.&lt;/p&gt;
&lt;h2 id="best-practices-for-ai-model-cards"&gt;Best Practices for AI Model Cards&lt;/h2&gt;
&lt;p&gt;These tips apply across all stages.&lt;/p&gt;
&lt;h3 id="tip-1-keep-model-cards-concise-but-evidence-linked"&gt;Tip 1: Keep model cards concise but evidence-linked&lt;/h3&gt;
&lt;p&gt;A model card should be readable. It should also point to deeper evidence where needed.&lt;/p&gt;
&lt;p&gt;Implementation tip: Keep the main card concise and link out to evaluation reports, fairness analyses, data documentation, and monitoring plans. This balances readability and depth.&lt;/p&gt;
&lt;h3 id="tip-2-update-model-cards-as-part-of-release-management"&gt;Tip 2: Update model cards as part of release management&lt;/h3&gt;
&lt;p&gt;Stale model cards quickly lose value.&lt;/p&gt;
&lt;p&gt;Implementation tip: Make model card review a required step for material model updates, new data sources, new deployment contexts, or major performance changes.&lt;/p&gt;
&lt;h3 id="tip-3-use-model-cards-in-real-governance-workflows"&gt;Tip 3: Use model cards in real governance workflows&lt;/h3&gt;
&lt;p&gt;Model cards should not live only in a documentation repository.&lt;/p&gt;
&lt;p&gt;Implementation tip: Require model cards in approval reviews, audits, risk assessments, and post-launch evaluations. Documents gain quality when people actually use them.&lt;/p&gt;
&lt;h3 id="tip-4-write-for-multiple-readers-without-losing-precision"&gt;Tip 4: Write for multiple readers without losing precision&lt;/h3&gt;
&lt;p&gt;Different stakeholders need different levels of detail, but they all need accuracy.&lt;/p&gt;
&lt;p&gt;Implementation tip: Use plain-language summaries at the top of each section, followed by more technical detail where needed. That structure works well across mixed audiences.&lt;/p&gt;
&lt;h2 id="references-for-ai-model-cards"&gt;References for AI Model Cards&lt;/h2&gt;
&lt;p&gt;If you want a stronger model card practice, anchor it in recognized AI governance and transparency standards.&lt;/p&gt;
&lt;p&gt;Here are the references I would use.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001, AI management systems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42005, information to include in an AI impact assessment&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 23894, AI risk management&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework 1.0&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Existing model card and system card research from leading academic and industry sources&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Internal documentation, model governance, and audit standards&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Security, privacy, fairness, and transparency requirements relevant to the deployment context&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If your organization already uses model registries, risk reviews, and architecture records, connect model cards into those systems. That makes them easier to maintain and more likely to be used.&lt;/p&gt;
&lt;h2 id="why-model-cards-fail-when-treated-as-static-documentation"&gt;Why Model Cards Fail When Treated as Static Documentation&lt;/h2&gt;
&lt;p&gt;When teams treat model cards as static documentation, they produce a neat artifact once, store it, and move on. The model changes. The deployment context changes. The risks change. The card does not. Over time it becomes less trusted, less used, and less worth maintaining.&lt;/p&gt;
&lt;p&gt;When teams treat model cards as living operational records, the document improves transparency, sharpens governance, supports audits, guides users, and helps teams manage change responsibly. That is when model cards become genuinely valuable.&lt;/p&gt;
&lt;p&gt;A strong AI model card works because it helps people understand not only what the model is, but how it should be used, watched, and questioned over time.&lt;/p&gt;
&lt;p&gt;If you reviewed your current AI documentation today, which gap would likely show up first: unclear intended use, weak data disclosure, thin risk documentation, weak fairness evidence, or stale update and monitoring guidance?&lt;/p&gt;</description></item></channel></rss>