<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ai-Impact-Assessment |</title><link>https://hwyler.github.io/tags/ai-impact-assessment/</link><atom:link href="https://hwyler.github.io/tags/ai-impact-assessment/index.xml" rel="self" type="application/rss+xml"/><description>Ai-Impact-Assessment</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 12 Mar 2026 00:00:00 +0000</lastBuildDate><image><url>https://hwyler.github.io/media/icon_hu_cd51c91342a84ed6.png</url><title>Ai-Impact-Assessment</title><link>https://hwyler.github.io/tags/ai-impact-assessment/</link></image><item><title>Implementation Tips for ISO 42005 AI Impact Assessments</title><link>https://hwyler.github.io/blog/implementation-tips-for-iso-42005-ai-impact-assessments/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/implementation-tips-for-iso-42005-ai-impact-assessments/</guid><description>&lt;h2 id="why-the-iso-42005-ai-impact-assessment-structure-matters"&gt;Why the ISO 42005 AI Impact Assessment Structure Matters&lt;/h2&gt;
&lt;p&gt;Most AI impact assessments fail before the first risk is even discussed.&lt;/p&gt;
&lt;p&gt;They fail in the form itself. Teams rush through fields, paste in vendor language, skip foreseeable misuse, and treat ISO 42005 as a documentation exercise instead of a decision tool. Then the assessment gets approved with gaps large enough to drive a regulatory inquiry through. I have seen this happen in hiring, fraud, customer service, and internal productivity tools. The pattern is always the same. The template exists, but nobody has turned it into an operational workflow.&lt;/p&gt;
&lt;p&gt;That is why this post matters. If you want an AI impact assessment that actually helps governance, you need more than a list of ISO 42005 fields. You need a working method for what to write, who owns each section, what evidence should sit behind it, and where common failure points show up. This guide gives you that method.&lt;/p&gt;
&lt;p&gt;Suggested visual: A one-page lifecycle view showing ISO 42005 fields mapped to intake, design review, testing, approval, deployment, and monitoring.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/modern-industrial-engineers-at-work.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="understanding-the-core-concept-for-iso-42005-ai-impact-assessment-fields"&gt;Understanding the Core Concept for ISO 42005 AI Impact Assessment Fields&lt;/h2&gt;
&lt;p&gt;ISO 42005 gives structure to an AI impact assessment. That structure is useful because AI projects drift fast. Functionality changes. Users change. Risk changes. Jurisdictions change. If the assessment does not capture those moving parts clearly, governance loses the thread.&lt;/p&gt;
&lt;p&gt;Here is the mental model I use. Every good ISO 42005 AI impact assessment should answer four questions.&lt;/p&gt;
&lt;p&gt;What is the system?&lt;/p&gt;
&lt;p&gt;Why does it exist?&lt;/p&gt;
&lt;p&gt;Who can it affect?&lt;/p&gt;
&lt;p&gt;What evidence shows the risks were taken seriously?&lt;/p&gt;
&lt;p&gt;Those four questions map directly to the field groups in the standard. General information tells you what document you are looking at and whether it is current. System description and purpose explain the tool and the claimed value. Data, model, deployment, and parties sections reveal who and what are in scope. Benefits, harms, failures, and misuse force teams to confront consequences.&lt;/p&gt;
&lt;p&gt;Most organizations struggle because they fill out fields one by one without connecting them. That creates contradictions. The “basic description” says the model offers recommendations only, while the intended use says it can auto-route claims, and the harms section forgets due process entirely. I have reviewed assessments where three different teams described the same AI system in three different ways. Nobody noticed until the approval meeting.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Start every ISO 42005 AI impact assessment with a 30-minute alignment session across product, engineering, legal, privacy, and the business owner. Put the core use case on one page before anyone touches the template. This cuts inconsistency fast.&lt;/p&gt;
&lt;h3 id="the-five-field-groups-that-matter-most"&gt;The five field groups that matter most&lt;/h3&gt;
&lt;p&gt;You should complete every section. Still, five groups carry most of the practical weight.&lt;/p&gt;
&lt;h3 id="1-identity-and-governance-fields"&gt;1. Identity and governance fields&lt;/h3&gt;
&lt;p&gt;These include AI system name or ID, lifecycle stage, revision history, reviewer, and approver fields. They sound administrative. They are not.&lt;/p&gt;
&lt;p&gt;These fields tell you whether the document is current, whether the system being assessed is the actual system going live, and whether the right people stood behind the review. In one client review, the version approved by governance was two model versions behind the one engineering deployed. The mismatch only surfaced because the revision dates were inconsistent.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Tie the AI system ID in the assessment to the product registry, model registry, and procurement record. If those IDs do not match, stop the review until they do.&lt;/p&gt;
&lt;h3 id="2-scope-and-use-fields"&gt;2. Scope and use fields&lt;/h3&gt;
&lt;p&gt;These include the system description, functionalities, purpose, intended uses, unintended uses, and dependencies. This is where teams often understate what the system does.&lt;/p&gt;
&lt;p&gt;A chatbot may summarize, infer sentiment, draft responses, detect abuse patterns, and pass outputs into another workflow. A hiring tool may rank candidates, reject applicants, generate recruiter notes, and capture video data. If only one of those functions is named, the assessment underestimates impact.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Require every functionality field to begin with an action verb such as classify, predict, rank, generate, summarize, identify, or recommend. Vague descriptions hide risk.&lt;/p&gt;
&lt;h3 id="3-data-and-model-evidence-fields"&gt;3. Data and model evidence fields&lt;/h3&gt;
&lt;p&gt;These cover datasets, data quality, algorithm suitability, model evaluation, drift, retraining, and bias or harms testing. This is where technical evidence enters the impact assessment.&lt;/p&gt;
&lt;p&gt;Weak assessments use placeholders here. Strong ones provide actual data lineage, performance metrics, subgroup testing, and retraining criteria tied to operating conditions. If you do not know what data shaped the model or how well it performs on the populations you will affect, the rest of the assessment is guesswork.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Add a rule that no field in this section can be answered with “standard process followed.” Ask for specifics, dates, metrics, and sign-off sources.&lt;/p&gt;
&lt;h3 id="4-deployment-and-affected-party-fields"&gt;4. Deployment and affected-party fields&lt;/h3&gt;
&lt;p&gt;These include geography, legal requirements, culture, at-risk groups, languages, deployment constraints, and relevant interested parties. This is the section that grounds the system in the real world.&lt;/p&gt;
&lt;p&gt;I once reviewed a language model deployment where the product team had tested English well and Spanish moderately, but the planned deployment included Arabic support by default in the interface settings. Nobody had validated it. The deployment field forced the issue. That one line probably prevented a bad launch.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Treat every new geography, language, and user group as a change in risk, not a scaling detail. Reopen the impact assessment when any of those variables expands.&lt;/p&gt;
&lt;h3 id="5-benefits-harms-failures-and-misuse-fields"&gt;5. Benefits, harms, failures, and misuse fields&lt;/h3&gt;
&lt;p&gt;These are the fields teams fear because they force honesty. Good. That is their job.&lt;/p&gt;
&lt;p&gt;If your AI system could expose personal data, reinforce discrimination, suppress lawful speech, create unsafe recommendations, or be repurposed for surveillance or fraud, say so clearly. A useful AI impact assessment is not a sales deck.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Ask teams to write one foreseeable harm that would make the project sponsor uncomfortable. If every harm sounds minor and generic, the assessment is not mature enough.&lt;/p&gt;
&lt;h2 id="stage-1-complete-the-general-information-fields-like-they-matter-because-they-do"&gt;Stage 1: Complete the General Information Fields Like They Matter, Because They Do&lt;/h2&gt;
&lt;p&gt;The first section of ISO 42005 is usually treated as setup. That is a mistake.&lt;/p&gt;
&lt;p&gt;The responsible parties here are the business owner, product manager, governance team, and document owner. The accountable person should be the system owner, not a rotating project coordinator who cannot answer questions later.&lt;/p&gt;
&lt;p&gt;The key artifacts are the AI system registry entry, lifecycle record, approval workflow, and document control log. These should all connect to the impact assessment fields for name, ID, lifecycle stage, revision history, review, and approval.&lt;/p&gt;
&lt;p&gt;What to implement: For AI System Name or ID, use the same identifier that appears in procurement, architecture, model ops, and incident management records. For AI System Life Cycle Stage, use a controlled list such as concept, design, development, validation, pilot, production, material change, retirement. For review and approval fields, record named roles and dates, not generic team labels alone.&lt;/p&gt;
&lt;p&gt;This is where many governance programs quietly break. A draft assessment gets copied from an earlier version. Dates remain old. Reviewer names remain wrong. The document looks complete, but nobody can prove who assessed the live version.&lt;/p&gt;
&lt;p&gt;I made this mistake early in my consulting work. We had a clean-looking impact assessment packet for a vendor tool. During a later incident review, we discovered the “approved” file belonged to the pilot, not the scaled deployment with new features. Same product family. Different risk. We had to reconstruct the review trail by hand. It took days.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Add one field internally that ISO 42005 does not spell out but every program needs, “Material change since last assessment.” If the answer is yes, force a short summary of what changed and whether prior approvals still apply.&lt;/p&gt;
&lt;h2 id="stage-2-write-a-system-description-that-exposes-real-scope"&gt;Stage 2: Write a System Description That Exposes Real Scope&lt;/h2&gt;
&lt;p&gt;The AI system description, functionalities, purpose, intended uses, unintended uses, and dependencies form the backbone of the assessment. If this section is weak, every later section becomes distorted.&lt;/p&gt;
&lt;p&gt;Responsible parties include product, engineering, enterprise architecture, procurement for vendor tools, and governance. Legal and privacy should review wording for scope and consequence, but product and engineering must own the factual details.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the architecture diagram, user flow, API map, vendor documentation, and intended use statement. These artifacts should support every field in this section. If the description says the system does not make decisions, the user flow should not show auto-rejection or auto-escalation without human review.&lt;/p&gt;
&lt;p&gt;What to implement: The Basic AI System Description should answer five plain questions. What input goes in. What output comes out. Who uses it. What decisions it influences. What other systems it sends information to. For functionalities, separate current features from planned ones and include estimated dates only when there is actual roadmap evidence.&lt;/p&gt;
&lt;p&gt;For intended uses, describe the end user, setting, and boundaries. “Customer support summarization for trained internal agents in English-language email workflows” is strong. “Support automation” is weak. For unintended uses, list both malicious misuse and predictable overreach. A sentiment model used for employee wellness may later be repurposed for performance management. That risk belongs in the form.&lt;/p&gt;
&lt;p&gt;Dependencies matter more than teams expect. If your AI output triggers another model, a business rule engine, a human review queue, or an external API, say so. Dependencies create hidden failure chains.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Add one internal control question under dependencies, “If this dependent system fails, what does the AI system do next?” Quiet fallback logic causes real harm. A ranking tool that defaults to a raw score when an explanation service fails can confuse reviewers and distort outcomes.&lt;/p&gt;
&lt;h2 id="stage-3-treat-data-information-and-quality-as-an-evidence-section-not-a-narrative-section"&gt;Stage 3: Treat Data Information and Quality as an Evidence Section, Not a Narrative Section&lt;/h2&gt;
&lt;p&gt;This section is where ISO 42005 gets serious. Dataset names, ownership, access rights, provenance, bias risks, quality processes, DPIA need, and data quality characteristics all belong here.&lt;/p&gt;
&lt;p&gt;The responsible parties are data engineering, data governance, privacy, security, machine learning teams, and the business owner. If a vendor provides the model or training data, procurement and vendor risk teams should support the response.&lt;/p&gt;
&lt;p&gt;The critical artifacts are data inventories, lineage records, access control logs, data use approvals, privacy assessments, quality reports, and retention schedules. A mature program can point to each one within minutes.&lt;/p&gt;
&lt;p&gt;What to implement: For each dataset, document the owner, version, size, collection period, geography, whether data is real or synthetic, who collected it, under what authority, and whether its use for AI has been approved. Then document known bias risks and the exact quality checks performed. If a DPIA is required, mark it and link the reference.&lt;/p&gt;
&lt;p&gt;For data quality characteristics met, name the characteristic and explain why it matters to the system. Completeness, representativeness, timeliness, label reliability, and class balance are common examples. For planned characteristics, do not write aspirations like “improve diversity.” Write the specific gap, why it matters, and the date by which the gap will be addressed.&lt;/p&gt;
&lt;p&gt;I have seen teams write “dataset is representative” with no evidence. Then you look closely and find the data over-indexes one region, one user segment, or one language. The assessment should force teams to confront those limits, not glide past them.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Make teams state one thing the dataset is bad at. This sounds small, but it changes the tone of the whole assessment. Honest limitations produce better controls than polished claims.&lt;/p&gt;
&lt;h2 id="stage-4-use-the-algorithms-and-models-section-to-show-decision-quality-evidence"&gt;Stage 4: Use the Algorithms and Models Section to Show Decision-Quality Evidence&lt;/h2&gt;
&lt;p&gt;This is the most technical part of the ISO 42005 AI impact assessment. It is also where non-technical reviewers often get lost.&lt;/p&gt;
&lt;p&gt;The solution is simple. Write technical truth in plain language.&lt;/p&gt;
&lt;p&gt;Responsible parties here are data science, machine learning engineering, model risk, security, privacy engineering, and domain experts. Governance should review for completeness and clarity, not rewrite the science.&lt;/p&gt;
&lt;p&gt;The critical artifacts are experiment logs, validation reports, model cards, bias assessments, robustness tests, red team outputs, retraining standards, and compute or environmental records. If these artifacts do not exist, the fields will become vague. That is the signal to stop and fix the process.&lt;/p&gt;
&lt;p&gt;What to implement: For algorithm suitability, explain why the chosen method fits the business task and the decision stakes. For validity and real-world performance, include prior deployments, known limitations, and evidence from published research or internal testing. For susceptibility to undesirable outcomes, name issues such as overfitting, spurious correlations, instability, proxy discrimination, hallucination, or prompt injection risk.&lt;/p&gt;
&lt;p&gt;For model fields, document training, validation, and testing data. Explain how you kept datasets disjoint. Describe feature selection criteria. List performance metrics with thresholds tied to use case risk. Include generalization testing on production-like data. Add bias and harm evaluations, PII leakage checks, robustness measures, drift detection methods, retraining triggers, and impacts from continuous learning if used.&lt;/p&gt;
&lt;p&gt;One practical point. Do not flood the form with every metric the team has. Pick the metrics that matter for the use case. For a classifier, that may be false positives and false negatives by subgroup. For a recommender, ranking quality and harmful amplification indicators may matter more. For generative AI, factuality, refusal consistency, privacy leakage, and unsafe output rates may be central.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Require every model section to include one sentence beginning with “This model should not be used when…” That sentence often reveals more practical governance value than two pages of metrics.&lt;/p&gt;
&lt;h2 id="stage-5-ground-the-assessment-in-deployment-reality-and-affected-people"&gt;Stage 5: Ground the Assessment in Deployment Reality and Affected People&lt;/h2&gt;
&lt;p&gt;A model can perform well in testing and still fail in deployment because the geography, language, legal setting, or user population changes.&lt;/p&gt;
&lt;p&gt;This section includes current and planned deployment areas, geo-specific legal requirements, cultural considerations, marginalized groups, languages, human traits relevant to the system, deployment method, and deployment constraints. It also includes internal and external interested parties.&lt;/p&gt;
&lt;p&gt;Responsible parties include product, legal, privacy, public policy, regional operations, accessibility specialists, and frontline operational leaders. If the tool affects workers, patients, students, claimants, or citizens, the relevant operational function needs to be in the room.&lt;/p&gt;
&lt;p&gt;What to implement: For geo areas, do not list countries only. List states, provinces, or cities when local law matters. For legal requirements, include labor law, data protection rules, sector rules, biometrics restrictions, consumer protection, and language access obligations where relevant. For marginalized groups, name the groups likely to be affected in that deployment context and explain why.&lt;/p&gt;
&lt;p&gt;For interested parties, separate those who use the system from those subject to its outputs. A customer service agent using an AI assistant is not the same as the customer whose case is summarized and routed. An HR recruiter using a ranking tool is not the same as the applicant filtered by it.&lt;/p&gt;
&lt;p&gt;I once worked on a case where the internal party list was detailed and the external party list was almost blank. That told us everything we needed to know about the maturity of the review. The team had thought about internal workflow efficiency and barely considered the people outside the company who would bear the impact.&lt;/p&gt;
&lt;p&gt;Original implementation tip: If you cannot identify at least one external party who could be harmed, the assessment is probably too shallow. Nearly every deployed AI system affects someone beyond the immediate operator.&lt;/p&gt;
&lt;h2 id="stage-6-write-benefits-harms-failures-and-misuse-with-operational-honesty"&gt;Stage 6: Write Benefits, Harms, Failures, and Misuse with Operational Honesty&lt;/h2&gt;
&lt;p&gt;This section is where the ISO 42005 AI impact assessment stops being descriptive and becomes evaluative.&lt;/p&gt;
&lt;p&gt;The fields cover accountability, transparency, fairness and discrimination, privacy, reliability, safety, explainability, and environmental impact. Then they move into failures and misuse. This is where the assessment should show that the team has looked past the happy path.&lt;/p&gt;
&lt;p&gt;Responsible parties include governance, legal, privacy, security, product, trust and safety, domain experts, and the business owner. If the use case is high impact, escalation to a risk committee makes sense.&lt;/p&gt;
&lt;p&gt;What to implement: For each benefit field, describe a realistic gain tied to actual operations. For each harm field, describe a reasonably foreseeable downside with enough specificity to inform controls. Then document at least two failures and two misuses with impacts on interested parties.&lt;/p&gt;
&lt;p&gt;A good example. For fairness and discrimination harms in a hiring tool, write that historical training data may reduce interview rates for women returning from caregiving gaps or for disabled applicants whose career patterns differ from prior hires. For misuse, write that recruiters may use the ranking score as a rejection tool despite policy saying it is advisory. That is a foreseeable misuse because people under time pressure take shortcuts.&lt;/p&gt;
&lt;p&gt;This section should connect directly to approval conditions. If you identify a privacy harm, where is the retention control. If you identify explainability harm, where is the user notice or appeal workflow. If you identify misuse risk, where is the training or restriction.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Ask the frontline operators what misuse they fear. They usually know before governance does. The people who work the queue see where the shortcuts, workarounds, and pressure points really are.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/watermark-free-gemini_generated_image_1tsv5t1tsv5t1tsv.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="tips-for-iso-42005-ai-impact-assessment"&gt;Tips for ISO 42005 AI Impact Assessment&lt;/h2&gt;
&lt;p&gt;These tips apply across the whole assessment. They keep the form useful over time.&lt;/p&gt;
&lt;h3 id="tip-1-do-not-let-one-team-write-the-whole-assessment-alone"&gt;Tip 1: Do not let one team write the whole assessment alone&lt;/h3&gt;
&lt;p&gt;Single-author assessments look neat and miss reality. Product sees value. Engineering sees architecture. Legal sees obligations. Operations sees failure conditions.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Assign section ownership by expertise, then run one editor across the final document for consistency. Shared drafting with single-point editing works well.&lt;/p&gt;
&lt;h3 id="tip-2-use-evidence-links-not-long-pasted-explanations"&gt;Tip 2: Use evidence links, not long pasted explanations&lt;/h3&gt;
&lt;p&gt;Teams often turn impact assessments into bulky documents full of copied text. That slows review and hides gaps.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Keep field answers concise and link to source artifacts such as DPIAs, test reports, architecture diagrams, or validation files. Short answers with evidence age better than long prose.&lt;/p&gt;
&lt;h3 id="tip-3-reopen-the-assessment-at-known-trigger-points"&gt;Tip 3: Reopen the assessment at known trigger points&lt;/h3&gt;
&lt;p&gt;An AI impact assessment is not a one-time event. It should reopen when the system changes in material ways.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Set mandatory reassessment triggers for new data sources, new model versions, new geographies, new user groups, new decision rights, major incidents, or a shift from advisory use to automated action.&lt;/p&gt;
&lt;h3 id="tip-4-separate-unknown-from-not-applicable"&gt;Tip 4: Separate “unknown” from “not applicable”&lt;/h3&gt;
&lt;p&gt;These are not the same thing. One means you have a gap. The other means the field genuinely does not apply.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Ban blank fields. Use a controlled response set such as completed, not applicable, unknown pending evidence. Unknown items should feed a tracked action list before approval.&lt;/p&gt;
&lt;h2 id="references-for-building-an-iso-42005-ai-impact-assessment-process"&gt;References for Building an ISO 42005 AI Impact Assessment Process&lt;/h2&gt;
&lt;p&gt;If you want your ISO 42005 AI impact assessment process to stand up in practice, build it against well-known standards and governance sources.&lt;/p&gt;
&lt;p&gt;Here are the references I would use.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42005, information to include in an AI system impact assessment&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001, AI management systems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 23894, AI risk management&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 22989, AI concepts and terminology&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 23053, framework for AI systems using machine learning&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 27701, privacy information management&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework 1.0&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;OECD AI Principles&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;UNESCO Recommendation on the Ethics of Artificial Intelligence&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;GDPR and Data Protection Impact Assessment guidance&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Sector-specific guidance for health, employment, financial services, public sector decision-making, and consumer protection&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If your organization already uses model risk, privacy impact, or security review processes, map ISO 42005 fields into those workflows instead of creating a totally separate bureaucracy. That saves time and improves consistency.&lt;/p&gt;
&lt;h2 id="why-iso-42005-becomes-useless-when-treated-as-a-form-filling-exercise"&gt;Why ISO 42005 Becomes Useless When Treated as a Form-Filling Exercise&lt;/h2&gt;
&lt;p&gt;When teams treat ISO 42005 as paperwork, the AI impact assessment becomes a polished archive of half-truths. Current and planned uses blur together. Data quality gets overstated. Bias risks are softened. Misuse is ignored because it feels uncomfortable. Reviewers sign off on a document that looks complete while the actual system keeps changing underneath it.&lt;/p&gt;
&lt;p&gt;When teams use ISO 42005 properly, the assessment becomes a living operating record. It tells you what the system does today, what it may do next, who can be affected, what evidence supports trust, where the risk sits, and what conditions must hold before launch or expansion. That changes governance from reactive to usable.&lt;/p&gt;
&lt;p&gt;ISO 42005 works when each field forces a real answer, backed by evidence, owned by the right people, and revisited when the system changes.&lt;/p&gt;
&lt;p&gt;If you reviewed one of your current AI impact assessments today, which section would show the biggest gap first: system scope, data quality, model evidence, deployment context, or foreseeable misuse?&lt;/p&gt;</description></item><item><title>Practical AI Assessments</title><link>https://hwyler.github.io/blog/practical-ai-assessments/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/practical-ai-assessments/</guid><description>&lt;h2 id="the-9-stage-ai-assessment-framework-that-answers-three-questions-every-project-must-face"&gt;The 9-Stage AI Assessment Framework That Answers Three Questions Every Project Must Face&lt;/h2&gt;
&lt;p&gt;Every AI project, regardless of industry, budget, or technology, must answer three questions at the right time. Can we build this? Are we ready to deploy it? Did it actually succeed?&lt;/p&gt;
&lt;p&gt;Most organizations answer the first question with enthusiasm, rush past the second, and never systematically address the third. The result is predictable. Projects that were technically feasible but operationally unready get pushed into production. Systems that are deployed never get measured against the business case that justified them. And organizations accumulate AI systems they can&amp;rsquo;t confidently say are delivering value.&lt;/p&gt;
&lt;p&gt;A structured AI assessment framework creates defined evaluation gates across the full project lifecycle. Nine assessments, grouped into three phases, ensure that every critical question gets asked at the point where the answer can still influence decisions. Skip an assessment and you&amp;rsquo;re making downstream commitments based on untested assumptions. Complete each one rigorously and you build a chain of evidence that supports every decision from concept through sustained operation.&lt;/p&gt;
&lt;p&gt;This post walks through all nine assessments, explains what each one evaluates, and provides the practical guidance that determines whether these assessments produce real decisions or decorative documentation.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/computer-cooling-system.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="phase-1-can-we-build-this"&gt;Phase 1: Can We Build This?&lt;/h2&gt;
&lt;p&gt;The first three assessments determine whether an AI project should proceed into development. They evaluate the problem, the technical viability, and the data foundation. Getting clear answers at this stage is the cheapest form of risk management available. Stopping a non-viable project during Phase 1 costs days of analysis time. Stopping it during development costs months of engineering effort.&lt;/p&gt;
&lt;p&gt;These three assessments are interdependent. A strong use case with poor data readiness shouldn&amp;rsquo;t proceed. Strong data readiness without a clear use case produces a solution looking for a problem. Strong feasibility without either produces a technology demonstration with no business value.&lt;/p&gt;
&lt;p&gt;All three assessments should be completed within the same evaluation window, typically two to four weeks, and reviewed together in a single go/no-go decision meeting.&lt;/p&gt;
&lt;p&gt;Implementation tip: Assign ownership of each Phase 1 assessment to a different team member or function. The use case assessment should be owned by a business stakeholder who understands the problem. The feasibility analysis should be owned by a technical lead who can evaluate architecture, skills, and infrastructure honestly. The data readiness assessment should be owned by a data engineer who can verify data quality empirically, not theoretically. When one person or team owns all three, assessments tend to confirm the conclusion the owner has already reached. When different people with different perspectives own different assessments, the combined evaluation produces a more honest picture of project viability.&lt;/p&gt;
&lt;h2 id="assessment-1-use-case-assessment"&gt;Assessment 1: Use Case Assessment&lt;/h2&gt;
&lt;p&gt;The use case assessment identifies the specific business problem the AI system will solve and quantifies the value of solving it. This is where most AI projects either build a strong foundation or begin accumulating the vague objectives that eventually undermine them.&lt;/p&gt;
&lt;p&gt;Three activities define a thorough use case assessment.&lt;/p&gt;
&lt;p&gt;Describe how users interact with AI to achieve a specific goal. This goes beyond describing what the AI system does. It describes the human workflow that the AI system fits into: who triggers the AI, what input they provide, what output they receive, what they do with that output, and how the AI-assisted workflow differs from the current process. A use case that describes only the AI component without describing the human workflow will produce a system that works in isolation and fails in practice.&lt;/p&gt;
&lt;p&gt;Map current processes to quantify inefficiencies, expected value, and improvement areas. Before you can measure improvement, you need a documented baseline of how the process works today. Map each step in the current process, measure the time each step takes, identify where errors occur most frequently, and calculate the cost of the current approach. This map becomes the reference point against which all future performance measurements are compared.&lt;/p&gt;
&lt;p&gt;Define measurable success metrics and align AI goals with user needs. Every use case should specify what success looks like in numbers: processing time targets, accuracy thresholds, cost reduction goals, and user satisfaction benchmarks. These metrics should reflect what users actually need, not what the technology can most easily deliver. A system that achieves 98% accuracy on a metric users don&amp;rsquo;t care about while achieving 70% accuracy on the metric they depend on has failed its use case regardless of the headline number.&lt;/p&gt;
&lt;p&gt;What to document: The use case assessment should produce a single document containing: the problem statement, the current process map with baseline measurements, the proposed AI-assisted process, identified user roles and their interactions with the system, success metrics with numerical targets, and a preliminary estimate of business value.&lt;/p&gt;
&lt;p&gt;Implementation tip: The process mapping step reveals hidden complexity that interviews and requirements documents miss. Documented processes and actual processes frequently diverge. Employees develop workarounds, skip steps that seem unnecessary, and add informal quality checks that aren&amp;rsquo;t in any procedure manual. Map the actual process by observing it, not by reading the documentation. The discrepancies between documented and actual processes often identify the real bottlenecks and the real opportunities for AI assistance, which may differ substantially from what the initial problem statement assumed.&lt;/p&gt;
&lt;p&gt;The use case assessment identifies the specific business problem AI can solve. This is where the project gets anchored in an actual business need.&lt;/p&gt;
&lt;p&gt;The responsible parties are the business owner, process owner, product lead, and AI governance or transformation lead. Legal, privacy, security, and compliance should be consulted where the use case touches regulated data or sensitive decisions.&lt;/p&gt;
&lt;p&gt;The critical artifacts are the use case statement, user interaction description, current process map, pain point summary, value hypothesis, and success metrics. These should be clear enough that a reviewer can understand the business problem without needing a demo.&lt;/p&gt;
&lt;p&gt;What to implement: Describe how users will interact with the AI system to achieve a specific goal. Map the current process to quantify inefficiencies, delays, rework, or quality problems. Identify the expected value and improvement areas. Define measurable success metrics and align the AI goal with user needs, not just management enthusiasm.&lt;/p&gt;
&lt;p&gt;This assessment should answer a basic question. Is this a real business problem with a plausible AI role, or just a technology idea looking for a use case?&lt;/p&gt;
&lt;p&gt;Implementation tip: Require one “current state” metric and one “target state” metric in the use case review. If there is no measurable gap, the value case is too weak.&lt;/p&gt;
&lt;h2 id="assessment-2-feasibility-analysis"&gt;Assessment 2: Feasibility Analysis&lt;/h2&gt;
&lt;p&gt;The feasibility analysis assesses whether the proposed AI solution can be built, deployed, and maintained within the organization&amp;rsquo;s technical, financial, and regulatory constraints. A viable use case that isn&amp;rsquo;t feasible should be shelved until constraints change, not forced into development.&lt;/p&gt;
&lt;p&gt;Four evaluation areas define feasibility.&lt;/p&gt;
&lt;p&gt;Evaluate tech stack, data, and skill availability. Does your current infrastructure support the proposed AI system&amp;rsquo;s compute, storage, and networking requirements? Does the team possess demonstrated experience with the required model architectures, development frameworks, and deployment patterns? Are gaps addressable within the project timeline through hiring, training, or partnerships? Honest answers to these questions prevent the common pattern of approving projects that require capabilities the organization doesn&amp;rsquo;t have and can&amp;rsquo;t acquire fast enough.&lt;/p&gt;
&lt;p&gt;Estimate business ROI and strategic fit. Calculate projected return on investment using conservative assumptions. Include all costs: development, infrastructure, data preparation, training, deployment, and ongoing maintenance and monitoring. Compare projected value against projected cost over a 3-year horizon. Separately assess strategic fit: Does this project align with organizational AI strategy? Does it build capabilities that support future AI initiatives? Strategic value can justify projects with marginal ROI, but that tradeoff should be made explicitly, not by default.&lt;/p&gt;
&lt;p&gt;Check regulatory and market readiness. Identify every regulation that applies to the proposed AI system in every geography where it will operate. Evaluate whether the system can meet compliance requirements. Assess whether the market context, including customer expectations, competitive dynamics, and industry norms, supports the proposed AI application. A technically feasible system that violates regulatory requirements isn&amp;rsquo;t feasible regardless of its other merits.&lt;/p&gt;
&lt;p&gt;Gauge time and budget constraints. Compare the estimated development timeline against business deadlines. If the business need expires before the AI system can be deployed, the project isn&amp;rsquo;t feasible in its current form. Consider whether a reduced-scope version could deliver partial value within the available timeline.&lt;/p&gt;
&lt;p&gt;Implementation tip: The most common feasibility analysis failure is evaluating each dimension independently and missing interactions between them. A project might be technically feasible (right skills, right infrastructure), financially feasible (positive ROI), and regulatorily feasible (compliant design) but still infeasible because the combination of regulatory compliance requirements and technical architecture decisions drives the cost above the ROI threshold. Evaluate feasibility dimensions in combination, not in isolation. Build a single feasibility summary that shows how constraints in one dimension affect assessments in others. This integrated view catches projects that pass each individual test but fail the combined evaluation.&lt;/p&gt;
&lt;h2 id="assessment-3-data-readiness"&gt;Assessment 3: Data Readiness&lt;/h2&gt;
&lt;p&gt;The data readiness assessment determines whether the data required for the AI system exists, is accessible, is of sufficient quality, and can be used within governance and privacy requirements. Data readiness issues are the most common cause of AI project delays and failures, and the most frequently underassessed.&lt;/p&gt;
&lt;p&gt;Four evaluation areas define data readiness.&lt;/p&gt;
&lt;p&gt;Confirm data volume and quality. Does enough data exist to train the proposed model effectively? Is the data accurate, complete, and representative of the scenarios the AI system will encounter in production? Quality assessment should include specific measurements: missing value rates, error rates verified against ground truth samples, consistency of formats across records, and demographic or segment representation compared to target population distributions.&lt;/p&gt;
&lt;p&gt;Check accessibility and format fit. Can the required data be accessed by the development team within security and governance requirements? Is the data in formats that the proposed model architecture can consume, or does significant transformation work stand between raw data and usable training sets? Data that exists but isn&amp;rsquo;t accessible, or that&amp;rsquo;s accessible but requires months of reformatting, changes the project timeline and cost significantly.&lt;/p&gt;
&lt;p&gt;Assess labeling effort required. If the proposed approach uses supervised learning, does labeled data exist? If not, how much labeling effort is required, who will do it, and how long will it take? Data labeling is one of the most underestimated costs in AI project planning. A model that requires 50,000 labeled examples, at an average labeling rate of 200 examples per day per labeler, needs approximately 250 person-days of labeling effort before model training can begin.&lt;/p&gt;
&lt;p&gt;Ensure data governance, provenance, and privacy compliance. Document the origin of each data source. Verify that the data can legally be used for the proposed purpose. Confirm that privacy requirements, including consent, anonymization, retention limits, and data subject rights, can be met. Identify whether a data protection impact assessment is required and, if so, complete it before development begins.&lt;/p&gt;
&lt;p&gt;Implementation tip: Data readiness assessments that rely solely on metadata and documentation consistently overestimate readiness. The data catalog says the dataset contains 500,000 records. The actual dataset contains 500,000 rows, of which 80,000 are duplicates, 35,000 have critical fields missing, and 12,000 contain values outside valid ranges. After deduplication and quality filtering, the usable dataset is 373,000 records, which may or may not be sufficient. Always run a quantitative data profile as part of the readiness assessment: record counts after deduplication, null rates per field, value distribution analysis, and sample-based accuracy verification against source systems. The gap between documented data quality and measured data quality is almost always larger than expected.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/watermark-free-gemini_generated_image_fn28r6fn28r6fn28-1.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="phase-2-are-we-ready-to-deploy"&gt;Phase 2: Are We Ready to Deploy?&lt;/h2&gt;
&lt;p&gt;The next three assessments determine whether the AI system is ready for real-world operation. They progress from controlled validation (proof of concept) through limited real-world testing (pilot program) to full deployment preparation (production readiness).&lt;/p&gt;
&lt;p&gt;Phase 2 assessments are inherently iterative. A proof of concept that reveals technical limitations feeds back into design changes. A pilot that surfaces user experience issues feeds back into interface refinement. A production readiness assessment that identifies security gaps feeds back into hardening work. This feedback is the point. Phase 2 exists to find problems while they&amp;rsquo;re still cheap to fix.&lt;/p&gt;
&lt;p&gt;The transition between Phase 1 and Phase 2 should be a formal gate. Only projects that pass all three Phase 1 assessments should enter Phase 2. Projects that pass Phase 1 with conditions (such as &amp;ldquo;proceed if data labeling is completed by date X&amp;rdquo;) should have those conditions tracked and verified.&lt;/p&gt;
&lt;h2 id="assessment-4-proof-of-concept"&gt;Assessment 4: Proof of Concept&lt;/h2&gt;
&lt;p&gt;The proof of concept tests core AI functionality with a minimal viable prototype using actual company data. It validates that the proposed approach works in practice, not just in theory.&lt;/p&gt;
&lt;p&gt;Two evaluation priorities define the proof of concept.&lt;/p&gt;
&lt;p&gt;Validate algorithm performance against defined metrics and baseline requirements. Using the success metrics defined in the use case assessment and the baseline measurements captured during process mapping, test whether the AI system meets, approaches, or falls short of targets. This validation must use actual company data, not public datasets or synthetic examples. Performance on generic data tells you whether the algorithm works in general. Performance on your data tells you whether it works for your problem.&lt;/p&gt;
&lt;p&gt;Identify technical limitations and data quality issues before major investment. The proof of concept is designed to surface problems early. Does the model struggle with certain input categories? Does data quality degrade for specific subsets? Are inference times acceptable under realistic conditions? Are there edge cases that produce clearly wrong outputs? Document every limitation discovered. Each one represents a decision: fix it before proceeding, accept it as a known limitation, or determine that it disqualifies the approach entirely.&lt;/p&gt;
&lt;p&gt;The proof of concept should be time-boxed. Two to four weeks is typical. The goal is to gather enough evidence to make a confident proceed/pivot/stop decision, not to build a polished system. Feature completeness is not the objective. Evidence-based confidence in the approach is the objective.&lt;/p&gt;
&lt;p&gt;Implementation tip: Define proof of concept success criteria before building the prototype, and make those criteria the basis for the proceed decision. Without predefined criteria, proof of concept evaluations become subjective. The data science team sees promising results and wants to continue. The business stakeholder sees limitations and has concerns. Without agreed-upon criteria, the discussion becomes a negotiation rather than an evidence-based evaluation. Specify: &amp;ldquo;The proof of concept succeeds if the model achieves at least 80% of the target accuracy metric on a representative sample of production data, with inference times below 2x the production latency requirement.&amp;rdquo; Clear criteria produce clear decisions.&lt;/p&gt;
&lt;h2 id="assessment-5-pilot-program"&gt;Assessment 5: Pilot Program&lt;/h2&gt;
&lt;p&gt;The pilot program deploys the AI solution with a limited user group to gather real-world performance data. It bridges the gap between controlled testing and full production by exposing the system to actual users, actual workflows, and actual operational conditions.&lt;/p&gt;
&lt;p&gt;Four evaluation areas define the pilot.&lt;/p&gt;
&lt;p&gt;Gather user feedback and pain points. The pilot is the first time real users interact with the system in their actual work context. Their feedback reveals usability issues, trust barriers, workflow friction, and output quality concerns that no amount of internal testing can replicate. Collect feedback through structured channels: in-application feedback mechanisms, weekly survey check-ins, and direct observation sessions where a team member watches users interact with the system.&lt;/p&gt;
&lt;p&gt;Assess operational integration ease. Does the AI system fit into existing workflows without creating disruption? Do users need to switch between multiple applications? Does the system&amp;rsquo;s output arrive at the right point in the process and in a format users can act on? Integration friction that seems minor in a demo becomes a major adoption barrier in daily use.&lt;/p&gt;
&lt;p&gt;Refine the implementation approach based on pilot findings. The pilot exists to generate the evidence needed to improve the system before full deployment. Plan for at least one refinement cycle between pilot completion and production rollout. Address the most common user complaints, fix the most impactful technical issues, and adjust the workflow integration based on observed usage patterns.&lt;/p&gt;
&lt;p&gt;Estimate preliminary ROI and impact. Using pilot data, project the business impact of full deployment. If 20 pilot users processed 500 cases with 82% automation rate and 3.2x speed improvement, extrapolate what full deployment across 200 users would deliver. Compare this projection against the ROI estimate from the feasibility analysis. If the pilot suggests significantly lower returns than projected, reassess before committing to full deployment.&lt;/p&gt;
&lt;p&gt;Implementation tip: Select pilot users deliberately, not randomly. Include enthusiastic early adopters (who will push the system&amp;rsquo;s capabilities and provide detailed feedback), skeptical experienced users (who will identify where the AI falls short of expert human judgment), and typical average users (who represent how the majority will interact with the system). A pilot group composed entirely of enthusiasts will produce optimistic results that don&amp;rsquo;t generalize. A pilot group composed entirely of skeptics will produce pessimistic results that discourage investment. A balanced group produces realistic data that supports honest deployment decisions.&lt;/p&gt;
&lt;h2 id="assessment-6-production-readiness"&gt;Assessment 6: Production Readiness&lt;/h2&gt;
&lt;p&gt;The production readiness assessment validates that the AI system meets all technical, operational, and governance requirements for full deployment. This assessment should confirm that every requirement identified during feasibility analysis has been met or explicitly accepted as a known limitation with documented mitigation.&lt;/p&gt;
&lt;p&gt;Four evaluation areas define production readiness.&lt;/p&gt;
&lt;p&gt;Validate algorithm performance against defined metrics and baseline requirements at production scale. Proof of concept and pilot performance may not extrapolate to production volumes. Test the system at projected production load with realistic data volumes and concurrent user counts. Verify that performance metrics hold under stress conditions, not just average conditions.&lt;/p&gt;
&lt;p&gt;Confirm that monitoring, alerting, and incident response mechanisms are operational. Before the system goes live, verify that production monitoring dashboards are functioning, automated alerts are configured for key performance thresholds, the incident response team knows their roles and procedures, and escalation paths are documented and tested.&lt;/p&gt;
&lt;p&gt;Verify compliance and governance readiness. Confirm that all regulatory requirements identified during feasibility analysis have been addressed. Verify that required documentation, including model cards, impact assessments, and data processing records, is complete and current. Confirm that access controls, audit logging, and data handling procedures meet security and privacy standards.&lt;/p&gt;
&lt;p&gt;Confirm operational support readiness. Verify that the support team knows how to triage AI-specific issues. Confirm that retraining procedures are documented and the team knows when and how to execute them. Verify that the rollback procedure, the process for reverting to the previous system if the AI deployment fails, has been tested and works.&lt;/p&gt;
&lt;p&gt;Implementation tip: Run the production readiness assessment as a formal checklist review with sign-off from every responsible function: engineering, operations, security, compliance, and the business owner. Each function signs off on the criteria within their domain. The system enters production only when all functions have signed. This process prevents the common pattern where one function, usually engineering, declares the system &amp;ldquo;ready&amp;rdquo; based on technical criteria while operational, security, or compliance readiness gaps remain unaddressed. The sign-off requirement forces every function to evaluate readiness through their own lens and take accountability for their determination.&lt;/p&gt;
&lt;h2 id="phase-3-did-we-succeed"&gt;Phase 3: Did We Succeed?&lt;/h2&gt;
&lt;p&gt;The final three assessments evaluate whether the AI system delivers the value it promised. These assessments occur before launch (final validation), shortly after launch (performance review), and on an ongoing basis (value tracking).&lt;/p&gt;
&lt;p&gt;Phase 3 is where most AI assessment frameworks end too early or never begin. Organizations invest heavily in determining whether they can build a system and whether they&amp;rsquo;re ready to deploy it, then stop measuring once it&amp;rsquo;s live. This creates a gap where systems operate without evidence of value, consuming resources indefinitely because nobody has the data to justify either continued investment or shutdown.&lt;/p&gt;
&lt;p&gt;The transition from Phase 2 to Phase 3 should be seamless. Production readiness completion should automatically trigger the pre-launch validation timeline. Post-launch review should be scheduled before launch occurs. Value tracking cadence should be defined in the project plan, not established retroactively.&lt;/p&gt;
&lt;h2 id="assessment-7-pre-launch-validation"&gt;Assessment 7: Pre-Launch Validation&lt;/h2&gt;
&lt;p&gt;Pre-launch validation ensures the AI system is technically robust, secure, and optimized before going live in production. This assessment occurs after production readiness approval and before the system is made available to all users.&lt;/p&gt;
&lt;p&gt;Four validation activities define this assessment.&lt;/p&gt;
&lt;p&gt;Stress-test for scalability and speed. Push the system beyond projected peak loads to identify breaking points. If normal production load is 1,000 predictions per hour, test at 3,000 and 5,000 predictions per hour. Determine where performance degrades, where errors begin, and where the system fails entirely. This information enables capacity planning and defines operational boundaries.&lt;/p&gt;
&lt;p&gt;Validate security and privacy controls. Conduct security testing specific to the AI system: test API endpoints for input validation and authentication, verify that model artifacts and training data are protected against unauthorized access, test for AI-specific vulnerabilities including prompt injection and data leakage, and confirm that privacy controls including data anonymization, consent verification, and retention enforcement function correctly.&lt;/p&gt;
&lt;p&gt;Run performance and load tests. Beyond stress testing, conduct sustained performance testing that simulates realistic production usage patterns over extended periods, typically 24 to 72 hours. This testing reveals issues that short-duration tests miss: memory leaks that accumulate over hours, gradual performance degradation under sustained load, and resource contention with other systems sharing infrastructure.&lt;/p&gt;
&lt;p&gt;Fix all bugs identified during validation before production launch. Every defect discovered during pre-launch validation must be classified, prioritized, and resolved or explicitly accepted before the system goes live. Critical and major bugs must be fixed. Minor bugs may be accepted with documented justification and a scheduled fix date. Do not launch with known critical defects.&lt;/p&gt;
&lt;p&gt;Implementation tip: Pre-launch validation should include a &amp;ldquo;chaos test&amp;rdquo; that simulates the failure of key dependencies. What happens when the database connection drops? What happens when the model serving endpoint becomes unavailable? What happens when input data arrives in an unexpected format? Systems that handle dependency failures gracefully, by queuing requests, falling back to default behaviors, or alerting operators, are production-ready. Systems that crash or produce silently wrong outputs when a dependency fails are not. These failure scenarios are inevitable in production. Testing for them before launch ensures the system responds safely when they occur rather than creating incidents.&lt;/p&gt;
&lt;h2 id="assessment-8-post-launch-review"&gt;Assessment 8: Post-Launch Review&lt;/h2&gt;
&lt;p&gt;The post-launch review measures actual business outcomes against initial projections and success criteria. This assessment should occur at defined intervals after launch: 30 days, 90 days, and 6 months are typical checkpoints.&lt;/p&gt;
&lt;p&gt;Three evaluation areas define the post-launch review.&lt;/p&gt;
&lt;p&gt;Assess user adoption and satisfaction. Measure what percentage of target users are actively using the system, how frequently they use it, and how satisfied they are with its outputs. Compare adoption rates against the targets set during use case definition. If adoption is below target, investigate whether the gap is caused by usability issues, trust concerns, training gaps, or workflow friction. Low adoption negates all other performance metrics because a system nobody uses delivers no value regardless of its technical capabilities.&lt;/p&gt;
&lt;p&gt;Monitor system performance, data drift, and model accuracy over time. Production performance should be measured against the same metrics used during proof of concept, pilot, and pre-launch validation. Track these metrics continuously, not just at review checkpoints. Watch for data drift, where the statistical properties of production data diverge from training data, causing model accuracy to degrade gradually. Establish automated alerts for accuracy drops, latency increases, and anomalous output distributions.&lt;/p&gt;
&lt;p&gt;Identify operational lessons learned to refine future AI strategies. Every deployment teaches lessons that improve subsequent projects. Document what worked well, what didn&amp;rsquo;t work as expected, what risks materialized that weren&amp;rsquo;t anticipated, and what controls proved effective or ineffective. These lessons should be captured formally and shared with teams planning future AI initiatives.&lt;/p&gt;
&lt;p&gt;Implementation tip: Schedule the 30-day post-launch review before the system launches, with a specific date, attendee list, and agenda template already established. Post-launch reviews that aren&amp;rsquo;t pre-scheduled get postponed indefinitely because the team moves on to the next project. The 30-day review is the most critical because it catches early problems while they&amp;rsquo;re still small and while the deployment team still has the context to diagnose them. By the 90-day review, team members may have rotated to other assignments and institutional memory about deployment decisions starts fading. The 30-day review window is the highest-leverage moment for identifying and correcting post-deployment issues.&lt;/p&gt;
&lt;h2 id="assessment-9-value-tracking"&gt;Assessment 9: Value Tracking&lt;/h2&gt;
&lt;p&gt;Value tracking evaluates whether the AI system delivers the promised business value over time. This is an ongoing assessment, not a one-time review. It answers the question that ultimately determines the system&amp;rsquo;s fate: is this worth what we&amp;rsquo;re paying for it?&lt;/p&gt;
&lt;p&gt;Three evaluation areas define value tracking.&lt;/p&gt;
&lt;p&gt;Calculate true ROI and cost benefits. Compare actual costs (infrastructure, maintenance, support, model retraining, monitoring) against actual benefits (time saved, errors prevented, revenue generated, cost avoided). Use the same methodology that was used to project ROI during the feasibility analysis, applied to actual data rather than estimates. This comparison reveals whether the business case has held up, exceeded expectations, or fallen short.&lt;/p&gt;
&lt;p&gt;Identify optimization opportunities. Production operation reveals inefficiencies and improvement opportunities that weren&amp;rsquo;t visible during development. Perhaps the model could be retrained on recent data to improve accuracy. Perhaps certain features could be simplified to reduce compute costs. Perhaps the system could be extended to adjacent use cases that share the same data and infrastructure. Value tracking should identify these opportunities and prioritize them based on expected incremental value.&lt;/p&gt;
&lt;p&gt;Assess long-term business impact. Beyond direct ROI, evaluate the system&amp;rsquo;s broader effects on the organization. Has it changed how teams make decisions? Has it created new capabilities that enable other initiatives? Has it affected employee satisfaction or customer perception? These broader impacts are harder to quantify but often represent more durable value than direct cost savings.&lt;/p&gt;
&lt;p&gt;Implementation tip: Value tracking should include a &amp;ldquo;continuation decision&amp;rdquo; at regular intervals, typically annually. At each interval, explicitly decide whether the system should continue operating, be enhanced, be maintained without further investment, or be retired. This decision requires comparing the ongoing cost of operation against the ongoing value delivered. Without a formal continuation decision, AI systems persist indefinitely by institutional inertia, consuming infrastructure costs, maintenance effort, and monitoring attention long after their value has diminished. The continuation decision forces the organization to treat every AI system as an investment that must justify its ongoing costs, not as a permanent fixture that operates until something breaks.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/retro-ai-televisions.png?w=713" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="implementation-of-ai-assessments"&gt;Implementation of AI Assessments&lt;/h2&gt;
&lt;p&gt;These principles apply across all nine assessments and all three phases.&lt;/p&gt;
&lt;p&gt;Implementation tip on assessment documentation standards: Use a consistent template across all nine assessments. Each assessment document should include: assessment name, date, assessor, the system being assessed, the criteria evaluated, the findings for each criterion, the overall determination (pass/conditional pass/fail), any conditions or actions required, and the date of the next scheduled assessment. Consistent formatting enables comparison across assessments and across projects. When your tenth AI project uses the same assessment templates as your first, organizational learning compounds because patterns become visible across projects. Teams spot recurring failure modes, common data readiness issues, and consistent integration challenges that project-specific documentation would never reveal.&lt;/p&gt;
&lt;p&gt;Implementation tip on assessment independence: The person or team conducting an assessment should not be the same person or team whose work is being assessed. Data scientists should not assess their own model&amp;rsquo;s production readiness. Project managers should not assess their own project&amp;rsquo;s feasibility. Business owners should not assess their own use case&amp;rsquo;s viability without external challenge. This principle creates tension that many organizations find uncomfortable. But self-assessment consistently produces optimistic evaluations because the assessor has a personal interest in the outcome. Independent assessment, whether from a dedicated governance function, a peer team, or an external party, produces more honest evaluations and catches issues that self-assessment misses.&lt;/p&gt;
&lt;p&gt;Implementation tip on connecting assessments across phases: Each assessment should explicitly reference findings from previous assessments. The pilot program assessment should reference proof of concept findings and document whether identified limitations were addressed. The post-launch review should reference production readiness findings and verify that accepted risks are being monitored. The value tracking assessment should reference the ROI projections from the feasibility analysis and document variance. This cross-referencing creates a continuous evidence chain that supports governance, demonstrates due diligence, and prevents the common pattern where each assessment exists as an isolated document disconnected from the assessments before and after it.&lt;/p&gt;
&lt;p&gt;Implementation tip on assessment cadence after Phase 3: Once all nine assessments are complete for a given AI system, the assessment cycle doesn&amp;rsquo;t end. Post-launch reviews should recur quarterly for the first year and semi-annually thereafter. Value tracking should recur annually at minimum. Any significant system change, such as model retraining, scope expansion, infrastructure migration, or regulatory change, should trigger reassessment of production readiness. Define this ongoing cadence in your AI governance framework so that it applies automatically to every deployed system rather than depending on individual project teams to remember.&lt;/p&gt;
&lt;h2 id="authoritative-frameworks"&gt;Authoritative Frameworks&lt;/h2&gt;
&lt;p&gt;Your AI assessment framework should align with these established standards:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001:2023, AI Management System (planning, evaluation, and improvement requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 5338, AI System Life Cycle Processes (stage-gate processes across AI development)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42005, AI Impact Assessment (assessment methodology and documentation)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 23894:2023, AI Risk Management (risk assessment across lifecycle stages)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework, Map, Measure, and Manage functions&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act, Articles 9-15 for high-risk AI system assessment requirements&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 25010, Systems and Software Quality Requirements (quality criteria for system evaluation)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;IEEE 2801-2022, Recommended Practice for Quality Management of Datasets (data readiness criteria)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 27001:2022, Information Security Management (security assessment requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;PMBOK Guide stage-gate methodology adapted for AI project governance&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you conduct AI assessments as paperwork exercises, filling in templates to satisfy governance requirements without allowing findings to influence decisions, you will approve projects that should have been stopped, deploy systems that aren&amp;rsquo;t ready, and operate AI that may or may not be delivering value. Each unchecked assumption compounds risk. Each skipped assessment creates a blind spot. The assessments will exist in your document management system. The problems they should have caught will exist in your production environment.&lt;/p&gt;
&lt;p&gt;When you treat each assessment as a genuine decision point, where findings lead to actions, where criteria determine outcomes, and where the answer &amp;ldquo;no, not yet&amp;rdquo; is valued as much as &amp;ldquo;yes, proceed,&amp;rdquo; you create a governance framework that protects both the organization and the people affected by its AI systems. The nine assessments answer three simple questions. Can we build this? Are we ready? Did it work? Organizations that answer these questions honestly, with evidence rather than optimism, build AI systems that earn the trust they require and deliver the value they promise.&lt;/p&gt;
&lt;p&gt;An AI system that passes every assessment on evidence earns confidence. An AI system that skips assessments borrows confidence it may never repay.&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and consulting work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative risk modeling, predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and globally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance landscapes.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item></channel></rss>