<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Iso-31000 |</title><link>https://hwyler.github.io/tags/iso-31000/</link><atom:link href="https://hwyler.github.io/tags/iso-31000/index.xml" rel="self" type="application/rss+xml"/><description>Iso-31000</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 28 Mar 2026 00:00:00 +0000</lastBuildDate><image><url>https://hwyler.github.io/media/icon_hu_cd51c91342a84ed6.png</url><title>Iso-31000</title><link>https://hwyler.github.io/tags/iso-31000/</link></image><item><title>How to Actually Use ISO/IEC 23894 for AI Risk Management</title><link>https://hwyler.github.io/blog/how-to-actually-use-iso-iec-23894-for-ai-risk-management/</link><pubDate>Sat, 28 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/how-to-actually-use-iso-iec-23894-for-ai-risk-management/</guid><description>&lt;h2 id="practical-isoiec-23894-implementation-for-ai-risk-management-without-turning-it-into-shelf-decoration"&gt;Practical ISO/IEC 23894 Implementation for AI Risk Management (Without Turning It Into Shelf Decoration)&lt;/h2&gt;
&lt;p&gt;Most AI risk programs fail before the first risk is ever scored.&lt;/p&gt;
&lt;p&gt;They fail because teams treat AI risk management as a
exercise, a model review checklist, or a late-stage legal sign-off. Then the first serious issue hits. Training data rights were unclear. A model drifts in production. A vendor changes an API. An automated decision harms a customer group nobody mapped. The organization scrambles, and trust evaporates fast.&lt;/p&gt;
&lt;p&gt;This is why ISO/IEC 23894 matters. It gives organizations a practical structure for AI risk management that fits how AI is actually built, bought, deployed, and used. In this post, I’ll show you how to turn ISO/IEC 23894 into an
with governance approval gates, clear role ownership, and an implementation checklist that avoids the common failure points I keep seeing in audits, design reviews, and board briefings.&lt;/p&gt;
&lt;p&gt;Here is why it fails: ISO/IEC 23894 is not a checklist. It is a guidance document built on top of ISO 31000, the general risk management standard, with AI-specific extensions layered in. If you treat it like a form to fill out, you will produce documentation that looks complete but protects nobody.&lt;/p&gt;
&lt;p&gt;This post walks through the standard&amp;rsquo;s actual structure, explains what each section demands in practice, and gives you the field-tested implementation tips I have gathered from helping organizations build AI risk management programs that survive contact with real AI systems.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/chatgpt-image-sep-11-2026-10_45_14-pm.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="why-isoiec-23894-exists-and-what-problem-it-solves"&gt;Why ISO/IEC 23894 Exists and What Problem It Solves&lt;/h2&gt;
&lt;p&gt;Before 2023, organizations managing AI risk had to improvise. They would borrow bits from information security frameworks, add some data governance controls, and hope the combination covered enough ground. It rarely did.&lt;/p&gt;
&lt;p&gt;ISO/IEC 23894:2023 was created by ISO/IEC JTC 1/SC 42 to provide a structured approach for any organization that develops, deploys, or uses AI systems. The standard applies the well-established ISO 31000 risk management framework to the specific challenges AI introduces. Think of it as a translation layer. It takes proven risk management principles and shows you exactly where AI creates new wrinkles.&lt;/p&gt;
&lt;p&gt;The standard covers three domains: principles that guide your thinking, a framework for embedding AI risk management into your organization, and processes for actually identifying, assessing, and treating AI-specific risks.&lt;/p&gt;
&lt;p&gt;The biggest mistake organizations make is treating ISO/IEC 23894 as a standalone. It explicitly references and extends ISO 31000:2018. If your team has not read ISO 31000 first, they will misunderstand the guidance in 23894 because they will lack the foundational context. Buy both standards. Read 31000 first. Then read 23894 as the AI-specific annotation layer it was designed to be.&lt;/p&gt;
&lt;h2 id="the-three-part-architecture-you-need-to-understand"&gt;The Three-Part Architecture You Need to Understand&lt;/h2&gt;
&lt;p&gt;ISO/IEC 23894 mirrors the clause structure of ISO 31000 deliberately. This is not an accident. The authors wanted organizations that already use ISO 31000 to integrate AI risk management without rebuilding everything from scratch. The three parts work together as a system.&lt;/p&gt;
&lt;h3 id="part-one-principles-clause-4"&gt;Part One: Principles (Clause 4)&lt;/h3&gt;
&lt;p&gt;The principles define the foundational values that should shape every AI risk decision your organization makes. ISO 31000 defines eight principles. ISO/IEC 23894 adds AI-specific guidance to five of them: Inclusive, Dynamic, Best Available Information, Human and Cultural Factors, and Continual Improvement.&lt;/p&gt;
&lt;p&gt;The &amp;ldquo;inclusive&amp;rdquo; principle is where most organizations stumble first. AI systems affect a wider set of stakeholders than traditional software. The standard explicitly calls out that stakeholders can help identify risks in data collection, define fairness criteria, identify bias, and determine where human oversight is needed. This is not a suggestion. If your risk management process does not include diverse stakeholder input, you are missing risks that will surface later in the worst possible way.&lt;/p&gt;
&lt;p&gt;The &amp;ldquo;dynamic&amp;rdquo; principle matters more for AI than for almost any other technology domain. AI systems based on machine learning can change their behavior through continuous learning. Customer expectations shift quickly. Regulatory requirements are updating constantly. Your risk management process needs to account for a system that is itself a moving target.&lt;/p&gt;
&lt;p&gt;When I first helped a financial services firm apply the &amp;ldquo;Inclusive&amp;rdquo; principle, they interpreted &amp;ldquo;stakeholder involvement&amp;rdquo; as sending a survey to the compliance team. That is not what the standard means. You need a structured dialog with people who will be affected by the AI system&amp;rsquo;s decisions. For a credit scoring model, that means talking to loan officers, applicants from different demographic groups, and consumer advocacy organizations. Map your stakeholders before you start the risk assessment, not after.&lt;/p&gt;
&lt;h3 id="part-two-framework-clause-5"&gt;Part Two: Framework (Clause 5)&lt;/h3&gt;
&lt;p&gt;The framework section describes how to embed AI risk management into your organizational structure. It covers leadership commitment, integration with existing management systems, organizational design, resource allocation, and communication.&lt;/p&gt;
&lt;p&gt;Two sub-clauses deserve special attention.&lt;/p&gt;
&lt;p&gt;Clause 5.2 on leadership and commitment adds an AI-specific requirement that many organizations overlook. Because trust and accountability are especially important for AI, the standard says top management should consider issuing public statements about their commitment to AI risk management. This is not corporate PR. It creates an accountability anchor. Once your CEO has publicly committed to responsible AI, the organization has real pressure to follow through.&lt;/p&gt;
&lt;p&gt;Clause 5.4.3 on assigning roles is where the framework becomes
. The standard requires that top management and oversight bodies allocate resources and identify specific individuals with authority to address AI risks and responsibility for monitoring AI risk processes. Not committees. Not shared inboxes. Named people with clear authority.&lt;/p&gt;
&lt;h3 id="part-three-processes-clause-6"&gt;Part Three: Processes (Clause 6)&lt;/h3&gt;
&lt;p&gt;This is where the standard gets specific about what you actually do. The risk management process follows a sequence: define scope and context, assess risks (identify, analyze, evaluate), treat risks, then monitor and report. Each step has AI-specific extensions.&lt;/p&gt;
&lt;p&gt;The process section is the longest part of the standard for good reason. AI risk assessment requires you to think about assets, risk sources, events. Do not try to build your AI risk management process on a blank sheet of paper. Clause 6.4.1 specifically recommends using the catalogue of AI-related risk sources in Annex B as a baseline for organizations performing risk assessment for the first time. I have seen teams spend three months trying to brainstorm AI risk sources when the standard already provides a structured catalogue. Start there. Customize from there.&lt;/p&gt;
&lt;h2 id="stage-1-establishing-scope-context-and-criteria"&gt;Stage 1: Establishing Scope, Context, and Criteria&lt;/h2&gt;
&lt;p&gt;This stage determines what your AI risk management process covers and how it connects to your broader organizational context. Get this wrong and everything downstream is compromised.&lt;/p&gt;
&lt;p&gt;The standard requires you to build an inventory of where AI systems are being developed or used in your organization. This sounds straightforward. It is not. In every organization I have worked with, the initial AI inventory missed at least 30% of actual AI usage. Teams embed machine learning models in spreadsheet macros, use AI-powered SaaS tools without formal procurement, or inherit AI components through acquisitions.&lt;/p&gt;
&lt;p&gt;For external context, the standard provides Table 2 with specific considerations. You need to track relevant legal requirements for AI, ethical guidelines from government and industry groups, domain-specific AI frameworks, technology trends, and societal implications of AI deployment. For internal context, Table 3 adds considerations about how AI affects organizational culture, the availability of AI expertise, intellectual property implications, and data quality constraints.&lt;/p&gt;
&lt;p&gt;Defining risk criteria for AI requires you to understand uncertainty across the entire AI system. The standard calls out data, software, mathematical models, physical extensions, and human-in-the-loop aspects. This is a broader scope than most organizations initially consider.&lt;/p&gt;
&lt;p&gt;What to do: Build your AI system inventory first. Document every AI system or component, its purpose, its data sources, who built it, who operates it, and who is affected by its outputs. Then map the external and internal context factors from Tables 2 and 3. Only then define your risk criteria.&lt;/p&gt;
&lt;p&gt;The standard warns that &amp;ldquo;AI is a fast-moving technology domain&amp;rdquo; and that measurement methods should be &amp;ldquo;consistently evaluated according to their effectiveness.&amp;rdquo; I learned this the hard way when a client&amp;rsquo;s risk criteria for a natural language processing system became obsolete within eight months because the underlying model was replaced with a fundamentally different architecture. Build a review trigger into your risk criteria. Any time the AI model architecture, training data source, or deployment context changes, the risk criteria should be re-evaluated. Do not wait for the annual review.&lt;/p&gt;
&lt;h2 id="stage-2-risk-assessment-the-core-of-the-process"&gt;Stage 2: Risk Assessment, the Core of the Process&lt;/h2&gt;
&lt;p&gt;Risk assessment has three sub-stages: identification, analysis, and evaluation. The standard treats each with specific AI guidance.&lt;/p&gt;
&lt;h3 id="risk-identification"&gt;Risk Identification&lt;/h3&gt;
&lt;p&gt;The standard breaks identification into five activities: identifying assets and their value, risk sources, potential events and outcomes, existing controls, and consequences. Each requires AI-specific thinking.&lt;/p&gt;
&lt;p&gt;For assets, you need to consider three levels: organizational (data, models, the AI system itself, reputation, trust), individual (personal data, privacy, health, safety), and societal (environment, socio-cultural values, educational equity). This three-level approach is one of the most important contributions of the standard. Most organizations only think about organizational assets when they identify AI risks. The standard forces you to consider who bears the consequences.&lt;/p&gt;
&lt;p&gt;For risk sources, Annex B provides categories including complexity of environment, lack of transparency and explainability, level of automation, machine learning specific risks, hardware issues, system life cycle issues, and technology readiness. Each category contains specific risk scenarios.&lt;/p&gt;
&lt;p&gt;For consequences, the standard makes a critical distinction that many teams miss. It instructs you to &amp;ldquo;identify any differences between the groups who experience the benefits of the technology and the groups who experience negative consequences.&amp;rdquo; This is not theoretical. A predictive policing system might benefit a city&amp;rsquo;s police department while disproportionately harming specific communities. A hiring algorithm might benefit an HR team&amp;rsquo;s efficiency while systematically disadvantaging certain applicant groups.&lt;/p&gt;
&lt;p&gt;What to do: For each AI system in your inventory, work through all five identification activities. Use Annex B as your starting checklist for risk sources. Document consequences at all three levels: organization, individual, and society.&lt;/p&gt;
&lt;p&gt;The standard lists methods for identifying potential events, including published standards, scientific papers, market data, incident reports, field trials, stakeholder reports, and expert interviews. When I run risk identification workshops, I always start with incident reports on similar systems. Nothing focuses a risk identification session like showing the team a real-world failure of a system similar to theirs. Search for published incidents, regulatory enforcement actions, and academic case studies related to your specific AI application domain before the workshop begins.&lt;/p&gt;
&lt;h3 id="risk-analysis"&gt;Risk Analysis&lt;/h3&gt;
&lt;p&gt;Risk analysis requires you to assess both consequences and likelihood. The standard distinguishes between three types of impact assessment: business impact, individual impact, and societal impact.&lt;/p&gt;
&lt;p&gt;For individual impact assessment, the standard specifies a detailed list of considerations: types of data used, intended impact, potential bias impact, potential impact on fundamental rights, fairness impact, safety implications, and the jurisdictional and cultural environment of the individual. This last point is easy to overlook. An AI system&amp;rsquo;s impact on an individual can vary dramatically depending on the legal and cultural context in which that individual lives.&lt;/p&gt;
&lt;p&gt;For likelihood assessment, the standard includes a nuanced warning that many organizations miss. It states that &amp;ldquo;there can be significant technical, economic and heuristic issues with decision-making based on likelihoods, particularly when the likelihood either can&amp;rsquo;t be calculated or where the calculation has a large margin of error.&amp;rdquo; In plain language: if you cannot reliably estimate how likely an AI failure is, do not pretend you can. Focus instead on consequence severity and your ability to detect and respond to failures.&lt;/p&gt;
&lt;p&gt;What to do: Run separate impact assessments for business, individuals, and society. Do not collapse them into a single score. For each risk, decide whether a meaningful likelihood estimate is possible. If it is not, shift your analysis to focus on consequence severity and control effectiveness.&lt;/p&gt;
&lt;p&gt;I once watched a team assign a &amp;ldquo;low likelihood&amp;rdquo; score to a bias risk in a hiring algorithm because the model had performed well in testing. Six months after deployment, the model was producing biased outcomes because the production data distribution had drifted from the test data. The team&amp;rsquo;s likelihood estimate was based on a snapshot that was already stale. For AI systems, especially those using machine learning, likelihood estimates decay faster than for traditional systems. Reassess likelihood every time the model is retrained, the data source changes, or the deployment population shifts.&lt;/p&gt;
&lt;h3 id="risk-evaluation"&gt;Risk Evaluation&lt;/h3&gt;
&lt;p&gt;Risk evaluation compares the analyzed risks against your established criteria to determine which risks need treatment and what priority they receive. The standard defers to ISO 31000:2018 here without adding AI-specific guidance, which tells you something important. The evaluation step is about organizational judgment, not technical analysis. You need decision-makers at the table who understand both the technology and the business context.&lt;/p&gt;
&lt;h2 id="stage-3-risk-treatment-and-implementation"&gt;Stage 3: Risk Treatment and Implementation&lt;/h2&gt;
&lt;p&gt;Once risks are evaluated, you choose treatment options. The standard lists the same options as ISO 31000: avoid the risk, take or increase the risk to pursue opportunity, remove the risk source, change the likelihood, change the consequences, share the risk, or retain the risk by informed decision.&lt;/p&gt;
&lt;p&gt;The AI-specific addition here is the concept of a risk-benefit analysis for residual risks. If you cannot reduce negative consequences to an acceptable level through treatment, the standard requires you to perform a risk-benefit analysis. This is particularly relevant for AI because some AI risks (like model opacity in deep learning) cannot be fully eliminated. You need to decide whether the benefits justify the residual risk.&lt;/p&gt;
&lt;p&gt;What to do: For each risk that exceeds your tolerance thresholds, select a treatment option and document it in a risk treatment plan. For residual risks that remain above tolerance after treatment, conduct a formal risk-benefit analysis. Record the rationale for accepting any residual risks.&lt;/p&gt;
&lt;p&gt;
for AI systems need to be version-aware. Traditional risk treatment plans assume relatively stable systems. AI systems, especially those using continuous learning, change over time. Your treatment plan should specify which version of the model it applies to and include trigger conditions for re-evaluation. I recommend tagging each treatment plan entry with the model version, training data date, and deployment configuration it was validated against. When any of these change, the treatment plan enters a mandatory review cycle.&lt;/p&gt;
&lt;h2 id="stage-4-monitoring-recording-and-reporting"&gt;Stage 4: Monitoring, Recording, and Reporting&lt;/h2&gt;
&lt;p&gt;The standard&amp;rsquo;s requirements for recording and reporting are more specific than many teams expect. Clause 6.7 requires organizations to establish a system for collecting and verifying information from both implementation and post-implementation phases, and to collect publicly available information on similar systems.&lt;/p&gt;
&lt;p&gt;This information must be assessed for relevance to the trustworthiness of the AI system. The standard specifically asks whether previously undetected risks exist or whether previously assessed risks are no longer acceptable. When either condition is true, you must perform a review of risk management activities and evaluate the effects on existing controls.&lt;/p&gt;
&lt;p&gt;The recording requirements include: system description and identification, methodology applied, intended use description, identity of assessors, terms of reference and date, release status, and degree to which objectives have been met. This is not optional documentation. It creates the audit trail that regulators and oversight bodies will examine.&lt;/p&gt;
&lt;p&gt;What to do: Build a risk management record template that captures all required fields. Establish a cadence for collecting and reviewing post-implementation data. Create triggers that automatically initiate risk reassessment when conditions change.&lt;/p&gt;
&lt;p&gt;The standard says risk management records &amp;ldquo;should allow the traceability of each identified risk through all risk management processes.&amp;rdquo; In practice, this means you need a risk register with unique identifiers for each risk that persist across assessment cycles. I have seen organizations create new risk registers for each assessment, losing the historical thread. Use a single, versioned risk register where each risk has a persistent ID, and track its status changes over time. This is the only way to demonstrate to auditors that your process is continuous, not episodic.&lt;/p&gt;
&lt;h2 id="implementation-tips"&gt;Implementation Tips&lt;/h2&gt;
&lt;p&gt;These apply across all stages and will determine whether your AI risk management program actually works.&lt;/p&gt;
&lt;h3 id="tip-1-align-risk-management-with-the-ai-system-life-cycle"&gt;Tip 1: Align Risk Management with the AI System Life Cycle&lt;/h3&gt;
&lt;p&gt;Annex C of the standard maps risk management activities to AI system life cycle stages: inception, design and development, verification and validation, deployment, operation and monitoring, continuous validation, re-evaluation, and retirement or replacement. This mapping is not decorative. It tells you which risk management activities should happen at each stage.&lt;/p&gt;
&lt;p&gt;Most organizations I work with front-load their risk management effort at the design stage and then drop attention during operation. The standard explicitly shows that risk assessment, treatment, monitoring, and recording continue through every life cycle stage, including retirement. When you decommission an AI system, you can lose decision expertise and information. Plan for that. Document what the system knew and how it made decisions before you turn it off.&lt;/p&gt;
&lt;h3 id="tip-2-handle-stakeholder-identification-seriously"&gt;Tip 2: Handle Stakeholder Identification Seriously&lt;/h3&gt;
&lt;p&gt;The standard provides a list of stakeholder categories: the organization itself, customers, partners, third parties, suppliers, end users, regulators, civil organizations, individuals, affected communities, and societies. That is nine categories. Most organizations consult two or three.&lt;/p&gt;
&lt;p&gt;Create a stakeholder map for each AI system at the inception stage. For each stakeholder category, document what information they need, how they are affected by the system, and how you will engage them. Update this map when the system&amp;rsquo;s scope or deployment context changes. The stakeholder categories that matter most are often the ones farthest from the development team. Affected communities and end users rarely have a voice in risk assessment unless you build a specific mechanism to include them.&lt;/p&gt;
&lt;h3 id="tip-3-document-your-risk-criteria-decisions-explicitly"&gt;Tip 3: Document Your Risk Criteria Decisions Explicitly&lt;/h3&gt;
&lt;p&gt;The standard&amp;rsquo;s Table 4 on risk criteria includes a requirement that organizations &amp;ldquo;take reasonable steps to understand uncertainty in all parts of the AI system.&amp;rdquo; This includes data, software, mathematical models, physical extensions, and human-in-the-loop aspects.&lt;/p&gt;
&lt;p&gt;When defining risk criteria, write down not just what your criteria are, but why you chose those thresholds. I worked with an organization that set a 95% accuracy threshold for an AI diagnostic tool. When a regulator asked why 95% and not 97% or 99%, nobody could answer. The threshold had been copied from a different project. Document the reasoning behind every criterion. Reference the clinical studies, industry benchmarks, or stakeholder consultations that informed your decision. This documentation is what separates a defensible risk management process from an arbitrary one.&lt;/p&gt;
&lt;h3 id="tip-4-treat-transparency-as-a-multi-audience-challenge"&gt;Tip 4: Treat Transparency as a Multi-Audience Challenge&lt;/h3&gt;
&lt;p&gt;Annex B of the standard discusses transparency and explainability as risk sources. It makes a point that &amp;ldquo;the kind and level of information that is appropriate strongly depends on the stakeholders, use case, system type and legislative requirements.&amp;rdquo; One-size-fits-all transparency does not work.&lt;/p&gt;
&lt;p&gt;Build a transparency framework that segments by audience. The standard&amp;rsquo;s Table 1 mentions tailoring transparency to &amp;ldquo;relevant personas&amp;rdquo; such as regulators, business owners, and model risk evaluators. In practice, I create three transparency tiers. Tier one is the public-facing description of what the AI system does and does not do. Tier two is the detailed technical documentation for internal reviewers and regulators. Tier three is the full model documentation, including training data provenance, architecture decisions, and test results. Each tier serves a different audience and contains different information. Trying to serve all audiences with one document produces a document that serves none of them.&lt;/p&gt;
&lt;h2 id="what-this-standard-actually-does-in-practice"&gt;What This Standard Actually Does in Practice&lt;/h2&gt;
&lt;h2 id="part-1-the-ai-risk-management-process"&gt;Part 1: The AI Risk Management Process&lt;/h2&gt;
&lt;p&gt;This is Clause 6 of the standard. It is the longest and most detailed section because it describes what you actually do.&lt;/p&gt;
&lt;h3 id="step-1-define-your-scope-context-and-criteria"&gt;Step 1: Define Your Scope, Context, and Criteria&lt;/h3&gt;
&lt;p&gt;Before you assess any risks, you need to answer three questions. What AI systems are we managing? What is the environment around them? And what criteria will we use to decide if a risk matters?&lt;/p&gt;
&lt;p&gt;The standard requires you to build an inventory of where AI is being developed or used in your organization. This inventory must be documented and included in your risk management process.&lt;/p&gt;
&lt;p&gt;Practical example: A mid-size insurance company I worked with discovered during this step that 14 different teams were using AI-powered tools, but only 3 had been formally identified by the IT governance team. Six were SaaS products with embedded ML models. Two were spreadsheet-based models built by actuaries. Three were
claims processing systems. The inventory step alone changed their understanding of their AI exposure.&lt;/p&gt;
&lt;p&gt;For context, the standard requires you to consider nine categories of stakeholders:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Your own organization&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Customers, partners, and third parties&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Suppliers&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;End users&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Regulators&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Civil organizations&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Individuals affected by the AI system&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Affected communities&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Societies broadly&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That last category is not abstract. If your AI system makes lending decisions, &amp;ldquo;societies&amp;rdquo; includes the economic communities shaped by those decisions over time.&lt;/p&gt;
&lt;p&gt;The standard also requires you to consider whether your AI systems can harm human beings, deny essential services, infringe human rights through biased automated decisions, or contribute to environmental harm.&lt;/p&gt;
&lt;p&gt;For risk criteria, the standard adds an AI-specific requirement that matters enormously in practice: you must understand uncertainty across all parts of the AI system. That includes the data, the software, the mathematical models, any physical components, and the human-in-the-loop aspects like data labeling. Most organizations define risk criteria only around the model itself and miss everything upstream and downstream.&lt;/p&gt;
&lt;p&gt;Practical insight for risk managers: Your AI risk appetite should account for your organization&amp;rsquo;s actual AI capacity and knowledge level. The standard says this directly. If your team has limited ML expertise, your risk appetite for complex deep learning systems should be lower than an organization with a mature data science function. This sounds obvious, but I have seen organizations approve high-risk AI projects with the same risk appetite they use for rule-based automation.&lt;/p&gt;
&lt;p&gt;Practical insight for data scientists: The standard warns that &amp;ldquo;AI is a fast-moving technology domain&amp;rdquo; and that measurement methods should be &amp;ldquo;consistently evaluated according to their effectiveness.&amp;rdquo; That means the metrics you use to evaluate model performance today might not be appropriate six months from now. Build metric review into your model monitoring cadence.&lt;/p&gt;
&lt;h3 id="step-2-identify-risks"&gt;Step 2: Identify Risks&lt;/h3&gt;
&lt;p&gt;Risk identification in ISO/IEC 23894 covers five distinct activities. Each one matters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identify assets and their value.&lt;/strong&gt; The standard requires you to think about assets at three levels:&lt;/p&gt;
&lt;p&gt;Organizational assets include your data, your trained models, the AI system itself (tangible), plus your reputation and stakeholder trust (intangible).&lt;/p&gt;
&lt;p&gt;Individual assets include personal data (tangible), plus privacy, health, and safety (intangible).&lt;/p&gt;
&lt;p&gt;Community and societal assets include the environment (tangible), plus socio-cultural beliefs, educational access, and equity (intangible).&lt;/p&gt;
&lt;p&gt;Practical example: When a healthcare AI startup assessed assets for their diagnostic imaging tool, they initially listed only their model and training data. The three-level framework forced them to also consider patient safety (individual intangible), the hospital&amp;rsquo;s reputation for diagnostic accuracy (organizational intangible), and equitable access to accurate diagnosis across demographic groups (societal intangible). Each of these surfaced risks that the initial asset list would have missed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identify risk sources.&lt;/strong&gt; The standard provides categories: organizational factors, processes, personnel, physical environment, data, AI system configuration, deployment environment, hardware and software, and dependence on external parties. Annex B expands these into detailed scenarios (covered in Part 2 below).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identify potential events and outcomes.&lt;/strong&gt; The standard lists specific methods for finding these: published standards and papers, market data on similar systems, incident reports on similar systems, field trials, usability studies, stakeholder reports, expert interviews, and simulations.&lt;/p&gt;
&lt;p&gt;Practical insight for AI security analysts: Incident reports on similar systems are the most underused source on this list. Before running any risk identification workshop, search for publicly reported failures, adversarial attacks, and regulatory actions against AI systems similar to yours. The
the
, and
CVE databases are good starting points. Nothing sharpens a risk identification session like a real-world failure story from your domain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identify existing controls.&lt;/strong&gt; Document what controls already exist and assess whether they actually work. The standard specifically calls out the importance of identifying control failures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identify consequences.&lt;/strong&gt; This is where the standard adds its most valuable AI-specific guidance. It instructs you to identify differences between the groups that benefit from the technology and the groups that bear negative consequences. A resume screening tool might benefit the HR department (faster processing) while systematically disadvantaging applicants from certain educational backgrounds or geographic regions.&lt;/p&gt;
&lt;p&gt;Consequences to organizations include investigation and repair time, lost opportunities, reputational damage, regulatory penalties, and litigation.&lt;/p&gt;
&lt;p&gt;Consequences to individuals and societies are harder to quantify but often more severe: threats to health, violations of privacy, infringement of fundamental rights.&lt;/p&gt;
&lt;p&gt;The standard makes a practical point that experienced practitioners already know: consequences for individuals and societies almost always affect the organization as well. A safety incident creates liability claims. A bias scandal damages the brand. Map these cascading effects explicitly.&lt;/p&gt;
&lt;h3 id="step-3-analyze-risks"&gt;Step 3: Analyze Risks&lt;/h3&gt;
&lt;p&gt;Risk analysis requires three separate impact assessments:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Business impact assessment.&lt;/strong&gt; How badly does this risk affect the organization? Consider criticality, tangible versus intangible impacts, and your established criteria.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Individual impact assessment.&lt;/strong&gt; How does this risk affect the people whose data is used or whose lives are influenced by the AI system? The standard lists specific factors: types of personal data used, potential bias impact, potential impact on fundamental rights, fairness impact, safety of the individual, existing protections against bias, and the jurisdictional and cultural environment of the individual.&lt;/p&gt;
&lt;p&gt;That last factor is easy to overlook but critical. An AI system&amp;rsquo;s impact on an individual in the EU (where GDPR applies) differs from its impact on an individual in a jurisdiction with no data protection law. The same system creates different risk profiles in different markets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Societal impact assessment.&lt;/strong&gt; How broadly does the AI system reach into different populations? The standard specifically asks you to consider whether the system amplifies or reduces pre-existing patterns of harm to different social groups.&lt;/p&gt;
&lt;p&gt;Example: A government agency deploying a predictive policing model would face very different societal impact conclusions than a private company using the same technology for retail theft prevention. The government use case reaches more broadly and carries the authority of the state, which amplifies both benefits and harms.&lt;/p&gt;
&lt;p&gt;For likelihood assessment, the standard includes a warning that practitioners should take seriously: &amp;ldquo;There can be significant technical, economic and heuristic issues with decision-making based on likelihoods, particularly when the likelihood either can&amp;rsquo;t be calculated or where the calculation has a large margin of error.&amp;rdquo; In plain language: if you cannot meaningfully estimate how likely something is, do not force a number. Focus on consequence severity and your ability to detect and respond instead.&lt;/p&gt;
&lt;p&gt;Practical insight for data scientists: Likelihood estimates for AI system failures decay faster than for traditional software. A model&amp;rsquo;s failure probability changes every time the data distribution shifts, the model is retrained, or the user population changes. If you assign a likelihood score, attach an expiration date to it.&lt;/p&gt;
&lt;h3 id="step-4-evaluate-risks"&gt;Step 4: Evaluate Risks&lt;/h3&gt;
&lt;p&gt;Compare the analyzed risks against your criteria. Prioritize. Decide which risks need treatment. The standard defers to ISO 31000 here without AI-specific additions, which tells you this step is about organizational judgment, not technical analysis. Get decision-makers in the room who understand both the technology and the business.&lt;/p&gt;
&lt;h3 id="step-5-treat-risks"&gt;Step 5: Treat Risks&lt;/h3&gt;
&lt;p&gt;The standard provides seven treatment options, consistent with ISO 31000:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Avoid the risk (stop or do not start the activity)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Accept increased risk to pursue an opportunity&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Remove the risk source&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Change the likelihood&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Change the consequences&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Share the risk (contracts, insurance)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Retain the risk by informed decision&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The AI-specific addition: if you cannot reduce negative consequences to an acceptable level through any treatment option, you must perform a risk-benefit analysis for the residual risk. This matters because some AI risks cannot be fully eliminated. The opacity of a deep learning model, for example, is inherent to the technology. You need to decide if the benefits justify the residual risk, and you need to document that decision.&lt;/p&gt;
&lt;p&gt;Practical insight for risk managers: Each treatment measure must be verified for effectiveness and recorded. Do not just document the plan. Document whether the treatment actually worked. I have seen organizations with detailed treatment plans and zero follow-up on whether the treatments reduced the risk as expected.&lt;/p&gt;
&lt;h3 id="step-6-monitor-record-and-report"&gt;Step 6: Monitor, Record, and Report&lt;/h3&gt;
&lt;p&gt;The standard&amp;rsquo;s recording requirements are specific. Your risk management record must include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Description and identification of the analyzed system&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Methodology used&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Intended use of the AI system&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Who performed the risk assessment&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Terms of reference and date&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Release status of the assessment&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Whether and to what degree objectives were met&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The standard also requires that you collect information from post-implementation phases and review publicly available information about similar systems. You must assess whether previously undetected risks exist or whether previously accepted risks are no longer acceptable.&lt;/p&gt;
&lt;p&gt;The records must allow traceability of each identified risk through all risk management processes. This means persistent risk IDs, version-controlled risk registers, and a clear audit trail from identification through treatment and monitoring.&lt;/p&gt;
&lt;p&gt;Practical insight for all three audiences: When the standard says &amp;ldquo;collect and review publicly available information on similar systems on the market,&amp;rdquo; it means this is an ongoing obligation, not a one-time activity. Set up alerts for incident reports, regulatory actions, and published research about AI systems similar to yours. The best early warning system for your own AI risks is someone else&amp;rsquo;s AI failure.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="part-2-risk-sources-and-objectives-your-identification-checklists"&gt;Part 2: Risk Sources and Objectives (Your Identification Checklists)&lt;/h2&gt;
&lt;p&gt;Annexes A and B of the standard provide catalogs that serve as starting points for risk identification. The standard itself says these catalogs have &amp;ldquo;shown value&amp;rdquo; for organizations performing AI risk assessment for the first time. Use them as baselines, not as exhaustive lists.&lt;/p&gt;
&lt;h3 id="ai-related-objectives-to-protect-annex-a"&gt;AI-Related Objectives to Protect (Annex A)&lt;/h3&gt;
&lt;p&gt;These are the things that can go wrong or right with AI systems. For each objective, the standard provides context on why it matters for AI specifically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Accountability.&lt;/strong&gt; AI changes who is responsible for decisions. When a person made a lending decision, that person was accountable. When an AI system makes that decision, accountability becomes unclear. Regulators worldwide are still working out who bears responsibility. You need to know the legislation in every market where your AI system operates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Expertise.&lt;/strong&gt; Building AI systems requires interdisciplinary specialists, not just software engineers. The standard also extends this to end users: they need enough understanding of the AI system to detect and override erroneous outputs.&lt;/p&gt;
&lt;p&gt;Practical example: A manufacturing company deployed an AI-powered quality inspection system but did not train the floor operators on how the system made decisions or what its failure modes looked like. When the system began missing defects due to a lighting change in the factory, operators trusted the system&amp;rsquo;s &amp;ldquo;pass&amp;rdquo; decisions for three weeks before someone escalated the rising customer complaint rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Training and Test Data Quality.&lt;/strong&gt; Training and test data must be validated for currency, relevance, diversity, and consistency. If you source data externally, data quality is still your responsibility. The amount of data required varies with the functionality and complexity of the environment.&lt;/p&gt;
&lt;p&gt;Practical insight for data scientists: The standard specifically calls out that training and test datasets should be independent when applicable. This is a basic ML practice, but the standard elevates it to a risk management concern. If your test set leaks into your training data, you have not just a technical problem but a governance failure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Environmental Impact.&lt;/strong&gt; AI can help the environment (optimizing energy use, reducing emissions) or hurt it (massive compute requirements for training). Both sides must be considered.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fairness.&lt;/strong&gt; Unfair outcomes can come from biased objective functions, imbalanced datasets, human biases in training data, biased product concepts, or decisions about when and where to deploy AI systems. The standard references ISO/IEC TR 24027 for deeper guidance on bias.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Maintainability.&lt;/strong&gt; ML-based systems are trained, not programmed. Modifying them to fix defects or adapt to new requirements is fundamentally different from patching traditional software. Understand the implications before deploying.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Privacy.&lt;/strong&gt; AI systems that depend on large datasets create privacy risks through data collection, through inference of sensitive information, and through model personalization. The standard notes that AI can infer sensitive personal data even when that data was not directly provided. A data protection impact assessment (per ISO/IEC 29134) is recommended.&lt;/p&gt;
&lt;p&gt;Practical insight for AI security analysts: The standard highlights that protecting privacy in AI includes protecting access to models personalized for individuals or models that can be used to infer characteristics of similar individuals. This means model extraction attacks are a privacy risk, not just a security risk. If an attacker can replicate your model, they can potentially infer characteristics of your training data subjects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
.&lt;/strong&gt; Can the system maintain performance under unexpected conditions? Neural networks are particularly challenging here because their nonlinear nature can produce unexpected behavior. Characterizing neural network robustness remains an open research problem.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Safety.&lt;/strong&gt; AI systems in vehicles, manufacturing, robotics, and medical devices introduce safety risks that must be evaluated against domain-specific safety standards. The standard does not replace those domain standards. It adds to them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Security.&lt;/strong&gt; Beyond classical information security, AI introduces new attack surfaces: data poisoning (corrupting training data), adversarial attacks (crafted inputs that fool the model), and model stealing (extracting the model through query access). ISO/IEC 27005 covers general information security risk management. The AI-specific threats require additional consideration.&lt;/p&gt;
&lt;p&gt;Practical example for security analysts: A financial institution&amp;rsquo;s fraud detection model was trained on transaction data that included a small number of poisoned records inserted by an insider. The poisoned data taught the model to classify certain fraudulent transaction patterns as legitimate. Classical information security controls (access management, encryption) did not prevent this because the insider had authorized access to the data pipeline. AI-specific controls (data provenance tracking, statistical anomaly detection on training data) would have caught it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;
&lt;/strong&gt; Transparency is about what the organization communicates. Explainability is about what the system can reveal about its own decision-making. Both matter, and they serve different purposes. Transparency builds trust with stakeholders. Explainability enables validation and verification by the organization itself.&lt;/p&gt;
&lt;p&gt;The standard also notes a tension: excessive transparency can create privacy, security, and intellectual property risks. You need to find the right level for each stakeholder group.&lt;/p&gt;
&lt;h3 id="ai-related-risk-sources-annex-b"&gt;AI-Related Risk Sources (Annex B)&lt;/h3&gt;
&lt;p&gt;These are the places where risks originate. Use this as a checklist during risk identification.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Complexity of environment.&lt;/strong&gt; The more complex the operating environment, the harder it is to ensure your training data covers all possible situations. For autonomous driving, you cannot guarantee coverage of every scenario. For a chatbot answering questions about a fixed product catalog, coverage is more achievable. Assess how well-understood your system&amp;rsquo;s environment is, because partial understanding creates uncertainty that is itself a risk source.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lack of transparency and explainability.&lt;/strong&gt; If you cannot explain why your model made a specific decision, you cannot fully validate it. This affects trustworthiness, accountability, safety, security, fairness, and robustness. The standard emphasizes that explainability matters for internal validation, not just external communication.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Level of automation.&lt;/strong&gt; Systems range from fully human-controlled to fully automated. Higher automation means less human oversight, which amplifies both the efficiency gains and the risk exposure. For systems where a human must be &amp;ldquo;ready to take over when necessary,&amp;rdquo; the handover itself is a risk source. Think about response time, operator attention, and situation awareness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Machine learning specific risks.&lt;/strong&gt; Data quality directly affects system behavior. Data collection processes are a risk source that is &amp;ldquo;especially hard to diagnose and detect.&amp;rdquo; Data can become unrepresentative over time. Data sourcing creates
risks. Failing to secure the data pipeline opens the door to adversarial manipulation. Continuous learning systems can change their behavior in production in ways that were not anticipated at deployment.&lt;/p&gt;
&lt;p&gt;Practical example: An e-commerce recommendation engine trained on user interaction data gradually learned to recommend increasingly sensationalized products because those generated more clicks. The continuous learning loop optimized for the engagement metric without any check on whether the recommendations were appropriate. The behavior drift was subtle enough that no one noticed for months.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;System hardware issues.&lt;/strong&gt; Hardware errors, soft errors from radiation, constraints when transferring models between different hardware platforms, and network issues for systems requiring remote processing. These are easier to overlook in AI because teams focus on model performance and forget about the physical infrastructure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;System life cycle issues.&lt;/strong&gt; Risks exist at every stage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Design: failing to anticipate deployment contexts&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Verification and validation: inadequate testing causing regressions&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Deployment: misconfigured resources&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Maintenance: unsupported but still-running systems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Reuse: using a system in a context it was not designed for (the standard gives the example of a social media face detection system repurposed for criminal suspect identification, a far more demanding use case)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Decommissioning: losing the decision expertise embedded in the retired system&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Technology readiness.&lt;/strong&gt; Less mature technologies carry unknown risks. More mature technologies create complacency and technical debt. Both ends of the maturity spectrum require attention.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="part-3-the-organizational-framework"&gt;Part 3: The Organizational Framework&lt;/h2&gt;
&lt;p&gt;This is Clause 5. It describes how to embed AI risk management into your o
.&lt;/p&gt;
&lt;h3 id="leadership-and-commitment"&gt;Leadership and Commitment&lt;/h3&gt;
&lt;p&gt;Top management, supported by
, must do two things the standard calls out specifically for AI:&lt;/p&gt;
&lt;p&gt;First, consider issuing public statements about the organization&amp;rsquo;s commitment to AI risk management. This creates accountability and builds stakeholder confidence.&lt;/p&gt;
&lt;p&gt;Second, recognize that AI risk management requires specialized resources and allocate them. An AI risk program staffed by people without AI expertise will produce documentation that misses the actual risks.&lt;/p&gt;
&lt;h3 id="organizational-context"&gt;Organizational Context&lt;/h3&gt;
&lt;p&gt;The standard provides two detailed tables for understanding your external and internal context.&lt;/p&gt;
&lt;p&gt;For external context, AI-specific considerations include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;AI-related legal requirements in your markets&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ethical guidelines from government groups, regulators, standardization bodies, civil society, academia, and industry associations&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Domain-specific AI frameworks&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Societal implications of your AI deployments&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How continuous learning might affect your ability to meet contractual obligations&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ownership and usage rights for training data provided by third parties&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For internal context, AI-specific considerations include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;How AI changes organizational culture by creating new roles and responsibilities&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Deskilling risks where human decision-making is increasingly replaced by AI&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The availability of AI tools that enable development without full understanding of the technology&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Intellectual property implications of AI systems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Additional data quality constraints imposed by AI&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The need to educate stakeholders on AI capabilities, failure modes, and failure management&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Practical insight for risk managers: The external context requirement to track &amp;ldquo;
and design of AI and automated systems issued by government-related groups, regulators, standardization bodies, civil society, academia and industry associations&amp;rdquo; is broad. Create a regulatory and guidance tracker specific to AI. Assign someone to update it quarterly. The landscape is changing fast enough that annual reviews will miss significant developments.&lt;/p&gt;
&lt;h3 id="roles-and-accountabilities"&gt;Roles and Accountabilities&lt;/h3&gt;
&lt;p&gt;The standard requires top management and oversight bodies to identify specific individuals with:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Authority to address AI risks&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Responsibility for establishing and monitoring processes to address AI risks&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Notice the standard says &amp;ldquo;individuals,&amp;rdquo; not &amp;ldquo;committees.&amp;rdquo; Named accountability matters.&lt;/p&gt;
&lt;p&gt;Practical insight: If the person accountable for AI risk management does not have authority over the AI development teams, the role is ceremonial. Verify that the accountability chain has teeth.&lt;/p&gt;
&lt;h3 id="communication-and-consultation"&gt;Communication and Consultation&lt;/h3&gt;
&lt;p&gt;The standard notes that stakeholders affected by AI systems can be &amp;ldquo;larger than initially foreseen, can include otherwise unconsidered external stakeholders and can extend to other parts of a society.&amp;rdquo; Plan your stakeholder engagement broadly from the start, because discovering a critical stakeholder group after deployment creates reactive crises.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="part-4-the-guiding-principles"&gt;Part 4: The Guiding Principles&lt;/h2&gt;
&lt;p&gt;Clause 4 defines eight risk management principles from ISO 31000. Five of them get AI-specific guidance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inclusive.&lt;/strong&gt; AI systems affect more stakeholders than traditional systems. Engage diverse internal and external groups. Stakeholders help identify data risks, define fairness criteria, spot bias, determine where human oversight is needed, and shape transparency and explainability approaches. The standard suggests segmenting transparency frameworks by stakeholder persona (regulators, business owners, model risk evaluators) when a one-size-fits-all approach does not work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic.&lt;/strong&gt; AI systems change through continuous learning. Customer expectations shift rapidly. Regulations update frequently. Your risk management process must anticipate, detect, and respond to these changes in real time, not annually.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Best available information.&lt;/strong&gt; Historical data about AI failures may be limited because the technology is relatively new. Future expectations change quickly. Track how your AI systems are used after deployment. Be aware that tracking external usage may be limited by IP, contractual, or market restrictions, and capture those limitations explicitly in your risk process.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human and cultural factors.&lt;/strong&gt; Monitor how your AI systems interact with pre-existing societal patterns that affect equitable outcomes, privacy, freedom of expression, fairness, safety, security, employment, the environment, and human rights.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Continual improvement.&lt;/strong&gt; Monitor the AI ecosystem for performance successes, shortcomings, lessons learned, and new research findings. Feed previously unknown risks back into the improvement cycle.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="part-5-risk-management-across-the-ai-system-life-cycle"&gt;Part 5: Risk Management Across the AI System Life Cycle&lt;/h2&gt;
&lt;p&gt;Annex C maps risk management activities to the AI system life cycle defined in ISO/IEC 22989:2022. The life cycle stages are:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Inception&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Design and development&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Verification and validation&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Deployment&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Operation and monitoring&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Continuous validation&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Re-evaluation&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Retirement or replacement&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;At the organizational level, the governing body sets risk appetite, establishes general criteria, and builds catalogs of risk criteria, risk sources, mitigation measures, monitoring techniques, and reporting formats. These catalogs improve over time as feedback flows up from individual AI system risk processes.&lt;/p&gt;
&lt;p&gt;At the project level, each AI system goes through its own risk management cycle at each life cycle stage. Risk criteria, assessments, and treatment plans are established at inception, continuously updated during design and development, verified during testing, potentially adjusted during deployment, and monitored throughout operation.&lt;/p&gt;
&lt;p&gt;The key takeaway: risk management is not a phase. It happens at every stage. The standard explicitly shows risk assessment, treatment, monitoring, and recording activities at every life cycle stage, including retirement.&lt;/p&gt;
&lt;p&gt;Practical insight for data scientists: The re-evaluation stage is often skipped. The standard requires that existing risk sources be examined for relevance, criteria re-evaluated against changes in scope or purpose, and regulatory updates incorporated. Build re-evaluation triggers into your model governance process. Any change in model purpose, data source, or regulatory environment should start a re-evaluation cycle.&lt;/p&gt;
&lt;p&gt;Practical insight for AI security analysts: The retirement stage creates specific risks. When you decommission an AI system, you can lose decision expertise and institutional knowledge embedded in that system. If a replacement system is deployed, the way the organization processes information and makes decisions changes. Both of these transitions create attack surface changes and knowledge gaps that need to be assessed as security risks.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="what-to-do-monday-morning"&gt;What to Do Monday Morning&lt;/h2&gt;
&lt;p&gt;If you are a risk manager: start with the AI system inventory. You cannot manage risks you do not know exist. Then map your stakeholders using the nine-category list. Then compare your existing risk criteria against the AI-specific factors in Table 4 of the standard.&lt;/p&gt;
&lt;p&gt;If you are a data scientist: read Annex B on risk sources, particularly section B.5 on machine learning risks. Then review your current model documentation against the recording requirements in Clause 6.7. The gap between what you document today and what the standard requires is likely significant.&lt;/p&gt;
&lt;p&gt;If you are an AI security analyst: start with Annex A sections on security (A.11), privacy (A.8), and robustness (A.9). Map your current threat model against the AI-specific attack surfaces the standard identifies: data poisoning, adversarial attacks, model stealing. Then check whether your security controls cover the full AI system life cycle or only the deployment and operation stages.&lt;/p&gt;
&lt;p&gt;The standard gives you the structure. The work is in applying it honestly to your specific systems, your specific organization, and your specific stakeholders. That is where the real risk management happens.&lt;/p&gt;
&lt;h2 id="key-references-and-related-standards"&gt;Key References and Related Standards&lt;/h2&gt;
&lt;p&gt;ISO/IEC 23894 does not exist in isolation. It references and connects to a broader ecosystem of standards that together form a comprehensive AI governance framework.&lt;/p&gt;
&lt;p&gt;ISO 31000:2018, Risk management, Guidelines. This is the foundational standard. ISO/IEC 23894 is built directly on top of it.&lt;/p&gt;
&lt;p&gt;ISO/IEC 22989:2022, Artificial intelligence, Concepts and terminology. Provides the AI-specific definitions and the system life cycle model referenced throughout 23894.&lt;/p&gt;
&lt;p&gt;ISO Guide 73:2009, Risk management, Vocabulary. Defines the risk management terms used in both 31000 and 23894.&lt;/p&gt;
&lt;p&gt;ISO/IEC 38507:2022, Governance implications of the use of artificial intelligence by organizations. Covers the governance layer that sits above risk management.&lt;/p&gt;
&lt;p&gt;ISO/IEC TR 24028:2020, Overview of trustworthiness in artificial intelligence. Provides background on AI trustworthiness that informs several sections of 23894.&lt;/p&gt;
&lt;p&gt;ISO/IEC TR 24027:2021, Bias in AI systems and AI-aided decision making. Essential reading for the fairness-related risk identification and analysis sections.&lt;/p&gt;
&lt;p&gt;ISO/IEC 29134:2017, Guidelines for privacy impact assessment. Directly relevant when AI systems process personal data.&lt;/p&gt;
&lt;p&gt;ISO/IEC 27005:2022, Guidance on managing information security risks. Covers the security dimension of AI risk, including threats like data poisoning and adversarial attacks.&lt;/p&gt;
&lt;p&gt;NIST AI Risk Management Framework (AI RMF 1.0). While not referenced in the standard, this U.S. framework maps closely to ISO/IEC 23894 and is increasingly expected by American regulators and enterprise customers.&lt;/p&gt;
&lt;p&gt;EU AI Act (Regulation 2024/1689). The European regulation that makes AI risk management a legal obligation for high-risk AI systems. ISO/IEC 23894 provides a structured path toward many of its requirements.&lt;/p&gt;
&lt;h2 id="supporting-ai-iso-standards-by-topic"&gt;Supporting AI ISO standards by Topic&lt;/h2&gt;
&lt;h3 id="core-ai-management--governance-standards"&gt;&lt;strong&gt;Core AI Management &amp;amp; Governance Standards&lt;/strong&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 42001:2023&lt;/strong&gt; – Information technology — Artificial intelligence — Management system (AIMS).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 38507:2022&lt;/strong&gt; – Governance implications of the use of AI by organizations.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 23894:2023&lt;/strong&gt; – Guidance on risk management for AI.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 42005:2025&lt;/strong&gt; – AI system impact assessment.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 42006:2025&lt;/strong&gt; – Requirements for auditing bodies providing AI management system certification.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="foundational-frameworks--terminology"&gt;&lt;strong&gt;Foundational Frameworks &amp;amp; Terminology&lt;/strong&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 22989:2022&lt;/strong&gt; – AI concepts and terminology.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 23053:2022&lt;/strong&gt; – Framework for AI systems using Machine Learning.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 5338:2023&lt;/strong&gt; – AI system life cycle processes.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 5339:2024&lt;/strong&gt; – Guidance for AI applications.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="trustworthiness-ethics--quality"&gt;&lt;strong&gt;Trustworthiness, Ethics &amp;amp; Quality&lt;/strong&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 24028:2020&lt;/strong&gt; – Overview of trustworthiness in artificial intelligence.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 24368:2022&lt;/strong&gt; – Overview of ethical and societal concerns.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 25059:2023&lt;/strong&gt; – Quality model for AI systems (SQuaRE).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 12791:2024&lt;/strong&gt; – Treatment of unwanted bias in ML tasks.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 12792:2025&lt;/strong&gt; – Transparency taxonomy for AI systems.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 42119-2:2025&lt;/strong&gt; – Testing of AI systems — Part 2: Test data and results.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 4213:2022&lt;/strong&gt; – Assessment of machine learning classification performance.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="data-quality--analytics"&gt;&lt;strong&gt;Data Quality &amp;amp; Analytics&lt;/strong&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 24668:2022&lt;/strong&gt; – Process management framework for big data analytics.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 5259-1:2024&lt;/strong&gt; – Data quality for analytics and ML — Part 1: Overview and terminology.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 5259-3:2024&lt;/strong&gt; – Data quality management requirements and guidelines.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ISO/IEC 5259-4:2024&lt;/strong&gt; – Data quality process framework.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-choice-in-front-of-you"&gt;The Choice in Front of You&lt;/h2&gt;
&lt;p&gt;Organizations that treat ISO/IEC 23894 as a compliance artifact will produce binders full of risk assessments that nobody reads, risk registers that go stale within weeks, and governance structures that exist on paper but have no operational authority. When something goes wrong, and with AI systems something eventually does go wrong, they will discover that their documentation protected nobody. Not the organization, not the individuals affected by the AI system, and not the communities that bore the consequences.&lt;/p&gt;
&lt;p&gt;Organizations that treat ISO/IEC 23894 as a living operational tool will build AI risk management into their development pipelines, their deployment decisions, and their ongoing monitoring. They will have named individuals with real authority, risk criteria grounded in evidence, stakeholder engagement that surfaces risks early, and records that tell a coherent story from inception through retirement. When something goes wrong, they will know about it faster, respond more effectively, and demonstrate to regulators that they took reasonable steps.&lt;/p&gt;
&lt;p&gt;The standard gives you the blueprint. What you build with it depends entirely on whether you treat AI risk management as paperwork or as practice.&lt;/p&gt;
&lt;p&gt;To maximize &lt;strong&gt;SEO authority&lt;/strong&gt; and drive high-value traffic back to your core assets, this &amp;ldquo;About the Author&amp;rdquo; section is restructured to emphasize your specific expertise in the &lt;strong&gt;EU AI Act&lt;/strong&gt;, &lt;strong&gt;Quantitative Risk&lt;/strong&gt;, and &lt;strong&gt;ISO 42001&lt;/strong&gt;. I have humanized the tone to move from a standard bio to a &amp;ldquo;Partnership Invitation,&amp;rdquo; while ensuring the links are prominent and descriptive.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="about-the-author-prof-hernan-huwyler-mba-cpa-caio"&gt;&lt;strong&gt;About the Author: Prof. Hernan Huwyler, MBA, CPA, CAIO&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The frameworks, taxonomies, and implementation toolkits shared in this article are part of the ongoing applied research and executive advisory work of &lt;strong&gt;Prof. Hernan Huwyler&lt;/strong&gt;. These materials are designed for operational realism and are freely available for adaptation in your own &lt;strong&gt;AI Governance, Risk Management, and Compliance (GRC)&lt;/strong&gt; programs under proper attribution.&lt;/p&gt;
&lt;h3 id="bridging-the-gap-between-ai-theory-and-production-controls"&gt;&lt;strong&gt;Bridging the Gap Between AI Theory and Production Controls&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;As an &lt;strong&gt;AI GRC Strategy Director&lt;/strong&gt; and &lt;strong&gt;Quantitative Risk Lead&lt;/strong&gt;, Prof. Huwyler works with global organizations in financial services, healthcare, and the public sector to build frameworks that survive both production demands and regulatory scrutiny. His expertise is focused on tje risk-adjusted AI adoption, ensuring that innovation remains defensible through rigorous &lt;strong&gt;algorithmic auditing&lt;/strong&gt; and &lt;strong&gt;automated compliance protocols&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="global-thought-leadership-and-executive-education"&gt;&lt;strong&gt;Global Thought Leadership and Executive Education&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area with a professional presence in Zurich, Geneva, Madrid, and Berlin, Prof. Huwyler operates where AI development is most active. He serves as an &lt;strong&gt;Executive Advisor&lt;/strong&gt; and &lt;strong&gt;Academic Director at IE Law School&lt;/strong&gt;, delivering specialized corporate training on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;EU AI Act Compliance Strategy&lt;/strong&gt; and &lt;strong&gt;ISO 42001&lt;/strong&gt; integration.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Risk Modeling&lt;/strong&gt; using Python and Monte Carlo simulations.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Predictive Risk Automation&lt;/strong&gt; for Board-level decision-making.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="explore-open-source-grc-tools-and-insights"&gt;&lt;strong&gt;Explore Open-Source GRC Tools and Insights&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Prof. Huwyler maintains a public repository of &lt;strong&gt;Python-based AI governance tools&lt;/strong&gt;, risk model templates, and automated compliance scripts to support the global practitioner community.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Technical Tools &amp;amp; Code:&lt;/strong&gt; Access the &lt;strong&gt;
&lt;/strong&gt; for risk model templates.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Expert Commentary:&lt;/strong&gt; Read the latest on GRC and Internal Audit at &lt;strong&gt;
&lt;/strong&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Professional Network:&lt;/strong&gt; Connect on &lt;strong&gt;
&lt;/strong&gt; to follow real-time updates on the evolving AI regulatory landscape.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h3 id="lets-turn-ai-governance-into-a-competitive-advantage"&gt;&lt;strong&gt;Let’s Turn AI Governance into a Competitive Advantage&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;If you are ready to move beyond &amp;ldquo;check-the-box&amp;rdquo; compliance and start capturing the &lt;strong&gt;positive ROI of AI&lt;/strong&gt;, let’s connect. True governance isn&amp;rsquo;t about slowing down; it’s about creating the structural discipline needed to achieve &lt;strong&gt;massive savings through automation&lt;/strong&gt; and to build &lt;strong&gt;new, resilient revenue streams&lt;/strong&gt; that regulators and customers can trust.&lt;/p&gt;</description></item><item><title>Implementation Tips for Expert Calibration and AI-Augmented Risk Estimation</title><link>https://hwyler.github.io/blog/implementation-tips-for-expert-calibration-and-ai-augmented-risk-estimation/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/implementation-tips-for-expert-calibration-and-ai-augmented-risk-estimation/</guid><description>&lt;h1 id="why-expert-calibration-matters-for-grc-professionals"&gt;Why Expert Calibration Matters for GRC Professionals&lt;/h1&gt;
&lt;p&gt;Most risk assessments rely on expert judgment. When historical loss data is absent, limited, or conflicting, you ask knowledgeable people to estimate probabilities and impacts. The problem is that unstructured expert judgment is unreliable. Experts overestimate rare events, underestimate common ones, anchor to previous numbers, and conform to dominant opinions in group settings.&lt;/p&gt;
&lt;p&gt;Expert calibration is a quantitative technique that measures and improves the accuracy of expert predictions over time. It treats expert judgment as data, subject to the same scientific principles of review, critical appraisal, and repeatability that you&amp;rsquo;d apply to any other data source in your risk assessment.&lt;/p&gt;
&lt;p&gt;The difference between a calibrated risk assessment and an uncalibrated one is the difference between a defensible estimate and an educated guess. Regulators, auditors, and boards increasingly expect the former.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/purposeful-stride-in-minimalist-setting.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="the-core-mechanism-how-expert-calibration-works"&gt;The Core Mechanism: How Expert Calibration Works&lt;/h2&gt;
&lt;h3 id="the-basic-cycle"&gt;The Basic Cycle&lt;/h3&gt;
&lt;p&gt;Expert calibration follows a straightforward cycle. Ask experts to estimate potential losses or probabilities of events occurring. Compare actual outcomes to their estimates. Use multiple data points over time to determine whether an expert tends to overestimate or underestimate. Feed this information back to improve future estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Group estimated probabilities into bands (for example, events the expert rated as 10-20% likely, 20-30% likely, and so on). Compare these bands to actual occurrence rates. A well-calibrated expert who assigns 20% probability to events should see roughly 20% of those events actually occur.&lt;/p&gt;
&lt;p&gt;Calculate each expert&amp;rsquo;s overall accuracy by averaging multiple estimates. A perfectly calibrated expert&amp;rsquo;s estimates should, on average, match what you&amp;rsquo;d expect from a uniform distribution across probability bands.&lt;/p&gt;
&lt;p&gt;Very low probability events present a challenge. If an expert estimates a 2% probability, you need 50 or more observations to determine whether 2% is accurate. For rare events, combine calibration data across similar event categories to build a sufficient sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Start building calibration histories now, even if you don&amp;rsquo;t plan to use them for six months. Every time your organization conducts a risk assessment, record each expert&amp;rsquo;s estimate alongside the question, the date, and eventually the actual outcome. Most organizations can&amp;rsquo;t calibrate their experts because they never retained the historical estimates. They have last year&amp;rsquo;s risk register but not the individual predictions that went into it. Store individual expert estimates in a structured database with fields for expert name, question, estimated probability, estimated impact range, date of estimate, and actual outcome when known. After 12 months of accumulation, you&amp;rsquo;ll have enough data points to calculate meaningful calibration scores for your most active experts. Without this history, calibration is impossible and you&amp;rsquo;re permanently stuck with uncalibrated judgment.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="two-approaches-to-aggregating-expert-opinions"&gt;Two Approaches to Aggregating Expert Opinions&lt;/h2&gt;
&lt;h3 id="behavioral-aggregation-the-workshop-method"&gt;Behavioral Aggregation: The Workshop Method&lt;/h3&gt;
&lt;p&gt;Behavioral aggregation brings experts together in face-to-face meetings to reach shared judgment through discussion and consensus. Experts exchange and debate their knowledge, potentially producing more informed and balanced decisions.&lt;/p&gt;
&lt;p&gt;This method is familiar. Most risk workshops use some version of it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Behavioral aggregation is vulnerable to well-documented biases. Group thinking causes experts to conform to the majority view even when they disagree. The halo effect allows a dominant expert&amp;rsquo;s opinion to unduly influence others. Anchoring causes experts to gravitate toward the first number mentioned. Polarization can prevent consensus even with skilled facilitation. And forced consensus, when imposed despite genuine disagreement, masks important differences in opinion and reduces the quality of the final judgment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; If you must use behavioral aggregation, implement three structural safeguards. First, collect individual written estimates before any group discussion begins. This prevents anchoring to the first number spoken aloud. Second, give equal time to every expert, actively drawing out quiet participants and managing dominant voices. Third, never force consensus. If experts genuinely disagree after discussion, document the disagreement and the range of estimates rather than artificially converging on a single number. A documented range of expert opinion is more honest and more useful than a false consensus that nobody actually believes. I&amp;rsquo;ve facilitated dozens of risk workshops where the &amp;ldquo;consensus&amp;rdquo; estimate was the number the most senior person in the room stated first. Everyone else adjusted toward it. The estimate reflected hierarchy, not expertise.&lt;/p&gt;
&lt;h3 id="algorithm-calibration-the-mathematical-method"&gt;Algorithm Calibration: The Mathematical Method&lt;/h3&gt;
&lt;p&gt;Algorithm calibration limits expert interaction to training and briefing sessions. Consensus is not achieved through discussion but through mathematical aggregation of individual expert opinions.&lt;/p&gt;
&lt;p&gt;This approach makes the aggregation process explicit and auditable. The Classical Model, developed by Roger Cooke, uses a linear combination of judgments weighted by each expert&amp;rsquo;s past performance in estimating risk impacts and probabilities. Better-calibrated experts receive higher weights. Poorly calibrated experts receive lower weights or zero weight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The tradeoff:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Algorithm calibration can be less effective when experts strongly disagree and receive little feedback from their peers. The mathematical aggregation may miss contextual nuances that discussion would surface. But it eliminates group biases entirely, produces reproducible results, and creates an auditable record of exactly how the final estimate was derived.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Use algorithm calibration as your primary method and behavioral discussion as a supplementary input. Collect individual estimates first using the structured elicitation protocol described below. Aggregate them mathematically using calibration weights. Then, if the weighted estimates show extreme divergence among high-weight experts, convene a focused discussion limited to understanding why those experts disagree. The discussion informs whether the divergence reflects genuine uncertainty (which should be preserved in the final estimate as a wider distribution) or a misunderstanding of the scenario (which should be corrected). This sequence, individual estimation first, mathematical aggregation second, targeted discussion third, captures the benefits of both approaches while minimizing the biases of each.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="mathematical-aggregation-methods"&gt;Mathematical Aggregation Methods&lt;/h2&gt;
&lt;h3 id="bayesian-updating"&gt;Bayesian Updating&lt;/h3&gt;
&lt;p&gt;Use each expert&amp;rsquo;s opinion to update your existing knowledge about the risk. Start with a prior estimate based on historical data or organizational experience. Then adjust that estimate based on each expert&amp;rsquo;s input, weighted by how confident you are in both your prior and in each expert&amp;rsquo;s judgment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Define your prior distribution based on available data. For each expert opinion, update the distribution using Bayes&amp;rsquo; theorem. The result is a posterior distribution that incorporates both your historical knowledge and the experts&amp;rsquo; collective judgment. Experts whose opinions align with strong historical evidence reinforce the estimate. Experts whose opinions diverge from historical patterns shift the estimate only if their track record or the strength of their reasoning justifies it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The Bayesian approach works best when you have a meaningful prior, meaning real historical data to start from. If your prior is purely a guess, the Bayesian update is just averaging guesses with extra mathematical notation. Before choosing this method, honestly assess whether your prior distribution is based on data or assumption. If it&amp;rsquo;s based on data, Bayesian updating is powerful. If it&amp;rsquo;s based on assumption, opinion pooling or the Cooke method may be more appropriate because they don&amp;rsquo;t pretend you have knowledge you don&amp;rsquo;t have.&lt;/p&gt;
&lt;h3 id="opinion-pooling-weighted-average"&gt;Opinion Pooling (Weighted Average)&lt;/h3&gt;
&lt;p&gt;Assign each expert a specific weight reflecting their relative expertise and trustworthiness. Combine their opinions as a weighted average. The result is a blended estimate that reflects how much you value each expert&amp;rsquo;s input.&lt;/p&gt;
&lt;p&gt;The Cooke method is a specific form of opinion pooling where weights are determined empirically by each expert&amp;rsquo;s past accuracy, not by subjective assessment of their credentials.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In the Cooke method, give more weight to experts who have been more accurate in the past, measured through calibration questions with known answers. Experts who consistently predict historical outcomes correctly receive higher weights. Experts who consistently miss receive lower weights or zero weight.&lt;/p&gt;
&lt;p&gt;Calculate weights by scoring each expert&amp;rsquo;s responses to calibration questions against known correct answers. The simplest scoring method assigns 1 for correct and 0 for incorrect, totals the scores, and converts them to percentages. More sophisticated scoring uses proper scoring rules that evaluate the full probability distribution each expert provides, not just point estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The weight assignment step is where most implementations fail. Organizations resist giving zero weight to experts with impressive titles or seniority. But the entire point of calibration is that credentials don&amp;rsquo;t guarantee accuracy. An expert with 20 years of experience who consistently overestimates by 300% should receive less weight than a junior analyst who consistently hits within 20% of actual outcomes. If you can&amp;rsquo;t bring yourself to weight experts by demonstrated accuracy rather than organizational rank, don&amp;rsquo;t use the Cooke method. You&amp;rsquo;ll corrupt it by overriding the calibration data with political judgments, and the result will be worse than simple averaging because it will carry a false veneer of scientific rigor.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="the-structured-elicitation-protocol"&gt;The Structured Elicitation Protocol&lt;/h2&gt;
&lt;h3 id="step-by-step-implementation"&gt;Step-by-Step Implementation&lt;/h3&gt;
&lt;p&gt;The structured elicitation protocol reduces biases and improves accuracy through a disciplined process. It treats expert judgments with the same rigor you&amp;rsquo;d apply to operational risk data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 1: Preparation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Identify relevant experts from various disciplines. Note that domain expertise doesn&amp;rsquo;t guarantee unbiased or error-free judgment. Gather relevant information about the problem, including historical data, regulatory context, and comparable cases. Prepare easy-to-understand data presentations. Share information with attendees before the meeting so they arrive informed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Select experts with diverse perspectives. For a GDPR fine estimation, you might include a data protection officer, a legal privacy advisor, a compliance officer, a privacy consultant, a head of compliance, and a head of data governance. Diversity of viewpoint is more valuable than depth in a single perspective.&lt;/p&gt;
&lt;p&gt;Prepare calibration questions with known answers related to the risk domain experts will predict. These questions test each expert&amp;rsquo;s accuracy before you ask them to estimate unknowns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The quality of your calibration questions determines the quality of your entire process. Calibration questions must be from the same domain as the prediction you&amp;rsquo;re asking experts to make, must have objectively verifiable correct answers, must span a range of difficulty levels, and must not be so obvious that every expert gets them right (which provides no differentiation). I typically prepare five to seven calibration questions per session. Three questions is the minimum for meaningful differentiation. Fewer than three doesn&amp;rsquo;t provide enough signal to separate well-calibrated experts from lucky guessers. For the GDPR fine estimation case, calibration questions might ask about the most common fine amount, the 75th percentile fine, and the probability of exceeding a specific threshold, all based on published regulatory data that can be verified.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 2: Workshop Opening&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Explain the workshop objectives and outline the problem structure and key uncertainties. Emphasize that exact probability knowledge isn&amp;rsquo;t required. Highlight how distributions allow for uncertainty expression. Present prepared data and information, encouraging open dialogue about variability and uncertainty. Discuss the logical structure and potential correlations, exploring scenarios that could lead to extreme outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Spend at least 20 minutes on training experts to think in distributions rather than point estimates. Most professionals are trained to give single numbers: &amp;ldquo;the fine will be €100,000.&amp;rdquo; Calibrated estimation requires ranges: &amp;ldquo;I&amp;rsquo;m 90% confident the fine will fall between €30,000 and €400,000.&amp;rdquo; This is a skill that must be taught. Use a simple warm-up exercise: ask experts to estimate something they can verify immediately, like the distance between two cities or the population of a country, as a 90% confidence interval. Then reveal the answer. Most people&amp;rsquo;s first confidence intervals are far too narrow, capturing the true answer less than 50% of the time instead of 90%. This exercise demonstrates overconfidence viscerally and motivates experts to widen their ranges appropriately. Run this exercise at the start of every calibration session.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 3: Workshop Facilitation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Encourage experts to develop their own opinions based on group discussion, giving equal prominence to quiet and dominating experts. Allow time for private consideration and explanation of parameter uncertainty. Emphasize that distributions don&amp;rsquo;t require more knowledge than point estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 4: Individual Estimations&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Conduct one-on-one interviews with each expert using three-point estimates: minimum (best case), most likely, and maximum (worst case). Gather individual estimates without group influence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The three-point estimate captures the expert&amp;rsquo;s uncertainty range. The minimum represents the lowest plausible outcome. The most likely represents the mode of their mental distribution. The maximum represents the highest plausible outcome. These three points can be fitted to a distribution (triangular, PERT, or beta) for further analysis.&lt;/p&gt;
&lt;p&gt;Collect estimates individually to prevent anchoring and conformity bias. Even after a group discussion phase, the actual numerical estimates must be provided privately.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; When collecting three-point estimates, ask for the minimum and maximum first, then the most likely value. If you ask for the most likely value first, experts anchor to it and set their minimum and maximum too close, producing artificially narrow ranges. By asking for extremes first, you force the expert to think about what could go wrong (maximum) and what the best realistic outcome looks like (minimum) before settling on their central estimate. This simple sequencing change consistently produces wider, more realistic ranges. I&amp;rsquo;ve tested both sequences with the same expert groups and the extremes-first approach produces ranges that are 30 to 50% wider, which better reflects genuine uncertainty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 5: Calibration Feedback&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Compare past estimates to actual outcomes to assess biases or patterns. Identify experts who consistently estimate accurately. Identify large differences in expert opinions and reconvene if necessary to discuss discrepancies. Provide feedback on estimation performance and discuss techniques for improving future estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 6: Consensus Building&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Facilitate a discussion to reach a shared understanding of risks and uncertainties, avoiding forced agreement on specific numbers. Summarize key points and insights. Outline next steps for using the gathered information in the risk analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 7: Follow-Up&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Document workshop outcomes and distribute results to participants. Plan for future calibration sessions to track improvement over time. Allow for estimate revisions as new information becomes available.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The follow-up phase is where most organizations drop the ball. They conduct the workshop, produce the aggregated estimate, use it in the risk assessment, and never revisit it. Without follow-up, there&amp;rsquo;s no learning. Schedule a calibration review six months and twelve months after each session. At the review, compare the aggregated estimate to any actual outcomes that have materialized. Update expert calibration scores. Share the results with the experts. Over time, this feedback loop demonstrably improves estimation accuracy. The Good Judgment Project documented that calibration feedback improved forecasting accuracy by 10 to 15% within the first year. Without feedback, accuracy stays flat or degrades. The feedback loop is what transforms expert judgment from a static input into an improving instrument.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="case-study-estimating-gdpr-fines-for-a-spanish-bank"&gt;Case Study: Estimating GDPR Fines for a Spanish Bank&lt;/h2&gt;
&lt;h3 id="step-1-gather-historical-data-for-calibration"&gt;Step 1: Gather Historical Data for Calibration&lt;/h3&gt;
&lt;p&gt;Before asking experts to estimate anything, gather objective data to calibrate their accuracy and provide context.&lt;/p&gt;
&lt;p&gt;For GDPR fines related to processing personal data without legal grounds (Article 6(1)) in Spain over the past two years, the data shows 87 fines ranging from €240 to €1,200,000 with a mean of €72,941, a median of €20,000, and a mode of €10,000 (appearing 8 times). The standard deviation of €154,431 indicates a wide spread. The 25th percentile is €6,000, the 75th percentile is €70,000, and the 90th percentile is €200,000. Banking sector fines tend to be higher: €1,200,000, €200,000, and €70,000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The statistical analysis of historical data serves two purposes. First, it provides the correct answers for calibration questions. Second, it gives experts an empirical foundation for their estimates. Share the summary statistics with experts before the session. Don&amp;rsquo;t hide the data to &amp;ldquo;test&amp;rdquo; their knowledge. The goal isn&amp;rsquo;t to trick experts. It&amp;rsquo;s to produce the most accurate possible estimate of future fines. Informed experts produce better estimates than uninformed ones. However, share the summary statistics, not the raw dataset. Experts who review 87 individual fine records will anchor to memorable outliers. Experts who see percentile distributions develop more balanced mental models. Present the data as distributions and percentiles, not as a list of cases.&lt;/p&gt;
&lt;h3 id="step-2-design-calibration-questions"&gt;Step 2: Design Calibration Questions&lt;/h3&gt;
&lt;p&gt;Prepare calibration questions based on the known statistics. Each question has a correct answer derived from the historical data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Question 1:&lt;/strong&gt; What is the most likely (mode) fine for processing personal data without legal grounds in Spain? Options: €10,000 / €70,000 / €200,000 / €1,200,000. Correct answer: €10,000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Question 2:&lt;/strong&gt; What do you estimate as the 75th percentile fine for GDPR violations related to insufficient legal grounds? Options: €20,000 / €70,000 / €200,000 / €500,000. Correct answer: €70,000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Question 3:&lt;/strong&gt; What is the probability a fine will exceed €200,000 for violating Article 6(1) GDPR? Options: 0-10% / 11-30% / 31-50% / 51-70% / 71-90% / 91-100%. Correct answer: 0-10% (the 90th percentile is €200,000, so approximately 10% of fines exceed this level).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Design calibration questions that test different aspects of the expert&amp;rsquo;s understanding: central tendency (mode or median), distribution shape (percentiles), and tail risk (probability of exceeding a threshold). An expert who correctly identifies the most common fine but overestimates tail risk has a specific bias pattern that the calibration can address. An expert who gets the percentiles right but misidentifies the mode has a different pattern. Three well-designed questions that test different distribution characteristics provide more differentiation than ten questions that all test the same type of knowledge. Also, use multiple-choice format for calibration questions rather than open-ended responses. Open-ended responses are harder to score consistently and create ambiguity about whether a &amp;ldquo;close&amp;rdquo; answer should receive partial credit.&lt;/p&gt;
&lt;h3 id="step-3-collect-expert-responses"&gt;Step 3: Collect Expert Responses&lt;/h3&gt;
&lt;p&gt;Six experts across different roles respond to the three calibration questions. Their responses are compared to the correct answers.&lt;/p&gt;
&lt;p&gt;The Data Processing Officer answers €10,000 (correct), €200,000 (incorrect), 0-10% (correct). The Legal Privacy Advisor answers €70,000 (incorrect), €200,000 (incorrect), 0-10% (correct). The Compliance Officer answers €10,000 (correct), €70,000 (correct), 0-10% (correct). The Privacy Consultant answers €200,000 (incorrect), €500,000 (incorrect), 11-30% (incorrect). The Head of Compliance answers €70,000 (incorrect), €70,000 (correct), 0-10% (correct). The Head of Data Governance answers €70,000 (incorrect), €200,000 (incorrect), 31-50% (incorrect).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Notice that the Compliance Officer scored 100% on calibration questions while the Privacy Consultant and Head of Data Governance scored 0%. This is a common pattern. Domain expertise and seniority don&amp;rsquo;t predict calibration accuracy. The Privacy Consultant may have deep knowledge of privacy law but poor calibration on quantitative estimates. The Head of Data Governance may understand data governance frameworks but have no feel for regulatory penalty distributions. The Cooke method handles this elegantly by assigning zero weight to experts who demonstrate poor calibration, regardless of their title. The hardest part of implementation is presenting these results to the experts themselves. Do it with transparency and respect. Frame it as &amp;ldquo;calibration accuracy for this specific question set&amp;rdquo; rather than &amp;ldquo;you don&amp;rsquo;t know what you&amp;rsquo;re talking about.&amp;rdquo; Calibration scores measure estimation skill, not domain knowledge. A poorly calibrated expert may still contribute valuable qualitative insights during the discussion phase.&lt;/p&gt;
&lt;h3 id="step-4-assign-weights-based-on-calibration-performance"&gt;Step 4: Assign Weights Based on Calibration Performance&lt;/h3&gt;
&lt;p&gt;Score each expert&amp;rsquo;s responses (1 for correct, 0 for incorrect) and calculate calibration weights.&lt;/p&gt;
&lt;p&gt;The Data Processing Officer scores 2 out of 3 (67%), assigned weight 25%. The Legal Privacy Advisor scores 1 out of 3 (33%), assigned weight 12%. The Compliance Officer scores 3 out of 3 (100%), assigned weight 37%. The Privacy Consultant scores 0 out of 3 (0%), assigned weight 0%. The Head of Compliance scores 2 out of 3 (67%), assigned weight 25%. The Head of Data Governance scores 0 out of 3 (0%), assigned weight 0%.&lt;/p&gt;
&lt;p&gt;Assigned weights are calculated by dividing each expert&amp;rsquo;s percentage by the total of all non-zero percentages (267%), producing the final weight distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The weight calculation is simple arithmetic, but its implications are profound. Two of six experts receive zero weight. Their estimates will not influence the final aggregated prediction at all. In a traditional workshop, these two experts would have equal voice with everyone else, potentially pulling the estimate toward their incorrect mental models. The Cooke method eliminates this influence mathematically. When presenting the methodology to stakeholders, emphasize that zero weight doesn&amp;rsquo;t mean the expert&amp;rsquo;s opinion is worthless. It means their quantitative estimation accuracy, as measured by the calibration questions, doesn&amp;rsquo;t support giving their numerical estimates influence over the final aggregate. They can still contribute qualitative context during discussions. But when it comes to the number, calibrated experts drive the result.&lt;/p&gt;
&lt;h3 id="step-5-aggregate-the-weighted-responses"&gt;Step 5: Aggregate the Weighted Responses&lt;/h3&gt;
&lt;p&gt;Multiply each expert&amp;rsquo;s estimate by their assigned weight and sum the results.&lt;/p&gt;
&lt;p&gt;Using the calibration question responses for the most common fine, the aggregated estimate is €32,472. This is significantly lower than a simple average of €71,667 because the two experts who estimated high values (Privacy Consultant at €200,000 and Head of Data Governance at €70,000) received zero weight.&lt;/p&gt;
&lt;p&gt;For a more accurate bank-specific estimate, ask experts to provide a revised estimate for the specific bank scenario. The aggregated bank-specific estimate is €106,236, driven primarily by the Compliance Officer (37% weight, €100,000 estimate) and the Head of Compliance (25% weight, €150,000 estimate).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Always collect both a general estimate and a scenario-specific estimate. The general estimate calibrated against historical data tells you how accurate each expert is at reading the base rate. The scenario-specific estimate applies their judgment to the actual case you care about, weighted by their demonstrated accuracy. The general estimate acts as a sanity check. If the scenario-specific aggregated estimate is dramatically different from the historical base rate, you need to understand why. In this case, the bank-specific estimate of €106,236 is higher than the general most-common estimate of €32,472 because experts appropriately adjusted for the banking sector&amp;rsquo;s higher fine profile. That&amp;rsquo;s a reasonable, explainable deviation. If the bank-specific estimate were €5,000,000, you&amp;rsquo;d need to investigate whether the experts are incorporating genuine sector-specific factors or simply overreacting to headline cases.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="using-ai-as-expert-estimators"&gt;Using AI as Expert Estimators&lt;/h2&gt;
&lt;h3 id="the-method"&gt;The Method&lt;/h3&gt;
&lt;p&gt;Large language models can serve as additional &amp;ldquo;experts&amp;rdquo; in the calibration process. The approach treats each LLM as an independent estimator whose predictions are weighted by demonstrated accuracy, just like human experts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Use multiple LLMs with diverse training data to estimate potential fines or impacts. Develop standardized prompts that provide consistent information about the risk scenario, relevant regulations, historical data, and the required output format. Calibrate LLM outputs using the same Cooke method applied to human experts: test them against known historical data and assign weights based on accuracy. Combine predictions from multiple LLMs using weighted averaging.&lt;/p&gt;
&lt;p&gt;The prompt structure should specify the role the LLM should adopt (such as a Data Protection Officer at a financial institution), the specific regulation and article at issue, the three scenarios to estimate (best case, most common, worst case), the factors to consider (severity, intent, cooperation, mitigation actions), and the requirement to reference historical cases and regulatory guidelines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The prompt design is critical. Inconsistent prompts across LLMs make comparison meaningless. Build a standardized prompt template that you use identically across all models. The template should include the exact same scenario description, the exact same historical context, and the exact same output format requirements. The only variable should be the LLM itself. I structure prompts with four sections: role definition, scenario description with specific regulatory context, action steps specifying the required outputs, and outcome expectations specifying the format and evidence requirements. Test the prompt on one model first to verify it produces the expected output structure. Then deploy it across all models simultaneously.&lt;/p&gt;
&lt;h3 id="calibrating-ai-estimates-against-reality"&gt;Calibrating AI Estimates Against Reality&lt;/h3&gt;
&lt;p&gt;In the GDPR fine case study, five LLMs produced dramatically different estimates for the most common fine.&lt;/p&gt;
&lt;p&gt;Llama estimated €200,000. Claude estimated €400,000. Mistral estimated €3,000,000. Gemini estimated €220,000. GPT-4o estimated €60,000.&lt;/p&gt;
&lt;p&gt;When calibrated against the actual most common fine of €10,000, GPT-4o was closest (still off by a factor of six), while Mistral was off by a factor of 300.&lt;/p&gt;
&lt;p&gt;Using the Cooke method, each LLM&amp;rsquo;s responses were scored against known historical data (best case, most common, worst case). Claude-3.5-sonnet achieved the best calibration (50% assigned weight) because its estimates had the lowest total absolute error percentage. Mistral received 32% weight. Llama received 9%. Gemini received 8%. GPT-4o received only 3% weight despite having the most accurate most-common estimate, because its best-case and worst-case estimates were significantly off.&lt;/p&gt;
&lt;p&gt;The aggregated AI estimate for the bank-specific scenario was €1,182,291, compared to the human expert estimate of €106,236.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; The AI estimates in this case were dramatically higher than human expert estimates, with the AI aggregate more than 10x the human aggregate. This divergence itself is valuable information. It suggests either that LLMs are poorly calibrated for regulatory fine estimation in specific jurisdictions (likely, given their training data includes global cases that may skew distributions upward), or that human experts are underestimating tail risk and the LLMs are capturing something the humans miss, or that the LLMs are anchoring to the maximum possible fine under GDPR (4% of global turnover or €20 million) rather than to actual enforcement patterns in Spain. Don&amp;rsquo;t automatically prefer the human estimate or the AI estimate. Investigate the divergence. In this case, the historical data strongly supports the human estimate range: the actual 90th percentile of Spanish GDPR fines is €200,000, making an aggregate estimate above €1 million an outlier relative to enforcement history. The AI models appear to be poorly calibrated for jurisdiction-specific fine estimation. Document this finding and adjust your methodology accordingly.&lt;/p&gt;
&lt;h3 id="when-to-use-ai-estimators"&gt;When to Use AI Estimators&lt;/h3&gt;
&lt;p&gt;AI estimation is most valuable when you need rapid preliminary estimates across many scenarios before investing in human expert time, when you want to identify the range of plausible outcomes to inform your calibration question design, when you&amp;rsquo;re looking for scenarios or factors that your human experts might not have considered, and when you want to stress-test human estimates by comparing them to an independent source.&lt;/p&gt;
&lt;p&gt;AI estimation is least reliable when jurisdiction-specific enforcement patterns differ significantly from global averages (as in the Spain case), when the scenario involves novel regulatory frameworks with limited enforcement history, when contextual factors (organizational size, cooperation level, remediation speed) heavily influence outcomes, and when you need defensible estimates for regulatory or board reporting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Use AI estimates as one input to your calibration process, not as a replacement for it. Include LLM estimates alongside human expert estimates in your aggregation. Apply the same Cooke method to assign weights based on calibration accuracy. In the Spain GDPR case, the AI estimates would receive low aggregate weight because their calibration accuracy was poor relative to the human experts. In a domain where LLMs demonstrate better calibration, perhaps because there&amp;rsquo;s more training data or less jurisdiction-specific variation, they might receive higher weight. Let the calibration data determine the weighting, not your assumptions about whether humans or machines are &amp;ldquo;better.&amp;rdquo; The Cooke method doesn&amp;rsquo;t care whether the estimator is human or artificial. It cares whether the estimator is accurate.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="building-a-repeatable-calibration-program"&gt;Building a Repeatable Calibration Program&lt;/h2&gt;
&lt;h3 id="institutional-calibration-infrastructure"&gt;Institutional Calibration Infrastructure&lt;/h3&gt;
&lt;p&gt;Individual calibration sessions are valuable. A sustained calibration program is transformative.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Maintain a calibration database that records every expert&amp;rsquo;s estimates, every calibration question and correct answer, every weight assignment, every aggregated result, and every actual outcome when it materializes.&lt;/p&gt;
&lt;p&gt;Track each expert&amp;rsquo;s calibration score over time. Identify experts who are improving (the feedback loop is working) and those who aren&amp;rsquo;t (they may need additional training or should receive lower weights).&lt;/p&gt;
&lt;p&gt;Build a library of calibration questions organized by risk domain: regulatory fines, cybersecurity incidents, operational losses, project overruns, market events. As you accumulate questions with known answers, your calibration testing becomes more robust and differentiated.&lt;/p&gt;
&lt;p&gt;Schedule calibration sessions quarterly for your most critical risk domains. Use shorter calibration exercises (three to five questions) as part of regular risk committee meetings to keep estimation skills sharp.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Measure and report your organization&amp;rsquo;s aggregate calibration improvement over time. If you&amp;rsquo;re running quarterly sessions with calibration feedback, your expert pool&amp;rsquo;s average accuracy should improve measurably within 12 months. Track two metrics. First, the average Brier score across all experts and all questions, which should decrease over time (lower is more accurate). Second, the percentage of experts whose 90% confidence intervals actually contain the true outcome 90% of the time, which should approach 90% from below as calibration training takes effect. Present these metrics to the risk committee as evidence that your risk assessment process is improving in measurable, auditable terms. This is how you move from &amp;ldquo;we think our risk estimates are reasonable&amp;rdquo; to &amp;ldquo;we can demonstrate that our estimation accuracy has improved by X% over the past four quarters.&amp;rdquo; The second statement is what boards and regulators want to hear.&lt;/p&gt;
&lt;h3 id="brier-scores-for-ongoing-accuracy-tracking"&gt;Brier Scores for Ongoing Accuracy Tracking&lt;/h3&gt;
&lt;p&gt;A Brier score measures the accuracy of probabilistic predictions. It ranges from 0 (perfect accuracy) to 1 (complete inaccuracy). For each prediction, the Brier score is calculated as the squared difference between the predicted probability and the actual outcome (1 if the event occurred, 0 if it didn&amp;rsquo;t).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For each expert&amp;rsquo;s probability estimate, record the predicted probability and the actual outcome. Calculate the Brier score for each prediction. Average Brier scores across multiple predictions to get each expert&amp;rsquo;s overall accuracy metric.&lt;/p&gt;
&lt;p&gt;Use Brier scores as an alternative or supplement to the simple correct/incorrect scoring used in the Cooke method. Brier scores capture nuance that binary scoring misses: an expert who assigns 80% probability to an event that occurs is more accurate than one who assigns 51%, even though both would be scored as &amp;ldquo;correct&amp;rdquo; under binary scoring.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Report Brier scores to experts individually and confidentially. Show them how their score compares to the group average without identifying other experts. Competitive benchmarking against an anonymous group average motivates improvement more effectively than abstract accuracy metrics. Frame it as a professional development tool: &amp;ldquo;Your Brier score this quarter was 0.21 versus the group average of 0.18. Here are the questions where your estimates diverged most from outcomes.&amp;rdquo; This is the same feedback mechanism that the Good Judgment Project used to develop superforecasters. It works because it provides specific, measurable, actionable feedback tied to actual outcomes, which is exactly what most professional development programs lack.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="common-implementation-failures-and-how-to-avoid-them"&gt;Common Implementation Failures and How to Avoid Them&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Failure: Skipping calibration and going straight to estimation.&lt;/strong&gt; Without calibration questions, you have no basis for weighting experts. Every expert gets equal weight, which means poorly calibrated experts have as much influence as accurate ones. Always include calibration questions, even if you only have three.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failure: Using the same experts for every assessment.&lt;/strong&gt; Expert fatigue reduces accuracy over time. Rotate experts across sessions. Bring in fresh perspectives. Maintain a pool of qualified experts for each domain rather than relying on the same three people for every risk assessment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failure: Not providing feedback.&lt;/strong&gt; Calibration without feedback is just measurement. Feedback is what drives improvement. Share calibration results with experts after every session. Show them where they were accurate and where they weren&amp;rsquo;t. Discuss techniques for improving (widening confidence intervals, adjusting for known biases, considering base rates before estimating).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failure: Treating the aggregated estimate as a point value.&lt;/strong&gt; The Cooke method produces a weighted point estimate, but the underlying expert distributions contain information about uncertainty. Report the aggregated estimate as a distribution (using the three-point estimates from each expert, weighted by calibration scores) rather than as a single number. A single number implies false precision.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failure: Allowing political override of calibration weights.&lt;/strong&gt; When a senior executive receives zero weight because their calibration accuracy was poor, organizational pressure to &amp;ldquo;adjust&amp;rdquo; the weights is inevitable. Resist this. Document the calibration methodology before the session and commit to applying it without modification. If you allow political overrides, you&amp;rsquo;ve destroyed the method&amp;rsquo;s value and you&amp;rsquo;re back to hierarchy-driven estimation with extra steps.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; Build the calibration methodology into a formal procedure document that your risk committee approves before the first session. The document should specify how calibration questions are selected, how scoring works, how weights are calculated, and that weights are applied mathematically without subjective adjustment. Get this approval once. Then reference it every time someone challenges the weights. The pre-approved procedure document prevents ad hoc political interventions because overriding the weights now requires overriding a committee-approved methodology, which creates its own accountability.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="integrating-calibrated-estimates-into-your-risk-framework"&gt;Integrating Calibrated Estimates Into Your Risk Framework&lt;/h2&gt;
&lt;h3 id="connecting-to-enterprise-risk-management"&gt;Connecting to Enterprise Risk Management&lt;/h3&gt;
&lt;p&gt;Calibrated expert estimates should feed directly into your quantitative risk assessment process, not sit in a separate workstream.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to implement:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Use the three-point estimates from calibrated experts to parameterize loss distributions in your risk models. The weighted minimum, most likely, and maximum values define a PERT or triangular distribution that can be input to Monte Carlo simulations.&lt;/p&gt;
&lt;p&gt;Report calibrated estimates alongside their uncertainty ranges. The board shouldn&amp;rsquo;t see &amp;ldquo;€106,236.&amp;rdquo; They should see &amp;ldquo;€106,236 weighted mean estimate from calibrated experts, with a 90% range of €40,000 to €300,000 based on the distribution of individual estimates.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Track the accuracy of your calibrated estimates against actual outcomes and report the tracking results to the risk committee. This creates a continuous improvement loop that raises confidence in your risk assessment process over time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original implementation tip:&lt;/strong&gt; When presenting calibrated estimates to the board, lead with the methodology&amp;rsquo;s credibility, not just the number. Explain that the estimate comes from X experts whose accuracy was tested against Y calibration questions with known answers, that experts were weighted by demonstrated accuracy, and that the method is based on the Cooke Classical Model used by regulators and international agencies for structured expert judgment. This framing differentiates your estimate from the typical &amp;ldquo;we asked some people and averaged their guesses&amp;rdquo; approach. Boards increasingly expect quantitative rigor in risk assessment. Calibrated expert judgment, properly documented, meets that expectation. Uncalibrated workshop consensus does not.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="key-references"&gt;Key References&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Expert Calibration Methods:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Cooke, R.M. (1991). &amp;ldquo;Experts in Uncertainty: Opinion and Subjective Probability in Science.&amp;rdquo; Oxford University Press.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Tetlock, P.E. (2015). &amp;ldquo;Superforecasting: The Art and Science of Prediction.&amp;rdquo; Crown Publishers.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Kahneman, D. (2011). &amp;ldquo;Thinking, Fast and Slow.&amp;rdquo; Farrar, Straus and Giroux.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Structured Expert Judgment:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;OECD/NRC (2018). &amp;ldquo;Expert Judgement in Risk and Decision Analysis.&amp;rdquo; (Guidance on the Cooke Classical Model)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;European Food Safety Authority (EFSA) guidance on expert knowledge elicitation (2014, updated 2019)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Scoring and Accuracy:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Brier, G.W. (1950). &amp;ldquo;Verification of Forecasts Expressed in Terms of Probability.&amp;rdquo; Monthly Weather Review.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Good Judgment Project documentation (goodjudgment.com)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;AI Risk Estimation:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;NIST AI RMF 1.0 (2023), Measure function&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 23894:2023 (AI Risk Management)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Regulatory Data:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;AEPD (Agencia Española de Protección de Datos) enforcement decisions database&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;GDPR Enforcement Tracker (enforcementtracker.com) for cross-jurisdictional fine data&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act, Regulation (EU) 2024/1689, Article 99 (penalties)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;The organizations that treat expert judgment as data, measure its accuracy, and improve it over time will consistently produce better risk estimates than those relying on unstructured workshops and colorful matrices.&lt;/p&gt;
&lt;p&gt;The math isn&amp;rsquo;t complex. The discipline is. Calibration requires admitting that credentials don&amp;rsquo;t guarantee accuracy, that feedback is essential for improvement, and that mathematical aggregation produces more defensible results than consensus driven by hierarchy.&lt;/p&gt;
&lt;p&gt;The choice between calibrated estimation and uncalibrated guessing is the choice between a risk function that can demonstrate its value quantitatively and one that relies on institutional trust to justify its existence. In an environment where regulators, auditors, and boards increasingly demand evidence, only one of those approaches survives scrutiny.&lt;/p&gt;</description></item><item><title>Quantitative Risk Assessment Using Monte Carlo Simulations and Convolution Methods in R</title><link>https://hwyler.github.io/blog/quantitative-risk-assessment-using-monte-carlo-simulations-and-convolution-methods-in-r/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/quantitative-risk-assessment-using-monte-carlo-simulations-and-convolution-methods-in-r/</guid><description>&lt;h1 id="why-probabilistic-risk-modeling-matters-for-grc-professionals"&gt;Why Probabilistic Risk Modeling Matters for GRC Professionals&lt;/h1&gt;
&lt;p&gt;Picture a risk committee meeting. Someone points at a heat map and says, &amp;ldquo;Vendor concentration risk is High.&amp;rdquo; Twenty minutes of discussion follow. Nobody asks the question that actually matters: how much money are we talking about, and how much should we set aside for it? Nobody can answer it, because a color on a grid was never built to answer it.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the quiet failure at the center of most enterprise risk programs. A 3x3 or 5x5 matrix takes a likelihood rating and an impact rating, both invented on the spot, multiplies them together, and calls the result a risk score. The math doesn&amp;rsquo;t hold up. Ordinal numbers, &amp;ldquo;3&amp;rdquo; for likely, &amp;ldquo;4&amp;rdquo; for severe, aren&amp;rsquo;t real quantities. You can&amp;rsquo;t multiply them any more than you can multiply two zip codes and get a meaningful address. Risk researchers have been pointing this out for close to two decades, and the finding holds up every time someone tests it: matrices routinely rank smaller risks above bigger ones, compress genuinely different exposures into the same box, and give false confidence to numbers nobody can defend in front of a CFO.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s a way out, and it doesn&amp;rsquo;t require a data science degree or a six-figure software license. Monte Carlo simulation lets you describe uncertainty as a probability distribution instead of a guess, run that distribution through tens of thousands of possible futures, and read off a statistically grounded answer. Pair it with convolution, a technique that combines how often something happens with how bad it is when it does, and you get a full loss curve instead of a single number. That curve is what finance teams actually need for reserve setting, capital allocation, and insurance decisions, because it speaks their language: probability and dollars, not colors and adjectives.&lt;/p&gt;
&lt;p&gt;The barrier used to be cost and complexity. Enterprise risk simulation platforms carry real license fees, and statistical programming isn&amp;rsquo;t a skill most GRC professionals picked up in their compliance training. That barrier is mostly gone. An
runs Monte Carlo simulation with convolution in a matter of seconds for 100,000 scenarios, is free to use, and runs in a browser through Google Colab with no local installation at all.&lt;/p&gt;
&lt;p&gt;This guide walks through how the method works, how to set it up, how to choose the right distributions, and how to turn the output into something a board will actually act on.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/chatgpt-image-aug-19-2026-05_56_17-pm.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; A risk matrix gives you a color. Monte Carlo simulation with convolution gives you a probability-weighted range of dollar outcomes you can reserve against, defend to an auditor, and use to price the ROI of a new control. It runs for free, in seconds, in your browser.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="the-two-building-blocks-monte-carlo-simulation-and-convolution"&gt;The Two Building Blocks: Monte Carlo Simulation and Convolution&lt;/h2&gt;
&lt;h3 id="what-monte-carlo-simulation-actually-does"&gt;What Monte Carlo Simulation Actually Does&lt;/h3&gt;
&lt;p&gt;Monte Carlo simulation generates thousands of random scenarios drawn from probability distributions you define for each risk variable. Instead of handing you one &amp;ldquo;expected loss&amp;rdquo; figure, it hands you a full population of possible outcomes, showing you the range, the shape, and how likely each level of loss actually is.&lt;/p&gt;
&lt;p&gt;In practice, you need two inputs for any risk you&amp;rsquo;re modeling:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Frequency&lt;/strong&gt;: how many times the event is likely to happen in a given period. This is a discrete quantity (you can&amp;rsquo;t have 2.3 breaches), so it&amp;rsquo;s typically modeled with a &lt;strong&gt;Poisson distribution&lt;/strong&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Severity&lt;/strong&gt;: how much each event costs when it happens. This is a continuous quantity, and for most operational losses it&amp;rsquo;s modeled with a &lt;strong&gt;lognormal distribution&lt;/strong&gt;, because losses tend to be right-skewed: plenty of small ones, a handful of very large ones.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The simulation then runs thousands of iterations. In each one, it draws a random number of events from the frequency distribution and a random loss amount from the severity distribution, then combines the two. Do that 100,000 times and you have a dataset of possible total losses you can analyze statistically instead of a single guess you have to defend on faith.&lt;/p&gt;
&lt;p&gt;Speed is not a real obstacle here. Ten thousand iterations complete in about half a second, plenty for an exploratory pass or a workshop where you&amp;rsquo;re testing assumptions live. A hundred thousand, the standard for most assessments, finishes in a few seconds. A million, reserved for regulatory capital calculations or board-level reserve recommendations where precision earns its keep, takes well under a minute. The accuracy gain from a hundred thousand to a million runs is marginal for everyday work, so there&amp;rsquo;s no reason to sit through a longer run every time you want to test an assumption during a live session.&lt;/p&gt;
&lt;p&gt;If you want the full quantitative framework behind everything described above, including the complete distribution taxonomy, the open-source Python Monte Carlo engine, and domain-specific applications across AI risk, cyber exposure, compliance debt, and financial risk, &lt;strong&gt;The Risk Management Blueprint&lt;/strong&gt; by me, Hernan Huwyler, builds it chapter by chapter for practitioners who are ready to move past the color grid for good.&lt;/p&gt;
&lt;p&gt;The book covers 26 chapters under one unified probabilistic methodology, with over 70 percent of its pages dedicated to applied quantitative methods rather than governance theory. You can start with the first four chapters for free and decide whether the rest is worth your time before spending a dollar. Preview the first four chapters of The Risk Management Blueprint here:
, or get the full book directly on Amazon at
&lt;/p&gt;
&lt;figure&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/the-risk-management-blueprint-for-quantitative-and-predictive-models-by-hernan-huwyler.jpg?w=683" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;figcaption&gt;
&lt;p&gt;The Risk Management Blueprint for Quantitative and Predictive Models by Hernan Huwyler&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="why-convolution-beats-simple-multiplication"&gt;Why Convolution Beats Simple Multiplication&lt;/h3&gt;
&lt;p&gt;The naive approach to quantifying risk is to take an expected frequency, multiply it by an expected severity, and call that the risk exposure. Four expected events times a $20,000 average loss gives you $80,000. That number is not wrong, exactly. It&amp;rsquo;s just almost useless, because it&amp;rsquo;s a single point with no sense of how much that number could vary, and variation is precisely what a reserve or a capital buffer exists to cover.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Convolution&lt;/strong&gt; is the mathematical operation that properly combines two full probability distributions instead of two single numbers. It preserves the shape of both the frequency distribution and the severity distribution, so the output isn&amp;rsquo;t a point estimate, it&amp;rsquo;s an entire curve. Two risks with the identical expected loss can have very different tail behavior: one might cluster tightly around its average, the other might have a long, thin tail of rare catastrophic outcomes. Simple multiplication treats them as identical. Convolution tells them apart, which is exactly the distinction that matters when you&amp;rsquo;re deciding how much capital to hold against each one.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s a nuance worth flagging here, because it trips people up the first time they run this. If your organization has been using deterministic &amp;ldquo;worst case&amp;rdquo; scenario planning, where someone picks a single pessimistic number and treats it as the ceiling, convolution&amp;rsquo;s output at high percentiles will usually come in lower than that old worst case, because a true worst case assumes the bad outcome happens with certainty, which is almost never realistic. But if your baseline has been simple expected-value multiplication, convolution&amp;rsquo;s tail percentiles will come in noticeably higher than that single center-of-mass number, because a plain average was never designed to show you the tail in the first place; it can&amp;rsquo;t, since it&amp;rsquo;s just one number. Neither of these is a contradiction, and neither is an error in the new model. It&amp;rsquo;s the difference between measuring the middle of a distribution and measuring the whole thing. When you make this switch, document it, and tell your stakeholders plainly: the earlier numbers weren&amp;rsquo;t wrong, they were incomplete, and the shift is a gain in precision, not a change in your risk appetite.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="getting-set-up-two-ways-to-run-this-today"&gt;Getting Set Up: Two Ways to Run This Today&lt;/h2&gt;
&lt;h3 id="google-colab-zero-installation-zero-it-ticket"&gt;Google Colab: Zero Installation, Zero IT Ticket&lt;/h3&gt;
&lt;p&gt;Google Colaboratory gives you a cloud-based notebook that runs R without touching your local machine, which quietly solves the single biggest adoption barrier in most companies: getting IT approval to install anything. Go to
, start a new notebook, switch the runtime to R, paste in the script, and run it cell by cell. You need a Google account and an internet connection. That&amp;rsquo;s the entire prerequisite list.&lt;/p&gt;
&lt;p&gt;One practical wrinkle: Colab sessions time out after inactivity and don&amp;rsquo;t save your data between sessions, so get in the habit of saving your customized script to Google Drive or downloading it locally when you&amp;rsquo;re done for the day. If you&amp;rsquo;re running assessments regularly, it&amp;rsquo;s worth building one template notebook per risk domain, operational, compliance, cyber, with your organization&amp;rsquo;s typical distribution types and parameter ranges already filled in. Customizing a pre-built template for a new assessment takes about five minutes. Building one from a blank notebook takes closer to half an hour. That difference compounds fast once you&amp;rsquo;re running quarterly assessments across a dozen risk categories.&lt;/p&gt;
&lt;p&gt;The full walkthrough, with every code block laid out step by step, is published on
, and the source scripts live in his
, including the convolution model under &lt;code&gt;PythonMinMaxConvMCS&lt;/code&gt; and a compliance-specific variant under &lt;code&gt;PythonTComplianceImpacts&lt;/code&gt;.&lt;/p&gt;
&lt;h3 id="rstudio-for-teams-that-want-this-in-their-workflow"&gt;RStudio: For Teams That Want This in Their Workflow&lt;/h3&gt;
&lt;p&gt;For regular use integrated into an organization&amp;rsquo;s existing tooling, install R locally: download R 4.3.2 or later from
, install RStudio as your development environment, and add the handful of required libraries. R runs cleanly on Windows, macOS, and Linux, and every piece of it, base install and libraries alike, is free and open source.&lt;/p&gt;
&lt;p&gt;If your organization pushes back on installing new software, the cost comparison makes the case for you. Commercial risk simulation platforms with this kind of capability typically run into five figures per user, per year, in enterprise licensing. This script produces statistically equivalent output, mean, median, percentiles, loss exceedance curves, for the specific job of Monte Carlo simulation with convolution, at zero license cost. It won&amp;rsquo;t give a non-technical user a polished GUI, and it doesn&amp;rsquo;t carry the full feature set of a commercial platform. But for the core task, quantifying a loss distribution and setting a defensible reserve, it gets you there. Bring that comparison, along with a quick note on R&amp;rsquo;s open-source licensing, to your procurement conversation.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="configuring-the-model-five-inputs-that-do-all-the-work"&gt;Configuring the Model: Five Inputs That Do All the Work&lt;/h2&gt;
&lt;p&gt;The entire model runs on five parameters, and every one of them should trace back to historical loss data or a properly calibrated expert estimate. None of them should be a number someone typed in because it &amp;ldquo;seemed about right.&amp;rdquo;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Simulations&lt;/strong&gt; — how many scenarios to run. Start at 100,000 for a standard assessment.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Events&lt;/strong&gt; — the expected number of loss events per year, feeding the Poisson distribution. Pull this from your incident log, near-miss records, or a structured expert elicitation if you have no internal data yet.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Loss&lt;/strong&gt; — the expected average financial loss per event, feeding the lognormal distribution.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Mean (Standard Deviation)&lt;/strong&gt; — the spread of losses around that average, expressed as a proportion. A value of 0.2 means losses typically vary by about 20% around the mean; push it to 0.4 and you&amp;rsquo;re describing a much wider, heavier-tailed world.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Reserve&lt;/strong&gt; — the percentile at which you want your reserve set. 0.8 covers 80% of simulated scenarios; 0.95 covers 95%. Your organization&amp;rsquo;s risk appetite statement should be the thing that sets this number, not a habit.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;r&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Simulations &amp;lt;- 100000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Events &amp;lt;- 4
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Loss &amp;lt;- 20000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Mean &amp;lt;- 0.2
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Reserve &amp;lt;- 0.8
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;set.seed(123)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;set.seed(123)&lt;/code&gt; is a small line that does a lot of quiet work. It forces the random number generator to produce the same sequence every time, which means anyone re-running your script with the same seed gets identical results. That&amp;rsquo;s not a nice-to-have. It&amp;rsquo;s what makes the output defensible in an audit trail and reproducible in a peer review, two things a color-coded matrix never had to worry about.&lt;/p&gt;
&lt;p&gt;The standard deviation parameter deserves more attention than it usually gets, because it has an outsized effect on the tail. Moving it from 0.2 to 0.4 doesn&amp;rsquo;t just widen the distribution modestly, it materially increases both the probability and the size of the worst outcomes. Before you commit to a final number, run the model five times with standard deviation values of 0.1, 0.2, 0.3, 0.4, and 0.5, holding everything else fixed, and plot the 95th percentile loss from each run. That sensitivity check takes about five minutes and tells you exactly how much your reserve calculation is riding on an assumption you may not be fully sure of. It&amp;rsquo;s remarkable how often a risk team locks in a round-number standard deviation without ever checking what happens to the output if that number is off by even 10%.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="choosing-the-right-distributions"&gt;Choosing the Right Distributions&lt;/h2&gt;
&lt;p&gt;Getting the shape right matters as much as getting the numbers right. A model built on the wrong distribution will produce confident, precise-looking output that&amp;rsquo;s quietly wrong.&lt;/p&gt;
&lt;h3 id="frequency-the-poisson-distribution"&gt;Frequency: The Poisson Distribution&lt;/h3&gt;
&lt;p&gt;The &lt;strong&gt;Poisson distribution&lt;/strong&gt; models how many times an event occurs in a fixed period, assuming events happen independently and at a roughly constant average rate. It&amp;rsquo;s a solid default for most operational event counts: fraud incidents per year, breaches per quarter, compliance violations per period.&lt;/p&gt;
&lt;p&gt;It works well when you have a reasonable estimate of the average rate, events don&amp;rsquo;t cluster or trigger one another, and the chance of an event in any small window is roughly steady. It stops working well when events cluster (one breach raising the odds of the next), when the rate is visibly trending up or down over time, or when the average frequency climbs above roughly 30 events per period, at which point a normal distribution often fits better.&lt;/p&gt;
&lt;p&gt;Pull the Events parameter from at least three years of incident history if you have it. A single year can be an outlier in either direction. If you logged 2 events last year, 6 the year before, and 3 the year before that, your average is roughly 3.7, and that&amp;rsquo;s the number to use, not last year&amp;rsquo;s count in isolation. When an auditor eventually asks why you assumed 4 events a year, you want a documented, evidence-based answer on hand, not &amp;ldquo;it felt reasonable.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="severity-the-lognormal-distribution"&gt;Severity: The Lognormal Distribution&lt;/h3&gt;
&lt;p&gt;The &lt;strong&gt;lognormal distribution&lt;/strong&gt; models positive-only values with a long right tail: most losses land in a moderate range, but a few run far larger. That pattern shows up consistently across operational, compliance, and cybersecurity losses, which is why lognormal is the default choice for financial impacts, fines, and remediation costs.&lt;/p&gt;
&lt;p&gt;r&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Impact &amp;lt;- rlnorm(n = Simulations, meanlog = log(Loss), sdlog = Mean)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;meanlog = log(Loss)&lt;/code&gt; converts your dollar figure onto the log scale the distribution requires, and &lt;code&gt;sdlog = Mean&lt;/code&gt; controls how wide that distribution spreads.&lt;/p&gt;
&lt;p&gt;Before you trust the choice, check it against your actual data. Plot your historical losses as a histogram. If it&amp;rsquo;s right-skewed with a long tail, lognormal fits. If your losses cluster around two clearly separate values, say, small procedural fines in one cluster and rare, large enforcement actions in another, a single lognormal curve will flatten that pattern into something that isn&amp;rsquo;t really there. In that case, build a mixture of two lognormal distributions, one per cluster, weighted by how often each type occurs. It&amp;rsquo;s a small code change, a handful of lines, and it materially improves the fit for any risk with a genuinely bimodal loss pattern.&lt;/p&gt;
&lt;h3 id="beyond-poisson-and-lognormal"&gt;Beyond Poisson and Lognormal&lt;/h3&gt;
&lt;p&gt;The two defaults cover most operational risk work, but they&amp;rsquo;re not the only tools available, and swapping them in only takes changing one function call:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;rnorm()&lt;/code&gt;&lt;/strong&gt; for a normal distribution, when losses are genuinely symmetric around the average rather than skewed.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;rgamma()&lt;/code&gt;&lt;/strong&gt; for a gamma distribution, when you want more flexible control over skewness than lognormal offers.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;rweibull()&lt;/code&gt;&lt;/strong&gt; for a Weibull distribution, standard in reliability engineering for time-to-failure and equipment breakdown risk.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;runif()&lt;/code&gt;&lt;/strong&gt; for a uniform distribution, when all you genuinely know is a floor and a ceiling with nothing in between.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;rbinom()&lt;/code&gt;&lt;/strong&gt; for a binomial distribution, when you&amp;rsquo;re modeling a fixed number of independent trials, each with the same probability of a &amp;ldquo;bad&amp;rdquo; outcome (for example, the odds that any one of 40 vendors has a material failure this year).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;rnbinom()&lt;/code&gt;&lt;/strong&gt; for a negative binomial distribution, when your frequency data is more erratic than Poisson assumes, some periods clustering with several events, others with none, a pattern statisticians call overdispersion.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Don&amp;rsquo;t pick a distribution because it&amp;rsquo;s the one you remember from a textbook. Pick it because it fits your data, and prove that fit rather than assert it. R&amp;rsquo;s &lt;code&gt;fitdistrplus&lt;/code&gt; library exists for exactly this: run &lt;code&gt;fitdist(your_data, &amp;quot;lnorm&amp;quot;)&lt;/code&gt; and &lt;code&gt;fitdist(your_data, &amp;quot;gamma&amp;quot;)&lt;/code&gt; side by side and compare their AIC (Akaike Information Criterion) scores, where a lower AIC signals a better-fitting model relative to its complexity. Write down the fit statistics along with your choice. &amp;ldquo;We selected lognormal based on goodness-of-fit testing against three years of loss history&amp;rdquo; is a sentence that survives a board meeting or a regulatory exam. &amp;ldquo;We used lognormal because that&amp;rsquo;s what people usually use for operational risk&amp;rdquo; is not.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="inside-the-convolution-engine"&gt;Inside the Convolution Engine&lt;/h2&gt;
&lt;p&gt;Here&amp;rsquo;s what&amp;rsquo;s actually happening under the hood once you hit run. For each of your 100,000 iterations, the script draws one random event count from the Poisson distribution and one random loss amount from the lognormal distribution, then convolves them, mathematically combining the two so the interaction between &amp;ldquo;how many&amp;rdquo; and &amp;ldquo;how much&amp;rdquo; is preserved rather than flattened into an average.&lt;/p&gt;
&lt;p&gt;r&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;combined_distribution &amp;lt;- lapply(1:Simulations, function(i) {
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; conv &amp;lt;- numeric(length(Prob[i]) + length(Impact[i]) - 1)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; for (j in seq_along(Prob[i])) {
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; for (k in seq_along(Impact[i])) {
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; conv[j + k - 1] &amp;lt;- conv[j + k - 1] + Prob[i] * Impact[i]
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; conv
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;})
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;x &amp;lt;- sapply(1:Simulations, function(i) sum(combined_distribution[[i]]))
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The output, &lt;code&gt;x&lt;/code&gt;, is a vector of 100,000 total-loss values, one per simulated scenario. That vector is your aggregate loss distribution, and it&amp;rsquo;s the raw material for every statistic and chart that follows.&lt;/p&gt;
&lt;p&gt;Run the naive calculation alongside it and the difference becomes concrete fast. Simple multiplication of Events × Loss gives 4 × $20,000 = $80,000. In a representative run of the model, the simulated mean lands close to that, around $81,599, which is reassuring; the center of the distribution roughly agrees with the naive estimate. But the 80th percentile comes in at $115,867, about 44% above the mean, and the 95th percentile sits higher still. The simple multiplication gave you the middle of the story. The simulation gives you the whole thing, tails included, and the tails are where the actual risk decisions live. When you present results, show the full distribution, not just the average. The mean tells a committee that everything looks manageable. The 95th percentile tells them what happens on a bad year. Both matter, and leaving either one out of the room is a mistake.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="reading-the-output-like-a-risk-committee-not-a-statistician"&gt;Reading the Output Like a Risk Committee, Not a Statistician&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;summary(x)&lt;/code&gt; hands you the core statistics. Using the illustrative example above, four expected events, a $20,000 average loss, and a 20% standard deviation, a representative run produces something like this:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Statistic&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Minimum&lt;/td&gt;
&lt;td&gt;$0 (scenarios with zero events)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25th Percentile&lt;/td&gt;
&lt;td&gt;$49,383&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median&lt;/td&gt;
&lt;td&gt;$75,715&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mean&lt;/td&gt;
&lt;td&gt;$81,599&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;75th Percentile&lt;/td&gt;
&lt;td&gt;$107,206&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;80th Percentile (Reserve)&lt;/td&gt;
&lt;td&gt;$115,867&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum&lt;/td&gt;
&lt;td&gt;$408,113&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Here&amp;rsquo;s how each of those numbers translates into something a business decision can be built on:&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;median&lt;/strong&gt; is the most typical single outcome, half of all simulated scenarios land below it. The &lt;strong&gt;mean&lt;/strong&gt; sitting above the median confirms the right skew: a handful of high-loss scenarios are pulling the average up above what actually happens most often, which is the standard signature of operational risk data. The &lt;strong&gt;interquartile range&lt;/strong&gt;, roughly $49,000 to $107,000 here, is your &amp;ldquo;normal range,&amp;rdquo; the band your baseline planning should comfortably absorb. The &lt;strong&gt;reserve figure&lt;/strong&gt;, set at your chosen percentile, tells you what you&amp;rsquo;d need to set aside to cover that share of possible outcomes, and by definition leaves the remaining share uncovered; at the 80th percentile, that&amp;rsquo;s a 20% chance actual losses exceed what you&amp;rsquo;ve reserved. The &lt;strong&gt;maximum&lt;/strong&gt; is your single worst simulated draw, low-probability but not zero, and it&amp;rsquo;s the number that should be informing your insurance conversations and catastrophic-loss planning even though you&amp;rsquo;ll never hold a full reserve against it.&lt;/p&gt;
&lt;p&gt;When you report the reserve number, always attach the coverage probability out loud. Don&amp;rsquo;t say &amp;ldquo;the reserve should be $115,867." Say: "A reserve of $115,867 covers 80% of simulated scenarios. There&amp;rsquo;s a 20% chance actual losses exceed that. Covering 95% would require $X instead.&amp;rdquo; Then let the committee choose the coverage level they&amp;rsquo;re comfortable holding capital against. Building a standing reserve table, dollar figures at the 50th, 75th, 80th, 90th, and 95th percentiles, turns this into a menu with clear risk-reward tradeoffs instead of a single number handed down from the model. Setting the reserve is a business decision. The model&amp;rsquo;s job is to lay out the honest options; leadership&amp;rsquo;s job is to pick one and own the tradeoff.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="turning-numbers-into-pictures"&gt;Turning Numbers Into Pictures&lt;/h2&gt;
&lt;h3 id="the-histogram"&gt;The Histogram&lt;/h3&gt;
&lt;p&gt;r&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;hist(x, main = &amp;#34;Histogram of Expected Losses&amp;#34;, xlab = &amp;#34;Total Loss&amp;#34;, ylab = &amp;#34;Frequency&amp;#34;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A histogram shows the shape of your simulated outcomes at a glance: the most common loss range, the right-tail skew stretching toward extreme values, and the overall spread. This is the single most effective way to make the point that risk isn&amp;rsquo;t a number, it&amp;rsquo;s a distribution, to an audience that&amp;rsquo;s used to thinking in single figures.&lt;/p&gt;
&lt;p&gt;For a board deck rather than a technical committee, dress it up a little. Mark the mean and the reserve line explicitly, and color the tail beyond the reserve so the uncovered scenarios are visually obvious rather than buried in the data.&lt;/p&gt;
&lt;p&gt;r&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;hist(x, main = &amp;#34;Distribution of Potential Losses&amp;#34;, xlab = &amp;#34;Total Loss ($)&amp;#34;, col = &amp;#34;lightblue&amp;#34;)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;abline(v = quantile(x, 0.8), col = &amp;#34;red&amp;#34;, lwd = 2)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;abline(v = mean(x), col = &amp;#34;blue&amp;#34;, lwd = 2)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The red line marks your reserve level. The blue line marks the mean. Everything to the right of the red line is the 20% of scenarios your current reserve doesn&amp;rsquo;t cover. One chart like this communicates more about real exposure than a thirty-page qualitative risk report, because it makes the gap visible instead of describing it in adjectives.&lt;/p&gt;
&lt;h3 id="the-loss-exceedance-curve"&gt;The Loss Exceedance Curve&lt;/h3&gt;
&lt;p&gt;r&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;number_sequence &amp;lt;- seq(0.01, 1, by = 0.001)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;y &amp;lt;- sapply(number_sequence, function(i) quantile(x, probs = i))
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;plot(number_sequence, y, type = &amp;#34;l&amp;#34;, xlab = &amp;#34;Percentile&amp;#34;, ylab = &amp;#34;Loss&amp;#34;, main = &amp;#34;Loss Exceedance Curve&amp;#34;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A &lt;strong&gt;loss exceedance curve&lt;/strong&gt; plots the probability of exceeding a given loss threshold across the full distribution, showing exactly how coverage level and required reserve trade off against each other. It&amp;rsquo;s the standard tool for insurance analysis, reserve calibration, and comparing risk tolerance across different scenarios on the same chart.&lt;/p&gt;
&lt;p&gt;This is also where you can put a real dollar figure on the value of a control. Run the model twice, once with your current parameters, once with the parameters you&amp;rsquo;d expect after implementing a proposed control, reduced event frequency, reduced average severity, or both, and overlay the two curves. The gap between them at any percentile is the financial value of that control. That&amp;rsquo;s the calculation behind a sentence like: &amp;ldquo;Implementing this control shifts the 95th percentile loss from $X to $Y, a $Z reduction in potential exposure. The control costs $W. Net return: $Z minus $W.&amp;rdquo; No qualitative matrix produces that sentence. A pair of loss exceedance curves does, directly.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="where-this-gets-used-four-domains-four-playbooks"&gt;Where This Gets Used: Four Domains, Four Playbooks&lt;/h2&gt;
&lt;h3 id="financial-risk"&gt;Financial Risk&lt;/h3&gt;
&lt;p&gt;Model potential losses from market moves, credit defaults, or liquidity events by setting Events to the expected count of adverse events per period and Loss to the average financial impact per event. For credit risk specifically, pull historical default rates and loss-given-default figures to parameterize the model, run it separately by risk grade across your portfolio, and aggregate the results into a portfolio-level credit loss estimate. Compare that against your current loan loss provisions. If your simulated 90th percentile meaningfully exceeds what you&amp;rsquo;re currently holding, you now have a quantitative, defensible basis for recommending an increase, not just a hunch.&lt;/p&gt;
&lt;h3 id="compliance-and-regulatory-risk"&gt;Compliance and Regulatory Risk&lt;/h3&gt;
&lt;p&gt;Estimate potential fines, remediation costs, and enforcement expenses by building a database of enforcement actions in your jurisdiction and industry for the specific regulation in question. Most regulators publish this data. Use it to set your Events parameter (how many enforcement actions per year hit organizations comparable to yours) and your Loss parameter (the average fine size), with the standard deviation pulled from the spread in that same dataset. A compiled set of GDPR enforcement actions against Spanish organizations, for instance, shows an average fine in the tens of thousands of euros but a standard deviation several times larger than the mean, evidence of just how lopsided regulatory penalties actually are, with a handful of large fines pulling the whole distribution far past what a &amp;ldquo;typical&amp;rdquo; fine would suggest. That kind of variability is precisely why lognormal, not a flat average, is the right shape here. Present the output to a compliance committee as: &amp;ldquo;Based on historical enforcement patterns, there&amp;rsquo;s an X% chance a fine exceeding €Y gets imposed. Recommended reserve at the 90th percentile: €Z.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="cybersecurity-risk"&gt;Cybersecurity Risk&lt;/h3&gt;
&lt;p&gt;Set Events to the expected number of breaches, ransomware incidents, or data loss events per year, and Loss to the average all-in cost per incident, response, remediation, notification, legal fees, and business interruption combined. Widely cited industry breach-cost research (annual reports from major cybersecurity and insurance research groups) gives you a reasonable starting point when internal data is thin, but treat those benchmarks as a starting shape, not a final answer. Adjust them for your organization&amp;rsquo;s size, data volume, regulatory footprint, and incident response maturity; a global bank&amp;rsquo;s breach profile and a regional retailer&amp;rsquo;s are not the same distribution wearing different labels. Let external data inform the shape of the curve and your own incident history calibrate its scale.&lt;/p&gt;
&lt;h3 id="operational-and-project-risk"&gt;Operational and Project Risk&lt;/h3&gt;
&lt;p&gt;Apply the same model to equipment failure, supply chain disruption, process breakdowns, or project overruns wherever you can estimate a frequency and a severity. For project risk specifically, it often makes more sense to break the single Loss parameter into separate models for cost overrun, schedule delay, and quality failure, run each one, and combine the output vectors with &lt;code&gt;c()&lt;/code&gt; into a single project-level aggregate. That gives you a picture that respects how differently those three failure modes actually behave instead of flattening them into one generic &amp;ldquo;project risk&amp;rdquo; number.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="back-testing-proving-the-model-isnt-just-precise-looking-fiction"&gt;Back-Testing: Proving the Model Isn&amp;rsquo;t Just Precise-Looking Fiction&lt;/h2&gt;
&lt;p&gt;A model is only worth trusting once it&amp;rsquo;s been checked against reality. &lt;strong&gt;Back-testing&lt;/strong&gt; means comparing what the model predicted against what actually happened, and using the gap to recalibrate.&lt;/p&gt;
&lt;p&gt;After each assessment period, quarterly or annually, record the actual total loss and find where it lands in your simulated distribution. If actual outcomes keep showing up in the extreme tails, above the 95th percentile or below the 5th, the model is miscalibrated somewhere upstream. Track this over time: for a well-calibrated model, roughly 50% of actual outcomes should fall inside the interquartile range, about 90% inside the 90th percentile band, and about 95% inside the 95th. Those aren&amp;rsquo;t arbitrary benchmarks; they&amp;rsquo;re just what &amp;ldquo;calibrated&amp;rdquo; means by definition, so persistent deviation from them is your signal to go back and adjust.&lt;/p&gt;
&lt;p&gt;Keep a running back-testing log: date, risk assessed, the parameters used (Events, Loss, standard deviation), the predicted statistics, and the actual outcome once it materializes. After eight to twelve periods of data, you can calculate real calibration metrics. If actual losses keep exceeding your 80th percentile prediction, you&amp;rsquo;re underestimating risk and need to raise your input parameters. If actuals keep landing below the 25th percentile, you&amp;rsquo;re over-reserving. Bringing back-tested accuracy to a risk committee earns a kind of credibility a brand-new, unproven model simply can&amp;rsquo;t claim yet, and it&amp;rsquo;s the same core validation logic that supervisory guidance on model risk management has long required of financial models, applied here to operational and compliance risk instead of credit models.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="from-model-to-boardroom-reserves-scenarios-and-control-roi"&gt;From Model to Boardroom: Reserves, Scenarios, and Control ROI&lt;/h2&gt;
&lt;h3 id="reserve-setting-and-capital-allocation"&gt;Reserve Setting and Capital Allocation&lt;/h3&gt;
&lt;p&gt;Build a reserve table for each material risk showing the dollar figure at the 50th, 75th, 80th, 90th, and 95th percentiles, and bring it to the risk committee with a recommended confidence level tied to your organization&amp;rsquo;s stated risk appetite, regulatory obligations, and capital position.&lt;/p&gt;
&lt;p&gt;Connect that table directly to the risk appetite statement rather than treating them as separate documents. If the statement says reserves should cover 90% of potential scenarios, the model&amp;rsquo;s 90th percentile output is your target reserve, full stop. If your current reserve sits below that, you&amp;rsquo;ve just converted a vague concern into a specific funding gap: &amp;ldquo;Our stated appetite requires reserves covering 90% of scenarios, which this model puts at $X. Current reserve is $Y. The gap is $X minus $Y.&amp;rdquo; That&amp;rsquo;s a very different conversation from &amp;ldquo;we probably need more reserves,&amp;rdquo; and it&amp;rsquo;s the version that actually gets funded, because it names a number instead of a feeling.&lt;/p&gt;
&lt;p&gt;For portfolio-level aggregation across several material risks, resist the temptation to just add the individual reserves together. Simple addition assumes every risk hits its worst case simultaneously, which overstates the true combined exposure. Either run a joint simulation that accounts for correlation between the risks, or apply a documented diversification factor to the summed total, and explain your reasoning for whichever approach you pick.&lt;/p&gt;
&lt;h3 id="scenario-analysis-and-the-financial-case-for-controls"&gt;Scenario Analysis and the Financial Case for Controls&lt;/h3&gt;
&lt;p&gt;Run the baseline model with today&amp;rsquo;s parameters, then change one input at a time and compare the outputs. What happens to the 80th percentile if event frequency doubles? If average severity rises 50%? If a proposed control cuts frequency from 4 events a year to 2? Document each variant side by side against the baseline so the comparison is visible at a glance, not buried in separate reports.&lt;/p&gt;
&lt;p&gt;This is the mechanism behind quantifying a control&amp;rsquo;s value in dollars rather than adjectives. Run the model once with current parameters and once with the parameters you&amp;rsquo;d expect post-control, then look at how much the reserve requirement shrinks at your chosen percentile. That shrinkage is the control&amp;rsquo;s financial value. Set it against the control&amp;rsquo;s cost and you get a return figure: a $50,000-a-year control that cuts the 90th percentile reserve requirement by $200,000 delivers a 4x return. That reframes the pitch from &amp;ldquo;we should do this because it reduces risk,&amp;rdquo; which is easy to defer, to &amp;ldquo;this delivers a 4x return on investment in reduced reserve requirements,&amp;rdquo; which tends to get approved.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="six-ways-quantitative-models-go-wrong"&gt;Six Ways Quantitative Models Go Wrong&lt;/h2&gt;
&lt;p&gt;Even a well-built simulation fails if you fall into one of these habits:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Using assumed parameters instead of data.&lt;/strong&gt; The model produces confident-looking output regardless of whether the inputs are grounded in evidence or invented on the spot. A simulation built on made-up numbers is just computational fiction with better production values. Document the source and evidence behind every input.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ignoring whether the distribution actually fits.&lt;/strong&gt; Defaulting to lognormal without checking it against your real loss history bakes in a systematic bias. Test the fit whenever you have the data to do it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Reporting only the mean.&lt;/strong&gt; The mean is the least useful number in the whole output for risk decisions. The tails are where decisions actually get made. Always pair the mean with percentile-based statistics.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Running it once and filing the report.&lt;/strong&gt; Risk profiles shift as the business, its controls, and the threat landscape all evolve. Re-run the model quarterly with updated parameters and track how the results move over time.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Skipping &lt;code&gt;set.seed()&lt;/code&gt;.&lt;/strong&gt; Without a fixed seed, every run of the model produces slightly different numbers, which makes runs impossible to compare cleanly and creates an audit trail headache nobody needs. Set it, and record it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Treating the output as a prophecy.&lt;/strong&gt; The model&amp;rsquo;s output is only as good as its inputs and assumptions. Present it as &amp;ldquo;given these assumptions, the model estimates,&amp;rdquo; not &amp;ldquo;the loss will be $X.&amp;rdquo; Uncertainty in, uncertainty out, and a sensitivity analysis is how you show your audience exactly how much of that uncertainty is riding on which assumption.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;One habit worth adding on top of all six: build a documentation template once and reuse it for every assessment, the risk assessed, data sources for each parameter, the distribution chosen and why, the simulation count, the seed, the software and version, the date, the author, the statistics, the sensitivity results, and the back-testing history. Treat it as a model card for your risk simulations. When an auditor asks how you got to a number, you hand them the template instead of reconstructing your reasoning from memory under pressure.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="beyond-r-python-and-where-ai-actually-fits"&gt;Beyond R: Python, and Where AI Actually Fits&lt;/h2&gt;
&lt;p&gt;A refactored Python version of the same methodology lives alongside the R code in the
, including a full convolution build under &lt;code&gt;PythonMinMaxConvMCS&lt;/code&gt;. If your data science team already works in Python, or you want to plug this into an existing machine learning pipeline or a web application, start there instead of forcing an R detour just to match the original methodology. The underlying math is identical regardless of language, and a tool your team already knows and will actually keep using beats a theoretically superior one that quietly falls out of use. If your team already lives in R for statistical work, there&amp;rsquo;s no reason to switch.&lt;/p&gt;
&lt;p&gt;Layering AI and machine learning on top of this foundation is a real and growing extension, not a replacement for it. Predictive models can forecast frequency parameters from leading indicators before they show up in a loss log. Natural language processing can pull structured loss data out of unstructured incident reports to feed the severity distribution automatically. Reinforcement learning can help optimize which combination of controls to fund given a simulated loss curve. But sequence matters here. Prove the basic Monte Carlo model&amp;rsquo;s value first, produce reserve recommendations, back-test them, show they hold up, and only then layer AI capability on top. Organizations that skip straight to AI-driven risk prediction without ever validating a basic quantitative foundation end up with sophisticated-looking output built on assumptions nobody has tested. The simulation is the foundation. AI is refinement on top of it, not a substitute for it.&lt;/p&gt;
&lt;p&gt;For a walkthrough of the same convolution logic built out in Python with a step-by-step presentation format, the
covers the same operational, compliance, and cyber use cases with the Python implementation front and center.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="making-the-switch-getting-your-organization-off-red-yellow-green"&gt;Making the Switch: Getting Your Organization Off Red-Yellow-Green&lt;/h2&gt;
&lt;p&gt;Moving an organization from matrices to probability distributions is a change management project as much as a technical one, and it goes better in stages than as a mandate.&lt;/p&gt;
&lt;p&gt;Start with one risk domain where your historical loss data is strongest, financial risk and cybersecurity usually have the most complete records. Run the model there, produce results, and set them side by side with the previous qualitative assessment. Let the gap speak for itself, especially in the tails and in reserve figures, rather than arguing the case in the abstract.&lt;/p&gt;
&lt;p&gt;Don&amp;rsquo;t rip out every matrix at once. Run the quantitative model in parallel with the existing qualitative process for two or three assessment cycles and let stakeholders watch both outputs land against real outcomes. The case for the quantitative approach tends to make itself once actual losses fall neatly inside the simulated range while sitting outside whatever the old matrix predicted.&lt;/p&gt;
&lt;p&gt;Invest in training. A two-day program covering basic R or Python, probability distributions, and how to interpret statistical output is generally enough to get a risk analyst running and customizing this model on their own. That&amp;rsquo;s a modest investment that pays off across every risk domain you touch afterward, not just the first one.&lt;/p&gt;
&lt;p&gt;The resistance you&amp;rsquo;ll hit is rarely about technical difficulty. It&amp;rsquo;s about the loss of subjective control. A matrix lets a senior risk officer set the rating wherever judgment points. A quantitative model lets the data drive the output, with judgment applied only to the documented, testable inputs. Some people experience that as a loss of influence. It&amp;rsquo;s worth reframing out loud: this is an upgrade in credibility, not a demotion. The risk professional&amp;rsquo;s role shifts from rating things subjectively to choosing the right distribution, interpreting the output, designing the scenarios, and translating the numbers into a business decision, work that commands more respect from finance and the executive table than a colored square ever did. A CFO who has never once acted on a red-yellow-green matrix will engage immediately with a probability-weighted loss curve, because it&amp;rsquo;s the same language they already use for every other financial decision they make.&lt;/p&gt;
&lt;p&gt;This shift also happens to be exactly what frameworks like ISO 31000 and COSO ERM have been asking for all along, quantified risk analysis tied to real decisions, rather than an ordinal scoring exercise that satisfies an audit checkbox and stops there. The method described here doesn&amp;rsquo;t compete with those frameworks. It&amp;rsquo;s how you actually execute the &amp;ldquo;risk analysis&amp;rdquo; step they&amp;rsquo;ve always called for, instead of substituting a color for it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="frequently-asked-questions"&gt;Frequently Asked Questions&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What is convolution in risk management?&lt;/strong&gt; Convolution is the mathematical operation that combines a frequency distribution (how often a risk event happens) with a severity distribution (how large the loss is each time) into a single, full probability distribution of total loss. It preserves the shape of both inputs instead of collapsing them into one averaged number, which is what lets it show the tail risk that simple multiplication misses entirely.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How many Monte Carlo simulations do I actually need?&lt;/strong&gt; Ten thousand iterations are enough for a quick exploratory pass. A hundred thousand is the standard for a full assessment and typically finishes in a few seconds. A million is worth the extra runtime only for high-stakes work like regulatory capital calculations, where the marginal precision gain matters more than the extra wait.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is Monte Carlo simulation actually better than a risk matrix?&lt;/strong&gt; For any decision that requires a dollar figure, reserve setting, capital allocation, insurance purchasing, control ROI, yes, decisively. A matrix can rank risks relative to each other in a rough, ordinal way, but it was never built to answer &amp;ldquo;how much should we reserve,&amp;rdquo; and the math behind multiplying two ordinal scores together doesn&amp;rsquo;t produce a meaningful quantity in the first place.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Which distribution should I use for loss severity?&lt;/strong&gt; Lognormal is the right default for most financial losses, fines, and remediation costs, because it&amp;rsquo;s right-skewed and can&amp;rsquo;t go negative, matching how real losses actually behave. Switch to a mixture of two lognormal curves if your data is genuinely bimodal, to gamma if you need more flexible control over skew, or to a normal distribution only if your losses are genuinely symmetric, which is rare for operational risk.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can I run this without paying for software?&lt;/strong&gt; Yes. The full methodology, in both R and Python, is published as an open-source script that runs for free in Google Colab with no local installation required.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="go-deeper"&gt;Go Deeper&lt;/h2&gt;
&lt;p&gt;For readers who want to run this themselves or dig into the full technical detail behind the method:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;
&lt;/strong&gt; — the full methodology paper, with the mathematics behind combining Poisson frequency and lognormal severity through convolution.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;
&lt;/strong&gt; — every code block from setup to reserve table, explained in sequence.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;
&lt;/strong&gt; — the full R and Python source, including the convolution model and a compliance-specific impact variant.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;
&lt;/strong&gt; — the same framework built out in Python, covering operational, compliance, and cyber risk.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A risk matrix tells a committee that something is &amp;ldquo;High.&amp;rdquo; A Monte Carlo simulation with convolution tells them there&amp;rsquo;s a 20% chance losses exceed $115,867 next year, and that reserving at the 95th percentile instead would cost more but close most of that gap. The first statement starts a conversation. The second one ends with a decision, a dollar figure, and a documented rationale an auditor can actually follow. The tools to make that switch are free, published, and run in under a minute. The only thing left standing in the way is the habit of reaching for the familiar color chart instead.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="key-references"&gt;Key References&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Methodology:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Huwyler, H. (2025). &amp;ldquo;Quantitative Risk Assessment in R: An Open-Source Convolutional Framework for Modeling Uncertainty and Reserves.&amp;rdquo; Quantitative Finance and Risk Management, Volume 10.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cox, A.L. (2008). &amp;ldquo;What&amp;rsquo;s Wrong with Risk Matrices?&amp;rdquo; Risk Analysis, 28(2), 497-512.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Krisper, M. (2021). &amp;ldquo;Problems with Risk Matrices Using Ordinal Scales.&amp;rdquo; arXiv:2103.05440.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Thomas, P., Bratvold, R., Bickel, E. (2014). &amp;ldquo;The Risk of Using Risk Matrices.&amp;rdquo; SPE Economics &amp;amp; Management, 6(2), 56-66.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Monte Carlo Methods:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Ferrero, A. et al. (2023). &amp;ldquo;General Monte-Carlo Approach to Consider a Maximum Admissible Risk in Decision-Making Procedures.&amp;rdquo; Acta IMEKO, 12(4).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Burtescu, E. (2012). &amp;ldquo;Decision Assistance in Risk Assessment: Monte Carlo Simulations.&amp;rdquo; Informatica Economică, 16(4), 86-92.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Young, H.K., Ingall, L. (2009). &amp;ldquo;Exploring Monte Carlo Simulation Applications for Project Management.&amp;rdquo; IEEE Engineering Management Review, 37(2).&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Convolution in Risk Management:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Yam, W.S. (2022). &amp;ldquo;Convolution Approach for Value at Risk Estimation.&amp;rdquo; Review of Pacific Basin Financial Markets and Policies.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Giuseppina Bruno, M., Tomassetti, A. (2006). &amp;ldquo;On the Calculation of Convolution in Actuarial Applications.&amp;rdquo; ACM.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Code Repository:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;GitHub: github.com/hwyler/Paper2024/blob/main/RBaseModel&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Published under open-source license for free use&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Software:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;R: cran.rstudio.com (free, open source)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Google Colaboratory: colab.research.google.com (free, cloud-based)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;The gap between qualitative risk assessment and quantitative risk assessment is not a matter of sophistication. It&amp;rsquo;s a matter of utility. A risk matrix tells you a risk is &amp;ldquo;high.&amp;rdquo; A Monte Carlo simulation tells you there&amp;rsquo;s a 15% probability that losses will exceed $250,000 in the next 12 months and that reserving $180,000 covers 90% of scenarios. The first statement informs a discussion. The second statement informs a decision.&lt;/p&gt;
&lt;p&gt;The tools to make this transition are free, the methodology is published, and the code runs in under five seconds. The only remaining barrier is the willingness to replace familiar but flawed methods with unfamiliar but accurate ones. The organizations that make this transition build risk functions that speak the language of finance, earn board-level credibility, and produce assessments that survive regulatory scrutiny. The ones that don&amp;rsquo;t will continue filling out colorful matrices and wondering why nobody uses them for actual decisions.&lt;/p&gt;</description></item><item><title>The 45 AI Threat Vectors That Your Security Team Probably Isn't Tracking</title><link>https://hwyler.github.io/blog/the-45-ai-threat-vectors-that-your-security-team-probably-isnt-tracking/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/the-45-ai-threat-vectors-that-your-security-team-probably-isnt-tracking/</guid><description>&lt;p&gt;A Practitioner&amp;rsquo;s Field Guide&lt;/p&gt;
&lt;p&gt;Most AI threat models are incomplete. Not slightly incomplete. Fundamentally incomplete.&lt;/p&gt;
&lt;p&gt;Last year I reviewed the threat model for a financial services company deploying a credit decisioning AI. Their security team had identified seven threat vectors. Seven. They covered the obvious ones: data breaches, unauthorized access, denial of service. They missed 38 others, including 15 that were specific to AI systems and had no equivalent in their traditional IT threat catalog.&lt;/p&gt;
&lt;p&gt;Three months after deployment, the model started producing subtly biased outputs. Not because of an external attack. Because a data scientist on the team had inadvertently introduced a feature engineering flaw that created a proxy variable for a protected characteristic. It was a negligence threat, an internal one, and it was not on anyone&amp;rsquo;s radar because the threat model only considered adversarial external actors.&lt;/p&gt;
&lt;p&gt;This pattern repeats across almost every organization I work with. Security teams build threat models based on their experience with traditional systems. They focus on external attackers, malicious intent, and system-level exploits. They miss the internal negligence threats that cause the majority of real-world AI failures. They miss model-specific attack vectors that have no parallel in conventional cybersecurity. They miss the human and organizational threats that create the conditions for technical failures.&lt;/p&gt;
&lt;p&gt;This post catalogs 45 distinct AI threat vectors organized across a two-dimensional taxonomy: intent (adversarial versus negligent) and target category (data, model, system, and human). Each vector includes a clear explanation, practical context, and implementation guidance. Use this as a working reference to audit the completeness of your own AI threat models.&lt;/p&gt;
&lt;h2 id="why-traditional-threat-models-fail-for-ai"&gt;Why Traditional Threat Models Fail for AI&lt;/h2&gt;
&lt;p&gt;Traditional threat modeling frameworks like STRIDE, PASTA, and even MITRE ATT&amp;amp;CK were designed for conventional information systems. They handle network attacks, authentication bypasses, privilege escalation, and data exfiltration well. They were not designed for systems where the &amp;ldquo;logic&amp;rdquo; is learned from data rather than written in code, where the attack surface includes the training pipeline itself, and where some of the most damaging threats come from well-intentioned internal teams making honest mistakes.&lt;/p&gt;
&lt;p&gt;AI systems introduce three categories of threat that traditional frameworks handle poorly.&lt;/p&gt;
&lt;p&gt;First, the model itself is an attack surface. Traditional systems have deterministic logic. If you protect the infrastructure and the data, the system behaves as designed. AI models are different. An attacker can manipulate the model&amp;rsquo;s behavior by carefully crafting inputs, without ever breaching the perimeter or accessing the infrastructure. They can extract sensitive training data by querying the model&amp;rsquo;s API. They can clone the model&amp;rsquo;s functionality through systematic probing. None of these attacks require the kind of infrastructure compromise that traditional threat models focus on.&lt;/p&gt;
&lt;p&gt;Second, the training pipeline is a persistent vulnerability. Traditional systems are vulnerable during operation. AI systems are vulnerable during development. Poisoned training data, biased labels, flawed feature engineering, and compromised pre-trained models all introduce vulnerabilities before the system ever reaches production. By the time the model is deployed, the damage is already embedded in its weights.&lt;/p&gt;
&lt;p&gt;Third, negligence threats cause more cumulative damage than adversarial threats. In traditional cybersecurity, the adversary is the primary concern. In AI, the internal team building and operating the system creates more risk through oversight, insufficient testing, poor documentation, and inadequate monitoring than external attackers do through deliberate exploitation. A threat model that only considers malicious actors misses the majority of the threat landscape.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When I first expanded an organization&amp;rsquo;s AI threat model beyond traditional categories, the security team pushed back hard. &amp;ldquo;We already cover insider threats,&amp;rdquo; they said. They did, but their insider threat model focused on malicious insiders who steal data or sabotage systems. It did not cover the data scientist who chooses features without considering proxy discrimination, the ML engineer who skips robustness testing under deadline pressure, or the architect who fails to build monitoring into the deployment pipeline. These are not insider &amp;ldquo;threats&amp;rdquo; in the traditional security sense. They are negligence risks that require fundamentally different controls. Treat them as separate categories in your threat model, not as subcategories of &amp;ldquo;insider threat.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-taxonomy-structure"&gt;The Taxonomy Structure&lt;/h2&gt;
&lt;p&gt;This threat taxonomy organizes 45 vectors across two dimensions.&lt;/p&gt;
&lt;p&gt;The first dimension is intent. Adversarial threats involve deliberate, intentional actions designed to compromise the AI system. Negligent threats involve unintentional actions or omissions that create vulnerabilities or cause harm. This distinction matters because adversarial and negligent threats require different controls. You defend against adversaries with detection, deterrence, and response. You defend against negligence with process, training, governance, and automation.&lt;/p&gt;
&lt;p&gt;The second dimension is agent. External agents operate outside the organization&amp;rsquo;s boundary, including hackers, competitors, nation-state actors, and third-party vendors. Internal agents operate within the organization, including developers, data scientists, architects, operators, and end users.&lt;/p&gt;
&lt;p&gt;Within each combination of intent and agent, threats target one of four categories: data (the information the system processes and learns from), model (the learned representations, algorithms, and parameters), system (the infrastructure, APIs, and operational environment), and human (the people who build, operate, and interact with the system).&lt;/p&gt;
&lt;p&gt;The result is a comprehensive matrix that surfaces threats most organizations overlook because they fall outside the traditional &amp;ldquo;external attacker targeting our infrastructure&amp;rdquo; frame.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/minimal-workspace-scene.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="adversarial-external-threats-the-ones-you-expect"&gt;Adversarial External Threats: The Ones You Expect&lt;/h2&gt;
&lt;p&gt;This quadrant contains the threats most security teams already think about, plus several AI-specific vectors they likely do not.&lt;/p&gt;
&lt;h3 id="data-threats-from-external-adversaries"&gt;Data Threats from External Adversaries&lt;/h3&gt;
&lt;p&gt;Two vectors target the data layer from outside the organization.&lt;/p&gt;
&lt;p&gt;Data exfiltration occurs when attackers extract sensitive data from the AI system, either during training or inference. This differs from traditional data theft because the AI model itself can leak data. An attacker who gains access to model outputs, gradients, or confidence scores can sometimes reconstruct training data without ever accessing the database directly. The model becomes an unintentional data disclosure channel.&lt;/p&gt;
&lt;p&gt;Data poisoning occurs when attackers introduce malicious or corrupted data into the training dataset. This is uniquely dangerous because the corruption happens before the model is deployed. If an attacker can influence any data source that feeds into the training pipeline, whether through compromised public datasets, manipulated web scraping sources, or tampered third-party data feeds, they can embed biases or backdoors that persist through every subsequent version of the model until the poisoned data is identified and removed.&lt;/p&gt;
&lt;p&gt;Data poisoning is the adversarial data threat I worry about most because it is the hardest to detect and the longest-lasting in impact. Traditional data validation checks look for obvious anomalies: missing values, out-of-range entries, format errors. Poisoned data is designed to look normal. The individual data points are plausible. The corruption is statistical, not syntactic. Defending against it requires comparing training data distributions across time windows to detect subtle shifts, validating training data provenance to verify that sources have not been compromised, and testing model behavior on held-out validation sets from trusted sources. If your training data comes from any source you do not fully control, you have exposure to this vector.&lt;/p&gt;
&lt;h3 id="model-threats-from-external-adversaries"&gt;Model Threats from External Adversaries&lt;/h3&gt;
&lt;p&gt;Six vectors target the model itself, and these are the threats most traditional security teams have the least experience with.&lt;/p&gt;
&lt;p&gt;Deception involves crafting inputs designed to make the AI model produce incorrect predictions or classifications. Think of carefully modified images that cause a computer vision system to misclassify objects, or subtly altered text inputs that cause a natural language model to produce wrong outputs. The modifications are often imperceptible to humans but effective against the model.&lt;/p&gt;
&lt;p&gt;Evasion is deception&amp;rsquo;s cousin, focused specifically on bypassing detection or classification mechanisms. An attacker modifies malicious inputs to evade an AI-based security system, such as altering malware signatures to bypass AI-powered threat detection or modifying fraudulent transactions to avoid AI-based fraud scoring.&lt;/p&gt;
&lt;p&gt;Exploitation targets weaknesses in the AI system&amp;rsquo;s implementation rather than the model&amp;rsquo;s learned behavior. Insecure APIs, poor input validation, missing authentication, and misconfigured endpoints are the entry points. This vector bridges traditional cybersecurity and AI-specific risk because the vulnerabilities are conventional but the assets being targeted (models, training pipelines, inference endpoints) are AI-specific.&lt;/p&gt;
&lt;p&gt;Inversion attacks reverse-engineer the model to infer sensitive information about training data. By carefully analyzing the model&amp;rsquo;s outputs across many queries, an attacker can deduce characteristics of the data the model was trained on. For models trained on medical records, financial data, or personal information, this vector directly threatens privacy even if the underlying database is perfectly secured.&lt;/p&gt;
&lt;p&gt;Membership inference is related but distinct. Instead of reconstructing training data, the attacker determines whether a specific known data point was part of the training set. &amp;ldquo;Was this person&amp;rsquo;s medical record used to train your diagnostic model?&amp;rdquo; is a question that, if answerable through the model&amp;rsquo;s API, creates privacy and compliance exposure regardless of whether the attacker can reconstruct the full record.&lt;/p&gt;
&lt;p&gt;Oracle attacks (also called model extraction or model stealing) involve an attacker querying the model extensively to understand its decision boundaries, then building a replica model that reproduces the original&amp;rsquo;s behavior without access to the training data. The attacker essentially steals the intellectual property embedded in the model through its public-facing API.&lt;/p&gt;
&lt;p&gt;Transfer learning attacks exploit vulnerabilities in pre-trained models or the transfer learning process itself. Many organizations build their AI systems on top of pre-trained foundation models. If the foundation model contains embedded vulnerabilities, biases, or backdoors, every downstream model inherits them. This vector is growing in importance as more organizations build on third-party foundation models they did not train and cannot fully audit.&lt;/p&gt;
&lt;p&gt;Of these six model-level threats, oracle attacks and transfer learning attacks are the two most underestimated. Oracle attacks are underestimated because organizations assume their model&amp;rsquo;s logic is protected by keeping the code proprietary. It is not. If the model is accessible through an API, its behavior can be replicated through systematic querying. Rate limiting helps but does not eliminate the risk. Transfer learning attacks are underestimated because organizations treat pre-trained foundation models as trusted components without verifying what is in them. I worked with a team that fine-tuned a publicly available language model for customer service automation. Nobody audited the base model for embedded biases or vulnerabilities. The assumption was &amp;ldquo;it&amp;rsquo;s from a reputable provider, so it&amp;rsquo;s safe.&amp;rdquo; That assumption is not supportable. If you use pre-trained models, document your trust assumptions about those models explicitly, test for bias and adversarial vulnerability in the fine-tuned model, and include the base model in your threat surface.&lt;/p&gt;
&lt;h3 id="system-threats-from-external-adversaries"&gt;System Threats from External Adversaries&lt;/h3&gt;
&lt;p&gt;Eight vectors target the infrastructure and operational environment.&lt;/p&gt;
&lt;p&gt;Advanced persistent threats involve state actors or organized crime groups using sophisticated tools to compromise the AI system over an extended period. These are the most resource-intensive attacks and typically target high-value AI systems in critical infrastructure, financial services, or defense applications.&lt;/p&gt;
&lt;p&gt;API-based attacks exploit vulnerabilities in the APIs used to interact with the AI system. This includes injecting malicious data through API endpoints, exploiting authentication weaknesses, or abusing API rate limits to conduct model extraction.&lt;/p&gt;
&lt;p&gt;Denial of service overwhelms the AI system with excessive traffic or resource demands. AI systems can be particularly vulnerable because inference workloads on complex models consume significant computational resources, making resource exhaustion attacks more efficient than against simpler services.&lt;/p&gt;
&lt;p&gt;Model freezing attacks target the AI system&amp;rsquo;s update mechanism, preventing the model from receiving updates. A model that cannot be updated remains vulnerable to known threats and cannot adapt to data drift, effectively turning a dynamic system into a static one.&lt;/p&gt;
&lt;p&gt;Model parameter poisoning introduces malicious updates or perturbations directly into the model&amp;rsquo;s parameters (weights, biases) rather than through training data. This vector is particularly relevant for federated learning systems where multiple parties contribute model updates.&lt;/p&gt;
&lt;p&gt;Poorly designed APIs represent a threat vector that straddles adversarial and negligent categories. While the design flaw is internal, external attackers exploit it. Insecure API design that exposes model internals, lacks input validation, or provides excessive information in error messages creates the conditions for multiple other attacks.&lt;/p&gt;
&lt;p&gt;Side-channel attacks exploit non-functional characteristics of the AI system, such as timing differences in inference responses, power consumption patterns during computation, or electromagnetic emissions, to extract information about the model or its data. These attacks do not target the model&amp;rsquo;s logic directly but extract information through observable physical or computational characteristics.&lt;/p&gt;
&lt;p&gt;Supply chain compromise involves attackers tampering with hardware, software components, or third-party services in the AI system&amp;rsquo;s supply chain. Compromised ML libraries, poisoned pre-trained models distributed through public repositories, or tampered GPU firmware all fall into this category.&lt;/p&gt;
&lt;p&gt;Supply chain compromise is the system-level external threat I spent the most time helping organizations address last year. The AI supply chain is broader and less controlled than most organizations realize. A typical ML pipeline might include open-source libraries (PyTorch, TensorFlow, scikit-learn), pre-trained models from public repositories (Hugging Face, GitHub), data from third-party providers, cloud services for training and inference, and container images from public registries. Each component is a potential entry point. The most practical defense is maintaining a software bill of materials (SBOM) for your AI systems that includes not just code dependencies but also model provenance, training data sources, and infrastructure components. When a vulnerability is discovered in any component, the SBOM tells you immediately which AI systems are affected. Without it, you are guessing.&lt;/p&gt;
&lt;h3 id="human-threats-from-external-adversaries"&gt;Human Threats from External Adversaries&lt;/h3&gt;
&lt;p&gt;One vector targets the human element from outside.&lt;/p&gt;
&lt;p&gt;Social engineering involves attackers manipulating developers, data scientists, or users into revealing sensitive information or interacting with the AI system in ways that compromise security. AI teams are often targeted because they have access to valuable IP (models, training data, feature engineering pipelines) and may not have received the same security awareness training as traditional IT staff. A data scientist who shares a model architecture diagram on a conference poster may not realize they have disclosed information useful for a model extraction attack.&lt;/p&gt;
&lt;h2 id="adversarial-internal-threats-the-ones-you-underestimate"&gt;Adversarial Internal Threats: The Ones You Underestimate&lt;/h2&gt;
&lt;p&gt;This quadrant contains only two vectors, but both are high-impact.&lt;/p&gt;
&lt;h3 id="human-threats-from-internal-adversaries"&gt;Human Threats from Internal Adversaries&lt;/h3&gt;
&lt;p&gt;Data sabotage occurs when insiders intentionally alter or destroy data used for training or inference. Unlike external data poisoning, an insider has legitimate access to data systems and can make changes that appear routine. A disgruntled data engineer who subtly modifies preprocessing scripts or alters label distributions can compromise model performance in ways that are extremely difficult to trace.&lt;/p&gt;
&lt;p&gt;Subversion involves authorized developers or contractors intentionally sabotaging, exfiltrating, or manipulating the AI system. This goes beyond data sabotage to include embedding backdoors in model code, exfiltrating trained model weights for competitors, or introducing vulnerabilities that can later be exploited. The insider&amp;rsquo;s authorized access makes traditional perimeter defenses irrelevant.&lt;/p&gt;
&lt;p&gt;Internal adversarial threats against AI systems are harder to detect than their equivalents in traditional IT for one specific reason: the normal behavior of an AI developer already includes activities that would be flagged as suspicious in other contexts. A data scientist routinely downloads large datasets, modifies data processing logic, changes model parameters, and deploys updated models. These are their job functions. Distinguishing between a legitimate model update and a sabotage event requires understanding what the model should be doing, not just what the developer is doing. The most effective control I have found is mandatory peer review for all changes to training data, feature engineering code, and model parameters before they reach production. Not automated testing, though that helps too. Human review by a second qualified person who can assess whether the change makes sense in context. This catches both intentional sabotage and unintentional errors.&lt;/p&gt;
&lt;h2 id="negligent-external-threats-your-vendors-and-dependencies"&gt;Negligent External Threats: Your Vendors and Dependencies&lt;/h2&gt;
&lt;p&gt;This quadrant covers unintentional risks introduced by parties outside your organization.&lt;/p&gt;
&lt;h3 id="human-threats-from-external-negligence"&gt;Human Threats from External Negligence&lt;/h3&gt;
&lt;p&gt;Supply chain negligence occurs when third-party vendors introduce vulnerabilities through insecure libraries, dependencies, or tools used during AI system development, deployment, or maintenance. Unlike supply chain compromise (which is intentional), this vector reflects genuine negligence: a vendor fails to patch a library, releases an update with a security flaw, or provides tooling that does not meet security standards. The impact on your AI system is the same whether the vulnerability was introduced deliberately or carelessly.&lt;/p&gt;
&lt;p&gt;Third-party data risk arises when organizations rely on external data sources that may be of poor quality, biased, or inadvertently altered. The third party is not acting maliciously. They simply do not maintain the data quality standards your model requires. Training on degraded external data produces degraded model performance, and the organization consuming the data may not detect the quality decline until outputs start failing.&lt;/p&gt;
&lt;h3 id="system-threats-from-external-negligence"&gt;System Threats from External Negligence&lt;/h3&gt;
&lt;p&gt;Outdated dependencies represent the use of unsupported software libraries in AI systems that introduce known, exploitable vulnerabilities. This is technically a negligence issue, an external provider stops maintaining a library, but the security impact is the same as an adversarial exploit because attackers actively scan for systems using deprecated dependencies.&lt;/p&gt;
&lt;p&gt;Third-party data risk is the negligent external threat I encounter most frequently in practice. Organizations build models on external data feeds and assume the data quality will remain stable. It does not. I worked with a company whose fraud detection model degraded over four months because a third-party transaction data provider changed their data formatting without notification. The change was minor, a modification to how categorical fields were encoded, but it silently corrupted the feature engineering pipeline. The model&amp;rsquo;s accuracy dropped from 91% to 78% before anyone noticed. The fix was straightforward but the damage was done. For every external data dependency, establish a data quality SLA with the provider that specifies format, completeness, timeliness, and quality metrics. Monitor incoming data against those SLAs automatically. When a deviation occurs, alert before the data enters your training pipeline, not after your model degrades.&lt;/p&gt;
&lt;h2 id="negligent-internal-threats-where-most-ai-failures-actually-originate"&gt;Negligent Internal Threats: Where Most AI Failures Actually Originate&lt;/h2&gt;
&lt;p&gt;This is the largest quadrant in the taxonomy, containing 24 of the 45 vectors. That distribution is not an accident. It reflects reality. The majority of AI failures in production stem from internal negligence, not external attacks.&lt;/p&gt;
&lt;h3 id="data-threats-from-internal-negligence"&gt;Data Threats from Internal Negligence&lt;/h3&gt;
&lt;p&gt;Two vectors target the data layer through internal oversight.&lt;/p&gt;
&lt;p&gt;Inaccurate data labeling occurs when developers fail to properly label training data during preparation. Labels are the ground truth the model learns from. If a medical imaging dataset contains mislabeled scans, the model learns to associate the wrong visual patterns with the wrong diagnoses. Unlike external data poisoning, this is not malicious. It is the predictable result of insufficient quality assurance in the annotation process, often driven by time pressure, undertrained annotators, or ambiguous labeling guidelines.&lt;/p&gt;
&lt;p&gt;Bias in data occurs when data scientists or developers incorporate biased data during training, producing discriminatory or unfair outputs. This can result from historical bias embedded in the data itself (past lending decisions that reflected discriminatory practices), selection bias in how data was collected (underrepresenting certain populations), or measurement bias in how variables were recorded. The developers are not trying to create discriminatory outcomes. They are training on data that reflects existing inequities.&lt;/p&gt;
&lt;p&gt;Bias in data is the negligent data threat with the highest regulatory and reputational impact. It is also the one where I see the most dangerous misconception. Teams believe that removing protected characteristics like race or gender from the training data eliminates bias. It does not. Other variables in the dataset frequently serve as proxies for protected characteristics. Zip code correlates with race. Job title correlates with gender. Part-time employment status correlates with caregiving responsibilities. Removing the protected characteristic while leaving the proxy variables in the feature set gives the appearance of fairness while producing the same discriminatory outcomes. The effective control is to test model outputs for disparate impact across protected groups, regardless of whether protected characteristics appear in the input features. Test the outputs, not the inputs.&lt;/p&gt;
&lt;h3 id="model-threats-from-internal-negligence"&gt;Model Threats from Internal Negligence&lt;/h3&gt;
&lt;p&gt;Five vectors target the model through internal oversight.&lt;/p&gt;
&lt;p&gt;Data and model drift occurs when developers fail to account for changes in data distribution or underlying concepts over time. A model trained on data from 2022 may not perform well on data from 2025 if customer behavior, market conditions, or the relationships between variables have shifted. This is not a one-time risk. It is an ongoing degradation that accelerates the longer a model operates without retraining or recalibration.&lt;/p&gt;
&lt;p&gt;Feature engineering flaws result from unintentionally introducing vulnerabilities through poor feature selection. Selecting features that are easily manipulated by adversaries, including features that leak future information (data leakage), or failing to consider how feature distributions might shift in production all fall into this category.&lt;/p&gt;
&lt;p&gt;Overfitting occurs when developers create models that fit too closely to the training data, learning noise and idiosyncrasies rather than genuine patterns. An overfit model shows excellent performance in testing and poor performance in production. From a security perspective, an overfit model is also more predictable to an adversary who understands its training data, making it easier to craft adversarial inputs.&lt;/p&gt;
&lt;p&gt;Overfitting to noise is a specific variant where the model learns from irrelevant data rather than meaningful signal. The model performs well on training metrics but produces unreliable results in real-world deployment because it has memorized artifacts in the training data rather than learning the underlying relationship.&lt;/p&gt;
&lt;p&gt;Unexplainability results when developers create AI systems too complex and opaque to understand or explain. This is a threat vector because opacity prevents detection of errors, biases, and vulnerabilities. A model whose decisions cannot be explained cannot be audited, cannot be debugged when it fails, and cannot satisfy regulatory requirements for explainability. Opacity does not cause harm directly, but it creates the conditions under which every other threat vector becomes harder to detect and address.&lt;/p&gt;
&lt;p&gt;Data and model drift is the negligent model threat that causes the most cumulative financial damage because it is slow, silent, and continuous. I have never worked with an organization that detected drift proactively on their first AI deployment. They always detected it reactively, after business outcomes degraded enough for someone to notice. The detection lag ranged from three weeks to nine months depending on how closely business stakeholders monitored the model&amp;rsquo;s downstream effects. The fix is statistical monitoring of input feature distributions and output prediction distributions, compared against baseline distributions from the training period. When statistical tests detect a significant shift, trigger an alert. Do not wait for business outcome metrics to degrade, because by then the model has been making suboptimal decisions for weeks or months. Population Stability Index (PSI) is a good starting metric. Monitor it weekly at minimum for production models.&lt;/p&gt;
&lt;h3 id="human-threats-from-internal-negligence"&gt;Human Threats from Internal Negligence&lt;/h3&gt;
&lt;p&gt;This is the richest subcategory in the entire taxonomy, containing 12 vectors. Each represents a different way that well-intentioned people create AI risk through oversight, insufficient skill, or organizational failure.&lt;/p&gt;
&lt;p&gt;Inadequate documentation occurs when teams provide insufficient documentation for AI models, data sources, and decision-making processes. Without documentation, models cannot be maintained by anyone other than their original developer, security reviews cannot assess the system&amp;rsquo;s design assumptions, and regulatory compliance cannot be demonstrated.&lt;/p&gt;
&lt;p&gt;Inadequate monitoring results from failing to build proper monitoring mechanisms to detect anomalies, adversarial activity, or model degradation in real time. An unmonitored model is a model whose failures go undetected until they manifest as business losses, customer complaints, or regulatory actions.&lt;/p&gt;
&lt;p&gt;Inadequate maintenance occurs when teams fail to regularly update and maintain AI models, leaving them vulnerable to known attacks, data drift, and exploits as the system ages. Models, like all software, require ongoing maintenance. Unlike traditional software, model maintenance includes retraining, recalibration, and feature re-evaluation in addition to patching.&lt;/p&gt;
&lt;p&gt;Inadequate testing results from failing to sufficiently test and validate models before deployment. This includes insufficient unit testing of data pipelines, absence of adversarial robustness testing, incomplete validation against held-out datasets, and failure to test for bias and fairness.&lt;/p&gt;
&lt;p&gt;Inadequate training affects end-users and operators who receive insufficient instruction on how to interact with or manage AI systems. A model that is technically sound can still produce harmful outcomes if the humans using it do not understand its limitations, do not know when to override its recommendations, or cannot recognize when its outputs are unreliable.&lt;/p&gt;
&lt;p&gt;Insecure design results from architects or developers failing to build secure AI systems from the start. This includes objective functions susceptible to manipulation, model architectures that leak information through their outputs, and deployment configurations that expose internal model details.&lt;/p&gt;
&lt;p&gt;Insider threat (unintentional) covers authorized team members who inadvertently cause harm. A developer who accidentally pushes a model trained on test data to production. A data engineer who modifies a preprocessing script that breaks feature normalization. An operations team member who changes a configuration parameter without understanding its downstream effects.&lt;/p&gt;
&lt;p&gt;Insufficient access control results from poorly managed access permissions that allow unauthorized access to models, data, or systems. This is a process failure, not a technology failure. The tools to enforce access control exist. The organization simply has not implemented them with sufficient rigor for AI-specific assets.&lt;/p&gt;
&lt;p&gt;Lack of governance occurs when organizations fail to establish proper governance frameworks for AI development and deployment. Without governance, teams operate independently, security practices are inconsistent, accountability is undefined, and risk accumulates without visibility.&lt;/p&gt;
&lt;p&gt;Over-reliance on AI results from decision-makers depending on AI outputs without sufficient human oversight or fail-safes. When users treat AI predictions as infallible, they stop applying the human judgment that catches model errors. This vector is particularly dangerous in high-stakes domains like healthcare, criminal justice, and financial services.&lt;/p&gt;
&lt;p&gt;Unclear AI accountability occurs when nobody is clearly defined as accountable for AI-related decisions, model performance, or security. When accountability is unclear, risks go unmanaged because everyone assumes someone else is responsible.&lt;/p&gt;
&lt;p&gt;Of these 12 human negligence vectors, inadequate monitoring and unclear accountability are the two that create the most cascading damage. They amplify every other threat in the taxonomy. An adversarial attack against an inadequately monitored system succeeds for longer. Data drift in a system with no accountable owner goes unaddressed for months. Bias in a model that nobody monitors against fairness metrics persists indefinitely. If you can only address two human negligence vectors immediately, address these two. Assign a named individual, not a team, as accountable for each production AI system. Then build monitoring that runs at the same speed as the model&amp;rsquo;s decision-making. Everything else becomes more manageable once you have visibility and ownership in place.&lt;/p&gt;
&lt;h3 id="system-threats-from-internal-negligence"&gt;System Threats from Internal Negligence&lt;/h3&gt;
&lt;p&gt;Five vectors target the system infrastructure through internal oversight.&lt;/p&gt;
&lt;p&gt;Inadequate incident response results from failing to develop or implement effective response plans for AI-specific incidents. AI incidents differ from traditional IT incidents. A model producing biased outputs is not a &amp;ldquo;system down&amp;rdquo; event. It does not trigger the same alerts. It requires different diagnostic procedures and different remediation steps. If your incident response playbook does not include AI-specific scenarios, your response will be improvised when it matters most.&lt;/p&gt;
&lt;p&gt;Inadequate logging results from insufficient logging and monitoring infrastructure that makes it difficult to detect, investigate, or respond to incidents. If you cannot see what the model received as input, what it produced as output, and how its behavior has changed over time, you cannot diagnose problems or provide evidence for regulatory investigations.&lt;/p&gt;
&lt;p&gt;Insecure data storage results from failing to implement secure storage for AI-specific data assets. Training data, model weights, feature engineering code, and hyperparameter configurations all contain sensitive intellectual property and potentially personal data. Storing them with the same (or lesser) security controls as general-purpose data creates exposure.&lt;/p&gt;
&lt;p&gt;Insufficient redundancy results from failing to build AI systems with adequate fail-safes. A single point of failure in the model serving infrastructure, the data pipeline, or the monitoring system can take down the entire AI capability. AI systems often have complex dependency chains that create hidden single points of failure.&lt;/p&gt;
&lt;p&gt;Misconfiguration results from incorrectly configuring security settings, environments, or system components. Cloud environment misconfigurations, exposed model endpoints, overly permissive IAM roles, and unencrypted data storage are the most common variants.&lt;/p&gt;
&lt;p&gt;Inadequate incident response for AI systems is the system-level negligence threat I have spent the most time remediating. Traditional incident response plans categorize incidents by severity and system type, but they almost never include AI-specific incident categories. What do you do when a model starts producing outputs that are statistically different from its validation period behavior? What do you do when a fairness audit reveals disparate impact? What do you do when you discover that training data was contaminated three months ago and every model version since then is potentially compromised? These are AI incidents that require AI-specific response procedures. Build an AI incident response appendix for your existing plan. Include at minimum: model rollback procedures, retraining triggers, bias investigation protocols, adversarial attack containment steps, and stakeholder notification procedures. Then tabletop exercise these scenarios. The first time you run through an AI incident scenario, you will discover gaps in your response capability that are fixable before a real incident occurs.&lt;/p&gt;
&lt;h2 id="cross-cutting-implementation-tips"&gt;Cross-Cutting Implementation Tips&lt;/h2&gt;
&lt;p&gt;These apply across all four quadrants of the threat taxonomy.&lt;/p&gt;
&lt;p&gt;Assess your taxonomy coverage quarterly. AI threat vectors evolve as attack research advances, new model architectures emerge, and regulatory requirements expand. A taxonomy that was comprehensive six months ago may have gaps today. Assign someone to monitor AI security research (MITRE ATLAS updates, conference proceedings from NeurIPS and USENIX Security, regulatory guidance from NIST and the EU AI Office) and flag new vectors that should be added.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When reviewing your taxonomy coverage, do not just ask &amp;ldquo;have any new threat vectors emerged?&amp;rdquo; Also ask &amp;ldquo;have any of our existing vectors changed in severity or likelihood?&amp;rdquo; The relative importance of threat vectors shifts as your AI systems mature. Early in deployment, development negligence threats (inadequate testing, insecure design) are most relevant because the system is new and untested. Six months into production, operational negligence threats (inadequate monitoring, data drift, inadequate maintenance) become dominant. Twelve months in, adversarial threats increase as your AI system becomes a known, valuable target. Reassess priority rankings at each quarterly review.&lt;/p&gt;
&lt;p&gt;Balance adversarial and negligence controls in your budget. Security teams naturally gravitate toward adversarial controls because they are more dramatic and more familiar. Adversarial robustness testing, penetration testing, red-teaming: these feel like &amp;ldquo;real&amp;rdquo; security work. Process controls, governance frameworks, training programs, and documentation standards feel like bureaucracy. In practice, the negligence controls prevent more incidents.&lt;/p&gt;
&lt;p&gt;Original implementation tip: I track a simple metric with every organization I advise: the ratio of AI incidents caused by adversarial action versus negligence. Across 14 organizations over three years, the ratio has consistently been approximately 15% adversarial and 85% negligence. Yet budget allocation for adversarial controls versus negligence controls is typically inverted: 60-70% on adversarial defenses and 30-40% on process and governance. Reallocate to match the actual threat distribution. This does not mean reducing adversarial defenses. It means increasing investment in monitoring, documentation, testing processes, training, and governance until the budget reflects where incidents actually originate.&lt;/p&gt;
&lt;p&gt;Map each vector to specific assets and controls. A threat vector without a corresponding asset mapping tells you what could happen but not where it could happen to you. A threat vector without a corresponding control tells you what to worry about but not what to do. For each vector in this taxonomy that applies to your AI systems, document three things: which specific assets are exposed, which controls currently mitigate the vector, and what residual exposure remains.&lt;/p&gt;
&lt;p&gt;The most effective way I have found to operationalize a threat taxonomy is to create a traceability matrix with four columns: threat vector, exposed assets, current controls, and residual risk rating. Populate this matrix for every production AI system. When a new asset is deployed, add rows. When a new threat vector is identified, add rows. When a control is implemented or modified, update the current controls column and reassess residual risk. This matrix becomes the working document for your AI security program. It tells you at any point what threats you have addressed and what gaps remain. Without it, the taxonomy is an intellectual exercise. With it, the taxonomy drives action.&lt;/p&gt;
&lt;h2 id="key-references-and-standards"&gt;Key References and Standards&lt;/h2&gt;
&lt;p&gt;This threat taxonomy draws from and aligns with the following authoritative frameworks.&lt;/p&gt;
&lt;p&gt;MITRE ATLAS (Adversarial Threat Landscape for Artificial Intelligence Systems) provides the primary reference taxonomy for AI-specific adversarial threats.&lt;/p&gt;
&lt;p&gt;NIST AI RMF (AI 100-1) provides the risk management framework for identifying and addressing AI threats across the lifecycle.&lt;/p&gt;
&lt;p&gt;ISO/IEC 27005:2022 provides the information security risk management process for integrating AI threats into enterprise risk assessment.&lt;/p&gt;
&lt;p&gt;ISO/IEC 23894:2023 provides AI-specific risk management guidance including threat identification.&lt;/p&gt;
&lt;p&gt;ISO/IEC 42001:2023 provides AI management system requirements for governance of the threat landscape.&lt;/p&gt;
&lt;p&gt;OWASP Machine Learning Security Top 10 provides a practitioner-focused list of common ML security threats.&lt;/p&gt;
&lt;p&gt;EU AI Act (Regulation 2024/1689) establishes regulatory requirements that make several of these threat vectors compliance-relevant.&lt;/p&gt;
&lt;p&gt;NIST SP 800-30 Rev. 1 provides guidance on threat identification and risk assessment methodology.&lt;/p&gt;
&lt;p&gt;ENISA Threat Landscape for AI provides European regulatory perspective on AI-specific threats.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/modern-office-meeting-with-colorful-glass-panes.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="using-this-taxonomy-to-find-your-gaps"&gt;Using This Taxonomy to Find Your Gaps&lt;/h2&gt;
&lt;p&gt;Organizations that treat this taxonomy as a reference list will read it, nod, and return to their existing threat models unchanged. They will continue to overweight adversarial external threats because those threats are familiar and dramatic. They will continue to underweight internal negligence threats because those threats feel mundane and uncomfortable to discuss. When an AI failure occurs, and it will, they will discover that the vector was sitting in a taxonomy they read but never operationalized.&lt;/p&gt;
&lt;p&gt;Organizations that treat this taxonomy as an audit tool will do something different. They will take each of their production AI systems and map every applicable threat vector to the specific assets, existing controls, and residual risk for that system. They will discover gaps, mostly in the negligent internal quadrant, and they will prioritize closing them. They will update their incident response plans to include AI-specific scenarios. They will build monitoring that catches negligence-driven failures before they reach customers. They will assign accountability for each production model to a named individual who cannot hide behind a team name.&lt;/p&gt;
&lt;p&gt;The threat your AI system faces tomorrow is almost certainly already in this taxonomy. The question is whether you have mapped it to your assets, built controls for it, and assigned someone to watch for it.&lt;/p&gt;
&lt;p&gt;Which quadrant of this taxonomy has the least coverage in your current AI threat model? For most organizations, the answer is negligent internal. That is where your biggest gap probably lives.&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and consulting work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative risk modeling, predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and globally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance landscapes.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item><item><title>The 49 AI Quality Characteristics That Define Whether Your System Actually Works</title><link>https://hwyler.github.io/blog/the-49-ai-quality-characteristics-that-define-whether-your-system-actually-works/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/the-49-ai-quality-characteristics-that-define-whether-your-system-actually-works/</guid><description>&lt;h2 id="a-practitioners-guide-to-isoiec-25059"&gt;A Practitioner&amp;rsquo;s Guide to ISO/IEC 25059&lt;/h2&gt;
&lt;p&gt;A model with 95% accuracy that nobody can explain, nobody can maintain, and nobody trusts is not a quality AI system. It is a liability waiting to surface.&lt;/p&gt;
&lt;p&gt;I learned this the hard way. Three years ago, I helped deploy a classification model for a European insurer. The accuracy metrics were excellent. The data science team celebrated. Six weeks later, the project was in crisis. The model could not be updated without breaking downstream integrations (maintainability failure). Users did not understand why it made specific recommendations and stopped trusting it (transparency failure). The system consumed three times the expected cloud resources during peak periods (performance efficiency failure). And when a regulator asked how the model made decisions about claims, nobody could provide an adequate explanation (accountability failure).&lt;/p&gt;
&lt;p&gt;The model worked. The system did not.&lt;/p&gt;
&lt;p&gt;That distinction, between a model that produces correct outputs and a system that delivers quality across its full operational lifecycle, is exactly what ISO/IEC 25059:2023 addresses. This standard defines 49 quality characteristics for AI systems across 11 requirement domains. It builds on the established software quality model of ISO/IEC 25010 but adapts it for the specific challenges of AI: opacity, learned behavior, data dependency, drift, fairness, and the unique ways AI systems interact with human judgment.&lt;/p&gt;
&lt;p&gt;This post walks through all 49 characteristics with practical implementation guidance for each. Use it as a checklist for AI system design, a framework for quality assurance, and a reference for identifying which characteristics create risk when they are absent.&lt;/p&gt;
&lt;h2 id="why-software-quality-models-are-not-enough-for-ai"&gt;Why Software Quality Models Are Not Enough for AI&lt;/h2&gt;
&lt;p&gt;ISO/IEC 25010 is the standard quality model for software products and systems. It has served the industry well for conventional software. It defines characteristics like reliability, security, maintainability, and usability that apply to any software system.&lt;/p&gt;
&lt;p&gt;AI systems need more.&lt;/p&gt;
&lt;p&gt;A traditional software system does what its code tells it to do. If the code is correct, the system is correct. An AI system does what its training data and learned parameters tell it to do. Correctness is probabilistic, not deterministic. The system can be &amp;ldquo;correct&amp;rdquo; on average while failing catastrophically for specific populations or edge cases.&lt;/p&gt;
&lt;p&gt;Three gaps in traditional software quality models become critical for AI.&lt;/p&gt;
&lt;p&gt;First, transparency and explainability are not optional quality attributes for AI. They are functional requirements. A user who cannot understand why an AI system made a particular decision cannot verify it, trust it, or correct it. Traditional software quality models treat transparency as a nice-to-have. For AI, it is a prerequisite for accountability and regulatory compliance.&lt;/p&gt;
&lt;p&gt;Second, AI systems degrade in ways traditional software does not. Data drift, concept drift, and model decay cause AI system quality to deteriorate over time even without any code changes. A quality model that only evaluates the system at deployment misses the ongoing quality challenges that define AI operations.&lt;/p&gt;
&lt;p&gt;Third, AI systems create societal risks that traditional software rarely produces. Bias, discrimination, loss of autonomy, environmental impact, and ethical harms are quality concerns specific to AI that require explicit quality characteristics and measurement approaches.&lt;/p&gt;
&lt;p&gt;ISO/IEC 25059 fills these gaps by extending the traditional quality model with AI-specific characteristics across every domain. The result is a comprehensive framework for evaluating whether an AI system is genuinely fit for purpose, not just whether it produces accurate outputs.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When I introduce ISO 25059 to organizations, the most common initial reaction is overwhelm. Forty-nine characteristics feels like an impossibly large quality surface to manage. The practical approach is to prioritize. Not every characteristic is equally relevant for every AI system. A customer-facing recommendation engine needs strong transparency, user controllability, and fairness characteristics. An internal process automation system needs strong reliability, maintainability, and robustness characteristics. Map the 49 characteristics to your specific system&amp;rsquo;s risk profile and context of use. Identify the 10 to 15 that are most critical. Focus your quality assurance resources there. Then expand coverage over time.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/engaged-professional-at-a-coffee-strewn-workstation.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="domain-1-functional-suitability"&gt;Domain 1: Functional Suitability&lt;/h2&gt;
&lt;p&gt;Four characteristics define whether the AI system does what it is supposed to do.&lt;/p&gt;
&lt;h3 id="functional-completeness"&gt;Functional Completeness&lt;/h3&gt;
&lt;p&gt;The degree to which the system&amp;rsquo;s functions cover all specified tasks and user objectives. The AI system provides all necessary functions for its intended purpose and fulfills all explicitly stated and implied user needs.&lt;/p&gt;
&lt;p&gt;This sounds basic. It is the characteristic most often violated in AI projects because teams focus on the model&amp;rsquo;s prediction function and neglect the surrounding functions that make the prediction useful: data preprocessing, output formatting, error handling, user feedback mechanisms, and integration with downstream workflows.&lt;/p&gt;
&lt;p&gt;To get this right, conduct rigorous requirements gathering that maps all user tasks to system functions. Validate coverage through user acceptance testing and traceability matrices that link requirements to implemented functions.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The functional completeness gap I find most often is the absence of a &amp;ldquo;decline to predict&amp;rdquo; function. Most AI systems are built to produce an output for every input. But there are inputs where the system should not produce a prediction because it does not have sufficient confidence, because the input falls outside its training distribution, or because the decision requires human judgment. Build the ability to abstain into your system. A credit scoring model that says &amp;ldquo;I cannot score this application with sufficient confidence, route to a human underwriter&amp;rdquo; is more functionally complete than one that produces a low-confidence score that a loan officer treats as definitive.&lt;/p&gt;
&lt;h3 id="functional-correctness"&gt;Functional Correctness&lt;/h3&gt;
&lt;p&gt;The degree to which the AI system provides correct results with the needed degree of precision. Outputs and effects are accurate and yield the right result.&lt;/p&gt;
&lt;p&gt;For AI systems, &amp;ldquo;correct&amp;rdquo; is probabilistic. A model with 90% accuracy is wrong 10% of the time. The question is not whether the system is perfect but whether its error rate falls within acceptable bounds for its application context, and whether errors are distributed fairly across populations.&lt;/p&gt;
&lt;p&gt;Establish ground truth datasets and validate continuously against predefined accuracy metrics such as F1-score, precision, and recall. Use adversarial testing to challenge model outputs and expose weaknesses.&lt;/p&gt;
&lt;h3 id="functional-appropriateness"&gt;Functional Appropriateness&lt;/h3&gt;
&lt;p&gt;The degree to which functions facilitate the accomplishment of specified tasks and objectives. Functions are suitable for the user&amp;rsquo;s stated goals and context of use.&lt;/p&gt;
&lt;p&gt;An AI system can be functionally complete and correct while still being inappropriate. A sentiment analysis model that classifies customer feedback into three categories (positive, negative, neutral) is functionally correct but functionally inappropriate if the customer service team needs to distinguish between 12 specific complaint types to route tickets effectively.&lt;/p&gt;
&lt;p&gt;Conduct task analysis and user studies to ensure functions align with actual user goals and workflows. Prioritize features based on user value, not technical feasibility.&lt;/p&gt;
&lt;h3 id="functional-adaptability"&gt;Functional Adaptability&lt;/h3&gt;
&lt;p&gt;The degree to which the AI system can be adapted for different specified tasks and environments. The system can be modified or configured for new purposes or contexts.&lt;/p&gt;
&lt;p&gt;AI systems that cannot adapt become obsolete quickly. Business requirements change, data distributions shift, and new use cases emerge. A system designed for a single, fixed purpose delivers diminishing value over time.&lt;/p&gt;
&lt;p&gt;Design systems with configurable parameters and hooks for retraining. Use feature flags and modular architecture to enable adaptation to new tasks without full redevelopment.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Functional adaptability is the suitability characteristic that determines long-term ROI, and it is almost always underinvested in during initial development because the pressure is on delivering the first use case. I worked with a logistics company that built a demand forecasting model tightly coupled to a single product category. When they wanted to extend it to two additional categories, they discovered the data pipeline, feature engineering, and model architecture were all hard-coded for the original category. Extending took nearly as long as building from scratch. Build adaptability into the architecture from day one, even if you are deploying for a single use case. Parameterize data sources, feature definitions, and model configurations. The marginal cost during initial development is small. The cost of retrofitting adaptability later is enormous.&lt;/p&gt;
&lt;h2 id="domain-2-performance-efficiency"&gt;Domain 2: Performance Efficiency&lt;/h2&gt;
&lt;p&gt;Three characteristics define whether the system uses resources appropriately.&lt;/p&gt;
&lt;h3 id="time-behaviour"&gt;Time Behaviour&lt;/h3&gt;
&lt;p&gt;The degree to which response and processing times meet requirements. The system delivers results within required time constraints.&lt;/p&gt;
&lt;p&gt;For AI systems, time behavior is more variable and harder to predict than for traditional software. Inference latency depends on model complexity, input size, hardware availability, and concurrent load. Training time depends on dataset size, model architecture, and compute resources.&lt;/p&gt;
&lt;p&gt;Profile system components to identify bottlenecks. Set Service Level Objectives for latency and throughput. Optimize models through quantization, pruning, or distillation for target deployment environments.&lt;/p&gt;
&lt;h3 id="resource-utilisation"&gt;Resource Utilisation&lt;/h3&gt;
&lt;p&gt;The degree to which resource usage meets requirements. The system uses appropriate amounts of processing capacity, memory, and network bandwidth.&lt;/p&gt;
&lt;p&gt;AI workloads consume significantly more resources than traditional applications. GPU costs for training, memory requirements for large models, and storage demands for training data can all exceed initial estimates.&lt;/p&gt;
&lt;p&gt;Monitor compute, memory, and network usage during both inference and training. Right-size infrastructure and use auto-scaling. Prefer efficient model architectures for resource-constrained deployment environments.&lt;/p&gt;
&lt;h3 id="capacity"&gt;Capacity&lt;/h3&gt;
&lt;p&gt;The degree to which maximum limits of system parameters meet requirements. The system handles the specified maximum number of items, users, or data volume.&lt;/p&gt;
&lt;p&gt;Perform load and stress testing to determine system limits across users, transactions, and data volume. Design architecture to scale horizontally. Build in rate limiting and graceful degradation so that exceeding capacity reduces performance rather than causing failure.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The performance efficiency characteristic that catches organizations off guard is resource utilization during retraining, not during inference. Teams size their infrastructure for inference workloads and then discover that monthly retraining jobs require 10 times the compute resources. The retraining job competes with inference for GPU capacity, degrading production performance during the retraining window. Separate training and inference infrastructure, or schedule retraining during off-peak periods with dedicated resource allocation. Monitor resource utilization during both operational modes separately.&lt;/p&gt;
&lt;h2 id="domain-3-compatibility"&gt;Domain 3: Compatibility&lt;/h2&gt;
&lt;p&gt;Two characteristics define how the system coexists with its environment.&lt;/p&gt;
&lt;h3 id="co-existence"&gt;Co-existence&lt;/h3&gt;
&lt;p&gt;The degree to which the AI system performs its functions while sharing a common environment and resources with other products. The system operates without negatively impacting other systems.&lt;/p&gt;
&lt;p&gt;Test the AI system in a staging environment that mirrors production, including all other applications that share resources. Ensure the AI system does not monopolize shared CPU, memory, or network bandwidth during peak inference or training periods.&lt;/p&gt;
&lt;h3 id="interoperability"&gt;Interoperability&lt;/h3&gt;
&lt;p&gt;The degree to which systems can exchange and use information. The AI system effectively communicates with other specified systems.&lt;/p&gt;
&lt;p&gt;Adopt standard data formats like ONNX and PMML for model exchange, and standard API protocols like REST and gRPC for communication. Implement rigorous schema validation for all data exchanges. Use API gateways for consistent management of interfaces.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Interoperability failures are among the most common reasons AI projects fail during the transition from development to production. The model works perfectly in the data science team&amp;rsquo;s environment but cannot consume data from the production pipeline because formats, schemas, or encoding conventions differ. Test interoperability between the development environment and the production environment early, ideally within the first two weeks of development. Discovering format mismatches at deployment is expensive. Discovering them during initial development is cheap.&lt;/p&gt;
&lt;h2 id="domain-4-usability"&gt;Domain 4: Usability&lt;/h2&gt;
&lt;p&gt;Seven characteristics define the human experience of interacting with the AI system. This is the largest traditional usability domain and includes two AI-specific additions: user controllability and transparency.&lt;/p&gt;
&lt;h3 id="appropriateness-recognisability"&gt;Appropriateness Recognisability&lt;/h3&gt;
&lt;p&gt;The degree to which users can recognize whether the system is appropriate for their needs. The system&amp;rsquo;s capabilities and limitations are clear to potential users.&lt;/p&gt;
&lt;p&gt;Provide clear documentation of capabilities, limitations, and intended use cases. Create a Model Card or similar fact sheet that communicates what the system does, what it does not do, what data it was trained on, and where it performs well or poorly.&lt;/p&gt;
&lt;h3 id="learnability"&gt;Learnability&lt;/h3&gt;
&lt;p&gt;The degree to which the system enables users to learn how to use it effectively. The system supports users in acquiring operational knowledge.&lt;/p&gt;
&lt;p&gt;Develop intuitive interfaces, comprehensive documentation, and interactive tutorials. Incorporate contextual help. Conduct usability testing to measure the learning curve across different user populations.&lt;/p&gt;
&lt;h3 id="operability"&gt;Operability&lt;/h3&gt;
&lt;p&gt;The degree to which the system is easy to operate and control. User effort for operation is minimized.&lt;/p&gt;
&lt;p&gt;Design clear and consistent interfaces and APIs. Provide effective error messages and status indicators. Automate complex operational tasks where possible.&lt;/p&gt;
&lt;h3 id="user-error-protection"&gt;User Error Protection&lt;/h3&gt;
&lt;p&gt;The degree to which the system protects users against making errors. The system prevents, detects, and helps users recover from mistakes.&lt;/p&gt;
&lt;p&gt;Implement input validation, confirmation dialogs for critical actions, and undo functionality. Use constraints to prevent invalid inputs. Guide users through complex tasks with clear step-by-step workflows.&lt;/p&gt;
&lt;h3 id="user-interface-aesthetics"&gt;User Interface Aesthetics&lt;/h3&gt;
&lt;p&gt;The degree to which the interface enables pleasing interaction. The design is visually and interactively appealing.&lt;/p&gt;
&lt;p&gt;Apply established design systems for visual consistency. Ensure a clean, uncluttered interface. Conduct user research on aesthetic perception to ensure the design supports rather than hinders the user&amp;rsquo;s task.&lt;/p&gt;
&lt;h3 id="accessibility"&gt;Accessibility&lt;/h3&gt;
&lt;p&gt;The degree to which the system can be used by people with the widest range of characteristics and capabilities. The system accommodates diverse user needs including disabilities.&lt;/p&gt;
&lt;p&gt;Follow WCAG 2.1 guidelines. Test with screen readers, ensure keyboard navigation, provide alt text for images, and support high contrast modes. Accessibility is not optional. It is a quality requirement and increasingly a legal one.&lt;/p&gt;
&lt;h3 id="user-controllability-ai-specific"&gt;User Controllability (AI-Specific)&lt;/h3&gt;
&lt;p&gt;The degree to which users can control the AI system&amp;rsquo;s behavior. Users can initiate, adjust, or stop the system&amp;rsquo;s operations.&lt;/p&gt;
&lt;p&gt;Provide settings to adjust system behavior such as confidence thresholds and filters. Allow users to start, stop, and correct operations. Ensure humans can always override AI decisions. This characteristic is directly tied to the EU AI Act&amp;rsquo;s requirements for human oversight of high-risk AI systems.&lt;/p&gt;
&lt;h3 id="transparency-ai-specific"&gt;Transparency (AI-Specific)&lt;/h3&gt;
&lt;p&gt;The degree to which the system&amp;rsquo;s functions, decisions, and outputs are understandable to the user. The system provides explanations for its behavior and results.&lt;/p&gt;
&lt;p&gt;Implement Explainable AI techniques like LIME and SHAP to provide output explanations. Document the model&amp;rsquo;s purpose, training data, and algorithms. Be explicit about the system&amp;rsquo;s AI nature. Transparency is not a single feature. It is a quality that must be designed into every interaction between the system and its users.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Of the seven usability characteristics, user controllability is the one most often missing from AI system designs, and it is the one regulators are asking about most frequently. I reviewed an AI system for a healthcare provider that provided diagnostic recommendations to physicians. The system had no mechanism for a physician to adjust the confidence threshold, no way to request an alternative recommendation, and no clear process for overriding the system&amp;rsquo;s output when clinical judgment disagreed. The system treated every recommendation as a final answer rather than an input to human decision-making. When the EU AI Act&amp;rsquo;s human oversight requirements were mapped against the system&amp;rsquo;s capabilities, the gap was significant. Build controllability from the start. Provide clear controls for adjusting, overriding, and stopping AI behavior. Document how these controls work and verify that users know how to use them.&lt;/p&gt;
\[Suggested image placement: A visual showing all 11 quality domains arranged in a wheel or grid, with the number of characteristics per domain indicated, highlighting the AI-specific additions in a distinct color\]&lt;h2 id="domain-5-reliability"&gt;Domain 5: Reliability&lt;/h2&gt;
&lt;p&gt;Five characteristics define whether the system performs consistently and recovers from failures. This domain includes one critical AI-specific addition: robustness.&lt;/p&gt;
&lt;h3 id="maturity"&gt;Maturity&lt;/h3&gt;
&lt;p&gt;The degree to which the system meets reliability needs under normal operation. The system is stable with a low failure rate in its standard operating environment.&lt;/p&gt;
&lt;p&gt;Establish a robust CI/CD pipeline with automated testing. Track mean time between failures. Use canary deployments to gradually roll out updates and catch stability issues before full deployment.&lt;/p&gt;
&lt;h3 id="availability"&gt;Availability&lt;/h3&gt;
&lt;p&gt;The degree to which the system is operational and accessible when required. The system has minimal downtime.&lt;/p&gt;
&lt;p&gt;Design for redundancy with failover mechanisms across availability zones. Monitor uptime and establish Service Level Agreements. Implement health checks and graceful degradation so that partial failures do not cause total outages.&lt;/p&gt;
&lt;h3 id="fault-tolerance"&gt;Fault Tolerance&lt;/h3&gt;
&lt;p&gt;The degree to which the system operates as intended despite hardware or software faults. The system continues functioning during component failures.&lt;/p&gt;
&lt;p&gt;Build systems that handle component failures without total collapse. Use retries with exponential backoff, circuit breakers, and fallback mechanisms. Design stateless services where possible to simplify recovery.&lt;/p&gt;
&lt;h3 id="recoverability"&gt;Recoverability&lt;/h3&gt;
&lt;p&gt;The degree to which the system can recover data and re-establish desired state after failure. The system restores service and data quickly.&lt;/p&gt;
&lt;p&gt;Implement automated backup and restore procedures for models and data. Define and test a Disaster Recovery plan. Track mean time to recovery. Ensure recovery points are consistent, meaning the model, its configuration, and its data are all restored to the same point in time.&lt;/p&gt;
&lt;h3 id="robustness-ai-specific"&gt;Robustness (AI-Specific)&lt;/h3&gt;
&lt;p&gt;The degree to which the system functions correctly despite invalid inputs, stressful conditions, or adversarial attacks. The system maintains performance under perturbation.&lt;/p&gt;
&lt;p&gt;This is the reliability characteristic most specific to AI and most critical for security. Test with noisy, out-of-distribution, and adversarial inputs. Use data augmentation, adversarial training, and defensive distillation to improve resilience. Monitor for data drift that degrades robustness over time.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Robustness is the reliability characteristic that creates the most direct link between quality and security. A model that is not robust against adversarial inputs is both a quality failure and a security vulnerability. Yet robustness testing is consistently treated as a security activity performed by the security team rather than a quality activity performed by the development team. This separation creates gaps. The security team tests for adversarial attacks. The development team tests for accuracy. Nobody tests for the space in between: inputs that are not adversarial but are unexpected, noisy, or from a different distribution than the training data. These &amp;ldquo;natural&amp;rdquo; robustness failures are more common than adversarial attacks and cause more cumulative damage. Include robustness testing in your development quality assurance process, not just in your security testing program.&lt;/p&gt;
&lt;h2 id="domain-6-security"&gt;Domain 6: Security&lt;/h2&gt;
&lt;p&gt;Six characteristics define the system&amp;rsquo;s security posture. This domain includes one AI-specific addition: intervenability.&lt;/p&gt;
&lt;h3 id="confidentiality"&gt;Confidentiality&lt;/h3&gt;
&lt;p&gt;The degree to which data are accessible only to those authorized. The system protects data from unauthorized disclosure.&lt;/p&gt;
&lt;p&gt;Encrypt data at rest and in transit. Implement strict role-based access controls and the principle of least privilege. Anonymize or pseudonymize training data. Consider secure multi-party computation for sensitive applications.&lt;/p&gt;
&lt;h3 id="integrity"&gt;Integrity&lt;/h3&gt;
&lt;p&gt;The degree to which the system prevents unauthorized modification of data or functions. The system ensures data and system accuracy and completeness.&lt;/p&gt;
&lt;p&gt;Use hashing and digital signatures to verify data and model artifacts have not been tampered with. Maintain an immutable audit trail. Validate inputs to prevent injection attacks, including adversarial inputs designed to manipulate model behavior.&lt;/p&gt;
&lt;h3 id="non-repudiation"&gt;Non-repudiation&lt;/h3&gt;
&lt;p&gt;The degree to which actions can be proven to have taken place. The system provides evidence for transactions that cannot be denied later.&lt;/p&gt;
&lt;p&gt;Implement secure logging and auditing for all significant actions and decisions. Use digital signatures to ensure actions can be attributed to a specific entity or user. For AI systems making consequential decisions, non-repudiation is essential for regulatory compliance and dispute resolution.&lt;/p&gt;
&lt;h3 id="accountability"&gt;Accountability&lt;/h3&gt;
&lt;p&gt;The degree to which actions can be traced to the entity that bears responsibility. The system enables assignment of responsibility.&lt;/p&gt;
&lt;p&gt;Maintain clear ownership of models and system components. Establish audit trails that log system decisions, data sources, and user interactions. Define clear lines of responsibility for every component and every decision the system produces.&lt;/p&gt;
&lt;h3 id="authenticity"&gt;Authenticity&lt;/h3&gt;
&lt;p&gt;The degree to which the identity of a subject or resource can be proved. The system verifies that entities are genuine.&lt;/p&gt;
&lt;p&gt;Implement strong authentication mechanisms including multi-factor authentication for system access. Verify the provenance of training data and model packages to prevent tampering or supply chain attacks. For AI systems, authenticity extends beyond user identity to include data authenticity and model authenticity.&lt;/p&gt;
&lt;h3 id="intervenability-ai-specific"&gt;Intervenability (AI-Specific)&lt;/h3&gt;
&lt;p&gt;The degree to which the system allows human intervention in its operation. Authorized humans can oversee and interrupt the system&amp;rsquo;s functions.&lt;/p&gt;
&lt;p&gt;Design human-in-the-loop processes for critical decisions. Provide clear interfaces for oversight, intervention, and manual override. Ensure the system can be paused or stopped safely at any point without data loss or inconsistent state. This characteristic is a direct requirement of the EU AI Act for high-risk AI systems.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Accountability and intervenability work together and fail together. An AI system that logs every decision (accountability) but provides no mechanism for a human to intervene when they see a problematic pattern in those logs (intervenability) creates awareness without agency. A system that allows human override (intervenability) but does not log who overrode which decision and why (accountability) creates agency without traceability. Design these two characteristics as a pair. The logging system should inform the intervention interface, and every intervention should be logged with the rationale for the override.&lt;/p&gt;
&lt;h2 id="domain-7-maintainability"&gt;Domain 7: Maintainability&lt;/h2&gt;
&lt;p&gt;Five characteristics define whether the system can be changed, fixed, and improved over time.&lt;/p&gt;
&lt;h3 id="modularity"&gt;Modularity&lt;/h3&gt;
&lt;p&gt;The degree to which the system is composed of discrete components where changes to one have minimal impact on others.&lt;/p&gt;
&lt;p&gt;Architect as loosely coupled components: separate data processing, training, and inference services. Use well-defined interfaces between components. This enables updating one component without risking the stability of others.&lt;/p&gt;
&lt;h3 id="reusability"&gt;Reusability&lt;/h3&gt;
&lt;p&gt;The degree to which components can be used in more than one system or context.&lt;/p&gt;
&lt;p&gt;Develop and package model components, feature pipelines, and datasets as reusable assets. Create shared libraries with clear documentation. Use containerization to make components portable across environments.&lt;/p&gt;
&lt;h3 id="analyzability"&gt;Analyzability&lt;/h3&gt;
&lt;p&gt;The degree to which the impact of an intended change can be assessed. The system can be diagnosed for deficiencies or failure causes.&lt;/p&gt;
&lt;p&gt;Implement comprehensive logging and monitoring for all components. Use distributed tracing to follow requests through the system. Maintain detailed documentation of architecture and data lineage so that when something fails, the cause can be traced efficiently.&lt;/p&gt;
&lt;h3 id="modifiability"&gt;Modifiability&lt;/h3&gt;
&lt;p&gt;The degree to which the system can be changed without introducing defects or degrading quality.&lt;/p&gt;
&lt;p&gt;Write clean, well-documented code. Avoid tight coupling. Use version control for all artifacts including code, data, and models. Implement feature toggles for controlled rollout of changes.&lt;/p&gt;
&lt;h3 id="testability"&gt;Testability&lt;/h3&gt;
&lt;p&gt;The degree to which test criteria can be established and tests can be performed effectively.&lt;/p&gt;
&lt;p&gt;Design systems with testing in mind from the start. Create isolated test environments. Automate unit, integration, and regression tests for both models and code. Monitor test coverage and maintain it as the system evolves.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Analyzability is the maintainability characteristic that determines how quickly you can respond to AI incidents, and it is the one most organizations invest in only after their first major incident. When a model starts producing unexpected outputs in production, the first question is always &amp;ldquo;what changed?&amp;rdquo; Without comprehensive data lineage, model versioning, and input/output logging, answering that question can take days. With them, it takes minutes. Build analyzability into your system before you need it. The cost of implementing logging, tracing, and lineage tracking during initial development is a fraction of the cost of retrofitting them during a production incident investigation.&lt;/p&gt;
&lt;h2 id="domain-8-portability"&gt;Domain 8: Portability&lt;/h2&gt;
&lt;p&gt;Three characteristics define how the system moves between environments.&lt;/p&gt;
&lt;h3 id="installability"&gt;Installability&lt;/h3&gt;
&lt;p&gt;The degree to which the system can be successfully deployed and removed in a specified environment.&lt;/p&gt;
&lt;p&gt;Package using standard tools like Docker containers, Helm charts, or pip packages. Automate deployment scripts. Provide clear installation documentation and dependency lists. AI systems often have complex dependency chains that make installation significantly harder than traditional software.&lt;/p&gt;
&lt;h3 id="replaceability"&gt;Replaceability&lt;/h3&gt;
&lt;p&gt;The degree to which the system can substitute for another product in the same environment.&lt;/p&gt;
&lt;p&gt;Adopt standard interfaces and protocols to avoid vendor lock-in. Ensure data and models are exportable in standard formats. Document APIs and dependencies thoroughly so that replacement is feasible when needed.&lt;/p&gt;
&lt;h3 id="adaptability"&gt;Adaptability&lt;/h3&gt;
&lt;p&gt;The degree to which the system can be adapted for different environments without custom modification.&lt;/p&gt;
&lt;p&gt;Use configuration files to manage environment-specific parameters. Avoid hard-coding values. Design the system to be environment-agnostic, sourcing configuration externally. Follow the Twelve-Factor App methodology for environment-independent design.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Portability characteristics collectively determine your vendor lock-in risk. I worked with an organization that deployed an AI system on a single cloud provider&amp;rsquo;s proprietary ML platform, using provider-specific data formats, training APIs, and deployment tools. When they needed to move to a multi-cloud architecture for resilience, the migration cost exceeded the original development cost. The system scored zero on all three portability characteristics. Before committing to a platform, evaluate your system against these three characteristics. If you score poorly on all three, you have accepted significant lock-in risk. Make that acceptance explicit and documented rather than accidental and discovered later.&lt;/p&gt;
&lt;h2 id="domain-9-quality-in-use"&gt;Domain 9: Quality in Use&lt;/h2&gt;
&lt;p&gt;This is the most expansive domain, containing 14 characteristics that evaluate the system&amp;rsquo;s quality from the perspective of actual use by real users in real contexts. Several of these characteristics are specific to AI and address societal and ethical dimensions that have no equivalent in traditional software quality models.&lt;/p&gt;
&lt;h3 id="effectiveness"&gt;Effectiveness&lt;/h3&gt;
&lt;p&gt;The degree to which accurate and complete results are achieved. The system helps users achieve specified goals with precision and comprehensiveness.&lt;/p&gt;
&lt;p&gt;Define clear metrics for accuracy and completeness aligned to user goals, not just model performance metrics. Implement robust validation and user acceptance testing. Track task success rates to measure real-world effectiveness.&lt;/p&gt;
&lt;h3 id="efficiency"&gt;Efficiency&lt;/h3&gt;
&lt;p&gt;The degree to which results are achieved with appropriate resources. The system minimizes user time, effort, and resource expenditure.&lt;/p&gt;
&lt;p&gt;Measure time-on-task and steps to completion for key user journeys. Optimize workflows and system performance to reduce user effort. Efficiency in the quality-in-use sense is about the user&amp;rsquo;s experience, not the system&amp;rsquo;s computational efficiency.&lt;/p&gt;
&lt;h3 id="usefulness"&gt;Usefulness&lt;/h3&gt;
&lt;p&gt;The degree to which the system is capable of achieving specified goals. The system serves a practical purpose and delivers tangible benefits.&lt;/p&gt;
&lt;p&gt;Conduct task analysis and user research to ensure the system solves a real problem. Prioritize features that deliver the highest value. Continuously validate usefulness through feedback and usage metrics.&lt;/p&gt;
&lt;h3 id="trust"&gt;Trust&lt;/h3&gt;
&lt;p&gt;The degree to which users have confidence that the system will behave as intended. The system is reliable, dependable, and predictable.&lt;/p&gt;
&lt;p&gt;Design for reliability, transparency, and fairness. Provide explanations for outputs and allow human oversight. Be clear about system limitations to build appropriate trust, not excessive trust that leads to overreliance.&lt;/p&gt;
&lt;h3 id="pleasure"&gt;Pleasure&lt;/h3&gt;
&lt;p&gt;The degree to which users obtain satisfaction from using the system. The experience is positive and enjoyable.&lt;/p&gt;
&lt;p&gt;Apply user-centered design principles. Conduct usability testing to identify and eliminate frustration points. Reward user actions positively through clear feedback and smooth interactions.&lt;/p&gt;
&lt;h3 id="comfort"&gt;Comfort&lt;/h3&gt;
&lt;p&gt;The degree to which users are satisfied with physical comfort during interaction. The system minimizes physical strain such as eye fatigue or repetitive stress.&lt;/p&gt;
&lt;p&gt;Design interfaces that adhere to ergonomic principles. Ensure readable text, comfortable interaction patterns, and support for assistive technologies.&lt;/p&gt;
&lt;h3 id="transparency-in-use"&gt;Transparency in Use&lt;/h3&gt;
&lt;p&gt;The degree to which users can understand the system&amp;rsquo;s functions, decisions, and outputs in practice. The system provides clarity on operations and reasoning.&lt;/p&gt;
&lt;p&gt;This extends the usability transparency characteristic into actual use contexts. Implement Explainable AI techniques suitable for end-users, such as natural language explanations rather than technical feature importance scores. Ensure explanations are actionable and understandable by non-technical users.&lt;/p&gt;
&lt;h3 id="economic-risk-mitigation"&gt;Economic Risk Mitigation&lt;/h3&gt;
&lt;p&gt;The degree to which the system mitigates potential economic risks. The system protects users and stakeholders from financial loss and wasted investment.&lt;/p&gt;
&lt;p&gt;Conduct cost-benefit and ROI analyses. Implement safeguards against errors that could lead to significant financial loss. Ensure transparency in automated financial decisions. This characteristic is particularly relevant for AI systems that make or influence financial decisions at scale.&lt;/p&gt;
&lt;h3 id="health-and-safety-risk-mitigation"&gt;Health and Safety Risk Mitigation&lt;/h3&gt;
&lt;p&gt;The degree to which the system mitigates health and safety risks. The system prioritizes human well-being above all else.&lt;/p&gt;
&lt;p&gt;Perform rigorous risk assessments such as Failure Mode and Effects Analysis for safety-critical applications. Implement fail-safes, human-in-the-loop controls, and continuous monitoring for hazardous situations. Comply with relevant safety standards such as IEC 61508 for functional safety.&lt;/p&gt;
&lt;h3 id="environment-risk-mitigation"&gt;Environment Risk Mitigation&lt;/h3&gt;
&lt;p&gt;The degree to which the system mitigates environmental risks. The system minimizes negative environmental impacts.&lt;/p&gt;
&lt;p&gt;Monitor and optimize computational efficiency and energy footprint. Prefer cloud regions powered by renewable energy. Consider the full lifecycle environmental impact including training, inference, and data storage.&lt;/p&gt;
&lt;h3 id="societal-and-ethical-risk-mitigation"&gt;Societal and Ethical Risk Mitigation&lt;/h3&gt;
&lt;p&gt;The degree to which the system mitigates societal and ethical risks. The system avoids causing harm, promotes fairness, and upholds ethical principles.&lt;/p&gt;
&lt;p&gt;Establish an AI Ethics board and guidelines. Proactively test for and mitigate biases. Ensure fairness, accountability, and transparency throughout the AI lifecycle. Conduct impact assessments for high-risk applications.&lt;/p&gt;
&lt;h3 id="context-completeness"&gt;Context Completeness&lt;/h3&gt;
&lt;p&gt;The degree to which the system can achieve goals in all specified contexts of use. The system functions effectively across all intended situations, environments, and user profiles.&lt;/p&gt;
&lt;p&gt;Identify and test all specified contexts during development. Use diverse datasets that represent all intended environments and user groups. Monitor for context drift in production where the system encounters situations outside its training distribution.&lt;/p&gt;
&lt;h3 id="flexibility"&gt;Flexibility&lt;/h3&gt;
&lt;p&gt;The degree to which the system can achieve goals in contexts beyond those initially specified. The system adapts to unanticipated situations.&lt;/p&gt;
&lt;p&gt;Design with modular and adaptable architectures. Allow configuration and customization. Use techniques like transfer learning to enable adaptation to new contexts without full redevelopment.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The quality-in-use characteristics that organizations most consistently neglect are the four risk mitigation characteristics: economic, health and safety, environmental, and societal/ethical. These characteristics feel like &amp;ldquo;someone else&amp;rsquo;s job.&amp;rdquo; The AI development team focuses on effectiveness and efficiency. The compliance team handles economic risk. The safety team handles health and safety. Nobody owns environmental or societal risk mitigation as a quality characteristic of the AI system itself. But ISO 25059 places these squarely within the quality model. They are quality attributes of the system, not external governance requirements. Treat them as you would any other quality characteristic: define metrics, set targets, test against them, and monitor in production. The system&amp;rsquo;s quality is incomplete without them.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/watermark-free-gemini_generated_image_fn28r6fn28r6fn28.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="using-this-framework-for-risk-identification"&gt;Using This Framework for Risk Identification&lt;/h2&gt;
&lt;p&gt;The 49 characteristics in ISO/IEC 25059 serve a dual purpose. They define what quality looks like for an AI system, and they identify where risk lives when quality is absent.&lt;/p&gt;
&lt;p&gt;Every characteristic that scores poorly represents a risk. Low functional correctness means the system produces errors. Low robustness means the system is vulnerable to adversarial inputs. Low transparency means decisions cannot be explained to regulators. Low intervenability means humans cannot stop the system when it malfunctions.&lt;/p&gt;
&lt;p&gt;Map this directly to your AI risk assessment. For each characteristic, ask three questions. How does our system perform against this characteristic? What is the consequence if this characteristic fails? What controls do we have in place to maintain this characteristic over time?&lt;/p&gt;
&lt;p&gt;The answers populate your risk register with specific, measurable, and controllable risks rather than generic categories like &amp;ldquo;model risk&amp;rdquo; or &amp;ldquo;AI quality issues.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Original implementation tip: I use the 49 characteristics as a structured interview guide during AI risk assessments. For each production AI system, I walk through every characteristic with the development team, operations team, and business owner. Each conversation takes about two hours. The output is a quality profile for the system with a red/amber/green rating for each characteristic. Red-rated characteristics map directly to risks in the risk register. Amber-rated characteristics map to watch items with monitoring requirements. This approach produces a more comprehensive risk identification than any brainstorming-based approach I have used, because the characteristics serve as prompts that surface risks the team would not think of on their own. &amp;ldquo;How does your system handle invalid inputs?&amp;rdquo; (robustness) and &amp;ldquo;Can a user override the system&amp;rsquo;s decision?&amp;rdquo; (intervenability) consistently uncover risks that open-ended risk identification sessions miss.&lt;/p&gt;
&lt;h2 id="key-references-and-standards"&gt;Key References and Standards&lt;/h2&gt;
&lt;p&gt;This quality framework draws from and aligns with the following authoritative sources.&lt;/p&gt;
&lt;p&gt;ISO/IEC 25059:2023 for the primary AI system quality model that defines the 49 characteristics described in this post.&lt;/p&gt;
&lt;p&gt;ISO/IEC 25010:2023 for the foundational software product and system quality model that ISO 25059 extends.&lt;/p&gt;
&lt;p&gt;ISO/IEC/IEEE 29148 for requirements engineering practices that support functional suitability assessment.&lt;/p&gt;
&lt;p&gt;ISO 9241-210 for human-centered design principles that support usability assessment.&lt;/p&gt;
&lt;p&gt;NIST AI RMF (AI 100-1) for the AI risk management framework that connects quality characteristics to risk management.&lt;/p&gt;
&lt;p&gt;NIST AI 100-2 for adversarial machine learning guidance that supports robustness assessment.&lt;/p&gt;
&lt;p&gt;EU AI Act (Regulation 2024/1689) for regulatory requirements that make several quality characteristics legally mandatory for high-risk AI systems.&lt;/p&gt;
&lt;p&gt;W3C WCAG 2.1 for web accessibility guidelines that support the accessibility characteristic.&lt;/p&gt;
&lt;p&gt;ISO/IEC 27001 for information security management standards that support security characteristics.&lt;/p&gt;
&lt;p&gt;IEC 61508 for functional safety standards relevant to health and safety risk mitigation.&lt;/p&gt;
&lt;h2 id="the-difference-between-accurate-and-good"&gt;The Difference Between Accurate and Good&lt;/h2&gt;
&lt;p&gt;Organizations that evaluate their AI systems only on accuracy metrics will continue deploying systems that work in testing and fail in production, that produce correct outputs nobody trusts, that cannot be maintained by anyone other than their original developer, and that create regulatory exposure because they cannot explain their decisions. High accuracy on a test set is one characteristic out of 49. Treating it as the only one that matters is how quality failures happen.&lt;/p&gt;
&lt;p&gt;Organizations that evaluate their AI systems across the full quality model will build systems that are not only accurate but explainable, robust, maintainable, fair, controllable, and recoverable. They will catch quality gaps during development rather than discovering them through production incidents. They will satisfy regulatory requirements because the quality characteristics regulators care about, transparency, accountability, intervenability, fairness, were designed in from the start.&lt;/p&gt;
&lt;p&gt;An AI system that scores well on one quality characteristic and poorly on 48 others is not a quality system. It is a model with infrastructure around it. The infrastructure is where quality lives or dies.&lt;/p&gt;
&lt;p&gt;Which of the 49 characteristics is weakest in your most important AI system? If you cannot answer that question, start with the structured interview approach described above. Two hours will reveal gaps that months of operation have hidden.&lt;/p&gt;</description></item><item><title>The AI Loss Taxonomy Your Risk Assessments Are Missing</title><link>https://hwyler.github.io/blog/the-ai-loss-taxonomy-your-risk-assessments-are-missing/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/the-ai-loss-taxonomy-your-risk-assessments-are-missing/</guid><description>&lt;h3 id="incident-types-and-direct-loss-categories-that-define-real-exposure-for-ai-projects"&gt;Incident Types and Direct Loss Categories That Define Real Exposure for AI Projects&lt;/h3&gt;
&lt;p&gt;Here is a question that reveals whether your AI risk program is mature or performative: Can you name the specific types of losses your AI systems could produce?&lt;/p&gt;
&lt;p&gt;Not vague categories like &amp;ldquo;financial impact&amp;rdquo; or &amp;ldquo;reputational damage.&amp;rdquo; Specific, measurable loss types with clear boundaries between them. The difference between a regulatory fine and a legal compensation payment. The difference between algorithm remediation costs and data regeneration costs. The difference between customer churn and business disruption.&lt;/p&gt;
&lt;p&gt;I asked this question to the risk committee of a healthcare AI company two years ago. The room went quiet. They had a risk register with 20 AI risks, each rated on a five-point scale for likelihood and impact. But when I asked &amp;ldquo;what kind of impact?&amp;rdquo; nobody could decompose their generic &amp;ldquo;high impact&amp;rdquo; ratings into the specific loss types that would actually appear on a financial statement or in a regulatory action.&lt;/p&gt;
&lt;p&gt;That gap matters. You cannot quantify what you cannot classify. And you cannot prioritize controls, calculate return on investment, or purchase appropriate insurance if you cannot distinguish between the types of losses your AI systems might generate.&lt;/p&gt;
&lt;p&gt;This post provides two complementary taxonomies. The first catalogs 37 distinct AI-related incident types across eight categories, each classified by whether it creates internal losses (relevant to risk assessments) or external losses (relevant to impact assessments) or both. The second catalogs 15 direct loss types across five domains that map to specific financial line items. Together, they give you the vocabulary and structure to make your AI risk assessments financially precise.&lt;/p&gt;
&lt;h2 id="why-generic-loss-categories-fail"&gt;Why Generic Loss Categories Fail&lt;/h2&gt;
&lt;p&gt;Most AI risk assessments use three to five impact categories: financial, operational, reputational, regulatory, and strategic. These categories are so broad that they obscure more than they reveal.&lt;/p&gt;
&lt;p&gt;When a risk assessment says an AI system has &amp;ldquo;high financial impact,&amp;rdquo; does that mean the organization will pay regulatory fines? Lose customers? Write off a failed project? Pay for emergency model remediation? All of these are &amp;ldquo;financial impact,&amp;rdquo; but they involve different stakeholders, different timescales, different control strategies, and different insurance coverage. Lumping them together makes the risk assessment useless for decision-making.&lt;/p&gt;
&lt;p&gt;The same problem applies to incident classification. &amp;ldquo;AI bias&amp;rdquo; is not a single incident type. It manifests as biased outputs, unequal performance across groups, unfair discrimination, and lack of diversity in development teams. Each manifestation has different causes, different controls, and different loss profiles. Treating them as one incident type produces controls that are too generic to be effective.&lt;/p&gt;
&lt;p&gt;The solution is granularity. Not complexity for its own sake, but sufficient decomposition to enable specific, actionable analysis. The taxonomies in this post provide that granularity.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When I first introduced a granular loss taxonomy to a financial services client, their initial reaction was that it added unnecessary complexity. They were managing 15 AI risks with five impact categories and felt that was sufficient. I asked them to take their highest-rated risk, &amp;ldquo;model produces biased outputs,&amp;rdquo; and trace it to specific financial consequences. They identified regulatory fines quickly. Then I asked about legal compensation payments to affected customers, algorithm remediation costs for retraining the model, control remediation costs for fixing governance gaps found during investigation, customer churn from affected populations, and reputation damage from media coverage. The total potential exposure across these six loss types was four times their original &amp;ldquo;high impact&amp;rdquo; estimate. Granularity did not add complexity. It revealed exposure they had been underestimating.&lt;/p&gt;
&lt;h2 id="part-1-ai-related-incident-types"&gt;Part 1: AI-Related Incident Types&lt;/h2&gt;
&lt;p&gt;The incident taxonomy organizes 37 distinct incident types across eight categories. Each incident is classified as producing internal losses (considered in risk assessments), external losses (considered in impact assessments), or both.&lt;/p&gt;
&lt;p&gt;This distinction matters for assessment methodology. Internal losses affect the organization directly through operational disruption, remediation costs, and control failures. External losses affect individuals, communities, or society through harm, discrimination, or rights violations. Many incidents produce both, requiring assessment from both perspectives.&lt;/p&gt;
&lt;h3 id="category-1-cognitive-degradation"&gt;Category 1: Cognitive Degradation&lt;/h3&gt;
&lt;p&gt;Three incident types address AI&amp;rsquo;s impact on human cognitive and decisional capacity.&lt;/p&gt;
&lt;p&gt;Addiction and digital wellness (external only) occurs when AI systems contribute to addictive behaviors and negative impacts on digital wellness. Recommendation algorithms that maximize engagement metrics can create patterns of compulsive use. AI-driven content curation that prioritizes emotional arousal over informational value degrades the quality of users&amp;rsquo; information environment. This is an external loss because the harm falls on users, not the organization, but regulatory attention to digital wellness is increasing, which creates secondary compliance exposure.&lt;/p&gt;
&lt;p&gt;Loss of autonomy (internal and external) occurs when AI systems make decisions that diminish user control. This happens when automated decision-making replaces human judgment in contexts where individuals should retain meaningful choice. Internally, this manifests when employees lose the ability to exercise professional judgment because AI systems override their input. Externally, customers or citizens experience reduced agency in decisions affecting their lives, such as credit, employment, or healthcare.&lt;/p&gt;
&lt;p&gt;Overreliance on AI (internal and external) occurs when users anthropomorphize, trust, or depend on AI systems beyond what the system&amp;rsquo;s capabilities warrant. Internally, decision-makers who treat model outputs as infallible stop applying critical judgment. Externally, users develop inappropriate emotional or material dependencies on AI systems, or form expectations the system cannot meet.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Overreliance on AI is the cognitive degradation incident type that creates the most immediate organizational risk, and it is almost never included in AI risk assessments. I worked with a lending organization where loan officers had become so accustomed to following the AI&amp;rsquo;s credit recommendations that they stopped reviewing the underlying data. When the model began producing anomalous scores due to a data pipeline issue, officers approved loans they would have questioned under manual review. The model was technically malfunctioning, but the actual failure was human. The loan officers had ceded their judgment to the system. The control is not technical. It is procedural: require documented human rationale for a sample of AI-supported decisions, and audit whether the rationale demonstrates independent judgment or simply restates the AI&amp;rsquo;s recommendation.&lt;/p&gt;
&lt;h3 id="category-2-discrimination"&gt;Category 2: Discrimination&lt;/h3&gt;
&lt;p&gt;Five incident types address unfair or unequal treatment produced by AI systems.&lt;/p&gt;
&lt;p&gt;Bias in AI outputs (internal and external) occurs when models produce systematically biased predictions or recommendations. This is the broadest discrimination incident type and encompasses statistical bias embedded in model outputs that disadvantages specific groups.&lt;/p&gt;
&lt;p&gt;Exposure to toxic content (external only) occurs when AI systems expose users to harmful, abusive, unsafe, or inappropriate content. Content recommendation systems, generative AI outputs, and AI-moderated platforms all carry this risk. The loss is borne by the affected users, but regulatory and reputational consequences flow back to the organization.&lt;/p&gt;
&lt;p&gt;Lack of diversity in AI development (internal and external) occurs when homogeneous development teams build systems that reflect their own perspectives and blind spots. This is a root cause incident type. It does not produce harm directly but creates the conditions for bias, unfair discrimination, and unequal performance across groups.&lt;/p&gt;
&lt;p&gt;Unequal performance across groups (internal and external) occurs when AI systems deliver different levels of accuracy, reliability, or quality for different user populations. A facial recognition system that works well for some skin tones and poorly for others. A speech recognition system that understands some accents and fails on others. The performance disparity itself is the incident, regardless of whether it results from intentional design or data limitations.&lt;/p&gt;
&lt;p&gt;Unfair discrimination (internal and external) occurs when AI systems treat individuals or groups unfairly in consequential decisions. This goes beyond statistical bias in outputs to encompass the downstream effects: denied loans, rejected applications, misclassified individuals, or misrepresented groups.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The discrimination incident type that is hardest to detect is unequal performance across groups, because standard accuracy metrics can mask it completely. A model with 92% overall accuracy might have 97% accuracy for the majority population and 74% accuracy for a minority group. The aggregate metric looks fine. The disaggregated metrics reveal a serious problem. When I audit AI systems for discrimination risk, I require performance metrics disaggregated by every protected characteristic available in the data. If protected characteristics are not in the data, which is common, I require proxy analysis using correlated variables. The first time you disaggregate your model&amp;rsquo;s performance metrics, you will almost certainly find disparities you did not know existed.&lt;/p&gt;
&lt;h3 id="category-3-disinformation-warfare"&gt;Category 3: Disinformation Warfare&lt;/h3&gt;
&lt;p&gt;Three incident types address AI&amp;rsquo;s role in the information environment.&lt;/p&gt;
&lt;p&gt;Disinformation and influence at scale (internal and external) occurs when AI systems enable large-scale manipulation of public opinion. This includes using AI to generate convincing fake content, automate social media manipulation, or conduct targeted influence campaigns. Internally, organizations face risk when their AI tools are misused for this purpose. Externally, society bears the cost of degraded public discourse.&lt;/p&gt;
&lt;p&gt;False or misleading information (internal and external) occurs when AI systems generate or spread incorrect or deceptive information. This includes hallucination in large language models, inaccurate summaries, fabricated citations, and confidently stated falsehoods. Unlike deliberate disinformation, this often results from model limitations rather than malicious intent, but the impact on users who rely on the information is the same.&lt;/p&gt;
&lt;p&gt;Pollution of information ecosystem (external only) occurs when AI-generated misinformation accumulates at sufficient scale to undermine shared reality. Filter bubbles, echo chambers, and the displacement of human-created content by AI-generated content of unknown reliability all contribute to this systemic effect.&lt;/p&gt;
&lt;p&gt;Original implementation tip: False or misleading information is the disinformation incident type with the most immediate organizational liability, particularly for companies deploying generative AI in customer-facing applications. I advised a professional services firm that deployed a generative AI assistant to help clients navigate regulatory requirements. Within the first month, the assistant fabricated a regulation that did not exist and cited it confidently to a client. The client made a business decision based on the fabricated guidance. The firm&amp;rsquo;s liability exposure from that single incident exceeded the entire annual budget for their AI program. The control that would have prevented this is output verification: for any generative AI system providing factual information to external users, implement a verification layer that checks generated claims against an authoritative source before presenting them. This adds latency and cost. It also prevents lawsuits.&lt;/p&gt;
&lt;h3 id="category-4-economic-displacement"&gt;Category 4: Economic Displacement&lt;/h3&gt;
&lt;p&gt;Nine incident types address AI&amp;rsquo;s macroeconomic and organizational effects, making this the largest incident category.&lt;/p&gt;
&lt;p&gt;Changes in employment patterns (internal and external) covers reduced quality of employment and increased exploitation of workers as AI reshapes job roles. Competitive dynamics (internal only) addresses the organizational risk from racing to deploy AI systems before they are safe, a pattern that increases the probability of releasing error-prone systems. Disruption of traditional industries (internal and external) covers economic instability when AI displaces established business models.&lt;/p&gt;
&lt;p&gt;Economic and cultural devaluation of human effort (internal and external) occurs when AI-generated output reduces the perceived or actual value of human-created work. This affects pricing, employment, and professional identity across creative, analytical, and service industries.&lt;/p&gt;
&lt;p&gt;Environmental harm (external only) covers the energy consumption, water usage, and carbon emissions from training and operating large AI systems. Governance failure (internal and external) occurs when regulatory frameworks cannot keep pace with AI development, creating gaps in oversight.&lt;/p&gt;
&lt;p&gt;Increased inequality and decline in employment quality (internal and external) addresses the broader societal pattern of AI benefits accruing to capital owners while labor bears displacement costs. Job displacement and economic disruption (internal and external) covers direct job losses and industry disruption. Power centralization and unfair distribution of benefits (external only) addresses the concentration of AI capabilities and their economic benefits among a small number of organizations.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Of the nine economic displacement incident types, governance failure is the one that creates the most direct and immediate organizational risk, because it applies to every organization deploying AI, regardless of industry or scale. Governance failure is not just about regulators failing to keep pace with technology. It is also about your organization failing to build internal governance that compensates for regulatory gaps. I worked with a technology company that was deploying AI across 14 use cases with no centralized governance body, no standardized risk assessment process, and no consistent documentation requirements. Each team made independent decisions about model deployment, monitoring, and retirement. When the EU AI Act requirements became concrete, the company had no way to determine which of their systems qualified as high-risk, what documentation existed for each system, or who was accountable for compliance. They spent 11 months and significant resources building governance retroactively that would have cost a fraction to build proactively. If your organization deploys AI and does not have a governance framework, this is your highest-priority incident type to address. Not because governance failure is the most dramatic risk, but because its absence makes every other risk harder to manage.&lt;/p&gt;
&lt;h3 id="category-5-exploitation"&gt;Category 5: Exploitation&lt;/h3&gt;
&lt;p&gt;Two incident types address deliberate misuse of AI for harm.&lt;/p&gt;
&lt;p&gt;AI weaponization (external only) covers the use of AI systems to develop cyber weapons or tools capable of mass harm. This is primarily a societal risk but creates organizational exposure when an organization&amp;rsquo;s AI tools or models are repurposed for weaponization by third parties.&lt;/p&gt;
&lt;p&gt;Fraud, scams, and targeted manipulation (external only) covers the use of AI to conduct fraud, run scams, or manipulate individuals through personalized deception. AI-generated deepfake voices used in CEO fraud, AI-crafted phishing messages personalized from scraped data, and AI-assisted identity theft all fall here.&lt;/p&gt;
&lt;h3 id="category-6-malicious-actors-and-misinformation"&gt;Category 6: Malicious Actors and Misinformation&lt;/h3&gt;
&lt;p&gt;Three incident types address AI-enabled attacks and synthetic media.&lt;/p&gt;
&lt;p&gt;AI-powered phishing and social engineering (external only) covers the use of AI to create sophisticated, personalized phishing attacks and social engineering campaigns. AI enables attackers to generate convincing communications at scale, personalized to each target using publicly available information.&lt;/p&gt;
&lt;p&gt;Use of AI for social engineering (external only) is a related but broader category covering all uses of AI to manipulate human behavior for unauthorized access or information disclosure.&lt;/p&gt;
&lt;p&gt;Deepfakes and AI-generated content (external only) covers AI-generated synthetic media used to spread misinformation, impersonate individuals, or manipulate public opinion. This includes fake video, audio, images, and text that are increasingly difficult to distinguish from authentic content.&lt;/p&gt;
&lt;h3 id="category-7-privacy-infringement"&gt;Category 7: Privacy Infringement&lt;/h3&gt;
&lt;p&gt;Five incident types address AI&amp;rsquo;s impact on personal data and privacy.&lt;/p&gt;
&lt;p&gt;AI system security vulnerabilities and attacks (external only) covers exploitation of vulnerabilities in AI systems leading to unauthorized access, data breaches, or system manipulation causing unsafe outputs.&lt;/p&gt;
&lt;p&gt;Collection of personal data (external only) covers AI systems that collect personal data without adequate consent. This includes passive data collection through AI-powered sensors, inference of personal characteristics from behavioral data, and collection that exceeds stated purposes.&lt;/p&gt;
&lt;p&gt;Compromise of privacy (external only) occurs when AI systems memorize and leak sensitive personal data, or infer private information about individuals without consent. This is distinct from data breaches because the privacy compromise occurs through the model&amp;rsquo;s normal operation, not through a security failure.&lt;/p&gt;
&lt;p&gt;Data breaches and unauthorized access (external only) covers traditional security incidents applied to AI contexts, including unauthorized access to training data, model weights, or inference logs containing personal information.&lt;/p&gt;
&lt;p&gt;Surveillance and monitoring (external only) covers AI-powered surveillance that erodes trust and creates unease among individuals and communities. Facial recognition in public spaces, behavioral monitoring in workplaces, and predictive policing systems all carry this risk.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Compromise of privacy is the privacy incident type that is most specific to AI and least covered by traditional privacy controls. A large language model can memorize and reproduce fragments of its training data, including personal information, in its outputs. This is not a data breach in the traditional sense. No attacker exploited a vulnerability. The model simply learned its training data too well and reproduces it when prompted in certain ways. Traditional privacy controls focus on securing data at rest and in transit. They do not address data that is encoded in model weights. The control for this risk is differential privacy during training (adding noise to prevent memorization of individual data points) combined with output filtering that detects and blocks personal information in model responses. If your AI system was trained on data containing personal information, this incident type applies to you.&lt;/p&gt;
&lt;h3 id="category-8-value-misalignment"&gt;Category 8: Value Misalignment&lt;/h3&gt;
&lt;p&gt;Seven incident types address fundamental alignment between AI systems and human values.&lt;/p&gt;
&lt;p&gt;AI possessing dangerous capabilities (external only) covers AI systems that develop or access capabilities increasing their potential for mass harm. This is an emerging and contested risk category, but it is increasingly relevant as AI systems become more capable.&lt;/p&gt;
&lt;p&gt;AI pursuing its own goals in conflict with human goals (external only) covers AI systems acting contrary to the intentions of their designers or users. This ranges from reward hacking in reinforcement learning systems (achieving the stated objective through unintended means) to more speculative scenarios of advanced AI systems developing emergent goals.&lt;/p&gt;
&lt;p&gt;AI system reliability and maintainability (internal and external) covers systems that are not reliable or maintainable, leading to errors and failures with significant consequences. This is particularly critical in applications requiring moral reasoning or operating in safety-critical environments.&lt;/p&gt;
&lt;p&gt;Lack of accountability (internal and external) occurs when AI decision-making processes have no clear accountable party, leading to situations where harmful outcomes cannot be attributed, corrected, or prevented from recurring.&lt;/p&gt;
&lt;p&gt;Lack of capability or robustness (internal and external) covers AI systems that fail under varying conditions. A model that works in testing but fails in production, a system that degrades when input distributions shift, or an application that produces errors under edge cases all represent this incident type.&lt;/p&gt;
&lt;p&gt;Lack of explainability (internal and external) occurs when AI systems cannot explain their decisions to stakeholders who need to understand them, whether those stakeholders are regulators, affected individuals, or internal decision-makers.&lt;/p&gt;
&lt;p&gt;Lack of transparency or interpretability (internal and external) covers broader challenges in understanding AI decision-making processes, leading to difficulty enforcing compliance, holding actors accountable, and identifying errors.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The value misalignment incident type I find most practically relevant for organizations today, the one that is neither speculative nor distant, is lack of accountability. Every AI failure I have investigated has had an accountability gap at its root. Not the absence of a responsible person in an organizational chart, but the absence of a person who knew they were responsible, had the authority to act, and had the information needed to act in time. The control is deceptively simple: for every production AI system, publish an accountability card that names the individual accountable for model performance, the individual accountable for data quality, the individual accountable for compliance, and the individual accountable for incident response. Post these accountability cards where the operations team can see them. Update them when people change roles. Test them by calling the named individuals during a tabletop exercise and verifying they know they are accountable and know what to do.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/professional-man-at-modern-workspace.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="part-2-direct-loss-types"&gt;Part 2: Direct Loss Types&lt;/h2&gt;
&lt;p&gt;The incident taxonomy tells you what can happen. The direct loss taxonomy tells you what it costs. These 15 loss types map to specific financial line items that appear in budgets, financial statements, and insurance claims. They give your risk quantification the precision needed for credible Monte Carlo simulation and ROI analysis.&lt;/p&gt;
&lt;p&gt;Five domains organize the 15 loss types.&lt;/p&gt;
&lt;h3 id="domain-1-compliance-losses"&gt;Domain 1: Compliance Losses&lt;/h3&gt;
&lt;p&gt;Four loss types address the financial consequences of regulatory and legal exposure.&lt;/p&gt;
&lt;p&gt;Regulatory fines cover penalties for violating AI regulations like the EU AI Act, privacy laws like GDPR, or sector-specific requirements. They also cover sanctions for data breaches, discriminatory outcomes, or copyright infringements produced by AI systems. These are typically the most visible AI losses because they are public, quantifiable, and reported.&lt;/p&gt;
&lt;p&gt;Legal compensations cover settlement payments to affected parties for harm caused by AI malfunctions or decisions. This includes attorney fees and court costs for defending lawsuits from individuals or groups. Unlike regulatory fines, which are imposed by authorities, legal compensations arise from private litigation. They can be larger than fines and take longer to resolve.&lt;/p&gt;
&lt;p&gt;Contractual credits cover service credits issued to customers when AI performance falls below guaranteed levels. Refunds and discounts applied for missed availability or accuracy commitments. These losses are often overlooked in risk assessments because they are managed by commercial teams, not risk teams, but they can be significant for organizations selling AI-powered services.&lt;/p&gt;
&lt;p&gt;Legal response costs cover external legal counsel fees for investigating and responding to AI-related claims, as well as internal legal team costs for compliance reviews and regulatory correspondence. These costs are incurred regardless of whether the organization is ultimately found liable.&lt;/p&gt;
&lt;p&gt;Control remediation covers costs to fix governance gaps identified in failed AI audits. This includes documentation, implementation, and certification expenses for new compliance controls and frameworks. This loss type often surprises organizations because it represents the cost of building governance they should have built proactively.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When estimating compliance losses for risk quantification, the most common error is using historical fine amounts as the basis for estimates. Historical data underestimates future exposure for two reasons. First, AI-specific regulations like the EU AI Act establish fine structures that far exceed previous penalties: up to 35 million euros or 7% of global annual turnover for certain violations. Second, regulatory enforcement of AI is in its early stages. The fines imposed in 2025 and 2026 will set precedents that do not yet exist in historical data. For AI compliance loss estimation, use the maximum penalty structures defined in applicable regulations as the upper bound of your range, not historical fine amounts. Your calibrated experts should estimate the probability of enforcement action and the likely penalty within the regulatory range, but the range itself should reflect the legal maximum, not past experience.&lt;/p&gt;
&lt;h3 id="domain-2-ittechnical-losses"&gt;Domain 2: IT/Technical Losses&lt;/h3&gt;
&lt;p&gt;Three loss types address the costs of technical remediation and infrastructure.&lt;/p&gt;
&lt;p&gt;Data regeneration covers costs to rebuild training datasets when data becomes corrupted, poisoned, or drifted beyond usability. This includes expenses for new data collection, labeling, cleaning, and validation. Data regeneration is expensive because high-quality training data is the most time-consuming and labor-intensive component of AI development. Rebuilding a corrupted training dataset can take months and cost more than the original data preparation.&lt;/p&gt;
&lt;p&gt;Algorithm remediation covers engineering costs to retrain models that produce biased or inaccurate predictions. This includes compute resources for retraining, testing expenses for validation, and the data science team time required to diagnose the root cause, design the fix, and verify the corrected model&amp;rsquo;s performance. For complex models, remediation can require multiple retraining cycles.&lt;/p&gt;
&lt;p&gt;Infrastructure overruns cover unexpected cloud computing and storage costs from inefficient AI resource usage. Emergency scaling expenses when systems face performance bottlenecks or capacity issues. AI workloads are computationally intensive and unpredictable. A model retraining job that runs longer than expected, a sudden spike in inference requests, or an unoptimized training pipeline can generate infrastructure costs that significantly exceed budget.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Algorithm remediation is the technical loss type most consistently underestimated in risk assessments. Teams estimate the compute cost of retraining but forget the human costs: the data science team time to diagnose the root cause (which can take weeks for complex model failures), the opportunity cost of pulling those data scientists off other projects, the testing and validation time for the remediated model, and the business cost of operating with a degraded model during the remediation period. When I help organizations estimate algorithm remediation costs, I use a formula that includes compute costs (typically the smallest component), data science team labor at fully loaded cost for the estimated remediation duration, lost productivity for the business processes that depend on the model during remediation, and any expedited procurement costs for additional compute resources or external expertise. The total is typically three to five times the compute cost alone.&lt;/p&gt;
&lt;h3 id="domain-3-operational-losses"&gt;Domain 3: Operational Losses&lt;/h3&gt;
&lt;p&gt;Five loss types address the business impact of AI failures on operations.&lt;/p&gt;
&lt;p&gt;Decision errors cover financial losses from incorrect AI-driven business decisions made at scale. This includes costs of resource misallocation in operations, investments, or strategic planning based on flawed AI recommendations. The defining characteristic of decision error losses is scale. An AI system making thousands of decisions per day can accumulate significant losses before the error pattern is detected.&lt;/p&gt;
&lt;p&gt;Operational inefficiency covers manual intervention costs when staff must correct or override AI outputs. Lost productivity from rework and staff time diverted to address AI failures. This loss type captures the ongoing drag on organizational performance that occurs when an AI system works poorly but not badly enough to take offline.&lt;/p&gt;
&lt;p&gt;Development waste covers write-offs of failed AI projects that never reach production deployment. Sunk costs in licenses, development efforts, and procurement that yield no value. Industry estimates suggest that between 60% and 85% of AI projects fail to reach production. Each failed project represents development waste that should be included in the organization&amp;rsquo;s AI loss profile.&lt;/p&gt;
&lt;p&gt;Business disruption covers revenue loss during downtime when AI-dependent processes stop functioning. Emergency replacement costs and lost transactions from service interruptions. This loss type is particularly relevant for organizations where AI systems sit in the critical path of revenue-generating processes.&lt;/p&gt;
&lt;p&gt;Provider switching covers contract termination fees and cancellation penalties with current AI vendors. Migration costs, integration expenses, and negotiation time for new provider onboarding. This loss type is often triggered by other incidents, such as a vendor&amp;rsquo;s quality declining, a security breach at the vendor, or a strategic decision to reduce vendor dependency, but the switching costs themselves represent a distinct financial impact.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Development waste is the operational loss type with the highest aggregate financial impact across most organizations I work with, and it is almost never included in AI risk assessments because it is treated as a project management issue rather than a risk management issue. When I aggregate the fully loaded costs of failed AI projects across an organization, including salaries, compute resources, license fees, and opportunity costs, the total frequently exceeds the organization&amp;rsquo;s estimated exposure from all other AI risk scenarios combined. Include development waste in your loss taxonomy. Estimate it by multiplying the average fully loaded cost of an AI project by the historical failure rate. If you do not track your AI project failure rate, start. That number alone will change how your organization evaluates AI investments.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/urban-tech-fusion.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h3 id="domain-4-revenue-losses"&gt;Domain 4: Revenue Losses&lt;/h3&gt;
&lt;p&gt;Two loss types address top-line financial impact.&lt;/p&gt;
&lt;p&gt;Customer churn covers lost revenue from customers leaving after negative AI experiences or failures. Acquisition costs for replacing churned clients and margin erosion from retention efforts. This loss type has a compounding effect because the cost of acquiring a new customer is typically several times the cost of retaining an existing one.&lt;/p&gt;
&lt;p&gt;Reputation damage covers brand value decline and crisis management costs following publicized AI incidents. Lost business opportunities and reduced market position from negative media coverage. This is the loss type most organizations acknowledge but least effectively quantify.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Reputation damage is the loss type I spent the most time helping organizations quantify, because it is the one where calibrated estimation is most valuable and most difficult. The approach that works is decomposition. Do not try to estimate &amp;ldquo;reputation damage&amp;rdquo; directly. Instead, estimate its measurable downstream effects. How many deals in the pipeline would be delayed or lost? (Estimate the pipeline value at risk.) How much would customer acquisition costs increase, and for how long? (Estimate the increment times the acquisition volume times the duration.) How much additional spending on PR and crisis management would be required? (Get a range from your communications team.) What revenue from existing contracts would be at risk of non-renewal? (Estimate the percentage of contracts with reputation-sensitive renewal decisions.) Add these components together. The total is more defensible than any direct estimate of &amp;ldquo;reputation damage&amp;rdquo; and more useful for risk quantification.&lt;/p&gt;
&lt;h2 id="connecting-incidents-to-losses-the-traceability-requirement"&gt;Connecting Incidents to Losses: The Traceability Requirement&lt;/h2&gt;
&lt;p&gt;The two taxonomies in this post are designed to work together. Each incident type produces one or more direct loss types. Mapping these connections creates the traceability needed for effective risk quantification.&lt;/p&gt;
&lt;p&gt;Take a concrete example. The incident type &amp;ldquo;bias in AI outputs&amp;rdquo; (Discrimination category, internal and external) can produce the following direct losses: regulatory fines (if the bias violates the EU AI Act or fair lending laws), legal compensations (if affected individuals or groups file lawsuits), algorithm remediation (costs to diagnose and fix the biased model), control remediation (costs to build governance controls that should have prevented the bias), customer churn (if the affected population includes customers who leave), and reputation damage (if the bias becomes public).&lt;/p&gt;
&lt;p&gt;Each of these loss types has a different magnitude, different timing, and different probability. Regulatory fines are large but require a regulatory investigation, which may take months. Legal compensations can exceed fines but require plaintiffs to organize and file. Algorithm remediation costs are incurred immediately but are typically the smallest component. Reputation damage may or may not materialize depending on media attention.&lt;/p&gt;
&lt;p&gt;Without this incident-to-loss mapping, your risk quantification combines everything into a single &amp;ldquo;impact&amp;rdquo; number that is neither precise enough for Monte Carlo simulation nor useful enough for control investment decisions.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Build an incident-to-loss mapping matrix for every AI system in your portfolio. Down the left side, list every applicable incident type from this taxonomy. Across the top, list every applicable direct loss type. In each cell, indicate whether the incident could produce that loss type, and if so, provide a rough magnitude range. This matrix becomes the foundation for your FAIR-based risk quantification. When you estimate the impact component of a risk scenario, you are not estimating a single number. You are estimating the aggregate of all applicable loss types for the specific incident. This granularity dramatically improves the quality of Monte Carlo simulation inputs and the credibility of the outputs.&lt;/p&gt;
&lt;h2 id="internal-versus-external-why-the-distinction-matters"&gt;Internal Versus External: Why the Distinction Matters&lt;/h2&gt;
&lt;p&gt;The taxonomy classifies each incident type as producing internal losses, external losses, or both. This classification is not academic. It determines which assessment methodology applies.&lt;/p&gt;
&lt;p&gt;Internal losses are costs borne by the organization. They are addressed through risk assessments that quantify exposure to the organization and inform control investment decisions. When you run a Monte Carlo simulation to calculate annualized loss exposure, you are modeling internal losses.&lt;/p&gt;
&lt;p&gt;External losses are harms borne by individuals, communities, or society. They are addressed through impact assessments that evaluate potential harm to affected parties and inform responsible AI decisions. External losses may or may not create financial exposure for the organization (through fines, lawsuits, or reputation damage), but they matter independently because they represent real harm to real people.&lt;/p&gt;
&lt;p&gt;Some incident types produce only internal losses. Competitive dynamics, for example, creates risk for the organization through unsafe AI deployment but does not directly harm external parties. Some produce only external losses. Surveillance and monitoring, for example, harms individuals and communities but may not create direct financial losses for the organization until it triggers regulatory action or public backlash.&lt;/p&gt;
&lt;p&gt;Most incident types produce both. Bias in AI outputs, for example, creates internal losses through remediation costs and external losses through discriminatory harm to affected individuals.&lt;/p&gt;
&lt;p&gt;Mature AI risk programs assess both dimensions for every applicable incident type. Immature programs assess only internal losses and are surprised when external harms generate regulatory, legal, or reputational consequences they did not anticipate.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The practical implication of the internal/external distinction is that you need two different assessment processes, and they should involve different people. Internal loss assessment is a financial exercise led by risk managers, using techniques like FAIR quantification and Monte Carlo simulation. External impact assessment is an ethical and societal exercise that should involve ethicists, affected community representatives, legal experts, and domain specialists, not just risk managers. I have seen organizations try to combine both assessments into a single process run by the risk team. The financial analysis crowds out the impact analysis every time. When a risk manager and an ethicist are in the same room, the conversation gravitates toward quantifiable financial exposure because that is what the risk manager knows how to discuss. Keep the assessments separate. Conduct them with different teams. Then combine the findings in a governance review where both perspectives inform the decision.&lt;/p&gt;
&lt;h2 id="cross-cutting-implementation-tips"&gt;Cross-Cutting Implementation Tips&lt;/h2&gt;
&lt;p&gt;Four principles apply across both taxonomies.&lt;/p&gt;
&lt;p&gt;Use the incident taxonomy to audit your risk register. Take every risk in your current AI risk register and map it to the incident types in this taxonomy. If a risk in your register maps to multiple incident types, decompose it. If incident types in this taxonomy have no corresponding risk in your register, you have a gap. This audit typically reveals that existing risk registers are too coarse and miss 40% to 60% of applicable incident types.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When I conduct this audit with organizations, the most common gaps are in the cognitive degradation and value misalignment categories. Risk teams are comfortable identifying bias, security, and privacy risks. They are much less comfortable identifying risks related to overreliance on AI, loss of autonomy, lack of explainability, or accountability gaps. These &amp;ldquo;softer&amp;rdquo; incident types are not soft in their consequences. Lack of accountability contributed to more AI incidents I have investigated than any specific technical failure. Include the full taxonomy in your audit, not just the categories that feel comfortable.&lt;/p&gt;
&lt;p&gt;Use the direct loss taxonomy to improve your quantification. For every risk scenario you quantify, decompose the impact into the specific direct loss types that apply. Estimate each loss type separately using calibrated ranges. Then aggregate them for the total impact distribution. This produces more accurate estimates than a single &amp;ldquo;impact&amp;rdquo; range because subject matter experts can estimate specific loss types more credibly than they can estimate total impact.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When conducting estimation workshops, present loss types one at a time, not all at once. Ask experts to estimate regulatory fine exposure, then legal compensation exposure, then algorithm remediation costs, then customer churn impact, and so on. This prevents anchoring, where the first estimate influences all subsequent estimates, and produces wider, more honest ranges. The first time I tried this approach, the aggregate impact estimate was 2.3 times higher than the single &amp;ldquo;total impact&amp;rdquo; estimate the same experts had provided before decomposition. Decomposition reveals exposure that aggregation hides.&lt;/p&gt;
&lt;p&gt;Update both taxonomies as the AI landscape evolves. New incident types emerge as AI capabilities expand. Generative AI created incident types like hallucination and prompt injection that did not exist five years ago. Autonomous agents will create new incident types that do not exist today. Review and update your taxonomies at least annually, and whenever a significant new AI capability is deployed within your organization.&lt;/p&gt;
&lt;p&gt;Align your taxonomy with regulatory requirements. The EU AI Act, NIST AI RMF, ISO 42001, and ISO 23894 each reference specific types of AI-related harms and losses. Map your taxonomy to the categories used by the regulations that apply to your organization. This ensures that your risk assessments address every category a regulator will ask about and that your documentation uses consistent terminology.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/some-photos-of-googles-new-ironwood-tpu-based-ai-superpods-v0-lhuqos1mfxzf1.webp?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="key-references-and-standards"&gt;Key References and Standards&lt;/h2&gt;
&lt;p&gt;This loss taxonomy draws from and aligns with the following authoritative frameworks.&lt;/p&gt;
&lt;p&gt;ISO/IEC 42001:2023 for AI management system requirements covering governance and accountability for AI-related incidents and losses.&lt;/p&gt;
&lt;p&gt;ISO/IEC 23894:2023 for AI risk management guidance, including classification of AI-specific risks and impacts.&lt;/p&gt;
&lt;p&gt;ISO/IEC 27005:2022 for the information security risk management process, including loss event classification.&lt;/p&gt;
&lt;p&gt;EU AI Act (Regulation 2024/1689) for the regulatory framework defining prohibited practices, high-risk requirements, and penalty structures for AI systems.&lt;/p&gt;
&lt;p&gt;NIST AI RMF (AI 100-1) for the AI risk management lifecycle including harm categorization.&lt;/p&gt;
&lt;p&gt;FAIR (Factor Analysis of Information Risk) for quantitative loss modeling taxonomy and methodology.&lt;/p&gt;
&lt;p&gt;OECD AI Principles for the international framework addressing AI-related societal impacts.&lt;/p&gt;
&lt;p&gt;UNESCO Recommendation on the Ethics of Artificial Intelligence for the broader ethical framework covering cognitive, social, and economic impacts.&lt;/p&gt;
&lt;p&gt;MIT AI Risk Repository for the comprehensive academic catalog of AI risk incident types that informed several categories in this taxonomy.&lt;/p&gt;
&lt;h2 id="making-these-taxonomies-operational"&gt;Making These Taxonomies Operational&lt;/h2&gt;
&lt;p&gt;Organizations that file these taxonomies as reference documents will continue making the same mistakes. Their risk assessments will use generic impact categories that obscure actual exposure. Their incident response plans will not cover incident types they have not named. Their loss estimates will undercount by factors of two to five because they have not decomposed generic &amp;ldquo;impact&amp;rdquo; into specific loss types. When an AI incident occurs, they will discover that they cannot quantify their exposure because they never built the vocabulary to describe it precisely.&lt;/p&gt;
&lt;p&gt;Organizations that operationalize these taxonomies will build risk assessments that distinguish between 37 distinct incident types and 15 direct loss categories. They will estimate exposure with the granularity needed for credible Monte Carlo simulation. They will map incidents to losses to controls, creating traceability that survives regulatory scrutiny. Their boards will understand AI risk in specific financial terms because the risk team can articulate exactly what kinds of costs would appear and on which financial lines.&lt;/p&gt;
&lt;p&gt;The precision of your AI risk management cannot exceed the precision of your loss taxonomy. Name the losses specifically, or accept that your risk numbers are wrong.&lt;/p&gt;
&lt;p&gt;Which loss types in this taxonomy are missing from your current AI risk assessments? Start with the ones you have never estimated. Those are where your biggest quantification gaps live.&lt;/p&gt;</description></item><item><title>The AI Risk Taxonomy Most Organizations Never Build</title><link>https://hwyler.github.io/blog/the-ai-risk-taxonomy-most-organizations-never-build/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/the-ai-risk-taxonomy-most-organizations-never-build/</guid><description>&lt;h1 id="top-risk-scenarios-and-controls-that-actually-protect-your-ai-project"&gt;Top Risk Scenarios and Controls That Actually Protect Your AI Project&lt;/h1&gt;
&lt;p&gt;A risk register with 15 vaguely worded AI risks and a color-coded heat map is not a taxonomy. It is a liability.&lt;/p&gt;
&lt;p&gt;I reviewed an organization&amp;rsquo;s AI risk assessment last year that listed &amp;ldquo;AI bias&amp;rdquo; as a single risk with a &amp;ldquo;medium-high&amp;rdquo; rating. That was it. No decomposition into the dozen distinct ways bias manifests. No distinction between bias in training data, bias from proxy variables, bias from temporal misalignment, or bias from feedback loops. No specific controls mapped to specific failure modes. When their credit model produced discriminatory outcomes six months later, nobody could trace the failure to a gap in their controls because their taxonomy was too shallow to reveal where the gaps were.&lt;/p&gt;
&lt;p&gt;The difference between organizations that manage AI risk effectively and those that just talk about it comes down to granularity. You need a taxonomy that decomposes AI risk into specific, actionable scenarios, each linked to a named control with concrete activities. This post provides exactly that: a structured taxonomy of 100 AI risk scenarios across 14 domains, with recommended controls mapped to COBIT 2019 governance objectives. Every scenario follows a consistent structure: what can go wrong, why it matters, and what to do about it.&lt;/p&gt;
&lt;p&gt;This is a long reference piece. Use it as a working document, not a single-sitting read.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/copenhagen.jpg?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="why-most-ai-risk-taxonomies-fail"&gt;Why Most AI Risk Taxonomies Fail&lt;/h2&gt;
&lt;p&gt;The typical AI risk taxonomy fails for three reasons.&lt;/p&gt;
&lt;p&gt;First, it operates at the wrong altitude. &amp;ldquo;Model risk&amp;rdquo; is not a scenario. It is a category that contains dozens of scenarios, each with different causes, different impacts, and different controls. When you treat a category as a scenario, your controls become generic and your residual risk unmeasurable.&lt;/p&gt;
&lt;p&gt;Second, it ignores organizational and process risks. Most AI taxonomies obsess over technical risks like adversarial attacks and data poisoning while overlooking the governance, people, and operational risks that cause the majority of real-world AI failures. A model that degrades because nobody owns monitoring in production is not a technical failure. It is a governance failure.&lt;/p&gt;
&lt;p&gt;Third, it lacks traceability from risk to control. Identifying a risk without mapping it to a specific, implementable control activity is an academic exercise. The taxonomy must create a direct line from &amp;ldquo;what could go wrong&amp;rdquo; to &amp;ldquo;what are we doing about it&amp;rdquo; to &amp;ldquo;how do we verify it is working.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The taxonomy presented here addresses all three failures. It spans 14 domains from strategy through business continuity, covers 100 distinct scenarios, and links each one to a named control with specific activities. I have organized it to follow the natural lifecycle of AI in an enterprise, from strategic planning through development, deployment, operations, and ongoing governance.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When I first built an AI risk taxonomy for a European bank, I started with the technical risks because that is where the AI team&amp;rsquo;s attention naturally went. We ended up with 40 technical scenarios and 5 organizational ones. After the first major incident, which was caused by unclear model ownership between data science and IT operations, we realized our taxonomy was inverted. The organizational and governance risks caused more actual damage than the technical ones. Start your taxonomy with strategy, governance, and people risks. Then layer in the technical domains. This sequencing forces the right conversations early.&lt;/p&gt;
&lt;h2 id="domain-1-business-value-risks"&gt;Domain 1: Business Value Risks&lt;/h2&gt;
&lt;p&gt;Strategy risks sit at the top of the taxonomy because every other risk domain inherits from them. If your AI strategy is flawed, your technical controls cannot compensate.&lt;/p&gt;
&lt;p&gt;Two scenarios define this domain.&lt;/p&gt;
&lt;p&gt;The first is strategy deficiency. Wasted resources and reputational damage may occur when an organization lacks a clear enterprise-wide AI strategy, leading to inefficient investments and potential misuse of AI. This is a Priority 1 risk.&lt;/p&gt;
&lt;p&gt;The recommended control is an enterprise AI strategy. Develop and put in place a comprehensive AI strategy aligned with overall business objectives. Create clear guidelines for AI adoption and integration across departments. Establish governance structures with defined roles, responsibilities, performance metrics, and risk management protocols. Document policies and maintain evidence of governance through reports and records. Review and update the strategy regularly to reflect changes in technology, business needs, and regulatory requirements. Communicate the strategy across the organization to achieve alignment and stakeholder buy-in.&lt;/p&gt;
&lt;p&gt;The second scenario is misaligned strategy. Missed opportunities may occur when insufficient stakeholder engagement leads to AI systems that do not support business goals or expose the organization to unacceptable risks. Also Priority 1.&lt;/p&gt;
&lt;p&gt;The recommended control is stakeholder alignment. Establish stakeholder engagement processes to ensure AI systems align with business goals. Develop communication protocols and document policies for continuous alignment. Collect and maintain meeting records, stakeholder feedback, and communication logs as evidence.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The strategy risk I see most often is not the absence of a strategy. It is the presence of multiple competing strategies. The data science team has a roadmap. The IT department has an automation strategy. The business units each have their own AI wishlists. These strategies contradict each other in ways nobody notices until budget conflicts or architectural incompatibilities surface months later. Before you write a strategy document, conduct a strategy reconciliation exercise. Collect every existing AI-related plan, roadmap, and initiative list across the organization. Map them on a single page. The conflicts will be immediately visible. Resolve those conflicts first. Then write the unified strategy.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/firefly_small-blooming-azalea-flowers-with-many-small-flotating-dollar-coins-portrayed-in-neo-885492.jpg?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="domain-2-governance-risks"&gt;Domain 2: Governance Risks&lt;/h2&gt;
&lt;p&gt;Governance is where principles become operational. Eight scenarios span this domain, and most organizations have gaps in at least half of them.&lt;/p&gt;
&lt;p&gt;Misaligned ethics is a Priority 1 scenario. Reputational damage may occur when AI decisions conflict with organizational cultural and ethical values, leading to poor decisions, negative public perception, and legal repercussions. The control is responsible AI principles: develop AI ethics guidelines, establish an ethics review board, and integrate ethical considerations into the design, development, and deployment of AI systems.&lt;/p&gt;
&lt;p&gt;Overconfidence in automation is equally critical. Wasted resources result from unrealistic expectations about AI capabilities, leading to disappointment, wasted investments, and erosion of trust. The control is capability and limitations communication: communicate the limitations and potential risks of AI technologies and avoid overstating capabilities to ensure realistic expectations and informed decision-making.&lt;/p&gt;
&lt;p&gt;Governance erosion occurs when AI negatively impacts existing governance mechanisms, reducing control over data processing and increasing breach risk. The control is control integration: update existing governance frameworks to incorporate AI-specific considerations and ensure alignment with established policies and risk management protocols.&lt;/p&gt;
&lt;p&gt;Compliance failure carries the most immediate financial consequences. Regulatory penalties and reputational damage result from non-compliance with internal or external AI requirements. The control is compliance audit: regularly test, audit, monitor, and assess AI system compliance with internal policies, external regulations, and ethical guidelines, and report findings to relevant stakeholders.&lt;/p&gt;
&lt;p&gt;Vendor lock-in limits flexibility when exit strategies for AI systems are absent. The control is exit planning: include exit strategy considerations in the design and procurement of AI systems, ensuring the ability to migrate to alternative providers.&lt;/p&gt;
&lt;p&gt;Three Priority 2 governance scenarios round out this domain. Trust deficit limits innovation when organizations lack trust in AI technologies. The control is an AI framework that documents and communicates limitations and capabilities, provides clear explanations of AI decisions, and establishes processes for independent verification. Communication breakdown results from a lack of common language for AI concepts. The control is AI glossary management: develop and maintain a glossary of AI terms and ensure consistent terminology across the organization. Ownership vacuum leads to unauthorized AI development and security breaches when ownership and operating models are undefined. The control is operating model definition: establish clear roles, responsibilities, and accountabilities for AI initiatives, including appropriate segregation of duties.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The governance risk that causes the most silent damage is the ownership vacuum. I worked with a technology company where three separate teams claimed partial ownership of a production ML model. Data science owned the algorithm. Platform engineering owned the infrastructure. The business unit owned the use case. Nobody owned the model in production. When performance degraded, each team assumed another team was monitoring it. The model served degraded predictions for 11 weeks before a customer complaint triggered investigation. Fix this by creating a RACI matrix for every production AI system that names one individual, not a team, as the accountable party for model performance in production. One name. Not a distribution list.&lt;/p&gt;
&lt;h2 id="domain-3-decision-making-risks"&gt;Domain 3: Decision-Making Risks&lt;/h2&gt;
&lt;p&gt;Two Priority 1 scenarios address how AI risk integrates into enterprise decision-making.&lt;/p&gt;
&lt;p&gt;Risk integration gap occurs when organizations fail to integrate risk assessment and controls into the AI framework. Unidentified or unmitigated risks result, along with potential compliance violations and reputational damage. The control is mandatory risk and impact assessments: create and enforce a comprehensive policy mandating systematic identification, evaluation, and management of AI-related risks. Include quantitative model impact assessments with statistical analyses of threat prevalence and potential losses. Integrate these processes with existing framework protocols covering confidentiality, integrity, availability, compliance, contracts, and responsible AI principles.&lt;/p&gt;
&lt;p&gt;AI model risk exposure results from outdated quantitative model risk management practices. The control is a quantitative risk model: develop and maintain up-to-date practices for AI model risk management, conduct impact assessments and statistical analyses to evaluate accuracy and reliability, address model metrics and acceptance criteria, and incorporate regular reviews of data, operational practices, and cybersecurity controls.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Most organizations attempt to integrate AI risk into their existing risk framework by adding a few AI scenarios to their enterprise risk register. This approach fails because the existing register was designed for risks that behave differently. AI risks are dynamic. A model that was within tolerance last quarter may be outside tolerance this quarter because the underlying data distribution shifted. Instead of simply adding AI rows to your existing register, create a parallel cadence of AI-specific risk reviews that feed into the enterprise register. Monthly AI risk reviews that update quarterly enterprise risk reports. This gives AI risks the attention frequency they require while maintaining integration with enterprise governance.&lt;/p&gt;
&lt;h2 id="domain-4-people-risks"&gt;Domain 4: People Risks&lt;/h2&gt;
&lt;p&gt;Three scenarios cover the human element of AI risk.&lt;/p&gt;
&lt;p&gt;Resource misalignment (Priority 1) occurs when unclear resourcing requirements in the AI strategy lead to staffing inefficiencies. The control is staff planning: define and document human resource requirements, including recruitment, role profiles, training, retention strategy, and third-party involvement, in alignment with the AI strategy and roadmap.&lt;/p&gt;
&lt;p&gt;Talent flight (Priority 1) results when poor development and retention of human talent produces AI solutions misaligned with organizational values. The control is talent alignment: establish HR processes to recruit, develop, and retain talent aligned with the AI strategy, including continuous professional development and performance evaluations.&lt;/p&gt;
&lt;p&gt;Knowledge deficit (Priority 1) creates ineffective AI operations and poor incident response when IT knowledge is not retained and developed. The control is knowledge continuity: assign and document specific individuals to fulfill business-as-usual roles and sustainment functions. Ensure ongoing knowledge retention through formal knowledge management practices, continuous training, and documentation of key processes and incidents.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The knowledge deficit risk is particularly dangerous with AI systems because the knowledge required is specialized and often held by a single individual. I have seen organizations where one data scientist understood the feature engineering pipeline, and when that person left, nobody could retrain the model. The entire production system became fragile overnight. For every critical AI system, maintain a &amp;ldquo;bus factor&amp;rdquo; register. For each key knowledge area, list how many people can perform the function. If the number is one, you have a Priority 1 risk that requires immediate cross-training or documentation. Yeah, this sounds obvious. But count how many of your production AI systems depend on a single person&amp;rsquo;s knowledge. The number will concern you.&lt;/p&gt;
&lt;h2 id="domain-5-architecture-risks"&gt;Domain 5: Architecture Risks&lt;/h2&gt;
&lt;p&gt;Architecture risks span seven scenarios across three priority levels.&lt;/p&gt;
&lt;p&gt;Unexplainability (Priority 1) is the inability to understand or explain AI decisions due to missing functionality. The control is explainability by design: integrate explainability as a functional requirement in design, build, and testing phases. Ensure explanations are clear and accessible, with documentation of explainability features and traceability of decision-making processes.&lt;/p&gt;
&lt;p&gt;Incompatibilities (Priority 1) cause operational issues from integration, scalability, and compatibility problems. The control is compatibility testing: develop testing procedures ensuring the AI model is compatible with the production environment, scalable to meet business needs, and integrated with other systems. Perform thorough compatibility testing across software, hardware, and network environments. Conduct scalability assessments including stress testing. Develop standardized integration protocols covering data formats, API usage, and security requirements.&lt;/p&gt;
&lt;p&gt;Misaligned architecture (Priority 2) prevents unified automation when AI architecture is undefined. The control is architecture alignment: define and document an enterprise AI architecture including preferred technologies, design concepts, logging protocols, security controls, and monitoring requirements.&lt;/p&gt;
&lt;p&gt;Segregation deficiency (Priority 2) creates security and data integrity losses in cloud or multi-tenant environments. The control is architectural segregation: define IT architecture principles enforcing segregation of AI system components and data from other infrastructure.&lt;/p&gt;
&lt;p&gt;Unavailability (Priority 2) disrupts business operations due to insufficient AI system resilience. The control is high availability: define and monitor availability metrics, establish redundancy plans, and test for system reliability.&lt;/p&gt;
&lt;p&gt;Ineffective security (Priority 3) results from failing to embed security by design. The control is security by design: incorporate security principles into the development methodology, ensuring all components adhere to established security standards.&lt;/p&gt;
&lt;p&gt;License noncompliance (Priority 3) creates legal and financial exposure. The control is license management: establish a license management system ensuring appropriate licenses and timely renewals for all AI system components.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Explainability by design is the architecture control most frequently treated as an afterthought. Teams build complex ensemble models or deploy large language models, and only when a regulator or auditor asks &amp;ldquo;how does this model make decisions&amp;rdquo; do they realize explainability was never a requirement. Retrofitting explainability onto a deployed model is expensive and sometimes impossible. Add explainability to your definition of done for model development. If the development team cannot demonstrate how the model produces its outputs before deployment, the model does not deploy. This one requirement, enforced consistently, prevents an entire category of compliance and trust risks.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/547784213_3082409608585447_5872836174410763975_n.jpg?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="domain-6-lifecycle-risks"&gt;Domain 6: Lifecycle Risks&lt;/h2&gt;
&lt;p&gt;This is the largest domain, spanning 10 scenarios, because the AI lifecycle from data to deployment contains the most failure points.&lt;/p&gt;
&lt;p&gt;Poor hypothesis (Priority 1) produces unreliable outcomes from inadequate governance around hypothesis development. The control is hypothesis testing: establish governance controls requiring approval of hypotheses based on predefined criteria, with regular reviews for ongoing relevance.&lt;/p&gt;
&lt;p&gt;Poor algorithms (Priority 1) leads to ineffective performance. The control is algorithmic controls: develop governance policies for algorithm development and maintenance, including validation and testing procedures with regular reviews.&lt;/p&gt;
&lt;p&gt;Flawed logics (Priority 1) produces unreliable outputs from inaccurate model parameters. The control is logic validation: establish rigorous logic validation and testing procedures, with approval requirements for any changes to model logic.&lt;/p&gt;
&lt;p&gt;Data accuracy failures (Priority 1) cause erroneous outputs. The control is data accuracy verification: develop and enforce data accuracy verification standards including validation, error detection, and correction before and during model use, with continuous monitoring and automated alerts.&lt;/p&gt;
&lt;p&gt;Six Priority 2 scenarios complete this domain. Unallocated roles create regulatory and operational failures through unclear data governance responsibilities. The control is data ownership: establish roles including data owners and stewards with regular audits. Data corruption results from unintended interactions between AI and other systems. The control is data integrity monitoring: implement controls to monitor data interactions with automated alerts for corruption incidents. Integration failure occurs from corrupted data inputs or outputs between systems. The control is integration testing: develop strict testing procedures with continuous monitoring for data anomalies. Poor data results from inadequate governance over learning and production data. The control is data governance: enforce quality checks with documented standards and periodic audits. Incomplete inputs lead to incorrect outcomes. The control is data completeness control: implement validation processes with protocols for handling incomplete datasets.&lt;/p&gt;
&lt;p&gt;At Priority 3, inaccurate results from models not reflecting underlying parameters are addressed by model validation: comprehensive validation protocols including sensitivity analysis and performance benchmarks.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The lifecycle risk that consistently surprises organizations is data accuracy failure during the transition from development to production. The training data has been cleaned, validated, and verified. The model performs beautifully in testing. Then in production, the live data feed introduces formats, edge cases, and quality issues that never appeared in the training set. I watched a fraud detection model go from 94% accuracy in testing to 71% in the first week of production because the live transaction data contained encoding inconsistencies that the training data had been cleaned of. Build a data reconciliation step between your training pipeline and your production pipeline. Compare distributions, formats, and quality metrics between the data the model was trained on and the data it receives in production. Do this before go-live and continuously afterward.&lt;/p&gt;
&lt;h2 id="domain-7-development-risks"&gt;Domain 7: Development Risks&lt;/h2&gt;
&lt;p&gt;Sixteen scenarios cover the development phase, making it the second largest domain.&lt;/p&gt;
&lt;p&gt;At Priority 1, four scenarios demand immediate attention. Design flaw occurs when poor methodology is not consistently applied. The control is development standards: establish and maintain AI development standards integrated with broader development standards. Inaccurate model results from undefined model universe definition. The control is model universe: define and document the AI model universe including data sources, quality, transformations, and assumptions, updated regularly. Variable misalignment causes incorrect results from mistaking correlation for causality. The control is relationship modeling: establish quality controls ensuring relationships between variables are defined correctly, including interdependencies. Overfitting causes loss of reliability when models perform well on training data but poorly on new data. The control is overfitting mitigation: design algorithms for flexibility with documented testing and validation.&lt;/p&gt;
&lt;p&gt;Learning bias (Priority 1) deserves special attention. Loss of accuracy and reliability occurs from data bias producing discriminatory outcomes. The control is bias mitigation: implement controls considering sensitivities across ethical, political, ethnic, racial, gender, and cultural groups, with documented evaluation processes and evidence of bias checks.&lt;/p&gt;
&lt;p&gt;At Priority 2, five scenarios address operational development risks. To-be inaccuracy results from poor knowledge of desired processes. The control is to-be analysis: maintain documentation of user stories and end-to-end process flows with program sponsor approval. Insufficient segregation occurs when testing environments do not match production. The control is environment segregation: maintain separate development, QA/test, and production environments. Temporal misalignment causes accuracy loss when data time scales conflict. The control is synchronization verification: establish controls ensuring data source alignment with the AI system&amp;rsquo;s time scale. Data duplication produces inflated insights from processing duplicate data. The control is duplication mitigation: implement file and data validation checks with documentation. AutoML issues create complexity and lack of transparency. The control is AutoML guides: develop guidelines for automated machine learning use with regular complexity assessments and explainability tool integration.&lt;/p&gt;
&lt;p&gt;At Priority 3, six scenarios cover remaining development risks. As-is ignorance results from poor knowledge of current processes. The control is as-is analysis: document pre-automation process narratives during the design phase. Undefined controls create vulnerabilities. The control is a control matrix covering all key areas. Control gap occurs when controls are not implemented in the developed solution. The control is control testing to verify processes align with design. Weak traceability compromises logging effectiveness. The control is bot identification with unique identifiers. Improper testing results from insufficient go-live strategy. The control is testing execution with comprehensive documentation. User acceptance deficiency results from inadequate business input. The control is test approvals with documented feedback and sign-off.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Variable misalignment, the correlation-versus-causation problem, is the development risk I find most often in production AI systems. And it is rarely caught by automated testing because the model&amp;rsquo;s statistical metrics look fine. A model might achieve high accuracy by using a variable that correlates with the target in training data but has no causal relationship. When the correlation breaks, which it eventually does, the model fails silently. The most effective countermeasure I have found is a mandatory &amp;ldquo;causal review&amp;rdquo; step in the development process where a domain expert, not a data scientist, reviews the feature set and challenges each variable&amp;rsquo;s causal relationship to the outcome. Data scientists are trained to find patterns. Domain experts are trained to question whether those patterns make sense. You need both perspectives before deployment.&lt;/p&gt;
&lt;h2 id="domain-8-project-risks"&gt;Domain 8: Project Risks&lt;/h2&gt;
&lt;p&gt;Five scenarios cover project-level risks.&lt;/p&gt;
&lt;p&gt;Operational misalignment (Priority 2) results from lacking strategic alignment between AI initiatives and organizational strategy. The control is a business case: establish a strategic alignment framework with formal approval and periodic review by stakeholders.&lt;/p&gt;
&lt;p&gt;Problem mismatch (Priority 2) occurs when model design does not match the business problem. The control is iterative development: adopt approaches like Agile for continuous testing and refinement, with prototyping to identify mismatches early.&lt;/p&gt;
&lt;p&gt;Management gap (Priority 2) results from poor project management methodology. The control is program management: implement project timelines, resource allocation, stakeholder engagement plans, and continuous alignment with business requirements.&lt;/p&gt;
&lt;p&gt;Poor benefits (Priority 3) occurs when benefits management fails to track ROI. The control is impact value: develop a benefits management framework with metrics for short, medium, and long-term tracking.&lt;/p&gt;
&lt;p&gt;Assurance deficit (Priority 3) results from lacking independent assurance. The control is assurance: engage an independent function to evaluate AI program setup, with regular reports on quality, costs, benefits, compliance, and internal control.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Problem mismatch is the project risk that wastes the most money. I worked with a retail organization that spent eight months building a demand forecasting model to solve what turned out to be a supply chain visibility problem. The model was technically excellent but solved the wrong problem. Their forecast accuracy improved by 15%, but the real issue was that they could not see inventory positions across warehouses in real time. The fix required a dashboard, not a model. Before approving any AI project, require the project team to answer one question in writing: &amp;ldquo;Why does this problem require machine learning, and what would the non-ML alternative look like?&amp;rdquo; If they cannot articulate why ML is necessary, there is a good chance a simpler solution would be more effective.&lt;/p&gt;
&lt;h2 id="domain-9-operations-risks"&gt;Domain 9: Operations Risks&lt;/h2&gt;
&lt;p&gt;Nine scenarios cover the operational phase where most AI failures actually manifest.&lt;/p&gt;
&lt;p&gt;Performance drift (Priority 2) is the operational risk with the highest real-world impact. Model accuracy degradation from data drift and concept drift occurs when stability checks are insufficient. The control is stability monitoring: implement model stability checks requiring ongoing validation, benchmarking, and performance evaluation to detect drift.&lt;/p&gt;
&lt;p&gt;Resource laxity (Priority 2) results from inadequate control over IT resource usage given AI&amp;rsquo;s unpredictable demands. The control is project monitoring: implement controls to monitor IT resource demands more closely than other systems.&lt;/p&gt;
&lt;p&gt;Error oversight (Priority 2) leads to unauthorized changes and incidents from undetected errors. The control is incident management: establish a consistent approach with clear procedures, timely resolution, and integration with regular incident management.&lt;/p&gt;
&lt;p&gt;Undetected error (Priority 2) causes delayed resolution from lacking procedures. The control is error resolution: perform timely exception processing with issue and performance monitoring.&lt;/p&gt;
&lt;p&gt;Unsupported jobs (Priority 2) results from insufficient job monitoring. The control is job monitoring: monitor system jobs and interfaces ensuring completeness and timeliness.&lt;/p&gt;
&lt;p&gt;Capacity issues (Priority 2) arise when availability and capacity management cannot meet evolving demand. The control is capacity management: implement availability and capacity management with scalability embedded in design.&lt;/p&gt;
&lt;p&gt;At Priority 3, shadow AI is the scenario most organizations underestimate. Inability to ensure AI aligns with strategy and risk appetite occurs when the organization lacks an inventory of all AI solutions. The control is AI inventory: maintain a complete, up-to-date inventory of all AI platforms, solutions, and use cases, including dependencies and ownership.&lt;/p&gt;
&lt;p&gt;IP loss (Priority 3) occurs when AI system intellectual property held by third parties is at risk. The control is IP protection: establish a repository of relevant IP, accessible in-house, secured with regular backups.&lt;/p&gt;
&lt;p&gt;AI component blindness (Priority 3) results from lacking understanding of IT components and relationships. The control is configuration management: establish a configuration management database fed through change management.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Shadow AI is growing faster than most governance teams realize. Every time an employee uses ChatGPT to draft a customer response, builds a quick predictive model in a Jupyter notebook, or connects a third-party AI tool to company data through a browser extension, they create shadow AI. I conducted a shadow AI audit at a financial services firm last year. The governance team believed they had 12 AI systems in production. We found 47 AI tools and models being used across the organization, most without any risk assessment, data governance, or access controls. The 35 unknown systems included four that processed customer PII. Start your shadow AI inventory not by asking teams to self-report, which underestimates the problem, but by auditing network traffic, SaaS subscriptions, cloud resource usage, and API calls for AI-related activity.&lt;/p&gt;
&lt;h2 id="domain-10-monitoring-risks"&gt;Domain 10: Monitoring Risks&lt;/h2&gt;
&lt;p&gt;Four scenarios address the monitoring function that keeps deployed AI systems safe.&lt;/p&gt;
&lt;p&gt;Outcome blindness (Priority 1) occurs when AI system behavior is not monitored against business and ethical requirements. The control is outcome monitoring: implement regular review of AI system outcomes using data analytics to ensure performance aligns with requirements. Maintain audit trails and ensure controls operate at the same pace as monitored activities.&lt;/p&gt;
&lt;p&gt;Monitoring ineffectiveness (Priority 1) reduces operational effectiveness from inadequate monitoring. The control is operational monitoring: develop a real-time monitoring and alerting framework to detect anomalies, establish KPIs and KRIs as the basis for effective monitoring, and trigger alerts followed by documented follow-ups.&lt;/p&gt;
&lt;p&gt;Undetected issues (Priority 1) cause financial losses and compliance fines from lacking post-deployment monitoring. The control is post-deployment monitoring: develop a monitoring framework defining specific metrics, thresholds, and alerts. Implement automated tools for continuous real-time tracking. Define key performance indicators and review them regularly against business objectives.&lt;/p&gt;
&lt;p&gt;Control override (Priority 2) leads to financial loss when automated stop/loss controls fail. The control is automated stop/loss: design controls to halt unintended AI behavior with an override process for exceptions, assessing exceptions against risk appetite and business impact.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The monitoring risk that catches organizations off guard is the gap between monitoring cadence and AI decision speed. I worked with an organization that monitored their AI system&amp;rsquo;s output quality weekly. The system made 50,000 decisions per day. By the time they detected a quality degradation in their weekly review, the system had already made 350,000 decisions at reduced quality. Match your monitoring frequency to your decision frequency. If your model makes real-time decisions, you need real-time monitoring. If your model runs daily batch predictions, daily monitoring may suffice. But weekly monitoring for a real-time system is a control that exists on paper but provides no actual protection.&lt;/p&gt;
&lt;h2 id="domain-11-security-risks"&gt;Domain 11: Security Risks&lt;/h2&gt;
&lt;p&gt;Seven scenarios span security from Priority 1 through Priority 3.&lt;/p&gt;
&lt;p&gt;Lack of auditability (Priority 1) prevents validation of AI outcomes. The control is auditability: securely store and ensure timely retrieval of data and algorithms, comply with data privacy regulations, prevent data context loss, and apply the vault principle.&lt;/p&gt;
&lt;p&gt;Unauthorized access (Priority 2) leads to inappropriate changes to AI learning and processing data. The control is data access: securely configure AI input datasets to prevent unauthorized changes with completeness and accuracy checks.&lt;/p&gt;
&lt;p&gt;At Priority 3, five scenarios address specific security concerns. Security breach results from inconsistent security management. The control is cyber security: apply a consistent approach integrated with regular security processes, aligned with ISO 27001. Malware attack affects AI environment integrity. The control is malware protection: implement protection systems and monitor patches, protecting self-learning components against malicious attacks. Data breach occurs from insecure handling of temporary files. The control is encryption: encrypt code, data storage, and network communications. Vulnerability blindness results from undetected security weaknesses. The control is vulnerability testing: conduct periodic penetration tests and red-team reviews.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The security risk unique to AI that most security teams miss is the attack surface created by the model itself. Traditional security teams protect the infrastructure around the model, the servers, networks, APIs, and databases. But the model is an attack surface too. An adversary who can query a production model thousands of times can extract information about the training data through model inversion attacks. They can find decision boundaries through systematic probing. They can manipulate outputs through carefully crafted inputs. Your security testing must include model-specific attack scenarios, not just infrastructure penetration testing. If your red team does not include someone who understands adversarial machine learning, your testing has a blind spot.&lt;/p&gt;
&lt;h2 id="domain-12-access-control-risks"&gt;Domain 12: Access Control Risks&lt;/h2&gt;
&lt;p&gt;Twelve scenarios cover access management for both human users and automated bots. All are Priority 3, but their aggregate effect is significant.&lt;/p&gt;
&lt;p&gt;These scenarios cover compromised bot accounts, compromised user accounts, excessive bot access, excessive user access, inadequate account provisioning, inadequate access revocation, undetected bot access, undetected user access, excessive privileged access, segregation of duties conflicts, weak authentication, and unauthorized third-party access.&lt;/p&gt;
&lt;p&gt;The controls follow a consistent pattern: bot control and user control for accountability, bot access authorization and user access authorization for least-privilege enforcement, account provisioning for formal approval processes, access revocation for timely deprovisioning, bot access review and user access review for periodic validation, privileged access authorization for restricting powerful accounts, access segregation for preventing conflicts, authentication for strong credential management, and third-party control for extending security standards to external users.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The access control risk specific to AI that most organizations handle poorly is bot account management. When a bot, an automated process, accesses systems, it typically uses a service account. These service accounts often accumulate privileges over time as the bot&amp;rsquo;s functions expand, but nobody conducts the same periodic access reviews for bot accounts that they do for human accounts. I audited one organization where a bot account for a data preprocessing pipeline had accumulated database administrator privileges, access to the production model repository, and write access to the training data store. Nobody had reviewed the bot&amp;rsquo;s access in 18 months. Treat bot accounts with the same access governance rigor as human accounts. Include them in quarterly access reviews. Apply least-privilege principles. Document and approve every privilege.&lt;/p&gt;
&lt;h2 id="domain-13-change-management-risks"&gt;Domain 13: Change Management Risks&lt;/h2&gt;
&lt;p&gt;Seven scenarios address how changes to AI systems introduce risk.&lt;/p&gt;
&lt;p&gt;IT impact assessment (Priority 1) is the most critical. Disruptions to other IT services may occur from AI system changes with insufficient impact analysis. The control is IT impact assessment: mandate thorough impact analysis for all AI changes, focusing on effects on related IT services, requiring integration testing with documented results.&lt;/p&gt;
&lt;p&gt;Inadequate ongoing testing (Priority 1) causes missed defects. The control is testing protocol: establish comprehensive testing protocols for ongoing AI validation with pre- and post-implementation tests executed by independent teams.&lt;/p&gt;
&lt;p&gt;At Priority 2, undetected errors result from inadequate automated monitoring. The control is automated error monitoring: deploy tools that continuously validate the AI system after changes, detecting anomalies in real time.&lt;/p&gt;
&lt;p&gt;Five Priority 3 scenarios cover unauthorized changes, untracked modifications, poor change control, and insufficient validations. Controls include formal change management processes, modification logging, change control procedures, and validation procedures requiring pre-deployment tests across functional, security, and performance criteria.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The change management risk specific to AI that organizations consistently underestimate is the cascading impact of retraining. When a model is retrained on new data, the outputs change. Sometimes subtly, sometimes dramatically. If downstream systems or business processes depend on the model&amp;rsquo;s output characteristics, for example expected score ranges, output distributions, or decision thresholds, retraining can break those dependencies without triggering any traditional change management alerts. Treat model retraining as a change that requires the same impact assessment, testing, and approval as a code deployment. Because functionally, it is one.&lt;/p&gt;
&lt;h2 id="domain-14-third-party-and-business-continuity-risks"&gt;Domain 14: Third-Party and Business Continuity Risks&lt;/h2&gt;
&lt;p&gt;The final domain covers four third-party scenarios and five business continuity scenarios.&lt;/p&gt;
&lt;p&gt;For third parties, black box solution (Priority 2) creates business disruption when the organization cannot understand the AI system&amp;rsquo;s logic. The control is contract review: define intellectual property ownership, include escrow agreements, ensure right to audit, and outline roles and responsibilities. Third-party default (Priority 3) exposes the organization to lower control maturity. The control is due diligence: subject third parties to at least the same level of control as internal operations. Third-party dependency (Priority 3) creates concentration risk. The control is third-party segmentation: identify and categorize suppliers by criticality with contingency plans. Shadow third-party (Priority 3) results from lacking an updated vendor inventory. The control is third-party management: develop a comprehensive inventory integrated into risk and continuity planning.&lt;/p&gt;
&lt;p&gt;For business continuity, inability to recover (Priority 1) causes prolonged disruptions when rollback mechanisms are absent. The control is roll-back: establish mechanisms to identify and recover the last known good AI state, with processes, algorithms, and cleansed data available for rapid retraining. Ineffective backups (Priority 1) results from inability to restore AI services. The control is backup restoration: implement appropriate backup and snapshot procedures, including frequent snapshots of learning data, with ability to roll back completely.&lt;/p&gt;
&lt;p&gt;Ineffective fallback (Priority 2) results from lacking alternative processing facilities. The control is fallback facility: establish alternative processing capabilities with regular risk assessments. Fragility (Priority 3) and ineffective response (Priority 3) address business continuity planning and testing, with controls for continuity planning aligned with ISO 22301 and continuity testing through regular BCP simulations.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The business continuity risk most specific to AI is the inability to recover the model&amp;rsquo;s learned state. Traditional systems can be restored from backups because their logic is deterministic. An AI model&amp;rsquo;s &amp;ldquo;logic&amp;rdquo; is its trained weights, which are the product of specific training data processed in a specific sequence with specific hyperparameters. If you lose the trained model and do not have the exact training data, preprocessing pipeline, and training configuration documented and backed up, you cannot recreate it. I have seen an organization lose a production model to a storage failure and spend six weeks recreating it because they had backed up the model artifacts but not the training pipeline configuration. Back up everything: the model, the training data, the preprocessing code, the feature engineering pipeline, the hyperparameter configuration, and the training environment specification. Test restoration by actually rebuilding the model from backups at least annually.&lt;/p&gt;
&lt;h2 id="implementation-tips"&gt;Implementation Tips&lt;/h2&gt;
&lt;p&gt;These four principles apply across all 14 domains and 100 scenarios.&lt;/p&gt;
&lt;p&gt;First, prioritize by actual exposure, not by perceived sophistication. The scenarios rated Priority 1 in this taxonomy are not necessarily the most technically interesting. They are the ones that cause the most organizational damage when they materialize. Strategy deficiency, ownership vacuum, and compliance failure cause more real-world harm than adversarial machine learning attacks. Fund controls for Priority 1 scenarios before you invest in exotic defenses against lower-probability technical attacks.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When presenting this taxonomy to leadership, resist the temptation to lead with the technically impressive scenarios like adversarial attacks or model inversion. Lead with the governance and strategy scenarios that connect to business outcomes leadership already cares about. &amp;ldquo;We lack a defined owner for our production AI models&amp;rdquo; resonates more with a board than &amp;ldquo;we are vulnerable to model extraction attacks.&amp;rdquo; Start with the risks they can feel, then build toward the ones they need to understand.&lt;/p&gt;
&lt;p&gt;Second, map controls to your existing control framework. This taxonomy aligns to COBIT 2019 objectives across four domains: Evaluate, Direct and Monitor (EDM), Align, Plan and Organize (APO), Build, Acquire and Implement (BAI), and Deliver, Service and Support (DSS), plus Monitor, Evaluate and Assess (MEA). If your organization uses a different framework, map these controls to your existing structure. The worst outcome is creating a parallel AI control framework that nobody integrates into operational governance.&lt;/p&gt;
&lt;p&gt;Original implementation tip: When mapping these controls to your existing framework, do not create 100 new control activities. Many of these AI controls are extensions of controls you already have. Data access controls for AI systems should be managed through the same access management processes you use for other systems. Change management for AI should follow the same change management framework with AI-specific additions. Identify which controls are genuinely new (explainability by design, bias mitigation, stability monitoring) and which are extensions of existing controls. New controls need new processes. Extensions need updated procedures. The distinction matters for implementation cost and adoption speed.&lt;/p&gt;
&lt;p&gt;Third, document decisions and rationale, not just outcomes. For every control in this taxonomy, maintain evidence that demonstrates not just that the control exists, but why specific decisions were made. When an auditor or regulator asks why you accepted a particular residual risk, &amp;ldquo;because we assessed it and decided it was within tolerance&amp;rdquo; is insufficient. They need to see the assessment, the alternatives considered, and the governance approval.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Create a standard decision record template with five fields: the decision, the alternatives considered, the rationale for selection, the assumptions that must remain valid, and the conditions that would trigger reassessment. Use this template for every Priority 1 and Priority 2 control decision. It takes five minutes per decision and saves hours of reconstruction during audits. I have seen organizations that adopted this practice clear regulatory examinations in half the time of those that relied on informal documentation.&lt;/p&gt;
&lt;p&gt;Fourth, reassess at a cadence that matches your risk velocity. AI risks change faster than traditional IT risks. Models degrade, data drifts, new attack techniques emerge, and regulations evolve. A taxonomy that is reviewed annually is a taxonomy that is wrong for 11 months of the year. Review Priority 1 controls quarterly, Priority 2 semi-annually, and Priority 3 annually at minimum. Update the taxonomy itself whenever a new risk scenario materializes that is not covered.&lt;/p&gt;
&lt;h2 id="references-and-standards"&gt;References and Standards&lt;/h2&gt;
&lt;p&gt;This taxonomy draws from and aligns with the following authoritative frameworks.&lt;/p&gt;
&lt;p&gt;ISO/IEC 27005:2022 for the information security risk management process structure.&lt;/p&gt;
&lt;p&gt;ISO/IEC 23894:2023 for AI-specific risk management guidance.&lt;/p&gt;
&lt;p&gt;ISO/IEC 42001:2023 for AI management system requirements covering governance, ethics, and accountability.&lt;/p&gt;
&lt;p&gt;COBIT 2019 for IT governance and management objectives, providing the control mapping framework used throughout this taxonomy.&lt;/p&gt;
&lt;p&gt;NIST AI RMF (AI 100-1) for the AI risk management lifecycle framework.&lt;/p&gt;
&lt;p&gt;EU AI Act (Regulation 2024/1689) for risk-based regulatory requirements governing AI systems in EU markets.&lt;/p&gt;
&lt;p&gt;ISO 22301 for business continuity management systems referenced in the continuity domain.&lt;/p&gt;
&lt;p&gt;ISO/IEC 27001 for information security management systems referenced in the security domain.&lt;/p&gt;
&lt;p&gt;MITRE ATLAS for the adversarial threat landscape specific to AI and machine learning systems.&lt;/p&gt;
&lt;p&gt;FAIR (Factor Analysis of Information Risk) for quantitative risk analysis methodology when assessing the scenarios in this taxonomy.&lt;/p&gt;
&lt;h2 id="making-this-taxonomy-work"&gt;Making This Taxonomy Work&lt;/h2&gt;
&lt;p&gt;Organizations that treat this taxonomy as a reference document to satisfy an audit requirement will miss its value entirely. They will have a comprehensive list of 100 scenarios that nobody operationalizes, controls that exist in policy but not in practice, and a false sense of security that evaporates at the first real incident. The taxonomy becomes shelfware, and the organization remains exposed to the same risks it cataloged so carefully.&lt;/p&gt;
&lt;p&gt;Organizations that treat this taxonomy as a living operational tool will use it differently. They will map their existing AI systems against these 100 scenarios to identify gaps. They will prioritize control implementation based on the priority ratings and their own risk appetite. They will assign named owners to each applicable control. They will review and update the taxonomy as new AI capabilities are deployed, new threats emerge, and new regulations take effect. Their risk conversations will be specific, traceable, and grounded in concrete scenarios rather than abstract categories.&lt;/p&gt;
&lt;p&gt;A taxonomy that names 100 things that can go wrong is only useful if it drives 100 decisions about what to do right.&lt;/p&gt;
&lt;p&gt;Which of these 14 domains has the biggest gaps in your organization right now? If you are honest with yourself, I suspect the answer is not the technical domains. It is strategy, governance, or people. Start there.&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and consulting work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative risk modeling, predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and globally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance landscapes.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item></channel></rss>