<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Iso-5338 |</title><link>https://hwyler.github.io/tags/iso-5338/</link><atom:link href="https://hwyler.github.io/tags/iso-5338/index.xml" rel="self" type="application/rss+xml"/><description>Iso-5338</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 12 Mar 2026 00:00:00 +0000</lastBuildDate><image><url>https://hwyler.github.io/media/icon_hu_cd51c91342a84ed6.png</url><title>Iso-5338</title><link>https://hwyler.github.io/tags/iso-5338/</link></image><item><title>Practical Post-Market Monitoring for AI Systems</title><link>https://hwyler.github.io/blog/practical-post-market-monitoring-for-ai-systems/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/practical-post-market-monitoring-for-ai-systems/</guid><description>&lt;h2 id="how-to-build-a-control-program-that-catches-problems-after-launch"&gt;How to Build a Control Program That Catches Problems After Launch&lt;/h2&gt;
&lt;p&gt;Most AI governance programs are strongest before launch and weakest after it. That is backwards. The real risk starts when the system meets live users, changing data, edge cases, workarounds, and business pressure. I have seen AI systems pass pre-launch review cleanly, then drift into risky territory within weeks because usage changed, configs changed, user behavior changed, or the model simply behaved differently at scale. The project team thought governance was done. The hard part had just started.&lt;/p&gt;
&lt;p&gt;That is why post-market monitoring matters. It is the operating discipline that tells you whether the AI system is still performing, still lawful, still useful, and still within its approved boundaries. This post shows you how to turn post-market monitoring into a real workflow using both developer controls and user-side controls, not a passive collection of dashboards and incident tickets.&lt;/p&gt;
&lt;p&gt;Post-market monitoring for AI is the structured practice of tracking, evaluating, and acting on an AI system&amp;rsquo;s behavior once it&amp;rsquo;s operating in the real world. It requires two distinct sets of controls: developer controls managed by internal staff and user controls managed by external parties. This post covers both sets, with the practical guidance I&amp;rsquo;ve gathered from monitoring AI systems across financial services, healthcare, and enterprise technology.&lt;/p&gt;
&lt;h2 id="why-post-market-monitoring-is-where-ai-governance-gets-real"&gt;Why Post-Market Monitoring Is Where AI Governance Gets Real&lt;/h2&gt;
&lt;p&gt;Pre-deployment testing tells you how a model performs under controlled conditions. Post-market monitoring tells you how it performs under real ones. These are very different things.&lt;/p&gt;
&lt;p&gt;Real-world data drifts. User behavior changes. Infrastructure degrades. Access privileges accumulate. New vulnerabilities emerge. Regulatory requirements evolve. Business objectives shift. None of these changes announce themselves. Without structured monitoring, they compound silently until something breaks visibly.&lt;/p&gt;
&lt;p&gt;The EU AI Act mandates post-market monitoring for high-risk AI systems. ISO/IEC 42001 includes ongoing monitoring as a core management system requirement. NIST&amp;rsquo;s AI Risk Management Framework positions monitoring as a continuous function, not a periodic activity. These aren&amp;rsquo;t aspirational recommendations. They reflect hard-won understanding that AI systems degrade in ways traditional software doesn&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;A traditional software application does the same thing on day 1,000 that it did on day 1, assuming no code changes. An AI system doesn&amp;rsquo;t. Its relationship with real-world data means its behavior shifts even when nothing in the system itself has changed. The world changes around it, and its performance changes with it.&lt;/p&gt;
&lt;p&gt;Post-market monitoring catches this drift before it causes harm. It operates through two complementary control sets: developer controls managed by the team that built and maintains the system, and user controls managed by the organizations and individuals who deploy the system in their business contexts. Both are necessary. Neither is sufficient alone.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Establish your post-market monitoring framework before deployment, not after. I know this sounds obvious. But on four of the six AI deployments I&amp;rsquo;ve supported, monitoring was designed after the system went live because &amp;ldquo;we&amp;rsquo;ll figure out monitoring once we see how it behaves in production.&amp;rdquo; That approach guarantees a blind period where the system operates without oversight. On one project, that blind period lasted 47 days. During those 47 days, a data pipeline error caused 6% of inference requests to receive default outputs instead of model predictions. No user complained because the default outputs were plausible. No alarm fired because no alarm existed. Design your monitoring controls during the development phase, test them in staging, and deploy them alongside the model. The monitoring system should go live the same day the model goes live.&lt;/p&gt;
&lt;h2 id="the-two-party-monitoring-framework"&gt;The Two-Party Monitoring Framework&lt;/h2&gt;
&lt;p&gt;Post-market monitoring requires controls from two distinct parties because each has visibility into different aspects of system behavior.&lt;/p&gt;
&lt;p&gt;The developer, your internal team, has visibility into model internals. They can track algorithmic metrics, monitor infrastructure performance, analyze system logs, review access controls, and test for vulnerabilities. They see the system from the inside.&lt;/p&gt;
&lt;p&gt;The user, your external stakeholders, has visibility into real-world impact. They see incident reports from end users, observe scope drift in how the system is being applied, experience contractual performance gaps, and can measure business value delivery. They see the system from the outside.&lt;/p&gt;
&lt;p&gt;Gaps in post-market monitoring almost always occur at the boundary between these two perspectives. The developer sees that the model is performing within technical parameters. The user sees that business outcomes are declining. Both are looking at the same system and drawing different conclusions because they&amp;rsquo;re measuring different things.&lt;/p&gt;
&lt;p&gt;A complete post-market monitoring program bridges this boundary with shared metrics, regular communication cadences, and defined escalation paths.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Create a shared monitoring dashboard that both developer and user controls feed into. I worked with an organization where the AI vendor tracked 14 internal metrics and the business unit tracked 8 external metrics. Neither party saw the other&amp;rsquo;s metrics. The vendor reported that model performance was stable. The business unit reported that customer complaints about AI-assisted decisions had increased 40%. It took six weeks of finger-pointing before someone put both datasets side by side and discovered that while overall model accuracy was stable, accuracy for a specific product category had degraded by 23%. The vendor&amp;rsquo;s aggregate metrics masked a localized problem that only the user&amp;rsquo;s complaint data could pinpoint. One shared dashboard, reviewed jointly on a biweekly call, would have surfaced this in days rather than weeks.&lt;/p&gt;
&lt;h2 id="developer-controls-tracking-algorithmic-metrics-against-acceptance-objectives"&gt;Developer Controls: Tracking Algorithmic Metrics Against Acceptance Objectives&lt;/h2&gt;
&lt;p&gt;The first and most important developer control is tracking algorithmic metrics against predefined acceptance objectives. This is where your deployment criteria become your monitoring criteria.&lt;/p&gt;
&lt;p&gt;Every AI system should have documented acceptance objectives established before deployment. These typically include accuracy thresholds, precision and recall targets, false positive and false negative rate limits, and fairness metrics across protected demographic groups. Post-market monitoring means measuring these same metrics continuously on production data and comparing them against the predefined thresholds.&lt;/p&gt;
&lt;p&gt;What to track: Set up automated metric computation on production inference data. For a classification model, compute accuracy, precision, recall, F1 score, and demographic parity ratios daily. Compare each metric against its acceptance threshold. Generate automated alerts when any metric falls below threshold or shows a sustained downward trend over a rolling 7-day window.&lt;/p&gt;
&lt;p&gt;The challenge is that production data doesn&amp;rsquo;t come with ground truth labels the way test data does. In many applications, you won&amp;rsquo;t know whether a prediction was correct until days, weeks, or months later, when the actual outcome becomes observable. A loan default prediction isn&amp;rsquo;t validated until the loan either defaults or is repaid. A medical diagnosis isn&amp;rsquo;t confirmed until follow-up testing occurs.&lt;/p&gt;
&lt;p&gt;This means your algorithmic monitoring needs two tracks. A real-time track monitors input data distributions, output distributions, and prediction confidence scores for signs of drift. A delayed track computes accuracy metrics once ground truth becomes available and compares them against acceptance objectives.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Monitor input data distributions as aggressively as you monitor model outputs. The first sign of model degradation is almost always a shift in input data, not a shift in output quality. Output quality degrades as a consequence of input drift, but it degrades with a delay that can mask the problem for weeks. I set up a simple distribution monitoring system that computes the Kolmogorov-Smirnov statistic between today&amp;rsquo;s input distribution and the training data distribution for each feature, daily. When any feature exceeds a predefined divergence threshold, it triggers an investigation. On one deployment, this caught a data provider format change that shifted how income values were reported from annual to monthly figures. The model didn&amp;rsquo;t crash. It just started treating everyone as low-income. The input distribution alert fired on day one of the change. Without it, we would have discovered the problem through output degradation days or weeks later.&lt;/p&gt;
&lt;h2 id="developer-controls-system-health-and-infrastructure-monitoring"&gt;Developer Controls: System Health and Infrastructure Monitoring&lt;/h2&gt;
&lt;p&gt;Three developer controls address the operational health of your AI system: monitoring system uptime and availability, analyzing user activity in usage logs, and monitoring infrastructure and capacity usage.&lt;/p&gt;
&lt;p&gt;System uptime and availability monitoring measures whether the AI system is accessible and responding when users need it. This sounds like basic IT monitoring because it is. But AI systems have availability failure modes that traditional applications don&amp;rsquo;t. A model serving endpoint might be &amp;ldquo;up&amp;rdquo; in the sense that it accepts requests and returns responses, but &amp;ldquo;down&amp;rdquo; in the sense that it&amp;rsquo;s returning cached or default responses instead of actual model predictions because the model loading process failed silently.&lt;/p&gt;
&lt;p&gt;What to put in place: Monitor not just endpoint availability but model health. Include a health check that verifies the correct model version is loaded, that inference is producing outputs within expected ranges, and that the model is actually executing rather than returning fallback responses. A simple canary request, a known input with a known expected output, run every five minutes, catches model loading failures that HTTP health checks miss entirely.&lt;/p&gt;
&lt;p&gt;User activity analysis from usage logs reveals how the system is actually being used. This differs from how it was designed to be used. Usage logs show query volumes, query types, user segments, peak usage patterns, and interaction sequences. They reveal whether users are adopting the system as intended or developing workarounds that indicate usability problems or unintended uses.&lt;/p&gt;
&lt;p&gt;What to track: Log every inference request with a timestamp, user identifier, input summary (respecting privacy requirements), output summary, confidence score, and response time. Analyze these logs weekly for patterns. Look for users who submit the same query repeatedly (suggesting they don&amp;rsquo;t trust the output), users who consistently override model recommendations (suggesting accuracy concerns for their use case), and usage spikes from unexpected user groups (suggesting scope drift).&lt;/p&gt;
&lt;p&gt;Infrastructure and capacity monitoring tracks compute resources, memory usage, GPU utilization, and storage consumption during production operation. AI systems have different resource profiles than traditional applications. A model that runs efficiently on average may spike to 400% GPU utilization during batch processing windows. A vector database that grows with every interaction will eventually exceed storage limits if not monitored.&lt;/p&gt;
&lt;p&gt;Original implementation tip: The usage log analysis is the developer control that produces the most actionable insights per hour of effort invested. I spent years focusing primarily on algorithmic metrics and infrastructure monitoring. Then a colleague suggested we analyze usage patterns. Within the first week of systematic log analysis, we discovered that 34% of queries to our document classification model came from a department that wasn&amp;rsquo;t in our intended user list. They were using the model to classify customer complaints, a use case we&amp;rsquo;d never tested for and that the model wasn&amp;rsquo;t validated to handle. The model was producing classifications for these inputs, but with significantly lower confidence scores than for its intended document types. Without usage log analysis, this scope drift would have continued indefinitely, with a department making operational decisions based on unvalidated model outputs.&lt;/p&gt;
&lt;h2 id="developer-controls-access-security-and-compliance"&gt;Developer Controls: Access, Security, and Compliance&lt;/h2&gt;
&lt;p&gt;Three developer controls address the security and compliance dimensions of post-market monitoring: reviewing and certifying access privileges, performing regular penetration tests and red-team exercises, and conducting compliance audits with external auditors.&lt;/p&gt;
&lt;p&gt;Access privilege review ensures that the right people have the right access to AI system components over time. Access privileges accumulate. The data scientist who needed full model access during development may not need it during production operation. The contractor who was granted temporary API access for integration testing may still have that access six months later. The service account created for a one-time data migration may still have write access to the production training data store.&lt;/p&gt;
&lt;p&gt;What to put in place: Conduct quarterly access reviews for all AI system components. This includes model artifact repositories, training data stores, inference API credentials, monitoring dashboards, and model management interfaces. For each credential, verify that the person or service still needs the access, that the access level is appropriate for their current role, and that the credential hasn&amp;rsquo;t been shared or compromised. Certify active access and revoke everything else.&lt;/p&gt;
&lt;p&gt;Penetration testing and red-team exercises test your AI system&amp;rsquo;s security posture under adversarial conditions. Standard penetration testing covers infrastructure vulnerabilities. AI-specific red-team exercises cover model-specific attack vectors: prompt injection, model extraction, training data reconstruction, adversarial evasion, and safety filter bypassing.&lt;/p&gt;
&lt;p&gt;What to put in place: Schedule infrastructure penetration tests quarterly and AI-specific red-team exercises semi-annually. After any major model update, architecture change, or deployment expansion, run an additional targeted assessment. Track findings in a persistent tracker and verify remediation in subsequent assessments.&lt;/p&gt;
&lt;p&gt;Compliance audits by external auditors provide independent verification that your AI system meets regulatory and standards requirements. Internal monitoring tells you what you think your compliance posture is. External audits tell you what it actually is.&lt;/p&gt;
&lt;p&gt;What to put in place: Engage an external auditor with AI-specific expertise annually. The audit should cover data handling practices, bias and fairness assessments, documentation completeness (model cards, impact assessments, risk registers), incident response readiness, and regulatory compliance across deployment jurisdictions. Address findings within defined timeframes and track closure.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Access privilege accumulation is the security risk that everyone acknowledges and nobody consistently addresses. I&amp;rsquo;ve conducted access reviews on AI platforms where 40% of active credentials belonged to people who had changed roles or left the organization. One production model serving endpoint had 23 API keys with full access. Only 8 were actively used. The other 15 were orphaned credentials from previous integration projects. Any one of them could have been used to query the model, extract its behavior, or submit adversarial inputs. We revoked the 15 unused keys. Three teams immediately reported that their integrations broke, which meant they were using credentials we had no record of. That&amp;rsquo;s the part that keeps me up at night. The credentials you don&amp;rsquo;t know about are the ones that create real exposure. Build automated credential inventory that cross-references every active key against an approved integration registry. Flag any credential not in the registry for immediate investigation.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/financial-analyst-working-late.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="user-controls-incident-reports-and-risk-reassessment"&gt;User Controls: Incident Reports and Risk Reassessment&lt;/h2&gt;
&lt;p&gt;User controls provide the external perspective that developer controls cannot. Four user controls address reactive monitoring: reviewing incident and misuse reports, reassessing requirements based on new risks, collecting end-user feedback, and monitoring scope drift.&lt;/p&gt;
&lt;p&gt;Incident and misuse report review is the user&amp;rsquo;s primary mechanism for communicating AI system problems back to the developer. End users encounter system behaviors that internal monitoring may not flag: outputs that are technically within accuracy thresholds but practically wrong for a specific context, interactions that feel biased even if aggregate fairness metrics look acceptable, and use patterns that suggest the system is being misused by other users.&lt;/p&gt;
&lt;p&gt;What to put in place: Establish a structured incident reporting process. Every report should capture: what happened, when it happened, who was affected, what the user expected versus what the system delivered, and what action the user took in response. Categorize incidents by type (accuracy failure, bias concern, availability issue, misuse observation, safety concern) and severity (critical, major, minor). Review incidents weekly during the first 90 days post-deployment, then biweekly for established systems.&lt;/p&gt;
&lt;p&gt;Risk reassessment on new risks acknowledges that the risk landscape changes after deployment. New attack techniques emerge. Regulatory requirements evolve. The user population shifts. Competitive dynamics change how the system is used. The user organization should reassess AI system risks at defined intervals and whenever significant changes occur in the operating environment.&lt;/p&gt;
&lt;p&gt;What to put in place: Conduct formal risk reassessments quarterly. Each reassessment should ask: Have new vulnerabilities been disclosed for the underlying model or framework? Have regulatory requirements changed in any deployment jurisdiction? Has the user population or use case expanded beyond the original scope? Have any incidents revealed risks not covered in the original risk assessment?&lt;/p&gt;
&lt;p&gt;End-user feedback and survey analysis provides structured input from the people who interact with the AI system daily. In-app surveys, feedback widgets, and periodic structured surveys capture satisfaction, trust, usability, and perceived accuracy.&lt;/p&gt;
&lt;p&gt;What to put in place: Deploy an in-app feedback mechanism that allows users to rate each AI interaction as helpful, unhelpful, or harmful. Run a more detailed survey quarterly that assesses overall satisfaction, trust in AI outputs, perceived accuracy for the user&amp;rsquo;s specific use case, and suggestions for improvement. Analyze feedback for patterns that correlate with specific user segments, use cases, or time periods.&lt;/p&gt;
&lt;p&gt;Original implementation tip: End-user feedback is the most undervalued signal in post-market monitoring. I used to treat it as a customer satisfaction input, useful for product improvement but not critical for compliance or safety monitoring. I was wrong. On one project, a structured quarterly survey revealed that 28% of users in a specific department reported &amp;ldquo;often&amp;rdquo; disagreeing with the AI system&amp;rsquo;s recommendations but following them anyway because &amp;ldquo;the system is supposed to be better than my judgment.&amp;rdquo; That finding exposed an automation bias problem that no algorithmic metric could detect. The model&amp;rsquo;s accuracy for that department&amp;rsquo;s use case was actually lower than the users&amp;rsquo; own judgment, but the users had been trained to defer to the system. We restructured the interface to present the AI recommendation alongside the key factors driving it, allowing users to apply their own expertise. User override rates increased from 3% to 17%, and decision quality, measured by downstream outcomes, improved by 11%. Feedback surveys catch human-system interaction problems that technical monitoring is blind to.&lt;/p&gt;
&lt;h2 id="user-controls-contractual-and-business-value-monitoring"&gt;User Controls: Contractual and Business Value Monitoring&lt;/h2&gt;
&lt;p&gt;Two user controls address the commercial dimension of post-market monitoring: reviewing contractual performance against license and service contracts, and assessing return on investment in business value reviews.&lt;/p&gt;
&lt;p&gt;Contractual performance review verifies that the AI system delivers what the vendor promised. Service level agreements typically specify uptime guarantees, response time thresholds, support response standards, and update frequencies. Post-market monitoring means systematically measuring actual performance against these contractual benchmarks.&lt;/p&gt;
&lt;p&gt;What to track: Build a contractual compliance tracker that lists each SLA metric, the contractual threshold, the measured performance for each reporting period, and the variance. Review this tracker monthly. When performance falls below contractual thresholds, document the shortfall and raise it with the vendor through the defined escalation process. Don&amp;rsquo;t wait for quarterly business reviews to surface SLA violations. By then, you&amp;rsquo;ve accumulated months of substandard performance with limited recourse.&lt;/p&gt;
&lt;p&gt;What to look for beyond SLA metrics: Monitor for contractual obligations that are harder to measure but equally important. Is the vendor providing the promised frequency of model updates? Are security patches being applied within agreed timeframes? Is the vendor maintaining the data handling practices specified in the contract? Are reporting and documentation obligations being met?&lt;/p&gt;
&lt;p&gt;Return on investment assessment in business value reviews determines whether the AI system is delivering the value that justified its deployment. This is the control that connects technical performance to business outcomes and answers the question that executive sponsors actually care about: &amp;ldquo;Is this worth what we&amp;rsquo;re paying for it?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;What to track: Define business value metrics during the project planning phase. These might include: processing time reduction (measured in hours saved per week), cost reduction (measured in dollars saved per quarter), revenue impact (measured in additional revenue attributed to AI-assisted processes), error reduction (measured in rework hours eliminated), and customer satisfaction impact (measured through satisfaction scores for AI-assisted versus non-AI-assisted interactions).&lt;/p&gt;
&lt;p&gt;Conduct formal business value reviews quarterly. Compare actual business outcomes against the projections in the original business case. If the system is delivering 40% of projected value at the 12-month mark, you need to understand why and decide whether to continue, modify, or discontinue.&lt;/p&gt;
&lt;p&gt;Original implementation tip: Scope drift is the user-side monitoring challenge that causes the most damage over time. Scope drift happens when the AI system gradually gets used for purposes beyond its original intended use. A document classification model starts being used for sentiment analysis. A customer service chatbot gets directed at internal HR queries. A fraud detection model gets applied to a new product line it was never validated for. Each individual expansion seems minor. Collectively, they move the system far outside its validated operating envelope. Build a scope drift monitoring process: maintain a living document that records the system&amp;rsquo;s intended uses and approved use cases. Review actual usage patterns quarterly against this document. Any use that doesn&amp;rsquo;t match an approved use case triggers a validation assessment before it&amp;rsquo;s permitted to continue. I&amp;rsquo;ve seen scope drift turn a well-governed AI deployment into an ungoverned one over the course of a single year, one small expansion at a time, with nobody making a conscious decision to operate outside validated boundaries.&lt;/p&gt;
&lt;h2 id="tips-for-post-market-monitoring"&gt;Tips for Post-Market Monitoring&lt;/h2&gt;
&lt;p&gt;These principles apply across both developer and user control sets.&lt;/p&gt;
&lt;p&gt;Original implementation tip on monitoring cadence: Match your monitoring frequency to your risk level, not your convenience. High-risk AI systems making consequential decisions about individuals, such as lending, healthcare, or criminal justice applications, need daily automated monitoring of algorithmic metrics, weekly human review of monitoring outputs, and monthly cross-party review meetings between developer and user. Lower-risk systems, such as internal productivity tools, can operate on weekly automated monitoring, monthly human review, and quarterly cross-party meetings. I&amp;rsquo;ve seen organizations apply the same monitoring cadence to every AI system regardless of risk. Their high-risk systems were under-monitored, and their low-risk systems consumed monitoring resources that produced minimal value. Right-size your monitoring investment to the risk profile.&lt;/p&gt;
&lt;p&gt;Original implementation tip on version change monitoring: Every model version change, configuration change, and infrastructure change should trigger a monitoring verification cycle. Not a full reassessment. A targeted check that confirms monitoring systems are still capturing the right metrics on the right version of the model. I worked with one organization that updated their model from version 4.2 to version 5.0. The monitoring system continued reporting metrics from version 4.2 because the metric computation pipeline hadn&amp;rsquo;t been updated to point to the new model endpoint. For three weeks, the monitoring dashboard showed stable performance for a model that was no longer in production. The new model&amp;rsquo;s actual performance was significantly different. Build a version verification check into your deployment pipeline: after every model update, automatically verify that monitoring systems are connected to the correct model version and producing fresh metrics.&lt;/p&gt;
&lt;p&gt;Original implementation tip on the handoff between developer and user monitoring: Define explicitly what the developer monitors, what the user monitors, and what both parties are responsible for. Document this in a monitoring responsibility matrix (a RACI for monitoring activities) and include it in your vendor agreement or internal operating procedures. The most common post-market monitoring failure I encounter is the assumption gap: the developer assumes the user is monitoring business outcomes, the user assumes the developer is monitoring model fairness, and nobody is monitoring either one. One deployment went 14 months before anyone measured demographic performance disparities because the developer thought &amp;ldquo;that&amp;rsquo;s a business decision&amp;rdquo; and the user thought &amp;ldquo;that&amp;rsquo;s a technical measurement.&amp;rdquo; It was both. And it was nobody&amp;rsquo;s assigned responsibility. The monitoring responsibility matrix eliminates assumption gaps by making every monitoring activity someone&amp;rsquo;s explicit obligation.&lt;/p&gt;
&lt;p&gt;Original implementation tip on when to stop monitoring and retire a system: Post-market monitoring should include defined criteria for system retirement. When should you stop monitoring and decommission the AI system? When accuracy falls below acceptance thresholds and retraining cannot restore performance. When the business value assessment shows negative ROI for two consecutive quarters. When regulatory changes make the system&amp;rsquo;s approach non-compliant without feasible remediation. When the underlying model or framework reaches end-of-life from the vendor. Define these retirement triggers before deployment. Without them, organizations tend to keep underperforming AI systems running indefinitely because nobody has the authority or the criteria to pull the plug. I&amp;rsquo;ve encountered AI systems still in production three years after the team that built them disbanded, with no monitoring, no maintenance, and no documented owner. They continued making decisions that affected real people. Define retirement criteria. Assign someone the authority to enforce them. Monitor accordingly.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/modern-industrial-engineers-at-work-1.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="references-and-authoritative-frameworks"&gt;References and Authoritative Frameworks&lt;/h2&gt;
&lt;p&gt;Your post-market monitoring program should align with these established standards:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act, Article 72 on post-market monitoring obligations for high-risk AI systems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001:2023, AI Management System (monitoring and measurement requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42005, AI Impact Assessment (ongoing monitoring provisions)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework, Measure and Manage functions&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 5338, AI System Life Cycle Processes (post-deployment monitoring)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 27001:2022, Information Security Management (access review and audit requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;OECD AI Principles, particularly accountability and robustness provisions&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;FDA guidance on AI/ML-based Software as a Medical Device (post-market requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ECB guidance on AI in banking supervision (ongoing monitoring expectations)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST SP 800-137, Information Security Continuous Monitoring&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you treat post-market monitoring as a passive reporting exercise, generating dashboards that nobody reviews and filing metrics that nobody acts on, your AI system will degrade in ways you won&amp;rsquo;t detect until an incident forces attention. The model will drift. The access privileges will accumulate. The scope will expand beyond validated boundaries. The business value will erode. And when the regulator, the auditor, or the affected individual asks what you were monitoring and what you did about what you found, your dashboards full of green indicators won&amp;rsquo;t explain why the system was producing biased outputs for the last nine months.&lt;/p&gt;
&lt;p&gt;When you build post-market monitoring as an active, structured, dual-party discipline, with defined metrics tied to acceptance objectives, assigned responsibilities across developer and user organizations, automated alerts tied to action protocols, and regular human review that looks for the patterns automation misses, you create the feedback loop that keeps AI systems trustworthy over time. You catch drift before it becomes degradation. You catch misuse before it becomes a headline. You catch value erosion before it becomes a write-off.&lt;/p&gt;
&lt;p&gt;An AI system without post-market monitoring is a decision-making machine that nobody is watching. Eventually, it will make a decision that someone should have caught.&lt;/p&gt;
&lt;p&gt;Which of your deployed AI systems has the weakest post-market monitoring? Start building the monitoring framework for that system today.&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and consulting work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative risk modeling, predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and globally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance landscapes.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item><item><title>Problem Definition for AI Projects and Use Cases</title><link>https://hwyler.github.io/blog/practical-problem-definition-for-ai-projects/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/practical-problem-definition-for-ai-projects/</guid><description>&lt;h2 id="how-to-choose-the-right-use-case-before-you-waste-time-and-budget"&gt;How to Choose the Right Use Case Before You Waste Time and Budget&lt;/h2&gt;
&lt;p&gt;Most AI projects go wrong before anyone builds a model.&lt;/p&gt;
&lt;p&gt;They go wrong in the problem statement. The team says they want “an AI solution” when what they really have is a workflow delay, a reporting bottleneck, a quality issue, or a staffing constraint. Then they spend months testing tools against a vague ambition, only to discover they never defined the business problem tightly enough to judge whether the solution worked. That is expensive. It is also avoidable.&lt;/p&gt;
&lt;p&gt;A strong AI project starts with problem definition. Not vendor demos. Not model selection. Not prompt experiments. This post shows you how to define the problem properly, screen for feasibility, structure a use case analysis, and avoid the common failure points that lead teams into broad, fuzzy, low-value AI work.&lt;/p&gt;
&lt;p&gt;A RAND Corporation study found that approximately 80% of AI projects fail. The most common reason wasn&amp;rsquo;t technical. The projects failed because the problem they were solving was poorly defined, misaligned with business needs, or better solved without AI.&lt;/p&gt;
&lt;p&gt;This pattern plays out predictably. A team gets excited about a new AI capability. They build a solution. They deploy it. Then they discover that the business process they automated wasn&amp;rsquo;t the bottleneck, or that users don&amp;rsquo;t trust the output, or that a simpler tool would have worked better at a fraction of the cost. The technology worked. The problem definition didn&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;Defining the problem is the most important and most frequently rushed step in any AI project. It determines everything downstream: the data you need, the technology you select, the success metrics you track, and whether anyone actually uses what you build. This post covers the complete problem definition process, from initial business assessment through feasibility evaluation and use case documentation, with the practical controls that prevent the most common failure modes.&lt;/p&gt;
&lt;h2 id="why-problem-definition-fails-the-technology-is-the-first-trap"&gt;Why Problem Definition Fails: The Technology is the First Trap&lt;/h2&gt;
&lt;p&gt;Most AI problem definitions fail because they start with the technology instead of the problem. &amp;ldquo;We need to use generative AI&amp;rdquo; is not a problem statement. &amp;ldquo;We spend 1,200 hours per year manually responding to client due diligence questionnaires, with a 12% error rate and a 9-day average turnaround&amp;rdquo; is a problem statement.&lt;/p&gt;
&lt;p&gt;The difference matters because technology-first framing skips the analysis that determines whether AI is the right solution. When a team starts with &amp;ldquo;we need to use AI,&amp;rdquo; every problem looks like an AI problem. When a team starts with &amp;ldquo;we need to reduce due diligence response time from 9 days to 2 days,&amp;rdquo; they can objectively evaluate whether AI, workflow automation, template standardization, or some combination delivers the best result.&lt;/p&gt;
&lt;p&gt;This trap intensifies during hype cycles. Generative AI&amp;rsquo;s rapid adoption has created organizational pressure to &amp;ldquo;do something with AI&amp;rdquo; that often overrides disciplined problem analysis. Leadership wants AI initiatives on the roadmap. Teams respond by fitting AI to whatever problems are available rather than identifying problems where AI genuinely adds value.&lt;/p&gt;
&lt;p&gt;The antidote is a structured problem definition process with specific gates that force teams to justify why AI is the right approach before any development begins.&lt;/p&gt;
&lt;p&gt;Implementation tip: Before any AI project receives funding or staffing, require the proposing team to answer one question in writing: &amp;ldquo;What happens if we solve this problem without AI?&amp;rdquo; If the answer describes a viable, cost-effective alternative, that alternative should be the default approach. AI should be selected only when it offers a measurable advantage over non-AI solutions. This single gate eliminates a significant percentage of projects that would otherwise consume resources and fail. Many organizations skip this question because it feels like an obstacle to progress. In practice, it protects teams from investing months of effort into AI solutions for problems that a well-designed spreadsheet macro or workflow automation tool could handle in weeks.&lt;/p&gt;
&lt;h2 id="step-1-assess-business-needs-before-starting-ai-projects"&gt;Step 1: Assess Business Needs Before Starting AI Projects&lt;/h2&gt;
&lt;p&gt;Problem definition begins with a thorough assessment of business needs and challenges, conducted before any AI project work starts. This assessment requires input from management, employees, and potentially customers. Each group brings a different perspective on where problems actually exist.&lt;/p&gt;
&lt;p&gt;Management identifies strategic priorities, resource constraints, and organizational goals that AI projects should serve. Employees identify operational pain points, workflow bottlenecks, and repetitive tasks that consume excessive manual effort. Customers identify service quality gaps, response time issues, and unmet needs that affect their experience.&lt;/p&gt;
&lt;p&gt;Three categories of problems are strong candidates for AI solutions.&lt;/p&gt;
&lt;p&gt;First, repetitive tasks consuming excessive manual effort. These are processes where humans perform the same cognitive work hundreds or thousands of times with minimal variation. Document classification, data extraction from forms, standard report generation, and routine customer inquiry responses all fall into this category.&lt;/p&gt;
&lt;p&gt;Second, blockers in workflow initiation. These are bottlenecks where work stalls because it depends on a step that&amp;rsquo;s slow, scarce, or inconsistent. If a compliance review takes 5 days because one specialist must manually review every submission, that bottleneck may be addressable with AI-assisted triage.&lt;/p&gt;
&lt;p&gt;Third, skill bottlenecks requiring specialized capabilities. These are situations where the organization needs capabilities like data analysis, trend visualization, or code generation that require expertise that&amp;rsquo;s scarce or expensive. AI can augment existing team members by handling the technical execution while humans provide judgment and context.&lt;/p&gt;
&lt;p&gt;What to put in place: Build a structured intake process. Create a centralized repository for validated AI use case proposals. Every proposal should include the business problem, the current process, the expected improvement, and a preliminary assessment of whether AI is the right tool. Review proposals against your AI strategy and responsible AI principles before approving development.&lt;/p&gt;
&lt;p&gt;Implementation tip: Start your AI program by educating teams on foundational AI applications before soliciting use case proposals. Teams that don&amp;rsquo;t understand what AI can and cannot do will either propose nothing (because they don&amp;rsquo;t see opportunities) or propose everything (because they overestimate capabilities). Run workshops covering practical applications like research automation, document analysis, and code generation assistance. After education, use case proposals are more realistic and more actionable. Organizations that skip this step and go straight to &amp;ldquo;submit your AI ideas&amp;rdquo; typically receive proposals that are either too vague to evaluate or too ambitious to execute. Foundational education calibrates expectations, and calibrated expectations produce better problem definitions.&lt;/p&gt;
&lt;h2 id="step-2-write-problem-statements-that-are-specific-enough-to-act-on"&gt;Step 2: Write Problem Statements That Are Specific Enough to Act On&lt;/h2&gt;
&lt;p&gt;Vague problem statements produce vague solutions. &amp;ldquo;Improve customer experience with AI&amp;rdquo; gives a development team no actionable direction. &amp;ldquo;Reduce average customer inquiry response time from 48 hours to 4 hours for the 15 most common question categories, which represent 73% of total inquiry volume&amp;rdquo; gives them everything they need to start.&lt;/p&gt;
&lt;p&gt;Five rules produce actionable problem statements.&lt;/p&gt;
&lt;p&gt;Avoid broad or vague formulations. Every problem statement should identify the specific process, the specific pain point, the specific people affected, and the specific outcome desired.&lt;/p&gt;
&lt;p&gt;Clarify assumptions about the problem. Teams frequently carry assumptions that don&amp;rsquo;t align with reality. &amp;ldquo;Our manual process is too slow&amp;rdquo; might be true, but the root cause might be a staffing shortage, not a process design issue. Validate assumptions with data before committing to a solution.&lt;/p&gt;
&lt;p&gt;Break down the problem into manageable steps or processes. Large problems are composed of smaller tasks. Identify which specific tasks within the larger process are the best candidates for AI assistance. Not every step in a workflow needs AI. Some steps need better tooling. Some need process redesign. Some need additional staff.&lt;/p&gt;
&lt;p&gt;Investigate how similar problems were handled before AI. Look at manual processes, prior AI attempts, and published methods as potential starting points. This research prevents teams from reinventing solutions that already exist and reveals approaches that have already been tried and failed, along with why they failed.&lt;/p&gt;
&lt;p&gt;Focus on solving the problem, not on using the latest technology. Let the problem dictate the tools. The question is never &amp;ldquo;How can we use generative AI?&amp;rdquo; The question is always &amp;ldquo;What&amp;rsquo;s the best way to solve this problem?&amp;rdquo; Sometimes the answer is generative AI. Sometimes it&amp;rsquo;s a rules-based system, a database query, or a process change that requires no technology at all.&lt;/p&gt;
&lt;p&gt;Implementation tip: The most reliable way to test a problem statement&amp;rsquo;s quality is to hand it to someone outside the project team and ask them to describe what a successful solution would look like. If their description matches what the project team envisions, the problem statement is clear. If their description diverges significantly, the statement is ambiguous. This takes ten minutes and reveals gaps that days of internal discussion can miss. Ambiguity in problem statements is invisible to the people who wrote them because they share unspoken context. An outsider doesn&amp;rsquo;t have that context, so ambiguity becomes immediately apparent.&lt;/p&gt;
&lt;h2 id="step-3-choose-the-right-tool-for-the-problem"&gt;Step 3: Choose the Right Tool for the Problem&lt;/h2&gt;
&lt;p&gt;The temptation to use generative AI for everything is strong and should be actively resisted. Generative AI excels at specific task categories: natural language understanding and generation, content creation, summarization, and conversational interaction. It performs poorly at other tasks: precise numerical computation, deterministic logic, real-time data processing, and tasks requiring 100% accuracy.&lt;/p&gt;
&lt;p&gt;Consider hybrid solutions that combine generative AI with other tools. A due diligence questionnaire automation system might use generative AI to draft responses, a retrieval system to find relevant source documents, and a rules-based engine to flag questions requiring human review. This combination is often more effective than any single technology alone.&lt;/p&gt;
&lt;p&gt;Evaluate the capabilities of different technologies and choose the ones that best solve the specific problem. A classification task with clear categories and abundant labeled data might be better served by a traditional machine learning model than by a large language model. A data extraction task with structured input formats might be better served by template-based parsing than by AI of any kind.&lt;/p&gt;
&lt;p&gt;Keep customer demands in perspective. Customers and internal stakeholders may request &amp;ldquo;AI-powered&amp;rdquo; solutions because the technology sounds impressive. The priority is delivering a solution that works and meets their needs, regardless of what technology drives it. A non-AI solution that works reliably at lower cost is superior to an AI solution that works inconsistently at higher cost.&lt;/p&gt;
&lt;p&gt;Stay open to non-AI tools for certain aspects of the problem. Many successful &amp;ldquo;AI projects&amp;rdquo; are actually hybrid systems where AI handles 30-40% of the work and conventional software handles the rest. The AI component gets the attention, but the conventional components often deliver more of the value.&lt;/p&gt;
&lt;p&gt;Focus on the end product&amp;rsquo;s capabilities and performance. The success of an AI project is measured by whether it solves the stated problem within the stated constraints, not by how sophisticated its underlying technology is.&lt;/p&gt;
&lt;p&gt;Implementation tip: When evaluating whether to use generative AI, traditional machine learning, or conventional software for a specific task, apply a simple decision filter. Does the task require generating novel content or understanding unstructured language? Consider generative AI. Does the task require classifying, predicting, or scoring based on patterns in structured data? Consider traditional ML. Does the task require applying deterministic rules to structured inputs? Consider conventional software. Many projects that start as &amp;ldquo;generative AI projects&amp;rdquo; end up as hybrid systems because the problem contains tasks from all three categories. Starting with this filter during problem definition prevents the common pattern of forcing generative AI into tasks where it performs worse than simpler alternatives.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/modern-disconnection.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="step-4-feasibility-assessment-before-development-begins"&gt;Step 4: Feasibility Assessment Before Development Begins&lt;/h2&gt;
&lt;p&gt;Every problem definition must include a feasibility assessment that evaluates whether the proposed AI solution can actually be built, deployed, and maintained within organizational constraints. Feasibility covers two dimensions: requirements and assessments.&lt;/p&gt;
&lt;p&gt;Requirements establish the governance gates. The proposed use case must comply with responsible AI principles, the organization&amp;rsquo;s AI strategy, and applicable privacy, continuity, and cybersecurity regulations. This is a pass/fail evaluation. If the use case conflicts with any of these requirements, it should be redesigned or rejected before development resources are committed.&lt;/p&gt;
&lt;p&gt;Assessments evaluate four practical feasibility questions.&lt;/p&gt;
&lt;p&gt;First, is the projected return on investment positive? Estimate both the costs (development, data preparation, infrastructure, ongoing maintenance, monitoring) and the benefits (time savings, error reduction, revenue impact, compliance improvement). If the ROI case is negative or marginal, the problem may be real but the AI solution may not be justified.&lt;/p&gt;
&lt;p&gt;Second, can the complexity and scalability be supported by existing and future infrastructure, data, models, explanatory requirements, and skills? An AI solution that requires capabilities the organization doesn&amp;rsquo;t have and can&amp;rsquo;t reasonably acquire isn&amp;rsquo;t feasible regardless of how well the problem is defined.&lt;/p&gt;
&lt;p&gt;Third, can quality, compliance, and security controls be met? If the use case requires processing sensitive personal data, can data protection requirements be satisfied? If the use case makes decisions affecting individuals, can explainability requirements be met? If the use case requires integration with regulated systems, can compliance controls be maintained?&lt;/p&gt;
&lt;p&gt;Fourth, can the change be managed? This includes addressing both fear of job displacement among employees whose tasks may be automated and fear of missing out among leaders who want AI initiatives regardless of fit. Change management is a feasibility dimension that technical teams frequently overlook.&lt;/p&gt;
&lt;p&gt;Implementation tip: The feasibility dimension most often underestimated is skills availability. Organizations frequently approve AI projects assuming they can hire or train the necessary talent during the development timeline. Industry data consistently shows that AI talent acquisition takes longer and costs more than initial estimates. Assess your current team&amp;rsquo;s capabilities honestly before approving a project. If the project requires skills your team doesn&amp;rsquo;t have, include talent acquisition or training timelines in the project schedule and treat them as dependencies, not assumptions. A project that&amp;rsquo;s technically feasible but talent-infeasible will stall at the same rate as one that&amp;rsquo;s technically impossible.&lt;/p&gt;
&lt;h2 id="documenting-the-use-case-what-a-complete-analysis-form-looks-like"&gt;Documenting the Use Case: What a Complete Analysis Form Looks Like&lt;/h2&gt;
&lt;p&gt;A well-defined problem needs structured documentation. A use case analysis form captures every element required for informed decision-making. The following sections should be completed for every AI project proposal.&lt;/p&gt;
&lt;p&gt;Use case title and objective. Write a clear, specific title and a one-paragraph objective that states what the AI system will do, what manual effort it will reduce, and what quality improvements it will deliver. Example: &amp;ldquo;Automating due diligence questionnaire reporting with AI. Objective: To automate the generation of due diligence questionnaire reports using an AI agent, reducing manual effort and ensuring consistency and accuracy.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Business need. Describe the problem in operational terms. Quantify the pain where possible. Example: &amp;ldquo;We frequently receive due diligence questionnaires from clients, requiring detailed responses on security controls, policies, and procedures. The current manual process is time-consuming, prone to error, and inconsistent across different formats.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Expected user roles. Identify every role that will interact with the AI system and their specific responsibilities. For a due diligence automation system: Security analysts review and finalize AI-generated reports. Compliance officers ensure responses align with regulatory requirements. IT managers oversee integration with existing systems. Each role should be named specifically, not described generically.&lt;/p&gt;
&lt;p&gt;Expected reach. Quantify the internal and external populations affected. Example: &amp;ldquo;Internal teams: 7 employees in security, compliance, and IT departments. External stakeholders: 240 clients receiving due diligence confirmations per year.&amp;rdquo; These numbers establish the scale of impact and inform risk assessment.&lt;/p&gt;
&lt;p&gt;Expected needed data. List every data source the AI system will require, with specifics about volume and content. Example: &amp;ldquo;IT control matrix: 154 security controls and corresponding narratives. Internal policies: 12 security policy documents. Procedures: 23 SOPs with steps and processes followed by the organization.&amp;rdquo; This inventory determines data preparation effort and identifies potential gaps before development begins.&lt;/p&gt;
&lt;p&gt;Implementation tip: The &amp;ldquo;expected needed data&amp;rdquo; section is where use case proposals most frequently underestimate effort. Teams list the data sources they know about and skip the preparation work required to make that data usable by an AI system. A list of &amp;ldquo;12 security policy documents&amp;rdquo; doesn&amp;rsquo;t reveal that 4 of those documents are outdated PDF scans that require OCR processing, 3 contain conflicting information that needs reconciliation, and 2 haven&amp;rsquo;t been reviewed in over a year and may not reflect current practices. For every data source listed, add a data readiness assessment: Is the data current? Is it in a format the AI system can process? Is it complete? Is it consistent with other sources? Does it require any transformation? This assessment typically adds 2-4 weeks to the project timeline. Discovering these issues during development adds 2-4 months.&lt;/p&gt;
&lt;h2 id="documenting-process-changes-and-anticipated-challenges"&gt;Documenting Process Changes and Anticipated Challenges&lt;/h2&gt;
&lt;p&gt;The use case analysis form must capture how the process will change and what challenges are anticipated. These sections prevent the common pattern of documenting the happy path while ignoring the difficult parts.&lt;/p&gt;
&lt;p&gt;As-is process. Document the current process step by step, with enough detail that someone unfamiliar with it could understand the workflow. Example: &amp;ldquo;(1) Clients send due diligence questionnaires in various formats. (2) Security analysts manually review and respond to each questionnaire based on current practices. (3) Responses are reviewed and approved by a compliance officer before submission.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;To-be process. Document the proposed AI-assisted process with the same level of detail. Clearly indicate where AI handles tasks and where humans remain in the loop. Example: &amp;ldquo;(1) Clients send due diligence questionnaires in various formats. (2) The AI agent automatically reviews and responds to each questionnaire based on the control matrix, internal policies, and SOPs. (3) The AI agent&amp;rsquo;s responses are reviewed and validated by the security leader.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Expected changes. Describe the anticipated improvements in specific terms: &amp;ldquo;Significant reduction in time required to generate due diligence reports. Increased consistency and accuracy in responses. Improved efficiency, allowing employees to focus on higher-value tasks.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Expected challenges. Document known difficulties honestly. For a due diligence automation system, realistic challenges include: ensuring the AI agent accurately interprets and extracts relevant data from internal documents, fine-tuning the AI to understand different formats and client-specific requirements, and integrating the AI agent smoothly with existing systems and workflows.&lt;/p&gt;
&lt;p&gt;AI limitations. Document what the AI system will not do well. This section is critical for setting realistic expectations. Example limitations: &amp;ldquo;The AI may struggle with highly nuanced or complex questions requiring deep contextual understanding. Potential for errors if the AI misinterprets data or lacks sufficient context. Dependence on the quality and completeness of input data.&amp;rdquo; Teams that skip this section create an expectation gap between what stakeholders believe the AI will do and what it actually can do. That gap becomes a project risk.&lt;/p&gt;
&lt;p&gt;Implementation tip: Require every use case analysis form to include both the &amp;ldquo;expected challenges&amp;rdquo; and &amp;ldquo;AI limitations&amp;rdquo; sections before approval. These sections are the ones teams most want to skip because they feel like arguments against the project. In practice, they&amp;rsquo;re the opposite. A proposal that honestly documents challenges and limitations demonstrates that the team understands what they&amp;rsquo;re building. A proposal that claims no challenges and no limitations demonstrates that the team hasn&amp;rsquo;t thought carefully enough. Review committees should be more skeptical of proposals with empty limitation sections than proposals with detailed ones. The projects that fail most expensively are the ones where nobody documented what could go wrong.&lt;/p&gt;
&lt;h2 id="defining-success-metrics-that-prevent-ambiguity"&gt;Defining Success Metrics That Prevent Ambiguity&lt;/h2&gt;
&lt;p&gt;Every use case analysis must include success metrics with specific numerical targets. Without defined success criteria, a project can never conclusively succeed or fail. It exists in a permanent state of &amp;ldquo;we&amp;rsquo;re still working on it.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Four categories of success metrics cover the essential dimensions.&lt;/p&gt;
&lt;p&gt;Time saved measures the operational efficiency gain. Example: &amp;ldquo;85% reduction in hours spent generating due diligence reports.&amp;rdquo; This metric requires a documented baseline. If you don&amp;rsquo;t measure how long the current process takes before deploying AI, you can&amp;rsquo;t measure improvement after.&lt;/p&gt;
&lt;p&gt;Accuracy rate measures quality of AI outputs. Example: &amp;ldquo;95% of AI-generated responses pass human review without significant modification.&amp;rdquo; Define &amp;ldquo;significant modification&amp;rdquo; precisely. A typo correction is not significant. Rewriting a substantive response is. Without this definition, the metric becomes subjective and unreliable.&lt;/p&gt;
&lt;p&gt;Customer or stakeholder satisfaction measures the impact on the people receiving AI-assisted outputs. Example: &amp;ldquo;80% positive feedback from clients on quality and timeliness of responses.&amp;rdquo; This metric requires a feedback collection mechanism designed before deployment, not added as an afterthought.&lt;/p&gt;
&lt;p&gt;Adoption rate measures whether target users actually use the system. Example: &amp;ldquo;99% of due diligence questionnaires processed through the AI system within 6 months of deployment.&amp;rdquo; This metric is the ultimate test of whether the problem definition was correct. If users don&amp;rsquo;t adopt the system, either the problem wasn&amp;rsquo;t as painful as believed, the solution doesn&amp;rsquo;t address it adequately, or change management was insufficient.&lt;/p&gt;
&lt;p&gt;Implementation tip: Set success metric targets before development begins and resist the pressure to adjust them downward during the project. Target adjustment is sometimes legitimate, when new information reveals that initial targets were based on incorrect assumptions. But more often, targets get adjusted because the project is underperforming and the team wants to redefine success rather than address the gap. Protect against this by requiring any target adjustment to be approved by the original project sponsor with a documented justification for the change. If the original target was &amp;ldquo;85% reduction in processing time&amp;rdquo; and the team wants to adjust it to &amp;ldquo;50% reduction,&amp;rdquo; the sponsor should understand why and explicitly accept the reduced ambition. This governance prevents the common pattern where projects gradually redefine success until any outcome qualifies.&lt;/p&gt;
&lt;h2 id="piloting-before-scaling-the-sequence-that-works"&gt;Piloting Before Scaling: The Sequence That Works&lt;/h2&gt;
&lt;p&gt;Problem definition should include a deployment strategy. The most reliable approach follows a specific sequence: educate, pilot, validate, scale.&lt;/p&gt;
&lt;p&gt;Pilot solutions addressing repetitive tasks first to demonstrate quick wins. Quick wins build organizational confidence in AI, generate concrete data for ROI calculations, and reveal integration challenges at low risk. A pilot that automates 5% of due diligence responses teaches you more about data quality requirements, user trust dynamics, and accuracy thresholds than months of theoretical analysis.&lt;/p&gt;
&lt;p&gt;Scale validated AI workflows while maintaining audit trails for compliance accountability. Scaling should begin only after the pilot has met its success metrics and the team has documented lessons learned. The audit trail requirement ensures that as the system handles more volume and higher-stakes decisions, every AI-generated output can be traced back to its inputs, the model version that produced it, and the human who reviewed it.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/futuristic-data-display-1.png?w=724" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Implementation tip: Define &amp;ldquo;pilot success&amp;rdquo; criteria before the pilot starts, and make those criteria the gate for scaling. The most common pilot failure mode is indefinite extension. The pilot runs for its planned duration, produces mixed results, and instead of making a go/no-go decision, the team extends the pilot &amp;ldquo;to gather more data.&amp;rdquo; Pilots that get extended once tend to get extended repeatedly, consuming resources without producing a scaling decision. Set clear criteria: &amp;ldquo;The pilot will run for 8 weeks with 50 due diligence questionnaires. If accuracy exceeds 90% and processing time reduction exceeds 70%, we proceed to scaled deployment. If either metric falls short, we conduct a root cause analysis and make a continue/modify/stop decision within 2 weeks.&amp;rdquo; That specificity forces decisions instead of indefinite experimentation.&lt;/p&gt;
&lt;h2 id="cross-cutting-tips-for-ai-problem-definition"&gt;Cross-Cutting Tips for AI Problem Definition&lt;/h2&gt;
&lt;p&gt;These principles apply across every stage of the problem definition process.&lt;/p&gt;
&lt;p&gt;Implementation tip on stakeholder alignment: Present the problem definition document to every stakeholder group before development begins and get their explicit agreement that the problem statement, success metrics, and scope accurately reflect their needs. Misalignment between what the project team thinks the problem is and what stakeholders actually need is the single most common source of AI project failure. This alignment meeting should produce a signed-off document, not a verbal agreement. When priorities shift mid-project (and they will), the signed document provides a reference point for scope discussions. Without it, every stakeholder remembers the problem definition differently, and the project tries to solve multiple unstated problems simultaneously.&lt;/p&gt;
&lt;p&gt;Implementation tip on documenting what you chose not to do: Your use case analysis should include a section on alternatives considered and reasons for rejection. &amp;ldquo;We considered using a template-based system but rejected it because client questionnaire formats vary too widely for template matching. We considered hiring additional analysts but rejected it because the volume is seasonal and full-time hiring isn&amp;rsquo;t cost-effective.&amp;rdquo; This documentation serves two purposes. It demonstrates that the team evaluated alternatives, which satisfies governance requirements. And it creates institutional memory that prevents future teams from revisiting the same options without benefiting from the analysis already performed.&lt;/p&gt;
&lt;p&gt;Implementation tip on the relationship between problem definition and ongoing monitoring: Your success metrics from the problem definition phase should become your post-deployment monitoring metrics. If you defined success as &amp;ldquo;95% accuracy rate on AI-generated responses,&amp;rdquo; that same metric should be tracked continuously after deployment. If you defined success as &amp;ldquo;85% reduction in processing time,&amp;rdquo; that measurement should appear on your operational dashboard. Disconnection between how you defined success and how you monitor the deployed system creates a gap where degradation goes undetected. Design your monitoring framework during problem definition, not after deployment.&lt;/p&gt;
&lt;p&gt;Implementation tip on revisiting problem definitions as projects mature: Problem definitions should be treated as living documents during the early stages of a project. The pilot phase will reveal aspects of the problem that weren&amp;rsquo;t visible during initial analysis. User feedback will surface needs that weren&amp;rsquo;t captured in stakeholder interviews. Data quality assessment will reveal constraints that affect solution design. Schedule a problem definition review at the end of the pilot phase. Update the use case analysis form to reflect what you&amp;rsquo;ve learned. Adjust success metrics if the pilot revealed that initial targets were based on incomplete understanding. This review doesn&amp;rsquo;t weaken the problem definition process. It strengthens it by incorporating real-world evidence.&lt;/p&gt;
&lt;h2 id="references-and-frameworks"&gt;References and Frameworks&lt;/h2&gt;
&lt;p&gt;Your AI problem definition process should align with these established standards and guidelines:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001:2023, AI Management System (planning and context requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42005, AI Impact Assessment (pre-deployment analysis requirements)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework, particularly the Map function&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 5338, AI System Life Cycle Processes (requirements analysis phase)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;OECD AI Principles, particularly the robustness and accountability provisions&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act, Annex IV documentation requirements for high-risk AI system purpose and intended use&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;IEEE 2801-2022, Recommended Practice for Quality Management of Datasets&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;PMI guidance on project scope definition adapted for AI initiatives&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;COBIT 2019 for alignment of AI projects with business governance objectives&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 25010, Systems and Software Quality Requirements (for defining quality-based success metrics)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you treat AI problem definition as a formality, filling in a use case form with vague objectives and optimistic metrics to get budget approval, you set the project up for the most expensive kind of failure: the kind where everything works technically but nothing works practically. The model performs well. Nobody uses it. Or everyone uses it for the wrong thing. Or it solves a problem that wasn&amp;rsquo;t the real bottleneck. And the organization concludes that &amp;ldquo;AI doesn&amp;rsquo;t work for us&amp;rdquo; when the real issue was that the problem was never properly defined.&lt;/p&gt;
&lt;p&gt;When you treat problem definition as the most consequential decision in the AI project lifecycle, with structured assessment, honest feasibility evaluation, specific success metrics, and documented alternatives, you create the foundation for everything that follows. The right problem definition makes technology selection obvious, makes data requirements clear, makes success measurable, and makes the go/no-go decision at each phase defensible. Every hour invested in rigorous problem definition saves multiples of that time in avoided rework, scope creep, and failed deployments.&lt;/p&gt;
&lt;p&gt;The best AI projects don&amp;rsquo;t start with the best technology. They start with the clearest understanding of the problem they need to solve.&lt;/p&gt;
&lt;p&gt;What business problem in your organization are you currently considering for AI? Run it through the feasibility framework in this post before writing a single line of code.&lt;/p&gt;</description></item></channel></rss>