<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ai-Development |</title><link>https://hwyler.github.io/tags/ai-development/</link><atom:link href="https://hwyler.github.io/tags/ai-development/index.xml" rel="self" type="application/rss+xml"/><description>Ai-Development</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Fri, 14 Aug 2026 00:00:00 +0000</lastBuildDate><image><url>https://hwyler.github.io/media/icon_hu_cd51c91342a84ed6.png</url><title>Ai-Development</title><link>https://hwyler.github.io/tags/ai-development/</link></image><item><title>Data Quality Requirements That Decide Whether Your AI Ships or Sinks</title><link>https://hwyler.github.io/blog/data-quality-requirements-that-decide-whether-your-ai-ships-or-sinks/</link><pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/data-quality-requirements-that-decide-whether-your-ai-ships-or-sinks/</guid><description>&lt;p&gt;The validation accuracy means nothing if the training data is broken. I reviewed a production model with 92% validation accuracy. Training data passed schema checks at more than 99%. The missing percent covered one geography, one device type, and one age group. The model had never seen those records. Average quality scores lied to us.&lt;/p&gt;
&lt;p&gt;This article gives you the complete framework: 10 concrete data quality requirements drawn from ISO 5259, ISO 42001, ISO 19157, and NIST guidance. Each requirement includes the controls, metrics, and validation tests that disciplined AI teams run before a single model trains. You will also see exactly where hard gates replace aggregate scores, and why that distinction is the difference between a model that holds up and one that quietly degrades. That distinction saves production systems.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/08/d6d31621-0966-4be5-8ec1-274cfd3fe968.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="the-data-quality-contract-covers-four-stages"&gt;The Data Quality Contract Covers Four Stages&lt;/h2&gt;
&lt;p&gt;Your model inherits every defect in the data it touches. That includes data you never directly inspect. Training data shapes learning. Validation data shapes your confidence. Feedback data shapes future updates. Production usage data determines what actually happens after deployment.&lt;/p&gt;
&lt;p&gt;Treat these four stages as separate contract areas. One dataset can pass a training check and still destroy a production model. Each stage has its own failure modes, its own controls, and its own owner. Confusing them is how governance failures slip through.&lt;/p&gt;
&lt;h3 id="training-data"&gt;Training data&lt;/h3&gt;
&lt;p&gt;Training data is the set of examples, features, and labels used to teach the model. It must be correct, complete, representative, and free of leakage.&lt;/p&gt;
&lt;p&gt;My first rule is deduplication before splitting. Never do it afterward. Duplicate records crossing train and test boundaries create inflated scores. Use group-aware splits so all records from one customer, patient, device, household, author, or organization stay in the same split. A group-aware split forces the model to generalize to new entities rather than memorize familiar ones. If every user in your test set also appears in your training set, your offline evaluation is measuring memorization, not generalization.&lt;/p&gt;
&lt;p&gt;The second rule is leakage detection. A feature derived from information available after the prediction timestamp will produce outstanding offline accuracy and catastrophic production failure. Random splits conceal this problem entirely. Use time-based splits where deployment involves future cases. For a model intended to predict next-day equipment failure, do not include maintenance records created after the prediction timestamp, even if those records improve offline accuracy. Build a leakage review into your feature documentation process. The dangerous leaks are subtle: a field updated at the time of outcome recording, a derived feature that aggregates future events, or an identifier that correlates with outcome because of how data was collected.&lt;/p&gt;
&lt;p&gt;My third role is about representativeness. Draw from a wide range of sources spanning different patterns, perspectives, and scenarios. Use stratified splits so rare classes and important subgroups appear in training with sufficient volume to measure. Do not assume demographic balance proves fairness. Some operational datasets should not mirror population proportions. A model trained to detect rare equipment failures should oversample failure cases, not mirror the natural ninety-nine to one imbalance. Document why your target distribution is appropriate. That justification is both a governance artifact and a defense against audit challenge.&lt;/p&gt;
&lt;p&gt;Label quality deserves its own controls. Record who produced or reviewed each label, when they were created, and which labeling guideline version was active. Use a holdout audit sample that annotators never see during preparation. Double-label a statistically justified sample and calculate inter-annotator agreement using Cohen&amp;rsquo;s kappa or Fleiss&amp;rsquo; kappa. Require independent adjudication for any disputed label. Skipping this audit sample because it feels expensive leads to discovering systematic labeling errors after deployment and spending three times as long rebuilding.&lt;/p&gt;
&lt;h3 id="validation-and-test-data"&gt;Validation and test data&lt;/h3&gt;
&lt;p&gt;Validation and test data measure performance. They must remain independent from training data and cover the same subgroups, edge cases, and failure modes the system will face.&lt;/p&gt;
&lt;p&gt;The first rule is independence. Check for train-test overlap before you trust any score. For language or image models, near-duplicate examples in the test set produce artificially high evaluation scores even when the model has poor generalization. Hash-normalize records to catch exact duplicates. Use similarity matching, MinHash, or embedding similarity to catch near-duplicates. Investigate data augmentation that creates near-duplicates in the test set.&lt;/p&gt;
&lt;p&gt;The second rule is temporal validity. Use time-based splits when deployment involves future cases. Random splits leak future information into training and hide the exact drift that will appear in production. A high validation score from a random split gives false confidence. Use group-based splits where deployment involves new users, new sites, new organizations, or new devices. If every user in your test set also appears in your training set, you are measuring memorization.&lt;/p&gt;
&lt;p&gt;The third rule is subgroup coverage. A model can achieve ninety-two percent accuracy overall while performing at sixty-eight percent for a specific subgroup. That gap will not appear in any aggregate metric. Test intersectional groups where sample sizes permit. Age crossed with gender crossed with region can reveal failure modes that are invisible in single-dimension analysis. Evaluate model outcomes separately by group, subgroup, and intersection. Publish accuracy, false-positive rate, false-negative rate, and calibration separately for each subgroup.&lt;/p&gt;
&lt;p&gt;Build a challenge set from known incidents, complaints, adversarial examples, and expert-defined edge cases. Your standard test set reflects what happened. Your challenge set reflects what can happen. A model that passes the standard test set but fails the challenge set is not ready for production.&lt;/p&gt;
&lt;h3 id="feedback-data"&gt;Feedback data&lt;/h3&gt;
&lt;p&gt;Feedback data includes user corrections, thumb ratings, complaint logs, production labels, and reviewer decisions. It is often dirty, unaudited, and adversarial.&lt;/p&gt;
&lt;p&gt;Treat feedback as untrusted production input. Scan it for secrets, personal data, prompt injection, and toxic content before any reuse. User inputs can contain prompt injection attempts, personally identifiable information, credentials, and malicious content. None of that should enter a retraining pipeline unsanitized. Feedback pipelines are a significant attack surface. A single poisoned feedback item can corrupt an entire retraining cycle.&lt;/p&gt;
&lt;p&gt;Link every feedback item to the model version that generated the output, the specific input, the reviewer decision, and the final disposition. Feedback that cannot be traced to a model version cannot be used to evaluate that model or to construct valid retraining data. Without this linkage, you cannot distinguish between feedback that applies to the old model and feedback that applies to the new one.&lt;/p&gt;
&lt;p&gt;Separate feedback stores from raw data stores. The risk profile of user-generated feedback is fundamentally different from the risk profile of a curated training dataset. Apply purpose limitation. Feedback collected for one model should not automatically flow into another model&amp;rsquo;s training pipeline without explicit review.&lt;/p&gt;
&lt;h3 id="production-or-usage-data"&gt;Production or usage data&lt;/h3&gt;
&lt;p&gt;Production data is what the model receives at inference time. It may drift, fail schema checks, arrive late, or contain out-of-domain inputs.&lt;/p&gt;
&lt;p&gt;Compare production input distributions to the training baseline at least weekly. One distribution shift alert matters more than a quarterly aggregate accuracy report. Use Population Stability Index,
or Wasserstein distance to measure drift. Set retraining or review triggers when drift persists across multiple periods, not just when a single batch looks unusual. A single unusual batch may be noise. Sustained drift means your training distribution and production distribution have separated.&lt;/p&gt;
&lt;p&gt;Store event time and processing time separately. Confusing them hides late-arriving data. If your pipeline processes a transaction at 14:00 that occurred at 08:00, the six-hour gap is only visible if you captured both timestamps. Test late-arriving, duplicated, out-of-order, and replayed events explicitly. These are the failure modes that stress-test assumptions baked into most feature engineering pipelines.&lt;/p&gt;
&lt;p&gt;Monitor out-of-domain inputs. Use applicability-domain or embedding-distance checks to detect unfamiliar inputs that fall outside the distribution the model was trained on. A model deployed in a new geography or on a new device type may receive inputs it has never seen. Detecting that shift before it produces bad decisions is the difference between a controlled rollout and a production incident.&lt;/p&gt;
&lt;p&gt;Keep production data stores separate from training stores. Do not allow a single access control policy to govern both. Production data often contains live sensitive information. Training data should be a governed snapshot with purpose limitations applied. The separation also prevents accidental feedback loops where production outputs get reused as training inputs without the required screening and version linkage.&lt;/p&gt;
&lt;p&gt;Each stage has its own minimum validation focus. Training data requires accuracy, completeness, representativeness, leakage detection, deduplication, label quality, privacy controls, and lineage coverage. Validation and test data require independence, subgroup coverage, stable labels, temporal validity, and contamination checks. Feedback data requires authenticity, authorization, injection screening, label confidence, reviewer agreement, and version linkage. Production data requires schema validity, freshness, drift detection, out-of-domain detection, access control, and incident tracking. Running the same generic test suite across all four stages is how quality gates become theater.&lt;/p&gt;
&lt;h2 id="the-ai-data-requirements"&gt;The AI Data Requirements&lt;/h2&gt;
&lt;p&gt;Industry data readiness frameworks often organize quality around six factors: diversity, timeliness, accuracy, security, discoverability, and consumability. Those factors map directly to the ten requirements below. Representativeness and relevance cover diversity. Timeliness and currentness cover freshness. Accuracy, completeness, and validity cover correctness. Security and compliance cover protection. Traceability and lineage cover discoverability. Consistency and relevance cover consumability.&lt;/p&gt;
&lt;p&gt;The requirements below are ordered by how frequently they cause production failures, not by how often they appear in governance documents.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="1-accuracy"&gt;1. Accuracy&lt;/h3&gt;
&lt;p&gt;Accuracy from algorrithmic training, validation and usage data means values, labels, and annotations correctly represent reality or an authoritative reference. Errors here propagate directly into every prediction the model makes. A credit model trained on mislabeled repayment outcomes learns to approve the wrong borrowers. A medical imaging system trained on incorrect diagnoses becomes dangerous at exactly the moment clinicians trust it most.&lt;/p&gt;
&lt;p&gt;Measurement accuracy and label accuracy are different problems that require different controls. A dataset can contain correct sensor readings with wrong class labels attached to every record. A medical imaging dataset can have pixel-perfect scans with incorrect diagnoses. Treating accuracy as a single dimension means you catch one failure mode while missing the other entirely.&lt;/p&gt;
&lt;p&gt;Data scientists should profile source data before any other quality check. Exploratory data analysis reveals characteristics, completeness, distribution, redundancy, and shape that aggregate quality scores hide. Build data quality rules from that profiling and monitor their efficacy continuously. Do not profile once at project start and assume the source system holds constant.&lt;/p&gt;
&lt;p&gt;Define tolerances by use case before you profile anything. A two-percent rounding error is acceptable in demand forecasting. That same error is not acceptable in a drug dosage recommendation system or a financial reporting model subject to regulatory audit. Write the tolerance down. Make it part of your data quality gate so it cannot be overridden informally when timelines compress.&lt;/p&gt;
&lt;p&gt;The annotation process needs governance separate from technical validation. Record who produced or reviewed each label, when they created it, and which labeling guideline version was active at that time. Without that provenance, you cannot audit a disputed prediction or trace a labeling error back to its source. When a regulator asks which annotator produced a specific label, &amp;ldquo;we used a crowdsourcing platform&amp;rdquo; is not an acceptable answer.&lt;/p&gt;
&lt;p&gt;Use a holdout audit sample that annotators never see during data preparation. Double-label a statistically justified sample of records. Calculate inter-annotator agreement using Cohen&amp;rsquo;s kappa or Fleiss&amp;rsquo; kappa. Require independent adjudication for any label where agreement falls below your defined threshold. The adjudication audit sample feels expensive until you discover a systematic labeling error after deployment and spend three times as long rebuilding the dataset from scratch.&lt;/p&gt;
&lt;p&gt;Enable lineage and impact analysis so data engineers and scientists can see the downstream consequences of changes before they happen. When a source system changes a field definition, you need to know immediately which models depend on it, not six months later when performance unexpectedly shifts.&lt;/p&gt;
&lt;p&gt;For a classifier, one acceptance rule is: at least 98 percent of critical labels must agree with adjudicated expert labels, with no high-severity label error remaining unresolved. The specific percentage depends on your use case and risk profile. What cannot vary is having the rule written down and enforced at the gate, not estimated after training.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="2-completeness"&gt;2. Completeness&lt;/h3&gt;
&lt;p&gt;Completeness measures whether all required records, fields, labels, time periods, and classes are present. Missing data is not a neutral condition. Absence is information, but it is information you did not plan to use. Missing values skew distributions, distort feature importance, and bias model behavior toward the populations and scenarios where data happened to be collected. The model learns what was measured, not what matters.&lt;/p&gt;
&lt;p&gt;The most dangerous completeness failure is concentrated missingness, not distributed missingness. A five percent missing rate overall sounds manageable. A five percent missing rate concentrated entirely within a specific demographic group, geography, or outcome class is a bias problem disguised as a data quality score. Your completeness metrics must break down by subgroup, source, and time period, not just by field and dataset.&lt;/p&gt;
&lt;p&gt;Set stricter thresholds for critical fields than for optional ones. Zero tolerance for missing mandatory identifiers and labels is a reasonable hard gate. A documented, bounded missingness rate for noncritical features is acceptable when the missingness mechanism is understood and recorded. Undocumented missingness is never acceptable, regardless of the rate.&lt;/p&gt;
&lt;p&gt;Imputation does not eliminate the problem. When you fill a missing value, retain an indicator flag showing the value was imputed. That flag is itself a feature the model can use and a governance artifact showing you acknowledged the gap. Imputing without flagging hides the extent of the problem from downstream consumers of the data.&lt;/p&gt;
&lt;p&gt;Measure completeness separately for training, validation, test, feedback, and production data. An aggregate completeness score across all four stages hides the fact that your minority class in the test set might have fifty percent label coverage while your majority class has ninety-eight percent. That imbalance will not show in any headline number.&lt;/p&gt;
&lt;p&gt;Investigate records that appear complete but contain default values. Zero, &amp;ldquo;unknown,&amp;rdquo; &amp;ldquo;N/A&amp;rdquo;, and the Unix epoch date 1970-01-01 are common proxies for missing data that pass completeness checks while carrying no real signal. These records inflate your completeness rate while quietly degrading your model.&lt;/p&gt;
&lt;p&gt;Reconcile dataset counts with source system counts and event logs. If your pipeline received 1.2 million records and your source system logged 1.4 million events, that gap is not a rounding difference. It is a completeness failure with a specific cause that needs investigation before you train on anything.&lt;/p&gt;
&lt;p&gt;A typical data quality gate requires zero missing values for mandatory identifiers and labels, while allowing a documented, bounded rate of missingness in noncritical features. The key word is documented. Undocumented gaps become undocumented assumptions that become production failures.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="3-timeliness-and-currentness"&gt;3. Timeliness and Currentness&lt;/h3&gt;
&lt;p&gt;A weather forecast based on yesterday&amp;rsquo;s conditions is wrong for today&amp;rsquo;s trip. An AI model trained on outdated information produces inaccurate or irrelevant results for exactly the same reason. Timeliness and currentness are related but separate problems, and treating them as one is where most pipelines fail.&lt;/p&gt;
&lt;p&gt;Timeliness concerns whether data arrives quickly enough for its intended use. Currentness concerns whether the values still reflect present conditions. A pipeline can be timely, meaning data arrives on schedule, while the data itself is stale because the underlying conditions it describes changed months ago. Both require separate measurement with separate controls.&lt;/p&gt;
&lt;p&gt;Define freshness in business terms before setting any technical threshold. &amp;ldquo;Updated daily&amp;rdquo; means nothing without context. For fraud detection, daily updates mean your model is twelve to twenty-four hours behind attacker behavior at all times. That is an acceptable lag for some fraud patterns and a catastrophic lag for others. For quarterly financial reporting, daily updates are likely far more than required. The business use case determines the freshness requirement, not the pipeline&amp;rsquo;s default cadence.&lt;/p&gt;
&lt;p&gt;Use low-latency data pipelines for time-sensitive AI applications. Change data capture delivers timely data from relational database systems by propagating incremental changes rather than full refreshes. Stream capture handles data originating from IoT devices and other high-velocity sources that require low-latency processing. Once captured, downstream analytical and operational stores should be updated continuously rather than in scheduled batches that create artificial staleness windows.&lt;/p&gt;
&lt;p&gt;Store both event time and processing time for every record. Confusing them hides late-arriving data. If your pipeline processes a transaction at 14:00 that occurred at 08:00, the six-hour gap is only visible if you captured both timestamps independently. Relying on a single timestamp means you cannot detect late arrival, replay, or out-of-order delivery.&lt;/p&gt;
&lt;p&gt;Test late-arriving, duplicated, out-of-order, and replayed events as part of your standard pipeline validation. These are the failure modes that stress-test assumptions baked into most feature engineering pipelines. A pipeline that handles clean, on-time data correctly will often fail in ways that corrupt model inputs when events arrive late or out of sequence.&lt;/p&gt;
&lt;p&gt;Measure distribution drift on a cadence separate from pipeline latency checks. Use Population Stability Index, Jensen-Shannon divergence, or Wasserstein distance to compare current production data against your training baseline. Set retraining or review triggers when drift persists across multiple measurement periods, not just when a single batch looks unusual. A single anomalous batch is often noise. Sustained drift means your training distribution and production distribution have separated and your model is operating outside the conditions it learned from.&lt;/p&gt;
&lt;p&gt;Teams often set a single drift alert threshold and then mute it when it fires continuously. That continuous firing is the signal, not the noise. Establish escalation procedures for sustained drift that include a defined review timeline and a retraining decision process, not just a repeated alert that gets ignored.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="4-consistency"&gt;4. Consistency&lt;/h3&gt;
&lt;p&gt;Consistency means the same concepts are represented uniformly across sources, versions, time periods, and processing steps. Inconsistency is invisible until it damages your model, and by that point the damage is already embedded in learned weights that are difficult to inspect and harder to correct.&lt;/p&gt;
&lt;p&gt;You will not see a consistency failure in a simple data profile. You see it when your model learns that &amp;ldquo;1&amp;rdquo; means true in one data source and &amp;ldquo;1&amp;rdquo; means a product category code in another. You see it when temperature features from two sensors suddenly shift because one system reported in Celsius and the other in Fahrenheit, and nobody documented the difference. You see it when referential integrity fails silently and foreign keys resolve to deleted parent records.&lt;/p&gt;
&lt;p&gt;Maintain a canonical data dictionary and controlled vocabulary. Version both the dictionary and your schemas. Treat schema changes as deployment events that require review and approval, with the same rigor you apply to code changes. Silent schema changes, the kind where an upstream system silently renames a column or changes a data type, should trigger alerts and halt downstream processing until explicitly approved.&lt;/p&gt;
&lt;p&gt;Run contract tests between data producers and consumers. If an upstream system silently changes a column type from integer to string, your pipeline should fail loudly, not silently cast values and continue. Contract tests define the expected shape, types, ranges, and semantics of the data at each interface. When either side of the contract changes, the test fails and humans are notified.&lt;/p&gt;
&lt;p&gt;Check units across every numeric field, not just ranges. Kilograms versus pounds. Celsius versus Fahrenheit. Milliseconds versus seconds. Kilometers versus miles. These errors do not appear in range checks because the values are plausible within their own unit system. They appear as feature distributions that look normal but shift model behavior in ways that are extremely difficult to trace without unit metadata.&lt;/p&gt;
&lt;p&gt;Test cross-source agreement for shared keys and attributes. When your CRM and transaction system both carry a customer age field, check whether they agree. Systematic disagreement means you have a consistency failure and you need to decide which source is authoritative before any model uses that field. That decision cannot be left to the feature engineering pipeline to resolve implicitly.&lt;/p&gt;
&lt;p&gt;Extend consistency checks to multimodal data. Confirm that text metadata corresponds to the correct image, audio clip, or document. Misaligned multimodal pairs corrupt the cross-modal signal the model is trained to learn, and they are nearly impossible to detect through standard quality profiles because each modality passes its own checks independently.&lt;/p&gt;
&lt;p&gt;One consistency failure that rarely appears in quality checklists is timezone inconsistency in timestamps. Different source systems default to different timezones, or to UTC in some cases and local time in others. This creates apparent time-of-day patterns in your data that are artifacts of timezone handling, not real behavioral signals. A model trained on this data learns spurious time-of-day effects that disappear or reverse in production when the inference system uses a different timezone convention.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="5-representativeness-and-diversity"&gt;5. Representativeness and Diversity&lt;/h3&gt;
&lt;p&gt;Representativeness asks whether the data reflects the populations, environments, conditions, and failure modes where the system will actually operate. Diversity asks whether meaningful variation exists across patterns, perspectives, scenarios, sources, languages, contexts, and edge cases.&lt;/p&gt;
&lt;p&gt;Bias in AI systems occurs when applications produce results that reflect human biases, including social inequality. This happens when training data reflects a narrow band of attributes, perspectives, or populations. A credit risk model trained primarily on historical data from a specific geography or demographic group will not generalize fairly to the full population it serves. A hiring model trained on ten years of historical approvals learns to replicate the biases embedded in those approvals, not to identify the best candidates.&lt;/p&gt;
&lt;p&gt;Diverse data means drawing from a wide range of sources spanning different patterns, variations, and scenarios relevant to the problem domain. That data might be structured or unstructured, cloud-hosted or on-premises, originating from transaction systems, IoT devices, enterprise applications, software as a service platforms, mainframes, databases, files, or documents. Narrowing your data sources to what is most convenient to access is one of the most reliable ways to build a model that fails the people it was designed to serve.&lt;/p&gt;
&lt;p&gt;Do not assume that demographic balance alone proves fairness. Some operational datasets should not mirror population proportions. A model trained to detect rare equipment failures should oversample failure cases, not mirror a ninety-nine to one natural imbalance. A fraud detection model that mirrors the natural fraud rate will have almost no positive examples to learn from. The target distribution must be justified explicitly based on the learning objective, not assumed to be correct because it matches census statistics.&lt;/p&gt;
&lt;p&gt;NIST recommends disaggregating evaluations across demographic groups and intersecting subgroups. Aggregate accuracy is not a fairness metric. A model can achieve ninety-two percent accuracy overall while performing at sixty-eight percent for a specific subgroup. That gap will not appear in any headline number. It will appear in user complaints, regulatory reviews, and adverse outcomes.&lt;/p&gt;
&lt;p&gt;Test intersectional groups where sample sizes permit. Age crossed with gender crossed with region can reveal failure modes that are entirely invisible in single-dimension analysis. A model that performs equally well for women and equally well for younger users can still fail systematically for young women of a specific ethnicity. Intersectional testing requires adequate sample sizes in each cell, which is itself a representativeness requirement.&lt;/p&gt;
&lt;p&gt;Build a challenge set from known incidents, complaints, adversarial examples, and expert-defined edge cases. Your standard test set reflects what happened in historical data. Your challenge set reflects what can happen in the real world. Models that pass standard test sets while failing challenge sets are models that have learned to perform in controlled conditions and generalize poorly to the unexpected.&lt;/p&gt;
&lt;p&gt;Use stratified splits so rare classes and important subgroups appear in validation and test sets with sufficient volume to measure. A random split on an imbalanced dataset will often leave your minority class almost entirely in the training split, making it impossible to evaluate performance on exactly the cases that matter most.&lt;/p&gt;
&lt;p&gt;A practical enforcement rule is to require every high-impact subgroup to exceed a minimum sample size in the test set and to publish accuracy, false-positive rate, false-negative rate, and calibration separately for each subgroup before the model is approved for deployment. That requirement makes representativeness gaps visible before they cause harm.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="6-relevance-and-fitness-for-purpose"&gt;6. Relevance and Fitness for Purpose&lt;/h3&gt;
&lt;p&gt;Relevance measures whether the data supports the intended task, user group, operating environment, and decision horizon. Irrelevant data adds noise. Future data used as model features creates leakage. Data collected under different conditions than deployment produces a model that works in the lab and fails in the field.&lt;/p&gt;
&lt;p&gt;Write an intended-use statement before selecting any data. Define prohibited uses and known out-of-scope populations explicitly. This is not a formality. It forces decisions about what the model is for and, critically, what it is not for. Those decisions constrain which data sources are legitimate inputs. Without an intended-use statement, data selection decisions default to whatever is available and convenient, which is almost never the right answer.&lt;/p&gt;
&lt;p&gt;Feature usefulness should be tested empirically, not assumed. Remove or mask a feature and measure whether model performance changes meaningfully on your validation set. A feature that survives permutation importance testing and ablation testing is contributing real signal. A feature that does not survive those tests may be noise, a proxy for a protected attribute, or a leakage vector that inflates offline performance while failing in production.&lt;/p&gt;
&lt;p&gt;Temporal leakage is the most dangerous relevance failure because it produces results that look correct by every offline metric and fail completely in deployment. A feature derived from information available after the prediction timestamp will produce outstanding offline accuracy and catastrophic production failure. A maintenance record created at the time of a failure event, when used to predict that failure, tells the model something it cannot possibly know before the event occurs.&lt;/p&gt;
&lt;p&gt;Random splits conceal temporal leakage entirely. Use time-based splits where deployment involves future cases. Use group-based splits where deployment involves new users, new sites, new organizations, or new devices. If every user in your test set also appears in your training set, your offline evaluation measures how well the model memorizes individual patterns, not how well it generalizes to users it has never encountered.&lt;/p&gt;
&lt;p&gt;Include a &amp;ldquo;not relevant&amp;rdquo; rejection category in human labeling instructions. This forces annotators to flag content that does not belong in the dataset at all, rather than forcing an assignment to the nearest available class. Without this category, annotators assign ambiguous or irrelevant examples to whatever class seems closest, introducing noise that the model learns as signal.&lt;/p&gt;
&lt;p&gt;The leakage failures that hurt most are subtle, not obvious. Nobody accidentally includes the outcome label as a raw feature. The dangerous cases are a field updated at the time of outcome recording, a derived feature that aggregates events occurring after the prediction timestamp, or an identifier that correlates with outcome because of how data was collected rather than because of any real relationship. Build a leakage review into your feature documentation process as a required step, not an optional check.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="7-validity"&gt;7. Validity&lt;/h3&gt;
&lt;p&gt;Validity means data conforms to specified syntax, types, ranges, codes, formats, and business constraints. Invalid data corrupts feature engineering, breaks preprocessing pipelines, and introduces errors that propagate through every downstream transformation.&lt;/p&gt;
&lt;p&gt;Run validation before data enters any training, feedback, or feature store. Every batch. Without exception. Validation that runs only on initial data load misses every defect introduced by schema changes, pipeline updates, and source system modifications that occur after the initial check.&lt;/p&gt;
&lt;p&gt;Quarantine failed records rather than dropping them silently. Silent dropping hides failure rates from everyone downstream, including the model owners who need to know whether the training set shrank, the business owners who need to know whether records are being lost, and the governance team that needs to know whether a systematic problem exists upstream. When your pipeline drops five percent of records from a specific source without logging or alerting, that five percent is invisible to everyone who needs to act on it.&lt;/p&gt;
&lt;p&gt;Version your validation rules and retain failure reports. When a model&amp;rsquo;s performance degrades, you need to be able to answer whether the validation rules changed, not just whether the source data changed. A validation rule that became more permissive because a developer found it inconvenient is a governance failure that should appear in the audit trail.&lt;/p&gt;
&lt;p&gt;Distinguish between invalid data and valid but unusual data. An outlier is not automatically an error. A transaction amount in the ninety-ninth percentile may be entirely legitimate. An age of two hundred and forty is not. Your validation rules need to encode that distinction explicitly, with separate handling for impossible values and improbable but possible values.&lt;/p&gt;
&lt;p&gt;Test adversarially malformed inputs and encoding problems. Validation rules are almost always written against clean, well-formed examples. Real data pipelines receive corrupted files, misencoded characters, truncated records, malformed JSON, and inputs that exploit edge cases in parsing libraries. Your validation layer needs to handle those cases explicitly and fail safely, not pass them to the model to fail on in ways that produce silent incorrect outputs.&lt;/p&gt;
&lt;p&gt;The most expensive validity failure is one that produces values that pass individual type and range checks but violate business constraints. A date that is technically valid but falls before the product existed. A transaction amount that is within the allowed range but combined with a currency code that makes it implausible. A combination of age and account creation date that is jointly impossible. These require cross-field and cross-record constraint rules that most teams never write. Building cross-field validations into your quality gate as a required step, not an optional enhancement, catches an entire class of failures that field-level validation misses entirely.&lt;/p&gt;
&lt;p&gt;A strong validation pipeline reports both the percentage of records that passed and the exact rejection reasons for every batch, segmented by source, time period, and field. Percentage alone tells you the scale of the problem. Rejection reasons tell you the pattern. Patterns reveal systematic problems upstream that need to be fixed at the source, not repeatedly quarantined at the gate.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="8-uniqueness-and-deduplication"&gt;8. Uniqueness and Deduplication&lt;/h3&gt;
&lt;p&gt;Uniqueness ensures that records represent distinct entities or events when duplicates would distort learning or evaluation. Duplicate records inflate the influence of certain examples, bias learned representations toward overrepresented patterns, and, in the worst cases, contaminate evaluation with examples the model has already memorized.&lt;/p&gt;
&lt;p&gt;The deduplication sequencing error is where most teams go wrong. Deduplicate before splitting data, not afterward. If you split first and then deduplicate within splits, you can remove duplicates within each partition while leaving near-identical records across the training and test boundary. That produces train-test contamination that inflates every metric without improving actual generalization.&lt;/p&gt;
&lt;p&gt;Use group-aware splits for records belonging to the same customer, patient, device, household, author, or organization. If all records from a given customer land in both training and test sets, your evaluation measures how well the model memorizes customer-specific patterns. When that customer calls a month after deployment with a problem the model should have caught, you discover that your ninety-three percent test accuracy meant nothing because the model never actually generalized.&lt;/p&gt;
&lt;p&gt;Train-test contamination is the most damaging uniqueness failure for language and image models. Near-duplicate examples in the test set produce artificially high evaluation scores even when the model has genuinely poor generalization. This is not a theoretical risk. It has produced published benchmark results that failed completely when the models were applied to real tasks. The contamination is invisible until someone runs explicit overlap detection between splits.&lt;/p&gt;
&lt;p&gt;Exact duplicate detection requires hashing normalized records after stripping whitespace, punctuation, and case variations. Near-duplicate detection requires similarity matching techniques such as MinHash, locality-sensitive hashing, or embedding similarity. Both are necessary. Exact deduplication misses records that differ only by formatting, encoding, or minor variations that carry identical semantic content.&lt;/p&gt;
&lt;p&gt;Investigate whether data augmentation has introduced near-duplicates into your test set. Augmented training examples that are semantically identical to test examples compromise evaluation integrity in exactly the same way as contamination from an external source. The fact that you created the near-duplicates deliberately through augmentation does not make the contamination less real.&lt;/p&gt;
&lt;p&gt;The question of what constitutes uniqueness in your specific dataset is a business decision, not a technical one. A customer with two accounts is one entity for churn modeling and two entities for fraud detection. A document that appears in multiple collections is one document for deduplication and multiple entries for citation analysis. That decision must be made explicitly, documented in your quality scorecard, and retained with the matching thresholds used. Duplicate-label conflicts, where the same record received conflicting labels from different annotators or across different dataset versions, introduce contradictory training signal. Check for duplicate keys with conflicting labels before training begins and resolve them through adjudication.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="9-security-privacy-accessibility-and-compliance"&gt;9. Security, Privacy, Accessibility, and Compliance&lt;/h3&gt;
&lt;p&gt;AI systems frequently operate on sensitive data. Personal identifiers, financial records, health information, biometric data, proprietary business content. Leaving that data unsecured creates two distinct problems that most teams conflate.&lt;/p&gt;
&lt;p&gt;The first problem is privacy. Exposed data compromises the people the model was built to serve and creates legal liability under GDPR, CCPA, HIPAA, and sector-specific regulations. The second problem is model integrity. Manipulated or exposed training data biases outputs in ways that are difficult to detect and potentially permanent. An attacker who can inject records into a training dataset can shift model behavior without ever touching the model itself.&lt;/p&gt;
&lt;p&gt;Three controls work together. Data classification automatically detects, categorizes, and tags data by sensitivity level, including sensitive, confidential, and restricted designations. Data protection applies the appropriate controls: masking, tokenization, or encryption for fields that require obfuscation, and access restriction for entire datasets. Access control defines who can access which data under which conditions, enforced through role-based permissions with least privilege as the default.&lt;/p&gt;
&lt;p&gt;Apply purpose limitation consistently. Authorized access does not automatically mean authorized use. A data scientist with read access to a training dataset is not automatically authorized to export that dataset to a personal environment, use it for a different model, or share it with a vendor. Purpose limitation requires that every data access decision answers not just &amp;ldquo;can this person access this data&amp;rdquo; but &amp;ldquo;is this access consistent with the documented use of this data.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Separate raw, de-identified, feature, feedback, and production data stores. A single access control policy applied across all five layers cannot adequately protect any of them. The risk profile of a raw dataset containing identifiable health records is fundamentally different from the risk profile of an aggregated feature store derived from those records. Separation reduces blast radius and enables more precise access auditing.&lt;/p&gt;
&lt;p&gt;Scan feedback data before reuse in any capacity. Feedback pipelines are a significant and underappreciated attack surface. User inputs can contain prompt injection attempts, personally identifiable information, credentials, malicious content, and adversarial examples designed to corrupt retraining. None of that should enter a retraining pipeline without explicit screening, classification, and authorization.&lt;/p&gt;
&lt;p&gt;Test re-identification risk on datasets you intend to share, publish, or move between environments. Linkage attacks and singling-out techniques can recover individual identities from datasets that passed standard anonymization checks. Run those tests before sharing any de-identified dataset. The fact that you removed direct identifiers does not mean the dataset is anonymous.&lt;/p&gt;
&lt;p&gt;Log access, export, transformation, and deletion events comprehensively. These logs serve as both a security control and a lineage artifact. They enable you to reconstruct who accessed what data, when, from where, and for what declared purpose. They are also the first artifact a regulator or auditor will request following an incident.&lt;/p&gt;
&lt;p&gt;The most underestimated security control in AI data pipelines is encryption of temporary files and intermediate outputs. Teams apply strong encryption to source data and final model artifacts and then leave temporary files, cache directories, intermediate training checkpoints, and experiment logs completely unencrypted. Those files often contain sensitive training records in raw or partially processed form. They must be included in your encryption requirements and in your deletion procedures when their purpose is complete.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id="10-traceability-and-lineage"&gt;10. Traceability and Lineage&lt;/h3&gt;
&lt;p&gt;Traceability means you can reconstruct where data came from, how it changed at every step, which model used it, and which outputs or decisions it influenced. Lineage is the artifact that makes that reconstruction possible. Together they enable accountability, reproducibility, and compliance. Without them, you cannot demonstrate that a model was built responsibly, and you cannot investigate a failure after it occurs.&lt;/p&gt;
&lt;p&gt;Assign immutable identifiers to every dataset version and snapshot. Record cryptographic hashes for files, partitions, labels, and model inputs where practical. This is what makes reproducibility real rather than theoretical. Without it, you cannot confirm that the model you are investigating used the data you think it used. &amp;ldquo;We trained on last quarter&amp;rsquo;s data&amp;rdquo; is not traceability. A versioned dataset identifier with a hash is traceability.&lt;/p&gt;
&lt;p&gt;Maintain a lineage graph from source through every transformation to the model and its outputs. When a production prediction is disputed, you need to trace it back to the specific training record, the labeling guideline version, the annotator, and the source system. That trace is only possible if you captured it at every step in the pipeline. Lineage that covers the first and last mile but misses the transformations in between is documentation theater.&lt;/p&gt;
&lt;p&gt;The lineage gap that creates the most governance risk is between the feature store and source data. Teams often track lineage through ingestion and initial transformation pipelines and then lose it at the feature store boundary. The model trains on features, not raw data. If you cannot trace a feature value back to the source record that produced it, your lineage is incomplete in exactly the place where failures most often originate.&lt;/p&gt;
&lt;p&gt;Link every model artifact to exact training, validation, and test dataset versions, feature engineering code, configuration files, and labeling guideline version. Link user feedback to the model version that generated the response, the specific input, the reviewer decision, and the final disposition. Feedback that cannot be traced to a model version cannot be used to evaluate that model or to construct valid retraining data without introducing unknown confounders.&lt;/p&gt;
&lt;p&gt;Preserve rejected data and quality exception reports when legally and operationally appropriate. The records you excluded are as important as the records you included. They document the boundaries of your dataset and enable you to assess, months or years later, whether those boundaries were appropriate and whether they introduced gaps that affected model behavior.&lt;/p&gt;
&lt;p&gt;Build a business glossary that maps business terms to technical items in your datasets. Add semantic typing to provide additional meaning for automated systems. Index all metadata in a searchable catalog. Required catalog fields include source, owner, collection time, transformation history, version, intended use, retention period, sensitivity tags, and known bias issues. A dataset that cannot be found or understood is functionally unavailable, regardless of how complete or accurate it is.&lt;/p&gt;
&lt;p&gt;Test your lineage by asking an independent reviewer to trace a specific production prediction back to its source records without your help. If they cannot complete that trace in a timeframe appropriate to your risk level, your lineage system is not operational. Set a specific target, such as completing any audit trace within four hours for high-risk models. Measure against it. A lineage diagram that looks complete on a whiteboard but takes three weeks to navigate in practice protects nobody.&lt;/p&gt;
&lt;p&gt;NIST emphasizes evaluating data and content provenance, including original sources, transformations, and decision-making criteria. ISO-oriented guidance requires documentation of provenance, update dates, training and validation categories, labeling processes, intended use, quality requirements, retention policies, and known bias issues for every dataset used in an AI system.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/08/chatgpt-image-aug-14-2026-05_30_55-pm.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="how-to-set-quality-gates-that-actually-block-bad-data"&gt;How to Set Quality Gates That Actually Block Bad Data&lt;/h2&gt;
&lt;p&gt;Quality gates block defective data from entering training or feedback stores. The design of those gates determines whether they protect your model or simply generate paperwork.&lt;/p&gt;
&lt;p&gt;Use hard gates for safety, privacy, schema, leakage, and lineage. Hard gates are binary. The data either passes or it does not. A dataset that fails a hard gate does not enter the training pipeline regardless of its performance on other dimensions. Use monitoring thresholds with defined escalation procedures for drift, freshness, subgroup balance, and performance degradation. Monitoring thresholds trigger human review and a defined response timeline. They do not automatically halt processing.&lt;/p&gt;
&lt;p&gt;Never let a high composite quality score mask a single unacceptable hard-gate failure. A composite score of 88 percent looks strong. A security dimension score of 22 percent within that composite means the model has access to unmasked personal data. The composite hides the single failure that matters most. Hard gates must be reported separately and evaluated independently. They cannot be averaged into an aggregate score.&lt;/p&gt;
&lt;p&gt;Do not copy thresholds from other systems. ISO/IEC 5259-2 treats data quality measures as context-dependent, and ISO/IEC 42001 requires requirements appropriate to the system&amp;rsquo;s intended use. A ninety-five percent label accuracy rate is a catastrophic failure for a clinical AI system. It may be entirely acceptable for an internal content recommendation system. Define your thresholds before training begins, document the justification for each one, and revisit them at every major model version.&lt;/p&gt;
&lt;h2 id="requirements-that-apply-at-every-stage"&gt;Requirements That Apply at Every Stage&lt;/h2&gt;
&lt;p&gt;Several requirements cut across all ten dimensions and all four data stages. These are not additional checks. They are the operating conditions that make the ten requirements enforceable over time.&lt;/p&gt;
&lt;p&gt;Version everything that touches the data. Schemas. Validation rules. Transformation code. Labeling guidelines. Annotation team composition. Sampling plans. Quality thresholds. Build your versioning cadence around deployment events, not calendar dates. Every time a model is deployed, create a snapshot of every artifact that contributed to it. That snapshot is your reproducibility baseline. When something fails in production three months later, you are not guessing what the training environment looked like. You have a record.&lt;/p&gt;
&lt;p&gt;Separate data preparation from data approval. The person who prepares the data should not be the only person who certifies it. Data stewards approve remediation strategies. Model owners cannot self-approve test data independence. Separation of duties is a control, not a bureaucratic inconvenience. When the same person who selected the data also signs off on its quality, the approval is not independent and the governance is not real.&lt;/p&gt;
&lt;p&gt;Document every rejected decision, not just the accepted ones. Preserve rejected data and quality exception reports when legally and operationally appropriate. When a model failure occurs, those rejected records often contain the earliest visible signal. They also provide evidence for regulators that you discovered, quarantined, and escalated the issue through a defined process rather than ignoring it.&lt;/p&gt;
&lt;p&gt;Require explicit written approval for every new data source added to a retraining pipeline. Teams add new sources without formal review because each addition feels incremental. Collectively, those additions shift the training distribution, introduce new privacy considerations, and alter model behavior in ways that were never assessed. The retraining approval gate is where governance fails most quietly, and it is where a formal requirement has the most leverage.&lt;/p&gt;
&lt;p&gt;Make data consumable as a functional requirement, not an afterthought. Traditional machine learning workflows favor well-formed tabular structures and feature stores where SQL is a first-class language. Generative AI workflows require text from unstructured sources to be split into manageable chunks, converted into embeddings, and stored in a vector index. Each chunk must carry lineage back to the source document. Without that lineage, retrieval results cannot be audited and retrieved content cannot be verified. Consumability requirements must be specified alongside quality requirements, not treated as a separate concern belonging only to the engineering team.&lt;/p&gt;
&lt;h2 id="controls-metrics-and-validation-guidance"&gt;Controls, Metrics, and Validation Guidance&lt;/h2&gt;
&lt;p&gt;Use this table during data quality gate reviews. The metrics and tests are starting points calibrated to common practice. Adjust thresholds to match your specific use case, risk profile, and regulatory context. Do not treat any threshold here as universal.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;AI Data Requirement&lt;/th&gt;
&lt;th&gt;Key Metrics and Validation Tests&lt;/th&gt;
&lt;th&gt;Recommended Controls&lt;/th&gt;
&lt;th&gt;Guidance and Acceptance Criteria&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy&lt;/td&gt;
&lt;td&gt;Value accuracy rate: correct values divided by inspected values. Label accuracy and label error rate. Numeric error: mean absolute error, root mean squared error, percentage error. Entity resolution precision and recall. Annotation confidence and adjudication rate. Compare samples against authoritative systems, original documents, instruments, or expert-reviewed ground truth. Double-label a statistically justified sample and calculate Cohen&amp;rsquo;s kappa or Fleiss kappa. Reconcile records with source-of-record systems and investigate discrepancies above a defined tolerance. Require independent review for low-confidence or disputed labels.&lt;/td&gt;
&lt;td&gt;Separate measurement accuracy from label accuracy. Define tolerances by use case before profiling. Record who produced or reviewed labels, when they were created, and which labeling instructions were used. Use a holdout audit sample that annotators cannot see during data preparation. Profile source data with exploratory data analysis to understand characteristics, completeness, distribution, redundancy, and shape. Operationalize remediation strategies with data quality rules and monitor them continuously. Enable lineage and impact analysis to trace origins and prevent accidental modification.&lt;/td&gt;
&lt;td&gt;For a classifier, one acceptance rule is at least 98 percent of critical labels must agree with adjudicated expert labels, with no high-severity label error remaining unresolved. A small rounding error may be acceptable in forecasting but unacceptable in safety-critical control. The audit sample is not optional. Build it into the data preparation budget from day one.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Completeness&lt;/td&gt;
&lt;td&gt;Field completeness: 1 minus missing values divided by expected values. Record completeness: received records divided by expected records. Label coverage: labeled records divided by records intended for supervised learning. Time period coverage. Class coverage and minority group coverage. Missingness rate by subgroup, source, and time period. Profile every column by dataset version and compare with thresholds. Reconcile dataset counts with source-system counts, event logs, or control totals. Reject or quarantine unlabeled records unless an explicit missing-label policy exists. Check for missing dates, gaps in event sequences, and unexplained inactivity.&lt;/td&gt;
&lt;td&gt;Set stricter thresholds for critical fields than optional fields. Do not treat imputation as elimination of the problem. Retain indicators showing which values were imputed. Measure completeness separately for training, validation, test, feedback, and production data. Investigate complete records that contain default values such as 0, unknown, or 1970-01-01.&lt;/td&gt;
&lt;td&gt;A typical data quality gate requires zero missing values for mandatory identifiers and labels, while allowing a documented, bounded rate of missingness in noncritical features. Concentrated missingness is more dangerous than distributed missingness. Five percent missing overall can hide fifty percent missing in a specific demographic group, geography, or outcome class. Always break completeness metrics down by subgroup before accepting any overall figure.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timeliness and currentness&lt;/td&gt;
&lt;td&gt;Data age: current time minus event time or last update time. Pipeline latency: ingestion time minus source event time. Freshness compliance rate. Update frequency and interval variance. Staleness rate. Distribution drift: Population Stability Index, Jensen-Shannon divergence, Wasserstein distance, population mean or variance change. Concept drift indicators comparing delayed outcomes, error rates, and label distributions over time. Enforce maximum allowed age for each source and use case. Run end-to-end latency tests from source generation to model availability. Alert when records or batches exceed service-level objectives. Check whether feeds arrive according to documented schedules. Identify records unchanged beyond a defined period.&lt;/td&gt;
&lt;td&gt;Define freshness in business terms, not technical terms. Store both event time and processing time separately. Test late, duplicated, out-of-order, and replayed events. Establish retraining or review triggers when drift persists rather than reacting to a single unusual batch. Use change data capture for relational sources and stream capture for low-latency sources such as IoT devices. Update downstream stores continuously.&lt;/td&gt;
&lt;td&gt;Document the last update, intended use, data category, and quality requirements for each dataset. Updated daily may be adequate for demand planning but not fraud detection. Sustained drift is the signal, not noise. Establish escalation procedures for persistent drift, not just individual alerts.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consistency&lt;/td&gt;
&lt;td&gt;Cross-source agreement rate. Schema conformity rate. Unit consistency rate. Referential integrity rate. Business rule violation rate. Cross-version reproducibility. Contradiction rate. Compare shared keys and attributes across systems of record. Validate column names, types, units, encoding, and allowed nullability. Detect incompatible units such as kilograms versus pounds or Celsius versus Fahrenheit. Confirm foreign keys resolve to valid parent records. Test constraints such as shipment date greater than or equal to order date. Re-run transformations and compare hashes, counts, and summary statistics. Find records where two fields or sources assert incompatible facts.&lt;/td&gt;
&lt;td&gt;Maintain a canonical data dictionary and controlled vocabulary. Version schemas and transformation code. Use contract tests between producers and consumers. Treat silent schema changes as deployment failures, not merely warnings. Test consistency within multimodal data such as text metadata corresponding to the correct image or audio file.&lt;/td&gt;
&lt;td&gt;Example rules: every transaction must reference an existing account. Currency must be explicit. Timestamps must contain a timezone. Monetary values must use the declared currency and precision. The most common hidden consistency failure is timezone inconsistency across sources. Validate timezone handling explicitly in every pipeline.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Representativeness and diversity&lt;/td&gt;
&lt;td&gt;Population coverage by group, region, language, device, environment, and use case. Distribution distance measures: standardized mean difference, KL divergence, Population Stability Index, Wasserstein distance. Class balance metrics and minority class share. Group-specific label rates and outcome rates. Coverage of known failure modes and edge cases. Fairness metrics: demographic parity difference, disparate impact, equal opportunity difference, equalized odds difference, group-specific error rates. Compare dataset proportions with the target deployment population or a justified sampling frame. Compare training, validation, test, and production distributions. Test whether rare but important classes are sufficiently represented. Investigate unexplained differences before training. Build a challenge set from incidents, complaints, expert scenarios, and adversarial examples. Evaluate model outcomes separately by group, subgroup, and intersection.&lt;/td&gt;
&lt;td&gt;Do not assume demographic balance alone proves fairness. Document why the target distribution is appropriate. Some operational datasets should not mirror population proportions. Test intersectional groups such as age crossed with gender crossed with region where sample sizes permit. Include domain experts and affected communities when selecting fairness criteria. Use stratified splits so important groups and rare events appear in validation and test sets. Draw from structured and unstructured sources across cloud, on-premises, operational databases, enterprise resource planning systems, software as a service applications, files, and documents.&lt;/td&gt;
&lt;td&gt;A practical test is to require every high-impact subgroup to exceed a minimum sample size and to publish accuracy, false-positive rate, false-negative rate, and calibration separately for each subgroup. Diverse data means drawing from a wide range of sources spanning different patterns, perspectives, variations, and scenarios. Narrowing data sources to what is convenient reliably builds biased models.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relevance and fitness for purpose&lt;/td&gt;
&lt;td&gt;Feature usefulness: mutual information, permutation importance, or task-specific ablation impact. Label feature time alignment. Coverage of intended use cases. Out-of-domain rate using applicability domain or embedding distance checks. Leakage rate searching for post-outcome fields, future timestamps, duplicated labels, or target-derived variables. Signal-to-noise indicators measuring unusable, irrelevant, corrupted, or unrelated content. Remove or mask a feature and determine whether it contributes meaningful validated performance. Test that features were available before the prediction point. Map each record to a documented use case, workflow, or scenario.&lt;/td&gt;
&lt;td&gt;Write an intended use statement before selecting data. Define prohibited uses and known out-of-scope populations. Test temporal leakage rigorously because random splits can conceal it. Use time-based or group-based splits where deployment involves future cases, users, sites, or organizations. Include a not relevant rejection category in human labeling.&lt;/td&gt;
&lt;td&gt;A model intended to predict next-day equipment failure should not use maintenance records created after the prediction timestamp, even if those records improve offline accuracy. The dangerous leakage cases are subtle. Build a leakage review into the feature documentation process.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Validity&lt;/td&gt;
&lt;td&gt;Schema validation pass rate. Type conformance rate. Range conformance rate. Pattern conformance rate. Vocabulary conformance rate. Constraint violation rate. Parsing or tokenization failure rate. Validate every batch against a versioned schema. Check dates, numerics, booleans, categorical codes, encodings, and nested structures. Reject impossible values such as negative age or humidity above physical limits. Validate identifiers, email formats, country codes, and timestamp formats. Check categorical values against approved code lists. Apply domain rules and cross-field validations. Test whether documents, images, audio, and structured records can be consumed correctly.&lt;/td&gt;
&lt;td&gt;Run validation before data enters training or feedback stores. Quarantine failed records rather than silently dropping them. Version validation rules and retain failure reports. Distinguish invalid data from valid but unusual data. Outliers are not automatically errors. Test adversarially malformed inputs and encoding problems.&lt;/td&gt;
&lt;td&gt;A strong pipeline reports both the percentage that passed and the exact rejected record reasons, enabling remediation and audit. The most expensive validity failures pass type and range checks but violate business constraints. Build cross-field and cross-record constraint rules into the quality gate as a first-class requirement.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Uniqueness and deduplication&lt;/td&gt;
&lt;td&gt;Exact duplicate rate. Near-duplicate rate using similarity matching, MinHash, locality-sensitive hashing, or embedding similarity. Duplicate entity rate using deterministic and probabilistic rules on names, addresses, identifiers, images, or documents. Event uniqueness rate checking event IDs, timestamps, sequence numbers, and source offsets. Train test overlap rate searching identical or near-identical examples across splits. Duplicate label conflict rate identifying the same item receiving conflicting labels. Hash normalized records and count repeated hashes. Use similarity matching and entity resolution tests.&lt;/td&gt;
&lt;td&gt;Deduplicate before splitting data, not afterward. Use group-aware splits for records belonging to the same customer, patient, device, household, author, or organization. Avoid deleting legitimate repeated events before determining what constitutes uniqueness. Investigate data augmentation that creates near-duplicates in the test set. Retain duplicate decisions and matching thresholds.&lt;/td&gt;
&lt;td&gt;For language or image models, train test contamination can produce deceptively high evaluation scores even when the model has poor generalization. Deduplicate before splitting. Check for duplicate keys with conflicting labels before training begins and resolve them through adjudication, not arbitrary selection.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security, privacy, accessibility, and compliance&lt;/td&gt;
&lt;td&gt;Unauthorized access incidents and access denial rate. Encryption coverage at rest and in transit. Sensitive field discovery and masking coverage. Re-identification risk using linkage and singling-out tests. Consent and legal basis coverage. Data retention compliance rate. Dataset availability and recovery time. Data access latency and availability. Test role-based access control with least privilege and negative authorization cases. Verify encryption configuration, certificates, key rotation, backups, and temporary files. Scan for personal, financial, health, confidential, or credential data and verify masking or tokenization. Reconcile records with consent, purpose, retention, geographic, and contractual restrictions. Test automatic deletion, archival, and legal hold rules. Perform restore tests and measure recovery point and recovery time objectives. Verify that authorized training and inference jobs can reliably obtain required data.&lt;/td&gt;
&lt;td&gt;Apply purpose limitation. Authorized access does not automatically mean authorized use. Separate raw, de-identified, feature, feedback, and production stores. Scan feedback data for prompt injection, secrets, personal data, and malicious content before reuse. Log access, export, transformation, and deletion events. Test whether sensitive attributes can be inferred from supposedly anonymized records. Classify data by sensitivity tier such as sensitive, confidential, or restricted. Apply protection policies such as masking, tokenization, or encryption. Use access control policies based on least privilege.&lt;/td&gt;
&lt;td&gt;ISO/IEC 42001 is an organizational management system standard for developing, providing, and using AI systems, including managing data quality and system performance. The underestimated security control is encryption of temporary files, cache directories, and intermediate training checkpoints. Those files can contain sensitive training records. Include them in encryption and deletion procedures.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Traceability and lineage&lt;/td&gt;
&lt;td&gt;Lineage coverage: records or datasets with complete provenance divided by total. Transformation reproducibility rate. Metadata completeness rate for catalog fields, data dictionary entries, licensing, retention, and sensitivity tags. Label provenance coverage confirming annotator, guideline version, timestamp, confidence, and adjudication status. Version linkage coverage connecting each model artifact to exact training, validation, test, feature, code, and configuration versions. Audit query success rate and time to reconstruct. Change detection latency. Require source, owner, collection time, transformation, version, and intended use metadata. Rebuild a dataset from source snapshots and compare checksums and quality statistics. Ask an independent reviewer to trace a production prediction back to source records.&lt;/td&gt;
&lt;td&gt;Assign immutable dataset and snapshot identifiers. Record hashes for files, partitions, labels, and model inputs where practical. Maintain a data lineage graph from source through transformations to model and output. Link user feedback to the model version, prompt or input, response, reviewer decision, and final disposition. Preserve rejected data and quality exceptions when legally and operationally appropriate. Build a business glossary that maps business terms to technical items. Use semantic typing to provide extra meaning for automated systems. Index metadata in a searchable catalog.&lt;/td&gt;
&lt;td&gt;NIST emphasizes evaluating data and content provenance, including original sources, transformations, and decision-making criteria. ISO-oriented guidance highlights provenance, update date, training and validation and test and production categories, labeling processes, intended use, quality, retention, and known bias issues. The lineage failure that creates the most governance risk is the gap between feature store and source data. Extend the lineage graph through the feature engineering layer.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consumability for machine learning and generative AI&lt;/td&gt;
&lt;td&gt;For traditional ML: well-formed, high-quality, tabular data structures with SQL as a first-class language for data scientists. For generative AI: unstructured sources such as presentations, mail archives, text documents, PDFs, and transcripts split into manageable chunks, converted into embeddings, and stored in a vector database for similarity search. Each chunk must carry lineage back to the source file.&lt;/td&gt;
&lt;td&gt;For ML: use lakehouse-based feature stores and database-like structures. For GenAI: enforce quality controls upstream before ingestion. Trusted, secure, governed data becomes input to embedding pipelines.&lt;/td&gt;
&lt;td&gt;Retrieval augmented generation outputs cannot be audited without lineage from chunk back to source file. That is the only defensible order of operations.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Minimum validation focus by data stage&lt;/td&gt;
&lt;td&gt;Training data: accuracy, completeness, representativeness, leakage, duplication, label quality, privacy, and lineage. Validation and test data: independence from training data, representative coverage, stable labels, subgroup metrics, temporal validity, and contamination checks. Feedback data: authenticity, user authorization, toxicity and injection screening, label confidence, reviewer agreement, and linkage to the relevant model version. Production or usage data: freshness, schema validity, drift, out-of-domain inputs, access control, incident rates, and outcome-based accuracy once labels become available. Retraining data: provenance, change impact, regression tests, fairness comparison with the previous model, and approval of newly added sources.&lt;/td&gt;
&lt;td&gt;Use the minimum validation focus table as a starting point. Add tests that match the specific risk profile. For training data, always run leakage detection and group-aware split verification. For validation data, always check independence and subgroup coverage. For production data, always check schema drift and out-of-domain rates. For feedback data, always scan for injection and personal data.&lt;/td&gt;
&lt;td&gt;Pair data quality tests with model-level tests. A fresh, valid, complete dataset can still corrupt a retrained model if the source distribution shifted. Before retraining, run a regression suite against the incumbent model. Compare subgroup false-positive rates and calibration. Approve new sources only after they improve at least one intended outcome without degrading protected groups.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quality gate design and composite trust score&lt;/td&gt;
&lt;td&gt;Each dataset needs a versioned scorecard containing intended use, unacceptable uses, required dimensions, metric definitions, thresholds, tolerances, severity levels, test frequency, responsible owner, sampling method, confidence intervals, quarantine and escalation procedures, dataset and label and code and model versions, and retained audit evidence.&lt;/td&gt;
&lt;td&gt;Define hard gates for safety, privacy, schema, leakage, and lineage. Use monitoring thresholds for drift, freshness, subgroup balance, and performance degradation. Quarantine failed records instead of silently dropping them. Preserve rejected data when legally and operationally appropriate. Recalculate the composite score regularly. Treat any hard-gate failure as zero readiness even if the composite looks acceptable.&lt;/td&gt;
&lt;td&gt;Never approve average quality scores without per-subgroup evidence. A 99.7 percent schema pass rate can hide all records from one geography and one device type. A composite score of 82 percent can pass review while the security dimension sits at 22 percent. Hard gates must be reported separately and cannot be averaged away.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-cutting controls&lt;/td&gt;
&lt;td&gt;Version everything that touches the data: schemas, validation rules, transformation code, labeling guidelines, annotation team composition, sampling plans, and quality thresholds. Build versioning cadence around deployment events. Document decisions, not just outcomes. Separate data preparation from data approval. The person who prepares the data should not be the only person who certifies it. Data stewards approve remediation strategies. Model owners cannot self-approve test data independence. Protect feedback channels like production input. Scan feedback for prompt injection, secrets, personal data, and malicious content before reuse.&lt;/td&gt;
&lt;td&gt;Do not use universal thresholds. ISO/IEC 5259-2 treats data quality measures as context-dependent. ISO/IEC 42001 requires requirements appropriate to the intended use. A medical diagnosis system, a recommendation engine, and an internal search tool need different thresholds.&lt;/td&gt;
&lt;td&gt;A ninety-five percent label accuracy rate is catastrophic in clinical AI and acceptable for a recommendation engine. Define thresholds explicitly before training begins. Document the justification. Revisit them at each major model version.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="versioned-quality-scorecard-and-what-to-capture-for-every-dataset"&gt;Versioned Quality Scorecard and What to Capture for Every Dataset&lt;/h2&gt;
&lt;p&gt;For each dataset, maintain a versioned scorecard that includes the following fields.&lt;/p&gt;
&lt;p&gt;The intended use and explicitly excluded uses, written as specific statements rather than general descriptions. Required quality dimensions with metric definitions and the rationale for each inclusion. Thresholds and tolerances organized by severity level, with the justification for each threshold documented. Test frequency and responsible owner for each dimension. Sampling method, sample size, and confidence intervals for each metric. Quarantine, escalation, and remediation procedures with defined response timelines. Dataset version, label version, transformation code version, and model version references. Evidence retained for audit and reproducibility, including failure reports, adjudication records, and approval decisions.&lt;/p&gt;
&lt;p&gt;This scorecard is a living governance artifact. Update it at every dataset version. Review it at every model deployment gate. Audit it when something goes wrong. The scorecard that is never touched after initial completion is the strongest signal that data quality governance is not operational.&lt;/p&gt;
&lt;h2 id="recommended-additional-articles"&gt;Recommended Additional Articles&lt;/h2&gt;
&lt;p&gt;Read more about training, validation and usage data requirements including lifecycle-specific validation, measurable controls, hard deployment gates, monitoring thresholds, data leakage prevention, lineage, privacy, and common implementation failures. My following articles provide more related guidance on governing AI data, systems, vendors, and operational risk.&lt;/p&gt;
&lt;h3 id="practical-ai-assessments"&gt;Practical AI Assessments&lt;/h3&gt;
&lt;p&gt;This framework connects data readiness with formal AI lifecycle gates, requiring measured accuracy, completeness, representativeness, provenance, privacy compliance, testing, monitoring, and documented go/no-go decisions.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h3 id="practical-iso-42001-certification-guidance"&gt;Practical ISO 42001 Certification Guidance&lt;/h3&gt;
&lt;p&gt;Huwyler explains how to operationalize AI governance through data cards, provenance records, lifecycle-specific data controls, quality metrics, bias analysis, validation evidence, monitoring, and accountable approval processes.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h3 id="iso-42001-implementation-for-companies"&gt;ISO 42001 Implementation for Companies&lt;/h3&gt;
&lt;p&gt;This article shows why AI governance must separate training, validation, and testing data while continuously measuring provenance, quality, bias, production performance, and outputs outside approved operating conditions.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h3 id="ai-contract-clauses-and-data-controls"&gt;AI Contract Clauses and Data Controls&lt;/h3&gt;
&lt;p&gt;Huwyler translates data quality and governance principles into procurement requirements covering provenance, representativeness, labeling, bias, privacy, retention, supplier accountability, model updates, drift, audit rights, and exit controls.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h3 id="the-pren-18286-reality-check"&gt;The prEN 18286 Reality Check&lt;/h3&gt;
&lt;p&gt;This detailed regulatory analysis links AI quality management with data governance, dataset quality, traceability, verification, validation, lifecycle evidence, risk controls, post-market monitoring, and auditable conformity obligations.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h3 id="ai-isnt-coming-for-grc-jobs"&gt;AI Isn’t Coming for GRC Jobs&lt;/h3&gt;
&lt;p&gt;This executive perspective explains how weak data governance undermines AI-enabled risk management, compliance, audit analytics, monitoring, and control assurance across the organization.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h3 id="sap-s4hana-ai-analytics-and-continuous-monitoring"&gt;SAP S/4HANA AI, Analytics, and Continuous Monitoring&lt;/h3&gt;
&lt;p&gt;This article applies data-driven control testing to GRC, showing how organizations can identify risk patterns, design analytic tests, monitor evidence, and connect AI performance with governance controls.&lt;/p&gt;
&lt;p&gt;
&lt;/p&gt;
&lt;h2 id="external-references"&gt;External References&lt;/h2&gt;
&lt;p&gt;
provides a formal data quality model and measurable characteristics for analytics and machine learning, treating quality measures as context-dependent rather than universal.&lt;/p&gt;
&lt;p&gt;
addresses organizational processes for data quality in machine learning training and evaluation, defining process requirements for maintaining quality across the data lifecycle.&lt;/p&gt;
&lt;p&gt;ISO/IEC 42001:2023 is a management system standard for AI systems covering development, provision, and use. It requires organizations to define data quality requirements appropriate to the intended use of each AI system and verify them throughout the lifecycle.&lt;/p&gt;
&lt;p&gt;
provides foundational data quality dimensions for geographic information that have been adapted in broader AI data quality frameworks.&lt;/p&gt;
&lt;p&gt;NIST AI Risk Management Framework version 1.0 provides guidance on assessing representativeness, suitability, relevance, and fairness metrics across demographic groups and intersecting subgroups across different AI lifecycle stages.&lt;/p&gt;
&lt;p&gt;NIST SP 1270, Towards a Standard for Identifying and Managing Bias in Artificial Intelligence, defines statistical parity, error-rate equality, equal opportunity, and related measures as context-specific evaluation metrics.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;When organizations treat these ten requirements as compliance artifacts, the result is a folder full of scorecards and a production incident they cannot trace. Teams fill in fields to satisfy review boards. Data drifts between gate reviews. Labels lose their provenance when annotators turn over and guidelines are not versioned. A model retrains on duplicated, stale records from a source that was deprecated six months ago. The composite scorecard reads 91 percent. The production system misclassifies the users who needed it most. Nobody can explain why, because the lineage was incomplete and the thresholds were never calibrated to actual risk.&lt;/p&gt;
&lt;p&gt;Treat these requirements as an operational discipline and the result is different. Data owners know exactly which failure modes block a release and why. Subgroup gaps surface before training, not after complaints arrive. Audit questions take hours rather than weeks. Retraining follows versioned evidence, regression tests, and documented fairness comparisons. Every threshold has a recorded justification. Every rejected dataset has a preserved exception report. Quality becomes a measurable engineering practice with defined owners, defined gates, and defined escalation paths.&lt;/p&gt;
&lt;p&gt;The model is only as trustworthy as the data contract you can prove you kept.&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and consulting work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative
predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and internationally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance, technical and business requirements.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item><item><title>A Practical Guide for Engineers, Architects, and Governance Teams Who Need to Get It Right</title><link>https://hwyler.github.io/blog/a-practical-guide-for-engineers-architects-and-governance-teams-who-need-to-get-it-right/</link><pubDate>Thu, 30 Jul 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/a-practical-guide-for-engineers-architects-and-governance-teams-who-need-to-get-it-right/</guid><description>&lt;p&gt;Organizations shouldn´t treat AI security as an extension of their existing cybersecurity program. They run the usual penetration tests, validate API authentication, review access controls, and call it done. Then something breaks. A model starts returning outputs it was never designed to produce. A retrieval pipeline exposes data that should have stayed locked. An autonomous agent executes an action nobody authorized.&lt;/p&gt;
&lt;p&gt;The problem is not that organizations are careless. The problem is that AI systems fail in ways that traditional security frameworks were never built to catch. This guide covers the full picture: the threat landscape, the controls that actually work, the governance processes that hold everything together, and the specific decisions you need to make before your next AI system goes live.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/07/chatgpt-image-jul-30-2026-06_03_48-pm.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="why-ai-security-is-a-different-problem"&gt;Why AI Security Is a Different Problem&lt;/h2&gt;
&lt;p&gt;Traditional software is deterministic. Given the same inputs, it produces the same outputs. Its behavior is explicitly programmed and can be inspected through source code. Conventional security frameworks evolved around those assumptions, and they work well for software that behaves predictably.&lt;/p&gt;
&lt;p&gt;AI systems violate every one of those assumptions.&lt;/p&gt;
&lt;p&gt;A model does not execute instructions. It generates probabilistic outputs based on learned patterns. You cannot read its source code to understand what it will do next. Small changes to input can produce dramatically different outputs. The same model, given slightly different context, can behave in entirely different ways. And because AI systems learn from data rather than being explicitly programmed, the data itself becomes an attack surface that has no equivalent in traditional software.&lt;/p&gt;
&lt;p&gt;This is not a theoretical concern. It changes what you need to protect, who is responsible for protecting it, and how you verify that protection is working.&lt;/p&gt;
&lt;h2 id="the-three-delivery-models-you-need-to-account-for"&gt;The Three Delivery Models You Need to Account For&lt;/h2&gt;
&lt;p&gt;Before you can secure an AI system, you need to understand what kind of system you are actually running. There are three common delivery models, and each one carries a different set of responsibilities.&lt;/p&gt;
&lt;p&gt;The first is using a hosted model or AI service. A provider operates the model and its serving infrastructure. You own the security of your application, your prompts, the data you retrieve and inject, the identities with access, the tool permissions, output handling, and monitoring. The provider&amp;rsquo;s security posture matters, but it does not substitute for yours.&lt;/p&gt;
&lt;p&gt;The second is running an externally sourced model on your own infrastructure. In addition to everything in the first case, you now own model selection, artifact integrity, deployment hardening, isolation, patching, and capacity management. The origin and ongoing maintenance of the model become supply chain concerns that belong to you.&lt;/p&gt;
&lt;p&gt;The third is training or adapting a model yourself. On top of both previous cases, you additionally own the training data, the pipeline that processes it, the evaluation process, the resulting model artifacts, and every release decision. Fine-tuning a hosted model falls somewhere between the first and third options, because responsibilities are genuinely shared with the provider.&lt;/p&gt;
&lt;p&gt;Real systems often combine all three. A product might use a hosted general-purpose large language model, a self-hosted image classifier, and a fine-tuned embedding model in the same request path. The mistake organizations consistently make is assigning one security label to the whole product. Record responsibilities per component. That is the only way to know who actually owns each risk.&lt;/p&gt;
&lt;h2 id="the-five-steps-to-organize-ai-security"&gt;The Five Steps to Organize AI Security&lt;/h2&gt;
&lt;p&gt;Once you understand your delivery model, you need a structured approach to actually doing something about it. The most practical framework for this is five sequential steps that build on each other.&lt;/p&gt;
&lt;h3 id="govern-first"&gt;Govern First&lt;/h3&gt;
&lt;p&gt;You cannot secure what you have not inventoried. Start by building a clear picture of where AI is being used in your organization, who owns each system, and what the relevant policies are. This means an AI program that covers development, deployment, procurement, and retirement, with named owners for each system and documented responsibilities across security, engineering, privacy, and compliance.&lt;/p&gt;
&lt;p&gt;This step is not exciting. Organizations consistently underinvest in it because it feels like administrative overhead rather than technical work. But every governance failure that appears later in the lifecycle, unclear ownership during an incident, unreviewed AI systems procured by individual business units, models running in production with no documented
can be traced back to skipping this foundation.&lt;/p&gt;
&lt;p&gt;An original implementation tip: do not treat the AI inventory as a one-time exercise. Shadow AI is a real phenomenon. Employees find hosted AI tools, use them with company data, and create risks the security team does not know about. Build a lightweight intake process that lets teams register new AI use cases before they go into production, and make the barrier low enough that people actually use it. The alternative is discovering the shadow systems after an incident.&lt;/p&gt;
&lt;h3 id="understand-which-threats-actually-apply"&gt;Understand Which Threats Actually Apply&lt;/h3&gt;
&lt;p&gt;The threat landscape for AI systems is large, but not every threat applies to every system. A model used for internal reporting has a completely different risk profile from an autonomous agent with access to external APIs and the ability to send communications on behalf of users.&lt;/p&gt;
&lt;p&gt;The way to navigate this is threat modeling: the process of moving from a catalog of possible attacks to a specific, prioritized list of risks that apply to your system. Walk through each threat type and ask two questions. First, does this threat theoretically apply given the architecture? Second, if it materialized, what would the impact actually be?&lt;/p&gt;
&lt;p&gt;Consider a concrete example. You do not need to protect against model inversion attacks that attempt to reconstruct training data if your training data is not sensitive. It sounds obvious, but the pattern of applying controls without first checking whether the underlying threat is relevant wastes significant security budget.&lt;/p&gt;
&lt;p&gt;The threat types that matter most, and the questions that help you identify which ones apply to your system, fall into three broad areas.&lt;/p&gt;
&lt;p&gt;The first is threats through model inputs. This includes adversarial examples designed to force wrong classifications, prompt injection attacks that use crafted text or hidden instructions to manipulate model behavior, and attempts to extract information about training data or model behavior through systematic querying.&lt;/p&gt;
&lt;p&gt;The second is threats during development and training. This includes data poisoning, where malicious samples are introduced into training data to corrupt model behavior, direct manipulation of model artifacts, and supply chain attacks where a compromised third-party model or dataset introduces vulnerabilities before you even begin.&lt;/p&gt;
&lt;p&gt;The third is conventional security threats applied to AI-specific assets. Model weights, training datasets, prompt templates, and evaluation sets are all assets with significant value and
s. They need the same protection as any other sensitive business asset, and in many cases they need more.&lt;/p&gt;
&lt;h3 id="adapt-your-existing-security-practices"&gt;Adapt Your Existing Security Practices&lt;/h3&gt;
&lt;p&gt;AI security does not replace your existing security program. It extends it. The controls you already have for access management, change control, incident response, and supply chain management all remain relevant. What changes is that AI-specific assets need to be added to your asset inventory, AI-specific threats need to be added to your threat model, and your testing practices need to include AI-specific techniques.&lt;/p&gt;
&lt;p&gt;The most important adaptation is in how you handle the supply chain. If you are using a ready-made model, whether open source or from a commercial provider, that model&amp;rsquo;s training data, training process, and any fine-tuning that happened upstream are all outside your direct control. Proper supply chain management means evaluating provider security posture, understanding what evidence they provide for their controls, and documenting what you have verified and what you are accepting as residual risk.&lt;/p&gt;
&lt;p&gt;Document risk assessment decisions as you make them. This is required under the EU AI Act for high-risk AI systems and it is good practice regardless of regulatory jurisdiction. A risk assessment that exists only in the memory of the person who did it provides no value when that person leaves the organization or when a regulator asks for evidence.&lt;/p&gt;
&lt;h3 id="reduce-potential-impact"&gt;Reduce Potential Impact&lt;/h3&gt;
&lt;p&gt;This step deserves more attention than it typically gets. The underlying principle is simple: AI models can always be wrong or manipulated, so the architecture needs to limit what happens when they are.&lt;/p&gt;
&lt;p&gt;The most important controls here are least privilege for model actions, human oversight for high-impact decisions, and guardrails that constrain what the model can do regardless of what it outputs. In an agentic system where the model can trigger real-world actions, these controls are not optional enhancements. They are the difference between a model error that produces a bad response and a model error that sends an unauthorized communication, executes a financial transaction, or modifies production data.&lt;/p&gt;
&lt;p&gt;Confidential data minimization is equally important. A model that never had access to sensitive data cannot leak it. Apply data minimization before training, before retrieval, and before injecting context into prompts. Every piece of sensitive data that enters the model&amp;rsquo;s context window is data that the model could potentially reproduce in output or expose through inference.&lt;/p&gt;
&lt;h3 id="demonstrate-that-controls-are-working"&gt;Demonstrate That Controls Are Working&lt;/h3&gt;
&lt;p&gt;Governance processes and technical controls only provide value if they demonstrably work. The final step is establishing evidence: through testing, through monitoring, through documentation, and through communication to the stakeholders who need to know the AI systems they rely on are under control.&lt;/p&gt;
&lt;p&gt;This means AI-specific security testing, not just standard penetration testing applied to the API in front of the model. It means continuous validation of model behavior, not just a one-time evaluation before launch. It means monitoring that watches for behavioral drift, unusual query patterns, and resource consumption anomalies that could indicate abuse or attack.&lt;/p&gt;
&lt;h2 id="building-the-risk-case-for-ai-systems-from-quality-objectives-to-funded-decisions"&gt;Building the Risk Case for AI Systems from Quality Objectives to Funded Decisions&lt;/h2&gt;
&lt;p&gt;Organizations trying to govern an AI project make the same sequencing mistake. They start by listing threats, prompt injection, data poisoning, model theft, and then scramble to figure out which ones matter. That order is backwards. A threat only matters once you know what it&amp;rsquo;s threatening, and what it&amp;rsquo;s threatening only becomes clear once you&amp;rsquo;ve named the quality objective the system is supposed to protect in the first place. The working method below reverses that instinct: start with what the AI system needs to preserve, find where the architecture actually fails to preserve it, connect those failures to the ways an attacker or an accident could exploit them, size the resulting exposure in terms a finance or legal team can act on, and then choose, deliberately, whether to build, insure, outsource, reprice, or walk away. Each step depends on the one before it. Skip the first and every later number is a guess dressed up as analysis.&lt;/p&gt;
&lt;h3 id="name-the-quality-objective-before-you-name-a-threat"&gt;Name the Quality Objective Before You Name a Threat&lt;/h3&gt;
&lt;p&gt;Every AI system, whether it&amp;rsquo;s a fraud classifier, a customer support agent, or a document summarizer, exists to protect a small set of properties.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Confidentiality: the training data, the input, the model weights, and anything retrieved into a prompt should stay with the people entitled to see it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Integrity: the model should behave the way it was designed to behave, not the way an attacker or a corrupted dataset nudges it to behave.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Availability: the system should keep answering requests instead of collapsing under a flood of expensive queries.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Beyond those three classic security pillars, AI systems carry two more objectives that conventional software rarely has to worry about at the same intensity: an ethical objective, meaning the system shouldn&amp;rsquo;t produce biased, discriminatory, or harmful outputs even when nobody attacked it, and a business objective, meaning the system needs to actually do the job it was funded to do, accurately enough, often enough, to justify its cost.&lt;/p&gt;
&lt;p&gt;The reason this step has to come first is that it determines everything downstream. A vulnerability only becomes worth discussing once you can say which of these five objectives it threatens. A retrieval pipeline that pulls in unverified vendor documents is a confidentiality and integrity problem if those documents can carry hidden instructions. A fraud model trained eighteen months ago with no retraining trigger is a business-objective and ethical problem, because it silently drifts away from the population it&amp;rsquo;s supposed to be classifying fairly and accurately. Naming the objective at risk before you go looking for a vulnerability keeps the exercise from turning into an unstructured list of scary-sounding attack names that nobody can prioritize.&lt;/p&gt;
&lt;p&gt;A useful discipline here is to walk the system&amp;rsquo;s actual engineering lifecycle and ask, at each stage, which quality objective is on the line. When the team frames the use case and writes acceptance criteria, the question is whether AI should even be used for this task, and what the worst plausible outcome looks like if it&amp;rsquo;s wrong, that&amp;rsquo;s where the ethical and business objectives get defined in the first place. When the team sources or builds the model, the question shifts to trust in the supply chain: can you trust where this model or dataset came from, and what evidence does the vendor actually hand over versus what they simply claim. When the team adapts model behavior through system prompts, retrieval indexes, or fine-tuning data, the live question becomes which untrusted inputs could change how the model behaves, this is where integrity risk concentrates most heavily in modern generative systems. When the model gets wired into an actual product, with tool access, API calls, identities, and secrets attached, the objective at risk expands to include everything the model can now read, modify, or trigger, and under whose permissions it&amp;rsquo;s doing so. Evaluation and release is where you&amp;rsquo;d normally claim the risk is handled, but a test suite only characterizes behavior on the inputs you thought to test, it doesn&amp;rsquo;t prove correctness on the inputs you didn&amp;rsquo;t. And once the system is running, the objective at risk becomes whether you can even detect that something has drifted, been abused, or started failing, before a customer or a regulator notices first.&lt;/p&gt;
&lt;h3 id="find-where-the-architecture-actually-breaks"&gt;Find Where the Architecture Actually Breaks&lt;/h3&gt;
&lt;p&gt;With the objective named, the next step is to look for the specific, concrete weakness in the planned architecture, stack, and deployment circumstances that could let that objective fail. This is different from listing generic attack categories. A vulnerability is a property of your system, not a property of AI in general: weak isolation between trusted system instructions and untrusted retrieved text, a service account with payment permissions far broader than the task requires, a training pipeline with no automated check for population drift, a vector database storing sensitive documents without access control matched to the people who should actually see them.&lt;/p&gt;
&lt;p&gt;Four properties of AI systems make this hunt harder than it is in ordinary software, and worth keeping in mind explicitly while you do it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The model is not the system. A vendor&amp;rsquo;s safety testing on their base model tells you very little about whether your retrieval layer, your agent orchestration, or your output parser introduces a new weakness once that model is wired into your product.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Evaluation characterizes behavior, it does not prove correctness. A test result is only as good as the data, the threat assumptions, the model version, and the configuration it was run against, and all four of those need to travel with the result, not get lost after the fact.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In a generative AI system, data can double as instruction. Anything that lands in the prompt, whether it&amp;rsquo;s a user message, a retrieved PDF, a tool&amp;rsquo;s output, or something pulled from stored memory, can end up steering model behavior even when the engineers who built the pipeline intended it as pure content. That single property is responsible for a huge share of the vulnerabilities showing up in production AI systems today.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Small changes invalidate old evidence. Swap the model version, tweak the prompt template, add a new retrieval source, or adjust a detection threshold, and every piece of testing you did before that change stops being trustworthy until you rerun it.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A practical way to run this stage without missing anything is to draw the actual flow, on a whiteboard or in a diagram, of data, instructions, and actions moving through the system, and at every step, write down the artifact sitting there: which model, which dataset, which prompt template, which retrieval source, which tool, which piece of infrastructure, and who owns or supplies it. That inventory is what turns a vague sense of unease into a specific list of weaknesses you can actually work with.&lt;/p&gt;
&lt;p&gt;AI threat modeling efforts waste time working through a full menu of possible attacks, prompt injection, model inversion, membership inference, evasion, supply chain poisoning, and evaluating every single one regardless of whether it could actually occur given how the system was built. A decision-tree approach fixes that by treating architecture as the filter, not the checklist.&lt;/p&gt;
&lt;p&gt;The method works the way a differential diagnosis works in medicine: rather than asking about every disease in a textbook, a clinician asks about symptoms to eliminate whole categories at once. Applied to AI security, the equivalent questions are architectural, not symptomatic: is this a generative model or a classical predictive one, who trained it, who hosts it, does it pull in external data at inference time, can it trigger downstream actions.&lt;/p&gt;
&lt;p&gt;Each answer removes an entire branch of threats from consideration rather than adding one more item to assess. A classification model with no text generation capability has no exposure to output injection. A system running entirely on a vendor-hosted model with no fine-tuning has no development-time data poisoning surface, because that responsibility sits with the supplier&amp;rsquo;s engineering process, not yours. This narrowing is what separates a useful threat model from an exhaustive but unfocused inventory: it produces a short list of threats that are actually reachable given the system in front of you, not a long list of threats that are theoretically possible somewhere in the universe of AI systems.&lt;/p&gt;
&lt;p&gt;Illustrative example of a decision-tree framework mapping AI threat models across system types:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;If YES → threats to assess&lt;/th&gt;
&lt;th&gt;If NO →&lt;/th&gt;
&lt;th&gt;Next question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Does the system use a predictive or classification model (fraud, credit, medical, spam, etc.)?&lt;/td&gt;
&lt;td&gt;Evasion attacks, adversarial examples, label anddata poisoning&lt;/td&gt;
&lt;td&gt;Skip predictive-specific threats&lt;/td&gt;
&lt;td&gt;Go to 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Is it used for a high-stakes decision (safety, fraud, medical, credit, hiring)?&lt;/td&gt;
&lt;td&gt;Evasion attack severity escalates, treat as high priority&lt;/td&gt;
&lt;td&gt;Evasion risk still applies but lower priority&lt;/td&gt;
&lt;td&gt;Go to 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Is the system Generative AI?&lt;/td&gt;
&lt;td&gt;Direct prompt injection&lt;/td&gt;
&lt;td&gt;Skip all generative-specific threats below&lt;/td&gt;
&lt;td&gt;Go to 6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Does the system insert external or retrieved content into the prompt (RAG, system prompts, tool output, memory)?&lt;/td&gt;
&lt;td&gt;Indirect prompt injection, augmentation data manipulation&lt;/td&gt;
&lt;td&gt;Skip this branch&lt;/td&gt;
&lt;td&gt;Go to 5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Is that retrieved and augmentation data stored somewhere (vector DB, memory store)?&lt;/td&gt;
&lt;td&gt;Augmentation data leak, protect the store itself&lt;/td&gt;
&lt;td&gt;Skip&lt;/td&gt;
&lt;td&gt;Go to 6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Who trained or fine-tuned the model: you, or a supplier?&lt;/td&gt;
&lt;td&gt;You: training-data poisoning, dev-time model leak, model extraction risk. Supplier: supply-chain model poisoning, shift to contractual or supplier assurance&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Go to 7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Who hosts and runs the model: you, or a supplier?&lt;/td&gt;
&lt;td&gt;You: runtime model poisoning, direct runtime model leak, your infra is the attack surface. Supplier: shift to supplier SLA and hosting assurance&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Go to 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Was the training, fine-tuning and augmentation data sensitive?&lt;/td&gt;
&lt;td&gt;Model inversion, membership inference, disclosure-in-output&lt;/td&gt;
&lt;td&gt;Skip data-leak threats&lt;/td&gt;
&lt;td&gt;Go to 9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Is the model wired into an agent, can it invoke tools, APIs, or trigger other agents?&lt;/td&gt;
&lt;td&gt;Agentic threats begin here: excessive tool permissions, goal hijacking, unauthorized tool use, agent-to-agent manipulation&lt;/td&gt;
&lt;td&gt;Worst case bounded to text output, go to 12&lt;/td&gt;
&lt;td&gt;Go to 10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Can the agent&amp;rsquo;s tools send data outward (email, API call, external write, clickable link)?&lt;/td&gt;
&lt;td&gt;Combine with Q11 to test the &amp;ldquo;lethal trifecta&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Exfiltration path closed, lower agentic severity&lt;/td&gt;
&lt;td&gt;Go to 11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Does the agent or a tool it can reach have access to sensitive data?&lt;/td&gt;
&lt;td&gt;If YES to both 10 and 11 → lethal trifecta confirmed: manipulated behavior + data access + exfil path = treat as critical&lt;/td&gt;
&lt;td&gt;Trifecta not complete, de-escalate&lt;/td&gt;
&lt;td&gt;Go to 12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;Does the model or system generate text, code, or markup that gets rendered or executed downstream?&lt;/td&gt;
&lt;td&gt;Output injection (XSS, malicious HTML/JS, unsafe commands)&lt;/td&gt;
&lt;td&gt;Skip&lt;/td&gt;
&lt;td&gt;Go to 13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;Is user and system input sensitive (PII, financial, medical, proprietary)?&lt;/td&gt;
&lt;td&gt;Input data leak, applies regardless of predictive, generative or agentic&lt;/td&gt;
&lt;td&gt;Skip&lt;/td&gt;
&lt;td&gt;Go to 14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;Always evaluate, regardless of prior answers&lt;/td&gt;
&lt;td&gt;Resource exhaustion , denial-of-service, cost abuse, plus conventional app-security controls (identity, logging, patching, infra hardening)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;End&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The sequence itself follows a defensible logic that mirrors how established frameworks such as MITRE ATLAS and the OWASP guidance for LLM and agentic applications structure their own threat catalogs, by attack surface and lifecycle stage rather than by attacker motivation. Generative architecture gets asked first because it gates two of the most consequential threats in production systems today, direct and indirect prompt injection, neither of which applies to a traditional classifier. Training provenance comes next, splitting the analysis cleanly: a self-trained model inherits data poisoning risk during your own pipeline, while a supplier-trained model shifts the relevant question toward contractual assurance and verification of the vendor&amp;rsquo;s own security posture, since you cannot inspect what you didn&amp;rsquo;t build.&lt;/p&gt;
&lt;p&gt;Whether the system augments its input, through retrieval, system prompts, or injected context, determines whether an entirely separate category of threats, augmentation data manipulation and augmentation data leakage, even needs to be on the table. And whether the model can trigger actions rather than simply return text is the single question that most changes the severity ceiling, because a model that can only produce output text has a bounded worst case, while a model wired to send emails, call APIs, or invoke other agents has a worst case defined by whatever permissions those integrations carry.&lt;/p&gt;
&lt;p&gt;Red teams benefit from following this same ordering deliberately: attacking an architecture&amp;rsquo;s actual reachable surface produces findings a development team can act on, while attacking every theoretical LLM vulnerability regardless of whether the system exhibits the precondition produces a report full of noise that erodes the credibility of the genuine findings buried inside it.&lt;/p&gt;
&lt;p&gt;The step that most threat-modeling exercises skip, and that separates a technically complete assessment from an operationally useful one, is asking what happens after a threat is confirmed reachable: does the resulting bad behavior actually reach something worth protecting. A model that can be manipulated into a wrong output is a materially different risk depending on whether that output only displays on a screen or whether it triggers a payment, an email send, or a database write, and depending on whether the system has any path, an API call, an outbound message, a clickable link, capable of moving sensitive data to somewhere an attacker can retrieve it.&lt;/p&gt;
&lt;p&gt;This is the same discipline good penetration testing has always applied to conventional software, treating a vulnerability as inert until an actual exploitation path and consequence are demonstrated, but it matters more for AI systems because the temptation to over-scope is stronger: an LLM is theoretically vulnerable to dozens of named attack classes, and without the architecture-first filtering and the reachability check at the end, both engineering teams and red teams end up spending their limited time defending against threats the system was never actually exposed to, while the two or three threats that genuinely apply, and genuinely have a path to harm, get the same amount of attention as everything else on the list instead of the attention they actually deserve.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/07/modern-device-close-up.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h3 id="connect-the-weakness-to-a-threat-and-the-threat-to-a-real-scenario"&gt;Connect the Weakness to a Threat, and the Threat to a Real Scenario&lt;/h3&gt;
&lt;p&gt;A vulnerability by itself doesn&amp;rsquo;t tell you anything about how bad your day is going to get. It has to be connected to a threat vector, the mechanism an attacker, an insider, or plain negligence would actually use to exploit it, and from there to a concrete scenario involving a specific actor, a specific path, and a specific consequence. This is the step most governance programs skip, jumping straight from &amp;ldquo;we have a vulnerability&amp;rdquo; to &amp;ldquo;here&amp;rsquo;s a control&amp;rdquo;, without ever stating out loud who would exploit it and how.&lt;/p&gt;
&lt;p&gt;Take a retrieval pipeline that ingests vendor-uploaded documents without sanitizing them (the vulnerability) and connect it to indirect prompt injection (the threat vector): an attacker embeds a hidden instruction inside a policy PDF, the system retrieves it, and the model treats the embedded text as an authoritative command rather than as untrusted content, drafting a noncompliant customer communication or leaking information it should have withheld. That&amp;rsquo;s a scenario, not just a vulnerability-threat pairing, because it names the actor, the path, and the outcome.&lt;/p&gt;
&lt;p&gt;Agentic systems deserve special attention here because the scenario-building step gets sharper stakes once a model can take action instead of just producing text. Three conditions have to line up simultaneously for the worst version of this to happen: untrusted data has to be able to reach the model during a session, the model or a connected agent has to have access to sensitive information, and that same model or agent has to have some way of sending data back out, an email tool, an API call, a link a user might click. When all three are present at once, a single successful manipulation of model behavior turns directly into data leaving the organization, and no amount of confidentiality control on the data itself will help if the exfiltration path through the model was never closed.&lt;/p&gt;
&lt;p&gt;Not every theoretically possible threat deserves a scenario, and this is worth saying plainly because over-scoping wastes as much governance effort as under-scoping. If a classification model&amp;rsquo;s training data was never sensitive to begin with, there&amp;rsquo;s no meaningful scenario for someone stealing it through model inversion, the vulnerability might technically exist, but there&amp;rsquo;s no path to a consequence worth pricing. The discipline of building an actual scenario, actor plus path plus consequence, is what filters a long catalog of theoretical weaknesses down to the short list that actually deserves budget.&lt;/p&gt;
&lt;h3 id="size-the-exposure-and-choose-where-the-money-goes"&gt;Size the Exposure and Choose Where the Money Goes&lt;/h3&gt;
&lt;p&gt;Once you have real scenarios instead of abstract threat categories, the next step is to put a number on each one, or at least a defensible range, covering both how often it&amp;rsquo;s likely to happen and how much it costs when it does. This is where most AI governance documentation quietly gives up and reaches for a red, yellow, green heat map instead, which feels like an answer but isn&amp;rsquo;t one, because a color tells a board nothing about whether the exposure behind it is ten thousand dollars or ten million.&lt;/p&gt;
&lt;p&gt;The prioritization that follows from a properly sized exposure has more options on the table than most teams initially assume, and naming all of them explicitly changes the conversation from &amp;ldquo;how do we fix this&amp;rdquo; to &amp;ldquo;what&amp;rsquo;s the most economical way to handle this&amp;rdquo;. A project can be rejected outright, when the exposure is large, the mitigation is expensive or technically unproven, and the business case doesn&amp;rsquo;t survive the honest number. A project can be accepted as presented, when the exposure is genuinely small relative to the benefit, and forcing controls onto it would cost more than the risk itself. Risk can be financed rather than engineered away, through cybersecurity or professional liability insurance sized to the calculated exposure, or by outsourcing the riskiest components, model hosting, fine-tuning, or specialized data handling, to a vendor better positioned to carry that risk than you are. Contract terms can shift the exposure directly: tightening warranties on a vendor&amp;rsquo;s model behavior, changing the pricing of a service to reflect its actual risk profile, or negotiating indemnification clauses that put the cost of a failure where it&amp;rsquo;s cheapest to absorb it. And of course, the exposure can be reduced directly through internal technical and compliance controls, retraining triggers, output filtering, scoped service credentials, human review gates, each control chosen because its cost is smaller than the expected loss it prevents, not because it appeared on a generic best-practices list.&lt;/p&gt;
&lt;p&gt;The organizations that get real value out of this process are the ones that treat quantification as a discipline applied consistently, scenario by scenario, rather than as a one-time slide for a steering committee. A fraud-detection model with a known drift vulnerability, sized honestly, might show an expected loss in the tens of thousands of dollars if caught within two weeks and hundreds of thousands if it runs unnoticed for a quarter, numbers a finance team can reserve against, insure, or fund a control for. A vague &amp;ldquo;medium risk&amp;rdquo; rating on the same model tells that finance team nothing they can act on. The entire value of walking through quality objectives, vulnerabilities, threats, scenarios, and exposure in that specific order is that it ends, every time, at a number and a named decision, not at a color and a shrug.&lt;/p&gt;
&lt;h2 id="the-threat-landscape-in-detail"&gt;The Threat Landscape in Detail&lt;/h2&gt;
&lt;h3 id="what-can-go-wrong-with-model-inputs"&gt;What Can Go Wrong With Model Inputs&lt;/h3&gt;
&lt;p&gt;Input threats are attacks that happen through the normal operation of the model. The attacker provides input and reads the output. No special access to infrastructure is required.&lt;/p&gt;
&lt;p&gt;Prompt injection is the most widely discussed input threat, and for good reason. In a system where the model receives natural language instructions, any source of text that the model processes becomes a potential instruction channel. An attacker who can place content into a document, a web page, a database record, or any other source that gets retrieved and inserted into a prompt can potentially influence model behavior. This is called indirect prompt injection, and it is the key threat in most agentic AI systems because the model has no reliable built-in way to distinguish instructions it was given from data it was asked to process.&lt;/p&gt;
&lt;p&gt;Direct prompt injection, where a user tries to override system instructions through their own input, is the more visible version of the same problem. Both require defense in depth: model alignment to reduce susceptibility, filtering at the input and output layers, and critically, architectural controls that limit what the model can do even if the injection succeeds. If a successfully injected prompt cannot trigger a harmful action because the architecture does not permit that action, the attack&amp;rsquo;s blast radius is contained.&lt;/p&gt;
&lt;p&gt;Evasion attacks target classification models. The attacker crafts input, sometimes imperceptibly different from legitimate input, that forces the model to make an incorrect decision. The relevance of this threat depends entirely on whether there is a plausible attacker with a plausible benefit from fooling the model. A spam filter is a meaningful target. A skin disease diagnostic tool used by a patient with no obvious motive to manipulate the result is a much lower-risk target in most contexts.&lt;/p&gt;
&lt;p&gt;Model extraction happens when an attacker uses the model&amp;rsquo;s outputs to approximate the model&amp;rsquo;s behavior, effectively stealing its functionality through systematic querying. Rate limiting, output truncation, and monitoring for query patterns consistent with extraction are the relevant controls.&lt;/p&gt;
&lt;h3 id="what-can-go-wrong-during-development"&gt;What Can Go Wrong During Development&lt;/h3&gt;
&lt;p&gt;Development-time threats are often underestimated because they happen before the system goes live. But the vulnerabilities introduced during development follow the model into production.&lt;/p&gt;
&lt;p&gt;Data poisoning is the introduction of malicious samples into training data to corrupt model behavior. This can be a deliberate attack where an adversary gains access to the training pipeline, or it can happen through the use of external data sources that have been compromised without your knowledge. The controls are quality assurance on training data, anomaly detection for samples that look inconsistent with the rest of the dataset, and careful supply chain management for any data sourced externally.&lt;/p&gt;
&lt;p&gt;Model poisoning at the supply chain level means receiving a model artifact that has been manipulated before you acquired it. An open source model downloaded from a public repository could contain a backdoor that activates only under specific input conditions. Verifying artifact integrity and testing acquired models for unexpected behaviors are the relevant controls.&lt;/p&gt;
&lt;p&gt;The development environment itself is an attack surface. Model weights, training datasets, evaluation sets, and configuration files stored in development environments need access controls, encryption, and integrity verification just like production assets. Breaches of development environments often remain undetected for extended periods precisely because development environments have historically received less security attention than production.&lt;/p&gt;
&lt;h3 id="what-can-go-wrong-at-runtime"&gt;What Can Go Wrong at Runtime&lt;/h3&gt;
&lt;p&gt;Runtime threats beyond input attacks include the full range of conventional security threats applied to AI-specific assets.&lt;/p&gt;
&lt;p&gt;Model weights stored in production need protection from both disclosure and modification. A model that an attacker can read can be used to craft more effective evasion attacks. A model that an attacker can modify is a model that can be reprogrammed to behave in whatever way the attacker chooses. Encryption at rest, integrity verification, and strict access controls are the baseline.&lt;/p&gt;
&lt;p&gt;Augmentation data, which includes the content retrieved for retrieval-augmented generation systems and the system prompts that define model behavior, is a high-value target. If an attacker can modify what gets retrieved and injected into prompts, they effectively control part of the model&amp;rsquo;s context. Integrity protection for retrieval stores and system prompt management are therefore security controls, not just operational considerations.&lt;/p&gt;
&lt;p&gt;Resource exhaustion is a meaningful threat for large language model deployments because inference costs money. An attacker who can force the system to process large volumes of expensive requests can create significant cost and availability problems. Rate limiting, session budgets, and cost monitoring are the relevant controls.&lt;/p&gt;
&lt;h2 id="agentic-ai-when-the-stakes-get-higher"&gt;Agentic AI: When the Stakes Get Higher&lt;/h2&gt;
&lt;p&gt;Agentic AI systems deserve particular attention because they change the consequences of every other threat. When a model can trigger real-world actions rather than just produce text output, the impact of prompt injection, data poisoning, or any other successful attack is no longer limited to a bad response. It extends to whatever the agent is capable of doing.&lt;/p&gt;
&lt;p&gt;There is a useful concept called the lethal trifecta for understanding data exfiltration risk in agentic systems. You need three conditions to be simultaneously present for an attacker to exfiltrate data through a manipulated agent: the ability to inject malicious instructions into data the model processes, the model&amp;rsquo;s access to sensitive data within the session, and the model&amp;rsquo;s ability to send that data to an external destination. If any one of these three conditions is absent, the exfiltration attack fails. Removing one of the three through architecture is often more practical than trying to prevent the injection itself.&lt;/p&gt;
&lt;p&gt;Least model privilege is the foundational control for agentic systems. Assign only the permissions the agent needs for its specific task. Separate read and write permissions. Require explicit approval for high-impact actions. These principles are well-established in conventional software security, but they require conscious application to agentic architectures where developers often assign broad permissions for convenience during development and never revisit those decisions before production.&lt;/p&gt;
&lt;p&gt;Human oversight, meaning meaningful human review at decision points that matter, is a control, not just a policy preference. An agent that can take consequential actions without any human checkpoint in the path is an agent where model errors, manipulated behaviors, and unexpected outputs translate directly into real-world consequences with no opportunity to intervene.&lt;/p&gt;
&lt;h2 id="the-specific-risks-of-generative-ai"&gt;The Specific Risks of Generative AI&lt;/h2&gt;
&lt;p&gt;Generative AI systems share most of their threat landscape with other AI types, but several risks are materially higher or take different forms.&lt;/p&gt;
&lt;p&gt;System prompts, the instructions that define how a hosted model should behave, are both a security control and an attack surface. They represent sensitive intellectual property that should be protected from disclosure, and they are a target for prompt injection attacks trying to override their content. Organizations frequently treat system prompts as configuration files without applying the access controls and integrity verification they would apply to any other sensitive configuration.&lt;/p&gt;
&lt;p&gt;Retrieval-augmented generation systems introduce a particularly important input data risk. The content retrieved and injected into prompts often includes sensitive company information, personal data, or proprietary business logic. This content travels to the model provider&amp;rsquo;s infrastructure in clear text if the model is externally hosted, it may not respect the original access controls that governed who could read the source documents, and it exists in the model&amp;rsquo;s context window where it can potentially appear in outputs. Assess what is being retrieved, verify that the retrieval respects access controls, and apply data minimization to limit what sensitive content reaches the prompt.&lt;/p&gt;
&lt;p&gt;Training data memorization is a genuine risk for large language models. A model trained on sensitive data can sometimes reproduce specific examples from that training set in its outputs. Testing for memorization before deployment, applying data minimization during training, and using privacy-preserving techniques during fine-tuning are the relevant controls.&lt;/p&gt;
&lt;p&gt;Output injection is often overlooked. When model output is rendered in a browser or executed in some downstream process without proper encoding, it can contain content that performs injection attacks. This is a conventional security control applied to an unconventional output source, but organizations sometimes fail to apply their existing output encoding practices to AI-generated content.&lt;/p&gt;
&lt;h2 id="risk-assessment-moving-from-threats-to-decisions"&gt;Risk Assessment: Moving From Threats to Decisions&lt;/h2&gt;
&lt;p&gt;Identifying threats is necessary but not sufficient. Every identified threat needs to be evaluated for likelihood and impact in your specific context, and then treated through one of four options.&lt;/p&gt;
&lt;p&gt;Treatment means implementing controls to reduce the likelihood or impact of the risk. This is the most common approach and the bulk of what this guide covers.&lt;/p&gt;
&lt;p&gt;Transfer means shifting the risk to a third party, through insurance, contractual agreements, or using a provider who takes on the relevant security responsibilities. This only works when you have verified that the third party is actually managing the risk, not just accepting contractual liability.&lt;/p&gt;
&lt;p&gt;Termination means changing the approach to eliminate the risk entirely. Sometimes the right answer is not to use AI for a particular application because the risk cannot be adequately managed. Removing an unnecessary AI component eliminates all AI-related risks for that component.&lt;/p&gt;
&lt;p&gt;Tolerance means acknowledging a risk and deciding to bear the potential consequences without further action. This is appropriate when the cost of treatment exceeds the expected impact. It requires explicit documentation of who made the acceptance decision and why, because an undocumented accepted risk is indistinguishable from an overlooked risk.&lt;/p&gt;
&lt;p&gt;When assessing likelihood, consider the attacker&amp;rsquo;s realistic motivation. Would an attacker actually benefit from fooling your model? What would they need to do to succeed? What is their likely budget and capability? Threats that exist in theory but have no plausible attacker with a plausible motive can often be accepted or managed with light controls.&lt;/p&gt;
&lt;p&gt;When assessing impact, consider the full chain of consequences. Direct technical consequences like compromised data integrity are usually the most visible. Indirect consequences like regulatory penalties, reputational damage, and loss of customer trust often matter more to the organization. In regulated industries, a security incident affecting an AI system may trigger reporting obligations and regulatory scrutiny that dwarf the direct technical cost of the incident.&lt;/p&gt;
&lt;h2 id="the-controls-that-actually-work"&gt;The Controls That Actually Work&lt;/h2&gt;
&lt;p&gt;Selecting controls requires matching the control to the threat, the system type, and the level of risk. Here is the practical breakdown organized by what each control category addresses.&lt;/p&gt;
&lt;p&gt;For governance and accountability, the essential controls are an AI program that inventories all AI use and assigns ownership, a security program that includes AI-specific assets and threats, compliance checking against applicable regulations, and ongoing security education for everyone who builds and operates AI systems. These are not glamorous controls. They are the foundation that makes every other control meaningful.&lt;/p&gt;
&lt;p&gt;For the supply chain, the key control is treating every external model, dataset, and hosting provider as a potential source of inherited risk. Verify provider security posture before adoption. Test acquired models in your own context rather than relying solely on published benchmarks. Track and patch dependencies in AI infrastructure with the same discipline applied to application dependencies. This last point deserves emphasis: teams frequently delay patching AI infrastructure components because they fear breaking model reproducibility. That hesitation creates a predictable, accumulating vulnerability.&lt;/p&gt;
&lt;p&gt;For protecting sensitive data, apply data minimization consistently. The less sensitive data that enters training pipelines, retrieval systems, and prompts, the smaller the disclosure risk. Obfuscate or remove sensitive values from training data. Apply short retention periods for data that does not need to be kept. Test your de-identification approaches for realistic re-identification risk, not just surface-level masking.&lt;/p&gt;
&lt;p&gt;For model behavior integrity, the engineering controls during model development include adversarial training, model alignment techniques, ensemble approaches that reduce the impact of any single manipulated component, and continuous validation that tracks model behavior against approved baselines over time. At runtime, input filtering, output filtering, anomaly detection, and rate limiting form the monitoring and detection layer.&lt;/p&gt;
&lt;p&gt;For runtime protection, access controls on model endpoints, integrity verification of model artifacts before serving, encryption for model parameters and inference data, and monitoring that watches for behavioral patterns consistent with attack or abuse form the defensive layer.&lt;/p&gt;
&lt;h2 id="responsibility-assignment-who-owns-what"&gt;Responsibility Assignment: Who Owns What&lt;/h2&gt;
&lt;p&gt;For every threat you identify, someone needs to own the response. In AI systems with multiple components from multiple sources, responsibility is frequently unclear.&lt;/p&gt;
&lt;p&gt;When a component is hosted by a provider, you share responsibility for that component&amp;rsquo;s security with the provider. The division depends on the specific hosting arrangement. Use a responsibility matrix to document which controls you own, which the provider owns, and which are shared. Then verify that the provider is actually implementing the controls assigned to them. Provider attestations and third-party audits are more reliable than self-reported compliance.&lt;/p&gt;
&lt;p&gt;When a provider is not transparent about their security practices, you face three options. Accept the risk based on your assessment that the provider&amp;rsquo;s posture is adequate even without verification. Implement your own compensating controls to address the risks the provider may not be managing. Or avoid using that provider for the application in question. The worst outcome is assuming the provider has it covered without checking.&lt;/p&gt;
&lt;p&gt;For internally developed or fine-tuned models, your organization owns the entire stack. That means the training data pipeline, the model artifacts, the evaluation process, the deployment environment, the runtime controls, and the ongoing monitoring. The breadth of this responsibility is why organizations with limited AI security maturity are often better served by starting with externally hosted models for lower-risk applications while building internal capability.&lt;/p&gt;
&lt;h2 id="standardize-your-ai-assessments-with-hernan-huwylers-threat-modeling-toolkit"&gt;Standardize Your AI Assessments with Hernan Huwyler´s Threat Modeling Toolkit&lt;/h2&gt;
&lt;p&gt;You cannot secure an AI pipeline with a generic IT checklist. Traditional application security focuses heavily on the API wrapper, identity layers, and network configurations. It completely misses the attack surface unique to machine learning: poisoned training data, instruction overrides in system prompts, and unauthorized actions executed by autonomous agents. I built the 
 to give architects, risk managers, and security engineers a deterministic, repeatable way to move from abstract security theory to an actionable, architecture-specific threat model.&lt;/p&gt;
&lt;p&gt;The toolkit provides a highly structured methodology tailored specifically to the type of AI system you are actually building. A predictive fraud model requires fundamentally different security controls than a Retrieval-Augmented Generation (RAG) chatbot or a multi-agent workflow. The repository ships with a 
, allowing you to script, filter, and score vulnerabilities programmatically. By running the included Python script (&lt;code&gt;generate_checklist.py&lt;/code&gt;), your team can instantly generate a precise assessment scope customized to your system type and sourcing model (built vs. procured), ensuring you never waste time evaluating irrelevant risks.&lt;/p&gt;
&lt;p&gt;Every vulnerability and threat vector within this toolkit is firmly anchored to community consensus. Instead of relying on isolated opinions, the catalogs are 
, including MITRE ATLAS, the OWASP Top 10 for LLM and Agentic Applications, NIST AI 100-2, and ISO/IEC 42001. Whether you are building an 
 before a red-team engagement or mapping classic STRIDE trust boundaries to an AI context, this open-source repository provides the exact templates and technical guidance required to execute a rigorous, defensible assessment.&lt;/p&gt;
&lt;p&gt;The 
 links ISO/IEC 42001 Annex A controls directly to the vulnerability catalog, giving teams a traceable path from identified weakness to documented control requirement. For practitioners who need the full narrative behind each catalog entry, the 
 provides complete detail on every cataloged vulnerability without summarizing, and the 
 does the same for every threat vector, explaining the attack path, the system types most exposed, and the controls that address it. When an assessment moves from analysis into reporting, the 
 provides a fillable, questionnaire-driven structure designed for red-team engagements, covering system classification, asset inventory findings, threat modeling results, control gaps, and risk acceptance decisions in a format that holds up under audit review.&lt;/p&gt;
&lt;p&gt;The 
 cross-references every catalog entry against the frameworks it maps to, so the catalog stays anchored to community consensus rather than one team&amp;rsquo;s judgment. Assessment outputs go into the 
 and the 
, both designed to produce artifacts that hold up under audit review. The toolkit is a living document: new attack techniques against AI systems are documented on a rolling basis, and the 
 sets out how to propose new entries, update mappings, or correct citations as the field moves.&lt;/p&gt;
&lt;h2 id="what-testing-ai-security-actually-looks-like"&gt;What Testing AI Security Actually Looks Like&lt;/h2&gt;
&lt;p&gt;AI security testing is not just penetration testing applied to an AI API. It requires techniques specific to AI threats.&lt;/p&gt;
&lt;p&gt;Adversarial testing for input threats means systematically crafting inputs designed to force wrong decisions, expose training data, extract model behavior, or manipulate outputs in harmful ways. For prompt injection specifically, it means testing with a wide range of injection attempts across multiple input channels, including indirect injection through retrieved content. Red team exercises that simulate an attacker trying to achieve a specific harmful outcome through the model are more valuable than checklist-based assessments.&lt;/p&gt;
&lt;p&gt;Model behavior validation before release and continuously in production means maintaining a held-out evaluation set with known correct outputs and testing the model against it regularly. Any significant change to model behavior, whether from a model update, a prompt change, or a retrieval index update, should trigger revalidation. The evaluation set needs to include adversarial examples and edge cases, not just typical production inputs.&lt;/p&gt;
&lt;p&gt;Supply chain verification means testing acquired model artifacts for integrity, checking for known vulnerabilities in the model&amp;rsquo;s dependencies, and where possible, running behavioral tests designed to surface backdoors or unusual behaviors that would not appear in standard accuracy evaluation.&lt;/p&gt;
&lt;p&gt;Privacy testing means evaluating whether the model can reproduce specific training data examples, whether embeddings can be used to reconstruct sensitive information, and whether de-identification approaches hold up against realistic linkage attacks.&lt;/p&gt;
&lt;h2 id="documentation-monitoring-and-the-long-tail"&gt;Documentation, Monitoring, and the Long Tail&lt;/h2&gt;
&lt;p&gt;The security work done before deployment matters. The monitoring and response capability after deployment matters equally.&lt;/p&gt;
&lt;p&gt;Monitoring for AI systems needs to go beyond infrastructure metrics. Uptime and latency tell you whether the system is running. They do not tell you whether it is behaving as intended, whether it is being probed for vulnerabilities, whether its outputs are drifting in quality or safety, or whether its resource consumption is consistent with legitimate use. Build monitoring that watches model behavior and output characteristics alongside infrastructure health.&lt;/p&gt;
&lt;p&gt;Incident response procedures for AI systems need to account for the specific ways AI incidents differ from conventional software incidents. The relevant artifacts include logs of model inputs and outputs, records of which model version and which retrieval content were in use at the time, and behavioral validation results that can establish what the model was doing before and after the incident. If those logs do not exist or were not retained, incident reconstruction becomes extremely difficult.&lt;/p&gt;
&lt;p&gt;Documentation of risk assessments, control selections, and residual risk acceptance decisions creates the evidentiary record that regulators, auditors, and board committees will ask for. Under frameworks like the EU AI Act, this documentation is a legal requirement for high-risk AI systems. Even outside regulated contexts, documented decisions are the foundation for organizational learning. An organization that documents why it made a specific risk acceptance decision can revisit and update that decision as circumstances change. An organization that does not document its decisions is perpetually starting from scratch.&lt;/p&gt;
&lt;h2 id="key-standards-and-frameworks"&gt;Key Standards and Frameworks&lt;/h2&gt;
&lt;p&gt;The field has developed a body of standards and guidance that provide the technical foundation for AI security programs. ISO/IEC 42001 establishes requirements for AI management systems, providing the governance framework within which security controls operate. ISO/IEC 27090 addresses AI security specifically and is currently in development with substantial community contribution shaping its content. ISO/IEC 27091 addresses AI privacy. I
&lt;/p&gt;
&lt;p&gt;At the regulatory level, the EU AI Act establishes mandatory requirements for high-risk AI systems, including risk management, technical documentation, data governance, transparency, human oversight, and post-market monitoring. NIST&amp;rsquo;s AI Risk Management Framework provides a voluntary but widely adopted structure for identifying, assessing, and managing AI risks organized around four core functions. The UK NCSC and CISA joint guidelines for secure AI system development provide practical guidance organized around secure design, development, deployment, and operation.&lt;/p&gt;
&lt;p&gt;These frameworks are not mutually exclusive. ISO/IEC 42001 provides the management system. NIST AI RMF provides the risk management process. Sector-specific regulations like the EU AI Act establish mandatory baseline requirements. A mature AI security program typically draws on all of them, using each framework where it provides the most useful structure.&lt;/p&gt;
&lt;h2 id="the-difference-between-documentation-and-practice"&gt;The Difference Between Documentation and Practice&lt;/h2&gt;
&lt;p&gt;An AI security program built entirely around documentation produces governance artifacts that satisfy auditors and inform no one. Risk registers that record threats without owners. Control frameworks that describe practices nobody follows. Compliance checklists completed after decisions are made rather than before.&lt;/p&gt;
&lt;p&gt;The organizations that actually reduce AI security risk treat governance artifacts as operational tools, not as endpoints. The risk register is updated when new AI systems come online and when existing systems change. The threat model is revisited when the architecture changes or when new attack techniques emerge. Control effectiveness is verified through testing, not assumed through documentation. Residual risk acceptance decisions are made by people with the authority and information to make them, and those decisions are recorded with enough context that they can be revisited meaningfully when circumstances change.&lt;/p&gt;
&lt;p&gt;The technical controls matter. The governance processes that ensure those controls remain effective over time matter just as much. An AI system that was secure at launch and has drifted due to model updates, changing retrieval content, or evolving attack techniques is not a secure AI system. Continuous validation, ongoing monitoring, and periodic reassessment are not optional enhancements for organizations with extra budget. They are how security is maintained in a technology domain where the threat landscape and the systems themselves are both changing continuously.&lt;/p&gt;
&lt;p&gt;Getting AI security right requires understanding the specific ways AI systems fail, building the controls that address those failures, and maintaining the governance processes that keep those controls effective. Start with the inventory, do the threat modeling, assign the responsibilities, implement the controls proportional to the risk, test them, monitor them, and document the decisions. That is the full picture.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;p&gt;ISO/IEC 42001:2023 - Artificial Intelligence Management Systems&lt;/p&gt;
&lt;p&gt;
(in development, draft for approval)&lt;/p&gt;
&lt;p&gt;ISO/IEC 27091 - Privacy and AI (in development)&lt;/p&gt;
&lt;p&gt;ISO/IEC 27005:2022 - Information Security Risk Management&lt;/p&gt;
&lt;p&gt;ISO/IEC 23894:2023 - AI Risk Management Guidance&lt;/p&gt;
&lt;p&gt;NIST AI Risk Management Framework 1.0 (January 2023): 
&lt;/p&gt;
&lt;p&gt;EU Artificial Intelligence Act, Official Journal of the European Union (2024)&lt;/p&gt;
&lt;p&gt;UK NCSC / CISA Joint Guidelines for Secure AI System Development: 
&lt;/p&gt;
&lt;p&gt;DSIT Code of Practice for the Cyber Security of AI (UK): 
&lt;/p&gt;
&lt;p&gt;MITRE ATLAS - Adversarial Threat Landscape for AI Systems: 
&lt;/p&gt;
&lt;p&gt;OpenCRE - Common Requirements Enumeration for AI Security Standards: 
&lt;/p&gt;
&lt;p&gt;SANS Critical AI Security Guidelines: 
&lt;/p&gt;
&lt;p&gt;AI Security Verification Standard (AISVS): 
&lt;/p&gt;
&lt;p&gt;By Prof. Hernan Huwyler, CAIO MBA CPA&lt;br&gt;
&lt;br&gt;
&lt;br&gt;
&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and advisory work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution. If you like the content, please like the article and share it.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative
predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe and internationally.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance, technical and business requirements.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item><item><title>Managing AI Development and Deployment Projects</title><link>https://hwyler.github.io/blog/managing-ai-development-and-deployment-projects/</link><pubDate>Fri, 13 Mar 2026 00:00:00 +0000</pubDate><guid>https://hwyler.github.io/blog/managing-ai-development-and-deployment-projects/</guid><description>&lt;h2 id="the-10-best-practices-that-separate-ai-projects-that-ship-from-ai-projects-that-stall"&gt;The 10 Best Practices That Separate AI Projects That Ship From AI Projects That Stall&lt;/h2&gt;
&lt;p&gt;Managing AI development and deployment projects requires practices fundamentally different from traditional software project management. AI systems derive behavior from training data rather than human-written code. They exhibit opacity, drift, and emergent properties that deterministic software doesn&amp;rsquo;t. A model that performs well during testing may degrade in production as real-world data evolves. A system that&amp;rsquo;s technically accurate may still fail from a compliance, fairness, or adoption standpoint.&lt;/p&gt;
&lt;p&gt;Most AI projects fail because the project was managed like ordinary software, governed too late, monitored too lightly, or deployed before the organization was ready to support it. Teams rush from prototype to launch, then discover that the data does not hold up, the model drifts in production, the vendor changes core behavior, users do not trust the output, or compliance asks questions nobody planned to answer. By then, delivery slows, confidence drops, and the business case gets harder to defend.&lt;/p&gt;
&lt;p&gt;A strong AI project needs a management approach built for experimentation, risk, operational change, and continuous improvement. This post brings together the practical best practices from the material you provided, including governance, MLOps, risk-based lifecycle controls, third-party oversight, phased deployment, continuous monitoring, and value tracking. The goal is simple. Help teams build and deploy AI systems that actually work in the real world and keep working after launch.&lt;/p&gt;
&lt;p&gt;This post covers the ten best practices that address these challenges: from governance structure through lifecycle management, MLOps implementation, regulatory compliance, third-party risk, phased deployment, human oversight, continuous monitoring, organizational literacy, and value measurement.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/data-center-technician.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="why-ai-projects-require-different-management-than-software-projects"&gt;Why AI Projects Require Different Management Than Software Projects&lt;/h2&gt;
&lt;p&gt;AI projects differ from conventional software development in ways that demand adapted management approaches. Three characteristics make traditional project management insufficient.&lt;/p&gt;
&lt;p&gt;First, AI development is inherently experimental. Unlike software where requirements can be specified and development follows a predictable path, AI model performance cannot be guaranteed until training is complete and validation is run. A technically sound model may not achieve business objectives due to data limitations, feature interactions, or distribution mismatches. Project plans must account for this uncertainty rather than treating model development as a deterministic activity with fixed timelines.&lt;/p&gt;
&lt;p&gt;Second, AI systems change after deployment without anyone modifying code. Data drift, concept drift, and population shifts cause model performance to degrade over time. A software application behaves the same on day 500 as on day 1. An AI model does not. This means deployment is the beginning of the maintenance lifecycle, not the end of the development lifecycle.&lt;/p&gt;
&lt;p&gt;Third, AI systems create novel risk categories. Algorithmic bias, hallucination, adversarial vulnerability, training data leakage, and model opacity don&amp;rsquo;t exist in traditional software. Managing these risks requires specialized controls that traditional project management frameworks don&amp;rsquo;t include.&lt;/p&gt;
&lt;p&gt;These three characteristics mean that success criteria, timeline expectations, governance structures, and post-deployment plans all need to be designed specifically for AI rather than adapted from software development templates.&lt;/p&gt;
&lt;p&gt;Implementation tip: Build flexibility into every AI project plan by defining two types of milestones: fixed milestones (governance approvals, compliance checkpoints, deployment dates) and adaptive milestones (model performance targets, data quality thresholds, accuracy objectives). Fixed milestones maintain project structure and stakeholder accountability. Adaptive milestones acknowledge that model development is experimental and may require iteration. When a project plan treats accuracy targets as fixed milestones with hard deadlines, teams either compromise on validation rigor to meet the date or blow past the deadline repeatedly. When accuracy targets are adaptive milestones with defined evaluation criteria and go/no-go decision procedures, the project maintains momentum while accommodating the inherent uncertainty of model development.&lt;/p&gt;
&lt;h2 id="best-practice-1-establish-clear-governance-and-accountability-structures"&gt;Best Practice 1: Establish Clear Governance and Accountability Structures&lt;/h2&gt;
&lt;p&gt;Effective AI project management begins with defined governance roles and decision rights. Organizations should build a structured AI management system aligned with ISO/IEC 42001, establishing clear accountability for each AI system through three distinct roles.&lt;/p&gt;
&lt;p&gt;A business owner is accountable for outcomes and compliance. This person owns the business case, defines success metrics, and bears responsibility for the system&amp;rsquo;s impact on users and the organization. A technical lead is responsible for model performance. This person owns model architecture decisions, training methodology, validation results, and technical documentation. A risk owner manages ongoing monitoring. This person owns post-deployment surveillance, drift detection, incident response, and the decision to retrain, roll back, or retire the system.&lt;/p&gt;
&lt;p&gt;These three roles may be filled by different people or combined in smaller organizations, but the responsibilities must be explicitly assigned. Unassigned responsibilities don&amp;rsquo;t get fulfilled.&lt;/p&gt;
&lt;p&gt;Project managers should ensure that every AI initiative has documented approval gates, with an AI ethics or review board empowered to condition or reject use cases at key lifecycle stages. This governance structure should integrate with existing risk management frameworks rather than operate separately.&lt;/p&gt;
&lt;p&gt;Implementation tip: The governance structure must have the authority to stop a project, not just review it. Many AI governance boards operate as advisory bodies that provide recommendations but lack enforcement power. When the governance board recommends against deployment but the business sponsor overrides the recommendation, governance becomes performative. Grant your governance structure explicit authority over three decisions: use case approval (can we build this), deployment approval (can we launch this), and continuation approval (should we keep running this). Without authority over these three gates, governance provides commentary rather than control.&lt;/p&gt;
&lt;h2 id="best-practice-2-implement-risk-based-lifecycle-management"&gt;Best Practice 2: Implement Risk-Based Lifecycle Management&lt;/h2&gt;
&lt;p&gt;Organizations should adopt a risk-based approach that applies governance intensity proportional to potential harm. A low-risk internal productivity tool doesn&amp;rsquo;t need the same oversight as a high-risk system making decisions about individuals&amp;rsquo; access to credit, healthcare, or employment.&lt;/p&gt;
&lt;p&gt;The AI lifecycle should include five structured phases, each with documented governance decision points.&lt;/p&gt;
&lt;p&gt;Business case identification defines the problem, expected value, and success metrics before technical work begins. This phase prevents the common failure of building solutions before confirming they solve the right problem.&lt;/p&gt;
&lt;p&gt;Design and data preparation assesses data availability, quality, and potential bias. This phase documents data provenance and identifies representativeness gaps before model development commits to specific data sources.&lt;/p&gt;
&lt;p&gt;Development and testing evaluates model performance, fairness, and robustness against defined criteria. This phase produces the validation evidence that supports deployment decisions.&lt;/p&gt;
&lt;p&gt;Deployment ensures that integration, monitoring, and compliance controls are in place before the system goes live. This phase confirms operational readiness, not just model readiness.&lt;/p&gt;
&lt;p&gt;Ongoing monitoring tracks drift, performance degradation, and emerging risks continuously after deployment. This phase maintains the system&amp;rsquo;s trustworthiness over time rather than assuming that deployment-time performance persists.&lt;/p&gt;
&lt;p&gt;Higher-risk applications require more rigorous validation and oversight at each phase. A classification system for AI risk levels (following the EU AI Act&amp;rsquo;s risk tiers or an internal equivalent) determines the governance intensity applied at each gate.&lt;/p&gt;
&lt;p&gt;Implementation tip: Conduct regulatory classification during the planning phase, not after development. Discovering that a system falls under high-risk classification after months of development typically requires redesign and delays deployment. By early 2026, over 72 countries have launched more than 1,000 AI policy initiatives, with the EU AI Act imposing fines up to 35 million euros or 7% of global turnover for non-compliance. Map your AI systems against applicable regulations based on where systems are developed, deployed, and whose data they process. Use ISO 42001 as a common governance layer that can be mapped to multiple regional requirements, reducing duplication while maintaining defensibility across jurisdictions.&lt;/p&gt;
&lt;h2 id="best-practice-3-adopt-mlops-for-scalable-reproducible-ai-operations"&gt;Best Practice 3: Adopt MLOps for Scalable, Reproducible AI Operations&lt;/h2&gt;
&lt;p&gt;MLOps extends DevOps principles to machine learning, providing a structured approach to AI deployment that addresses the scalability, reproducibility, and governance challenges that manual AI operations can&amp;rsquo;t handle at scale.&lt;/p&gt;
&lt;p&gt;Five MLOps components deliver measurable operational improvements.&lt;/p&gt;
&lt;p&gt;Data engineering forms the foundation. Tools like Apache Airflow, Apache Kafka, and Apache Spark automate data collection, preprocessing, and feature engineering. Published studies indicate these practices can reduce data preparation time by up to 30% and improve data quality by 25%.&lt;/p&gt;
&lt;p&gt;Model development with version control and experiment tracking ensures reproducibility. Tools like Git, DVC (Data Version Control), and MLflow enable teams to track every experiment, reproduce results, and manage model iterations systematically. Organizations using these practices have reported a 40% reduction in time spent on experiment management. Currently, 89% of organizations use version control for ML models, leading to a 41% improvement in model reproducibility.&lt;/p&gt;
&lt;p&gt;CI/CD pipelines automate model testing and deployment. Automated pipelines continuously check model accuracy, latency, resource usage, and data drift on each deployment, with thresholds and alerts. Published data suggests CI/CD implementation can reduce deployment time by up to 70% and decrease production errors by 60%.&lt;/p&gt;
&lt;p&gt;Model serving and monitoring maintains production performance. Efficient serving infrastructure (Kubernetes, TensorFlow Serving) and continuous monitoring tools (Prometheus, Grafana) detect degradation early. Published studies indicate robust monitoring can reduce model performance degradation by up to 35% and improve mean time to resolution by 50%.&lt;/p&gt;
&lt;p&gt;Governance and security integration builds compliance into the pipeline. Regulatory compliance checks, model security against adversarial attacks, and bias monitoring run as automated steps in the deployment process rather than as manual reviews after the fact. Organizations report a 45% reduction in compliance-related incidents and a 30% improvement in model robustness from these practices.&lt;/p&gt;
&lt;p&gt;Implementation tip: Start MLOps adoption with version control for models, data, and configurations. This single practice, which costs minimal effort to implement, addresses the reproducibility crisis that undermines trust in AI systems. When a model in production behaves differently than expected, version control enables the team to identify exactly which model version is running, which data it was trained on, which configuration produced it, and what changed between the current and previous versions. Without version control, diagnosis relies on individual memory and informal records, which degrade rapidly as time passes and team members change. Version control is the foundation upon which every other MLOps practice builds.&lt;/p&gt;
&lt;h2 id="best-practice-4-build-modular-pipelines-with-automated-testing"&gt;Best Practice 4: Build Modular Pipelines With Automated Testing&lt;/h2&gt;
&lt;p&gt;Two MLOps practices deserve individual attention because they produce the largest operational impact: modular pipeline design and automated testing.&lt;/p&gt;
&lt;p&gt;Modular pipelines decompose the AI workflow into independent, reusable components: data ingestion, preprocessing, feature engineering, model training, validation, deployment, and monitoring. Each module can be developed, tested, updated, and debugged independently. Organizations using modular pipelines have reported a 28% reduction in model deployment time, improved collaboration across teams, and a 45% decrease in code duplication.&lt;/p&gt;
&lt;p&gt;Modularity also enables component-level reuse across projects. A data quality validation module built for one AI system can serve every subsequent system that uses similar data types. This compounding value accelerates each successive AI project.&lt;/p&gt;
&lt;p&gt;Automated testing extends beyond traditional software testing to include data validation, model performance testing, fairness testing, and drift detection. Comprehensive automated testing has been shown to reduce production incidents by 37% and detect data drift issues before they impact model performance.&lt;/p&gt;
&lt;p&gt;What to automate: Data integrity tests verify that incoming data matches expected schemas, ranges, and distributions. Model performance tests run the model against a standard validation dataset after every update and compare results against acceptance thresholds. Fairness tests compute demographic performance metrics and flag disparities exceeding defined limits. Integration tests verify that model outputs flow correctly to downstream systems. These tests should run automatically in the CI/CD pipeline, blocking deployment when any test fails.&lt;/p&gt;
&lt;p&gt;Implementation tip: The testing practice with the highest return is automated data validation at pipeline ingestion. Most AI production failures originate from data problems, not model problems: unexpected null values, changed field formats, shifted distributions, and corrupted data feeds. An automated data validation step that runs before every model training and inference cycle catches these problems at their source. Build validation rules for every input field: acceptable ranges, expected data types, maximum null rates, and distribution similarity to training data. When any rule is violated, the pipeline pauses and alerts the data engineering team. This single control prevents the cascade where bad data produces bad predictions that produce bad business decisions before anyone notices the data quality degradation.&lt;/p&gt;
&lt;h2 id="best-practice-5-manage-third-party-and-embedded-ai-rigorously"&gt;Best Practice 5: Manage Third-Party and Embedded AI Rigorously&lt;/h2&gt;
&lt;p&gt;Most organizations acquire more AI capabilities than they build. AI is embedded in vendor software ranging from procurement platforms to human resources systems, CRM tools, and enterprise resource planning systems. Each embedded AI component carries risks that the organization remains accountable for regardless of who built it.&lt;/p&gt;
&lt;p&gt;Third-party AI management requires four disciplines.&lt;/p&gt;
&lt;p&gt;Due diligence on vendor development practices and training data. Before procurement, evaluate the vendor&amp;rsquo;s model development methodology, training data provenance, bias testing practices, and performance validation approach. Request model cards or equivalent documentation for every AI component embedded in vendor software.&lt;/p&gt;
&lt;p&gt;Contractual provisions for transparency, liability allocation, and update notifications. Contracts should specify the vendor&amp;rsquo;s obligations regarding performance metrics, fairness standards, explainability requirements, drift management, and change notification procedures. Liability for AI-related harms should be explicitly allocated, and vendor obligations should include regular compliance audits.&lt;/p&gt;
&lt;p&gt;Monitoring vendor systems post-deployment for drift or changes. Vendor AI components change when the vendor retrains models or updates algorithms, often without customer notification. Build independent monitoring that tracks vendor AI performance on your data and your use case, detecting degradation regardless of whether the vendor reports it.&lt;/p&gt;
&lt;p&gt;Exit strategies addressing data portability. Before signing a contract, understand what happens to your data, your configurations, and any custom model components if the relationship ends. Data portability terms negotiated before commitment are always more favorable than those negotiated during exit.&lt;/p&gt;
&lt;p&gt;Shadow AI requires specific attention. When employees adopt AI tools outside formal channels, using personal ChatGPT accounts for work tasks, connecting unauthorized AI plugins to enterprise systems, or using AI-powered browser extensions that process company data, they create unmanaged risk. Detection mechanisms, clear acceptable use policies, and approved alternatives that meet security requirements address shadow AI more effectively than prohibition alone.&lt;/p&gt;
&lt;p&gt;Implementation tip: Build a third-party AI inventory that catalogs every vendor AI component operating in your environment, including AI embedded in SaaS platforms that may not be marketed as &amp;ldquo;AI products.&amp;rdquo; Many organizations discover during their first inventory that they have 3-5 times more third-party AI components than they knew about, because AI features were added to existing vendor products through routine software updates. Review the release notes and feature updates from your top 20 software vendors for the past 18 months. Many will have added AI-powered features (smart recommendations, automated classification, predictive analytics, chatbot capabilities) without prominently labeling them as AI. Each of these features is a third-party AI component that should be governed accordingly.&lt;/p&gt;
&lt;h2 id="best-practice-6-adopt-phased-implementation-with-clear-metrics"&gt;Best Practice 6: Adopt Phased Implementation With Clear Metrics&lt;/h2&gt;
&lt;p&gt;Successful AI adoption follows a staged approach rather than attempting comprehensive deployment at once. Three phases build capability and confidence progressively.&lt;/p&gt;
&lt;p&gt;Phase 1 automates repetitive administrative work to build trust and demonstrate quick wins. Targets include data entry automation, report generation, document processing, and routine classification tasks. These applications have well-defined inputs and outputs, clear success metrics, and low risk if they underperform. Success in Phase 1 generates the organizational support needed for more ambitious deployments.&lt;/p&gt;
&lt;p&gt;Phase 2 adds predictive analytics for decision support, using historical data to forecast trends, identify risks, and optimize resource allocation. This phase introduces AI into decision-making processes but maintains human judgment as the final authority. Success metrics shift from efficiency (time saved) to effectiveness (prediction accuracy, forecast reliability, decision quality improvement).&lt;/p&gt;
&lt;p&gt;Phase 3 deploys AI-powered optimization with intelligent matching, automated responses, and autonomous decision-making for appropriate use cases. This phase requires the most robust governance, monitoring, and human oversight mechanisms because the AI system is taking or heavily influencing consequential actions.&lt;/p&gt;
&lt;p&gt;Each phase should have defined success metrics measured against baselines established before deployment: time saved on reporting, improved forecast accuracy, reduced administrative burden, error reduction, or customer satisfaction improvement.&lt;/p&gt;
&lt;p&gt;Implementation tip: Define the metrics for each phase before beginning the phase, and measure against a baseline established from the current manual or non-AI process. Without a baseline, improvement claims are unverifiable. &amp;ldquo;The AI system processes documents in 3 minutes&amp;rdquo; sounds impressive until you learn that the manual process took 4 minutes. The improvement is real but marginal. Baselines enable honest ROI calculation: &amp;ldquo;The AI system processes documents in 3 minutes versus the manual process average of 47 minutes, representing a 94% reduction in processing time across approximately 400 documents per month, saving an estimated 293 hours monthly.&amp;rdquo; This specificity supports investment decisions, demonstrates value to stakeholders, and provides the evidence base for scaling to subsequent phases.&lt;/p&gt;
&lt;h2 id="best-practice-7-integrate-human-oversight-and-escalation-pathways"&gt;Best Practice 7: Integrate Human Oversight and Escalation Pathways&lt;/h2&gt;
&lt;p&gt;Despite AI&amp;rsquo;s capabilities, human judgment remains critical for high-risk decisions. Best practice requires documented human oversight mechanisms with defined triggers and response procedures.&lt;/p&gt;
&lt;p&gt;Human-in-the-loop processes ensure that consequential decisions receive human review before action. The design of human oversight matters as much as its presence. If the human reviewer sees the AI&amp;rsquo;s recommendation before reviewing the case independently, automation bias may cause them to defer to the AI even when their own judgment disagrees. If the reviewer is presented with the case facts first and asked for their independent assessment before seeing the AI recommendation, the oversight is more genuine.&lt;/p&gt;
&lt;p&gt;Escalation pathways define what happens when problems are discovered. When bias is detected, who gets notified, within what timeframe, and with what authority to act? When the model produces unexpected outputs, who investigates, and what actions can they take (pause the system, retrain the model, roll back to a previous version, shut down)? When a user reports that the AI system produced a harmful output, what&amp;rsquo;s the response procedure?&lt;/p&gt;
&lt;p&gt;These pathways should be documented before deployment, tested through tabletop exercises, and verified through periodic review of escalation logs.&lt;/p&gt;
&lt;p&gt;Implementation tip: Measure the actual override rate for human-in-the-loop processes. If the AI makes 10,000 recommendations per month and human reviewers override 12 of them (0.12% override rate), the human oversight may be functionally nonexistent. Reviewers may be rubber-stamping AI outputs because of time pressure, automation bias, or insufficient training. Published research consistently shows that human oversight degrades when reviewers process high volumes of AI outputs without adequate time, training, or incentive to exercise independent judgment. If your override rate is below 2-3%, investigate whether the low rate reflects genuine agreement (the AI is consistently correct) or passive acceptance (reviewers aren&amp;rsquo;t actively evaluating). Analyze override patterns: do overrides come from specific reviewers while others never override? Does the override rate vary with workload? These patterns distinguish active oversight from passive compliance.&lt;/p&gt;
&lt;h2 id="best-practice-8-monitor-continuously-and-plan-for-change"&gt;Best Practice 8: Monitor Continuously and Plan for Change&lt;/h2&gt;
&lt;p&gt;AI systems require ongoing monitoring because model performance degrades as real-world conditions change. Four types of drift require continuous surveillance.&lt;/p&gt;
&lt;p&gt;Data drift occurs when the statistical properties of production inputs diverge from training data. The model receives inputs it wasn&amp;rsquo;t trained to handle.&lt;/p&gt;
&lt;p&gt;Concept drift occurs when the relationship between inputs and outcomes changes. What predicted customer churn in 2023 may not predict it in 2026 because customer behavior has evolved.&lt;/p&gt;
&lt;p&gt;Model drift occurs when the model&amp;rsquo;s predictions shift over time even without changes to the model itself, typically as a consequence of data drift or concept drift.&lt;/p&gt;
&lt;p&gt;Performance degradation occurs when accuracy, fairness, or other performance metrics decline below acceptable thresholds.&lt;/p&gt;
&lt;p&gt;When monitoring identifies issues, organizations need documented retraining and update procedures that include re-validation before deployment. This ensures that changes don&amp;rsquo;t introduce new risks. The monitoring system should include defined thresholds for investigation, retraining, rollback, and retirement, with each threshold triggering a specific response procedure.&lt;/p&gt;
&lt;p&gt;Cloud-native deployment enables dynamic scaling of monitoring and retraining operations. Published data indicates that cloud-native solutions have led to a 62% improvement in model training speed and an average cost reduction of 35% in ML infrastructure expenses.&lt;/p&gt;
&lt;p&gt;Implementation tip: Build your monitoring system to detect problems in hours, not weeks. The most expensive monitoring failures are the slow ones, where performance degrades gradually over days or weeks without triggering any alert because each daily change is individually minor. Configure your monitoring to detect trends, not just threshold breaches. A model that drops 0.3 percentage points of accuracy per day doesn&amp;rsquo;t breach a 5-point accuracy threshold for 16 days. Trend detection that flags sustained directional movement over 5-7 days catches the same problem in one-third the time. Trend-based alerts supplement threshold-based alerts and catch the gradual degradation that threshold alerts miss.&lt;/p&gt;
&lt;h2 id="best-practice-9-build-ai-literacy-across-the-organization"&gt;Best Practice 9: Build AI Literacy Across the Organization&lt;/h2&gt;
&lt;p&gt;Effective AI governance depends on shared understanding across roles. Technical teams can&amp;rsquo;t govern AI systems alone because they lack regulatory and business context. Business teams can&amp;rsquo;t govern AI systems alone because they lack technical understanding. Governance requires both perspectives working together, which requires minimum AI literacy across the organization.&lt;/p&gt;
&lt;p&gt;Four audience-specific literacy programs address different needs.&lt;/p&gt;
&lt;p&gt;Executives need to understand strategic AI risk: what can go wrong at the organizational level, what the regulatory exposure looks like, and how to evaluate whether AI investments are delivering value.&lt;/p&gt;
&lt;p&gt;Business managers need to understand how to propose use cases responsibly, how to evaluate whether AI is the right tool for a specific problem, and how to set realistic expectations for AI capabilities.&lt;/p&gt;
&lt;p&gt;Operational staff need to understand how to interact with AI systems correctly, when to trust AI outputs, when to override them, and how to provide feedback that improves system performance.&lt;/p&gt;
&lt;p&gt;Technical teams need to understand governance requirements, regulatory constraints, and ethical considerations that affect model design, testing, and deployment decisions. Technical excellence without governance understanding produces systems that work technically but fail regulatory or ethical standards.&lt;/p&gt;
&lt;p&gt;Published data indicates that organizations considering ethical AI as a critical component of their AI operations increased from 54% in 2021 to 82% in 2023. Bias monitoring tools have led to a 39% reduction in biased outcomes in organizations that deploy them. These improvements require organizational literacy to sustain because tools alone don&amp;rsquo;t create responsible AI culture.&lt;/p&gt;
&lt;p&gt;Implementation tip: The most effective AI literacy investment is cross-functional workshop sessions where technical and business teams work through real scenarios together. A workshop where a data scientist explains a model card to a compliance officer, who then explains a regulatory requirement to the data scientist, produces more practical understanding than either person attending a separate training course. These workshops reveal the translation gaps between technical and business language that cause miscommunication in daily operations. Schedule quarterly cross-functional workshops covering a current AI system, its performance data, its governance documentation, and a hypothetical incident scenario. The shared experience of working through these materials together builds the mutual understanding that individual training cannot replicate.&lt;/p&gt;
&lt;h2 id="best-practice-10-measure-value-not-just-compliance"&gt;Best Practice 10: Measure Value, Not Just Compliance&lt;/h2&gt;
&lt;p&gt;While risk management is critical, successful AI programs also measure business value. A governance framework that prevents every possible risk but blocks every possible value creation isn&amp;rsquo;t serving the organization. Balance requires measuring both dimensions.&lt;/p&gt;
&lt;p&gt;Project managers should define success metrics that include both technical performance and business outcomes.&lt;/p&gt;
&lt;p&gt;Technical metrics include accuracy, precision, recall, F1-score, latency, throughput, and resource utilization. These metrics confirm that the AI system functions correctly.&lt;/p&gt;
&lt;p&gt;Business metrics include efficiency gains (time saved, manual effort reduced), revenue impact (increased conversion, reduced churn, optimized pricing), cost reduction (lower processing costs, reduced error remediation), and customer satisfaction (NPS improvement, resolution time reduction, service quality). These metrics confirm that the AI system creates value.&lt;/p&gt;
&lt;p&gt;Organizations implementing MLOps practices have reduced model deployment time by an average of 63%, from 45 days to 17 days. AI technologies, enabled by effective operations practices, could boost labor productivity by 0.8% to 1.4% annually through 2030. These gains materialize only when organizations measure and optimize for business outcomes alongside technical performance.&lt;/p&gt;
&lt;p&gt;Regularly evaluate the ROI of AI projects to guide future investments and technology decisions. A project that delivers strong technical performance but negative ROI may need scope adjustment, cost optimization, or retirement. A project that delivers modest technical performance but strong ROI may deserve additional investment to improve its technical foundation.&lt;/p&gt;
&lt;p&gt;Implementation tip: Create a balanced scorecard for each AI system that tracks four quadrants: technical performance (model accuracy, latency, reliability), business impact (ROI, efficiency gains, revenue contribution), risk and compliance (bias metrics, regulatory compliance, incident rates), and user adoption (adoption rate, satisfaction scores, override rates). Review all four quadrants quarterly. A system that scores well in three quadrants but poorly in one has a specific, identifiable problem to address. A system that scores well in technical performance and compliance but poorly in business impact and user adoption is a well-governed system that nobody uses, which means it&amp;rsquo;s not delivering value. The balanced view prevents the common pattern where technical teams celebrate model performance while business outcomes go unmeasured.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;&lt;img src="https://hernanhuwyler.wordpress.com/wp-content/uploads/2026/03/modern-professional-in-a-sunny-co-working-space.png?w=1024" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 id="implementation-tips-for-ai-project-management"&gt;Implementation Tips for AI Project Management&lt;/h2&gt;
&lt;p&gt;These principles apply across all ten best practices.&lt;/p&gt;
&lt;p&gt;Implementation tip on the prototype-to-production transition: The GreatAI framework, developed through design science research and evaluated with practitioners, identifies 33 specific best practices for transitioning AI from prototype to production. The research found that both ease of use and functionality are crucial factors for adopting deployment technologies. The most common failure point isn&amp;rsquo;t building a working prototype. It&amp;rsquo;s converting that prototype into a production system with proper data pipelines, monitoring, error handling, versioning, and governance. Budget the prototype-to-production transition as a separate project phase with its own timeline, resources, and success criteria. Teams that treat deployment as a simple step after development consistently underestimate the effort required.&lt;/p&gt;
&lt;p&gt;Implementation tip on managing stakeholder expectations: AI projects have a unique expectation management challenge because stakeholders often have inflated expectations about AI capabilities drawn from media coverage and vendor marketing. Set expectations during the planning phase using concrete examples from comparable deployments, not abstract capability descriptions. &amp;ldquo;Our customer churn model is expected to identify 75-85% of customers likely to leave within 30 days, based on results from similar models in our industry&amp;rdquo; is a manageable expectation. &amp;ldquo;AI will predict customer churn&amp;rdquo; invites the assumption that the model will identify 100% of churning customers with certainty. The specificity of the first statement protects both the team and the stakeholder from the disappointment that vague promises create.&lt;/p&gt;
&lt;p&gt;Implementation tip on documentation as a project deliverable: Treat documentation (model cards, risk assessments, compliance records, governance approvals) as project deliverables with the same status as code and model artifacts. Documentation completed as an afterthought after deployment is consistently lower quality than documentation completed as each phase concludes. Include documentation deliverables in your project plan with specific owners and due dates. Review documentation quality at each governance gate. A model that passes technical validation but lacks complete documentation should not proceed to deployment.&lt;/p&gt;
&lt;p&gt;Implementation tip on the relationship between AI project management and organizational change: Every AI deployment changes how people work. Processes that were manual become automated. Decisions that were intuitive become data-driven. Roles that centered on data gathering shift toward analysis and judgment. These changes require active management. Published data on AI implementation consistently shows that the most common deployment failure mode isn&amp;rsquo;t technical. It&amp;rsquo;s adoption. Users who don&amp;rsquo;t trust, understand, or know how to use the AI system revert to previous methods. Dedicate project management attention and budget to change management activities: user training, workflow redesign, communication, and adoption monitoring. Treat adoption rate as a first-class success metric alongside technical performance metrics.&lt;/p&gt;
&lt;h2 id="key-references-and-authoritative-frameworks"&gt;Key References and Authoritative Frameworks&lt;/h2&gt;
&lt;p&gt;Your AI project management practices should align with these established standards:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 42001:2023, AI Management System (governance, lifecycle, and performance evaluation)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;NIST AI Risk Management Framework 1.0, Govern-Map-Measure-Manage functions&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;IIA AI Auditing Framework and 2024 IIA Standards&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;EU AI Act (risk classification, compliance requirements, documentation obligations)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 5338, AI System Life Cycle Processes&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ISO/IEC 23894:2023, AI Risk Management&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;MLOps frameworks and practices for automation, monitoring, reproducibility, and governance&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;GreatAI and related deployment best-practice frameworks focused on prototype-to-production transition&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Internal PMO, change management, architecture review, security review, and product governance standards&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;ETSI TS 104 008, Continuous Auditing-Based Conformity Assessment&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;MLOps maturity model frameworks from Google, Microsoft, and AWS&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;GreatAI Framework for prototype-to-production best practices (Visser, 2023)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;MLOps integration research (Sachdeva, 2024; Kabbay, 2024)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;PMBOK Guide adapted for AI project lifecycle management&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;COBIT 2019 for IT governance of AI initiatives&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you manage AI projects using traditional software development practices, treating model development as deterministic, deployment as a one-time event, and post-deployment monitoring as optional, you will produce systems that work in testing environments and degrade in production. The model will drift without detection. The governance will exist without function. The business case will remain unverified because nobody measured the outcomes. And each failed project will make the next one harder to fund because the organization will have learned to distrust AI promises without learning the management practices that make AI promises deliverable.&lt;/p&gt;
&lt;p&gt;When you apply AI-specific project management practices, building governance structures with real authority, implementing MLOps for reproducibility and scale, managing the AI lifecycle as a continuous process rather than a one-time project, integrating human oversight that functions rather than merely exists, and measuring business value alongside technical performance, you create the conditions for AI projects to deliver sustained value. The model gets built with proper validation. It gets deployed with proper monitoring. It gets maintained with proper governance. And it gets measured against the business outcomes that justified its creation.&lt;/p&gt;
&lt;p&gt;An AI project managed like a software project is a project managed for its first 30 days. An AI project managed for its full lifecycle is a project managed for its full value.&lt;/p&gt;
&lt;p&gt;Which of these ten best practices is weakest in your current AI project management approach? Strengthen that practice before your next AI initiative kicks off.&lt;/p&gt;
&lt;h2 id="about-the-author"&gt;About the Author&lt;/h2&gt;
&lt;p&gt;The frameworks, tools, and implementation guidance described in this article are part of the applied research and consulting work of Prof. Hernan Huwyler, MBA, CPA, CAIO. These materials are freely available for use, adaptation, and redistribution in your own AI governance, risk management, and compliance programs. If you find them valuable, the only ask is proper attribution.&lt;/p&gt;
&lt;p&gt;Prof. Huwyler serves as AI GRC Consultancy Director, AI Risk Manager, and Quantitative Risk Lead, working with organizations across financial services, technology, healthcare, and public sector to build practical AI governance frameworks that survive contact with production systems and regulatory scrutiny. His work bridges the gap between academic AI risk theory and the operational controls that organizations actually need to deploy AI responsibly.&lt;/p&gt;
&lt;p&gt;As a Speaker, Corporate Trainer, and Executive Advisor, he delivers programs on AI compliance, quantitative risk modeling, predictive risk automation, and AI audit readiness for executive leadership teams, boards, and technical practitioners. His teaching and advisory work spans IE Law School Executive Education and corporate engagements across Europe.&lt;/p&gt;
&lt;p&gt;Based in the Copenhagen Metropolitan Area, Denmark, with professional presence in Zurich and Geneva, Switzerland, Madrid, Spain, and Berlin, Germany, Prof. Huwyler works across jurisdictions where AI regulation is most active and where organizations face the most complex compliance landscapes.&lt;/p&gt;
&lt;p&gt;His code repositories, risk model templates, and Python-based tools for AI governance are publicly available at 
. His ongoing writing on Governance, Risk Management and Compliance appears on his blogger website at 
(more than 500k views).&lt;/p&gt;
&lt;p&gt;Connect with Prof. Huwyler on LinkedIn at 
 to follow his latest work on AI risk assessment frameworks, compliance automation, model validation practices, and the evolving regulatory landscape for artificial intelligence.&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re building an AI governance program, standing up an AI risk function, preparing for EU AI Act compliance, or looking for practical implementation guidance that goes beyond policy documents, reach out. The best conversations start with a shared problem and a willingness to solve it with rigor.&lt;/p&gt;</description></item></channel></rss>