How ISO 24970 and prEN 18229-1 Turn Post-Deployment Chaos Into Auditable Evidence
When AI Systems Fail, Logs Tell the Story
Your AI system just flagged 300 legitimate transactions as fraud. A biometric authentication tool locked out half your workforce. A content moderation model started removing benign posts at twice the normal rate. In each case, the first question from your board, your regulator, or your customer is the same: what happened?
Without structured logs, you have no answer. Without a logging framework that captures the right events at the right resolution, you cannot reconstruct the failure, validate your risk controls, or prove you met your oversight obligations. This is the operational gap that ISO 24970 and the European pre-draft standard prEN 18229-1 were built to close.

Why AI Logging Is Different From Application Logging
AI systems generate decisions under uncertainty. A traditional application either executes correctly or throws an error. An AI model can produce a technically valid output that is still wrong, biased, unsafe, or out of scope. The system can drift over time as input distributions shift, adversarial patterns emerge, or model retraining introduces new failure modes.
Standard application logs capture exceptions and transactions. AI logs must capture context, decisions, inputs, outputs, model state, human interventions, and the conditions under which the system operated. They must support not only debugging but also compliance, human oversight, risk detection, bias monitoring, and post-market surveillance.
Designing AI logging architectures based solely on deterministic software practices guarantees blind spots. If your infrastructure fails to capture the exact input distribution and model version during an anomalous inference, you cannot reconstruct the failure or quantify the resulting model risk.
The challenge is that you cannot predict in advance which events will matter. A logged input that seems routine today can become the key evidence in a discrimination claim six months from now. A pattern of outlier detections that you ignored can signal the onset of adversarial attack or domain drift. Logging for AI is not just instrumentation. It is a form of institutional memory that lets you reconstruct what the system knew, what it decided, and what humans did or did not do in response.
Relevance is also not static. As the system interacts with users, encounters new data, or gets deployed in new contexts, the events worth logging can change. Some systems can adapt their logging behavior automatically. Others require human reconfiguration. The standards do not mandate one approach, but they do require that you document your triggers, justify your event selection, and ensure that your logs remain usable across the system lifecycle.
The operational payoff is clear. Logs support monitoring, troubleshooting, strategic planning, and continuous improvement. They feed risk management processes, inform retraining decisions, and provide the evidence base for regulatory filings. But the value depends entirely on log quality, governance, access controls, and the organizational capacity to interpret and act on the data. A poorly designed logging system creates compliance theater. A well-designed one turns operational telemetry into decision support and legal protection.

The Regulatory Context: ISO 24970 and prEN 18229-1
ISO 24970 is the international standard for AI system logging. It defines what to log, when to log it, how to structure log entries, and how to manage log storage and access. The standard is technically precise, format-agnostic, and applicable across sectors and jurisdictions.
prEN 18229-1 is the European counterpart, currently in pre-draft status under CEN-CENELEC Joint Technical Committee 21. It embeds logging into a broader trustworthiness framework that also covers transparency and human oversight. The standard is being developed to support compliance with the EU AI Act, particularly the logging obligations in Article 12 for high-risk systems and the enhanced requirements in Article 14 for remote biometric identification.
The two standards overlap heavily on technical content. Both require event-based logging, traceability through timestamps and identifiers, risk-driven event selection, and governance controls on access and retention. Both treat logs as evidence that must survive audits, support post-market monitoring, and enable deployer oversight.
The main difference is scope and regulatory intent. ISO 24970 is a general-purpose technical foundation. prEN 18229-1 wraps that foundation in a compliance layer designed for EU AI Act obligations, including explicit ties to legal requirements for transparency, human oversight, and post-market surveillance. For organizations deploying high-risk AI in Europe, prEN 18229-1 translates ISO 24970 into a regulatory checklist.
Because prEN 18229-1 is still in pre-draft status, the text is subject to change. The current draft is under enquiry within the European standardization process. It references Directive 2024/1689 (the AI Act) and is expected to be cited in the Official Journal of the European Union once finalized. Organizations building logging systems today should track both standards and design for convergence.
The practical approach is to start with ISO 24970 to define your logging architecture, then map those logs to the compliance and oversight requirements in prEN 18229-1. For high-risk systems, that means aligning your event triggers, log content, retention policies, and access controls with the AI Act from the beginning. Retrofitting logging after deployment is expensive and often incomplete.
EU AI Act Requirements for High-Risk Systems
Article 12 of the EU AI Act mandates automatic logging capabilities for all high-risk AI systems. The logs must capture events that indicate emerging risks or significant modifications to the system under Article 79. They must enable post-market monitoring under Article 72. They must support deployer oversight under Article 26.
Article 12: Record-Keeping: High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system. In order to ensure a level of traceability of the functioning of a high-risk AI system that is appropriate to the intended purpose of the system, logging capabilities shall enable the recording of events relevant for: (a) identifying situations that may result in the high-risk AI system presenting a risk within the meaning of Article 79(1) or in a substantial modification; (b) facilitating the post-market monitoring referred to in Article 72; and (c) monitoring the operation of high-risk AI systems referred to in Article 26(5). For high-risk AI systems referred to in point 1 (a), of Annex III, the logging capabilities shall provide, at a minimum: (a) recording of the period of each use of the system (start date and time and end date and time of each use); (b) the reference database against which input data has been checked by the system; (c) the input data for which the search has led to a match;(d) the identification of the natural persons involved in the verification of the results, as referred to in Article 14(5).
For remote biometric identification systems listed in Annex III, the logging requirements are more specific. You must log precise start and end timestamps for each usage session. You must record the reference database used during input validation. You must log input data that triggered search matches. You must identify the individuals responsible for verifying results, as required by Article 14.
These are not optional features. They are legal obligations. Failure to implement automatic logging, retain the required data, or make logs available to competent authorities can trigger enforcement action, including fines up to 3 percent of global annual turnover for severe violations.
The standards give you the technical blueprint to meet these obligations. But compliance also depends on governance. You need documented policies on what to log, how long to retain it, who can access it, and how to respond when logs reveal risks. You need processes to review logs, escalate anomalies, and update the system when logging reveals gaps or failures. And you need technical controls to prevent log tampering, ensure log integrity, and protect log confidentiality.
| Framework Feature | ISO/IEC 24970 | prEN 18229-1 (Pre-Draft) |
|---|---|---|
| Primary Scope | Global technical logging mechanism | European trustworthiness and compliance |
| Target Application | General AI system architecture | High-risk AI systems (EU AI Act) |
| Key Directives | Event triggers, data models, traceability | Post-market monitoring, deployer oversight |
| Operational Focus | Diagnostic telemetry and error handling | Legal accountability and transparency |
Core Concepts: Logs, Log Entries, and Logging Components
A log is a structured repository of log entries. Each log entry is a discrete record that captures a specific event, condition, state, input, output, or decision related to the AI system. Logging is the process of generating, capturing, and managing those entries.
A model is a representation of a system, entity, or process, whether physical, mathematical, or logical. In AI, the model is typically the trained artifact that produces predictions or decisions. But the AI system is larger than the model. It includes data pipelines, serving infrastructure, user interfaces, monitoring tools, and external integrations.
An audit is a systematic, independent process for obtaining and evaluating objective evidence to determine whether audit criteria are met. Internal audits are conducted by the organization. External audits are conducted by customers, regulators, or third-party certification bodies.
Auditability is the capability to collect and make available the evidence needed to conduct an audit. For AI systems, that evidence lives in logs. Without logs, you cannot prove what the system did, when it did it, or under what conditions.
An error is a discrepancy between a computed value and the true or specified value. Errors can be caused by component failures or by the activation of latent faults. In AI, errors also include incorrect predictions, misclassifications, or outputs that violate safety or fairness constraints.
Monitoring is the ongoing observation and assessment of system behavior, outputs, and context. Monitoring can be automated or manual. It detects deviations from expected operation, such as failures, malfunctions, cyberattacks, out-of-domain inputs, or abnormal usage.
A logging component is the part of the AI system, or a linked external system, that enables logging. It can consist of multiple subcomponents that generate, format, filter, or forward log entries. The logging component can be implemented in software, hardware, or a hybrid configuration.
A log user is the organization or entity that accesses, reviews, or analyzes logs. Log users include developers, testers, operators, auditors, deployers, and regulators. Each has different access rights and different purposes.
De-identification is the process of removing or altering data so that individuals or entities cannot be identified, directly or indirectly. De-identification is often required to comply with privacy regulations or to share logs with third parties.
A data principal is the entity to which data relates. This includes persons, organizations, devices, or software applications. The term is broader than personally identifiable information principal or data subject.
An organization is a person or group with its own functions, responsibilities, and objectives. This includes companies, government agencies, nonprofits, and partnerships.
An AI user is the organization or entity that uses AI products or services. A stakeholder is anyone who can affect, be affected by, or perceive themselves to be affected by the AI system. An AI developer is the organization involved in development.
Memory capacity is the maximum number of items that can be held in the logging component’s volatile memory, typically measured in bytes. Storage capacity is the maximum number of items that can be held in persistent storage.
A software error is an erroneous result produced by the use of a software product. This includes incorrect outputs, exceptions, or failures to execute.
A controller is an authorized human or external agent that performs control actions on the AI system. A control point is the part of the system interface where control can be applied, such as a function, switch, or signal receiver.
Control engagement is the process where a controller takes over control points. Control disengagement is when a controller releases control points. Control transfer is the handover of control points from one controller to another.
A governance scheme is the set of rules that defines how the system is managed and controlled. This can be a regulation, standard, guideline, convention, or social norm.
An AI provider is the organization that provides products or services using one or more AI systems.
These definitions matter because they set the boundaries of what must be logged, who has access, and what counts as evidence. If your logging system does not distinguish between a software error and a model prediction error, you cannot diagnose failures. If your logs do not capture control transfers, you cannot prove human oversight. If you do not de-identify logs before sharing them, you violate privacy law.
What Goes Into an AI System Log
An AI system log captures information related to operation, behavior, inputs, outputs, or context. The log can contain structured data like JSON objects, semi-structured data like annotated text, or unstructured data like screenshots. Logs can originate from the AI system itself, its internal components, interacting systems, users, or external observers.
Logs can be generated continuously, periodically, or in response to specific conditions. They serve multiple purposes including monitoring, debugging, auditing, compliance, human oversight, iterative improvement, and accountability.
AI system logs can include time-stamped events, which are recorded occurrences linked to a specific moment. Examples include when a model generates a prediction, an error occurs, or a user interaction takes place. They can include status snapshots, which are point-in-time captures of system conditions such as memory usage, model state, or active components.
Logs can include sensor or input data, meaning information received from external sources like camera images, user inputs, location data, or telemetry. They can include outputs such as classifications, recommendations, predictions, or generated content. They can include decisions, which are discrete choices or actions taken by the system, either autonomously or through human-in-the-loop mechanisms.
Logs can include error messages, which are alerts or diagnostic records indicating failures, exceptions, or issues. They can include environmental context such as network status, sensor readings, user load, or surrounding events. They can include annotations, which are supplementary notes or metadata added manually or automatically to describe behavior, flag anomalies, or provide interpretive context.
AI system logs can be stored persistently for long-term retention, inspection, or regulatory compliance. They can be processed in real time to support live monitoring, alerting, or adaptive behavior. They can be managed under data minimization or privacy constraints to avoid collecting unnecessary personal data, ensure user consent, or comply with legal frameworks.
Logs can be machine-readable, formatted for automated processing using standards like JSON or XML. They can be human-interpretable, presented in a way that allows developers, auditors, or analysts to understand the content without complex tooling.
Logging Components and Their Role
A logging component is the functional part of the AI system or an external system that supports the generation, capture, formatting, storage, or management of log data. It can consist of one or more subcomponents responsible for detecting events, recording log entries, applying data policies like filtering or redaction, or ensuring secure and reliable handling.
Logging components can be internal to the AI system, integrated into model-serving infrastructure or runtime environments. They can be external systems or services such as observability platforms, audit modules, or compliance loggers. They can operate independently or in coordination with other system parts. Complexity varies from a simple event logger to a distributed, multi-service logging pipeline.
The logging component does not assume a fixed structure, automation level, or deployment location. It can be implemented in software, hardware, or hybrid configurations. It is designed to meet different operational, analytical, or regulatory objectives.
The logging component and the storage used for logging are not necessarily part of the AI system itself. They can be separate infrastructure managed by third parties, provided that confidentiality, integrity, and availability are maintained according to applicable regulatory requirements.
Logging in Context: Operational and Management Integration
Management of an AI system in operation is naturally integrated with operation itself. Management and operation share the fundamental goal of navigating uncertainty to achieve purposes. This involves capitalizing on opportunities and mitigating risks through planning, monitoring, decision-making, and learning.
Components of the AI system, including monitoring systems, can use logs to better fulfill the intended purpose, including risk mitigation. A log user can collect and analyze possibly de-identified logs from multiple AI systems to create and improve AI systems.
The logging component logs behaviors of the AI system. Monitoring systems can use the logs to help the AI user assess potential benefits and harms of AI system activities. Logs can be used to assess the continuous fulfillment of various requirements such as accuracy, robustness, security, privacy, safety, and data quality. This assessment can inform the selection of actions.
Organizations can collect and analyze logs from similar AI systems or similar components on the market to support the creation, maintenance, and continuous improvement of the data, AI systems, and components they provide.
Structure and Content of Log Entries
AI system log entries are discrete, identifiable units of information within a log. Each entry captures a specific event, condition, state, input, output, decision, or contextual detail related to the functioning or environment of the AI system.
Log entries are typically composed of a combination of metadata such as timestamps, source identifiers, and severity levels, along with content-specific data relevant to the purpose of the log. Protection of confidential information must be taken into account.
Log entries can vary in structure and content depending on the type of information being recorded and the intended use of the log. Some entries are highly structured, such as a JSON object. Some are semi-structured, such as textual annotations. Some are free-form, such as screenshots.
Components of a log entry can include a timestamp, the date and time at which the logged event or condition occurred. They can include a source identifier, a label or address indicating which component, system, user, or external observer generated the entry. They can include an event or message code, a categorization or classification of the type of event such as an error event, inference event, or user override event.
They can include payload or data content, the core data being recorded such as input features, output values, error traces, or contextual metadata. They can include a severity or priority indicator, a label indicating the importance or criticality of the event, useful for filtering or alerting.
Log entries can be generated automatically by system components or instrumentation. They can be manually created by users, operators, or auditors, such as annotations or overrides. They can be derived from external systems such as monitoring tools or interacting AI components.
To be useful for downstream analysis, log entries should be recorded in a way that ensures traceability, interpretability, and data integrity over time.
The Process of AI System Logging
AI system logging is the process of generating, capturing, recording, and managing information related to the operation, behavior, decisions, or context of an AI system, for the purpose of creating one or more logs.
Logging can be performed automatically by system components, manually by users or operators, or through hybrid methods. It can occur during design, testing, deployment, or post-deployment operation.
Logging can involve the collection of data from a variety of sources including system components such as model execution, middleware, or infrastructure. It can involve user interactions such as input submissions or user overrides. It can involve external observers such as monitoring tools or regulatory systems. It can involve interacting AI systems such as decision handoffs or multi-model coordination.
Logging activities can be continuous, such as telemetry data. They can be event-driven, such as error occurrences. They can be scheduled, such as periodic health checks. They can be conditional, such as events triggered by threshold violations or policy rules.
Logging can include one or more sub-processes. Instrumentation is the implementation of tools, code, or mechanisms to monitor, extract, and capture data from software or hardware components during execution. Serialization is converting data structures or objects into a standardized format such as JSON, XML, or binary for logging, storage, or transmission.
Storage and retention is saving logs to appropriate storage systems with defined retention policies. De-identification processes are applied according to applicable legal or ethical standards. Validation and integrity checking ensures that log data is accurate, complete, and has not been tampered with.
Logging should be guided by clearly defined objectives such as performance monitoring, safety validation, auditability, transparency, compliance, user redress, or support for system improvement.
General Requirements for AI System Logging
AI system logging must provide specific capabilities. Logging functions must enable traceability between multiple events and log entries if necessary to manage risk, relevant to the intended purpose, and technically feasible given the inputs and outputs.
The organization must identify the security and privacy requirements for integrity and confidentiality protection. It must protect information taking into account different purposes of logging for different stakeholders. For example, an AI developer who implements, an AI tester who tests the system, and an AI service provider who monitors service during operation all have different purposes.
Security and privacy requirements include those based on applicable privacy and security regulatory requirements.
Logging functions must enable the recording of events relevant for identifying situations that can result in the AI system presenting a risk according to the risk management process. They must facilitate the monitoring of AI systems as a product or service, proportional to their risks, to enable collection, documentation, and analyzing performance data from initial development to the end of the retirement stage.
Technical Documentation for Logging
The technical documentation for the AI system must explain and justify the specific criteria for determining relevant events. It must explain and justify the specific criteria for logging relevant events. It must specify any interaction with human controllers. It must specify any interaction with automated monitoring.
It must recommend a frequency and scope of monitoring for relevant events. It must recommend a frequency and scope of logging relevant events. It must explain and justify the accuracy and precision of timestamps, where used.
It must explain and justify resource constraints such as memory capacity, storage capacity, and processing power. It must explain and justify constraints related to privacy. It must refer to related legal requirements related to data protection, system accountability, traceability, and transparency.
It must include appropriate information security considerations and data retention policies. It must include specification of failure handling, such as AI system reaction in case of log memory overloading. It must include interfaces with other systems. It must contain a specification of used log data structures.

AI System Logging Fields for Governance Professionals
The following table covers event types with five columns: what field to capture, what to actually record and why it matters, the governance purpose it serves, and whether it’s required, recommended, or optional, with a risk level for each.
Legends
Required, must be captured, non-negotiable
Recommended, best practice, capture when feasible
Optional, adds value in specific contexts
| Event type | Field to capture | What to record and why it matters | Governance purpose | Obligation and risk |
|---|---|---|---|---|
| Operational events, triggered by normal AI system activity | ||||
| Transaction initiation When a request enters the AI system | event_id | A globally unique ID for this specific request. Used to trace a single transaction through all downstream logs and audit trails. Without this, you cannot link what went in to what came out. | Traceability, audit | Required High risk |
| system_id | Identifies which AI system, as a governed unit, processed the request. In shared infrastructure where one platform serves multiple products, this field distinguishes which system the risk assessment applies to.e.g. “loan-approval-v2” not just “ml-cluster-3” | Accountability, risk scoping | Required High risk | |
| model_id | Full model name and version, as granular as the provider makes available. This is critical for post-incident investigation: if a model version introduced a bias or error, you need to know which transactions were affected.e.g. “gpt-4o-2024-08-06” not just “GPT-4” | Incident response, model versioning | Required High risk | |
| timestamp | Precise date and time the request was received, in a standardised format (ISO 8601). Include timezone explicitly. For systems without a real-time clock, record a relative counter (e.g. cycles since startup) to preserve ordering.e.g. “2025-11-14T09:32:11.482Z” | Sequencing, forensics | Required Medium risk | |
| input_ref | A pointer to where the full input is stored, not necessarily the input itself. Capture the source identity (which user, API endpoint, sensor, or system sent this). Enables tracing input provenance in multi-source environments. | Traceability, security audit | Recommended Medium risk | |
| input_payload | The actual content sent to the model, or a reference to retrieve it. Required when understanding the input is necessary to explain the output. For sensitive inputs, store a reference and apply access controls. Do not log raw personal data unnecessarily. | Explainability, redress | Recommended High risk | |
| metadata | Contextual parameters that shaped how the model processed the request, for example, temperature settings, prompt version, language of input, encoding type. Without these, reproducing or explaining a result is often impossible. | Reproducibility, debugging | Optional Lower risk | |
| Transaction outcome When the AI system returns a result | output_payload | The actual content returned by the model. Essential for auditing whether the AI system produced harmful, biased, or incorrect outputs. Store a reference if the payload is large or sensitive. | Accountability, bias detection | Required High risk |
| confidence_level | The model’s confidence or probability score for its output, where available. Helps identify cases where low-confidence outputs led to consequential decisions, a key signal for human review thresholds. | Risk calibration, oversight | Recommended Medium risk | |
| correlation_id | Links this outcome back to its originating transaction and any intermediate processing steps. Essential in multi-stage systems where a request passes through several components before a response is returned. | End-to-end traceability | Recommended High risk | |
| Transaction feedback When correctness of an output is established | ground_truth | The correct or intended output for a given input, provided after the fact by a human reviewer, test data, or authoritative source. Used to measure model accuracy over time and detect performance degradation. Not applicable to all AI system types. | Model performance, continuous improvement | Recommended Medium risk |
| feedback_source | Who or what provided the ground truth, including a named human reviewer, an automated test suite, or a regulatory authority. Allows weighting of feedback by source reliability and supports audit of the feedback process itself. | Accountability, audit quality | Recommended Medium risk | |
| Anormality and security events triggered by automated monitoring | ||||
| Software error When the system fails to process normally | error_code | A structured code classifying the error type, such as inference failure, timeout, component crash. Allows filtering and trending of error types across large volumes of logs without reading free-text descriptions. | Reliability monitoring, SLA | Required High risk |
| error_message | Human-readable description of what failed. Pair with error_code. Include severity level (critical / warning / informational) and the impact: did this affect the user’s outcome? Was a fallback triggered? | Incident response, debugging | Required High risk | |
| recovery_action | What the system did in response, retried, switched to fallback, notified the user, escalated to a human, or failed silently. Silent failures with no logged recovery action are a major governance gap. | Resilience, human oversight | Recommended High risk | |
| Outlier input detected When an input falls outside expected distribution | outlier_flag | Indicates the input deviated from the statistical profile of the training domain, e.g. a feature value outside established bounds. Log the specific metric or threshold that was breached, not just a boolean flag. | Risk detection, domain monitoring | Required High risk |
| Adversarial attack When a deliberate attempt to manipulate the model is detected | attack_type | Classify the detected pattern: prompt injection, model inversion attempt, data poisoning signature, unauthorised access pattern. Detection may be triggered by a single input or a pattern across multiple inputs, note which applies. | Security, incident response | Required High risk |
| detection_basis | Whether the attack was identified from a single request or inferred from a pattern across prior logged inputs. If pattern-based, reference the window of prior log entries that contributed to detection. | Forensics, alert calibration | Recommended High risk | |
| Bias detected When outputs show unwanted differential treatment | bias_indicator | The specific metric that triggered the alert, e.g. demographic parity gap, equalized odds differential, disparate error rates across groups. Log the measured value alongside the threshold that defines “unwanted” for this system. | Fairness, regulatory compliance | Required High risk |
| affected_population | Which groups or segments were identified as affected by the differential output. Required for meaningful impact assessment. Handle with care, this field may itself contain sensitive information requiring access controls. | Impact assessment, remediation | Recommended High risk | |
| Out-of-domain input When the system is used outside its intended scope | domain_violation | Describes how the input fell outside the operational domain, such as violated a feature boundary, represented an unseen data distribution, or triggered domain drift detection across recent inputs. Distinguish single-input violations from distributional drift. | Scope compliance, safety | Required High risk |
| Model drift When model behaviour shifts from its validated baseline | drift_metric | The performance or distributional metric that revealed drift, such as prediction shift, output distribution change, increasing error rate. Log the measured value and the baseline it is compared against. Only applicable to systems with updatable models. | Model governance, revalidation | Required High risk |
| drift_window | The time period or number of transactions over which drift was measured. Without this, a drift alert cannot be investigated, you need to know which inputs to review. | Forensics, retraining triggers | Recommended Medium risk | |
| Human oversight events triggered by human controllers acting on the system | ||||
| Human intervention When a person stops, overrides, or corrects the AI system | controller_id | Unique identifier of the person who intervened not a role or team name, but a specific individual. This is non-negotiable for accountability: if a human altered an AI decision, there must be a named person in the log. | Accountability, audit | Required High risk |
| intervention_reason | Why the person intervened, prevented a serious incident, corrected an erroneous output, responded to a user complaint. Record based on risk: for high-risk systems, the reason is always required. For lower-risk systems, assess and justify. | Accountability, learning | Recommended High risk | |
| intervention_outcome | What the intervention achieve, the AI output was blocked, modified, approved, or escalated. Captures the difference between what the AI system would have done and what was actually delivered to the user. | Oversight effectiveness | Recommended High risk | |
| Control transfer When operational control of the AI system changes hands | from_controller | The controller relinquishing control, who held authority before the transfer. Without this, you cannot reconstruct accountability chains for decisions made during the transition period. | Chain of custody, accountability | Required High risk |
| to_controller | The controller taking over. Log both the engagement (new controller accepts) and the disengagement (previous controller releases) as separate timestamped entries to capture the full handover. | Chain of custody | Required High risk | |
| Output validation When a human checks or approves an AI output | validator_id | Who performed the check. Distinct from the person who made the downstream decision, a validator confirms the AI output is fit for use, not necessarily that the final decision is correct. | Quality assurance, accountability | Required Medium risk |
| validation_result | Approved, rejected, or approved with modification. If modified, record what was changed and why. An unrecorded modification between AI output and human decision is a critical governance gap. | Oversight quality | Recommended High risk | |
| User interaction events triggered by actions involving the people using the system | ||||
| User complaint and feedback When a user challenges or disputes an AI output | complaint_ref | A unique reference linking the complaint to the original transaction it concerns. Without this link, you cannot investigate whether the system behaved correctly, the complaint is unverifiable. | Redress, accountability | Required High risk |
| complaint_outcome | The result of processing the complaint — upheld, rejected, referred. Log when each stage occurred and who was responsible. Regulators may require evidence that complaints were handled within defined timeframes. | Redress, regulatory compliance | Recommended High risk | |
| User information disclosure When privacy notices, disclaimers, or policy terms are communicated | disclosure_type | What was communicated, privacy notice, AI disclosure, terms of service, limitation of liability. Log each disclosure separately so you can prove which notice a user received at which point in time. | Legal compliance, consent | Required Medium risk |
| acknowledgement | Whether the user accepted, declined, or did not respond to the disclosure and when. This is your evidence of consent or its absence. Critical for GDPR, AI Act, and similar frameworks requiring informed consent. | Consent management, legal | Required High risk | |
| AI/ML model development events, for auditable training of machine learning models | ||||
| Training checkpoint At repeated intervals during model training | epoch_id | Identifies where in the training process this checkpoint was taken which iteration or epoch. Allows reconstruction of the training trajectory and selection of the best-performing model version post-training. | Model auditability, reproducibility | Required Medium risk |
| model_parameters | A snapshot of the model weights at this checkpoint. This is what allows you to restore and re-evaluate any historical model state. Store until the final model selection decision is made; retain for selected models per legal and business requirements. | Reproducibility, regulatory audit | Required High risk | |
| quality_metrics | Performance measures at this checkpoint, such as validation loss, accuracy, F1, or domain-specific metrics. Enables assessment of overfitting and helps justify the final model selection to auditors and regulators. | Model selection justification | Required High risk |
Risk levels are assigned based on the consequences of not capturing the field, not just the sensitivity of the field itself. For example, recovery_action on a software error is flagged high risk because a silent failure with no logged response is one of the most common and serious gaps in AI oversight programs.
Designing the Logging System with Risk as the Primary Driver
Risk is the primary driver for monitoring and controlling AI systems that are enabled by logging. Risk must be considered when determining which events are to be detected, determining which events are relevant, and determining which relevant events are to be logged.
Examples of risk management standards that can be applied include ISO 23894 or prEN 18228.
Events must be logged in relation to inputs or outputs and when caused or observed by the controllers or components of the AI system. Relevant events to be logged must be selected based on risk, including determining the most effective and efficient way to manage the risk.
Inputs or outputs relevant to event detection must be logged at a frequency that is technically feasible and allows risk to be managed in the context of the intended purpose. For example, events from streaming inputs or outputs can be logged at different frequencies based on the time resolution of the input, or can be logged at a frequency that is appropriate for monitoring a situation, such as at a higher frequency during a cyberattack.
Logging functions must be designed and configured to generate logs accurately representing such events.
Sources of information to be logged can include communication between end users and the AI system, communication between the AI system and its components, and acquisition and utilization of stored or external data.
Traceability Through Timestamps and Identifiers
Log entries about events should be timestamped. The timestamp must record the time of the event to an accuracy and precision appropriate for the type of event and its role with respect to the intended purpose of the AI system. Where technically feasible, the order of log entries should correspond to the order of the events logged.
Timestamps should be formatted according to ISO 8601-1. If the time zone is not included within the timestamp, a mechanism to determine the time zone which is applied for the timestamp must be specified in the technical documentation.
Log entries should include an information element that enables connection between the logged information and the AI system or its components, where appropriate.
This is not an academic concern. In a discrimination investigation, the order of decisions matters. In a safety audit, the timing of control transfers matters. In a breach investigation, the sequence of access events matters. Timestamps that are imprecise, inconsistent, or missing time zone information make logs unusable as evidence.
Additional Logging Functions for Specific Use Cases
Additional logging functions can be provided based on the nature of the system, the organization or entity role, and based on applicable legal requirements.
These can include recording the period of each system use such as start and end timestamp. They can include reference to external data source or database against which input data is checked, if applicable. They can include logging the relevant input data. They can include traceability at the level that enables identification of individuals involved in result verification.
For remote biometric identification systems under the EU AI Act, these functions are mandatory. For other high-risk systems, they depend on the risk assessment and the regulatory context.
Anomaly Monitoring of the Logging Component Itself
Logging functions must issue alerts when the integrity of log processing is violated, when the confidentiality of log storage has been compromised, when the integrity of stored logs is violated or can no longer be ensured for the full operational lifetime, or when log storage capacity is being reached or exceeded.
The frequency and monitoring of alerts must be justified.
This requirement recognizes that the logging system itself can fail. A logging component that silently drops entries, allows unauthorized access, or runs out of storage creates a false sense of compliance. The system must monitor itself and escalate failures to operators.
Triggers for Logging: Operation, Monitoring, and Oversight
A log entry can be triggered by the reception or processing of an input, by human actions, by specific software interactions, or by the automated or manual detection of certain events.
Events are at the center of certain processes within or around the AI system, including automated monitoring and human oversight. The underlying goal of logging is to keep records of relevant events occurring in relation to the AI system.
Events can pertain to the inputs, the outputs, the state of the AI system, or a combination. Relevant events can consist of a pattern of information, such as a change or a particular balance over a period of time, or they can correspond to a property of those inputs, outputs, and state, such as the presence of a particular feature. They can occur across multiple inputs or within a single input.
Detection of relevant events can occur through human oversight or automated monitoring and can involve consideration of past inputs and outputs and other information pertaining to the event.
Detected relevant events can be logged, including various information pertaining to the event and corresponding inputs and outputs. Human oversight or automated monitoring can change the outputs of the AI system.
Triggers From Operation
A log entry must be recorded when the AI system, or a component of it, encounters a software error. This can be caused by internal or external factors. The standard provides an information model for software errors that includes error codes, messages, severity level, impact level, and system context. It also includes detailed error handling information such as failed operation, retrying, switching to a fallback mechanism, notifying the user, logging the error, escalating the issue, and recovering if possible.
Outlier inputs must be detected based on statistical thresholds, domain-specific anomaly detection metrics, and contextual metadata. An outlier input is one that deviates significantly from the expected distribution or boundaries of the input space. Outliers can indicate data quality issues, adversarial inputs, or emerging use cases that were not anticipated during design.
Potential attack triggers must include unauthorized access patterns, data integrity violations, model inversion signatures, and poisoning signatures. Model inversion is an attack where an adversary uses outputs to reconstruct sensitive training data. Poisoning is an attack where an adversary manipulates training data to degrade model performance or introduce backdoors.
A log entry should be recorded when a user requests a review of a transaction, when a user submits a complaint or provides feedback, or when authorized personnel or systems process the user’s complaint or feedback. The standard provides an information model for human feedback.
A log entry should be recorded upon determination of the outcome of a user request, with a reference to the original user request. This creates an audit trail from complaint to resolution.
The communication of information to AI users or subjects can trigger a log entry. For example, communicating a privacy policy or disclaimer to a user, or their acceptance or non-acceptance of it, can be recorded in a log.
Triggers From Automated Monitoring
A log entry must be triggered when an adversarial attack is detected. This detection can occur on a single input or be inferred from a pattern over multiple inputs. In the latter case, it relies on prior logging of inputs ahead of event detection.
A log entry must be triggered when unwanted bias is detected in the outputs of the AI system. This detection is typically done over multiple inputs and corresponding outputs. It relies on prior logging of inputs ahead of event detection. Unwanted bias refers to systematic differences in outcomes across demographic groups that violate fairness constraints.
A log entry must be triggered when the AI system is detected to operate out of its domain. Depending on the domain and its defining characteristics, this detection can be made either on an individual input, such as violating feature boundaries, or it can be meaningful solely over multiple inputs, such as for domains that set distributional properties on certain features. In the latter case, it relies on prior logging of inputs ahead of event detection. Detection of domain drift must be considered as operating out of the domain.
For AI systems whose models are updated, either on a continuous basis or with another timescale or manual intervention, a log entry must be triggered when a model of the AI system is detected to have drifted. This detection is typically done over multiple inputs and corresponding outputs. It relies on prior logging of inputs ahead of event detection. Model drift occurs when the statistical properties of the model’s predictions change over time, often due to shifts in the underlying data distribution.
For AI systems containing machine learning models, the organization must determine if the models’ design and characteristics are required to be audited or auditable. Only if this is the case does the rest of this requirement apply. At repeated points during the training of a machine learning model, such as after each iteration over the whole training dataset, the logging system must log information to locate the current step within the training process such as identifier of epoch, the current model parameters also known as checkpoint, and any available information on the quality of the current model such as evaluation measures on a validation dataset, loss value on validation data, or accumulated training loss over the epoch.
This information enables assessment of overfitting characteristics of deployed models. Overfitting occurs when a model learns the noise in the training data rather than the underlying signal, resulting in poor generalization to new data.
The logs must be stored until a decision is made to select one or more trained models among the candidate ones, and are retained at least for the models selected, and more if there are legal requirements specifying otherwise or other business value.
Triggers From Human Oversight
A log entry must be recorded, including unique identification of the representative or controller, as appropriate, when a human controller has interrupted or intervened in the operation of an AI system to prevent or remediate a serious incident, checked or validated an output of an AI system, or engaged, transferred, or disengaged control of an AI system.
The standard provides an information model for control activities and human oversight.
The organization must assess and justify whether it is necessary to record the reason that these events occurred based on risk.
Where human actions occur outside the technical boundary of the AI system, the logging functions must record them based on applicable regulatory requirements.
This requirement reflects the reality that many AI systems operate under partial human control. A human operator can override a decision, pause the system, or hand off control to another operator. Those actions are not internal to the AI system, but they are part of the system’s operational history and must be logged to establish accountability.
Required Information in Log Entries
The log record must be linked to AI system version information, which enables the connection between each log entry and the version of the AI system. When the AI system is based on multiple models or a model that changes over time, then model identifier and version information must also be included.
The log entries must contain a unique reference to the log event, a timestamp of the log entries using a standardized time format such as ISO 8601-1 for systems that have access to clock time, and inputs and outputs if they are necessary for understanding and analysis of an event having triggered this log entry or for supporting the detection of future events.
AI systems that do not have access to clock time must include information to enable estimation of time since the start of the AI system, such as the number of clock cycles since startup or a numerical identifier giving an ordering to entries.
A unique reference to the inputs and outputs may be used in place of the inputs and outputs. For example, sensor values can be stored in another system specifically for that purpose, and a reference to each sensor value be included in the AI system log.
Recommended Information in Log Entries
The log entries should contain event types which affect the ability of an AI system to perform in accordance with its intended purpose, such as inputs received, output generated, or error encountered.
They should contain source identification, for scenarios where inputs are routed from multiple sources, input provenance, traceability, and security auditing purposes. In case of multiple sources, each source can correspond to a sensor and a single sensor value can be traced to each individual sensor. Examples include sensor identifier, API endpoint, or data stream identifier.
They should contain a correlation identifier that correlates related log entries across the system for traceability. In an AI system in which a request is processed in several stages before a response is returned, the system can include an identifier that connects the content of the log entries at each stage of the request and is unique to the request.
They should contain system status with respect to a situation or behavior when the event was logged. A system status does not have to be recorded through a logging component for the AI system. It may be recorded through other logging components.
They should contain error handling, which includes detailed information on any errors or exceptions, including error codes and descriptions. They should contain error information containing detailed information on any errors or exceptions that occurred, including error codes, error messages, severity level, impact level, and system context.
They should contain detailed error handling information, which includes detailed information on failed operation if any, retrying, switching to a fallback mechanism, notifying the user, logging the error, escalating the issue, and recovering if possible.
Storing and Access to Logs
Logs refer to all the log entries that are created at some point in the life cycle of the AI system. However, this does not imply that those log entries are kept forever or are necessarily accessible to all stakeholders.
Some log entries warrant long-term storage, for instance if they are required for fulfilling regulatory obligations on record keeping. Log entries warranting long-term storage must be stored in a persistent way for future access. If the obligation to store them comes from an external stakeholder to the organization itself or applicable regulatory requirements, then the logs must have backups.
Governance schemes can both promote and restrict data access in relation to logging. Legal requirements, for example about data portability and privacy, can expand or restrict the requirement to maintain logs, along with the ability to use and share them.
Requirements for Third-Party Access
The organization may refrain from transmitting the AI system logs or parts of AI system logs if the intended recipient of the log has no permission to access the otherwise included information.
The transmission of the AI system logs or parts of AI system logs may be rejected if the intended recipient does not ensure that logs are stored securely, including confidentiality, integrity, and availability according to applicable regulatory requirements and the state of the art, that logs, backups, and derived information are deleted when no longer needed or when legally required to delete, that logs are not transferred to third parties unless the organization agrees, and that results of the evaluation of logs by the recipient are made available to the organization upon request.
Specific considerations of the legal basis, such as data subject consent, can be relevant to take into account when deciding whether to reject.
If there are multiple logging components within a single AI system and logging can be aggregated from the logging components, then aggregated logs can be transmitted.
Access for AI Users and Providers
The persons performing human oversight must have access to log entries triggered by automated monitoring.
Access to logs by the AI provider can be useful, for instance for facilitating post-market monitoring of an AI system. This access is typically subject to limitations due to confidentiality, intellectual property, or privacy. Aggregated information from logs can be accessed instead of the logs themselves.
Practical Implementation, Start With Risk and Work Backward
The standards are not prescriptive about architecture. You can implement logging in software, hardware, or a hybrid configuration. You can use centralized log aggregation or distributed logging pipelines. You can store logs in relational databases, object storage, or time-series databases.
What matters is that your logging system meets the functional requirements and supports the risk management, compliance, and oversight needs of your organization.
Start by conducting a risk assessment. Identify the harms that the AI system could cause, the events that could indicate those harms, and the evidence you would need to detect, investigate, or prove those events. Map those events to log triggers. Define the content, frequency, and retention requirements for each event type.
Next, design the logging architecture. Decide where logging components will be deployed, how log entries will be structured, how logs will be stored, and who will have access. Document your design decisions, justify them in terms of risk and compliance, and embed them in your technical documentation.
Implement logging as part of the system build, not as an afterthought. Instrument your code to generate log entries at the defined triggers. Validate that timestamps are accurate, that log entries contain the required information, and that logs are stored securely.
Test your logging system under realistic conditions. Simulate failures, adversarial inputs, domain drift, and human interventions. Verify that the logging system captures the events, that the logs are interpretable, and that you can reconstruct what happened.
Establish governance processes for log review, retention, and access. Define who can read logs, who can write logs, who can delete logs, and under what conditions. Implement access controls, audit trails, and tamper detection. Train your staff on how to use logs for monitoring, troubleshooting, and compliance.
Monitor the logging system itself. Set up alerts for log processing failures, storage capacity limits, and integrity violations. Treat logging failures as system failures.
Finally, map your logging system to the requirements in the future prEN 18229-1 if you are deploying high-risk AI in Europe. Verify that your logs support post-market monitoring, deployer oversight, and the specific obligations in Article 12 and Article 14. Update your technical documentation to reference the standard and explain how your logging system meets it.
AI System Logging Control Matrix
Below is a structured compliance reference for AI governance practitioners mapping every logging requirement, obligation level, and implementation guidance drawn from the ISO/IEC 24970 standard on AI system logging.
| Control ID | Requirement | Obligation | Implementation guidance | Examples and notes | Governance purpose | Evidence and artefacts |
|---|---|---|---|---|---|---|
| 5.1.1 | An AI system log must represent information related to the operation, behavior, inputs, outputs, or context of an AI system, recorded to support current or future retrieval, analysis, oversight, or decision review. | Shall | Define the purpose of the log during the design phase. Logs must cover at minimum what the system did, what it received, and the context in which it operated. Retrieval must be possible for future audits rather than just live monitoring. | A loan decision AI must log the input features used, the credit score output, and the regulatory context, rather than only logging error events. | Accountability and audit readiness | Log schema documentation, and a data flow diagram showing log coverage |
| 5.1.2 | AI system logs can consist of structured, semi-structured, or unstructured data and can originate from the AI system itself, its internal components, interacting AI systems, users, or external observers. | May | Design logging to accommodate multiple data formats and sources. Do not restrict the logging architecture to a single format. Ensure your ingestion pipelines can handle structured JSON, semi-structured text annotations, and unstructured data such as screenshots or audio clips. | A multimodal AI processing both text and images can log text responses as JSON and image outputs as file references pointing to object storage. | Completeness and flexibility | Logging architecture diagram, and an ingestion pipeline specification |
| 5.1.3 | AI system logs can include time-stamped events, status snapshots, sensor or input data, outputs, decisions, error messages, environmental context, or annotations. | May | Use this list as a completeness checklist when designing log scope. For high-risk systems, you should consider all eight categories. For each category you exclude, document the justification in your technical files. | A fraud detection system should log the transaction data, the fraud probability score, the decision threshold applied, any model timeout, and the network latency at the time of the event. | Log completeness and risk coverage | Log content specification, and a gap analysis against this standard checklist |
| 5.1.4 | AI system logs can be stored persistently, processed in real time, managed under data minimization or privacy constraints, machine-readable, or human-interpretable. | May | Select storage and processing modes based on risk and specific use cases. High-risk systems with regulatory obligations require persistent storage. Systems needing live intervention require real-time processing. All logs must be interpretable by a human auditor because machine-readable data alone is insufficient for governance. | A medical AI must retain logs persistently for regulatory inspection, process alerts in real time for patient safety, and present logs in human-readable form to clinical auditors simultaneously. | Privacy compliance, auditability, and oversight | Retention policy, privacy impact assessment, and log viewer documentation |
| 5.2.1 | A logging component must be a functional part of an AI system, or an external system interacting with it, that supports the generation, capture, formatting, storage, or management of log data. | Shall | Formally identify logging components and document them in the system architecture. They can be internal or external, such as a separate observability platform or compliance logger. Regardless of location, they are subject to the same governance requirements as the AI system itself. | A cloud-based AI service can use the logging service of its cloud provider as an external logging component, but the organization remains accountable for what is logged and how it is secured. | System design and accountability | Architecture diagram identifying all logging components, and vendor contracts for external loggers |
| 5.2.2 | A logging component can consist of one or more subcomponents responsible for event detection, recording log entries, applying data policies such as filtering or redaction, or ensuring secure and reliable handling of log information. | May | Decompose complex logging needs into dedicated subcomponents. A redaction subcomponent should run before storage to prevent personal data from entering persistent logs. Event detection subcomponents should operate independently from storage subcomponents to avoid a storage failure silencing your detection capabilities. | A redaction pipeline strips patient names from medical AI logs before writing them to the audit store. A separate detection subcomponent receives the unredacted stream to assess anomalies but does not persist the sensitive data. | Privacy by design and resilience | Subcomponent design documentation, redaction policy, and a data flow diagram |
| 5.2.3 | Logging components must not assume a fixed structure, automation level, or deployment location. They can be implemented in software, hardware, or hybrid configurations. | Shall | Keep the logging system design flexible. Do not hard-code assumptions about where logs will be written or how automation is applied. This is critical for edge deployments, federated systems, or systems deployed in air-gapped environments. | An autonomous vehicle AI can log safety-critical events to onboard hardware storage and synchronize to a cloud audit store when connectivity becomes available. | Resilience and deployment flexibility | Logging design specification covering all intended deployment configurations |
| 5.5.1 | AI system logging must be the process of generating, capturing, recording, and managing information related to the operation, behavior, decisions, or context of an AI system. | Shall | Logging encompasses the full lifecycle including generation, capture, recording, and management. You must govern all four stages properly. | An organization that captures logs but never manages their retention or access controls has an incomplete logging governance program. | Governance completeness | Logging governance policy covering all four stages, alongside a retention schedule |
| 5.5.2 | Logging can be performed automatically by system components, manually by users or operators, or through hybrid methods. It can occur during design, testing, deployment, or post-deployment operation. | May | Establish which logging activities at each lifecycle stage are automated versus manual. Manual logging requires the same integrity controls as automated logging. Hybrid approaches are valid but require clear process documentation. | During testing, a developer manually annotates a log entry to flag unexpected model behavior. This annotation becomes a formal log entry subject to standard access controls and retention rules. | Lifecycle coverage and integrity | Logging procedures for each lifecycle stage, and an annotation policy |
| 5.5.3 | Logging activities can be continuous, event-driven, scheduled, or conditional. | May | Choose your logging frequency based on risk and operational needs. Document the chosen mode and its justification in your technical documentation. | During a cybersecurity incident, a system configured for event-driven logging switches to continuous logging of all inputs to capture the full attack pattern. | Risk proportionality and operational efficiency | Logging frequency policy, and technical documentation justifying the chosen modes |
| 5.6.1 | AI system logging must provide capabilities enabling traceability between multiple events and log entries if necessary to manage risk, relevant to the intended purpose, and technically feasible given the inputs and outputs. | Shall | End-to-end traceability is a core requirement. For any event requiring investigation, you must be able to reconstruct the full chain of events. Implement correlation identifiers to link related log entries across components and stages. | In a multi-stage content moderation AI, a single user post passes through language detection, toxicity scoring, and policy enforcement. A shared correlation ID links all three log entries so an auditor can reconstruct the full processing chain. | Accountability, forensics, and audit | Traceability design specification, and correlation ID implementation documentation |
| 5.7.1.a | The organization must identify the security and privacy requirements for integrity and confidentiality protection of logs. | Shall | Conduct a data protection impact assessment specific to logging. Identify what personal data might appear in logs, who has access, what encryption is applied, and which regulatory frameworks apply. | A healthcare AI must identify that patient identifiers can appear in input logs and require that these be pseudonymized before storage, encrypted at rest, and accessible only to authorized clinical staff. | Privacy compliance and data security | Data protection impact assessment, encryption specification, and an access control policy |
| 5.7.1.b | The organization must protect information in logs while taking into account different purposes of logging for different stakeholders. | Shall | Different stakeholders have different legitimate access needs and risks. Design role-based access controls reflecting these different purposes rather than giving all stakeholders access to all log content. | An AI developer needs full stack traces and model parameters, while a data protection officer only needs to see whether personal data was processed lawfully. | Privacy, role-based access, and proportionality | Stakeholder access matrix, and role-based access control documentation |
| 5.7.2.1 | Logging functions must enable the recording of events relevant for identifying situations that can result in the AI system presenting a risk according to the risk management process. | Shall | The risk register must drive log design. For every identified risk in the AI system risk assessment, there must be a corresponding log event or pattern that would surface that risk if it materialized. | If the risk register identifies biased outputs as a risk, the logging system must capture outputs with sufficient demographic context to detect this pattern through analysis. | Risk management integration | Risk register, and a mapping document linking risks to log events |
| 5.7.2.2 | Logging functions must facilitate monitoring of AI systems proportional to their risks, enabling collection, documentation, and analysis of performance data from initial development to end of retirement. | Shall | Logging must span the entire AI system lifecycle. Monitoring intensity should scale with the risk level, meaning higher-risk systems warrant more frequent and comprehensive logging. Performance data must be collected and actively analyzed. | A high-risk AI system used in employment decisions requires logging throughout development, testing, deployment, and decommissioning. | Lifecycle governance and proportionality | Lifecycle logging plan, and a risk-proportionate monitoring schedule |
| 5.8.a | Technical documentation must explain and justify the specific criteria for determining relevant events. | Shall | Document not just which events are logged, but why those events were chosen. The criteria must be traceable to risk assessments so an auditor can understand the decision logic for event selection. | Any transaction where the confidence score falls below a specific threshold is logged as a low-confidence event because the risk assessment identifies this as a driver of incorrect decisions. | Auditability and transparency | Technical documentation section on event selection criteria, and a risk-to-event mapping |
| 5.8.b | Technical documentation must explain and justify the specific criteria for logging relevant events. | Shall | Separate from determining which events are relevant, document the criteria for how and when they are logged. This includes thresholds, sampling rates, and triggering conditions, alongside justifications for any exclusions. | Inputs highly deviating from the mean are logged with full payloads, while minor deviations are logged with a flag only due to storage constraints and low incremental risk. | Auditability and design transparency | Technical documentation section on logging criteria, and a storage cost versus risk trade-off analysis |
| 5.8.c | Technical documentation must specify any interaction with human controllers. | Shall | Identify every point in the system where a human can observe, intervene, override, or validate the AI system behavior. Document how each interaction is logged to maintain accountability. | The documentation specifies control points such as a compliance officer pausing inference, a caseworker overriding decisions, and a data scientist retraining the model. | Human oversight and accountability | Human-in-the-loop design documentation, and a control point register |
| 5.8.d | Technical documentation must specify any interaction with automated monitoring. | Shall | Document what automated monitors are connected to the logging system, what conditions they detect, and what actions they trigger. Include the detection thresholds and the basis for setting them. | An automated bias detection module checks output distributions periodically and triggers an alert log entry if the demographic parity gap exceeds an internal policy threshold. | Automated oversight and transparency | Automated monitoring specification, and a threshold justification document |
| 5.8.e | Technical documentation must recommend a frequency and scope of monitoring for relevant events. | Shall | Specify how often each category of event is monitored and reviewed. Scope should define which aspects of the log are reviewed and by whom. | High-risk events such as adversarial attacks are monitored in real time by automated systems and reviewed by a human within hours, while routine events are reviewed in weekly batch analyses. | Operational oversight | Monitoring schedule, and escalation procedures |
| 5.8.f | Technical documentation must recommend a frequency and scope of logging relevant events. | Shall | Specify how often logging occurs for each event type and what information is captured each time. Document the trade-offs between observability, cost, and data volume. | Transaction initiation events are logged continuously, model drift metrics are logged hourly as a batch summary, and training checkpoints are logged after each epoch. | Design governance | Logging frequency specification per event type |
| 5.8.g | Technical documentation must explain and justify the accuracy and precision of timestamps where used. | Shall | Document the clock source, its synchronization mechanism, the precision used, and the timezone convention. For distributed systems, document how you manage clock skew between components. | The system uses UTC timestamps at millisecond precision, synchronized to a specific server. | Forensics and traceability | Clock synchronization specification, and a timestamp format definition |
| 5.8.h | Technical documentation must explain and justify resource constraints affecting logging, such as memory capacity, storage capacity, or processing power. | Shall | Document the physical and financial limits on logging. If resource constraints force explicit trade-offs, these trade-offs must be justified and reviewed periodically as risk levels change. | An edge deployment has limited onboard storage. Logs are compressed and streamed to cloud storage periodically. If connectivity is lost, the system overwrites the oldest entries first. | Risk management and design | Resource constraint analysis, and a fallback logging specification |
| 5.8.i | Technical documentation must explain and justify constraints related to privacy that affect logging. | Shall | Identify every privacy constraint that limits what can be logged based on data protection laws, contractual obligations, or ethical commitments. Document what data is excluded from logs and list any compensating controls. | Privacy constraints prevent logging raw user queries containing sensitive health data. As a compensating control, queries are classified and logged using category codes instead of full text. | Privacy compliance | Privacy constraint register, legal basis documentation, and compensating controls |
| 5.8.j | Technical documentation must refer to related legal requirements concerning data protection, system accountability, traceability, and transparency. | Shall | Maintain a live register of applicable legal requirements that intersect with logging, such as the GDPR or the EU AI Act. Update this register when regulation changes. | The legal requirements register includes rules around data minimization, retention limits, and specific logging mandates for high-risk AI systems. | Legal compliance | Legal requirements register, and a regulatory mapping document |
| 5.8.k | Technical documentation must include appropriate information security considerations and data retention policies for logs. | Shall | Logs are sensitive assets. Document the security controls applied, such as encryption, access controls, integrity protection, and retention periods. | Transaction logs are retained for several years due to regulatory requirements, encrypted at rest and in transit, and accessed only via multi-factor authentication. | Information security and compliance | Retention schedule, encryption specification, access control policy, and integrity protection specification |
| 5.8.l | Technical documentation must include a specification of failure handling detailing how the AI system reacts when log memory is overloaded. | Shall | Define what happens when storage is full, when the logging component crashes, or when network connectivity is lost. Failure modes must be designed to avoid silent data loss. | If log storage reaches capacity, an alert is raised and the system switches to emergency logging mode. If storage hits maximum capacity, the system halts new inference requests rather than proceeding unlogged. | Resilience and safety | Failure mode specification, and an incident response procedure for logging failures |
| 5.8.m | Technical documentation must include interfaces with other systems that affect logging. | Shall | Document every external system that sends data to or receives data from the logging system. Include APIs, data formats, authentication methods, and failure responses. | Logging interfaces include upstream model serving infrastructure pushing events via an internal API and downstream platforms pulling logs securely. | System architecture and completeness | Interface register, API specifications, and failure response documentation |
| 5.8.n | Technical documentation must contain a specification of the log data structures used. | Shall | Publish a formal schema for every log entry type to enable automated processing, consistent querying, and third-party audits. Specify field names, data types, and permissible values. | A transaction log entry schema requires specific fields like event IDs, system IDs, timestamps, and input references. | Interoperability and auditability | Log schema specification, data dictionary, and schema version control |
| 6.1.1 | Risk must be considered when determining which events are to be detected. | Shall | The AI system risk assessment must be the primary input to logging design. Every identified risk must map to at least one detectable event in the logging system. | A risk of geographic bias translates to a detectable event where region tags are logged with each transaction and analyzed in batch reviews. | Risk management integration | Risk-to-event mapping document, and a risk register |
| 6.1.2 | Risk must be considered when determining which events are relevant. | Shall | Relevance is determined by the risk context. Document the relevance criteria explicitly to avoid logging trivial events that create noise and degrade the quality of governance. | A model serving high request volumes logs transactions above a specific value threshold and samples a small percentage of routine transactions for monitoring purposes. | Risk proportionality | Relevance criteria documentation, and a risk-proportionality justification |
| 6.1.3 | Risk must be considered when determining which relevant events are to be logged. | Shall | Log events in order of risk severity. Safety-critical and compliance-critical events must always be logged, while lower-priority events can be subject to sampling or conditional logging. | Priority events like adversarial attacks or bias detections are always logged. Routine transaction metadata is sampled based on available resources. | Risk prioritization | Event priority matrix, and a logging resource allocation policy |
| 6.1.4 | Events must be logged in relation to inputs or outputs when caused or observed by controllers or components of the AI system. Relevant events to be logged must be selected based on risk. | Shall | Every logged event must be anchored to an observable input or output rather than an internal state that cannot be independently verified. Record the analysis that led to the selection of logged events. | When a human operator overrides a model output, the log entry captures the original output, the override action, the modified output, and the controller identity. | Accountability and verifiability | Event selection analysis, and a risk-based selection methodology |
| 6.1.5 | Inputs or outputs relevant to event detection must be logged at a frequency that is technically feasible and allows risk to be managed in the context of the intended purpose. | Shall | Set frequency based on the time horizon of the risk. If a risk could cause harm quickly, logging frequency must be sufficient to detect it within that window. | A trading AI with systemic risk implications logs transactions in real time, while a content recommendation AI logs aggregate bias metrics hourly. | Risk timeliness | Frequency justification per event type, and a risk time-horizon analysis |
| 6.1.6 | Logging functions must be designed and configured to generate logs accurately representing logged events. | Shall | Validate logging accuracy periodically by comparing logged data against ground truth from the AI system itself. Treat any logging inaccuracy as a governance defect. | During a validation test, any discrepancy between the submitted test inputs and the logged inputs requires immediate remediation. | Integrity and reliability | Logging accuracy validation procedure, and test results |
| 6.2.1 | Log entries about events should be timestamped. The timestamp must record the time of the event to an accuracy and precision appropriate for the type of event and its role with respect to the intended purpose. | Shall | Timestamp precision must match the risk horizon of the event. Real-time safety-critical systems require high precision, while compliance reporting systems can use lower precision. | A high-frequency trading AI requires microsecond timestamps to reconstruct event orders, while a monthly bias audit system requires only date-level timestamps. | Forensics and sequencing | Timestamp precision specification per event type, and a justification document |
| 6.2.2 | Where technically feasible, the order of log entries should correspond to the order of the events logged. | Should | Log entry ordering is critical for forensic reconstruction. Use sequence numbers or logical clocks where exact wall-clock ordering cannot be guaranteed, and document any known ordering limitations. | In a distributed AI system with latency, a logical sequence number is appended to each log entry to provide ordering within each node. | Forensics and traceability | Ordering mechanism specification, and known limitations documentation |
| 6.2.3 | Timestamps should be formatted in a standardized format. If the time zone is not included within the timestamp, a mechanism to determine the applicable time zone must be specified in technical documentation. | Shall | Use the ISO 8601 format with explicit UTC offsets for all timestamps. If system constraints prevent this, the technical documentation must provide an unambiguous method for determining the applicable timezone. | Timestamps use clear formatting with explicit UTC offsets. If the timezone cannot be included, documentation strictly defines the default timezone used by the system. | Interoperability and forensics | Timestamp format specification, and timezone documentation |
| 6.2.4 | Log entries should include an information element that enables connection between the logged information and the AI system or its components, where appropriate. | Should | Every log entry should carry an identifier linking it to the specific AI system that generated it, which is essential in shared infrastructure environments. | In a microservices environment, each log entry includes a system ID that identifies which governed AI system generated the entry. | Accountability and attribution | System reference specification, and an AI system register |
| 6.3.a | Additional logging functions can record the period of each system use, such as start and end timestamps. | May | Recording session boundaries enables you to calculate system utilization and identify unusually long sessions that may indicate misuse. | A medical AI logs session start and end times per user. Sessions longer than expected trigger a review. | Usage monitoring and security | Session logging specification |
| 6.3.b | Additional logging functions can reference an external data source or database against which input data is checked. | May | If the AI system validates inputs against an external reference like a sanctions list, log which version of that external source was used at the time of the check. | A financial crime AI records the specific sanctions database version identifier used for each check to enable retrospective reviews if the list updates. | Reproducibility and accountability | External reference logging specification, and a version management policy |
| 6.3.c | Additional logging functions can log the relevant input data. | May | Full input logging provides the richest basis for audit but carries high data volume and privacy costs. Log full inputs for high-risk decisions and log input references for routine transactions. | A credit decision AI logs the full feature vector for declined applications to enable explanations, while approved applications are logged by reference only. | Explainability and redress | Input logging policy, and a privacy impact assessment |
| 6.3.d | Additional logging functions can provide traceability at a level that enables the identification of individuals involved in result verification. | May | If human verification of AI outputs is part of the process, record exactly who performed the verification to establish individual-level accountability. | When a caseworker verifies a benefits assessment recommendation, the log records their employee ID, the timestamp, and whether they accepted or modified the output. | Human oversight accountability | Verification logging specification, and an individual identification mechanism |
| 6.4.a | Logging functions must issue alerts when the integrity of log processing is violated. | Shall | Monitor the logging pipeline itself. If log entries are dropped or corrupted, this is a critical governance failure. Implement checksums and processing integrity checks. | Alerts trigger immediately if the message queue depth exceeds expected thresholds or if checksum mismatches are detected. | Logging integrity and governance assurance | Log processing integrity monitoring specification, and alert configurations |
| 6.4.b | Logging functions must issue alerts when the confidentiality of log storage has been compromised. | Shall | Unauthorized access to log storage is a security incident. Implement access logging on the storage itself and alert on anomalous access patterns. | Alerts trigger if log storage is accessed by unauthorized accounts or if bulk downloads occur outside normal business hours. | Information security and privacy | Log storage access monitoring specification, and an incident response procedure |
| 6.4.c | Logging functions must issue alerts when the integrity of stored logs is violated or foreseeably can no longer be ensured for the full operational lifetime. | Shall | Implement integrity verification like cryptographic hashing to maintain log integrity for the entire retention period. Alert when checks fail or storage degrades. | Daily integrity verification runs compare stored log hashes against write-time hashes, triggering alerts upon any mismatch. | Long-term integrity | Integrity verification specification, storage health monitoring, and integrity check results |
| 6.4.d | Logging functions must issue alerts when log storage capacity is reached or exceeded. | Shall | Implement tiered capacity alerts to provide sufficient warning for remediation before capacity is reached, preventing silent data loss. | Capacity alerts trigger warnings at 85 percent capacity to initiate archiving processes, and critical alerts at 95 percent to trigger emergency expansion. | Operational resilience | Capacity monitoring specification, alert thresholds, and a capacity management procedure |
| 6.4.e | The frequency and monitoring of logging anomaly alerts must be justified. | Shall | Document the justification for each alert threshold to avoid alert fatigue while ensuring genuine issues are detected. Specify alert response times clearly. | The documentation justifies capacity alert thresholds based on lead times required for archiving processes and log growth rates. | Governance assurance | Alert justification document, alert fatigue reviews, and service level agreements for alert responses |
| 7.1.1 | A log entry can be triggered by the reception or processing of an input, human actions, specific software interactions, or the automated or manual detection of certain events. | May | Implement logging triggers across all categories to prevent unmonitored activity. Map triggers to the risk register to confirm that every risk has a corresponding detection trigger. | Triggers for a claims processing AI include new claims received, human overrides, completed model inferences, and automated detection of outlier values. | Coverage and risk management | Trigger inventory, and a risk-to-trigger mapping |
| 7.1.2 | Events can pertain to inputs, outputs, the state of the AI system, or a combination. Relevant events can consist of patterns over time or properties of individual inputs, outputs, and states. | May | Design monitoring to detect both instantaneous events and gradual temporal patterns. Pattern-based detection requires input logging to be active prior to pattern identification. | A single transaction with an unusually high value is an instantaneous event, while a gradual increase in high-value transactions over a month is a pattern event requiring historical logs. | Detection completeness | Event detection specification, and a pattern detection design |
| 7.2.1 | A log entry must be recorded when the AI system, or a component of it, encounters a software error. | Shall | Log all software errors affecting the AI system, including those caused by internal faults and external factors like malformed inputs. Ensure you capture errors that affect system outputs. | If a model request times out and the AI returns a fallback response, both the timeout error and the fallback mechanism used must be logged. | Reliability and incident response | Error event log entries, and incident reports |
| 7.2.2 | Outlier inputs must be detected based on statistical thresholds, domain-specific anomaly detection metrics, and contextual metadata. | Shall | Define and document the outlier detection methodology before deployment. Derive statistical thresholds from the training data and utilize contextual metadata to inform your assessments. | An outlier is detected if a pixel intensity distribution deviates significantly from the training set or if an image resolution falls below minimum diagnostic standards. | Anomaly detection and safety | Outlier detection specification, threshold justifications, and a review schedule |
| 7.2.3 | Potential attack triggers must include unauthorized access patterns, data integrity violations, model inversion, and poisoning signatures. | Shall | Implement detection for unauthorized access volumes, tampered inputs, systematic probing patterns, and malicious inputs. This requires comprehensive input logging. | If a system detects a sequence of queries from a single source showing systematic feature variation, it flags a potential model inversion attack and triggers a security alert. | Security and integrity | Attack detection specification, security monitoring configurations, and incident response procedures |
| 7.2.4.a | A log entry should be recorded when a user requests a review of a transaction. | Should | User review requests indicate potential algorithmic unfairness or error. Link each request to the original transaction log entry using a unique reference system. | When a user disputes a loan rejection, the complaints system generates a log entry referencing the original transaction ID from the AI decision log. | Redress and accountability | Complaint log entries linked to transaction logs, and complaints management system integrations |
| 7.2.4.b | A log entry should be recorded when a user submits a complaint or provides feedback. | Should | Log every instance of user feedback, including informal ratings. Aggregate feedback serves as a governance signal to identify systematic AI errors. | The system logs negative ratings on AI recommendations, including pseudonymized user IDs and timestamps, for weekly review by the product governance team. | User redress and quality monitoring | Feedback log entries, and aggregate feedback reporting |
| 7.2.4.c | A log entry should be recorded when authorized personnel or systems process a user complaint or feedback. | Should | Log the processing of complaints to create a complete audit trail. This enables you to assess if complaints are handled appropriately and within required timeframes. | Log entries track when a complaint is received, assigned to a handler, investigated, and ultimately resolved, along with handler IDs and timestamps. | Redress process accountability | Complaints processing logs, and a handler activity audit trail |
| 7.2.5 | A log entry should be recorded upon determination of the outcome of a user request, with a reference to the original user request. | Should | Every complaint requires a corresponding outcome log entry referencing the original AI decision to prove that the issue was resolved. | Outcome logs detail the final decision, any remedial actions taken such as reversing the decision, and the handler responsible for the resolution. | Redress completeness | Outcome log entries, and complaints closure reports |
| 7.2.6 | The communication of information to AI users or subjects can trigger a log entry. | May | Log when disclosures, privacy notices, or terms of service are communicated to users. Record what was disclosed, when, and the user response. | When an AI system informs a user they are interacting with an algorithm, the system logs the disclosure type, timestamp, user identifier, and the specific disclosure text version. | Legal compliance and consent management | Disclosure log entries, consent records, and disclosure text version control |
| 7.3.1 | A log entry must be triggered when an adversarial attack is detected. Detection can occur on a single input or be inferred from a pattern across multiple inputs. | Shall | Adversarial attack detection is mandatory. Pattern-based detection requires historical input logging to analyze probing campaigns across thousands of inputs. | Detecting a prompt injection attempt in a single API call or identifying coordinated queries across multiple IPs will both trigger log entries and security alerts. | Security and integrity | Adversarial attack log entries, security monitoring configurations, and attack detection methodologies |
| 7.3.2 | A log entry must be triggered when unwanted bias is detected in the outputs of the AI system. | Shall | Bias detection requires both input and output logging to identify patterns across multiple transactions. Define acceptable fairness metrics and document the thresholds that trigger alerts. | If demographic parity gaps exceed policy thresholds over a specific rolling window, the system creates a log entry and notifies the governance team. | Fairness and regulatory compliance | Bias detection log entries, fairness metric specifications, and threshold justifications |
| 7.3.3 | A log entry must be triggered when the AI system is detected to operate out of its domain. Detection of domain drift must be considered as operating out of the domain. | Shall | Out-of-domain operation occurs when inputs fall outside the training data distribution. Detect single violating inputs or gradual distributional shifts and flag the outputs as unreliable. | If the proportion of inputs from a new geographic region increases significantly and shifts the distribution outside training parameters, a drift event is logged. | Safety and model governance | Out-of-domain detection log entries, domain boundary specifications, and domain drift monitoring |
| 7.3.4 | A log entry must be triggered when a model of the AI system is detected to have drifted. | Shall | Model drift indicates behavioral changes from validated baselines. You must log historical performance baselines alongside current metrics to accurately detect and address drift. | When the divergence between current output distributions and the rolling baseline exceeds limits, a drift event is logged to initiate model revalidation. | Model governance and safety | Model drift log entries, baseline performance specifications, and drift detection methodologies |
| 7.3.5.a | For auditable ML models, the logging system must log information to locate the current step within the training process at repeated points during training. | Shall | If auditability is required, logging infrastructure must be active during training. Use epoch identifiers to allow the reconstruction of the training trajectory. | After each epoch, the log entry records the epoch ID, training run ID, dataset version, and exact start and end timestamps. | Model auditability and reproducibility | Training log entries, and a training run registry |
| 7.3.5.b | For auditable ML models, the logging system must log the current model parameters at repeated points during training. | Shall | Model checkpoints act as physical evidence of the model state during training, enabling restoration or investigation. Account for the significant storage requirements. | After each epoch, model weights are serialized and stored to a checkpoint registry with unique IDs stored in append-only storage. | Model auditability and reproducibility | Checkpoint storage, checkpoint registry, and checkpoint integrity controls |
| 7.3.5.c | For auditable ML models, the logging system must log any available information on the quality of the current model at each training checkpoint. | Shall | Quality metrics at each checkpoint enable auditors to detect overfitting and verify that the deployed model was appropriately validated. | Checkpoint logs record validation loss, validation accuracy, and training metrics to ensure gaps indicating overfitting are reviewed before deployment. | Model quality assurance and auditability | Quality metric log entries, and overfitting assessment procedures |
| 7.3.5.d | Training logs must be stored until a decision is made to select candidate models, and retained at least for the selected models. | Shall | Logs must persist through the model selection process. Afterward, retain logs for selected models based on legal requirements and document the deletion of rejected models. | Following a training run, logs for the selected model are retained for years, while logs for rejected epochs are deleted shortly after the decision is documented. | Retention compliance | Retention policy for training logs, model selection decision records, and deletion logs |
| 7.4.a | A log entry must be recorded, including the unique identification of the controller, when a human interrupts or intervenes in the operation of an AI system to prevent or remediate a serious incident. | Shall | Accountability requires a named individual in the log, not a generic team role. Define what constitutes a serious incident and record any human intervention addressing it. | When an operator halts an AI system due to unexpected behavior, the log captures their specific employee ID, the intervention type, and the reason for the halt. | Accountability and incident management | Intervention log entries, incident reports, and controller identity verification |
| 7.4.b | A log entry must be recorded, including the unique identification of the controller, when a human checks or validates an output of an AI system. | Shall | In human-in-the-loop systems, log every validation event with the validator’s identity to create an audit trail of who approved specific AI decisions. | When a radiologist reviews a diagnostic suggestion, the log records their ID, the timestamp, and whether they accepted or modified the AI output. | Human oversight accountability | Validation log entries, and validator identity management |
| 7.4.c | A log entry must be recorded, including the unique identification of the controller, when a human engages, transfers, or disengages control of an AI system. | Shall | Control transfers shift accountability. Create clear records showing who held control, when they relinquished it, and who took over to prevent governance gaps. | During a shift handover, the log details the controllers involved, the timestamp, the specific control points transferred, and any operational handover notes. | Chain of custody and accountability | Control transfer log entries, and a control chain reconstruction capability |
| 7.4.d | The organization must assess and justify whether it is necessary to record the reason for human controller actions based on risk. | Shall | explicitly decide and document whether recording the reason for human interventions is mandatory. For high-risk systems, reasons are typically essential for regulatory reporting. | An organization mandates reason recording for employment AI overrides to distinguish legitimate governance actions from biased interventions. | Accountability and risk management | Assessment documentation, and a reason-recording policy |
| 7.4.e | Where human actions occur outside the technical boundary of the AI system, the logging functions must record them based on applicable regulatory requirements. | Shall | Implement manual log entry capabilities to capture physical actions or verbal instructions related to the AI system that regulations require you to track. | A manager verbally instructs an operator to power down a server during an incident, and subsequently enters a manual log detailing the action and regulatory basis. | Regulatory compliance and completeness | Manual log entry procedures, out-of-system action records, and a regulatory requirements register |
| 8.1.1 | The log record must be linked to AI system version information, enabling connection between each log entry and the version of the AI system. | Shall | Specify version identifiers clearly to distinguish between releases. This is essential for identifying all log entries generated by a system version if a defect is discovered later. | Log entries include granular system version IDs such as hotfix tags, allowing you to isolate transactions processed between specific updates. | Incident investigation and version control | Version identifier in all log entries, and a release register |
| 8.1.2 | When the AI system uses multiple models or models that change over time, the specific model identifier and version information must be included. | Shall | Identify the precise model processing the transaction. For third-party foundation models, ensure you log the provider’s specific model version rather than just an internal reference. | Systems utilizing external LLMs must log the specific model release versions to distinguish behaviors before and after provider updates. | Model accountability and incident investigation | Model identifiers in all log entries, model version registers, and third-party model version tracking |
| 8.1.3.a | Log entries must contain a unique reference to the log event. | Shall | Assign a globally unique identifier to every log entry at the point of event occurrence. This identifier serves as the primary key for deduplication, correlation, and auditing. | The system generates a UUID for each event, returning it to the calling application so it can be included in future user communications regarding that transaction. | Traceability and reference integrity | Event ID generation specification, and uniqueness guarantees |
| 8.1.3.b | Log entries must contain a timestamp of when the event was observed by the logging function. Systems without clock access must enable estimation of time since system start. | Shall | Record when the logging function observed the event. For embedded systems lacking real-time clocks, use cycle counts and document the methodology to approximate wall-clock time. | Systems without clocks record the cycle count alongside a boot timestamp to allow accurate approximations of event times. | Forensics and sequencing | Timestamps in all log entries, and time source documentation |
| 8.1.3.c | Log entries must contain inputs and outputs, or unique references to them, if necessary for understanding the event or supporting future event detection. | Shall | You must capture inputs and outputs for governance-relevant events. Use pointers to external storage locations for large or sensitive payloads instead of embedding them directly. | Instead of embedding a massive sensor reading, the log includes a URI pointing to a secure object storage bucket where the data is kept. | Auditability and investigation | Input and output references in log entries, and storage system specifications |
| 8.2.a | Log entries should contain event types that affect the ability of an AI system to perform in accordance with its intended purpose. | Should | Classify entries using a controlled vocabulary to enable automated filtering and targeted alert rules. Define the classification scheme before deployment. | Use an event taxonomy featuring standardized terms like input received, human override, or bias alert to streamline analytics. | Monitoring and analytics | Event type taxonomy, and classification documentation |
| 8.2.b | Log entries should contain source identification for scenarios where inputs route from multiple sources to enable provenance, traceability, and security auditing. | Should | Knowing input origins is crucial for identifying attacks or faulty sensors. Use highly specific source identifiers rather than generic API labels. | IoT systems log specific sensor node IDs and physical locations to rapidly identify which device is generating anomalous readings. | Traceability and security auditing | Source identifiers in relevant log entries, and a source registry |
| 8.2.c | Log entries should contain a correlation identifier linking related log entries across the system for traceability. | Should | Generate a correlation ID when a request is received and propagate it to all downstream components to track the transaction through the entire processing chain. | An auditor can use a single correlation ID to query log entries from the API gateway, preprocessing service, model inference layer, and response handler simultaneously. | End-to-end traceability | Correlation ID implementation, and traceability query capabilities |
| 8.2.d | Log entries should contain system status context regarding situations or behaviors present when the event was logged. | Should | Include context such as system load, active maintenance windows, or recent configuration changes to distinguish normal variations from genuine incidents. | A bias alert log includes system context noting recent model weight updates, helping investigators determine the root cause of the alert. | Context and investigation | System status logging specification, and status data sources |
| 8.2.e | Log entries should contain detailed information on errors or exceptions, including error codes and descriptions. | Should | Standardize error codes and descriptions into human-readable formats. Raw stack traces are insufficient for governance as they require developer interpretation. | Error logs detail both the technical exception and a human-readable summary explaining that the user received a fallback response. | Incident management and auditability | Error taxonomy, and error code documentation |
| 8.2.f | Log entries should contain error information detailing severity levels, impact levels, and system context. | Should | Distinguish between technical severity and actual user impact. Both dimensions must be logged alongside system context to enable proportionate incident responses. | A core model failure triggers a high severity alert, but indicates low user impact because a fallback model successfully served the requests. | Incident prioritization and response | Error log entries containing severity and impact data, and impact classification methodologies |
| 8.2.g | Log entries should contain detailed error handling information such as retries, fallback switches, user notifications, escalations, and recovery. | Should | An error handled gracefully is entirely different from a silent failure. Log the complete response chain to enable assessments of your error handling procedures. | The log chain captures exactly when an error was detected, when retries failed, when fallback mechanisms activated, and when the primary model was restored. | Resilience and incident management | Error handling log entries, and incident timeline reconstruction capabilities |
| 9.1.1 | The organization is not required to keep all log entries forever or make them accessible to all stakeholders. | May | Establish a formal retention schedule specifying retention periods and deletion triggers for each log entry type based on legal obligations and business needs. | Transaction logs are kept for years due to financial regulations, while user complaint logs are retained based on limitation periods for legal claims. | Retention compliance and data minimization | Retention schedule, legal requirements register, and deletion procedures |
| 9.1.2 | Log entries warranting long-term storage must be stored persistently for future access. | Shall | Implement persistent storage specifically for log types requiring long-term retention. Ensure storage survives hardware failures, software faults, and deliberate deletion attempts. | High-risk AI logs are stored in append-only storage replicated across geographically separated data centers, with periodic recovery testing. | Regulatory compliance and resilience | Persistent storage specification, replication architecture, and recovery test results |
| 9.1.3 | If external stakeholders or regulatory requirements mandate log storage, the logs must have backups. | Shall | Backup copies are mandatory for regulated logs. Backups must be independent of primary storage, regularly tested for restorability, and subject to strict security controls. | Logs are backed up daily to separate sites, encrypted with independent keys, and tested monthly to ensure regulatory compliance. | Regulatory compliance | Backup specifications, backup test results, and an external obligation register |
| 9.1.4 | Governance schemes can both promote and restrict data access in relation to logging. | May | Document all schemes affecting log access, including regulatory inspection rights that promote access and confidentiality obligations that restrict it. Establish conflict resolution protocols. | If data subject access rights conflict with third-party confidentiality, the protocol dictates extracting and redacting the logs before sharing. | Governance and legal compliance | Governance scheme register, conflict resolution protocols, and access rights documentation |
| 9.2.1 | The organization can refrain from transmitting AI system logs if the intended recipient lacks permission to access the information. | May | Verify recipient authorizations against your access control policy before transmission. Redact logs to provide only the data the recipient is authorized to view. | A third-party auditor requesting full logs is provided a redacted extract containing decision outputs but excluding unauthorized raw input data. | Access control and privacy | Access permission assessment procedures, and recipient authorization records |
| 9.2.2.a | Log transmission can be rejected if the recipient does not ensure logs are stored securely according to regulatory requirements. | May | Obtain written assurances of security controls from third parties before sharing logs. Require encryption, access controls, and availability SLAs in data processing agreements. | Contract clauses mandate that recipients encrypt all received AI logs at rest and in transit while maintaining strict access logging. | Information security | Data processing agreements, recipient security assessments, and transmission refusal records |
| 9.2.2.b | Log transmission can be rejected if the recipient does not ensure logs and derived information are deleted when legally required. | May | Derived reports and analyses must be deleted alongside raw logs. Require recipients to provide evidence of deletion when retention periods end. | Contracts mandate that recipients delete all logs and derived analytical reports within specific timeframes and provide written certification of completion. | Data lifecycle management | Deletion obligation clauses in contracts, and deletion certificates from recipients |
| 9.2.2.c | Log transmission can be rejected if the recipient does not ensure logs are kept from third parties unless the organization agrees. | May | Control sub-processing by requiring prior written consent before a recipient transfers logs to their own vendors, ensuring sub-processors meet equivalent security standards. | Contract clauses strictly forbid recipients from sharing AI logs with third-party cloud providers without prior written consent. | Supply chain control | Sub-processing consent records, sub-processor registers, and onward transfer controls |
| 9.2.2.d | Log transmission can be rejected if the recipient does not share the results of their evaluation of the logs with the organization. | May | Maintain visibility into how recipients use your logs. Require auditors or regulators to share evaluation findings so you can improve your internal governance programs. | Audit agreements stipulate that external auditors must share summaries of their log review findings within a specific timeframe after completion. | Governance assurance | Evaluation results sharing clauses, and records of results received |
| 9.2.3 | If there are multiple logging components and logging can be aggregated, aggregated logs can be transmitted. | May | Transmitting aggregated data is a practical, privacy-preserving approach when raw logs contain sensitive information. Document the aggregation methods applied. | Instead of sharing millions of raw transaction logs, the organization transmits a monthly summary of error rates and bias metrics to a regulator. | Practical compliance and privacy | Aggregation methodology documentation, and aggregated log transmission records |
| 9.3.1 | Persons performing human oversight must have access to log entries triggered by automated monitoring. | Shall | Surface automated monitoring alerts to designated oversight personnel in a timely, interpretable format using role-controlled dashboards. | Governance officers utilize real-time dashboards to view bias alerts and adversarial attack detections, allowing them to drill down into the underlying log entries. | Human oversight effectiveness | Oversight dashboard specifications, access control records, and oversight personnel registers |
| 9.4.1 | AI providers can access logs for post-market monitoring purposes, subject to limitations regarding confidentiality, intellectual property, or privacy. | May | Govern provider access strictly to ensure they do not receive unfettered access to customer data. Define exactly what the provider can access and for what specific purposes. | An agreement allows an AI provider to view aggregated performance metrics monthly, but explicitly forbids access to individual transaction inputs or outputs. | Provider accountability and privacy | Provider access agreements, access scope documentation, and access logs |
| 9.4.2 | Aggregated information from logs can be accessed by AI providers instead of the logs themselves. | May | Define aggregation granularity in your provider agreements. Ensure it provides sufficient detail for monitoring while minimizing the exposure of sensitive customer data. | Providers receive monthly reports detailing latency percentiles and error rates without exposing any personal data or raw transaction content. | Privacy and provider governance | Aggregated access specifications, provider access agreements, and aggregation methodologies |
Logs Are Not Optional Anymore
The regulatory and technical landscape for AI has shifted. Logging is no longer a developer convenience or a debugging tool. It is a legal requirement, a risk control, and a source of institutional memory.
ISO 24970 and prEN 18229-1 give you the blueprint. They define what to log, when to log it, how to structure log entries, and how to manage log access and retention. They embed logging into risk management, human oversight, and post-market surveillance. They turn operational telemetry into auditable evidence.
If you are building, deploying, or operating high-risk AI systems, start designing your logging system now. Map your risks, define your triggers, implement your logging components, and establish your governance processes. Document your decisions, validate your implementation, and monitor your logs.
When your system fails, your logs will tell the story. Make sure the story you tell is one you can defend.
