On 2 August 2026, the EU AI Act's enforcement window for high-risk AI systems opened. Buried in the compliance checklist is Article 12, and it's the one that turns from paperwork into engineering: high-risk systems must automatically log events relevant to identifying risks and substantial modifications. This post decodes what that actually demands.
Three words that carry the weight
"Automatic" rules out best-effort logging that depends on engineers remembering to instrument. If the record exists only where someone thought to add a log line, it isn't automatic — it's aspirational. The practical test: when a team ships a new tool for an agent next quarter, does its activity enter the record because someone remembered, or because the recording layer catches it by construction? Article 12 is asking for the second answer.
"Events relevant to identifying risks" means the log must be useful for reconstruction: who acted, what was done, with what tools and inputs, with enough context that a risk pattern is findable afterwards. A latency metric doesn't meet that bar; a complete, ordered account of the system's actions does. And "record-keeping" implies the records survive, intact, for the period in which questions can arrive — a log that can be quietly edited satisfies the letter and defeats the purpose. Regulators drafting logging duties have tamper-evidence in mind even where they don't name it: a record whose integrity can't be demonstrated is testimony, not record-keeping.
Where an agent recorder fits
Wytness captures agent activity automatically — tool calls, model calls, policy decisions — and structures it into exactly the form an assessment needs. An EU AI Act Evidence Pack bundles agent inventories, tool capability descriptions, signed event samples, anomaly history, and control-by-control attestations into one report: printable for the file, and downloadable as a signed, dated JSON pack with the signed events attached.
Article 12's record-keeping also leans on its neighbour, Article 11's technical documentation — which is where the Register earns its place: risk tiers with written justifications, approval states, named operators, exportable as a document. The two articles are easier together than apart — the register says what each system is and who answers for it; the log says what it did. An assessor will read them side by side, and they were built to be read that way.
What we deliberately don't claim
The Act punishes overclaiming, and so does an assessment, so our boundaries are printed on the compliance page rather than discovered in diligence. Wytness is not a notified body and does not certify your system as conformant. Risk classification is the operator's call — we map events to the categories you identify. Bias testing is out of scope: we log activity, we don't evaluate fairness. And whether the evidence demonstrates compliance is a legal question for your counsel. We're explicit about which articles we evidence versus merely support — a distinction most vendors blur and every assessor checks.
The practical reading
If you operate agents that might be high-risk under the Act, the sequence is: classify them (your call, documented in the Register), ensure their activity is automatically recorded (the recorder's job), and be able to produce structured, verifiable records on request (the pack's job).
A useful this-quarter checklist for a team taking it seriously: enumerate the agents that plausibly touch a high-risk category; get them recording — SDK for the ones you build, collection for the platforms you operate; classify each in the register with the justification written down; and generate one pack as a dry run, because the first pack always surfaces the gap you'd rather find before an assessor does. None of those steps get easier retroactively — Article 12 describes a record that must have been running, and the only day to start one is before you need it.
Questions about your systems and the Act? Ask us.