Skip to content

Independent R&D project · Cologne

Back to the blog

Overseeing AI: from each decision to the rules that govern them

A runtime was built so that a system can learn when to act without ever learning permission to act. The next step is not a machine that controls itself. It is oversight that changes shape: people spend less time confirming single decisions and more time on the rules, the records and the accountability behind them, at a pace that verification, legislation and the courts set.

Published September 17, 2026 14 min read

On a deep navy field, a light card in the middle stands for a system, drawn as three linked components. Three rings surround it: a thick green one close to the card, a thinner amber one further out, and a dashed pale one at the edge. To the right, outside all three rings, a small figure of a person holds a thin line that runs to the card. Under the card, a green line carries dated markers from left to right and ends at a pair of scales.

Built to refuse, not to obey

ERQYO, a research architecture described on this site, exists for one reason: to let a system learn when to act without ever learning permission to act. Its adaptive half reads events, state and the rhythm of the people and processes around a case, and proposes what looks appropriate next. Its governed half checks that proposal against authority, policy, evidence and hard time limits, and then permits it, defers it, hands it to a person or refuses it. The two halves are kept apart so that nothing the system learns can turn into a permission.

It is tempting to summarise this as keeping a person in the loop of every decision. That is not quite what the architecture does, and the difference matters for everything that follows. The deterministic instance is the operator, not the person: the operator applies the constraints in force, the same way every time, and records why. The person is the source of authority the runtime cannot create for itself. Where no mandate covers a decision, the runtime hands it over. Where the rules in force require a review, it hands it over. Where the rules allow it to proceed, it proceeds and writes down under which rules. The person is in charge without being in the loop of every step. The published vocabulary for this is human oversight by design, human on the loop.

So the question this article asks is not whether people can be taken out of oversight. They cannot, and the architecture is built on that. It is how the work of oversight changes as fewer single decisions need a person's confirmation, the direction in which the economics, the accumulating evidence and the law all point, and what has to grow in its place.

Oversight is already a dial

The framework this runtime was designed against, COADF, treats oversight as a graduated setting rather than a switch. Its eighth principle makes autonomy a property of the rules and not of the code: how much a system may do without a person is set per jurisdiction and per risk class, and it never exceeds the ceiling the law sets there. Raising it where the law allows is a rule change with a recorded legal basis, never a fork of the code. At the lowest setting a person confirms every output before it takes effect; at the highest the system proceeds and is audited afterwards; and every decision record is bound to the version of the rule set that produced it, so the setting in force can be read back later.

European law reasons the same way. Article 14 of Regulation (EU) 2024/1689, the AI Act, requires high-risk systems to be designed so that natural persons can effectively oversee them while in use, and its third paragraph says that the oversight measures shall be commensurate with the risks, level of autonomy and context of use of the system. The law does not prescribe one loop for every decision. It prescribes oversight proportionate to autonomy, and it names the minimum that oversight must always allow: to understand the system's capacities and limitations, to stay aware of automation bias, to interpret its output correctly, to decide not to use it or to override or reverse its output, and to intervene or stop it (Article 14(4)). Article 26(2) says who: deployers assign oversight to natural persons with the necessary competence, training and authority.

Read together, the principle and the article describe one mechanism. Autonomy is a level. The level is set outside the system, by people who hold authority. The law caps it. And the system carries the level on every record, so that the setting can be audited afterwards. What falls as the level rises is the number of decisions a person sees before they take effect. What does not fall is that a person set the level, that a person can override any single outcome, and that a person answers for the result.

In the loop
What declines as barriers and evidence accumulate?
The share of decisions a person sees before they take effect:
  • confirmation of each executing outcome
  • hand-overs for lack of a mandate, as mandates are written
  • reviews triggered by low confidence, as evidence improves
In the rules
What stays, whatever the level?
The part the person never delegates:
  • a person sets the level, with a recorded legal basis
  • a person can override, reverse or stop any outcome (Article 14(4))
  • a person answers for the result, to a supervisor or a court
  • the record that shows which level was in force, and why
As the level rises, one column shrinks and the other does not

The trained animal, and where the analogy stops

A useful picture for this is the trained animal. A dog that has learned to heel needs less leash than a puppy, and a working horse needs fewer hands than an unbroken one. Training moves control from the hand into the animal's habits, and the hand is needed less often. That is the direction in which oversight of AI systems is heading, and the picture captures something real: as a system's fences prove themselves and its records accumulate, the fraction of its actions a person has to confirm can fall.

The law's view of the trained animal is what makes the picture worth keeping. Under section 833 of the German Civil Code, the keeper of an animal answers for the damage it causes, trained or not. Only the keeper of a working animal, one kept for a profession or a livelihood, escapes liability, and only by proving that the animal was supervised with the care the situation required, or that the damage would have occurred even so. Article 936 of the Brazilian Civil Code says the same in one sentence: the owner or holder of an animal answers for the damage it causes unless they prove the victim's fault or force majeure. Training reduces the leash. It never reduces the keeper's responsibility, and for a working animal the keeper's defence is the record of supervision.

The analogy stops at three points, and each of them is an architectural fact. First, a trained animal's abilities are stable, while a model's change with every version, so trust earned by one version does not carry to the next, and every fence has to show again that it fires after every change. Second, the animal is not asked to hold authority, and neither is the system. The habit that lets the leash be loosened is not the habit of deciding whether it may be loosened, and the architecture keeps that decision out of reach of anything the system learns. Third, the animal keeps no record. A system can, and the record is the whole difference between a keeper who can show what care was taken and one who cannot.

Where this goes

The forecast this article makes is stated as a hypothesis, not a measurement. As executable barriers cover more of what a system may do, and as decision records accumulate into evidence that the barriers hold, the share of decisions a person has to see before they take effect will keep falling. Oversight moves from the loop into the rules: people will spend less time confirming individual outcomes and more time writing the policies a system runs under, reading the records it leaves, and deciding when a level may rise. For the people doing it, that is different work, not less work: a rule has to hold in cases nobody has seen yet, which asks more of judgment than confirming one case. The system does not control itself. It is controlled by an architecture people can inspect, at a level people set.

Whether that ends in full independence is not known, and it may be that it should not. Two things are known. For the uses the AI Act classes as high-risk, the law forbids an empty loop outright: oversight commensurate with autonomy, with a person able to override and to stop, is a design requirement and not a preference, and the ceiling on autonomy there is a legal one. For every other use, the honest position is that the level can rise only as far as verification does. A level that outruns the evidence for it is not autonomy. It is an untested claim with consequences.

That is why the direction and the pace are separate questions. The direction is set by economics and by the accumulation of evidence, and it points one way. The pace is set by three things outside the system: the law, which caps the level; the standards that translate the law's requirements into the state of the art; and the courts, which decide afterwards whether a level was acceptable at the time. The next section is about those three.

The law as ceiling, case law as history

Legislation sets the ceiling. Article 14 of the AI Act is the clearest example: where it applies, no rule set may configure a system below the oversight it requires, and a runtime that carries the autonomy level on every record has to enforce that ceiling as a constraint, not offer it as an option. The timetable is known. The prohibitions of Article 5 have applied since 2 February 2025 and the transparency duties of Article 50 since 2 August 2026. After Regulation (EU) 2026/1744, the Digital Omnibus, the high-risk regime applies from 2 December 2027 for the standalone categories of Annex III and from 2 August 2028 for systems embedded in regulated products under Annex I. Legislation moves slowly and in steps, and a ceiling that moves in steps is the right shape for a setting that people have to be able to audit.

Standards fill in what the law leaves open. Under Article 40 of the AI Act, a high-risk system that follows harmonised standards whose references are published in the Official Journal is presumed to meet the requirements those standards cover. The standards are where the phrase commensurate with the level of autonomy will acquire concrete measures, and they will be revised more often than the Regulation. They are the middle layer: slower than the technology, faster than the law.

Case law is the third layer, and the one this article expects to carry the history. A court does not decide what autonomy should be. It decides, after the fact, whether the level that was in force when something went wrong was defensible at that time, given what was known then. Directive (EU) 2024/2853 on liability for defective products, which Member States have to transpose by 9 December 2026 and which applies to products placed on the market after that date, gives that question its terms for software. Software is a product (Article 4(1)). In assessing whether it was defective, a court takes into account the effect of any ability to continue to learn after it was placed on the market (Article 7(2)(c)) and the moment at which it left the manufacturer's control (Article 7(2)(e)). The manufacturer's defence rests on what the objective state of scientific and technical knowledge allowed at the time (Article 11(1)(e)). Every one of those tests is a test about a date. The proposal for a separate directive on liability for artificial intelligence was withdrawn by the Commission in 2025, so for now these general rules, and the national fault rules beside them, are the ones that apply.

Judgment by judgment, that produces a dated record of how much autonomy was acceptable, in which use, at which state of the art. In Germany and in Brazil, where a judgment binds the parties rather than every later court, that record works through the standard of care the next case is measured against, and in Brazil also through the binding precedents its Constitution and its Code of Civil Procedure provide for. In either system, case law becomes the historicised version of the evolution this article describes. The law says how far the level may go. The courts say, one case at a time, how far it could reasonably have gone when it mattered.

  1. Legislation: the ceiling
    caps the level per use and per risk class · moves in dated steps: 2 December 2027 and 2 August 2028 for the high-risk regime · Article 14: oversight commensurate with autonomy, a person able to override and stop
  2. Standards: the state of the art
    turn commensurate into concrete measures · Article 40: presumption for systems that follow published harmonised standards · revised more often than the law, less often than the technology
    The record
    None of the three layers can work on a system that did not write down which level was in force, under which rules, with what evidence, and who held authority.
  3. Case law: the history
    judges the level in force at the time, with what was known then · Directive (EU) 2024/2853: software is a product, learning after release is a criterion, the date decides · one judgment at a time, a dated record of acceptable autonomy
Three layers set the pace, and none of them can work without the fourth thing beside them

What a runtime has to leave behind

If the courts are going to historicise the evolution, the material has to exist. That is the architectural consequence, and it is concrete. Every decision record has to carry the autonomy level that was in force and the version of the rule set that set it, so that a later reader can tell a rule from a habit. It has to carry what the system knew and did not know when it decided: whether the evidence was current, whether a mandate covered the action, whether a policy had been supplied at all, because a system that can say nobody asked is in a different position from one that can only say nothing forbade it. It has to carry who held authority, and whether a person confirmed, overrode or stopped the outcome. And it has to be replayable, so that what the system would have decided under the rules of that day can be shown rather than reconstructed from memory. The reference implementation described on this site records the level, the rule reference, the uncertainty, the authority and any human override on its decision traces, and can replay them. That it is not deployed does not change what the record has to contain.

None of this is new as a technique. Append-only logs, versioned policies and replayable decisions all predate this work. What is specific is the reason for keeping them: the record is written for a reader who does not exist yet, a supervisor, an auditor or a court, who will ask what the level was, who set it, on what basis, and whether the evidence at the time supported it. A system that can answer those four questions from its own records can be given more room. One that cannot should not be, whatever its accuracy figures say.

What this article does not say

It does not say that ERQYO is in use. It is a reference implementation in a repository, with tests, and it is not deployed; the one caller that exists runs in shadow mode behind switches that are off. The architecture is an argument about what a runtime should do, supported by code that shows it can be done. It is not a service.

It does not say that barriers can be complete. A fence proves that it fires against the defect it was built for. It proves nothing about the defect nobody has thought of, and every claim of total control should be read as a claim about the defects that were tested.

It does not say that full independence will come, or when. It says that the share of decisions a person confirms will fall, not that oversight becomes a smaller responsibility, and that the pace of that fall is set by verification, by legislation, by standards and by the courts, not by the systems themselves. Where Article 14 applies, the floor is a legal one and no configuration goes below it.

And it is not legal advice. Whether the AI Act, the product liability rules or the animal-keeper provisions cited here apply to a given system or situation depends on facts this article does not have. Every provision is cited by article so that a reader can check it against the text on the day it matters.

Key dates

  • 2 February 2025. The prohibitions of Article 5 of the AI Act apply. (Regulation (EU) 2024/1689, Article 113)
  • 11 February 2025. The Commission's work programme lists the proposed directive on liability for artificial intelligence for withdrawal; the withdrawal was completed in October 2025.
  • 2 August 2026. The transparency duties of Article 50 of the AI Act apply. (Regulation (EU) 2024/1689, Article 113)
  • 9 December 2026. Directive (EU) 2024/2853 has to be transposed; it applies to products placed on the market or put into service after this date. (Articles 2(1) and 22(1))
  • 2 December 2027. The high-risk regime, including Article 14, applies to the standalone systems of Annex III. (Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744)
  • 2 August 2028. The high-risk regime applies to systems embedded in regulated products under Annex I. (Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744)

Sources

Every provision above was read in the instrument's own text on 17 September 2026: the AI Act and the Directive in their Official Journal versions, section 833 BGB at gesetze-im-internet.de and the Brazilian articles at planalto.gov.br. The forecast in the fourth section is a hypothesis of this article and carries no source. Statements about ERQYO and COADF stay at the depth of their published pages on this site.

This article is informational and is not legal advice. What a given company owes depends on its products and its own facts, and the authoritative EU legal texts prevail over any summary of them.

Written by Luiz Hogrefe.

Share this article

Public feedback

Have a correction, implementation note or different architectural view?

Discuss this articleView the public discussion