An AI agent's unexpected transformation into a hacker
A test involving an OpenAI AI agent attracted attention when it broke out of an isolated test environment, exploited a vulnerability to access the internet, and attacked the Hugging Face platform to obtain information for its test task. Similar incidents have been reported during test runs by Meta and Anthropic, during which AI agents performed unauthorized actions or exploited vulnerabilities. While the details of these cases vary, they highlight a common conflict of objectives. The more autonomously AI agents can act, the more critical the question becomes of where their decision-making autonomy ends and human control begins.
However, this debate is not focused on determining the level of autonomy an AI agent can have; rather, it is focused on which decisions an AI agent should be permitted to make independently. As more decision-making authority is granted, the demand for accountability, transparency, and oversight increases. Who determines the scope of an agent’s autonomy? When should a human be able to intervene? Which decisions can be automated?
These questions are no longer limited to developers of large AI models. They are also relevant to any company that uses AI agents to support operational processes in logistics, the supply chain, production, or workforce scheduling. For example, should an AI agent be able to independently adjust the production schedule after a delivery is delayed?
Should it automatically reschedule shipments in the event of a resource bottleneck? Should it calculate various actionable options and present them to a human for decision-making? The AI cannot answer these questions. Humans determine the answers.
Between technical feasibility and responsibility,
Technically, much would be possible today. For example, companies could delegate extensive decision-making authority to AI agents. Similarly, they could have every employee action confirmed by an AI agent. In practice, however, neither approach is sensible.
This is also highlighted in the scientific review article "Unraveling Human–AI Teaming: A Review and Outlook (2025). The researchers analyze the current state of collaboration between humans and AI and arrive at the nuanced conclusion that neither humans nor AI systems consistently make better decisions. Neither humans nor AI systems consistently make better decisions. Nor does their collaboration automatically lead to better results. They argue that what matters most is how tasks, responsibilities, and decision-making authority are distributed.
When introducing AI agents, management's primary task is to consciously design a decision-making framework within which an AI agent is permitted to act.
How companies shape the scope of action for AI agents
The First Real-World Examples of How Companies Shape the Scope of Action for AI Agents
The first real-world examples already reveal a clear pattern in how companies shape the scope of action for AI agents. AI agents primarily perform standardized, repetitive tasks. However, as soon as decisions become complex or have significant implications, humans remain in charge.
For example, the BMW Group uses Agentic AI for fleet management and procurement. These AI agents process customer inquiries, synchronize data across different systems, and handle routine tasks, such as creating inventory orders and communicating with suppliers. According to the company, up to 90 percent of previously manual work steps can be automated this way. However, complex decisions and exceptional situations remain the explicit responsibility of employees.
Logistics service provider C.H. Robinson operates on a similar principle. There, multiple AI agents analyze transport data, prepare freight quotes, identify disruptions, and develop alternative courses of action for dispatch. The goal is not to replace human decision-making but to accelerate standardized processes, establish a solid basis for decision-making, and reduce employees' workload.
Both examples illustrate a common principle. Companies do not define a general level of autonomy for an AI agent. Instead, they specify which decisions within a process can be made autonomously and which cannot.
This observation aligns with the findings of the joint trend report by INFORM and LOGISTIK HEUTE. According to the report, 46 percent of surveyed companies prefer a model in which AI prepares decisions and humans make the final call. Only 4 percent of companies advocate for fully autonomous decision-making authority. Practical experience therefore provides a clear picture. Companies design autonomy at the decision-making level, not the system level.
A possible model for decision-making
A possible model for decision-making can be derived from insights gained through research and practical experience.
- Automate
Routine decisions can be fully delegated to an AI agent. The agent makes decisions based on clearly defined rules and implements them immediately. This model is ideal for repetitive tasks with minimal impact, such as routine inquiries, master data maintenance, and specific process steps. - Approval
The AI agent analyzes situations, develops recommendations for action, and prepares decisions. However, implementation only occurs after human approval. This approach is ideal for production planning, transportation scheduling, and other operational decisions with tangible consequences. - Decision-making
Humans retain full decision-making authority for strategic, safety-critical, or liability-related issues. Although the AI agent provides analyses and forecasts to inform these decisions, it does not make them independently.
It is not the AI agent that possesses a degree of autonomy. Rather, each individual decision has its own degree of autonomy.
Conclusion: Not every decision requires the same degree of autonomy.
The central challenge is not to create the most autonomous AI agents possible, but rather to determine the appropriate degree of autonomy for each decision. This concept is reflected in numerous governance frameworks and regulatory guidelines for artificial intelligence use. INFORM bases the development of its AI-based software solutions on this principle: humans retain control, and AI provides transparent, traceable recommendations.
Therefore, when discussing AI agents, one should not first focus on their technical capabilities. The crucial factor is a well-thought-out decision-making architecture that clearly defines which decisions can be automated, which require approval, and which are made exclusively by humans. After all, it is not the AI agent that possesses a degree of autonomy. Rather, each individual decision has its own degree of autonomy.