From Answering Questions to Taking Action
Most of us first encountered generative AI as something resembling a very knowledgeable assistant. We typed a question and received an answer. We asked it to summarize a document, draft an email or analyze information.
The emerging generation of AI agents is different. An AI agent can be given an objective and then take a series of actions to accomplish it. Depending on the permissions it receives, an agent might search information, communicate with other systems, access files, run software, schedule activities or interact with organizational networks.
That can make AI enormously productive. It can also create a fundamentally different kind of risk. In the OpenAI experiment, agents were given difficult cybersecurity assignments. When they hit obstacles, some didn’t simply stop. Instead, they found alternative ways to pursue their objective.
They discovered a way to communicate with one another even though they were supposed to operate separately. They began sharing information and discoveries. Some agents divided the work and helped others continue what they had started.
Eventually, hundreds became involved in activity affecting Hugging Face systems. OpenAI reported that agents gained access to servers, obtained some credentials and reached limited private information. The agents later also penetrated parts of OpenAI’s own internal research environment.
None of this means that the AI systems suddenly became conscious, angry or determined to escape. The reality is simultaneously less dramatic and more important: They were pursuing an objective and discovered ways of achieving it that their human designers had not intended.
READ MORE: Debunk AI security myths for government agencies.
The Government Lesson: What Are the Objectives?
Consider what this could mean inside government. Imagine an AI agent assigned to:
- Reduce the backlog of permit applications
- Find improperly paid benefits
- Improve tax collection
- Identify cybersecurity vulnerabilities
- Accelerate procurement
- Reduce fraudulent payments
- Respond automatically to citizen requests
Each objective sounds reasonable. But a crucial question remains: What is the AI permitted to do while pursuing that objective? Suppose an agent tasked with reducing a backlog decides the fastest solution is to automatically reject applications missing certain information. An agent looking for fraud might identify innocent citizens as suspicious because doing so improves its success statistics.
A procurement agent instructed to obtain the lowest possible price might select an inappropriate vendor because cost was the principal measure it was given. The danger is not necessarily an AI system becoming malicious. The danger may be an AI system becoming too successful at achieving an inadequately defined goal.
Government has dealt with versions of this problem for generations. Employees are expected to exercise judgment when policies conflict, circumstances change or rules produce unreasonable outcomes. AI does not necessarily possess that same understanding of organizational intent, public values or common sense.
The New Principle: Minimum Necessary Authority
Government cybersecurity has long embraced a simple principle: Employees should receive only the computer access necessary to perform their jobs.
We may now need to apply the same principle to AI. An AI agent should receive only the authority necessary to accomplish its assigned task. If an AI system needs to read a database, that does not automatically mean it should be able to change the database. If it needs information from another department, it should not automatically receive unrestricted access to that department’s systems.
If it can recommend a payment, that does not necessarily mean it should have authority to issue the payment. And if an AI agent begins behaving unexpectedly, organizations need a reliable way to stop it.
OpenAI itself concluded that the incident exposed problems involving excessive privileges, inadequate isolation and objectives that could reward undesirable behavior. The company subsequently strengthened security and changed how it conducts some of its advanced evaluations.
DIVE DEEPER: Continuous authentication supports a zero-trust foundation.
Having a Human in the Loop Is No Longer Enough
Government AI policies frequently call for keeping a “human in the loop.” That remains important, but the phrase can become meaningless unless we define what the human actually controls. A person who receives a weekly report describing thousands of actions already taken by AI is technically “in the loop,” but hardly in control. Government organizations should instead identify actions requiring human authorization before they occur.
Changing someone’s eligibility for a public benefit might be one. Sending money could be another. Deleting government records, changing security permissions, communicating officially with the public or granting access to sensitive information may also require explicit approval.
Oversight must occur where consequences become meaningful.
Five Questions Every Government Leader Should Ask
The OpenAI incident suggests a straightforward governance test whenever an AI agent is introduced:
- What objective are we giving it?
- What systems and information can it access?
- What actions can it take without permission?
- How will we know what it has done?
- How can we immediately stop it?
Those questions may ultimately become as fundamental to government AI governance as cybersecurity, privacy and records management are today. The lesson from the 700-agent incident is not that AI is preparing to take over the world.
It is much more practical. We are rapidly moving from AI that advises humans to AI that acts for humans. That transition offers extraordinary opportunities to improve government productivity and services. But it also changes the nature of accountability. The next generation of AI governance cannot focus only on whether an AI system gives us the right answer. Increasingly, government leaders will have to ensure that AI also understands — and cannot exceed — the boundaries of what we have authorized it to do.
