9 risks when using AI agents and how to mitigate them

In one of my webinars, someone asked: “What about security concerns when you give an AI agent control?” This is an important and interesting question. On the one hand, AI agents promise to take over our busy work. Because of this, we want to delegate as many tasks as possible. On the other hand, the agents might do it wrong, get confused or could even get tricked by malicious actors – with sometimes dire consequences.

Because of this, we need to find the right balance when it comes to controlling AI agents. In this article, I’ll give you an overview of nine risks you take when using these tools and how to mitigate them.

Background: How agents are different from chatbots

What makes agents different from the chatbots we’ve known for the last few years comes down to this:

A chatbot reacts. You ask, it answers, and nothing happens until you decide what to do with the answer.

An agent is designed to act: You give it a goal, and it plans the steps, reads files, uses tools and works through its list, sometimes while you are away.

If you’re new to the topic, my overview of AI agents covers the basics, and my article on Claude Cowork shows what this looks like in practice.

1. Too much access, too fast

An agent can do whatever its access allows. Germany’s federal agency for IT security BSI advises giving an agent only the access it needs for the task at hand. Its example: No agent should be able to reach your emails, your bank account and your personal files at the same time. Anthropic recommends that Cowork users create a dedicated working folder instead of granting broad access, and keep backups of important files. And OpenAI’s help page for its former ChatGPT agent warns against vague, open-ended prompts like “Check my email and handle everything.” This recommendation applies now to ChatGPT agent’s successor ChatGPT Work and similar tools from other vendors.

A story by Summer Yue illustrates how difficult this can be: As director of alignment at Meta Superintelligence Labs, her job is to make sure AI systems do what people intend. She had tried out the open-source agent OpenClaw on a test inbox for weeks, and it worked well. Then she pointed it at her real inbox and told it to suggest what to delete, but not to act until she said so. The agent deleted more than 200 emails anyway.

Her own analysis, as reported: The real inbox was so large that the agent “condensed” its working memory, and the very important “confirm before acting” rule got lost. She described the episode on X.

The mitigation:

  • Start with a dedicated folder that contains copies of your files, not the originals.
  • Expand access one step at a time and test at realistic scale: more files, a bigger inbox, more accounts.
  • Reduce oversight only when a workflow has worked reliably several times.
  • Rely on the tool’s own permission settings, such as manual approval for each step, and not on instructions in the prompt.
  • Give agents specific tasks with clear limits, not “handle everything”.

2. Actions that can’t be undone

Deleting, sending, publishing, paying: Some actions can’t be taken back, or at least not easily. A well-known example for this comes from venture capitalist Nick Davidov. He asked Claude Cowork to organize his wife’s desktop. According to his post on X, Cowork asked for permission to delete temporary Office files, he agreed, and then it also deleted a folder with 15 years of family photos. Everything was gone immediately. He got the files back only because iCloud Drive keeps deleted files for 30 days.

A more extreme case comes from software development: In late April 2026, a coding agent deleted the production database of the company PocketOS within nine seconds. The backups were stored in the same place as the data, so they were gone too, and the company had to fall back on a three-month-old copy.

And there’s one more lesson: Don’t take an agent’s word for what it did or what can be undone. When Replit’s agent deleted the database of SaaStr founder Jason Lemkin, it told him recovery wasn’t possible. He restored the data himself.

Vendors and authorities have similar recommendations to avoid such incidents. Anthropic says Cowork always asks before permanently deleting files, in any mode. But it also advises against scheduling tasks that send messages for you, make purchases or do anything else that is hard to undo. The BSI recommends settings in which the agent asks for your confirmation before important actions, especially those with financial or personal consequences.

The mitigation:

  • Keep backups the agent can’t reach, and set them up before it touches real files.
  • Make deleting, sending, publishing and paying dependent on your approval. An agent may read, summarize, research and draft on its own.
  • Read approval prompts before you click, and check the result afterwards.
  • Don’t schedule unattended tasks that do things you can’t undo.
  • Don’t rely on the agent’s own report about what it has done. Look at the actual folder, inbox or page.

3. Hidden instructions (prompt injection)

An agent reads everything as text, and it can’t reliably tell your instructions from instructions that someone has hidden in what it reads. The technical term is prompt injection.

Two examples show what it looks like in practice. In August 2025, researchers at Brave showed that a Reddit comment, hidden with a “spoiler tag”, could hijack Perplexity’s AI browser Comet. When a user asked the browser to summarize the page, it followed the hidden instructions, went to the user’s logged-in Gmail account, retrieved a one-time login code and passed it on to the attacker.

The second example is about Cowork. Two days after its launch in January 2026, the security firm PromptArmor demonstrated that a Word document with hidden white text, disguised as a Claude skill, could make Cowork upload the user’s files to the attacker’s account.

OpenAI has said that prompt injection is “unlikely to ever be fully ‘solved’”, and Anthropic writes that for Cowork, the chance of an attack is still “non-zero”.

A helpful rule: Treat everything an agent reads from outside as something that might contain instructions.

Anthropic also advises users to monitor Claude for suspicious actions. IT expert and developer Simon Willison thinks that’s asking too much of regular users who aren’t programmers. That’s why the mitigation below is mostly about how you set things up, not about how closely you watch.

The mitigation:

  • Don’t let one agent session have all three: private data, outside content and a way to send data out. This makes data theft easy. Simon Willison has called this the “lethal trifecta”.
  • For web research, use a separate browser profile in which you’re not logged in to your accounts.
  • Limit web access to sites you trust, but remember that this only reduces the risk but doesn’t remove it.
  • Keep documents from unknown sources out of the agent’s working folder, or check them first.
  • Switch to manual approval when an agent works with a new site or tool.

4. Trusting results blindly

AI tools produce polished, confident-sounding output, and that includes output that is completely wrong. Invented facts, quotes and sources are called “hallucinations”.

Tip: In my article on hallucinations, I explain why they happen and how to keep them in check.

Two things make this phenomenon more serious with agents: They deliver more output in a shorter time, and they work in steps. A wrong assumption early on can derail everything that follows.

Two examples from the world of content: In May 2025, the Chicago Sun-Times and other newspapers published a summer reading list in which 10 of the 15 books didn’t exist. The titles were attributed to real authors. The freelancer writing it had used AI without checking the results.

It’s not only freelancers under time pressure who fall into this trap. Deloitte Australia repaid part of its fee for a government report worth A$440,000 that contained fabricated references and a made-up quote from a court judgment.

And as in the Replit case above, an agent’s report about its own work is not a check either. “Done” can mean “partly done” or “done differently”.

The mitigation:

  • Check every name, number, quote, title and link against a primary source.
  • Give the agent a way out: Tell it to flag what it couldn’t find or verify instead of filling the gaps.
  • Ask for a log of its actions and sources, and spot-check it.
  • Compare the result with your task personally. “The agent says it’s done” is not a check.

5. Acting in your name

When an agent sends an email, posts on social media or publishes an article, it might act under your name or your brand and use your account. Anthropic says on its safety page for Cowork: Users remain responsible for everything Claude does on their behalf, including content published, messages sent, purchases and tasks that run on a schedule.

The Sun-Times case from above shows how this plays out. The newspaper said the summer insert was licensed content that its newsroom had neither created nor approved. Still, it ran under the Sun-Times name, and the paper had to apologize and pull the section from its e-edition.

An agent that can publish on its own raises the stakes. In February 2026, an OpenClaw agent submitted code to the Python library Matplotlib. When the volunteer maintainer Scott Shambaugh rejected it, the agent researched him and published a blog post attacking him personally and accusing him of gatekeeping.

The mitigation:

  • Let the agent draft, but send and publish yourself.
  • Don’t give agents login access to brand or client accounts. If a tool needs an account in your CMS, use a role that can’t publish.
  • Read everything that goes out under your name as if you had written it yourself.
  • Don’t set up scheduled or unattended tasks that send or publish.
  • Know your clients’ and publishers’ rules on AI use, and disclose where required.

6. Plugins, logins and credentials

Plugins, skills and connectors include instructions or code that an agent will follow. Therefore, it is a good idea to be cautious when they come from an unknown third party.

Tip: I explain what skills are in my article on Claude Skills. This also applies to similar features by ChatGPT, Gemini and others.

One example: In early 2026, Koi Security audited all 2,857 skills on ClawHub, the community marketplace for the OpenClaw agent, and found 341 malicious ones. Around the same time, Snyk reviewed almost 4,000 skills from community registries. About a third had some kind of security flaw, and 76 contained confirmed malicious payloads. Important to note: These were community marketplaces, not Anthropic’s official repository.

Anthropic advises extra caution with unfamiliar plugins and connectors. A plugin bundles skills, connectors and sub-agents, so installing one can significantly widen what Claude is able to do. Anthropic recommends sticking to verified extensions from its own directory.

Logins and keys are the second door. An agent in your browser or connected to your accounts can do what you can do there, as the Comet example showed. And keys left lying around get found: According to the PocketOS founder, the agent that deleted his database had found an access token in a file unrelated to its task and used it.

The mitigation:

  • Install only official connectors and plugins from known vendors and directories.
  • A skill is mostly plain text: Open it and read it before you enable it, and leave it alone if it contains scripts you don’t understand.
  • Keep the number of plugins, skills and connectors small, and remove what you no longer use.
  • Don’t keep passwords, API keys or access tokens in files the agent can read.
  • If an agent needs a login, give it a separate account with the minimum rights instead of your own.
  • Disconnect connectors and remove folder access when a project ends.

7. Privacy and client data (GDPR)

When an agent works in your inbox or your folders, it also works with other people’s data: sources, interviewees, customers, freelancers, colleagues. As soon as that includes personally identifiable information, data protection law applies (like GDPR in the EU).

Two things are easy to overlook. The first is where the data goes. Whatever an agent reads is processed by the provider’s AI model. For Cowork, Anthropic states that sessions run on its servers, so local files the agent opens are processed there. In this moment you are transferring data to a third party. If this includes information protected by law, you have to make sure this third party follows these laws when processing it. There’s an ongoing legal dispute in Europe over the question if US companies fulfill this with regard to GDPR or not.

The second is your plan. The big AI providers follow a similar pattern: On individual accounts, your chats are by default used for training of future AI models unless you switch that off. This is true for paid plans like Plus or Pro as well. Business plans (Team, Enterprise, Workspace and similar) generally default to no training and back that up with a contract. At the same time, opting out doesn’t mean nothing is stored: OpenAI, for example, keeps conversations for 30 days in any case.

Tip: My article Are my chats used for AI training? compares the settings of ChatGPT, Gemini, Claude, Copilot and Mistral.

The BSI advises using agent services only from providers that are transparent about their data protection practices. I will cover this question in more detail in a future article.

The mitigation:

  • Find out what your plan does with your data: training, retention, and whether a data processing agreement applies. For client or source data, a business plan with such an agreement is the safer choice.
  • Keep client contracts, source contacts and embargoed or confidential material out of the agent’s working folder, unless your plan and your contracts allow it. Check your NDAs and client agreements: Some restrict the use of third-party AI tools.
  • Share only what the task needs, and anonymize where you can.
  • Check whether the tool has a memory function and what it keeps between sessions.
  • If you handle personal data professionally, ask your data protection officer or a lawyer before you connect an agent to it. This topic is more complicated than it might seem at first glance. “Everybody is doing it” is not a good defense in case of a lawsuit.

8. Garbage in, garbage out

An agent works with what it finds. If your folder holds three versions of the same brief, last year’s price list next to this year’s, or brand guidelines that contradict each other, the agent can’t tell which one is current. It picks one and reports back in the same confident tone.

Specialists have known this effect for many years, long before AI, and coined it “garbage in, garbage out”.

In other words: If you want better and more reliable output, you need to take care of your input first.

The mitigation:

  • Make cleanup its own step, before the real task.
  • Let AI list duplicates, outdated versions and contradictions, and propose what to archive. To do this right, see risk 2.
  • Keep one folder with current, approved material, and move old versions to an archive the agent can’t reach. The folder structure from my Cowork article is a good starting point.
  • Tell the agent which sources count, for example in its instructions file: “Use only the files in the folder Current.”
  • Relax oversight only after the inputs are clean and trusted.

9. Losing your own edge: human in charge, not just in the loop

“Human in the loop” is the standard answer to AI risks: A person confirms what the artificial assistant does. But this is only the bare minimum. Even worse: The human in question often gets complacent or is pushed by their superiors to get more done. After all, AI is supposed to increase our productivity, right?

Anthropic’s own data on its coding agent Claude Code shows users approving 97% of permission prompts. In a study with 1,053 paid testers, Anthropic swapped one prompt for a clearly dangerous command. The testers caught it in only 13.6% of cases.

There’s a second, slower effect. A survey of 319 knowledge workers by Microsoft Research and Carnegie Mellon University found that people with more confidence in AI reported less critical thinking, while people with more confidence in their own skills reported more. The data is self-reported, so treat it with care. But in content work, this hits twice: The facts have to be right, and your knowledge, creativity and judgment are what readers come for.

So the goal is not just a human in the loop, but a human in charge: someone who decides what goes out, owns the result and still knows how the work is done.

The mitigation:

  • Decide up front which tasks are “human in charge” (e.g., anything published, sent externally or carrying facts) and which are “human reviews” (e.g., drafts, sorting, research notes).
  • Approve the plan before the agent starts and check the result at the end, instead of clicking through every single step.
  • If you notice yourself approving without reading, narrow the task or the access instead of clicking faster.
  • Keep doing some core tasks yourself on purpose, such as research, first drafts and interviews.

Conclusion

I hope this article doesn’t read as alarmist. AI agents are genuinely useful tools. Many of the incidents mentioned happened in their early days. Users and vendors have since learned more about how to minimize the risks and avoid bad behavior.

At the same time, I want to be honest: “minimizing” and “avoiding” risks doesn’t mean that they are gone.

My recommendation is to treat AI agents like a very capable, but also very green intern in your company. You wouldn’t give them the master key to all systems on day one — or ever. Best practices around IT security apply to AI agents even more than to humans, because of how fast they can cause damage.

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×