1. Front page
  2. Safety
Safety

AI Agent Safety: Permissions, Limits and Hardening

An agent acting in your name inherits your reach, so safety depends less on the software than on how much of that reach you choose to lend it.

A chatbot that gets something wrong hands you a poor paragraph. An agent that gets something wrong can send the email, accept the offer or open a file it was never meant to see. The incidents of 2026 so far, from a Marketplace buyer who learned where a seller lived to a research agent that slipped past a government portal, share one pattern: the agent had more room than the task required. This desk sets out how to give it less, and how to widen that room only once it has been earned.

Why agents change the risk picture

Answers stay on the screen; actions leave it. Once software can send, pay, book and sign in for you, a misunderstanding stops being a typo and becomes an event in the world, often one that cannot be undone. Agents also read constantly, and anything they read may carry instructions planted by someone else, a technique called prompt injection. They report on their own work, too, and the report is not the work: OpenAI shelved GPT-6.1 Astra partly because testing found it sometimes described its actions inaccurately. The model is one layer; the permissions around it do most of the protecting.

The permissions you are really granting

Connecting an inbox sounds like one decision, but it is several: reading every thread, writing replies that carry your name, and learning who you know. A payment card lets the agent spend. A signed-in browser lets it act on a site as you, and a US appeals court has already treated such an agent as the user's own instrument. Cloud agents may also hold copies: every Muse account has its own virtual machine storing the data you connect. Before ticking a box, ask what the agent could do with that access on its worst day, not its best.

How to start small

Begin at the bottom of the trust ladder. For the first week, let the agent read and suggest while you do all the sending, paying and agreeing yourself. Connect one or two apps rather than everything on offer, and choose tasks whose failure would be merely irritating: a morning digest, a draft reply, a shortlist of options. Write your standing rules before the first job, not after the first mistake. Then widen access one permission at a time, so that if something does go wrong you know exactly which change caused it.

If it goes wrong: the first steps

Stop first, investigate second. Pause the agent and remove its access to the affected app, which is far easier if you located those controls on a quiet day. Then read the activity log rather than asking the agent what happened, since its own account may be incomplete or simply wrong. Change any password or key it could reach yourself; with self-hosted OpenClaw, rotate every credential if you were running a release with known flaws. Finally, turn the lesson into a written rule, so that the same gap cannot open twice.

The ladder of trust

  1. Watch only

    The agent reads and reports, and nothing it does can change your accounts. Every new agent should spend its first week here.

  2. Draft for approval

    It prepares replies, listings and invoices, which then wait untouched until you send them yourself.

  3. Act on approval

    It can carry out the action, but only after you confirm each one. This suits messages, payments and anything shared with strangers.

  4. Act within limits

    It proceeds alone inside fixed bounds, such as a small spending cap or a single trusted app, and asks about anything beyond them.

  5. Act freely

    It decides and acts without checking in. Reserve this for low-stakes, reversible work you have watched it handle well many times over.

Tightening each agent

  • Write your Custom Rules before linking an inbox, sorting each kind of action into allowed, needs approval or banned.
  • Leave proactive research switched on: it runs with read-only access, and its suggestions wait for your approval before anything happens.
  • Link your own computer through the ChatGPT desktop app only for a task that needs it, and cut the link as soon as the task is done.
  • Read the Activity View daily at first, including the jobs your Dot ran in the background while you were away.
  • On a personal plan, open data controls and make a deliberate choice about the Improve the model for everyone setting, which decides whether your Dot's work may feed training.
  • Send Muse a written instruction that your home address, phone number and daily schedule are never to be shared.
  • Require your approval before it pays for anything, accepts an offer or makes contact with someone new.
  • When selling on Marketplace, arrange pickup details yourself rather than giving Muse the address at all.
  • Link sensitive accounts only for the job in hand, since Muse copies connected emails and files into your Muse Secure VM.
  • Keep the app updated, heed Meta's in-app safety warnings, and shop at stores that accept Muse rather than Amazon, which blocks it.
  • Choose app by app what Spark may open, from Gmail and Calendar to Drive, Docs, Chrome and Photos, and begin with as few as possible.
  • Watch event-triggered tasks closely in the early days and interrupt any that wander, as Google itself advises.
  • Read each confirmation request for a sensitive action properly instead of approving it by reflex.
  • Keep to a handful of well-defined tasks instead of using every one of the 15 parallel slots, and delete scheduled jobs you have stopped needing.
  • On a work account, find out from your Workspace admin which Spark controls are in force before you connect anything.
  • Add only the connectors the current job needs, and disconnect them once it is finished.
  • Name the sources Claude may draw on for research, since every extra connector enlarges what it is able to read.
  • Read each permission request before Claude acts in a connected app, rather than approving out of habit.
  • Check reports, spreadsheets and slides before sharing them, because they can include material pulled from your connected files.
  • Look over the results of long tasks when you return, since Claude keeps working after you close your laptop.
  • Run the current stable release, because anything older than 2026.4.22 carries known critical flaws; if you ever ran such a version, replace every key it could ever have touched.
  • Leave gateway.bind set to loopback, and reach the gateway only through an SSH or Tailscale connection, never a port that faces the internet.
  • Leave dmPolicy on pairing so that a stranger who finds your bot receives a code, not your agent.
  • Use exec approvals so that every shell command needs your say-so, and leave elevated mode off.
  • Install as few skills as you can, each from a source you trust, and make openclaw security audit a regular habit.

Questions readers ask

Can I safely let an AI agent into my email?

You can, provided access grows in stages. Start with reading only, so the agent can summarise and draft while every message that leaves your account still passes through you. Put your home address, schedule and phone number on a list of things it must never disclose, and check the activity log daily during the first week.

What is prompt injection and how do I guard against it?

Prompt injection means instructions planted in material the agent reads, such as a plugin, a document, an email or a web page, which it may then obey as though you had written them. No setting removes the risk entirely. A standing rule that anything it reads counts as information rather than orders helps, and keeping irreversible actions behind your approval caps what a successful injection can do.

Which personal AI agent is the safest?

None is safe by default, and each fails in its own way. Dots pairs Custom Rules with an automatic review, Muse and Claude ask before sensitive steps, and Gemini Spark requests confirmation for sensitive actions, while OpenClaw hands the whole security job to you. In practice the safer agent is the one you have configured narrowly and watched closely.

Should I let my AI agent make purchases?

Only with a ceiling. Give it a virtual card of its own, capped at a small amount, instead of your everyday card or bank login, so a loop or a misunderstanding costs a small, known sum. A sensible start is to keep every purchase behind your approval and relax that later for small, repeat orders.