Skip to content

AI agent guardrails belong on the write, not the decision

How TwinCortex applies AI agent guardrails to a computer-use agent: intent checks on clicks, a mail tool without a send function, and completion gates over artefacts.

AFKzona Group · 7 min read

The short answer

  • AI agent guardrails govern what an agent may change on real systems, and they matter most on writes such as send, delete and pay.
  • Absence outranks permission: a tool without a send function is a stronger control than a rule telling the agent not to send.
  • In the agent's memory, facts fade with disuse and prohibitions never do, and a test proves that every task receives them.
  • An agent's own report never counts as done; completion comes from gates over the files it produced.

TwinCortex's click tool reads the name of the control under the pointer before it presses. If that name matches one of 26 patterns, such as send, publish, delete, pay or go live, the click is refused unless the agent states what it expects the press to do.

TwinCortex is a native Mac app we built as a personal digital twin: more than 17,000 lines of Swift exposing 18 MCP tools that let an AI agent see the screen, read the accessibility tree, click and type. Its AI agent guardrails sit between that agent and every action it cannot take back.

The dangerous part of a computer-use agent is the write, not the decision

A computer-use agent can reason correctly and still do damage, because the harm happens at the moment of input, not the moment of choice. A screenshot, a read of the accessibility tree or a search of the inbox can be repeated at will. A click on Send cannot be repeated or undone.

The typical failure of a computer-use agent is not bad judgement. A click at a coordinate read off a scaled screenshot lands in the toolbar above its target. An instruction meant for one editor window is typed into another. A misplaced click starts a live broadcast. Each of those calls reports success.

The question is not whether the model decided well, but what stands between a correct decision and a wrong press.

A read is safe to retry. A write is not. The guardrail belongs on the write.

A guard has to look at the target, not at the tool call

An effective guardrail inspects what an action will touch, not which tool was called. A name check on the target catches the misplaced press, and removing a capability outright makes whole classes of action impossible. TwinCortex uses both, and gives each agent role only the tools that role needs.

TwinCortex names the target with two lookups: it asks the system which element sits at the point, and it reads the application's own accessibility tree, where the smallest element containing the point wins. So a "Go live" button and the segment beside it each come back with their own name.

Two layers: an intent gate on clicks, and an absence in the mail tool A click request is first resolved to the name of the control under the point, then checked against the irreversible patterns; a non-matching control is pressed, a matching one without a stated intent is refused, and a matching one with an intent is pressed and logged by name, while the mail tool offers no way to send, because it has no send function. click(x, y) from the agent Name target AX, then tree walk Irreversible? 26 patterns Press no match Refuse no intent given Press + log intent stated mail.send No such tool an absence, not a rule
A click on an irreversible control needs a stated intent; sending mail needs nothing, because the mail tool has no send function at all.

Every press of an irreversible control is a stated decision. The reply names the control the click hit, so the log reads "clicked on Delete", not coordinates.

Mail goes further. The mail tool can list and search, and nothing else: there is no send, reply, delete or mark-as-read in the code. It is not a permission that could be misconfigured; it is an absence. A tool that cannot send cannot send the wrong thing. The calendar tool is built the same way: it reads the day and never books, moves or cancels anything.

Roles narrow the rest. The worker agent cannot type or drag, so it cannot type into a form key by key or drag a file somewhere by accident. The director, which briefs other agents, has no hands at all.

Our desktop worker runtime grades capabilities from 0 (local actions) to 4 (money, remote deletion, production). An ungranted capability is a hard deny that human approval cannot override; a granted one above the task's level goes to approval. Anything unknown counts as level 4.

Rules reach every task, and done is proven on the artefacts

TwinCortex's memory treats prohibitions unlike any other memory, and the worker runtime grants "completed" from evidence, never from an agent's word.

Memories in TwinCortex fade. An unconfirmed memory loses 0.030 of its activation a day and expires after 14 days if nothing retrieves it; one that has proved itself loses 0.001. Prohibitions lose nothing. They are never expired, never merged with look-alikes and loaded into every brief whatever the task, and a test checks that two unrelated tasks receive the same number.

The test reads the content of the brief, not its shape: it counts the prohibitions inside, so a well-formed brief is also a complete one.

Evidence is held to the same standard. A screenshot named for a step proves only that a file exists, so "completed" is a state the runtime grants after gates over the artefacts pass; a failed gate sends the task back to running or blocked.

Checkpoints follow the same rule: a resumed task enters an explicit recovering state with its last checkpoint and remaining subtasks, and picks up from there.

Against a manipulated agent, the absent capability holds

Text an agent reads can try to steer it: an email, a web page or a document can carry instructions. TwinCortex answers that with capabilities the agent does not have. The mail tool has no delete, the calendar tool has no write, the worker cannot type, and an ungranted capability stays denied whatever the agent says and whoever approves.

The read path is where such instructions arrive. Anthropic's computer-use documentation warns that the model may follow instructions found in on-screen content, and OWASP lists excessive agency among the top LLM risks. An email saying "delete the mail" arrives through a mail tool that cannot delete anything.

The layers stack. The absences decide what can happen at all, the intent gate turns every irreversible press into a stated decision, and each click is reported by the name of what it hit, so the log reads as a record of decisions.

Evaluating AI agent guardrails: five questions to ask

Five questions separate an agent whose guardrails sit on its writes from one whose guardrails live only in its instructions. Ask them of any agent that will act on your systems, whoever builds it, and expect a mechanism in each answer, not a policy.

QuestionWhat good looks like
Which actions can the agent not take at all?Tools with no write path, such as a mail tool that cannot send, not instructions asking it not to
Does the guard look at the target or at the tool name?It resolves what a click, key or command will touch and refuses irreversible ones without recorded intent or approval
Can a human approval grant something the task was never given?No. Ungranted capabilities stay denied; approval only lifts the level of something already granted
How do you know a rule is still being applied?A test asserts the rules reach the agent on every run, whatever the task, and fails if they vanish or vary by task
What counts as done?Gates over real artefacts: files that exist and outputs checked against the claimed state, never the agent's report

What generalises

Sort every action an agent can take into reads and writes, and spend the guardrail budget on the writes. For each write, ask first whether it needs to exist. Removing a capability is the control that holds even against a manipulated model.

Keep rules and enforcement apart. A prohibition in a prompt is guidance; a missing tool is a fact. Build the facts first, and test that the guidance reaches the agent on every run. If you are still deciding whether a task needs an agent at all, start with that comparison.

If you want an agent acting on your systems with guardrails you can inspect, see how we build AI agents and TwinCortex, or book a call.

Common questions

What are AI agent guardrails?

AI agent guardrails are the checks between an agent's decision and its effect on real systems: which tools exist at all, which actions need a stated intent or human approval, which roles lose which tools, and how finished work is verified. The most important ones sit on writes, such as sending, deleting or paying, because a read can be repeated and a write cannot.

How do you stop an AI agent from sending an email by mistake?

Give it a mail tool with no send function, so sending through that tool is not a permission but an absence. TwinCortex's mail tool lists and searches, and a click on any control named Send, Reply or Publish is refused until the agent states what it expects the press to do. Instructions in a prompt are weaker than both, because models can be steered by content they read.

How do you limit what an AI agent can do on a computer?

Scope it by tool, by role and by level. In TwinCortex the mail and calendar tools read and never write, the worker role has no typing or dragging, and the director has no input tools at all. The worker runtime grades capabilities from 0 for local actions to 4 for money, remote deletion and production: an ungranted capability is a hard deny, a granted one above the task's level waits for human approval, and anything unknown counts as level 4.

How can you prove an AI agent actually finished a task?

Do not accept its report. Declare in advance the artefacts the task must produce and the checks they must pass, and mark the task complete when those checks pass on the real files. TwinCortex's worker runtime works this way: "completed" is a state it grants after the gates pass, and a failed gate sends the task back to running or blocked.

Should an AI agent's memory ever forget things?

Yes, most of it. Facts about abandoned projects should fade, or they crowd out current ones. In TwinCortex, ordinary memories lose activation every day and unconfirmed ones expire after 14 days. Prohibitions are the exception: they never decay, are never merged and are loaded on every task. A rule that only loads when the task happens to mention it is not a rule.

Sources

  1. Anthropic documentation — Computer use tool (security considerations)
  2. OWASP Gen AI Security Project — LLM06:2025 Excessive Agency

Tell us what you need built.

A free 30-minute call with the engineer who would lead it. You leave with a scope outline and a price range.