AI agent guardrails belong on the write, not the decision
How TwinCortex applies AI agent guardrails to a computer-use agent: intent checks on clicks, a mail tool without a send function, and completion gates over artefacts.
AFKzona Group · 7 min read
- The short answer
- The dangerous part of a computer-use agent is the write, not the decision
- A guard has to look at the target, not at the tool call
- Rules reach every task, and done is proven on the artefacts
- Against a manipulated agent, the absent capability holds
- Evaluating AI agent guardrails: five questions to ask
- What generalises
- Questions
The short answer
- AI agent guardrails govern what an agent may change on real systems, and they matter most on writes such as send, delete and pay.
- Absence outranks permission: a tool without a send function is a stronger control than a rule telling the agent not to send.
- In the agent's memory, facts fade with disuse and prohibitions never do, and a test proves that every task receives them.
- An agent's own report never counts as done; completion comes from gates over the files it produced.
TwinCortex's click tool reads the name of the control under the pointer before it presses. If that name matches one of 26 patterns, such as send, publish, delete, pay or go live, the click is refused unless the agent states what it expects the press to do.
TwinCortex is a native Mac app we built as a personal digital twin: more than 17,000 lines of Swift exposing 18 MCP tools that let an AI agent see the screen, read the accessibility tree, click and type. Its AI agent guardrails sit between that agent and every action it cannot take back.
The dangerous part of a computer-use agent is the write, not the decision
A computer-use agent can reason correctly and still do damage, because the harm happens at the moment of input, not the moment of choice. A screenshot, a read of the accessibility tree or a search of the inbox can be repeated at will. A click on Send cannot be repeated or undone.
The typical failure of a computer-use agent is not bad judgement. A click at a coordinate read off a scaled screenshot lands in the toolbar above its target. An instruction meant for one editor window is typed into another. A misplaced click starts a live broadcast. Each of those calls reports success.
The question is not whether the model decided well, but what stands between a correct decision and a wrong press.
A read is safe to retry. A write is not. The guardrail belongs on the write.
A guard has to look at the target, not at the tool call
An effective guardrail inspects what an action will touch, not which tool was called. A name check on the target catches the misplaced press, and removing a capability outright makes whole classes of action impossible. TwinCortex uses both, and gives each agent role only the tools that role needs.
TwinCortex names the target with two lookups: it asks the system which element sits at the point, and it reads the application's own accessibility tree, where the smallest element containing the point wins. So a "Go live" button and the segment beside it each come back with their own name.
Every press of an irreversible control is a stated decision. The reply names the control the click hit, so the log reads "clicked on Delete", not coordinates.
Mail goes further. The mail tool can list and search, and nothing else: there is no send, reply, delete or mark-as-read in the code. It is not a permission that could be misconfigured; it is an absence. A tool that cannot send cannot send the wrong thing. The calendar tool is built the same way: it reads the day and never books, moves or cancels anything.
Roles narrow the rest. The worker agent cannot type or drag, so it cannot type into a form key by key or drag a file somewhere by accident. The director, which briefs other agents, has no hands at all.
Our desktop worker runtime grades capabilities from 0 (local actions) to 4 (money, remote deletion, production). An ungranted capability is a hard deny that human approval cannot override; a granted one above the task's level goes to approval. Anything unknown counts as level 4.
Rules reach every task, and done is proven on the artefacts
TwinCortex's memory treats prohibitions unlike any other memory, and the worker runtime grants "completed" from evidence, never from an agent's word.
Memories in TwinCortex fade. An unconfirmed memory loses 0.030 of its activation a day and expires after 14 days if nothing retrieves it; one that has proved itself loses 0.001. Prohibitions lose nothing. They are never expired, never merged with look-alikes and loaded into every brief whatever the task, and a test checks that two unrelated tasks receive the same number.
The test reads the content of the brief, not its shape: it counts the prohibitions inside, so a well-formed brief is also a complete one.
Evidence is held to the same standard. A screenshot named for a step proves only that a file exists, so "completed" is a state the runtime grants after gates over the artefacts pass; a failed gate sends the task back to running or blocked.
Checkpoints follow the same rule: a resumed task enters an explicit recovering state with its last checkpoint and remaining subtasks, and picks up from there.
Against a manipulated agent, the absent capability holds
Text an agent reads can try to steer it: an email, a web page or a document can carry instructions. TwinCortex answers that with capabilities the agent does not have. The mail tool has no delete, the calendar tool has no write, the worker cannot type, and an ungranted capability stays denied whatever the agent says and whoever approves.
The read path is where such instructions arrive. Anthropic's computer-use documentation warns that the model may follow instructions found in on-screen content, and OWASP lists excessive agency among the top LLM risks. An email saying "delete the mail" arrives through a mail tool that cannot delete anything.
The layers stack. The absences decide what can happen at all, the intent gate turns every irreversible press into a stated decision, and each click is reported by the name of what it hit, so the log reads as a record of decisions.
Evaluating AI agent guardrails: five questions to ask
Five questions separate an agent whose guardrails sit on its writes from one whose guardrails live only in its instructions. Ask them of any agent that will act on your systems, whoever builds it, and expect a mechanism in each answer, not a policy.
| Question | What good looks like |
|---|---|
| Which actions can the agent not take at all? | Tools with no write path, such as a mail tool that cannot send, not instructions asking it not to |
| Does the guard look at the target or at the tool name? | It resolves what a click, key or command will touch and refuses irreversible ones without recorded intent or approval |
| Can a human approval grant something the task was never given? | No. Ungranted capabilities stay denied; approval only lifts the level of something already granted |
| How do you know a rule is still being applied? | A test asserts the rules reach the agent on every run, whatever the task, and fails if they vanish or vary by task |
| What counts as done? | Gates over real artefacts: files that exist and outputs checked against the claimed state, never the agent's report |
What generalises
Sort every action an agent can take into reads and writes, and spend the guardrail budget on the writes. For each write, ask first whether it needs to exist. Removing a capability is the control that holds even against a manipulated model.
Keep rules and enforcement apart. A prohibition in a prompt is guidance; a missing tool is a fact. Build the facts first, and test that the guidance reaches the agent on every run. If you are still deciding whether a task needs an agent at all, start with that comparison.
If you want an agent acting on your systems with guardrails you can inspect, see how we build AI agents and TwinCortex, or book a call.
Common questions
What are AI agent guardrails?
AI agent guardrails are the checks between an agent's decision and its effect on real systems: which tools exist at all, which actions need a stated intent or human approval, which roles lose which tools, and how finished work is verified. The most important ones sit on writes, such as sending, deleting or paying, because a read can be repeated and a write cannot.
How do you stop an AI agent from sending an email by mistake?
Give it a mail tool with no send function, so sending through that tool is not a permission but an absence. TwinCortex's mail tool lists and searches, and a click on any control named Send, Reply or Publish is refused until the agent states what it expects the press to do. Instructions in a prompt are weaker than both, because models can be steered by content they read.
How do you limit what an AI agent can do on a computer?
Scope it by tool, by role and by level. In TwinCortex the mail and calendar tools read and never write, the worker role has no typing or dragging, and the director has no input tools at all. The worker runtime grades capabilities from 0 for local actions to 4 for money, remote deletion and production: an ungranted capability is a hard deny, a granted one above the task's level waits for human approval, and anything unknown counts as level 4.
How can you prove an AI agent actually finished a task?
Do not accept its report. Declare in advance the artefacts the task must produce and the checks they must pass, and mark the task complete when those checks pass on the real files. TwinCortex's worker runtime works this way: "completed" is a state it grants after the gates pass, and a failed gate sends the task back to running or blocked.
Should an AI agent's memory ever forget things?
Yes, most of it. Facts about abandoned projects should fade, or they crowd out current ones. In TwinCortex, ordinary memories lose activation every day and unconfirmed ones expire after 14 days. Prohibitions are the exception: they never decay, are never merged and are loaded on every task. A rule that only loads when the task happens to mention it is not a rule.