AI agents vs workflow automation vs custom code: when each fits
Workflow automation for fixed steps, custom code for core rules and volume, an AI agent only where judgement on unstructured input is needed. A decision table.
AFKzona Group · 6 min read
The short answer
- Agents earn their place on emails, documents and free-text requests, where the next step depends on what the text means.
- Most production systems combine all three: code for the rules, automation for the plumbing, an agent for the judgement step.
- Before an agent acts, decide four things: what it may do alone, what needs approval, how it is evaluated, and how a run is undone.
- An evaluation set agreed in advance is a better acceptance test for an agent than a demo.
Use workflow automation (n8n, Make, Zapier) when a process is a fixed sequence of steps between tools. Use custom code when the logic is your core business rules, needs tests or runs at high volume. Use an AI agent only for the step that needs judgement on unstructured input, such as reading an email or a document, and put approvals, evaluations and audit logs around it.
Most systems we build use all three. This article gives a decision table, the costs and risks of each, and the governance an agent needs before it touches real data.
Three options, defined
The three options differ in who decides the next step: you, when you design the flow; a programmer, in code; or a language model, at run time. That single difference drives cost, predictability and risk.
Workflow automation is a tool that runs predefined steps between applications, usually built in a visual editor: when a form is submitted, create a CRM record, send an email. n8n, Make and Zapier are common examples. n8n can also be self-hosted; Make and Zapier are hosted services.
Custom code is software written for your process, kept in version control, tested and deployed like any other application. The logic is explicit and reviewable line by line.
AI agent is software in which a language model decides, step by step, which tools to call and what to do next, based on the input and its instructions. It handles input that does not fit a form, and approvals and evaluations keep each action within bounds.
The decision table
Pick the simplest option that handles the input reliably. If the steps can be drawn in advance, an agent adds cost and risk without adding value. If the input is free text and the decision depends on its meaning, a fixed flow will break and an agent, or an agent step inside a flow, is the right tool.
| Question | Workflow automation | Custom code | AI agent |
|---|---|---|---|
| Best for | Fixed steps between tools with connectors | Core business rules, data integrity, volume | Judgement on unstructured input |
| Input | Structured: forms, webhooks, records | Anything you can specify | Free text, emails, documents |
| Predictability | High: same input, same path | High, and testable | Variable: needs evaluation |
| Build cost | Low for simple flows | Fixed price on a written scope | Moderate, plus evaluation work |
| Running cost | Tool subscription, usually priced by usage | Hosting | Model usage per run, plus hosting |
| Who maintains it | An operations person in a visual editor | Developers | Developers, plus someone who owns the evaluation set |
| Main risk | Sprawl: many flows nobody owns; silent failures | Low: rules are explicit, reviewed and tested | Wrong action taken confidently; prompt injection |
| Testing | Manual runs; limited automated tests | Unit and integration tests in CI | Evaluation set of real cases, re-run on every change |
| Rollback | Re-run or fix by hand | Deploy previous version; database migrations reversible | Undo each action from the run log; previous prompt and model version |
| Good example | New order → invoice → Slack message | Booking engine that must never double-book | Triage incoming emails and draft replies for approval |
When workflow automation fits
Workflow automation fits when the process is a short, stable sequence between tools that already have connectors, the volume is moderate, and the person who owns the process wants to change it without a developer. It is often the simplest way to connect a form, a CRM and an email tool.
It stops fitting when flows multiply and nobody knows which one does what, when a failure goes unnoticed because a step silently returned nothing, or when a flow starts holding business rules that need tests. At that point, move the rules into code and keep the tool for the plumbing, or replace it.
Check where the data goes. A hosted automation tool processes your data on its infrastructure; if data must stay in the EU or on your servers, check the vendor's hosting options or self-host.
When custom code fits
Custom code fits when the logic is the business: pricing rules, availability, invoicing, permissions, anything where a wrong result costs money or trust. Code can be reviewed, tested against a real database and rolled back to a previous version.
Two examples from systems we run. In our rental back office, a concurrency test proves that two customers can never book the same item, because the rule is enforced in the database. Invoices are immutable, with gap-free numbering. Neither guarantee belongs in a visual flow or in a model's judgement.
Code buys predictability: every change is reviewed, tested and reversible.
When an AI agent fits
An AI agent fits when the input is unstructured and choosing the next step requires understanding it: classifying and routing emails, extracting data from documents that do not share a layout, answering questions from your own content, or drafting a reply for a person to approve.
Where the steps are known in advance, automation or code does the job. An agent belongs where each action can be undone and someone owns the evaluation set. Evaluation set is a fixed collection of real inputs with the expected outcome for each, run against every new version of the agent before it goes live.
Keliox, the chat widget we build and run, shows the pattern: it answers visitors from the business's own website, collects contact details for call-backs and routes hard questions to a person.
Governance: what an agent needs before it acts
An agent is safe to run in production when four controls are in place: limited permissions with approvals for writes, an evaluation set as the release gate, a complete log of every run, and a way to undo a run. Decide all four before the first real input, not after the first incident.
OWASP's 2025 list of LLM application risks ranks prompt injection first (LLM01) and includes excessive agency (LLM06), meaning a model has more permissions or autonomy than the task needs. NIST's AI Risk Management Framework organises AI risk work into four functions: govern, map, measure and manage. The controls below apply both.
| Control | What it means in practice |
|---|---|
| Least privilege | The agent can read what it needs and write only through named tools; no general database or admin access |
| Approvals | Any action that sends, pays, deletes or changes a record goes to a person in Slack, Teams or email first |
| Evaluations | An agreed evaluation set must pass before a new prompt, model or tool version goes live |
| Audit log | Every run records input, model version, tool calls, outputs and who approved what |
| Rollback | Each run can be undone from its log; the previous prompt and model version can be restored |
| Budgets | Limits on model usage per run and per day, with alerts |
| Injection defence | Content from emails, documents and web pages is treated as data, never as instructions |
If your agent's use case falls into an area the EU AI Act treats as high-risk, such as employment decisions, stricter obligations apply. Check the Act and its current timeline before you build.
Combining the three
Most production systems combine the options. A typical pattern: automation or code receives an email and stores it; an agent classifies it and drafts a reply; a person approves; code writes the result to the CRM and logs it. The agent does only the judgement step, and everything around it is predictable.
We have built this pattern at several scales: a workflow builder with dozens of step types and approvals in chat tools, a ten-agent pipeline that generates complete product demos, and an autonomous multi-agent framework that writes, tests and reverts its own code.
Where to start
Start with one named workflow, a written description of what the agent may and may not do, and 30 to 50 real examples with the expected outcome. That becomes the evaluation set and the acceptance test.
- What we build: AI agents and workflow automation and AI governance and security
- Services: AI agents and automation from €6,000, integrations from €3,500
- All prices: pricing
- Book a free 30-minute call and bring one process you want to automate.
Common questions
What is the difference between an AI agent and workflow automation?
Workflow automation runs steps you defined in advance: when X happens, do Y, then Z. An AI agent uses a language model to decide which steps to take and which tools to call, based on the input. Automation is predictable and cheap to run; an agent handles messy input, with approvals, evaluations and logs keeping it safe.
When should I use n8n, Make or Zapier instead of custom code?
When the process is a sequence of steps between tools that already have connectors, the volume is moderate, and a person can maintain the flow in a visual editor. Move to custom code when the logic becomes your core business rules, when you need tests and version control, or when volume and error handling outgrow the tool.
How do you make an AI agent safe for business processes?
Give it the permissions its task needs, route write actions through human approval, log every run, and release each new version after it passes an evaluation set. That answers the risk OWASP calls excessive agency: a model holding more permissions or autonomy than the task needs.
How much does an AI agent cost to build?
AFKzona Group builds one named workflow as an AI agent from €6,000, with an evaluation set as the acceptance test. Integrations start from €3,500. We estimate running costs from model usage and volume during scoping and put the fixed build price in writing before work starts.
What is an evaluation set for an AI agent?
A fixed collection of real inputs with the expected outcome for each, run against every new version of the agent before it goes live. Thirty to fifty real examples from one named workflow make a strong start, and the same set serves as the acceptance test.