No. 08Build or buy5 min readOCTYN
AI agent vs workflow automation for small business: most of the work wants rules.
For a small business weighing an AI agent against workflow automation, most of the work belongs to workflow automation: rules with tests beside them, so the same input decides the same way every time. Keep the model to the narrow parts that need judgement, like reading a website or drafting a first email, and check what it returns.
Ask whoever answers new enquiries how they decide which ones get a quote that day, and the reply is usually a rule they have never written down. Anything from overseas goes on the other price list. A repeat customer skips the form, and a large order waits for the owner to see it first. The person at the desk applies all of this the same way on most days, and the days they do not tend to be the days somebody else was covering.
That is a workflow, even when it lives in one person's head. A workflow follows a path somebody drew in advance, and an AI agent is given a goal and a set of actions and chooses its own path as it goes. The pitch for agents in a small business usually comes down to handing the whole desk over, so a model reads each enquiry and decides what happens next. Most of what that desk does, though, is apply decisions that were settled long ago. Ask a model to make a settled decision twice and it can give two answers.
The split that holds up is narrower than agent or automation. Decisions that should come out the same way every time go to rules, each with its own tests. The model gets the parts that need reading or writing, and nothing else. OCTYN found where that line sits on a client's outreach system, and the clearest case was a dull question about where a company is based.
The question of where a company is
An independent craft retailer sells its handmade work to brands. OCTYN built the system that finds those brands and writes to them, and runs it every day. Before a rebuild on 29 July 2026, an earlier pipeline left two decisions to a model. Geography was one of them. Leads for the home market and leads from abroad landed in the wrong workspace, because where a company sat had been treated as a judgement call.
That sounds like a filing problem, and it is a little worse than one. The market is worked out early in the overnight run, ahead of the product match and the draft, so a wrong answer there is carried into every step after it. The case study gives the volume since 25 August as 25 first emails a day, 15 for the home market and 10 abroad, and which of those two piles a company joins depends on that answer.
After the rebuild, a gate decides home market or abroad from four signals: currency, phone prefix, tax registration and domain. None of the four needs interpreting. In the case study's own words, geography stopped being a judgement, and the first failure went away with it.
The emails also invented details the brands did not have, which was the other decision going wrong. Drafting stayed with a model after the rebuild.
Why the decisions want tests beside them
Each gate on the retailer's system is plain code with tests beside it. A test, for an owner's purposes, is a worked example with the right answer already written next to it. You feed the gate a company whose market is already known and check that it gives that answer, and the example stays, so every later change to the rule gets checked against it as well. On 29 July 2026 the retailer's system had 51 passing unit tests covering voice, cadence, geography, the catalogue and the email kit.
This is the whole argument for rules on the decisions. The same input decides the same way every time, and a change that would break that shows up as a failing test instead of a misfiled lead. A model asked to use judgement on the same lead twice can give two answers, and the second arrives in exactly the same tone as the first.
Tests also make a rule safe to change. A new geography signal on the retailer's system is one rule with a test. A new banned phrase is a line in a list. Changing a model's behaviour means rewording its instructions and finding out, one case at a time, what else moved.
Rules also make the order of the steps readable. That is most of what orchestration comes down to in practice: which step runs after which, and what happens when one of them fails.
What the model is still for
None of this leaves the model idle. On the retailer's system the heavy part runs overnight in a working session, which finds brands that fit, reads their sites, scores each one against the retailer's criteria, matches a product from the catalogue to what the brand makes and drafts the email. Working out what a brand makes from its own website is a reading job, and a model is decent at it. So is writing a first email that does not sound like a mail merge.
The difference is in what a second answer costs. Two drafts of the same first email can both be fine. Two answers about which market a company is in mean one of them is wrong, and nothing in the wording says which.
Even the reading and the writing have rules checking what comes out of them. Every draft passes a hard check on the retailer's voice before it can be queued. OCTYN also keeps a rule that nothing a model returns is trusted unless the text is literally present in the page it was given, which is what keeps an invented email address out of a send queue. The send itself has no model in it at all, and that story has its own post.
The same thinking runs through OCTYN's own planning system, Brain, which knows the state of every project and is not allowed to build or deploy anything. Its case study is plain about the alternative: letting it act would make it a different system, and the safer move is to let something else hold the permissions.
Sorting your own process
Take one process the business runs every week and write each step on its own line. Go down the page and mark every line where two careful people, given the same case, ought to reach the same answer. Those lines are rules. Each one wants writing down properly with a few worked examples beside it, so the rule gets checked whenever somebody changes it.
The lines left over tend to be reading or writing, such as making sense of a message written in someone's own words, or drafting the reply to it. That is the narrow part a model is for, and even there something should check what it hands back before anything leaves. Anything that cannot be undone goes to plain code a person can watch.
If most of the page turns out to be rules, the agent question has mostly answered itself, and what remains is a smaller question about which lines get a model. On the craft retailer's system, the model still drafts every first email. It no longer gets a say in where the company is.
Bring the version of this that is about your company.
The first conversation is with the people who would build it. If there is a fit, the Consult comes next.