# How to scope an AI project so it ships: write down what it will never do

To scope an AI project so it ships, tie it to one piece of work your team repeats every week, with your weekly hours on it, and then write down what the first version will never do. Feature lists grow with every meeting. The stated limits keep it small enough to finish and safe enough to leave running.

Date: 2026-09-29
Reading time: 5 min
Author: OCTYN
No. 05

The first scope for an AI project usually gets written in the week everyone likes the idea most, and it reads that way. The system will read the inbox, draft the replies, update the CRM, flag the client who sounds unhappy and write the Friday report. Each line is reasonable on its own. None of them says where the thing ends, so every meeting adds one more, and a few months later the project lives in a document people mention in the past tense.

Most of that sprawl comes from a scope that only lists features. A feature list has no natural end. There is always one more thing the system could plausibly do, and somebody in the room will think of it. What keeps a first version small enough to finish is a second list, usually missing, of what it will never do, and that same list is what makes it safe to leave running once it works.

## One piece of work, with your hours on it

Pick the work before you pick what the system does. It should be something your team does every week, where a person has to rebuild the same context by hand before they can act: scrolling back through a client's WhatsApp thread before a call, or opening the inbox and a sheet side by side to find out what was promised.

Then count it. The arithmetic OCTYN runs on a call is clients, times touchpoints per cycle, times the minutes spent rebuilding context before each touchpoint, which comes out as hours per week. OCTYN asks for those three inputs and does not fill them in for you. The number is yours, and it is usually larger than the person who owns it expected.

Say, as a made up example, that you have thirty clients, each gets two touchpoints a week, and each touchpoint costs ten minutes of reading back first. That is ten hours a week.



Those ten hours do two jobs for the scope. They tell you when the first version is finished, which is when the number has gone down and stayed down. Every idea that arrives in a meeting now has a test as well: if a feature does nothing to those hours, it waits for a later version, however good it sounds.

Write the number down with the date you measured it. That is a [baseline measurement](/glossary/baseline-measurement), and without one, "the system helped" is an opinion held mostly by whoever built the system. The [page on what a system gives back](/what-it-costs) runs the same arithmetic on your own figures, if you would rather not do it in the margin of a meeting agenda.

## The line that says what it will never do

OCTYN runs its own portfolio on a planning system called Brain. It reads the notes on every project, plans the next piece of work and writes what it learned back into those notes. Building and deploying are off the list, permanently, and that one rule is why it is safe to ask it anything. If Brain is wrong, somebody reads a wrong answer and ignores it.

The [Brain case study](/casestudies/octyn-brain) is plain about the alternative. Every property of the system depends on it not acting, so letting it act would make a different system that needs its own containment, and the safer move is to keep Brain as it is and let something else hold the permissions. That is worth copying into your own scope as a sentence, with your system's name in it.



The smallest thing OCTYN runs has a line like it. A relay takes every error from every project and puts it on a phone as one readable line, with the full trace attached. Once an error has gone out, the relay keeps no record of it. Somebody will eventually want it to remember: to mark an alert as seen, or to escalate one that sat unanswered. Both need state, and [its case study](/casestudies/sentry-telegram-relay) says adding state is the moment this stops being a relay. Replying to an alert to restart something goes further again, because a chat group that can act on production has become a credential.

In both systems the rest of the design rests on the limit. Take the limit away and you are drawing a new system, which is fine, as long as everyone knows that is what is being asked for.

For your own project, the limits tend to fall out of one question: what would be expensive if the system got it wrong? Anything irreversible goes near the top of the list. Two cascading deletes once removed data at OCTYN that had been expensive to collect, without a word to anyone, and both now stop with a message saying what they are blocking.

## What the limits decide after launch

An independent craft retailer's outreach system, which OCTYN built and runs every day, shows what a limit is worth once a thing is live. Sending is the part that can do damage, so sending has no intelligence in it. That rule came from a failure. A sender driven by a model was tried first, and it errored, sent the wrong thing, then reported a send that had not happened. The replacement checks for due emails every minute, marks each one as it goes and confirms every send.

Limits also sort the requests that arrive later. Most changes to that system are settings, and the targeting criteria sit in a file per company. A new geography signal is one rule with a test. A rebuild is kept for a single case, a daily volume far past 25, since the part that reasons is deliberately slow. So when somebody asks for more, the answer is one sentence about which kind of change it is, and the scope stays the size it was.

The same case study lists what OCTYN would do differently, and the first thing on it is the daily summary, because without it the operator has nothing to read. Nothing goes out in the morning, so the owner can read the day's queue and pull whatever they dislike before it moves. For most first versions, that page belongs in the scope from day one.

The first page of a scope, then, before any features:

- the one piece of work, named the way your team names it
- its hours per week, on your numbers, with the date you counted them
- what the first version will never do
- what would count as a different system
- the page a person reads every day

The [Consult](/consult) produces a version of this in writing, as a map, a baseline, the architecture, a roadmap and a fixed quote, and it is yours whether or not OCTYN builds anything.

The relay still remembers nothing it has sent, and it is still the system OCTYN would miss first.

## Related

**Systems this bears on**

- [Company Brain](https://octyn.co/systems/company-brain)

**Case studies cited**

- [Brain](https://octyn.co/casestudies/octyn-brain)
- [Error relay](https://octyn.co/casestudies/sentry-telegram-relay)
- [Brand sales for an independent craft retailer](https://octyn.co/casestudies/craft-retail-brand-sales)

**Terms used here**

- [Baseline measurement](https://octyn.co/glossary/baseline-measurement)

**The longer argument**

- [What it costs to change an AI workflow six months later](https://octyn.co/compare/cost-to-change-ai-workflow-six-months)

**Read next**

- [Build vs buy AI tools for operations: look for the side spreadsheet](https://octyn.co/blog/build-vs-buy-ai-tools-for-operations-look-for-the-side-spreadsheet)
- [What should a founder automate first? Start with the catching up](https://octyn.co/blog/what-should-a-founder-automate-first)
- [Should an AI agent send your emails? Ours did, once](https://octyn.co/blog/should-an-ai-agent-send-your-emails)
- [How to stop rebuilding client context before every call](https://octyn.co/blog/how-to-stop-rebuilding-client-context-before-every-call)

More on build or buy: [Build or buy](https://octyn.co/blog/topic/custom-ai-systems)

Canonical: https://octyn.co/blog/how-to-scope-an-ai-project-so-it-ships-write-down-what-it-will-never-do
