An independent craft retailer · brand salesRunning

Land brand deals from a short list of the right companies.

An independent craft retailer sells its handmade work to brands, and its outreach is a two-part system where a reasoning session researches and drafts and a plain, rule-bound sender is the only part allowed to put an email in front of a company.

Architecture

The system.

  1. 01FindBrands that fit the retailer’s criteria, from sources that show what they make and sell.
  2. 02ResearchEach brand’s site is read, scored against the criteria and placed in its market by rule.
  3. 03MatchA product from the catalogue is matched to what the brand makes, and the right person found.
  4. 04DraftThe first email is drafted and must pass a check on the retailer’s voice before it can be queued.
  5. 05SendA plain sender with no model in it sends what is due, within daily and per-inbox limits.
The boundary · a rule, not a modelA plain sender with no model in it is the only part allowed to send, within daily and per-inbox limits, and only after the owner’s daily window to read the queue and pull anything.

The client is an independent retailer that makes handmade goods and sells them to brands. OCTYN built its outreach system and runs it every day.

The situation

The system was rebuilt on 2026-07-29. The earlier pipeline left two decisions to a model, and both failed in ways a buyer notices. Leads for the home market and for abroad landed in the wrong workspace, because geography was a judgement call. And the emails invented details the brands did not have.

Why it was hard

The volume is small and the cost of one wrong email is large. A made-up detail, or an email dropped into a deal already under way, costs a relationship, not a percentage point. So every step that runs unwatched had to be safe to run unwatched, and every gate had to be written as a rule, not left to a model's opinion.

What was built

Two parts that never talk to each other directly. They share one database and nothing else.

The heavy part runs at night. It finds brands that fit, reads their sites, scores each one against the retailer's criteria, works out which market it belongs to, matches a product from the catalogue to what the brand makes, finds the right person, drafts the email and schedules it. The reasoning happens in a working session, not a single model call. The light part is always on, on a small server with no browser and no model, and does one thing: it sends what is due.

The gates are plain code with tests beside them. One decides home market or abroad from currency, phone prefix, tax registration and domain. Another is a hard check on the retailer's voice that every draft passes before it can be queued. Geography stopped being a judgement, and the first failure went away with it.

What changed

Under 300 companies contacted, and the retailer landed its brand deals from them. Volume is 25 first emails a day, split 15 home market and 10 abroad since 2026-08-25, under per-inbox limits for new emails and follow-ups. Follow-ups go out three and five days later under a daily cap of 60, and nothing sends at the weekend. As of 2026-07-29, 51 unit tests cover voice, cadence, geography, the catalogue, the email kit and the learning loop.

What happens when it breaks

Sending is the part that can do damage, so sending has no intelligence in it. A model-driven sender was tried first. It errored, sent the wrong thing, then reported a send that had not happened. The replacement is a small service that checks for due emails every minute, marks each one as it goes and confirms every send. The mail credentials are not stored where it runs; it fetches them when it needs them.

Every guard exists because something went wrong on a dated day, and is written into the operating notes beside what it prevents.

On 2026-08-10, 115 follow-ups came due across a weekend and nothing picked them up: the due-check only looked at today, and Monday already held 93 against a cap of 60. The scheduler now moves overflow forward when it creates a follow-up, and moves weekend dates to Monday.

On 2026-08-12, a company had replied naming the person to talk to, a proposal was already with that person, and a rerouting step saw the old shared inbox go quiet and queued a cold first email to a third person at the same company. It was cancelled sixteen minutes before it would have gone. The rule now is that the moment anyone at a company answers, that company belongs to the conversation.

Rollback is per email. A queued email can be edited, moved or cancelled up to its send time, and what sends is the edited version.

Who looks after it

One person, one window a day.

The send time sets the attention cost. Nothing goes out in the morning, so the owner can read the day's queue and pull anything they dislike before it moves. The system does not ask for approval on every email, on purpose: at 25 a day, a per-email approval becomes rubber-stamping within a week.

The recurring cost is judgement on replies, where a person takes over and should. The other is that the sender runs on its own server, so a change to the scheduling rules is not live until it is deployed there.

What changes when requirements change

Most changes are settings. The target criteria are a file per company: sources, weights, thresholds, categories, market. Daily volume, the geography split, inbox limits, send windows and follow-up timing are values. Other companies already run on the same engine with completely different criteria.

Some are small code changes. A new geography signal is one rule with a test. A new banned phrase is a line in a list. A new email layout is an entry in the kit, and the learning loop starts scoring it against the others on its own.

A new channel is larger. Everything here assumes email: threading, bounce detection, inbox limits, the reply check. A new channel reuses the finding, scoring and drafting, and needs its own sending part.

A rebuild is for one case only: a daily volume far past 25, where the reasoning session stops being the right engine, because it is deliberately slow.

What we would do differently

Build the daily summary first, because without it the operator has nothing to read. Check contact quality at the front, because seven leads in ten sat on shared inboxes before anyone counted, and a shared inbox cannot carry a follow-up sequence. And stop trusting a model with an irreversible action sooner.

Demo · synthetic dataOpen the demo

Something like this, for your company?

The first conversation is with the people who would build it. If there is a fit, the Consult comes next.

Map something similar