- MantraMindAI
- Blog
- AI Agents & Automation
Small-business AI automation: what to automate first
Jai Rao
August 22, 202619 min read
A shape test and a short piece of arithmetic for deciding which tasks in a small business are worth automating, the workflows that repay it, and what must stay manual.
The usual small-business AI project goes like this. Someone reads that AI is now good enough to be useful, signs up for a tool, and then goes looking for something to point it at. Six weeks later there is a half-configured thing that nobody owns, one person quietly stopped using it, and the supplier invoices are still being keyed in by hand on Thursday afternoons. The tool was fine. The choice of work was wrong.
Choosing the work is the whole game, and it is cheap to get right. Tasks that automate well have a recognisable shape, and you can check for that shape in ten minutes without hiring anyone. Once a task passes the shape test, a short piece of arithmetic tells you whether to do it. What follows is the test, the arithmetic with two worked examples — one that says go and one that says stop — the handful of workflows that genuinely repay the effort at five to fifty people, and the short list of things that must never run without a person signing off.
Pick work by its shape, not by how clever it looks
Four properties. A task needs all four, not three.
It happens often. Ten times a week is the floor; a hundred times a week is where this gets interesting. Frequency is what pays back the setup effort, and it is also what gives you enough examples to tell whether the thing is working. A task you do twice a month will never generate enough evidence to trust, and you will have forgotten how it behaves by the time it runs again.
It follows a pattern you can describe out loud. Try saying the rule to a new starter in three sentences. "Take the supplier name, the invoice number, the date and the total off the PDF and put them in these four fields" is a pattern. "Work out what this customer actually wants and how much we should charge them" is not — not because it is hard for software, but because you cannot write down what correct looks like, which means you cannot check the output or tell whether it improved.
A mistake surfaces quickly and costs little to fix. This is the property people skip, and it is the one that hurts. A wrong total on a keyed invoice gets caught at reconciliation and takes a minute to correct. A wrong delivery date in a quote you already sent surfaces in six weeks, in front of a customer, and costs you the goodwill and possibly the job. Same amount of automation, wildly different risk, and the difference is entirely about how fast the error becomes visible.
Someone does it by hand today and dislikes doing it. This sounds like the soft criterion. It is the most reliable one. Work people resent is usually repetitive rather than skilled, which is exactly the shape you want. It also hands you an ally: the person who hates the task will tell you the moment the output is wrong, whereas someone whose judgement you are trying to replace will find reasons the output is fine.
Notice which tasks fail. The exciting ones — pricing, hiring decisions, anything described as "strategic" — usually fail the second and third tests together. They have no describable pattern and their mistakes stay hidden for months. The boring ones pass easily. That is the whole insight, and it is why the least glamorous candidate on your list is usually the right one to start with.
Running the numbers on one workflow
Before spending anything, write the sum on one side of a page. Numbers below are plain — read them in whatever currency you pay wages in.
What you save is hours per week multiplied by a fully loaded hourly cost. Fully loaded is not the wage. It is the wage plus employer taxes plus the software and the desk and the share of everything else that person needs — call it a third more than the hourly wage, and use the real figure if you have it. What you pay is three separate things that people habitually collapse into one: a one-off setup cost in someone's time, a per-request running cost, and the ongoing cost of checking the output.
That third item is where most projects die, so say it plainly: automation replaces doing with checking. It only pays when checking is much faster than doing. If a person has to read every output as carefully as they would have done the work themselves, you have saved nothing and added a monthly bill. The good workflows have an enormous gap between the two — glancing at four extracted fields takes five seconds, typing them takes ninety. The bad ones have almost no gap at all.
Where the sum says go
A twelve-person building supplies firm receives about two hundred supplier invoices a week as PDF attachments. The bookkeeper keys supplier, invoice number, date, net, tax and total into the accounting system. Timed rather than guessed, it takes her six hours a week. Fully loaded, her time costs 25 an hour, so the task costs 150 a week, a little under 8,000 a year.
Automated, extraction pulls the six fields and posts them as draft entries. Per-invoice cost for a short extraction like this is small — assume 0.02 until you have a week of real usage to check it against, which puts running cost at 4 a week. Setup is two days of a competent person's time wiring it to the accounting system and getting the field mapping right: call it 700 once. Afterwards the bookkeeper reviews the flagged entries and spot-checks the rest, which measures out at an hour a week, or 25.
So 150 a week becomes 29 a week. That is 121 saved weekly, and the 700 setup pays for itself in six weeks. Budget an hour a month for maintenance, because a supplier will redesign their invoice template and accuracy on that supplier alone will drop. Even with that, the annual picture is strongly positive, and — this matters more than the money at twelve people — six hours a week comes back to a person who has other work worth doing.
Where the sum says no, and people automate anyway
The same firm quotes bespoke fabrication jobs. The owner does about eight quotes a month, forty-five minutes each: reading the drawing, checking material prices, judging site access, deciding how much he wants the work. That is six hours a month, an hour and a half a week. His time is worth 40 loaded, so the task costs 60 a week. Tempting — it looks like the same order of saving as the invoices.
Run it against the shape test and it fails three of the four. Eight a month is thin volume. The pattern is not describable: every job differs, and the part that takes the time is exactly the part that is judgement. And a mistake is both expensive and invisible — quote 15% low and you carry a loss-making job for six months; quote high and you lose the work without ever finding out why. There is no error signal at all, which means you cannot tell whether the automation is any good.
The arithmetic then finishes the argument. Setup is not two days, because the system needs material rates, labour assumptions and the history of what you have charged before: realistically three to four weeks, and the pricing tables need maintaining forever. And because a quote is a document a customer can hold you to, the owner will read every line of every quote regardless. Checking a quote carefully takes most of the time that writing it did. Best case he saves half an hour a week — 20 — against several thousand up front and a permanent maintenance job. It does not pay back this decade, and the failure mode is silent. This is the one to walk away from, and it is the kind of project that gets funded because it sounds impressive in a meeting.
| Question | Keying supplier invoices | Pricing a bespoke job |
|---|---|---|
| How often | 200 a week | 8 a month |
| Describable rule | Yes — six named fields | No — it is judgement |
| When a mistake shows up | At reconciliation, days | Months later, or never |
| Cost to fix a mistake | A minute of typing | The margin on the job |
| Checking vs doing | 5 seconds vs 90 | Nearly the same |
| Setup | Two days | Three to four weeks |
| Verdict | Do it this month | Do not do it |
Inbound enquiries: triage before drafting
Two different jobs get mashed together under "handle the inbox", and they carry very different risk. Triage means reading an incoming message and turning it into a record: what kind of enquiry it is, who it is from, what it refers to, how urgent it is, who should pick it up. Drafting means writing the reply. Triage is boring, safe, and worth more than people expect, because it converts an undifferentiated inbox into a queue you can count, route and measure. Do triage first and live with it for a while before you let anything write.
Here is a triage instruction that does real work. The point of it is not the wording; it is the fixed field list, the closed set of categories, and the two explicit refusals at the bottom.
You turn one customer email into one JSON record. Output JSON only.Fields: customer_name string, from the signature or the From line; "" if absent order_ref string, an order or invoice number if one appears; else "" category one of: quote_request, order_status, complaint, returns, billing, other urgency one of: normal, urgent urgent only if the sender names a deadline within 48 hours summary one sentence, under 25 words, no greeting, no sign-off needs_human true if the email mentions money owed, cancellation, legal action or the press, or if category is "other"Rules: Never guess an order_ref. If a field is not in the email, leave it empty. Do not write a reply.EMAIL:"""{{email_body}}"""Every enquiry now becomes a row with a category and a routing decision, and the two refusals do the safety work: an empty order_ref is honest where an invented one is a wrong lookup, and needs_human plus the other category is the escape hatch for everything you did not anticipate. Keep the original email attached to the record so anyone can check the extraction against the source in one click.
What good looks like: your categories match how the team already works, not a generic list; the large majority of enquiries land in a real category and the remainder land in other and actually get read; response times drop because nothing sits unnoticed. What to watch: category drift, where a new product line generates enquiries with no home and they all silently pile into other; invented references; and the temptation, after one good week, to switch drafts to auto-send. Keep drafts in the drafts folder. A human pressing send costs two seconds and removes almost all of the risk.
Invoices, receipts and forms: getting fields out of paper
This is the most reliable win available to a small business, because the task is narrow, the volume is real and the output is checkable in seconds. Anything that arrives as a document and ends up as fields in a system qualifies: supplier invoices, expense receipts, timesheets, delivery notes, paper enquiry forms, insurance renewals.
The thing that makes this genuinely safe is that you can check most of the work with plain arithmetic rather than judgement. Net plus tax should equal total. Line items should sum to net. An invoice date should not be in the future, and an invoice number you have already posted is a duplicate. Run those checks automatically and route only the failures to a person. Where you have a purchase order, match the total against what you expected and flag the gap. Deterministic checks like these catch a large share of extraction errors for free, which is what turns "review everything" into "review the eleven that failed a check".
What good looks like: a stack of documents in, draft entries out, a small flagged pile, and a person clearing the pile in minutes rather than hours. What to watch: track accuracy per supplier rather than overall, because the failure mode is one supplier redesigning their template while your average stays reassuring; photographs taken at an angle in a van and handwritten receipts will always be worse than clean PDFs, so send those down the manual path deliberately; and never let an extracted total trigger a payment. Extraction fills the form. A person still presses pay.
First-line answers from your own written policies
There is a precondition nobody mentions in the sales pitch: this only works if your policies are actually written down. If the returns window lives in the owner's head and the delivery promise depends on who is asking, there is nothing to answer from, and the model will fill the gap with something plausible and wrong. If that is your situation, the valuable project is not automation at all — it is a week spent writing down the twenty answers your team gives most often. Do that and you have improved the business whether or not you automate anything afterwards.
With the documents in place, the pattern is simple: the assistant answers only from those documents, shows which document and which line it used, and hands over to a person for anything else. The traceability is not decoration. It is how you audit the thing cheaply, and how a member of staff can tell in three seconds whether an answer was legitimate.
What good looks like: it declines confidently and often, and the handover is a real handover rather than a dead end — a named queue with a response time, not "please try again later". A high decline rate on your first pass is a good sign, not a failure; it usually means the documents have gaps worth filling. What to watch: answers that sound like your policy but are not in any document; stale content, which is the quiet killer, because a policy that changed in March will keep being answered as it was in February until someone updates the source; and repeat contacts, which are the honest measure — count how many people came back with the same question after being answered, not how many conversations ended without a human.
Meeting notes and listing copy: the two quiet wins
Notes into actions with a name attached
Automatic meeting summaries are oversold and undervalued at the same time. Oversold because a summary nobody reads is not worth a subscription. Undervalued because the useful output is not the summary — it is the list of decisions and actions, each with a person's name and a date, dropped into wherever your team actually tracks work. That is the part that changes behaviour, because the failure it fixes is real: things agreed in a meeting that nobody wrote down.
What good looks like: every action has an owner and a date, or is explicitly marked unassigned so someone assigns it before the meeting ends. What to watch: confident invention of an owner, which is worse than leaving it blank, so instruct it to leave the owner empty rather than guess; consent, because recording a call has rules and clients notice; and the fact that a transcript of a customer conversation is a customer record — it lives under the same retention and access rules as everything else you hold about them, and it should not sit in a personal folder forever.
Copy at volume, without publishing it blind
Four hundred product listings that need a description each is a textbook fit: high volume, a describable pattern, and a mistake that is visible on the page and cheap to fix. The same applies to rewriting inherited supplier copy, generating variants for different channels, or filling in the short descriptions your catalogue has been missing for two years.
What good looks like: descriptions built strictly from the spec fields you supply, in your tone, with a person reviewing a sample of every batch and every item in the top-selling tail. What to watch: claims, which is the only serious risk here. Never let generated copy assert something that is not in the source data — waterproof, hypoallergenic, certified, compatible with, suitable for children. List the fields it may use and forbid everything else. Watch out too for four hundred descriptions that read identically to a customer comparing two of them, and judge the result by what happens to search traffic and conversion over the following weeks rather than by how good the copy looks the day you publish it.
Nothing sends, commits, or pays without a person's name on it
Four families of action stay behind a human approval, permanently, regardless of how well the automation has been performing.
- Anything that moves money. Payments, refunds, credit notes, payroll, changing bank details on a supplier record. Automation prepares the payment run; a person approves it.
- Anything that commits you to a customer. Prices, discounts, delivery dates, contract terms, warranty decisions — anything in writing that a customer could reasonably hold you to. A generated reply that promises Thursday is a promise.
- Anything touching employment, health, or law. Hiring and firing, disciplinary notes, references, sickness or medical information, tax filings, regulatory returns, insurance claims. The consequences are not commercial, and confident-sounding wrongness is expensive.
- Anything hard to reverse. Deleting records, emailing your whole list, publishing to your site, cancelling a supplier account, sending a legal notice. The test is simple: if undoing it means phoning people and apologising, it needs a signature.
The mechanism costs almost nothing. The automation prepares the action and puts it in a queue; a person reviews and approves; the approval is recorded with who and when. Seconds of human time per item, and it removes nearly all of the tail risk — which is the only risk that can actually hurt a business of this size. A firm of twenty people can absorb a hundred small mistakes. It cannot absorb one automated email to every customer containing the wrong price.
One workflow, a fortnight, beside the person who does it
Take the single highest-scoring candidate from your shape test and do only that one. Not three, not a platform, not a plan for the whole business. One.
Then run it in parallel for two weeks. The automation produces its output, the person keeps doing the job exactly as before, and at the end of each day someone compares the two. It costs a fortnight of a modest amount of duplicated work and it buys you a real error rate instead of a hopeful one. Before day one, write down three things: how many items a week, how long the manual version actually takes when you time it, and what specifically counts as an error for this task. Guessing the baseline is how projects end up unable to prove anything.
Sort every item into three buckets: needed no change, needed a small edit, and wrong in a way that would have caused real harm if it had gone out. The first two numbers tell you whether it saves time. The third decides whether you proceed at all. If that bucket is empty across two weeks and the no-change rate is high, move to running it live with review. If it has anything in it, either add a specific check that catches that class of error and run another fortnight, or stop. Do not average the harmful bucket away against a good overall accuracy figure.
The owner, and the bill nobody puts in the plan
One person's name goes against the workflow. Not "the team" and not the person who happened to set it up if they are leaving in three months. The owner checks a sample regularly, fields complaints about it, and has the authority to switch it off.
Budget the ongoing cost honestly, because it is real and it is always left out: the per-request bill, an hour or two a month of sampling output, and the fix when something upstream changes — a supplier redesigns a form, your product data gets new fields, a provider updates a model and phrasing shifts slightly. None of these are emergencies, and all of them need somebody's afternoon. Write a single page while you still remember how it works: what it does, where it runs, what it costs, who to call, and how to turn it off and go back to the manual process within ten minutes. That last line is the most valuable sentence in the document, and you will be glad of it on the day something behaves oddly at nine in the morning.
What should be different at the end of the month
Judge it the way you would judge hiring a part-timer, with five questions.
- Is someone's week visibly lighter, and can they say what they did instead? If nobody can point to reclaimed hours and name what filled them, nothing happened, regardless of how the demo looked.
- Did a queue get shorter? Enquiries answered the same day, invoices posted within forty-eight hours, backlog counted on the same day of the week each time. Pick the one number you were embarrassed about a month ago and look at it again.
- Did errors go up? Check the specific failure you were worried about, not your general impression. Ask the person who checks the output, not the person who bought the tool.
- What did it cost, all in? Subscriptions, per-request charges, and the hours spent reviewing and fixing. Put that beside the sum you wrote before you started and see whether the sum was honest.
- Would the person doing the work fight to keep it? This is the strongest signal available to you. Staff do not defend tools that waste their time, and they will not tell you directly that one does.
Four yeses and a flat error rate: automate the next workflow, one at a time, with the same fortnight of parallel running. Anything else: turn it off and say so plainly. Stopping a two-week experiment costs two weeks. Running an unowned automation for a year costs you in ways that never appear on the subscription line — wrong records nobody caught, a customer promise nobody made, and a team that has quietly stopped trusting the next thing you bring them.