How-To Guide

How to Run an AI Pilot Project in Your Small Business (Without Wasting 90 Days)

A practical 30-day method for testing AI on one workflow — with a baseline, success criteria, and a decision at the end instead of a vague feeling

B Biztrategy Published 21 September 2026 · 8 min read
Close-up of a whiteboard with colorful sticky notes for task organization and planning.

Most small businesses do not fail at AI because they picked the wrong tool. They fail because they never ran a proper test. Someone buys a subscription, a few people poke at it for a fortnight, nothing obviously improves, and six months later the honest summary is "we tried AI and it did not really stick." There is no baseline, no success criteria, and no decision — just a direct debit nobody wants to be the one to cancel.

A pilot project fixes that. It is a small, time-boxed, measured test of AI on one workflow, with a written definition of what success looks like and a date on which you decide. This guide walks through the five steps, with the numbers and the prompts you need at each one. It takes about two hours to set up and 30 days to run.

What an AI pilot actually is (and what it is not)

A pilot is not "letting the team play with ChatGPT for a month." It has four properties that distinguish it from ordinary experimentation:

  • One workflow. Not one department, not one tool — one repeatable task with a clear start and finish.
  • A measured baseline. You know what that task costs today, in minutes and euros, before you change anything.
  • Written criteria. The numbers that mean "scale this" and the numbers that mean "stop" are agreed in advance, in writing.
  • A decision date. On day 30 somebody decides. Pilots that drift into month four are not pilots; they are unmanaged spending.

The reason to be strict about this is that AI results are genuinely easy to misread. It feels faster. Everyone says it is helping. Then you measure and discover the drafting time fell from 40 minutes to 15, but the editing time rose from 10 to 25, and the net saving is four minutes on a task you do twice a week. That is a real and common outcome, and it is far better to learn it in 30 days for €60 than in a year for €6,000.

Step 1: pick a workflow that is boring, frequent, and measurable

The instinct is to pilot AI on the most exciting thing you do. Resist it. The best first pilot is deliberately unglamorous, because unglamorous tasks are the ones where you can actually tell whether something improved.

Score your candidate workflows against four criteria:

  1. Frequency. It should happen at least weekly, ideally daily. A monthly task gives you three data points in a 30-day pilot — not enough to conclude anything.
  2. Measurability. You can count it: minutes per item, items per week, error rate, response time.
  3. Low blast radius. If the AI gets it wrong, you notice before a client does. Internal drafts, first-pass research, and meeting summaries qualify. Sending unreviewed output to customers does not.
  4. Owned by one person. A workflow with five owners becomes a coordination project, and you will end up measuring the coordination, not the AI.

Workflows that consistently make good first pilots for SMBs: turning meeting recordings into action items; drafting first-pass proposal sections from a brief; categorising and routing inbound enquiries; summarising supplier or client documents; writing product descriptions or job adverts; and producing weekly reporting commentary from a spreadsheet export.

Bad first pilots: anything customer-facing and unsupervised, anything involving regulated advice, and anything where the current process is undocumented.

Step 2: write down the numbers before you start

This is the step almost everyone skips, and skipping it is why so many businesses end up with a vague feeling instead of an answer. Spend one week — or one honest afternoon of recollection — establishing a baseline.

For the workflow you chose, record:

  • Volume: how many times per week does this happen?
  • Time per instance: how many minutes, end to end, including review and rework?
  • Who does it: and what is their loaded hourly cost to the business?
  • Quality today: how often does it need a second pass, get sent back, or contain an error?

Multiply the first three and you have your baseline cost. A concrete example: a five-person agency writes 12 proposal drafts a month, each taking 90 minutes of an account manager's time at a loaded cost of €45 an hour. That is 18 hours and roughly €810 a month on proposal drafting. Now you have a number that a 40% improvement can be measured against — €324 a month, against an AI subscription of perhaps €25. That is the kind of arithmetic that settles arguments.

If you want a fuller framework for this, our guide on how to calculate the ROI of AI implementation covers the full cost model, including the hidden costs most people forget.

Step 3: set success and kill criteria in advance

Write two sentences before the pilot starts, and get whoever holds the budget to agree to them.

We will scale this if, by day 30, it saves at least 30% of the baseline time with no increase in error rate. We will stop if it saves less than 15%, or if quality complaints rise at all.

The exact thresholds matter less than having them. Thirty per cent is a reasonable default for a drafting workflow; below about 15% the overhead of managing a new tool eats the gain. Adjust before you see results, not after.

The kill criterion is the important half. Without it, every pilot succeeds, because there is always a reason to give it another month. Pre-committing to a number is what converts an opinion into a decision.

Step 4: run the 30-day pilot

Keep the operating rules simple. One tool, one workflow, one person logging results.

Week 1: set up and calibrate

Build a single reusable prompt rather than improvising each time. A good working prompt for a drafting pilot has five parts — role, context, input, constraints, and format. For example:

"You are an account manager at a five-person marketing agency in Spain. Using the client brief below, write the Approach section of a proposal. Constraints: British English, 250–350 words, no bullet lists, no pricing, reference the client's stated goal at least twice. Format: three paragraphs with a bold one-line summary at the top. Brief: [paste]."

Run it five or six times in week one and refine the constraints each time something comes back wrong. By the end of the week the prompt should be stable and saved somewhere the whole team can reach it.

Weeks 2–3: run it for real and log every instance

A spreadsheet with five columns is enough: date, task, minutes taken, edits required (none / light / heavy), and a free-text note. Thirty seconds per instance — and the log is the entire evidential basis of your decision, so insist on it. Log the failures in detail too: "invented a statistic" and "tone was too American" are different problems with different fixes.

Week 4: stop changing things and measure

Freeze the prompt and the process for the final week so you are measuring a stable system. Then total the log and compare it honestly with your baseline from step 2.

Compute three numbers: time saved per week, the error or rework rate versus baseline, and the true monthly cost including the subscription and the time spent managing the pilot. That third one catches the pilots that look like wins until you count the 90 minutes a week someone spent babysitting them.

Step 5: decide — scale, adjust, or kill

On day 30, one of three things happens.

Scale. You beat the success criterion. Now document the workflow properly, train the rest of the team on the saved prompt, and — crucially — pick the next workflow to pilot. The compounding comes from the second, third, and fourth pilot, not from the first.

Adjust and re-run once. You landed between the kill and success thresholds, and the failure log points at a specific, fixable cause: the wrong tool for the task, a prompt missing context, or a workflow step that should have stayed manual. Fix that one thing and run a second 30-day cycle. One re-run — not three.

Kill it. You missed the kill criterion. Cancel the subscription, write two paragraphs on what you learned, and move to a different workflow. This is a successful pilot. You bought a reliable answer for about €30 and a month of mild inconvenience, and you will not spend the next two years wondering.

Common ways pilots fail

Five patterns account for most disappointing pilots, and all five are avoidable.

Piloting three things at once. When you test AI on proposals, customer service, and social media simultaneously, you cannot tell which result belongs to which workflow, and the team is learning three tools at once. One workflow.

No baseline. Covered above, and worth repeating because it is the single most common failure. Without a baseline, your day-30 conversation is a debate about feelings.

The enthusiast problem. If the pilot runs entirely through the one person on the team who loves AI, you have measured that person's aptitude, not the workflow's. Include at least one ordinary, mildly sceptical user.

Moving the goalposts. Rewriting the success criteria in week three because the results are disappointing. If you feel the urge, that urge is itself the finding.

Quietly extending forever. The pilot that never ends becomes an unmanaged subscription. Put the decision date in the calendar on day one. Our piece on the most common AI mistakes small businesses make covers the wider versions of these traps.

Where pilots fit in a wider AI strategy

A pilot answers a narrow question well: does AI improve this specific task enough to be worth keeping? It does not tell you which workflows to attack in what order, where AI gives you a durable advantage rather than a marginal time saving, or what your business should look like in three years. Those are strategy questions, and running pilots without them means you will optimise a series of small tasks while the important ones go untouched.

The sensible sequence is strategy first, then a queue of pilots that serve it. Our walkthrough on how to create an AI strategy for small business covers the framework, and the AI implementation roadmap template turns it into a sequence you can actually work through.

The bottom line

Pick one boring, frequent, measurable workflow. Write down what it costs today. Agree in writing what would count as a win and what would count as a failure. Run it for 30 days with a log. Then decide, and say so out loud. That is the entire method, and it is the difference between a business that knows what AI does for it and one that has six subscriptions and a shrug. The discipline is not in the technology — it is in the baseline, the kill criterion, and the date in the diary.

Where does your business stand on AI?

Take the free 3-minute AI Readiness Quiz and get a personalised score with your next steps.

Take the Free Quiz →