Do not start with ‘Where can we add AI?’
A better question is: where do we lose time, quality or control, and which kind of automation fits? Some problems need a deterministic rule. Others benefit from one language-model step for classification or summarisation. Only a smaller group needs an agent that plans and uses multiple tools.
OpenAI distinguishes predictable workflow automation, LLM-powered steps that add limited interpretation and agents that adapt actions to context. More autonomy requires more evaluation, monitoring and guardrails.
A prioritisation matrix
| Criterion | Strong candidate | Weak candidate |
|---|---|---|
| Volume | Daily or weekly | Rare one-off case |
| Repeatability | Stable input and expected output | Every case is fundamentally different |
| Data | Available, structured and owned | Scattered with unclear rights |
| Risk | Errors are detectable and reversible | Legal, financial or safety impact |
| Measure | Baseline for time, cost and quality | Success cannot be defined |
| Human review | Clear approval point | No reviewer for a critical outcome |
Decision framework: score the work and choose the top three
Do not choose a process because its demo looks impressive. List 10–15 recurring activities and score each from one to five on three factors: frequency, active human time and cost of error. Multiply the scores: frequency × time × error cost. A higher result signals greater potential value.
Frequency 1 means quarterly and 5 means daily. Time 1 means under 15 minutes each week and 5 means more than five hours. Error cost 1 is an easy correction with no customer impact; 5 means lost revenue, serious rework or risk. This is a management scale, not an accounting formula. Apply it consistently so processes are comparable.
Then apply a risk gate—a plain check that the process is suitable now. Work involving legal, financial, medical or safety decisions does not automatically enter the top three even with a high score. AI may prepare or check material, while a qualified person keeps the decision.
1. List the recurring work
For one week, record tasks that repeat, interrupt the team or involve copying information. Do not start with AI product names.
2. Measure the baseline
For each process, record monthly volume, active time, error rate, performer and what a correct result looks like.
3. Score frequency × time × error cost
Give each factor a 1–5 score and multiply them. Score as a team so one person's frustration does not distort the priority.
4. Check data, risk and ownership
A candidate needs accessible data, repeatable rules, detectable errors and a person responsible for its result.
5. Select three, then automate sequentially
Pilot the first, measure it and stabilise exceptions before building the second and third. Avoid three unfinished automations.
6. Reassess after 30 days
If time or quality has not improved, stop, narrow the scope or use a simpler rule. Do not keep an automation only because it has been built.
| Example process | Frequency | Time | Error cost | Score | Decision |
|---|---|---|---|---|---|
| Transfer data into an invoice | 4 | 4 | 4 | 64 | Top 3; draft plus human approval |
| Follow up a sent proposal | 5 | 4 | 3 | 60 | Top 3; rule plus AI draft |
| Classify support questions | 5 | 3 | 3 | 45 | Top 3; limited categories |
| Review a complex contract | 2 | 3 | 5 | 30 | AI assistance, not automatic decision |
| Quarterly strategy planning | 1 | 5 | 5 | 25 | Do not automate; human judgement |
A high score shows potential value, not permission for autonomy. Risk and human accountability remain a separate gate.
Ten useful starting use cases
- Classify and route incoming enquiries.
- Extract fields from proposals, contracts and forms.
- Summarise meetings and propose tasks for approval.
- Draft customer reports from verified data.
- Check required fields before sales-to-operations handoff.
- Search internal procedures with citations.
- Create a standard project and checklist after a won deal.
- Flag delivery or SLA risk.
- Categorise customer feedback.
- Draft personalised communication from an approved template.
The first AI project should be important enough to create value and narrow enough to measure.
Rule, LLM step or agent?
Do not use an agent for a process that can be expressed in five stable rules. Simpler systems are cheaper, easier to audit and more predictable.
| Type | Use when | Example |
|---|---|---|
| Rule | Conditions and actions are fully known | Create a project when a contract is signed |
| LLM step | One step requires language understanding | Classify an email by topic and priority |
| Agent | A goal requires selecting actions and tools | Research a case, gather data and propose a resolution |
How to design the pilot
Baseline
Measure current time, cost, volume, error rate and quality on a real sample.
Scope
Define the input, output, users, exceptions and explicit exclusions.
Evaluation set
Create normal, difficult and high-risk cases with expected outcomes.
Guardrails
Limit access and actions; require human approval before irreversible steps.
Rollout
Start with a small user group, monitor quality and cost, then expand.
Worked ROI: follow-up after a proposal
A small B2B company sends about 35 proposals each month. An employee checks the CRM and inbox, chooses who needs a reminder, copies text, personalises it and records the result. The work takes five hours a week, or 21.65 hours each month. At a fully loaded labour cost of €22 per hour, the process costs about €476 monthly.
The new workflow uses a rule—not AI—for the predictable part: two days after a proposal, the system checks whether the customer has replied and creates a follow-up task if not. AI uses approved proposal data and a template to suggest a short personalised draft. A person approves the first 50. A reply, rejection or special condition stops the sequence.
An example monthly stack is Make Core at $12, Brevo Starter from $9 and a small AI-usage budget. Plan approximately €25–30 per month depending on volume and exchange rate. If human work falls from five hours to 30 minutes each week, the saving is 4.5 × 4.33 = 19.5 hours per month.
19.5 hours × €22 = €429 of recovered labour capacity. With a €27 tool cost, net value is approximately €402 monthly. If configuration and testing cost €600, simple payback is about one and a half months: €600 ÷ €402. Extra sales from better follow-up are excluded, keeping the model conservative.
This does not prove ROI in advance. Measure proposal volume, time, replies and incorrectly sent messages for four weeks before and after the pilot. Automation succeeds only when it saves active time without reducing quality or trust.
| Measure | Before | After pilot |
|---|---|---|
| Active time | 5 hrs/week | 0.5 hrs/week |
| Monthly labour capacity | €476 | €48 |
| Tool cost | €0 additional | about €27/month |
| Recovered value | — | €429/month |
| Net value | — | about €402/month |
What not to automate yet
Sometimes the best first improvement is a checklist, a better form or a clear rule in an existing system. AI is useful when work requires understanding language, summarising context or proposing an option—not when five fixed conditions solve the task more reliably.
Moving from a working business to a business that works begins with a visible process and an owner. AI follows as a layer for speed and scale, not a substitute for management clarity.
- A process every person performs differently: agree and document the standard route first.
- A rare, low-volume task: build and maintenance time may exceed the saving.
- Work based on missing or unreliable data: AI does not automatically repair a poor CRM or incomplete customer records.
- Legal, tax, medical, safety or large financial decisions: use AI for search and drafts, while a qualified person decides.
- An angry key customer, pricing dispute or contract exception: the relationship and context require human judgement.
- A process without an owner: if nobody checks results and errors, the automation operates without control.
- A process without a baseline: without current hours, cost and quality, you cannot know whether the investment helped.
- An irreversible action without approval: payments, deletion, publishing and contract changes need human review.
What to measure
AI automation is an operating change, not a model demonstration. Process quality, data, team behaviour and risk management matter as much as the technology.
- Active time saved
- Accuracy and rework rate
- Cost per successful case
- Adoption and human override rate
- Escalations and policy breaches
- Impact on conversion, SLA, margin or retention
Sources and further reading
Product capabilities change. The links below are primary or official sources reviewed when this guide was published.