The cheapest way to implement AI is to solve one measurable business problem with an existing model, prove the value in a short pilot, and only then spend on custom work. Most overspending comes from doing those steps in the wrong order: buying platforms, hiring data scientists or training models before anyone has agreed what "success" looks like.
The economics have also shifted. The Stanford AI Index 2025 reports that the inference cost of a system performing at GPT-3.5 level fell more than 280-fold between November 2022 and October 2024. Capability that once needed a research budget is now available through a pay-as-you-go API. Yet smaller businesses still lag: in the EU, Eurostat (2025) found 20% of enterprises with 10 or more employees used AI, and only 17% of small enterprises did.
This guide walks SMB owners and department heads through eight steps, from choosing the problem to keeping governance proportionate.
Key Takeaways
- Start with one problem that has a number attached to it (hours, error rate, response time), not with "an AI strategy".
- Use off-the-shelf APIs and foundation models first. Custom training is a later decision, made with evidence.
- For most business knowledge tasks, retrieval-augmented generation (RAG) is cheaper to start and maintain than fine-tuning.
- Run a time-boxed pilot with a success metric and kill criteria agreed before any work starts.
- Control running costs with usage caps, smaller models, caching and batch processing.
- Match governance to risk: most SMB use cases are low-risk, but data protection always applies.
Why AI projects overspend
AI projects rarely go over budget because the technology is expensive. They overspend because the scope is vague, the data isn't ready, and nobody decided in advance when to stop.
The gap between adoption and value is well documented. McKinsey's State of AI 2025 found that 88% of respondents' organisations use AI regularly in at least one business function, but only around a third have begun to scale it, and just 39% report any EBIT impact at enterprise level.
For smaller firms, the barriers are mostly practical. The UK government's AI Adoption Research (DSIT, 2026) found that only 16% of UK businesses currently use at least one AI technology, with 71% of non-adopters saying they hadn't identified a use for it, and 60% of all businesses citing limited AI skills or knowledge as a barrier. Among EU enterprises that considered AI but didn't use it, Eurostat (2025) recorded lack of relevant expertise (71%), unclear legal consequences (53%) and data protection concerns (49%) as the top reasons.
Step 1: Pick one measurable problem
The goal here is a single use case, a baseline number and an owner. Start by finding the need rather than the tool. Look for work that is repetitive, text- or image-heavy, and easy to measure.
Good first candidates usually look like this:
- Triage and drafting: classifying inbound emails or support tickets and drafting first replies.
- Document extraction: pulling fields from invoices, delivery notes or application forms into your existing system.
- Internal knowledge search: answering staff questions from policies, manuals and past project notes.
- Forecasting from data you already hold: demand, staffing or stock levels.
Write down the baseline before you start, such as "a coordinator spends two days a week keying in supplier invoices". If you can't measure the current state, you can't prove the improvement. For more examples by function, see our guide to machine learning in business.
Step 2: Start with off-the-shelf APIs and foundation models
Hosted foundation models handle most language, summarisation and classification tasks out of the box, on pay-per-use pricing with no infrastructure to run. Open-weight models are a credible alternative if you need to host privately: the AI Index 2025 found the performance gap between open-weight and closed models narrowed from 8% to 1.7% on some benchmarks within a year.
A practical approach:
- Build a test set of 50 to 200 real examples (anonymised where needed) with the "right answer" for each.
- Run them through two or three models, including at least one smaller, cheaper model.
- Score the outputs against your baseline, and note cost per task.
If a general model reaches acceptable quality with good prompts, stop there. Custom training only makes sense when off-the-shelf options demonstrably fall short. The same logic applies to buying software in general, as covered in custom software vs off-the-shelf.
Step 3: Choose RAG or fine-tuning based on the problem
Foundation models don't know your prices, policies or product catalogue. There are two main ways to fix that, and they have very different cost profiles.
- Retrieval-augmented generation (RAG) searches your documents at query time and passes the relevant passages to the model. Updating it is as simple as adding a document, and answers can cite their source.
- Fine-tuning retrains a model on your examples so it learns a style, format or narrow task. It needs curated training data, repeat runs when things change, and careful evaluation.
As a rule of thumb, use RAG when the problem is "the model needs to know our facts", and consider fine-tuning when the problem is "the model needs to behave in a very specific, consistent way" that prompting can't achieve.
Step 4: Compare the approaches before you commit
The table below summarises the main options. Cost descriptions are relative, because actual figures depend on volume, model choice and provider pricing.
| Approach | Typical cost profile | Time to value | When to use |
|---|---|---|---|
| AI features in tools you already pay for | Lowest; often included or a per-seat add-on | Days | Generic tasks such as meeting notes, drafting, spreadsheet help |
| Foundation model API with prompts | Low; pay per use, no infrastructure | Days to a few weeks | Classification, summarisation, drafting, extraction |
| RAG over your own documents | Low to moderate; adds search, storage and integration work | A few weeks | Answers must reflect your policies, products or records |
| Fine-tuning an existing model | Moderate; data preparation, training runs, re-evaluation | Weeks to months | Consistent format or tone, or a narrow task prompts can't reach |
| Custom model trained on your data | Highest; data collection, labelling, ML engineering, ongoing monitoring | Months | Proprietary prediction or vision problems with no off-the-shelf fit |
Custom models still have a place, especially for forecasting and computer vision on your own data. Our Punjab air-quality dashboard combines machine learning with GIS to produce a 72-hour forecast, which no general chat model provides. The point is to reach that tier deliberately. Our machine learning development and computer vision pages cover what that work involves.
Step 5: Run a pilot with a success metric and kill criteria
A pilot is cheap only if it has an end date, and the plan should be written so anyone in the business can read it. Before any build work, agree four things:
- Success metric: the number from Step 1 and the target, such as "cut invoice keying time by half with no more errors than today".
- Duration: usually four to eight weeks of real use, not just a demo.
- Budget cap: a fixed amount for build effort and a monthly ceiling for usage costs.
- Kill criteria: the conditions under which you stop, such as accuracy below an agreed threshold after two rounds of prompt tuning, or staff rejecting the workflow.
Keep a human in the loop: staff review outputs, flag errors and record time saved. That gives you honest data and protects customers.
Stopping a pilot that misses its targets is a good outcome, not a failure.
Step 6: Get your data ready on a budget
Aim for just enough clean, accessible data for the pilot, and no more. Data readiness is a common reason for delay, and the OECD's 2025 report on AI adoption by SMEs notes that the share of large firms using AI is more than three times that of small firms, with skills and data readiness among the factors behind the gap. You don't need a data warehouse to start.
- Scope data to the use case. For a policy assistant, that's the current policies, not every file on the shared drive.
- Fix the obvious problems. Remove duplicates and outdated versions, and agree which document is the source of truth.
- Strip what you don't need. Personal data that isn't required for the task shouldn't go anywhere near a model.
If your processes still live in spreadsheets, a modest business application that captures structured data can be the step that makes AI useful later.
Step 7: Put cost controls in from day one
The aim is predictable monthly running costs. Usage-based pricing is cheap at pilot scale and can surprise you in production. These controls are cheap to set up:
- Usage caps and alerts. Set hard spending limits and alerts per project or API key with your provider.
- Right-size the model. Route simple tasks such as classification to a smaller, cheaper model and keep larger models for complex reasoning.
- Cache repeated context. Long system prompts and reference documents sent with every request are a hidden cost. With Anthropic's prompt caching, for instance, cache reads are billed at a fraction of the standard input price (0.1x on most models).
- Batch work that isn't urgent. Overnight document processing doesn't need instant replies. OpenAI's Batch API offers a 50% discount compared with synchronous requests, with results returned within 24 hours.
- Track cost per task. Report it next to the business metric, so you can see whether value grows faster than spend.
Step 8: Decide whether to build, partner or buy
By the end of this step you should know who will run the solution after the pilot.
- Buy when a SaaS product already solves the problem well and your data and workflow fit its assumptions.
- Build in-house when you have, or can hire, people who'll own the system long term, and the use case is core to how you compete.
- Partner when you need specialist skills for a defined period, such as integration, RAG design or model evaluation, without permanent hires.
Whatever the mix, insist on ownership of your prompts, data pipelines and evaluation sets, so you can switch models or suppliers later. Our guide to choosing a software development partner lists the questions to ask, and our AI and ML solutions page describes how we scope this kind of work.
Keep governance and risk proportionate
Governance should scale with the risk of the use case. An internal drafting assistant doesn't need the controls of a credit-scoring system, but every use case needs basic data protection.
The EU AI Act entered into force on 1 August 2024 and applies in stages, as set out on the European Commission's AI Act page. Prohibited practices and AI literacy duties have applied since 2 February 2025, and general-purpose AI model rules since 2 August 2025. Under the Digital Omnibus on AI (Regulation (EU) 2026/1744, in force since 27 July 2026), obligations for stand-alone high-risk systems (such as recruitment or credit scoring) now apply from 2 December 2027, and those for AI embedded in regulated products from 2 August 2028, according to the AI Act implementation timeline. Most SMB use cases, such as drafting, search and document extraction, fall outside the high-risk categories, but transparency duties can still apply, for example telling people when they're interacting with a chatbot.
Data protection applies regardless of the AI Act. In the UK, the ICO's guidance on AI and data protection covers lawful basis, fairness and transparency. A proportionate baseline for most SMBs:
- Keep a simple register of AI tools in use, what data they touch and who owns each one.
- Check provider terms on data retention and whether your inputs are used for training.
- Run a data protection impact assessment where personal data is involved at any scale.
- Keep a human review step for decisions that affect customers or employees.
Sector rules in areas such as healthcare and finance come on top of this.
Frequently asked questions
How much should a small business budget for a first AI pilot?
There's no reliable industry-wide figure, because costs depend on the use case, data and integration needed. The safer approach is to set your own cap: a fixed build budget and a monthly usage ceiling, linked to the value of the problem you measured in Step 1. If the potential saving is small, the pilot budget should be small too.
Is RAG always cheaper than fine-tuning?
Not always, but it usually is to start with and to maintain. RAG avoids training runs and updates when you add documents. Fine-tuning can reduce per-request costs for high-volume, narrow tasks, so compare both on your own test set once volumes are known.
Does the EU AI Act apply to a UK or US business?
It can. The Act applies to providers and deployers whose AI systems are placed on the EU market or whose outputs are used in the EU, wherever the business is based. If you serve EU customers, check how your use cases are classified.
Conclusion
A cost-effective AI strategy is mostly about sequence: one measurable problem, an existing model, a capped pilot with kill criteria, and cost controls before scale. Inference costs have fallen sharply, so the main risk for smaller businesses is no longer the price of the technology. It's spending without a clear target.
If you have a use case in mind and want a second opinion on scope, approach and budget, book a free 30-minute discovery session with our team.

