AI Implementation Services: What to Buy, What to Build, What to Skip
Most AI implementation invoices pay for slide decks and workshops. Here's the buyer's guide I wish every leader had before signing.

TL;DR
The AI implementation services market is flooded with vendors selling different products under the same label. This guide breaks down the four capabilities that actually make up an AI implementation engagement — discovery, delivery, integration, and enablement — tells you which to outsource and which to keep in-house, and shows the pricing shape of a real project versus a consulting-theater one. If a proposal is 60% "strategy workshops" and 10% shipped code, you're buying the wrong thing.
Why the market is confusing
Type "AI implementation services" into any search engine and you'll get three completely different vendor types on the first page:
- Big consulting firms selling multi-month "AI transformation" engagements — heavy on strategy decks, light on shipped software.
- Boutique AI shops selling model fine-tuning and RAG pipelines by the sprint — heavy on engineering, light on business outcomes.
- System integrators selling "AI-enabled" versions of the ERPs and CRMs they already implement — heavy on platform lock-in.
All three call themselves "AI implementation services." All three have valid use cases. But if you don't know which one you actually need, you'll sign the wrong contract and be six months into the wrong project before you find out.
The four capabilities you're actually buying
Every real AI implementation engagement is some mix of these four capabilities. Get clear on which ones you need before you get pitched.
1. Discovery
Turning "we want to do AI" into a prioritized list of use cases with a measurable ROI thesis for each. Includes user interviews, workflow mapping, data readiness audits, and a build-vs-buy decision per use case.
Insource or outsource? Outsource for the first engagement — an outside perspective catches assumptions your team can't see. Insource for the second and third — your team will be fine by then.
Red flag: a discovery deliverable that's a slide deck without numbers. Discovery should produce a use-case scorecard with dollar values, time estimates, and named owners. Not a "vision."
2. Delivery
Actually building the working system. Prompt engineering, model selection, evaluation harness, guardrails, the RAG pipeline, the fine-tune, the fallback logic when the model is wrong.
Insource or outsource? Outsource until you've shipped two production AI systems. Then in-source the third. The learning curve is real but not infinite — most engineering teams can own AI delivery within 6–9 months if they ship real projects, not tutorials.
Red flag: a "delivery" line item that's mostly meetings. Delivery is measured in commits, not calls. Ask for a weekly demo commitment in the SOW — if the vendor won't commit to it, they don't intend to ship.
3. Integration
Wiring the AI into your existing systems: your CRM, your ticketing tool, your data warehouse, your identity provider. Includes auth, rate limits, cost controls, PII redaction, and logging.
Insource or outsource? Insource by default. This is standard enterprise integration work — your team already does it for every other SaaS product. The AI part is trivial; the integration part is where your team's institutional knowledge matters.
Red flag: an integration line item that's 30%+ of the SOW. That's a signal the vendor is trying to insert themselves as a middleware layer. Push back.
4. Enablement
Change management. Training the people who will use the tool. Redesigning workflows. Establishing the trust zones (see Enterprise AI Adoption for what those are). This is the capability that most vendors under-scope and most buyers under-fund.
Insource or outsource? Blend. Outsource the first workshop, the trust-zone playbook, and the first cohort of change champions. Then run the ongoing enablement in-house.
Red flag: no enablement line item at all. If the SOW is 100% engineering, the project will ship on time and be abandoned within 90 days.
The pricing shape of a real engagement
Here's the rough allocation shape I see on projects that succeed:
| Capability | % of budget | Duration |
|---|---|---|
| Discovery | 10–15% | 2–4 weeks |
| Delivery | 45–55% | 8–14 weeks |
| Integration | 15–20% | overlaps delivery |
| Enablement | 15–25% | starts week 4, continues 90 days post-launch |
Compare this to a typical "big consulting" proposal shape:
| Capability | % of budget | Duration |
|---|---|---|
| Discovery | 40–60% | 6–10 weeks |
| Delivery | 15–25% | 4–6 weeks |
| Integration | 5–10% | often "customer responsibility" |
| Enablement | 15–20% | mostly slide-deck training |
The second shape is what I mean by "consulting theater." You're paying for the workshop and getting the software as a rounding error.
Questions to ask any vendor before signing
Steal these. Ask them in the same order.
- "Show me a working system you shipped in the last 90 days, not a case study. Screen-share the actual UI." If they can't, they build slides, not software.
- "What's the smallest, cheapest thing you could ship in 30 days that would prove your approach works?" Good vendors have a strong answer. Bad ones try to sell you a 6-month discovery.
- "Who on my team owns this at handover, and what documentation do they get?" If the answer is vague, you're being set up for a support-and-maintenance retainer.
- "What's your evaluation methodology? Show me a real eval report from a past project." Every serious AI vendor has an eval harness. If they don't, they ship on vibes.
- "What happens in month 4?" This is the honesty question. The right answer is: "You should have most of this in-house by then, and we should be phasing out." The wrong answer is a pitch for phase two.
What to expect from a good AI implementation partner
Beyond the four capabilities, a partner worth their invoice will:
- Say no to at least one use case you asked for. Not every problem is an AI problem. A vendor who agrees with every idea is optimizing for revenue, not outcomes.
- Bring their evaluation framework, not build one on your dime. Serious teams have evals from previous work. If they're building it from scratch, you're funding their tooling.
- Push for the smallest first release possible. If they want to ship "the full vision" in v1, they either don't understand the problem or they're padding hours.
- Have opinions about model choice, not just "we'll use whatever you prefer." Vendor-neutral is a red flag. Real practitioners have preferences based on evals, not politics.
- Hand off cleanly. The best sign of a good engagement is that you don't need them for phase two. Vendors who structure phase-one to guarantee phase-two are optimizing against your interests.
What this means for the RFP you're about to send
If you're evaluating AI implementation vendors this quarter, don't lead the RFP with a use case. Lead it with a capability breakdown. Ask each vendor to price out discovery, delivery, integration, and enablement separately, with a demo-per-week commitment in delivery.
The vendors who can price it that way have shipped before. The ones who send back a lump-sum "AI transformation" number have not.
Evaluating AI implementation vendors and want a technical second opinion on the proposals? Get in touch — we do this without a retainer.
Also published on my newsletter.
Read on SubstackKeep reading
GenAI Training for Teams: Which Format Fits Which Audience
A one-hour lunch-and-learn changes nobody's behaviour. Here are the four formats that do, and the audience each one is built for.
Read articleFrom AI Pilot to Production: The 18-Point Readiness Checklist
The gap between a pilot that impresses and a system people rely on is a list of unglamorous items. Here are the eighteen that matter.
Read articleTurn ideas into action
If any of this hit home — let's talk about applying it to your team.
Talk to us