Skip to main content

Business challenge

Getting past the AI demo and into production

Building something impressive with a language model takes an afternoon. Making it reliable, affordable, explainable and safe to put in front of customers is the actual project.

What this looks like

You will recognise at least one of these

  • 01

    The prototype works and the rollout does not

    It answers the demo questions well and falls apart on the long tail of real ones, because nothing was ever measured against a representative set of inputs.

  • 02

    Nobody can say whether it is right

    There is no evaluation set and no baseline, so 'it seems better' is the only available verdict and every change is a matter of opinion.

  • 03

    The data is not ready

    The knowledge the model needs is spread across a wiki, a shared drive and three inboxes, in inconsistent formats, some of it out of date and none of it labelled.

  • 04

    The cost is unpredictable

    Per-token pricing behaves nothing like a licence. Usage that is fine in a pilot becomes a material line item at organisational scale.

How we approach it

The order the work goes in

Sequence matters more than tooling here. Most of the expensive mistakes are made by doing the right things in the wrong order.

  1. 1

    Pick a task with a checkable answer

    Classification, extraction and routing can be measured. Open-ended generation is far harder to hold to a standard, and a much riskier first project.

  2. 2

    Build the evaluation before the feature

    A set of real inputs with known-good outputs turns 'seems better' into a number, and is what makes a model or prompt change safe to ship.

  3. 3

    Keep a human in the loop where it matters

    Draft-and-approve is a legitimate destination, not a stepping stone — especially anywhere a wrong answer carries regulatory or financial weight.

  4. 4

    Decide the governance before launch, not after

    What data may be sent to a third-party model, what is logged and for how long, and what a customer is told about how their information is used.

Where we help

The services that do this work

These are existing engagements rather than a new offering — each links to what it actually involves.

If an AI pilot has stalled somewhere between promising and shippable, the blocker is usually evaluation or data rather than the model.

Start with a conversation, not a proposal

Tell us what is not working. If we are not the right people for it, we will say so.