Kreštalica

Engineering, AI & Data

Moving AI from experimentation into operations

Pilots are easy to start and hard to finish. What separates AI that runs in an operation from AI that stays in a slide deck is rarely the model. It is the process, the data and the ownership around it.

Kreštalica6 min readEngineering, AI & Data
Close-up of a dark blue printed circuit board with gold contact pads.

Many organisations now have something they call an AI pilot. A team tried a language model on a document-heavy task, built a proof of concept, and showed encouraging results. Then the pilot sat there. Nobody owned it, the data it needed was not reliably available, and the people who would use it every day were never really part of the design. This pattern is common enough to be predictable, which also means it is avoidable.

Start from a task, not from a technology

Durable AI deployments tend to begin with a narrow, repetitive task that a specific team performs often and measures already: classifying incoming requests, extracting fields from invoices, drafting first responses, reconciling records. The task has a known volume, a known cost and a known error rate. That gives the project a baseline and a definition of success before a single model is chosen.

Projects framed the other way round, as a search for places to apply a particular tool, tend to produce demonstrations rather than operations.

Fix the input before improving the model

A model is only as consistent as what it receives. If documents arrive in six formats, if the master data has duplicates, or if the same field means different things in two systems, the output will be uneven regardless of the model quality. In practice, a large share of the work in a successful deployment is unglamorous: agreeing formats, cleaning reference data, and building the integration that feeds the model reliably.

Design for the exceptions

In a pilot, the interesting question is how often the model is right. In operations, the important question is what happens when it is wrong. A deployable solution needs a clear route for low-confidence cases: who reviews them, how quickly, and how their corrections flow back into the process. Systems that make exceptions visible and easy to handle earn trust. Systems that hide them lose it quickly.

Give it an owner and a budget line

  • A named business owner who is accountable for the outcome, not just the technology team.
  • A measured baseline and a small set of operational metrics that are reviewed monthly.
  • A maintenance plan: who monitors quality, updates prompts or models, and handles changes in the upstream data.
  • A decision rule for expanding scope only after the first task is stable.

Controls are part of the product

Access rights, logging, data retention and review steps should be designed in from the start, not added after a concern is raised. Responsible use of AI is mostly good engineering and good process discipline: knowing what data the system sees, what it produces, who checks it, and how mistakes are corrected.

The goal is not an impressive model. It is a task that gets done more reliably, at lower cost, with people who trust the result.

When those conditions are in place, expansion becomes straightforward. The second and third use cases reuse the integration, the review process and the governance that the first one established. That is the point at which AI stops being an experiment and becomes part of how the business operates.

Let’s move your next decision forward.

Tell us what you are working on. We will respond with a clear view of how we can help, and what it would involve.