Skip to content
CipherCruCipherCru

Menu

AI Development

AI Development Services

Applied AI with a measurable job to do. We find out whether a model can do that job on your data before anyone commits a year to building around it.

Typical duration
Three to nine months
Team shape
Two to four senior engineers
Starts with
A four-week evaluation

Evaluated before it is built

Most AI proposals are answerable in weeks. The expensive ones are those nobody checked before committing a year to them.

The hard part of an AI project is almost never the model. It is deciding what the model is allowed to be wrong about, and what happens when it is. That question has a cost attached, it can be answered on real data in a few weeks, and answering it first is the difference between a system that changes a business decision and a pilot that has been running for eleven months without deciding anything. So the first thing we deliver is the evaluation, and it is written so that it can honestly recommend not building.

A hand annotating measured results on a tablet
Scored against an agreed measure, on the data you actually receive.

Where AI projects quietly fail

Almost none of the AI work we are called in to rescue failed on the model.

What we are usually called in to fix

  • A demo that convinced everyone and never held up on real data
  • Accuracy nobody measured, because nobody agreed what accurate meant
  • A model that worked until the documents changed shape in March
  • Answers a regulator would ask about and nobody can explain
  • A pilot that has run for eleven months and decided nothing

How we work instead

  • Measure whether it can work before building anything around it
  • Agree what accurate means, in the business's own terms, in writing
  • Test on the data you actually receive, not on a cleaned sample
  • Keep a person in the loop wherever being wrong is expensive
  • Recommend stopping when the evaluation says stop

What that involves

The work that decides whether a model changes a business outcome or only impresses in a demo.

Evaluation before build

The first deliverable is a decision, not a model.

We measure whether a model can do the job on your own data, against a definition of accurate that you agreed in advance. The result is a written recommendation, and it is allowed to say no.

Generative interfaces

Assistants and drafting tools built on retrieval, so an answer can be traced to a source.

Predictive models

Scored against what actually happened, and attached to the decision they exist to change.

Document intelligence

Extraction that holds up on scans, photographs and forms filled in by hand.

Retrieval and vector search

Chunking, indexing and ranking, which is where most answer quality is won or lost.

Human-in-the-loop design

Where the model acts alone, where it drafts, and where a person still signs.

Evaluation in production

Accuracy tracked after launch, so a quiet regression is visible before a customer finds it.

What we build with

Chosen for what has to be explained afterwards, and for who has to run it once we have gone.

  • Anthropic
  • OpenAI
  • Gemini
  • Llama
  • Mistral

Where we have built this

Sectors where a wrong answer has a cost we can name, which is the only place this work is worth doing.

A banking hall desk where a laptop shows an account dashboard with a running balance, recent transactions and a monthly activity chart, a phone beside it

Fintech & BFSI

Secure, reliable systems for businesses built on trust.

We build and modernize financial platforms, workflows, integrations, and customer-facing applications where reliability, data integrity, auditability, security, and performance are fundamental to the business.

A clinician talking with a patient at the bedside while a monitor shows that patient's record - vitals, scans and recent reports - with the same case open on a phone

Healthcare

Technology that connects complex care and operational workflows.

We help build digital healthcare platforms, operational workflows, integrations, and data-driven applications designed around reliability, secure information handling, usability, and the realities of interconnected healthcare systems.

An open-plan office with a laptop showing a workforce dashboard - headcount, attendance, leave requests and approvals - and the same modules on a phone

HRMS

Workforce systems built around how organisations actually operate.

We build and improve workforce platforms spanning recruitment, onboarding, employee management, attendance, operational workflows, reporting, and integrations, reducing friction across the employee lifecycle.

What you end up with

Stated from your side rather than ours.

  • A decision on whether to proceed

    An evaluation that can honestly recommend not building, and sometimes does.

  • Measured accuracy, not a demo

    A figure taken on your own data, with the method written down beside it.

  • Answers you can explain

    Sources, thresholds and the human step, recorded for the people who will ask.

  • A team that can run the next version

    Your engineers run the evaluation and the deployment before we leave.

How an AI engagement runs

The same five stages whichever of the services above does the work.

Diagnose

What decision the model would change, and what it costs when it is wrong.

Decide

Whether to build at all, with the evaluation attached. Often the answer is a rule, not a model.

Prove

The hardest case tested on your own data, scored against the agreed measure.

Deliver

Built in slices, each one behind a switch and monitored from the day it is on.

Hand over

Your team runs it with us watching, then without us.

Questions about AI development

The ones we are asked most often before a first conversation.

How do we know it will work before we commit?
You do not, which is why the first piece of work is an evaluation rather than a build. We agree what accurate has to mean in your terms, assemble a sample of the data you actually receive, and measure against it. That takes weeks rather than months, and it ends in a written recommendation you can act on whichever way it goes.
Do we need to train our own model?
Usually not. For most business problems a general model with good retrieval over your own content outperforms a small model trained on a dataset you can realistically assemble, and it costs a fraction as much to keep current. Training your own becomes the right answer when the task is narrow, the data is genuinely yours, and the volume makes the running cost matter more than the build cost. That is a decision the evaluation is designed to settle.
Does our data end up training someone else's model?
Not unless you choose a service that does that, and it is a decision we put in front of you rather than make quietly. The major providers offer terms where prompts and content are not used for training, and where the region the data is processed in is fixed. Which provider and which terms apply is recorded with the architecture, so the answer is available when somebody asks for it later.
How do you measure accuracy?
Against a definition you agree before we start, because accuracy on its own is not a number. A document extractor and a fraud score fail in completely different ways, and the cost of a false positive is rarely the cost of a false negative. We write down which errors matter, build a scored test set from real cases, and report against it. The same test set then runs in production.
What happens when the model gets worse?
It will, and quietly. Documents change format, customer behaviour drifts, and providers retire model versions on their own schedule. The evaluation set runs continuously against live traffic so a fall in quality raises an alert rather than a complaint, and every model call sits behind a switch that reverts to the previous version or to the existing rules.
What happens after launch?
An AI system needs more attention after launch than a conventional one, because its inputs keep moving. Your engineers have been running the evaluation pipeline alongside ours since the first slice, so the monitoring is already theirs. What we agree separately is how much of our time you want for the retraining and provider changes that follow.

Have an AI idea worth testing?

Tell us the decision you want it to change. We will tell you whether it can.

Strictly necessaryEssential for the site to function: page navigation, security, session management, and remembering the cookie choices you make here.
Always on
FunctionalRemembers choices you make, such as language, region or display preferences, so the site opens the way you left it.
Performance and analyticsPerformance and analytics cookies show us how the Website is used: which pages are visited, how long is spent on them, where visitors came from, and what errors occur. They are set by Google Analytics and by HubSpot, whose cookies also link the pages you viewed to any enquiry you later send us.
Marketing and targetingTracks browsing activity to measure advertising and show relevant ads. We set none of these today, and will not without your opt-in.