Predictive Analytics
A forecast nobody acts on is a report. The engineering that matters is not the model — it is the connection between the prediction and the decision.
- Typical duration
- Three to nine months
- Team shape
- Two to four senior engineers
- Starts with
- A baseline, and a decision

Start from the decision, and from what you already do
Two questions decide whether a predictive project is worth running, and both come before any modelling. What decision changes as a result of this forecast? And how well does your current approach — a rule, a spreadsheet, or somebody's judgement — already do?
Without the first, you get a dashboard nobody opens. Without the second, you have no way to know whether a model that sounds impressive is actually better than the person it replaced. So we establish the baseline and the decision in the first weeks, and every model afterwards is judged against them rather than against a metric chosen because it flattered the result.
What usually goes wrong, and what we do instead
Predictive projects rarely fail on modelling. They fail on baselines, leakage and adoption.
Where predictive projects fail
- No baseline, so nobody can say whether the model is an improvement
- Training data containing information that would not exist at decision time
- A metric chosen after the results, because it made them look better
- A forecast delivered to a dashboard rather than into the workflow
- Model drift after launch that nobody is scoring for
How we build it instead
- The current approach measured first, and treated as the bar to beat
- Features validated against what is actually known at decision time
- The success metric agreed in writing before the first model is trained
- Predictions delivered into the tool where the decision is made
- Scored against outcomes continuously, so drift is visible early
What that involves
The work that decides whether a forecast changes a business outcome or only fills a report.
Baseline before model
You cannot improve on something you have not measured.
We measure how well your current rule, spreadsheet or judgement already performs, and agree the metric that would count as better. Every model afterwards is scored against that, and it is allowed to lose.
Feature engineering
Built from what is genuinely known at decision time, which is where leakage quietly gets in.
Forecasting and classification
Demand, churn, risk and failure, modelled at the horizon the decision actually needs.
Explainability
Why this prediction, in terms the person acting on it can check against what they know.
Decision integration
Delivered into the tool where the decision is made, not into a dashboard beside it.
Human-in-the-loop design
Where the forecast acts automatically, where it advises, and where a person overrides it.
Scoring in production
Predictions compared with outcomes as they arrive, so drift shows up before a quarter does.
What you end up with
Stated from your side rather than ours.
A model you can prove is better
Scored against your current approach on an agreed metric, so the improvement is a measurement rather than an impression.
A forecast people actually use
It arrives in the tool where the decision is made, which is the difference between a model that changes something and one that does not.
Predictions someone can challenge
Explanations in terms the person acting on it recognises, so a wrong prediction can be caught by the expert rather than obeyed.
Drift you find before your customers do
Predictions are scored against outcomes continuously, so a model degrading is a dashboard rather than a surprise.
How the pieces fit together
The model is one component among six, and the two on either side of it decide whether it is worth anything.

What each part does
- Sources: the systems the history comes from, with the timestamp that says when each fact became known.
- Feature pipeline: the same transformations at training and at inference, which is what stops leakage and skew.
- The model: retrained on a cadence, versioned, and always comparable against the baseline it has to beat.
- Explanation: the reasons behind a prediction, produced alongside it rather than reconstructed afterwards.
- Decision integration: the prediction delivered into the tool where somebody acts, with the override path recorded.
- Outcome feedback: what actually happened, matched back to what was predicted — the loop that makes drift visible.
How the work runs
The same five stages as every engagement, applied to a system that is right most of the time.
Diagnose
Which decision this changes, how well it is made today, and whether the data supports predicting it at all.
Decide
The baseline, the success metric and the decision horizon, agreed in writing before any model is trained.
Prove
One model scored against the baseline on held-out history, then run in shadow beside the current approach.
Deliver
The prediction wired into the decision, with the override path and the outcome feedback built at the same time.
Hand over
Your team retrains and scores the model with us watching, then without us.
What we build it with
Chosen for what has to be explained afterwards, and for who has to run it once we have gone.
Modelling
- Python
- PyTorch
- scikit-learn
- XGBoost
Data
- PostgreSQL
- Snowflake
- dbt
- Kafka
Services
- TypeScript
- Node.js
- Redis
- GraphQL
Platform
- AWS
- Azure
- Docker
- Datadog
What you get, and when
Handed over as it is produced, not assembled at the end.
- Your current approach measured, as the bar to beat
- The success metric and the decision horizon, agreed in writing
- An honest assessment of whether the data supports this at all
- Source in repositories you own
- One feature pipeline used at both training and inference
- A retraining process your team can run and interpret
- The human boundary: where it acts, advises, and is overridden
- Explanations produced alongside each prediction
- Drift dashboards scoring predictions against outcomes
The rest of our AI work
Most AI engagements need one of these underneath them, and some need two.
Related work
Engagements where a number had to change a decision to be worth anything.
Questions about predictive work
The ones we are asked most often before a first conversation.
How accurate will the forecast be?
Do we have enough data?
Will we be able to explain a prediction?
How do we get people to actually use it?
What happens when the world changes?
What if a model would not beat what we do now?
Have a decision that could be better informed?
Tell us the decision and how it is made today. We will tell you whether a model would beat it.

