Crosswind
AI & predictive models1 December 2025

Predictive models are cheap. Being wrong is not.

The model is now the commodity. The decision it feeds is the only thing with a price.

Training a competent forecast is a weekend. Wiring it into the moment where someone acts, with a fallback for when it is wrong, is a quarter. Guess which part gets the budget and which part gets the demo.

Redwind research
~80%
Share of project time that is data plumbing, not modelling
2 to 5%
Accuracy gain from a fancier architecture on typical business data
10 to 30%
Value gain from acting on an existing forecast one day earlier
1
Question that matters: what changes when this number appears?

The commoditisation already happened

Ten years ago, having a model was a differentiator. Today a gradient boosted forecast on clean tabular data is a solved problem, available to anyone with an afternoon and a laptop, and a competent language model is an API call. What is scarce is not prediction. It is a business process capable of absorbing a prediction and changing behaviour because of it.

We say this to clients who arrive wanting 'an AI model' the way people used to arrive wanting 'a website' in 2001. The model is not the project. The project is the three meetings nobody wants to schedule: who owns the decision the model feeds, what happens when the model is wrong, and who gets paged when the pipeline goes stale on a Sunday. Skip those meetings and you have built an expensive party trick.

A forecast that nobody is contractually obliged to act on is a very expensive opinion.

We have watched three separate teams spend six months tuning an architecture that moved accuracy from 84% to 87%, then discover the downstream process only used three buckets: low, medium, high risk. The extra three points of accuracy never crossed a bucket boundary for a single customer in the test set. That is not a modelling failure, it is a scoping failure that a modelling team was never positioned to catch.

Start from the decision, backwards

  • Name the decision. Who makes it, how often, and what are the options?
  • Price the error. What does a false positive cost? A false negative? They are almost never symmetric.
  • Set the threshold from the economics, not from an F1 score.
  • Design the fallback before the model. What happens on the day the pipeline is stale?
  • Instrument the outcome. If you cannot measure the decision improving, you have shipped a dashboard.

In risk and forecasting work, the asymmetry step is where most of the value hides. A fraud model tuned to a balanced accuracy metric is tuned to the wrong thing when a missed fraud costs sixty times a false alarm. The threshold, not the architecture, is the lever.

We have seen a demand forecasting engagement where the client's data science team optimised mean absolute percentage error for eight weeks, then handed the output to a planning process that only ever reordered stock in fixed weekly batches. The precision of the forecast was wasted on a decision cadence that could not use it. Once we moved the reorder decision to match the forecast's actual update frequency, stockouts fell by 18% with the same model, unchanged.

  • Ask what decision cadence the business can actually act on before building a forecast faster than that cadence.
  • Write the cost of a false positive and a false negative on the same page as the model spec, in currency, not in percentages.
  • Treat the threshold as a business parameter owned by finance or ops, reviewed quarterly, not a hyperparameter buried in a notebook.

What good looks like in numbers

<1 wk
From baseline model to first live decision, in a healthy setup
Daily
Drift and freshness checks, with an alert that reaches a human
2
Metrics per model: technical accuracy and business outcome
100%
Predictions logged with inputs, so you can explain any decision later

Every model we run in production ships with a simple, stupid baseline next to it: last week's number, a seasonal average, a fixed rule. If the sophisticated model cannot beat the stupid one on the business metric for a full quarter, the stupid one wins. That rule has retired more models than any budget cut.

The baseline is not a courtesy, it is a governor. Complex models decay quietly: a supplier changes a SKU code, a marketing campaign shifts the customer mix, and the model keeps producing confident numbers that are now wrong in a new way. A baseline sitting alongside it, recomputed the same way every week, gives you an early warning that costs almost nothing to maintain and catches roughly a third of the silent failures we have diagnosed in post mortems.

The governance nobody wants to write

Commodity components need boring rules: versioning, rollback, an owner, an audit trail, and a documented answer to 'why did the system decline this customer?'. None of that is exciting. All of it is what separates a predictive capability from a pilot that quietly stopped being retrained in November.

We keep a one page runbook for every model in production: what triggers a retrain, who signs off a threshold change, what the rollback path is, and how long a customer facing decision can be explained after the fact. It reads like a compliance document because in every regulated sector it eventually becomes one. Building it after a regulator asks costs ten times what it costs to build it on day one.

Key takeaways
  1. 01Budget the plumbing at four times the modelling. That is the real ratio.
  2. 02Derive thresholds from the cost of each error type, never from a leaderboard metric.
  3. 03Ship a dumb baseline alongside every model and let it compete.
  4. 04Log inputs and outputs for every prediction. Explainability is an ops feature.
  5. 05Write the retrain and rollback runbook before launch, not after the first incident.
ForecastingRisk scoringDrift monitoringDecision designBaselinesMLOps
The Redwind Briefing

One note. Every domain we run.

Infrastructure, blockchain, media, ventures, AI, marketing, expeditions and yacht deliveries. The real moves, four times a year, with zero filler.

Four times a year. No noise. Unsubscribe anytime.