Machine learning on horse racing looks like an easy problem and behaves like a hard one. Every race has a clear outcome, there are hundreds of thousands of them, and each runner comes with a long list of numbers. Then you train a model and find that it often barely beats a simple baseline. This primer is about why, written from the side of the person who supplies the data.

We do not sell predictions. We sell the data that people build their own models from, so what follows is about the data, not about any particular result.

Decide what you are modelling

The target comes first, and there are three common choices.

Will this horse win? That is a binary label, and it is very unbalanced. In our UK archive the average field has 9.5 runners and 99.8% of races have exactly one winner, so the base rate is about one in ten.

Will it finish in the first three? Less unbalanced, and less sensitive to the sort of luck that decides close finishes.

How fast will it run? A regression on finishing time has a continuous label, which is attractive. The trouble is availability. UK finishing times exist from 2017 and, as our coverage note shows, only for some races and no Irish ones. You can only train on the races that were timed.

Pick one, write it down, and resist changing it halfway through because the numbers look better.

What each runner comes with

For a UK runner in 2025, these fields are filled in:

FieldFlat AWFlat TurfChaseHurdle
Weight100%100%100%100%
Draw99.7%91.2%0.4%0.0%
Rating81.4%71.9%90.3%68.2%

Draw simply does not apply over jumps, so do not fill it in. The rating is absent for between 10% and 32% of runners, depending on the type of race. Whether you treat that as its own category or as a missing value is a modelling choice with real consequences.

Most horses have little history

This catches people out more than anything else. Across the 155,221 horses in our UK archive, which starts in 2011, the median horse has run 7 times. 8.7% have run once and 31.9% fewer than five times. Only 18.2% have run 20 times or more.

So for about a third of runners, any feature built from past form is built on a handful of observations. Rolling averages over the last five runs mean something very different for a horse with fifty runs behind it. Include the number of previous runs as a feature, and think about whether to shrink noisy averages towards a prior.

Split by date, never at random

A common error in racing models is a random train and test split. Horses run many times, so a random split puts a horse’s 2024 runs in the training set and its 2023 runs in the test set. The model has seen the future. Accuracy looks wonderful and means nothing.

Split by date: train on everything before a cutoff, test on what comes after, and repeat with a moving cutoff. It is slower and the numbers are worse, and they are the only ones that tell you anything.

Know what was known when

A racing record mixes information from before the race and after it. Mix them up and you get leakage that is easy to miss.

Known before the race: age, weight, draw, gear, jockey, trainer, the horse’s rating and its past runs.

Known only after: the finishing place, the margins, the finishing time and the sectional splits, and the final starting price, which is only settled at the off.

A model that includes the margin of victory in its inputs is not a model. It is a copy of the answer. Build your features from a racecard-style view of each runner: what could you have known at declaration time?

Missing data is not random

Sectional times are the clearest example. They exist for about 91% of British runners from 2024 and for none in Ireland. A model that uses them has quietly restricted itself to certain courses and certain years. Keep a flag for whether the field was present and test the model separately on each group. See the UK sectional times coverage for the full breakdown.

Keep a baseline

Before you trust any model, compare it with something dull. Ranking the runners by their official rating, or by the starting price, gives you a benchmark that already encodes a great deal of what is known. A complicated model that cannot beat it has not found anything.

Next steps

None of this makes machine learning on racing hopeless. It makes it a data problem first and a modelling problem second, and the order matters. Work through feature engineering for horse racing models next, then run the checks in a horse racing dataset for machine learning. The raw UK and Hong Kong tables are in the datasets and the API.

Categories:

Comments are closed

0
    0
    Your Cart
    Your cart is emptyReturn to Shop