Most of the time lost on racing models is lost before the model exists. A horse racing dataset for machine learning looks clean in the first few rows, then turns out to have duplicate runners, two horses sharing one identifier, or a column that is quietly empty for half the years. These are the five checks we would run first, in the order we would run them, whatever the source of the data.
1. Does a horse ID mean one horse?
Identity is the foundation. If two different horses share an ID, their records merge and every feature built from a horse’s history is wrong. If one horse has two IDs, its history is cut in half.
The classic cause is the name. Horses reuse names across years, and sometimes across countries. In September 2026 we found two different horses called All Good in the UK data: one by Equiano out of Love Me Tender, the other by Tasleet out of Mythical Spirit. Matched on name and country they look identical. Matched on the source’s own identifier they are two animals with two careers, and we store them as two IDs, 83173 and 156579.
Check any dataset you use by listing the most common names and asking how many distinct IDs sit behind each. If a source joins on name, be suspicious.
2. Are there duplicates or impossible rows?
Three tests find most problems.
The same horse appearing twice in one race is a duplicate. The same horse in two races on the same day is an error, because it cannot be in two places at once. And a race whose number of runners does not match its number of result rows has lost or gained a row.
A short pandas audit for a records table:
import pandas as pd
def audit(df):
out = {}
out["duplicate_race_horse"] = int(df.duplicated(["Race_ID", "Horse_ID"]).sum())
per_day = df.groupby(["Date", "Horse_ID"])["Race_ID"].nunique()
out["horse_in_two_races_same_day"] = int((per_day > 1).sum())
numeric = df[df["Place"].str.fullmatch(r"\d+")].copy()
numeric["place_n"] = numeric["Place"].astype(int)
numeric = numeric.sort_values(["Race_ID", "place_n"])
falls = numeric.groupby("Race_ID")["Distance_btn_total"].apply(
lambda s: bool((s.diff().dropna() < 0).any()))
out["races_where_margin_total_falls"] = int(falls.sum())
return out
The last check, that a running margin never decreases as the finishing place increases, catches mistakes in any column derived from the order of the runners. We run checks like these on every nightly update, along with ones for missing jockey and trainer identifiers, before any file is published.
3. What is the label really?
The place column is not an integer. It holds numbers and codes. In the UK, 78,682 runner records are PU (pulled up), 18,364 are F (fell) and 12,205 are U (unseated). Hong Kong adds withdrawals and dead heats such as 4 DH. If you cast the column to a number and drop what fails, you silently remove every horse that did not finish, and your model learns from a world where nothing ever goes wrong. Our result codes guide explains each one.
Decide on purpose what the target is for non-finishers, and write it down.
4. Where are the holes?
Some columns are empty for good reasons, and some for reasons you need to understand before training.
Draw is empty for jump races, because there is none. The rating is absent for between 10% and 32% of runners in 2025. And timing data is the biggest one: UK finishing times and sectionals start in 2017 and reach about 91% of British runners from 2024, with none for Irish racing. The pattern is by year and course, so a model that uses timing is trained on a particular slice of racing. Our coverage note has the figures.
Count the missing values by year and by course before you decide how to treat them.
5. Can you split it by time?
Racing needs a time-based split, so you need a reliable date on every row. Check that every race has one, that the range is what you expect, and that there are no unexplained gaps. The UK archive, for example, has between 11,800 and 13,300 races in every full year since 2011, with one dip: 10,287 in 2020. Anything stranger than that deserves a look.
Also check the order within a day. Two races on the same date need a time to be put in sequence, otherwise a horse’s earlier run on a given afternoon cannot be distinguished from its later one.
Formats that need parsing
A few columns look numeric and are not. Finishing times are text, such as 1m 12.45s. Distances are text, such as 7f 14y in the UK. Margins are a mix of fractions and abbreviations. We cover each in its own post: distances and surfaces and margins. Build the parsers once and keep the original text beside the number.
Where to start
Run the five checks before you train, not after you are disappointed. The UK and Hong Kong datasets come as CSV, SQL and JSON, the API serves the same tables, and the next step is feature engineering for horse racing models.

Comments are closed