people riding horses on a grassy field

Say “AI in horse racing” to ten people and you will get ten pictures: a model that picks winners, a robot trainer, a chatbot that knows every horse. Most of those are marketing. A smaller set of uses really does work, and almost all of them depend on having clean, structured data underneath.

Below are the uses that have held up for us on UK and Hong Kong racing data, and the ones we would treat with suspicion.

Language models need the table in front of them

Ask a general language model about a particular horse and it will answer fluently. It will also invent a form line, a sire and a trainer, because nothing in its training ties the name to a verified record. A name is not an identifier, and the model has no way to tell two horses called the same thing apart.

The reliable pattern is the opposite: hand the model the records and ask it to describe them. Give it a race, its runners, finishing order and margins from a database, and ask for a short report. The numbers come from the data and the model supplies the prose. We build our own content this way, and then check every figure against the source before anything is published.

The same applies to any question. “How did this horse run at Kempton?” is answerable from a query. A model should call the query, not recall an answer.

Classifying text that already exists

Racing data has more free text than people expect. Hong Kong publishes veterinary notes, and our archive holds 53,966 of them. The vocabulary is small: just ten phrases, such as “Unacceptable performance.” and “Castration.”, make up 43% of all notes. That is an ideal problem for a text classifier, and a modest one will group the notes into withdrawals, post-race findings, surgery and routine flags. The veterinary records post has a starter and the counts.

Nobody needs a large model for that. A few rules, or a small trained classifier, gets most of the way, and you can read every misclassification.

Pace and performance ratings

This is where the numbers matter most. With sectional times you can describe how a race was run, and with margins and class you can build a rating that compares horses across courses. These are statistical models, not artificial intelligence in any dramatic sense, and they are the established way of working.

Their limit is coverage. Sectional times exist for about 91% of British runners since 2024, for none in Ireland, and for Hong Kong runners from 2008. A model is only as broad as its inputs. See the coverage by year.

Data quality itself

An underrated use is checking the data. Anomaly detection on a nightly feed flags a margin that does not fit, a field of runners that does not match the declared number, or a horse with a birth year that contradicts its history. Most of these need rules, not learning, but a model is a good way to find the cases you did not think of.

What we would not promise

Predicting winners is the use everyone asks about, and it is the one with the weakest evidence. A field of nine or ten runners, one winner, and much of what decides a race, such as a horse’s fitness on the day, never appears in a dataset. We sell data, not forecasts, and we would be wary of anyone who promises results from a forecast.

If you do build a predictive model, treat it as a research problem. Split by time, keep a plain baseline and test for leakage. The primer on machine learning with racing data covers those habits, and feature engineering for horse racing shows how to build the inputs.

Starting with the data

If a vendor tells you AI picks winners, ask to see the results split by date. If they show you a dataset with stable identifiers and documented fields, you are talking to someone who has done the unglamorous part. Ours are in the horse racing datasets and the API, and the dataset checklist lists what to verify before you train anything.

Categories:

Tags:

Comments are closed

0
    0
    Your Cart
    Your cart is emptyReturn to Shop