The July Course, Newmarket - Horses pass the winning post

A Data-Rich Sport

Horse racing is one of the most data-intensive sports in the world. Every meeting at every course — from Happy Valley in Hong Kong to Ascot in the UK — generates a structured record of events: declarations before the race, results after it, and increasingly, granular sectional times captured at furlong intervals throughout.

The challenge is not the absence of data. It is the consistency, completeness and accessibility of that data at scale.

Decades of Racing History, One Consistent Schema

HRDB’s Hong Kong dataset covers race results back to 1979 — over four decades of racing history from the HKJC. UK and Ireland coverage extends to 2011. Each year in our archive conforms to the same relational schema: the same field names, the same foreign key conventions, the same encoding for gear, going, distance and class.

This matters enormously for longitudinal research. When you want to study course bias over ten seasons, or track a trainer’s record at a specific distance class, you need data that joins cleanly across time — not an archive of inconsistently formatted flat files.

What the Data Actually Contains

  • Meeting header — Venue, date, surface type, going condition, weather
  • Race entries — Every runner: horse, age, weight, barrier draw, gear changes, trainer, jockey
  • Result records — Finishing order, winning time, margins, disqualifications
  • Sectional times — Where available: split times per furlong from gates to finish line
  • Dividends — Win, place, exacta and combination pools (HK dataset)
  • Veterinary records — Veterinary notations for individual runners where disclosed

The Engineering Challenge of Making It Accessible

Collecting raw racing data is straightforward. Normalising it — resolving duplicate horse names across seasons, standardising course identifiers, handling mid-career trainer and owner changes, encoding going descriptions consistently across regulatory bodies — is where most of the engineering effort goes.

At HRDB, the data pipeline runs continuously. Each day’s results are ingested, validated, normalised and appended to the historical archive before delivery to clients. The API reflects the same canonical dataset as the bulk exports — there is no divergence between real-time and historical access.

Getting Access

Our datasets are available as daily-updated API feeds or as historical bulk exports in CSV, SQL, and direct database dump formats. Coverage spans Hong Kong and UK/IRE, with USA, Canada and global markets coming soon.

View available datasets or talk to us about enterprise data delivery.

No responses yet

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    0
      0
      Your Cart
      Your cart is emptyReturn to Shop