Comparison of artificial intelligence algorithms for predicting breeding value of Performance Traits in Turkmen Racing Horses

Document Type : Research Paper

Authors

1 Department of Animal Science, Faculty of Agriculture, University of Zanjan, Zanjan, Iran.

2 Department of Animal Science, University of Zanjan, Zanjan, Iran

3 Department of Psychology, University of Freiburg, Freiburg, Germany.

Abstract

Introduction: Accurate estimation of breeding values (EBVs) plays a crucial role in enhancing the genetic progress of livestock populations. Performance traits in racehorses, including race completion time (RCT), rank at the end of competition (REC), and average speed (AS), are directly associated with competition success, economic returns, and the overall efficiency of the horse racing industry. Traditional methods for predicting EBVs, such as single-trait animal models (STM) using linear mixed models, provide unbiased estimates but require extensive pedigree information, high computational capacity, and specialized statistical knowledge. With the increasing availability of large phenotypic and pedigree datasets, conventional approaches may be insufficient to capture the complex genetic and environmental interactions affecting performance traits. Machine learning (ML) algorithms, with their capacity to model nonlinear relationships and high-order interactions, offer an efficient alternative for predicting genetic merit and accelerating breeding programs. However, comprehensive comparisons of different supervised ML methods for predicting EBVs for performance traits in Iranian Turkmen racehorses have not yet been conducted. This study aims to evaluate and compare the performance of a single-trait animal model with a suite of supervised ML algorithms in predicting EBVs for RCT, REC, and AS in Turkmen racehorses.
Materials and Methods: Phenotypic and pedigree data were collected from the official database of the Equestrian Federation of the Islamic Republic of Iran over the period from 1999 to 2025, initially comprising 11,008 race records from 1,678 Turkmen horses (902 stallions and 776 mares). After data cleaning and outlier removal using the interquartile range method, complete performance records for RCT, REC, and AS were retained. Environmental factors, including sex, race distance, handicap, jockey, and race location, were recorded and accounted for. Breeding values for RCT, REC, and AS were estimated using eight variants of single-trait animal models incorporating additive genetic effects, permanent environmental effects, jockey effects, and Handicap–Race Condition (HC) effects. Gibbs sampling with 150,000 cycles, a burn-in period of 20,000, and a sampling interval of 10 was used to estimate variance components and EBVs. The best-fitting model for each trait was selected based on the Deviance Information Criterion (DIC). For ML model development, phenotypic variables and environmental factors were used as predictors. Features were standardized using Z-score normalization. EBVs estimated from the STM models were used as target labels for supervised learning. The dataset was split into training (75%) and testing (25%) subsets, with horses born before 2019 assigned to the training set and those born from 2019 onwards assigned to the testing set. Thirteen supervised ML algorithms were trained and evaluated, including linear regression (LR), multiple linear regression (MLR), Bayesian ridge regression (BRR), decision tree (DT), random forest (RF), extremely randomized trees (ERT), support vector regression (SVR), k-nearest neighbors (KNN), gradient boosting regression (GBR), XGBoost (XGB), CatBoost (CBR), LightGBM (LGBM), and artificial neural networks (ANN). Hyperparameters were tuned using grid search for ANN and Bayesian optimization for other models. Model performance was assessed using accuracy, correlation, bias, dispersion, mean absolute error, and mean squared error.
Results and Discussion: Variance component analysis indicated that models incorporating the Handicap–Race Condition random effect (particularly M8 for RCT, REC, and AS) provided the best fit. The HC effect explained approximately 79% of phenotypic variance for RCT, 76% for AS, and 3% for REC, while additive genetic variance was low across all traits (h² ≈ 0.003–0.06). The jockey effect explained approximately 7% of phenotypic variance for REC but only 0.3% for RCT and AS, highlighting the substantial influence of rider skill on competition ranking and the dominant effect of race conditions on race completion time and average speed. Performance comparison of ML models demonstrated that all ML algorithms substantially outperformed the STM. Artificial neural networks achieved the highest accuracy (RCT: 0.49, REC: 0.55, AS: 0.47) and correlation (RCT: 0.62, REC: 0.66, AS: 0.59) with minimal bias and the lowest mean squared error. Ensemble and gradient boosting methods, including XGBoost (RCT: 0.47, REC: 0.53, AS: 0.45), LightGBM (RCT: 0.46, REC: 0.52, AS: 0.44), and gradient boosting regression (RCT: 0.45, REC: 0.51, AS: 0.43), also demonstrated superior predictive ability compared to STM. Random forest provided stable improvements (RCT: 0.45, REC: 0.50, AS: 0.43). Linear models including LR, MLR, and BRR showed moderate improvements over STM, reflecting their limited ability to model nonlinear relationships. KNN and DT were among the weakest models. The low heritability estimates for all three traits suggest that performance traits in Turkmen horses are strongly influenced by environmental and management factors, and genetic improvement through selection may be slow unless environmental variance is carefully managed.
Conclusions: This study demonstrates that performance traits in Turkmen racehorses are influenced by both additive genetic effects and, more predominantly, by environmental factors including jockey skill and race conditions. While single-trait animal models provide unbiased EBV estimates, their predictive accuracy can be substantially enhanced using advanced ML algorithms. Artificial neural networks and ensemble methods, particularly XGBoost and LightGBM, achieved the highest predictive accuracy, correlation, and stability, highlighting their suitability for genetic evaluation in horse breeding programs. These findings indicate that integrating machine learning with traditional genetic evaluation can significantly improve the precision and efficiency of selection programs, enabling faster genetic progress and more effective management of racing performance traits in Turkmen horses.

Keywords

Main Subjects