N
Nica Furs
Guest
For sportsbooks that boast technical capability, evolved odds calculation is not just a value-add but also a fundamental point of operation. Understanding how much one team is better than the others provides the base for calculating that initial, game-wide betting line. Bookmakers who can get correct odds out early have a big advantage as they receive more volume, better manage their risk, and improve margin control.
But what happens when the sport's nature defies traditional modeling? That's precisely the case in esports, a dynamic, constantly evolving ecosystem that requires far more adaptive modeling than traditional sports can offer.
In football, the Poisson distribution is commonly used to model the number of goals in a match — statistically independent events that occur within a fixed timeframe. The model estimates scoreline probabilities based on the average goal count expected for each team, balancing attacking and defensive strength. Esports challenges this approach in several key ways.
First, win conditions vary drastically between titles. In MOBAs like Dota 2 or League of Legends, teams win by destroying the opponent's main structure. In FPS (First-person Shooter) games such as CS2 or Valorant, the goal is to win a specific number of rounds. In fighting games, it's about depleting an opponent's health bar across timed rounds. These are not easily reduced to a common, repeatable scoring event.
Second, even when events analogous to goals exist — such as kills in a shooter — they don't always correlate directly with winning. Strategic elements like map control, economy management, and objective execution are often more decisive. Modeling these factors via Poisson and aggregating them into a win probability would be highly complex and likely inaccurate.
Third, the nature of strength fluctuation in esports is fundamentally different from traditional sports. In football, team strength evolves gradually through player transfers, long-term form, and coaching decisions. In esports, change happens rapidly. Meta shifts introduced by game patches can instantly alter team effectiveness. Individual form can rise or fall quickly, and roster changes are frequent. This volatility demands models that are far more dynamic.
Given this environment, it's essential to apply models that reflect the structure and demands of esports. This is why rating systems like Elo, Glicko, and TrueSkill are so valuable. These systems are inherently designed to track relative strength and update it dynamically after each match, making them more responsive to fluctuations in player performance, team composition, and meta context. They also offer a more stable foundation for calculating pre-match odds.
The Elo system, developed in the 1960s, was initially designed for chess. Each player receives a numerical rating. Before a match, an expected outcome is calculated based on the rating difference. If a lower-rated player wins, they gain more points, and their opponent loses an equivalent amount. The K-factor determines how much a rating changes after a game. A higher K-factor results in larger fluctuations (typically used for new players), while a lower one limits volatility (used for experienced players).
Elo is relatively simple to understand and implement. It remains widely used and serves as a foundation for many other systems. However, it doesn't account for rating reliability. A rating earned after 10 games is treated the same as if it were based on 1,000 games. It also fails to address inactivity or team-based dynamics, as it was designed for individual competition.
The Glicko system, developed in the 1990s, builds on Elo by adding a measure of confidence in a player's rating — called Ratings Deviation (RD). The more matches a player completes, the lower their RD. Inactivity causes RD to rise, reflecting growing uncertainty.
Glicko-2 introduces an additional parameter: rating volatility (σ-sigma). This measures how erratic a player's performance is over time. Players with unpredictable results or rapidly changing form will have higher volatility. This makes the system more responsive to shifts in actual skill.
Compared to Elo, Glicko more accurately reflects a player's strength by considering how stable and recent the data is. It's better suited to environments with varying experience levels and inconsistent participation. However, its calculations are more complex.
TrueSkill, developed by Microsoft in the 2000s, was built specifically for multiplayer and team-based formats. It uses Bayesian inference to model skill as a probability distribution defined by a mean (μ-mu, analogous to the rating) and a variance (σ²-sigma squared, which reflects uncertainty), capturing both expected performance and uncertainty.
After each match, these distributions are updated for all players based on the outcome, team compositions, and prior skill estimates. TrueSkill handles more than two players or teams, accounts for draws, and adjusts even when players come and go between matches.
This system is well-suited for team-based formats, effectively manages uncertainty, and provides statistically robust predictions. However, it is the most computationally intensive of the three and requires a firmer grasp of Bayesian statistics to implement correctly.
Choosing the right rating system is only the beginning. Achieving accuracy in odds-making requires deeper work, and the data science team plays a critical role in this process. In CS2, for example, understanding how teams approach map control can be as crucial as round count. In Dota 2, hero drafts and balance patches introduce layers of complexity that must be factored in. These game-specific dynamics require tailored data features engineered with deep domain knowledge. At DATA.BET, these nuances are addressed through carefully engineered data features informed by domain expertise.
Data scientists begin by analyzing key mathematical metrics after tuning the optimal rating system for predicting team strength. Log Loss, calibration curves (which compare predicted probabilities to actual outcomes), and the width of probability distributions are especially important. These metrics show how well a model is calibrated and whether its predictions are both confident and fair. Ultimately, they reduce to two business-critical qualities: accuracy and stability - the foundation of financial performance.
Contrasting this automated system with the previous approach, which relied on manual trader evaluations, is crucial. Based on subjective expertise, trader predictions yielded a profit (Yield) in the 0-4% range and even resulted in losses in some sports. Our system, grounded in mathematics and precise calculation, provides a more reliable and superior performance baseline. Furthermore, it mitigates a key psychological bias: human traders often fear being overconfident, which leads them to underestimate firm favorites systematically. This results in a narrower and less decisive probability distribution of their predictions than an objective model.
To turn this theoretical advantage into a practical result, we went through several stages of development. We started by analyzing baseline models to understand their strengths and weaknesses, and gradually moved towards creating a more sophisticated system. Our research reveals significant differences between them. As shown in Figure 1, the baseline Elo system, built on team data, is far from ideal; its predicted probabilities deviate significantly from the actual outcomes represented by the y=x line. In contrast, the TrueSkill system provides a much better description of team strength, confirmed by its higher accuracy. However, the hybrid system we have developed delivers even stronger results, providing more precise predictions and greater confidence in those probabilities.
The key question for the business is: what is the financial return on these improvements? Based on simulated betting using the models' predictions, our analysis shows stark contrasts. Elo produces negative profits, while TrueSkill offers modest but steady returns of 3–5% depending on the discipline. The hybrid system changes the equation entirely, delivering a 20–30% increase in profitability over TrueSkill. Therefore, choosing the hybrid system is both mathematically sound and financially wise, providing a clear competitive advantage.
The foundation of rating systems is to combine Bayesian principles with machine learning to build hybrid models. The models can detect subtle, nonlinear patterns in performance and adjust to changing environments. They are not static tools but evolving systems, tested against historical data, deployed in production environments, and continuously refined.
This approach allows us to build predictive infrastructure that doesn't just calculate odds, but responds in real time to competitive shifts — maintaining accuracy across titles, formats, and metas. The result is a living data-driven ecosystem built to support best-in-class pre-match pricing - the true math behind odds.
This article was published under HackerNoon's Business Blogging program.
But what happens when the sport's nature defies traditional modeling? That's precisely the case in esports, a dynamic, constantly evolving ecosystem that requires far more adaptive modeling than traditional sports can offer.
Why Esports Demands More Than Football's Poisson Model
In football, the Poisson distribution is commonly used to model the number of goals in a match — statistically independent events that occur within a fixed timeframe. The model estimates scoreline probabilities based on the average goal count expected for each team, balancing attacking and defensive strength. Esports challenges this approach in several key ways.
First, win conditions vary drastically between titles. In MOBAs like Dota 2 or League of Legends, teams win by destroying the opponent's main structure. In FPS (First-person Shooter) games such as CS2 or Valorant, the goal is to win a specific number of rounds. In fighting games, it's about depleting an opponent's health bar across timed rounds. These are not easily reduced to a common, repeatable scoring event.
Second, even when events analogous to goals exist — such as kills in a shooter — they don't always correlate directly with winning. Strategic elements like map control, economy management, and objective execution are often more decisive. Modeling these factors via Poisson and aggregating them into a win probability would be highly complex and likely inaccurate.
Third, the nature of strength fluctuation in esports is fundamentally different from traditional sports. In football, team strength evolves gradually through player transfers, long-term form, and coaching decisions. In esports, change happens rapidly. Meta shifts introduced by game patches can instantly alter team effectiveness. Individual form can rise or fall quickly, and roster changes are frequent. This volatility demands models that are far more dynamic.
Given this environment, it's essential to apply models that reflect the structure and demands of esports. This is why rating systems like Elo, Glicko, and TrueSkill are so valuable. These systems are inherently designed to track relative strength and update it dynamically after each match, making them more responsive to fluctuations in player performance, team composition, and meta context. They also offer a more stable foundation for calculating pre-match odds.
Elo: The Foundation of Rating Systems
The Elo system, developed in the 1960s, was initially designed for chess. Each player receives a numerical rating. Before a match, an expected outcome is calculated based on the rating difference. If a lower-rated player wins, they gain more points, and their opponent loses an equivalent amount. The K-factor determines how much a rating changes after a game. A higher K-factor results in larger fluctuations (typically used for new players), while a lower one limits volatility (used for experienced players).
Elo is relatively simple to understand and implement. It remains widely used and serves as a foundation for many other systems. However, it doesn't account for rating reliability. A rating earned after 10 games is treated the same as if it were based on 1,000 games. It also fails to address inactivity or team-based dynamics, as it was designed for individual competition.
Glicko and Glicko-2: Adding Confidence and Volatility
The Glicko system, developed in the 1990s, builds on Elo by adding a measure of confidence in a player's rating — called Ratings Deviation (RD). The more matches a player completes, the lower their RD. Inactivity causes RD to rise, reflecting growing uncertainty.
Glicko-2 introduces an additional parameter: rating volatility (σ-sigma). This measures how erratic a player's performance is over time. Players with unpredictable results or rapidly changing form will have higher volatility. This makes the system more responsive to shifts in actual skill.
Compared to Elo, Glicko more accurately reflects a player's strength by considering how stable and recent the data is. It's better suited to environments with varying experience levels and inconsistent participation. However, its calculations are more complex.
TrueSkill: Designed for Teams and Complex Formats
TrueSkill, developed by Microsoft in the 2000s, was built specifically for multiplayer and team-based formats. It uses Bayesian inference to model skill as a probability distribution defined by a mean (μ-mu, analogous to the rating) and a variance (σ²-sigma squared, which reflects uncertainty), capturing both expected performance and uncertainty.
After each match, these distributions are updated for all players based on the outcome, team compositions, and prior skill estimates. TrueSkill handles more than two players or teams, accounts for draws, and adjusts even when players come and go between matches.
This system is well-suited for team-based formats, effectively manages uncertainty, and provides statistically robust predictions. However, it is the most computationally intensive of the three and requires a firmer grasp of Bayesian statistics to implement correctly.
DATA.BET’s Hybrid Approach
Choosing the right rating system is only the beginning. Achieving accuracy in odds-making requires deeper work, and the data science team plays a critical role in this process. In CS2, for example, understanding how teams approach map control can be as crucial as round count. In Dota 2, hero drafts and balance patches introduce layers of complexity that must be factored in. These game-specific dynamics require tailored data features engineered with deep domain knowledge. At DATA.BET, these nuances are addressed through carefully engineered data features informed by domain expertise.
Data scientists begin by analyzing key mathematical metrics after tuning the optimal rating system for predicting team strength. Log Loss, calibration curves (which compare predicted probabilities to actual outcomes), and the width of probability distributions are especially important. These metrics show how well a model is calibrated and whether its predictions are both confident and fair. Ultimately, they reduce to two business-critical qualities: accuracy and stability - the foundation of financial performance.
Contrasting this automated system with the previous approach, which relied on manual trader evaluations, is crucial. Based on subjective expertise, trader predictions yielded a profit (Yield) in the 0-4% range and even resulted in losses in some sports. Our system, grounded in mathematics and precise calculation, provides a more reliable and superior performance baseline. Furthermore, it mitigates a key psychological bias: human traders often fear being overconfident, which leads them to underestimate firm favorites systematically. This results in a narrower and less decisive probability distribution of their predictions than an objective model.
To turn this theoretical advantage into a practical result, we went through several stages of development. We started by analyzing baseline models to understand their strengths and weaknesses, and gradually moved towards creating a more sophisticated system. Our research reveals significant differences between them. As shown in Figure 1, the baseline Elo system, built on team data, is far from ideal; its predicted probabilities deviate significantly from the actual outcomes represented by the y=x line. In contrast, the TrueSkill system provides a much better description of team strength, confirmed by its higher accuracy. However, the hybrid system we have developed delivers even stronger results, providing more precise predictions and greater confidence in those probabilities.
The key question for the business is: what is the financial return on these improvements? Based on simulated betting using the models' predictions, our analysis shows stark contrasts. Elo produces negative profits, while TrueSkill offers modest but steady returns of 3–5% depending on the discipline. The hybrid system changes the equation entirely, delivering a 20–30% increase in profitability over TrueSkill. Therefore, choosing the hybrid system is both mathematically sound and financially wise, providing a clear competitive advantage.
The foundation of rating systems is to combine Bayesian principles with machine learning to build hybrid models. The models can detect subtle, nonlinear patterns in performance and adjust to changing environments. They are not static tools but evolving systems, tested against historical data, deployed in production environments, and continuously refined.
This approach allows us to build predictive infrastructure that doesn't just calculate odds, but responds in real time to competitive shifts — maintaining accuracy across titles, formats, and metas. The result is a living data-driven ecosystem built to support best-in-class pre-match pricing - the true math behind odds.
This article was published under HackerNoon's Business Blogging program.