Methodology

How the numbers are calculated

Plain explanations of every figure Sport Mule shows, and the rules behind them.

Where the numbers come from

Every figure on Sport Mule is calculated from stored historical matches, markets and prices by one deterministic engine. Nothing is estimated, modelled or generated. Pattern Discovery, Pattern Details and Backtesting all read the same stored results, so the same criteria always produce the same numbers.

Qualifying matches

A qualifying match is a completed historical match that matches the selected criteria and has a stored, settled market result. Matches without a settled result for the selected market are excluded rather than assumed.

Historical hit rate

Historical hit rate counts each settled market result once: wins divided by wins plus losses, with pushes kept separate rather than counted as either. It describes what happened in the stored sample. It is not a probability and not a forecast.

Pushes

A push is a settled result where the selection neither won nor lost — for example a handicap landing exactly on the line, or a drawn two-way market in leagues where a draw voids the market. Pushes are reported separately and never counted as wins.

Odds and historical ROI

Historical ROI treats every stored bookmaker price as its own observation and applies one reference unit to each. It is never derived from the displayed average odds, which is descriptive only. ROI is withheld below 30 settled observations, and a limited-sample note is shown below 30.

Backtesting order

Backtests walk qualifying matches in chronological order using only information available at the time of each match. There is no look-ahead, no re-ordering by result and no artificial smoothing of the cumulative return.

Sample size

Sample size is shown alongside every figure. Historical rates, frequencies and percentages are withheld below 30 observations. Large samples are 75 or more settled observations and moderate samples are 30 to 74; smaller samples keep their raw counts but do not present a percentage as a meaningful result.

Live prices

Current prices for upcoming fixtures are shown separately from stored historical data. They are reference prices for upcoming matches, are never mixed into historical figures and are never stored.

Research: Hot and Cold

Hot means a player or team has performed clearly above its own season average over its last 10 games: at least 1 unit and at least 15% higher, with at least 15 games played this season. Cold is the exact opposite — at least 1 unit and 15% below its own average. A low number on its own is never called Cold.

Research: Role Change

Role Change compares a player's last 10 games with the 10 games just before them, using a usage measure such as minutes, targets, carries, pass attempts, time on ice, plate appearances, pitching outs, hitouts or clearances. It needs all 20 games and a change of at least 20%, and always shows the before and after figures. The 20 games are counted, not taken from a fixed number of days: they are the player's 20 most recent completed games, even across an off-season (important for NFL, which plays about 17 games a year). If fewer than 20 qualifying games are saved, no result is shown — the sample is never padded. For Football minutes, both 10-game periods must average at least 45 minutes, so short substitute appearances alone cannot create a Role Change.

Research: Opportunity

Opportunity looks at upcoming saved fixtures only. It appears when a team has scored at least 10% above its competition average over its last 10 games and its opponent has conceded at least 10% above that average over its own last 10, with at least 20 games played in the competition. It describes an unusual matchup, not a recommendation.

Research: Market Value

Market Value compares a current saved price (checked within the last 4 hours) with how often that exact outcome at that exact line has happened in stored results. The implied probability is 1 divided by the decimal price; the historical hit rate counts wins out of wins plus losses, with pushes left out. A market appears when there are at least 20 decided results and the historical hit rate is at least 5 percentage points above the implied probability. It is separate from Best Bets.

What Sport Mule is not

Sport Mule is a research and analytics tool. It is not a sportsbook. No wagering, deposits or withdrawals are offered or facilitated, and nothing shown is a recommendation or a prediction. Historical results do not guarantee future outcomes.

Historical probability

Worked example from stored AFL games: Lachie Neale (Brisbane Lions), 25 or more disposals. The percentage is simply successful games ÷ qualifying games — no smoothing or adjustment.

2024 season

Calculating…

Last 15 games

Calculating…

Last 7 games

Calculating…

Last 4 games

Calculating…

Model prediction

A separate measurement from historical probability. The Elo baseline rates every team from stored results in date order, using only matches completed before kickoff. It is a first baseline, not a tuned model.

AFL — latest stored match

Calculating…

Premier League — latest stored match

Calculating…

Model evaluation

  • Historical probability — historical frequency of qualifying outcomes.
  • Model prediction — a model estimate generated using information available before the match.
  • Model evaluation — how those predictions performed historically.

Two baseline models are backtested in date order: each match is predicted from earlier matches only, the result is revealed afterwards, then the model moves on. Elo v1.0 rates team strength; Poisson v1.0 estimates expected goals from recent goals scored and conceded, then derives every scoreline. Accuracy counts how often the highest probability matched the result. Log loss is the average −ln of the probability given to the actual result (bounded away from 0 and 1). Brier score sums the squared gap between each outcome's probability and what happened. Lower log loss and Brier mean probabilities sat closer to results. Calibration compares predicted probabilities with how often those outcomes occurred.

Running backtest over stored matches…

Ensemble research · experimental

Research only, not a production prediction. Ensemble v1.0 combines probabilities already produced before kickoff by Elo v1.0, Poisson v1.0 and Logistic Regression v1.1 (v1.1A, monthly). The component models are not changed. Equal Weight averages the three and renormalises to 100%. Learned Weight picks non-negative weights summing to 1 (each at most 0.8) that gave the lowest log loss on earlier months only; below 380 earlier matches it falls back to equal weights. The calibrated variant learns a single sharpening exponent from earlier months only. Nothing uses the match being predicted or anything after it.

Not run on page load. Start an explicit run (about a minute); results are cached.

Feature-based models · Logistic Regression v1.0 · Gradient Boosting v1.0

A multinomial logistic regression estimates home, draw and away probabilities from point-in-time features: team ratings, recent form (last 5 and 10), home/away records, head-to-head, rest days and current-season results. Every feature uses only matches completed before kickoff. It is trained walk-forward: at the start of each month it is retrained on matches completed before that month, then predicts that month's matches. Missing features are never treated as zero — they take the training average and switch on a missing flag. No prediction is made until 760 earlier matches exist. Five feature variants (A–E) are measured on the same matches. Gradient Boosting v1.0 uses the same 41 features, the same monthly walk-forward retraining and the same 760-match minimum: 50 rounds of depth-3 trees, learning rate 0.1, 16 value bins per feature fitted on training data only, with missing values kept in their own bin. It has no random elements, so the same data always gives the same result. Measurements only — not a ranking.

Running walk-forward backtest over stored matches — this can take a little while the first time…

Football research · Premier League

Model diagnostics

Research and development information, not a recommendation. These experiments re-run the Logistic Regression v1.0 walk-forward backtest with one controlled change at a time — removing a feature or group, shuffling one feature, trying an alternative form calculation, or changing how often the model retrains — and report how the measured results moved. Every experiment uses only information available before each kickoff. The production models are not changed. A positive log-loss change means results were worse without the feature (lower log loss is better). Associations only — not causes.

Diagnostics are not run on page load. Start an explicit run to calculate them (several minutes); results are cached.

Logistic Regression v1.1 · experimental

Experimental research model, separate from Logistic Regression v1.0. v1.1 uses a simplified, less redundant feature set derived from the Model Diagnostics measurements: Elo difference plus one last-10 recent-form window. Same algorithm, missing-data treatment, minimum history and point-in-time walk-forward rules as v1.0. v1.0, Elo v1.0 and Poisson v1.0 are unchanged and shown as benchmarks. Measurements only — no model is selected.

Not run on page load. Start an explicit run (a few minutes); results are cached.

Temporal validation · research

Research only. Re-scores the existing walk-forward predictions of Elo v1.0, Poisson v1.0, Logistic Regression v1.1 and Ensemble v1.0 by chronological period to see whether measured results are consistent over time. Nothing is re-trained or tuned: every prediction was already made from information available before its kickoff. No model is selected.

Not run on page load. Start an explicit run; results are cached.

Recency & Model Adaptation Research · experimental

Experimental research only. Tests whether giving earlier matches less weight during training changes out-of-sample results. Logistic Regression v1.2 uses the v1.1 features unchanged; Poisson v1.1 weights past scoring by age; the adapted ensemble averages Elo v1.0 with both at equal weight. Existing models are unchanged and no model is selected.

Weighting schemes (fixed before evaluation)

weight = exp(−λ × age in days), λ = ln 2 ÷ half-life. LR: age measured from the start of the monthly retrain period. Poisson: age measured from the target kickoff. Only matches completed before kickoff receive a weight.

SchemeHalf-life (days)λ per dayWeight at 1 year
A · No weightingnone0.0000001.000
B · Mild7300.0009500.707
C · Moderate3650.0018990.500
D · Strong1800.0038510.245
Not run on page load. Start an explicit run; results are cached.

AFL research

AFL temporal validation · research

Research only. Applies the existing temporal validation to stored AFL matches using only models already validated for AFL. Every prediction uses information from before kickoff. No model is created, tuned or selected.

Not run on page load. Start an explicit run; results are cached.

AFL research

AFL Elo v2.0 — Draw Model · experimental

Experimental. A three-outcome AFL baseline (home win, draw, away win). Ratings are identical to AFL Elo v1.0, which stays the benchmark; v2.0 only adds a draw probability. Not a production model.

How the draw probability is calculated
  • e = v1.0 home-win expectation from pre-match ratings (home advantage 40, scale 400, K 32).
  • Closeness c = 4 × e × (1 − e): 1 for an evenly rated match, near 0 for a mismatch.
  • League draw rate r = (earlier draws + 1) ÷ (earlier matches + 100) — only matches completed before kickoff.
  • Draw = r × c ÷ (mean c of all earlier matches), capped at 5%. Home = (1 − draw) × e; away = (1 − draw) × (1 − e).
  • Shown to 0.1% and summing exactly to 100.0%, because AFL draw probabilities are around 1%.
Not run on page load. Start an explicit run; results are cached.

Prefer to explore? Browse historical trends or search stored patterns.

Research only · No wagering · Past results are not predictions.

© 2026 Sport Mule