A scoreboard tells you who won the last point, not who is in control or whether the momentum has turned. We built the models behind the US Open's prediction engine, turning years of match data and a live feed of every shot into a win probability that updates after every point.
Note on this write-up. The IBM–USTA partnership is public, but the internal model design, variables and figures shown here are sanitized recreations built for this portfolio, directionally accurate rather than the exact production system.
The same live win-probability line that fans watched during the 2015 Williams–Vinci upset: a story the scoreboard could not tell. Each point feeds serve speed, rally length, unforced errors and a hundred other signals back into the model, which re-reads the match and updates who is really in control.
During the US Open, a fan opening the app can see more than a dozen matches at once, and the score alone never answers the question they actually have: which match deserves my attention right now? A player can be down two sets and quietly building the biggest comeback of the tournament, and nothing in the scoreline shows it.
Coaches had the mirror-image problem. The information they needed to prepare for an opponent sat buried in hours of footage that a human had to watch and tag by hand. In both cases the raw data existed, and the intelligence did not. The work was to turn millions of data points into foresight, before and during the match, at tournament scale.
Before a ball is struck, the model already has a view. It learns from years of Grand Slam history, tens of millions of past points, blended with current ATP and WTA rankings, head-to-head records, and even the language of expert and press coverage read from thousands of articles. From that it forms a baseline: how this matchup tends to go, and how much of an edge each player carries into it.
"It is not a prediction of who will win. It is a read on what each player has to do well to earn it."
A win probability is interesting; a plan is useful. So the model distills each matchup into a short set of keys, the handful of performance targets that history says decide this particular contest. There are more than two dozen candidate keys, and only the few that actually move the outcome for these two players surface. As the match unfolds, each player's real performance is tracked against their keys in real time.
Bars show live performance; the tick marks the target the model set. Hitting the keys and the win probability follows.
Once play begins, a second model takes over. Every point produces more than a hundred and fifty signals, serve speed, aces and double faults, unforced errors, rally length, distance covered, return and volley patterns, and all of it flows in continuously. After each point the win probability is recomputed, so a momentum shift shows up the instant it starts rather than after the set is lost.
The models did not only serve fans. The same footage and telemetry were turned into scouting and performance analysis for players and their coaches: automatically tagging the key moments of a match and assembling the report a coach would otherwise spend hours cutting by hand. Preparation that used to take an afternoon of manual review could be ready in minutes, which is exactly the kind of edge a touring professional and their team plan around.
The win-probability line became one of the tournament's signature features and a reference point for how live sport uses prediction. Its most-watched moments belong to the players themselves, from Serena Williams' epic runs to the upsets nobody's scoreboard saw coming, and the work became a flagship of IBM Consulting's sports & entertainment portfolio.