PitWall AI

How PitWall AI works

A strategy engine built the way a team's would be — a transparent physical lap-time model, calibrated with uncertainty, solved exactly, stress-tested by simulation — and evaluated honestly on races it had not seen.

Lap-time model

tlap = basedriver + offsetcompound + wearcompound,team · age · (1 + φ · fuelremaining) + q · age² − fuelburn · lap + ε

Fuel burn is identified from race data because tyre age resets at each stop while lap number does not (fitted ≈ 0.04–0.07 s/lap in 2026).

Pipeline

  1. 1Ingest

    FastF1 timing for every 2026 session (practice, sprint, qualifying, race). Live: OpenF1 or F1's live-timing feed.

  2. 2Season prior

    Fit the lap-time model to every earlier race: driver pace, compound offset, wear rate, fuel burn. Between-race spread = prior uncertainty.

  3. 3Weekend evidence

    Detect long runs in FP/sprint (robust Theil–Sen slopes), short-run compound gaps, then map practice → race with transfer ratios learned on earlier weekends.

  4. 4Bayesian blend

    Conjugate Gaussian update per compound; team effects shrunk toward the field (hierarchical); compound order enforced by isotonic projection.

  5. 5Revealed preference

    Inverse optimisation: learn soft-tyre penalty, fuel-wear sensitivity and per-stop track-position cost that best reproduce teams' past choices.

  6. 6Optimise

    Exact dynamic programming over every 0–3 stop compound sequence and every pit lap, with FIA two-compound rule and stint-life limits.

  7. 7Simulate risk

    Monte Carlo over Safety Cars / VSCs (per-lap hazards) and posterior parameter draws, with a reactive pit-wall policy. Reports P(best), regret, CVaR.

  8. 8Live loop

    Every lap: update wear from all cars' green laps, re-estimate pit loss, re-optimise from the car's tyre state, check undercuts, rejoin slot and model drift.

Machine-learning techniques, and why each is here

Hierarchical Bayesian updating
Practice, season and live evidence are combined with explicit uncertainty; sparse teams borrow strength from the field instead of overfitting.
Walk-forward evaluation
Every round is predicted with models trained only on earlier rounds. No random splits, no leakage — the original project's random-split R² of 0.66 vs future-race R² < 0 is exactly why.
Split-conformal prediction intervals
First-stop intervals come from calibrated residual quantiles on earlier races, with the finite-sample correction; coverage is reported, not assumed.
Learned transfer functions
Practice wear is not race wear. The practice→race ratio and its error are regressed from earlier weekends, so the model knows how much to trust FP2.
Hybrid physics + ML
A walk-forward multinomial logistic regression learns how teams actually choose stop count and starting tyre from physics features (1- vs 2-stop gap, optimal stop lap, wear, grid slot); the optimiser then times the plan inside that class. 'Fastest' and 'likely' are reported separately.
Personal wear (partial pooling)
Every driver has their own wear multiplier. Variance components were measured on 2026 races; the per-stint weight learned on rounds 1–8 beat the field average on rounds 9–15 (MAE 0.531 vs 0.557), while raw per-driver numbers did worse. Pre-race history/practice did not beat the field, so personalisation comes from the driver's own race stints.
Weekend severity
Compounds without practice long runs inherit the weekend's overall severity (× typical compound ratio) instead of a low-wear season average.
Stop-count odds
Monte Carlo reports how often 1, 2 or 3 stops is fastest, so knife-edge decisions are shown as odds rather than a single confident plan.
Inverse optimisation
Teams' decisions are treated as demonstrations; behaviour parameters are fit so the optimiser reproduces them (inverse-RL in spirit, but tractable).
Exact DP instead of RL
The decision problem is small enough to solve optimally in ~10 ms, so no policy-gradient approximation is needed. Uncertainty lives in the parameters, handled by Monte Carlo.
Physics-informed constraints
Isotonic projection forces softer compounds to be faster when new and to wear faster — fixes inverted estimates caused by selection effects.
Drift monitoring
Live residuals of the last laps vs the predictive band raise an alert when the tyre leaves the model (cliff, graining, damage).
Tool-grounded LLM
The chat assistant can only cite numbers it obtained by calling the engine (function calling), so it cannot invent strategy figures.
Experiment tracking
Backtests are logged to MLflow (kept from the original pipeline) for reproducible comparisons.

Validation

Loading…

Known limitations

  • Track position is modelled only as a per-stop cost plus live rejoin/undercut checks — there is no full multi-car race simulation with overtaking probabilities.
  • Pre-race prediction of team behaviour is only modestly better than a naive baseline on some metrics and worse on others (see the table) — e.g. the first-stop lap is still predicted less accurately than the season median. Team orders, incidents and strategic covering are not observable in public timing data.
  • Wet races are excluded from dry-strategy metrics; intermediate/wet tyre strategy is not modelled.
  • Tyre wear is linear-plus-quadratic with a fuel-load interaction; real thermal degradation and cliffs are approximated through stint-life limits and the soft-tyre penalty.
  • Sepang (Round 16) has no recent F1 data; pit loss falls back to the 2026 season median until stops are observed live.