HalfCourt
An NBA halftime analyst with a strict division of labor
Give HalfCourt a game ID at halftime and it returns two things: a number and a story. A trained model predicts who wins and by how much. An LLM then writes the broadcast-style halftime report explaining why that outcome makes sense — and it is not permitted to disagree with the model.
That constraint is the whole project.
Why the constraint matters
Ask a language model to predict a basketball game and it will happily give you a percentage. It will sound authoritative. The percentage is fabricated — not in the sense of being random, but in the sense that nothing in the model's process actually estimated a probability. It pattern-matched the shape of confident sports analysis and filled in a plausible-looking number. You cannot calibrate that, you cannot audit it, and it will be wrong in ways that correlate with how famous the teams are.
But LLMs are genuinely excellent at the other half of the job. Turning "the home team is shooting 31% from three, their starting center has four fouls, and the bench has been outscored by fourteen" into three paragraphs that sound like a person who watches basketball — that's a real capability, and one that a gradient-boosted tree will never have.
So I split it. The model does arithmetic it's qualified to do. The LLM does language it's qualified to do. Neither one is asked to cover for the other, and the prompt gives the LLM no room to renegotiate the prediction it's been handed. Its job is explanation, not forecasting.
This sounds like a small architectural note. In practice it's the difference between a system whose predictions can be evaluated and a system that just sounds good.
How the pieces fit
The pipeline runs in one direction and the seams are deliberately narrow.
Collection. Per-player Q1 and Q2 stats pulled from the NBA API for any game, historical or in progress.
Feature engineering. Roughly 43 halftime features — score state, shooting efficiency splits, pace, bench contribution, foul trouble, star player availability, rest days. Most of the actual thinking in this project lives here rather than in the model architecture, which is usually how it goes.
Prediction. A trained classifier returns a winner and a win probability, and depending on which model you're using, a point margin.
Attribution. SHAP values are extracted and the top three are pulled out.
Narration. Model output plus SHAP attributions plus the raw halftime stats go to the LLM, which writes four sections: score and momentum, key players, X-factors and foul trouble, and a second-half outlook.
SHAP is the handoff
This is the join I'm most pleased with. Without it, the LLM would be handed a bare probability and asked to invent a reason for it — which puts you right back in fabrication territory, just one layer downstream. It would grab whatever stat looked most dramatic and build a story around it, and that story would have no relationship to what the model was actually responding to.
Passing the top-3 SHAP features means the explanation is anchored to the mechanism. When the report says the prediction is being driven by second-quarter bench scoring, that's not the LLM's editorial judgment — it's what moved the model. The prose becomes a readable surface over a real attribution rather than a plausible-sounding cover story.
Three models, three levels of honesty
The system supports an MLP, XGBoost, and logistic regression, and they are not equally capable — which the reports say out loud.
The MLP is the default and the only one that genuinely predicts margin: it has a dedicated margin head trained jointly with the win probability head.
XGBoost predicts win probability only. Margin gets estimated by mapping from that probability with a linear fit, which is a reasonable approximation and definitely worse than a model that was actually trained to predict margin. The report doesn't pretend otherwise.
Logistic regression is the baseline. Win probability, nothing else. When it's the active model, the generated report states plainly that margin prediction isn't available rather than quietly filling the field with something.
Building the limitation into the output rather than into a footnote in the docs was intentional. The report is the artifact people actually read; it should be the thing that tells the truth about what produced it.
Problems worth talking about
The data collection is the unglamorous majority of the work. Building the training set means hitting the NBA API for every game across multiple seasons — around 1,200 games a season, two to three hours of wall-clock time per season, subject to every rate limit and transient failure a public API can produce. The fix isn't clever, it's just necessary: cache aggressively at the raw layer so an interrupted run resumes instead of restarting. Nobody puts this in a demo. It's most of the project.
Halftime is exactly when the stats API is least useful. The natural moment to run this tool is during the break — and that's precisely when the historical stats endpoint returns incomplete data, because the game isn't finished and the aggregates haven't settled. The live path had to be rebuilt against the NBA CDN live endpoint instead, which is a different response shape feeding the same feature pipeline. Live snapshots get cached per game ID, and re-fetching means deleting the cache file, which is a rough edge I left in on purpose — during a real break you usually want the snapshot frozen, not silently changing under you.
Two data sources, one feature contract. The consequence of the above is that two quite different ingestion paths have to produce identical feature vectors, or the model is quietly being fed something it wasn't trained on. This is the part of the codebase where a bug would be hardest to notice and most damaging, because nothing would crash — you'd just get worse predictions.
What I'd push on next
The baseline number is the one I'm missing. Test accuracy sits around 62–67% across all three models, which is a real signal but not a self-explanatory one. The comparison that matters isn't published all-quarters models — those see the third quarter and are solving an easier problem, so beating them was never the goal. The comparison that matters is the naive rule: whoever leads at halftime wins. Until that number is computed on the same test split, 65% doesn't mean anything, and it's the first thing anyone who knows basketball will ask. It's a straightforward thing to measure from the dataset I already have and it belongs in the report header.
Calibration, not just accuracy. A win probability that reads 78% should be right about 78% of the time. Accuracy doesn't tell you that; a reliability curve does. For a system whose output is a probability shown to a reader, this matters more than another point of accuracy.
The LLM inherits the model's confidence, including when it's wrong. The design says the LLM never contradicts the model, and that's the right call — but it means a confidently wrong prediction becomes confidently wrong prose. The report will build a fluent, persuasive case for an outcome that doesn't happen. The mitigation is on the model side, in surfacing uncertainty the narration is then required to carry, rather than in letting the writer hedge freely.
Stack
Python, nba_api for collection, pandas and numpy for feature engineering, PyTorch / XGBoost / scikit-learn for the three models, SHAP for attribution, and a swappable LLM layer over OpenAI, Anthropic, or Gemini. Four CLI commands: collect, train, predict, predict-live.