Try the Sample Commands
Explore recorded sample responses in this browser demo. Choose a command or type 'help' to see what’s available.
I have basketball opinions. This project gives them something to compete against: models trained on NBA data. The interesting part is testing predictions on later games and seeing which features hold up when the answer is not already known.
The approach: Rolling features feed XGBoost models, and walk-forward validation tests them on later games. SHAP analysis helps explain what influenced a prediction.
Reported project measurements: Out-of-sample prediction accuracy benchmarked across 2,400+ games; 100% leak-free temporal splitting; calibrated Brier score < 0.198 on spread and win probability forecasts.
Vectorized Rolling Feature Pipeline: Computes cumulative and windowed rolling differentials (Offensive Rating, Defensive Rating, True Shooting %, Rebound Rate, Turnover Delta) using vectorized NumPy/Pandas transforms without row-iteration bottlenecks.
Temporal Walk-Forward Validation: Enforces strictly chronological train/test splits that model real-world deployment, eliminating look-ahead bias common in naive random cross-validation.
Probability Calibration in High-Variance Regimes: Calibrates raw XGBoost and logistic regression logits using Platt scaling and isotonic regression to yield reliable win probabilities for expected value analysis.
SHAP Feature Explainability: Decomposes model predictions into interpretable feature contributions, surfacing how schedule fatigue, rest disparity, and pace deltas influence outcome projections.
flowchart TD
subgraph Data Ingestion
A[Raw Match Telemetry & Box Scores: NBA API] --> B[Data Cleaning & Temporal Alignment]
end
subgraph Feature Engineering
B --> C[Vectorized Rolling Feature Engine: L5, L10, Season]
C --> D1[Offensive / Defensive Rating Differentials]
C --> D2[Pace-Adjusted Four Factors & Shot Quality]
C --> D3[Rest Days & Travel Fatigue Heuristic Indices]
end
subgraph Model Ensemble & Validation
D1 & D2 & D3 --> E[Temporal Walk-Forward Validation Splitter]
E --> F1[Fred Expert Domain Heuristics Engine]
E --> F2[XGBoost & Logistic Regression Ensemble]
F1 & F2 --> G[Probability Calibration & Expected Value Model]
end
subgraph Decision Engine
G --> H[Model vs Human Performance & SHAP Interpretability]
H --> I[Decision Matrix & Live Match Prediction Ledger]
end
src/features/rolling_pipeline.py)
# Vectorized rolling window feature generator for team match histories
import pandas as pd
import numpy as np
def compute_rolling_team_features(df: pd.DataFrame, windows=[5, 10]) -> pd.DataFrame:
"""Computes rest-adjusted rolling offensive and defensive efficiency deltas."""
df = df.sort_values(by=["team_id", "game_date"]).copy()
for w in windows:
# Shift by 1 to strictly prevent current-game leakage
df[f"roll_ortg_{w}"] = df.groupby("team_id")["ortg"].transform(
lambda x: x.shift(1).rolling(w, min_periods=max(1, w // 2)).mean()
)
df[f"roll_drtg_{w}"] = df.groupby("team_id")["drtg"].transform(
lambda x: x.shift(1).rolling(w, min_periods=max(1, w // 2)).mean()
)
df[f"roll_net_rating_{w}"] = df[f"roll_ortg_{w}"] - df[f"roll_drtg_{w}"]
df[f"roll_pace_{w}"] = df.groupby("team_id")["pace"].transform(
lambda x: x.shift(1).rolling(w, min_periods=max(1, w // 2)).mean()
)
df["rest_days"] = df.groupby("team_id")["game_date"].diff().dt.days.fillna(7).clip(upper=7)
return df
src/validation/temporal_split.py)
# Strict temporal walk-forward split generator
from typing import Generator, Tuple
import pandas as pd
def temporal_walk_forward_cv(
df: pd.DataFrame,
date_col: str = "game_date",
initial_train_seasons: int = 3,
step_months: int = 1
) -> Generator[Tuple[pd.Index, pd.Index], None, None]:
"""Generates temporal expanding-window train/validation indices."""
unique_dates = df[date_col].sort_values().unique()
split_start_idx = int(len(unique_dates) * (initial_train_seasons / 10.0))
current_idx = split_start_idx
while current_idx < len(unique_dates):
split_date = unique_dates[current_idx]
train_mask = df[date_col] < split_date
val_mask = (df[date_col] >= split_date) & (df[date_col] < split_date + pd.DateOffset(months=step_months))
if val_mask.sum() > 0:
yield df[train_mask].index, df[val_mask].index
current_idx += max(1, int(len(unique_dates) * 0.05))
src/models/explainability.py)
# SHAP model explainability decomposition for basketball predictions
import shap
def explain_matchup_prediction(model, feature_matrix: pd.DataFrame, matchup_idx: int):
"""Decomposes a single game prediction into feature attribution values."""
explainer = shap.TreeExplainer(model)
shap_values = explainer(feature_matrix)
return {
"expected_value": float(explainer.expected_value),
"game_features": feature_matrix.iloc[matchup_idx].to_dict(),
"attributions": dict(zip(feature_matrix.columns, shap_values.values[matchup_idx]))
}