Explore recorded sample responses in this browser demo. Choose a command or type 'help' to see what’s available.
A risk model needs more scrutiny than a good leaderboard score. This project looks at cardiovascular data through preprocessing, cross-validation, calibration, and feature explanations, with particular attention to leakage and changes in the data distribution.
The approach: Scikit-Learn preprocessing feeds XGBoost and LightGBM models. Cross-validation, probability calibration, and SHAP analysis provide different views of model behavior.
Reported project measurements: Out-of-fold AUROC of 0.894 with Brier calibration score < 0.088; 100% leak-free temporal & group stratified validation; sub-5ms single-patient risk scoring inference.
Strict In-Fold Preprocessing: Enforces zero feature leakage by executing robust quantile scaling, iterative multivariate imputation, and one-hot encodings strictly inside cross-validation splits.
Ensemble Model Stacking: Blends gradient-boosted decision trees (XGBoost, LightGBM, CatBoost) with regularized logistic regression baselines.
Probability Calibration Engine: Applies Platt scaling (logistic sigmoid calibration) and isotonic regression to convert raw logit margins into well-calibrated posterior probabilities suitable for clinical decision curves.
SHAP Attribution & Feature Importance: Computes TreeSHAP feature interactions to explain individual patient risk factor contributions (e.g. ST depression delta, exercise-induced angina, age-adjusted max heart rate).
flowchart TD
subgraph Data Ingestion & Preprocessing
A[Raw Patient Data: Clinical CSV] --> B[Exploratory Data Analysis & Outlier Truncation]
B --> C[Scikit-Learn ColumnTransformer Pipeline]
C --> D1[Continuous Numerical: RobustScaler / QuantileTransformer]
C --> D2[Categorical Features: OneHotEncoder]
end
subgraph Cross-Validation & Model Ensembling
D1 & D2 --> E[Stratified K-Fold CV Splitter]
E --> F1[XGBoost Classifier]
E --> F2[LightGBM Gradient Booster]
E --> F3[Calibrated Logistic Regression Baseline]
F1 & F2 & F3 --> G[Model Evaluation & Hyperparameter Search: Optuna]
end
subgraph Calibration & Explainability
G --> H[Platt Probability Calibration]
H --> I[SHAP Feature Importance & Attribution Engine]
I --> J[Actionable Clinical Risk Score Matrix]
end
src/pipeline/oof_validator.py)
# Strict In-Fold Preprocessing & Ensemble OOF Validator
import numpy as np
from sklearn.model_selection import StratifiedKFold
from sklearn.impute import IterativeImputer
from sklearn.calibration import CalibratedClassifierCV
import lightgbm as lgb
class LeakFreeClinicalPipeline:
def __init__(self, n_splits: int = 10, random_state: int = 42):
self.n_splits = n_splits
self.random_state = random_state
self.oof_predictions = None
self.models = []
def fit_predict_oof(self, X: np.ndarray, y: np.ndarray) -> np.ndarray:
skf = StratifiedKFold(n_splits=self.n_splits, shuffle=True, random_state=self.random_state)
self.oof_predictions = np.zeros(len(X))
for fold, (train_idx, val_idx) in enumerate(skf.split(X, y)):
X_train, y_train = X[train_idx], y[train_idx]
X_val, y_val = X[val_idx], y[val_idx]
# Imputation transformer fit strictly on train fold
imputer = IterativeImputer(max_iter=10, random_state=self.random_state)
X_train_imp = imputer.fit_transform(X_train)
X_val_imp = imputer.transform(X_val)
# Gradient Boosted Classifier with early stopping
clf = lgb.LGBMClassifier(
n_estimators=1000,
learning_rate=0.03,
num_leaves=31,
random_state=self.random_state + fold
)
clf.fit(
X_train_imp, y_train,
eval_set=[(X_val_imp, y_val)],
callbacks=[lgb.early_stopping(50, verbose=False)]
)
# Calibrate probabilities via Isotonic Regression / Platt Scaling
calibrator = CalibratedClassifierCV(clf, cv="prefit", method="isotonic")
calibrator.fit(X_val_imp, y_val)
self.oof_predictions[val_idx] = calibrator.predict_proba(X_val_imp)[:, 1]
self.models.append(calibrator)
return self.oof_predictions
src/explainability/shap_engine.py)
# Patient-level SHAP explainability engine
import shap
import pandas as pd
def compute_patient_shap_waterfall(model, background_data: pd.DataFrame, patient_features: pd.DataFrame):
"""Computes TreeSHAP attribution values and expected base value for a patient."""
explainer = shap.TreeExplainer(model, background_data)
shap_values = explainer(patient_features)
top_risk_drivers = pd.DataFrame({
"feature": patient_features.columns,
"patient_value": patient_features.iloc[0].values,
"shap_attribution": shap_values.values[0]
}).sort_values(by="shap_attribution", ascending=False)
return {
"base_risk": float(explainer.expected_value),
"predicted_risk": float(model.predict_proba(patient_features)[:, 1][0]),
"top_drivers": top_risk_drivers.to_dict(orient="records")
}