Case 02 · ML pipeline · Celebal Technologies internship

Lead Scoring System

Scoring 9,000+ sales leads by conversion probability — with a pipeline that can't quietly cheat.

Role
Built end-to-end during my Celebal internship
Stack
Python · Scikit-Learn · CatBoost · LightGBM · Optuna · SHAP · Streamlit
Links
Code ↗

The problem

Sales teams need to know which leads to call first. A model trained on every available column looks brilliant — but some of those columns only get filled in after a salesperson has already made contact. It's predicting the past.

What I built

  • Found and removed 8 post-contact features that were leaking the answer and inflating ROC-AUC by 0.12.
  • A reusable, leakage-free Scikit-Learn pipeline: every preprocessing step is fitted inside each cross-validation fold.
  • 7 classifiers benchmarked with stratified 5-fold CV; CatBoost / LightGBM tuned with Optuna and explained with SHAP.
  • A Streamlit app with batch scoring and an interactive, threshold-based prediction interface.
Leakage-free training & scoring
Raw leads9,000+ Leakage auditdrop 8 post-contact cols Stratified 5-fold CV — fit happens inside each fold Preprocessfit on train fold only Modelscore on held-out fold Benchmark7 classifiers CatBoost / LightGBMOptuna-tuned · 0.893 AUC SHAPwhy a lead scored high Streamlit appbatch scoring · threshold Leaky features inflated ROC-AUC by 0.12 — the honest number is the one reported.

Why the lower score is the better one

The leaky model scored higher. It would also have been useless on a new lead, because the columns it relied on don't exist yet at the moment you need a prediction. Proving the inflation (0.12 ROC-AUC) and rebuilding the pipeline so fitting can't see the held-out fold is what makes 0.893 a number you can trust.

0.893ROC-AUC, held-out
91.8%recall, held-out
9,000+leads scored
8leaky features removed
Next caseIntelligent Tourist Guide →