Case 02 · ML pipeline · Celebal Technologies internship
Lead Scoring System
Scoring 9,000+ sales leads by conversion probability — with a pipeline that can't quietly cheat.
The problem
Sales teams need to know which leads to call first. A model trained on every available column looks brilliant — but some of those columns only get filled in after a salesperson has already made contact. It's predicting the past.
What I built
- Found and removed 8 post-contact features that were leaking the answer and inflating ROC-AUC by 0.12.
- A reusable, leakage-free Scikit-Learn pipeline: every preprocessing step is fitted inside each cross-validation fold.
- 7 classifiers benchmarked with stratified 5-fold CV; CatBoost / LightGBM tuned with Optuna and explained with SHAP.
- A Streamlit app with batch scoring and an interactive, threshold-based prediction interface.
Why the lower score is the better one
The leaky model scored higher. It would also have been useless on a new lead, because the columns it relied on don't exist yet at the moment you need a prediction. Proving the inflation (0.12 ROC-AUC) and rebuilding the pipeline so fitting can't see the held-out fold is what makes 0.893 a number you can trust.
0.893ROC-AUC, held-out
91.8%recall, held-out
9,000+leads scored
8leaky features removed