Selected work
every number below traces to committed code or the résuméFive public repositories and two private ones. Where a project runs on simulated data, the card says so — simulation is how you grade an estimator against a known ground truth, and hiding it would defeat the point.
Showing all 7 projects
-
2026
Credit Intelligence Platform
real-time scoring · drift-gated retraining · regulated explanations
A loan-scoring service that watches its own accuracy decay and retrains itself.
Nine containerized services (FastAPI, Kafka, Redis, PostgreSQL, MLflow, Prometheus, Grafana) serving a stacking ensemble — XGBoost, LightGBM and logistic regression under a meta-learner trained on 5-fold out-of-fold predictions, evaluated on a strictly temporal holdout. FRED macro features join under a one-month publication lag, so a February loan only sees January's unemployment release; that single detail designs out the look-ahead leakage most credit models quietly ship with. Hand-built drift monitoring — PSI, KS and Wasserstein per feature, plus rolling-AUC concept-drift checks — feeds a scheduler that retrains and promotes through MLflow only when the challenger beats production. Denials generate SHAP-grounded, ECOA/FCRA-style adverse action notices through a swappable GPT-4o-mini or self-hosted Mistral-7B backend.
9docker services2.2Mloans trained on0.79auc, temporal holdout3drift statistics per featuresource → 3,863 lines · 16 tests -
2025–26
Research Pipeline Harness
lehigh RA work · public source
A backtesting framework built to answer one hard question honestly: is this skill, or is it luck?
Every component — feature family, model, baseline, performance metric — is a subclass behind one of five abstract base classes, resolved at runtime by dotted path from JSON config with subclass verification. A new strategy is one file plus a config edit; the orchestrator has never changed. The walk-forward backtest spans 30+ years of daily data, tuning only on trailing validation windows, with features constrained to be causal through expanding statistics shifted one row. Monte Carlo simulation of 100,000 bank-balance trajectories builds an empirical null, so each held-out period is scored against the strategy's own luck.
36classes, 5 abstractions100ktrajectories4naive baselines to beatsource → résumé claims verifiable here -
2026
Product Analytics Pipeline
cohorts · segmentation · ship-kill-iterate
Analytics that ends in a decision, not a dashboard.
Built on the real GA4 BigQuery schema and query path against Google's public Merchandise Store export, extending 92 days to 156 weeks through documented resampling that preserves event mix, device share and conversion rates — with a source flag on every row so charts always separate real from generated. On top: weekly cohort retention, behavioral segmentation with k chosen by silhouette score, and an experimentation pipeline running two-proportion z-tests, Wilson intervals, Cohen's h and observed power, ending in an explicit ship, iterate or kill call per experiment.
27tests, all passing156weeks of cohorts5notebooks, executedsource → real GA4 path · synthetic fallback without GCP -
2025
Variance Reduction for A/B Tests
CUPED · post-stratification · novelty decay
The statistics large platforms use to detect smaller wins without spending more traffic.
Implements and benchmarks four estimators head to head — Welch's t-test with hand-derived Satterthwaite degrees of freedom, CUPED (Deng et al., 2013), post-stratification (Miratrix et al., 2013) and regression adjustment — compared on standard error, variance reduction and minimum detectable effect. The telemetry is simulated on purpose: a seeded generator plants a known 2% lift with exponentially decaying novelty, so the novelty detector and every estimator are graded against ground truth rather than eyeballed. Unit tests verify CUPED's variance reduction empirically against its theoretical bound.
4estimators benchmarked1−ρ²bound verified by test11statistical unit testssource → simulated telemetry, by design -
2026
Matching Medieval Court Records
bipartite assignment on real archival data
Teaching a computer to recognize the same 15th-century person across transcription errors and variant spellings.
Reconciles AI-transcribed records from the KB27/799 King's Bench plea rolls against a hand-curated ground truth of 909 cases, then reconstructs the litigation network of the individuals it identifies. Matching is framed as two-slot bipartite assignment and solved globally with the Hungarian algorithm over a blended RapidFuzz cost, rather than greedy best-match-first, which misallocates records whenever several transcriptions plausibly map to one person. Fuzzy matching nearly doubles the strict baseline F1, and most of that gain came from grouping fragmented tokens into person entities before scoring — not from the matcher itself.
909ground-truth cases0.491f1 vs 0.262 baseline5testssource → real archival data -
2025
Controllable Music Generation
latent navigation in diffusion transformers
Text-to-music that follows directions: a language model plans the track, a diffusion model performs it.
An end-to-end generative audio pipeline pairing a 2B-parameter diffusion transformer with a 0.6B Qwen planner to synthesize 48kHz stereo audio, reaching a top-5 CLAP similarity of 0.528 against a baseline near 0.125. Inference is distributed across heterogeneous hardware — CUDA and Apple MPS — behind FastAPI microservices, cutting generation from roughly three minutes to one, with a multi-page Streamlit dashboard exposing fine-grained controls.
0.528clap vs ~0.125 baseline2B + 0.6Bparameter model pair3×faster generationprivate repository · figures as reported on résumé -
2026
Agentic RAG with Evaluation
query decomposition · tool use · graded answers
A question-answering agent that plans, uses tools, and gets graded on every answer.
A LangGraph-orchestrated RAG system that decomposes complex queries into sub-tasks, routes them through tool use and a ChromaDB vector store with OpenAI embeddings, then synthesizes answers from the pieces. The differentiator is the measurement: an evaluation harness scoring retrieval precision, faithfulness, hallucination rate and latency across 200+ test queries, deployed as a FastAPI service.
200+evaluation queries4quality dimensions15integration testsprivate repository · figures as reported on résumé
How I build
Four habits that show up in every repository above, each one visible in the source.
No look-ahead, ever
Time-ordered test splits. Macro data joined under a one-month publication lag, so a February loan only sees January's release. Backtest features computed from expanding statistics shifted one row, so week t never sees week t. Leakage is the quietest way an ML system lies, and it gets designed out at the data layer rather than caught in review.
fred_macro_pipeline.py · feature_calculator.py
Beat the null before believing the model
Every strategy is scored against a 100,000-trajectory Monte Carlo null and four naive baselines. Every A/B estimator is validated on telemetry with a planted, known effect. A result that cannot outperform its own luck is not a result.
monte_carlo.py · data/generate.py
Config in, components out
New models, features and metrics plug in as subclasses resolved by dotted import path from JSON, type-checked against their abstract base before any data loads. The orchestrator has not changed since the second component was added.
dynamic_loading.py · config_validator.py
Simulated data says so, on every row
The analytics pipeline stamps a source flag on each extended record so downstream charts always know real from resampled. The experimentation suite's generator exists precisely so estimators can be graded against ground truth. Simulation is a validation tool, not something to pass off as production scale.
data_extension.py
Experience
as_of 2026-09Bethlehem, PA
Research Assistant, Machine Learning Engineering
Lehigh University
- Architected a config-driven ML pipeline with 16 modular components across 2,100+ lines, using abstract base classes and runtime resolution via importlib so new extractors, models and metrics plug in through JSON config without touching orchestration code.
- Built a time-series-aware training harness retraining classifiers on rolling windows across 30+ years of daily data, searching 1,500+ hyperparameter combinations per fold, with causal feature construction and strict temporal splits to eliminate leakage.
- Developed an evaluation framework computing 30+ metrics with 100,000-iteration Monte Carlo simulation, four baseline models, and KS, chi-square and Wald-Wolfowitz significance testing.
The code behind this role is public. Line count, baseline classes, Monte Carlo iterations and all three statistical tests check out against the source.
India
Data Science Intern
CASHe — fintech / digital lending
- Built a hybrid company-search system pairing TF-IDF n-gram fuzzy matching for known entities with GPT-based semantic resolution for unseen names, improving search accuracy by 35%.
- Enriched 500K+ customer location records via the Google Maps API and applied geo-clustering to surface spatial demand patterns across 12 metro regions.
- Engineered cluster membership as a predictive feature in the loan propensity model, strengthening regional risk differentiation and improving downstream AUC.
Bethlehem, PA
M.S. Data Science
Lehigh University · GPA 3.94 / 4.0
India
B.E. Mechanical Engineering, Minor in Data Science
Birla Institute of Technology & Science, Pilani · GPA 3.5 / 4.0
Toolkit
things I would defend in an interview- languages
- Python · SQL · R · C++
- ml_modeling
- PyTorch · scikit-learn · XGBoost · LightGBM · SHAP · pandas · NumPy · SciPy · HuggingFace
- llms_genai
- GPT-4o-mini · Mistral-7B · Diffusion Transformers · RAG · LangChain · LangGraph · ChromaDB · FAISS · LoRA · LLM evaluation
- infrastructure
- FastAPI · Docker · Kafka · Redis · PostgreSQL · MLflow · Prometheus · Grafana · Streamlit · AWS (S3, EC2) · Git · CI/CD
- statistics
- Causal measurement · CUPED · bootstrap intervals · Monte Carlo · hypothesis testing · drift detection · Bayesian inference