作品集 · Selected work

Projects in bloom

A selection of projects spanning data analytics, machine learning and applied AI. Each one turns raw data into clear, actionable decisions.

購買

E-Commerce Funnel & Retention Analysis

Traced 300,000 user events through a multi-category store to explain why 88,000 monthly users and $1.68M in revenue converted under 5%. SQL analysis put the break at discovery rather than checkout. 93.2% of sessions never added to cart. Month-1 retention sat at 1.2% against an 8 to 15% benchmark and smartphones were the largest recoverable category. Delivered as a three-page Power BI report with cohort retention and RFM segmentation.

  • SQL
  • MySQL
  • Power BI
  • Python
  • Cohort & Retention Analysis
  • Funnel Analytics
不正

Healthcare Provider Fraud Risk Explorer

End-to-end ML workflow that turns real Medicare Part B billing data into a ranked and explainable worklist of providers most likely committing fraud. Built a 6M provider-year panel from CMS and OIG LEIE data with peer-relative features then trained Logistic Regression, Gradient Boosting and XGBoost to optimize precision-at-top-k for extreme class imbalance of about 0.02% fraud.

  • Python
  • XGBoost
  • scikit-learn
  • Feature Engineering
  • Imbalanced Classification
  • Healthcare Analytics
交通

NYC Congestion Pricing Causal Impact

Measured the real effect of NYC congestion pricing on 65M taxi trips with a difference-in-differences design. The naive before and after read says traffic inside the zone rose 13.2%. Controlling for citywide demand growth flips the sign. Trips fell 17.5% at p below 0.001 and speeds rose 6.5%. That is roughly 75,000 fewer trips a week. Validated with parallel-trends testing, a placebo toll date and robustness windows.

  • Python
  • statsmodels
  • Causal Inference
  • Difference-in-Differences
  • A/B Testing Methods
  • Policy Analytics
顧客

Customer Segmentation with RFM Analysis

Segmented 3,490 retail customers into 11 behavioural groups so a bike parts business could target the right 1,000 prospects. Cleaned four raw tables covering 20,000 transactions then scored every customer on recency, frequency and profit. Those quartile-ranked scores drive the segment map across $21.9M in sales and $10.8M in profit. Platinum Customers came out at just 4.8% of the base while buying 70% more often than average.

  • Python
  • pandas
  • RFM Segmentation
  • Tableau
  • Customer Analytics
  • Data Cleaning
公正

Algorithmic Fairness Audit (COMPAS)

Rebuilt the ProPublica and Northpointe dispute over the COMPAS recidivism score from raw data and showed both sides were right. Black defendants who never reoffended were flagged high risk 44.8% of the time against 23.5% for white defendants. The score still stayed calibrated for both groups. Proved the gap is forced by differing base rates and holds across 800 combinations of PPV and FNR. Equal precision and equal false positive rates cannot both be satisfied.

  • Python
  • pandas
  • Responsible AI
  • Model Fairness & Bias Auditing
  • Classification Metrics
  • Statistics