1. Jan 2026 – Aug 2026

    Data Scientist · GrowthGPT (LeapMind)

    Remote

    • Independently owned the data module for an agentic analytics capability, including data structure design, agent workflows, and its semantic layer.
    • Operationalized repeatable data analysis workflows through semantic layer routing and codified data retrieval; internal comparisons of similar tasks showed average processing times approximately 2–3 minutes shorter than for workflows that relied on runtime API documentation search and code generation.
    • Designed historical user interaction data structures for recurring review of live tasks and built an offline advertising database to improve data utilization and support ongoing capability development.
    • Partnered with the team on product roadmap and engineering implementation to launch the capability for users.
  2. Nov 2022 – Mar 2025

    Senior Data Analyst, User Growth · miHoYo

    Singapore

    • Adjusted SKAd trigger points and postback data across Genshin Impact iOS advertising, yielding a 13% relative improvement in ROAS.
    • Segmented traffic using device hardware and post-activation behavior, then suppressed postbacks for low-performing devices; one campaign in the experimental group improved ROI by 5% on an absolute basis after the strategy launched.
    • Classified creative assets using CPC, conversion, and purchase rate signals to guide creative direction; a new creative group delivered an 8% relative increase in install conversion rate versus historical creatives for the same campaign and period.
    • Built a traffic quality and fraud detection framework across online and offline attribution, combining click-to-install timing, click frequency, device metadata, IP clustering, and account matching to flag anomalies and strengthen measurement.
  3. Nov 2021 – Nov 2022

    Data Scientist, Tencent QQ · Tencent

    Shenzhen, China

    • Replaced a legacy rule-based recommendation system with an incrementally trained gradient boosted trees model that incorporated weekly user event rollups; A/B testing showed 3× CTR relative to the legacy system.
    • Combined profile attributes and behavioral time-series features in a Cox proportional hazards model, improving the C-index by 0.029.
    • Built a LightGBM model to segment users by next-day activity and translated lifecycle insights into product actions that reactivated 10% of silent users; separately, product-level A/B testing showed an approximately four-minute increase in session duration.
    • Led a department-wide data governance framework and automated data quality monitoring to improve data access, reuse, and the reliability of model inputs.
  4. Mar 2019 – Sep 2021

    Data Scientist, Claim Home Office · GEICO

    Washington, D.C.

    • Led the end-to-end development of a catastrophe risk system by integrating National Hurricane Center track data with internal policy data in GeoPandas, engineering spatiotemporal exposure features, and training logistic regression for ZIP-level loss probability and XGBoost for ZIP-level claim volume.
    • Operationalized the system to refresh hurricane data every six hours following new NHC alerts and distribute updated risk estimates through an internal API and automated department email updates for claims staffing and resource planning.
    • Designed a two-stage framework for third-party claims data in Hadoop without a shared join key: deterministic matching on available unique identifiers, followed by fuzzy/probabilistic matching for unresolved records; delivered the logic and linked dataset to engineering.
    • Developed a random forest driver analysis of NPS survey responses, using feature importance to identify key satisfaction drivers and translate findings into targeted quality assurance recommendations.
  5. Jan 2019 – Mar 2019

    Data Scientist, Intern · Waldron Inc

    New York City, NY

    • Deployed a contextual multi-armed bandit engine to dynamically optimize the exploration–exploitation trade-off for online shopping recommendations.
    • Engineered an end-to-end NLP pipeline to structure unstructured product descriptions, creating a unified feature matrix that improved downstream ranking precision.
  1. Aug 2026
    Seattle, WA

    Master of Science in Information Management (MSIM) · University of Washington

  2. Oct 2018
    New Jersey, USA

    Master of Engineering · Rutgers University

Languages & Data
Python, SQL, Spark, Hadoop, MySQL, Git
Machine Learning & Statistical Methods
scikit-learn, random forest, XGBoost, LightGBM, predictive modeling, recommender systems, survival analysis, feature engineering, model evaluation
Experimentation & Product Analytics
A/B testing, hypothesis testing, metric design, user segmentation, growth analytics, attribution and measurement
Data & AI Systems
ETL and data pipelines, data quality monitoring, API integration, semantic layers, agent workflows, Tableau, Excel