Howdy! I’m a Statistics Ph.D. candidate at Boston University, advised by Debarghya Mukherjee and Luis Carvalho. Before BU: M.A. in Statistics at Columbia, B.S. in Mathematics at Shandong University, and a year at AMSS, Chinese Academy of Sciences. Earlier I worked with Zhanxing Zhu and Yongshun Gong, whose research on spatio-temporal structure still shapes how I think about heterogeneous, evolving data.

I work on transfer learning and representation learning: optimal transport, graph methods, and multimodal models, with a focus on what holds up when data is scarce, high-dimensional, and non-IID. The question I keep coming back to is how you reuse what a model already knows when the world won’t sit still. Part of the answer is knowing when transfer provably works, through minimax rates, oracle inequalities, and safe-transfer criteria. The other part is the cases where structure itself is the obstacle: aligning graphs and manifolds without known correspondence, warm-starting policies in environments that keep moving, and specializing pretrained LLMs and VLMs without letting them overfit or drift out of alignment.

A gentler entry point, in slides: transfer learning · graph learning · optimal transport · LLMs for time series

🔥 News

  • 2026.05TESS (co-first) got an Oral at ICML 2026, top 0.5% of submissions!
  • 2026.05 — “Network Perturbation Aggregation for Graphon Estimation” (co-author) is in at SLADS.
  • 2026.04 — Heading to Amazon as an Applied Scientist this summer, Bay Area bound.
  • 2026.04 — Honored to receive the Dean’s Dissertation Fellowship from BU’s Graduate School of Arts and Sciences.
  • 2025.09GTrans (first author) accepted at NeurIPS 2025.
  • 2025.08 — “Cross-Domain Hyperspectral Image Classification” (co-author) accepted at IEEE TGRS.

📝 Publications

First author

GTrans
Transfer Learning on Edge Connecting Probability Estimation Under Graphon Model
Graphon-level transfer without node correspondence: aligns graphs via Gromov–Wasserstein, then transfers edge structure nonparametrically, with residual smoothing that unlocks small and sparse targets.
NeurIPS 2025 / Paper / Poster / Slides / Code
TESS
From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space
Translates free-form text into interpretable temporal primitives (distribution shift, volatility, shape, lag) and conditions a Transformer forecaster on them, reducing error by up to 29% under event-driven non-stationarity.
ICML 2026 (Oral) / Paper / Slides / Code
Phase Transition
Phase Transition in Nonparametric Minimax Rates for Covariate Shifts on Approximate Manifolds
Minimax theory for near-manifold covariate shift, exposing a sharp phase transition governed by the support gap, together with a ratio-free estimator that adapts to unknown intrinsic dimension.
Under review / arXiv / Poster / Slides / Code
SCOT
SCOT: Multi-Source Cross-City Transfer with Optimal-Transport Soft-Correspondence Objectives
Sinkhorn entropic-OT coupling gives many-to-many region alignment across cities with no node matching, paired with an OT-weighted contrastive objective that resists collapse under multi-source heterogeneity.
Under review / arXiv / Slides

Co-author

Net-Paging
Network Perturbation Aggregation for Graphon Estimation
Generates graphon-preserving networks from a single observed graph and aggregates them to cut estimation variance, with a closed-form bias correction and guarantees that the base estimator's convergence is preserved.
SLADS 2026 / Paper / Code
MKDNet
Cross-Domain Hyperspectral Image Classification via Mamba-CNN and Knowledge Distillation
Pairs a Mamba global spectral encoder with CNN local features, then transfers across domains through teacher–student distillation and OT-guided graph consistency under severe spectral mismatch.
IEEE TGRS 2025 / Paper / Slides
SSGP
What Papers Don't Tell You: Recovering Tacit Knowledge for Automated Paper Reproduction
Prunes dense scientific graphs into task-adaptive subgraphs by rank-based ensemble scoring, shrinking the search space for LLM agents and raising the success rate of automated paper reproduction.
Under review / arXiv

🤖 LLM & DS Projects

Alignment
AlignDPODPO, IPO, KTO under QLoRA
RAGAuditNLI, SelfCheckGPT, semantic entropy
Retrieval
GraphRAGdense + entity graph + CLIP
Adaptive RAGdifficulty-aware query routing
DraftVerifyspeculative decoding, draft-verifier
HQQ1–8 bit quantization, 12.7x
Causal inference
CausalLensDoWhy, double ML, causal forests
Congestion PricingCS-DiD on 12M+ NYC TLC trips
A/B Testingbootstrap CIs, power analysis
Language
Financial SentimentDistilBERT on financial news
Spam DetectionTF-IDF, Naive Bayes
Vision
Dog Breed ClassificationVGG16, ResNet50 transfer
Mask DetectionResNet50, Grad-CAM
Statistics
Bayesian Logisticspike-and-slab MCMC, RStan
Time SeriesSARIMA, ETS, Prophet
Credit RiskXGBoost, SMOTE
SegmentationK-Means, elbow and silhouette
Recommendation SystemALS, SVD matrix factorization

📖 Educations

Boston University  ·  Ph.D. in Statistics  ·  2021.09 – present
Columbia University  ·  M.A. in Statistics, Data Science Track  ·  2019.09 – 2020.05
Shandong University  ·  B.S. in Mathematics  ·  2015.09 – 2019.06
Chinese Academy of Sciences  ·  Jointly Supervised Talent Program, AMSS  ·  2018.05 – 2019.06

💻 Internships

Applied Scientist Intern · Amazon · Summer 2026
LangChain agent with LLM-based heuristic learning that turns request-level attribution into ranked traffic-blocking policies; MIMO forecasting on large-scale HTTP logs, benchmarking tabular foundation models against Chronos-2.

Data Scientist Intern · Plymouth Rock Insurance · Summer 2025
Multimodal property risk scoring with GPT-4o and Street View imagery; XGBoost Tweedie loss model on SageMaker (+4.3% Gini). Slides

🎖 Honors

Boston University  ·  Dean’s Dissertation Fellowship (2026) · Ralph B. D’Agostino Fellowship (2025) · Outstanding Teaching Award (2025)
Shandong University  ·  Outstanding Graduate (2019) · First-Class Scholarship (2018) · Outstanding Student Leader (2018)
National  ·  Hua Loo-Keng Scholarship (2018) · National Gold Award, Internet+ Innovation and Entrepreneurship Competition (2018)

📝 Service & Teaching

Presentations  ·  CIKM 2024, NeurIPS 2025, ICML 2026
Reviewer  ·  CIKM 2025, ICME 2026, ICML 2026, KDD 2026, KDD 2027
Instructor, Boston University  ·  Mathematical Statistics (MA 582), Elementary Statistics (MA 113)
Teaching Fellow  ·  Generalized Linear Models (MA 575), Data Science in R (MA 415), Applied Statistics (MA 214)

✨ My Apps

A quiet collection of cinematic, atmospheric, and emotionally resonant side projects — part digital keepsakes, part memory-keepers.  See all →

Wilderness
🌲 Wilderness
MBTI Vibe
MBTI Vibe
What If Cinema
🎬 What If Cinema
Letters from the Screen
✉️ Letters from Screen
If You Disappeared
✈️ If You Disappeared
Souvenirs
🎟️ Souvenirs
The Map of Me
🗺️ Map of Me
A Room in Macondo
🦋 A Room in Macondo
Say It Like a Classic
✒️ Say It Like a Classic
The Boston Archive
🏛️ Boston Archive

🎨 Interests

🎵 Mandarin R&B loyalist — Leehom Wang, David Tao, Khalil Fong 🦋, Dean Ting
🎹 Trained in piano, calligraphy, and ink painting
🏞️ National park lover · 🫧 lake admirer · 🌅 opacarophile — welcome to my Gallery