Portrait of Xiang Li

Xiang Li

I am currently a postdoctoral researcher at the University of Pennsylvania, working with Weijie J. Su and Qi Long. I received my B.S. and Ph.D. in Statistics from Peking University, where I was advised by Zhihua Zhang. I will join the Department of Statistics at Rutgers University as an Assistant Professor in January 2027.

My research interests lie at the intersection of statistics, optimization, machine learning, and AI. My current research focuses on the theoretical foundations of generative AI, including text watermarking and LLM evaluation, as well as convergence and statistical inference for stochastic approximation. My earlier work includes federated learning with heterogeneous data and decision-making under uncertainty.

At a broader level, my work combines statistical thinking with algorithm design to understand the behavior, reliability, and uncertainty of learning systems.

Prospective students (Fall 2027). I expect to recruit PhD students at Rutgers for Fall 2027. If you are interested in statistical foundations of AI, machine learning, or optimization, please email me with a brief introduction and your research interests.

News

  • Oct. 2026 — Our textGrain watermark has been adopted by OpenAI! I’m excited to share this collaboration between researchers at Penn, Yale, and OpenAI. textGrain uses optimal transport to embed a detectable signal in language-model text while allowing direct control over the trade-off between watermark strength and sampling diversity. Congratulations and many thanks to all my collaborators! Read the OpenAI blog and the technical report for methodological details. The research paper will follow a week later.
  • Nov. 2026 — I will attend the SIAM Conference on Mathematics of Data Science (MDS26) in Salt Lake City, November 16–20, 2026. Looking forward to connecting!
  • Sep. 2026 — Two papers accepted to NeurIPS 2026! Steer-to-Detect uses hidden representations to detect AI-generated text; Conformal Prediction with Paraphrase-Aware Scoring accounts for paraphrases when quantifying LLM uncertainty. Congratulations to all my collaborators!
  • Aug. 2026 — I will attend the 2026 IMS New Researchers Conference and present Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? at JSM 2026. Looking forward to connecting!
  • Jul. 2026 — Selective Disclosure Watermarking for Large Language Models was accepted to ICML 2026. arXiv
  • Mar. 2026 — I will present Optimal Detection for Language Watermarks with Pseudorandom Collisions at ENAR, IWSM, and ICSA.
  • Nov. 2025 — I gave an SDLS webinar on LLM API usage and watermarking. Slides
  • Sep. 2025 — Two papers were accepted to NeurIPS 2025 as spotlights: one on watermark detection and one on privacy in decentralized federated learning.
  • Aug. 2025 — I presented recent work on estimating watermark proportions at JSM 2025.
  • Jun. 2025 — At ICSA 2025, I presented work on robust watermark detection and taught a short course on LLM watermarking. Slides
  • Apr. 2025 — I received the IMS New Researcher Travel Award.
  • Aug. 2024 — I chaired a federated learning session at MOPTA 2024.
  • May 2024 — I presented recent watermarking work at the IMS-NUS workshop.