Foundations of Generative AI
Large language models (LLMs) and other generative models have achieved remarkable success, yet fundamental questions about their reliability, interpretability, and safety remain. I view statistics as a principled language for reasoning about data, dependence, and uncertainty, and algorithms as the bridge from these principles to practical learning systems.
Guided by this perspective, my work studies how generative AI systems can be made more reliable and better understood. It spans statistical watermarking, model evaluation, uncertainty quantification, and broader questions about their training and behavior.
LLM Watermarking
- A Statistical Framework of Watermarks for Large Language Models: Pivot, Detection Efficiency, and Optimal Rules
X. Li, F. Ruan, H. Wang, Q. Long, and W. J. Su. The Annals of Statistics, 2025. - Optimal Detection for Language Watermarks with Pseudorandom Collision
T. T. Cai, X. Li, Q. Long, W. J. Su, and G. G. Wen (Alphabetical). arXiv preprint arXiv:2510.22007, 2025. - Robust Detection of Watermarks in Large Language Models under Human Edits
X. Li, F. Ruan, H. Wang, Q. Long, and W. J. Su. Journal of the Royal Statistical Society: Series B, 2025. - On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
W. He*, X. Li*, T. Shang, L. Shen, W. J. Su, and Q. Long. NeurIPS, 2025 (Spotlight). - Debiasing Watermarks for Large Language Models via Maximal Coupling
Y. Xie, X. Li, T. Mallick, W. J. Su, and R. Zhang. Journal of the American Statistical Association, 2025. - Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
X. Li, G. G. Wen, W. He, J. Wu, Q. Long, and W. J. Su. arXiv preprint arXiv:2506.22343, 2025. - Optimal Watermark Localization in Mixed-Source Large Language Model Texts
J. H. Blanchet, T. T. Cai, X. Li, H. Liu, Q. Long, and W. J. Su (Alphabetical). arXiv preprint arXiv:2608.14906, 2026. - Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models
W. He*, X. Li*, L. Shen, W. J. Su, and Q. Long. ICLR, 2026. - Selective Disclosure Watermarking for Large Language Models
X. Chen, X. Li, Y. Xie, and Q. Long. ICML, 2026.
AI Content Detection
- Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts
L. Liang and X. Li. arXiv preprint arXiv:2605.12890, 2026. - Robust Spectral Watermark for Synthetic Tabular Data
Y. Zhao, X. Li, P. Song, Q. Long, and W. J. Su. Statistical Learning and Data Science, 2026.
LLM Evaluation
- Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
X. Li, J. Xin, Q. Long, and W. J. Su. arXiv preprint arXiv:2506.02058, 2025. - UCS: Estimating Unseen Coverage for Improved In-Context Learning
J. Xin, X. Li, E. Qiang, W. He, T. Shang, W. J. Su, and Q. Long. Findings of ACL, 2026.