LLM evaluation · post-training · agentic systems

Zheyuan Xiao

I study how language models are measured, how they learn, and how they act in real systems.

M.S. candidate in Computer Science at UT Austin and Software Engineer II at Resideo, based in Auckland, New Zealand.

New · Aug 2026 Student-Centric Answer Selection accepted to EMNLP 2026 Main Conference. Read paper

Research agenda

Measure. Learn. Act.

My work connects reliable evaluation, data-centric post-training, and agentic systems—three views of the same question: how do we build models whose behavior holds up beyond a benchmark?

01

Measure

When can we trust an LLM judge?

I study hidden factors in preference evaluation and design measurements that separate genuine quality from artifacts such as response length.

02

Learn

What supervision helps this model learn?

I investigate learning through cognitive frameworks and student-centric data selection, with the goal of matching training signals to the learner.

03

Act

How should model behavior meet the real world?

I build and evaluate agents and population-aligned simulations where evidence, representation, and end-to-end reliability matter.

Selected work

Research, with the signal kept visible.

114 citations across 4 papers Google Scholar · updated Sep 2026

  1. EMNLP 2026
    The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
    Zhengyu Hu*, Zheyuan Xiao*, Linxin Song, Fengqing Jiang, and 9 more authors
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing, 2026
    Main Conference, forthcoming. * Equal contribution.
    Selects verified teacher answers by student-specific learning cost instead of assuming the strongest teacher always teaches best.
  2. Findings 2025
    Explaining Length Bias in LLM-Based Preference Evaluations
    Zhengyu Hu, Linxin Song, Jieyu Zhang, Zheyuan Xiao, and 6 more authors
    In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025
    Explains length bias through desirability and information mass, then introduces a length-aligned evaluation protocol.
  3. arXiv 2025
    Population-Aligned Persona Generation for LLM-based Social Simulation
    Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Max Xiong, and 6 more authors
    arXiv preprint arXiv:2509.10127, 2025
    Builds representative persona sets by aligning LLM-generated profiles with real population-level psychometric distributions.
  4. NeurIPS 2025
    Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
    Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Seraphina Zhang, and 4 more authors
    In Advances in Neural Information Processing Systems 38, 2025
    Introduces a cognitive framework for testing how language models learn from instructors, concepts, and experience.
View full publication record

Recent

News & milestones

Aug 21, 2026 Our paper The Strongest Teacher Is Not Always the Best Teacher has been accepted to the EMNLP 2026 Main Conference.
Nov 04, 2025 Our work explaining length bias in LLM-based preference evaluations appears in Findings of EMNLP 2025.
Sep 18, 2025 Our paper Unveiling the Learning Mind of Language Models has been accepted to the NeurIPS 2025 Main Conference.
Sep 12, 2025 Our preprint on population-aligned persona generation for LLM-based social simulation is now available.

About

Research shaped by systems work.

I am an M.S. candidate in Computer Science at The University of Texas at Austin. I collaborate with researchers at HKUST(GZ) and Microsoft Research Asia on LLM evaluation, knowledge transfer, social simulation, and agentic systems.

Alongside research, I work as a Software Engineer II at Resideo, developing automated verification and validation systems for firmware and cloud-connected products. That systems perspective keeps my research grounded in behavior that is measurable, dependable, and useful in practice.

Before UT Austin, I received a First Class Honours B.E. in Software Engineering from The University of Auckland. See my CV for education and experience, or get in touch if you would like to discuss research or collaboration.