Weiyuan Ding

Ph.D. Student, Computer Science
North Carolina State University
Advisor: Prof. Bowen Xu

I am a Ph.D. student at NC State, working with Prof. Bowen Xu. My current research interests center broadly on causal inference, building on my earlier work on the robustness, evaluation, and safety of large language models and agentic systems. This past summer I was a Research Intern at IDeaS (a SAS company), where I built a multi-agent knowledge engine for hotel revenue management. I received my M.S. in Computer Science from NC State in May 2026, and my B.S. in Information and Computing Science from Peking University.

Portrait of Weiyuan Ding

News

Publications

See Google Scholar for the full list.

AEGIS: From Clues to Verdicts — Graph-Guided Deep Vulnerability Reasoning via Dialectics and Meta-Auditing

Sen Fang, Weiyuan Ding, Zhezhen Cao, Zhou Yang, Bowen Xu

Preprint, 2026. arXiv

A multi-agent framework that grounds LLM-based vulnerability detection in a closed factual substrate. AEGIS reconstructs per-variable dependency chains over a repository-level Code Property Graph, then uses a dialectical Verifier and an independent Audit Agent (with veto power) to suppress hallucinated verdicts. Sets a new state of the art on PrimeVul (122 pair-wise correct predictions — first to surpass 100), reducing false positives by up to 54.40% at $0.09 per sample with no task-specific training.

EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming

Sen Fang, Weiyuan Ding, Mengshi Zhang, Zihao Chen, Bowen Xu

Preprint, 2026. arXiv · Leaderboard · Code

A loop-based robustness evaluation that alternates code generation and summarization on the same task. Applied to 100+ LLMs (incl. Claude Opus 4, DeepSeek-V3), we observe a 2.65%–47.62% pass@1 drop over 10 loops, and find that robustness is not strictly tied to initial accuracy.

Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation

Sen Fang, Weiyuan Ding, Antonio Mastropaolo, Bowen Xu

Preprint, 2025. arXiv

A study of how quantization affects robustness of code LLMs across LLaMA, DeepSeek, CodeGen, and StarCoder (350M–33B). Counter-intuitively, quantized models show higher robustness to adversarial and noisy inputs (51.59% vs. 42.86%) than their full-precision counterparts.

Experience

Education