About Me
Hi, I’m Jaylen Jones, a fifth-year PhD student in Computer Science & Engineering at The Ohio State University (OSU) advised by Prof. Huan Sun and Prof. Eric Fosler-Lussier. Prior to joining OSU, I received my B.S. in Computer Science from the College of Informatics at Northern Kentucky University (NKU).
Research Interests
My research centers on understanding, evaluating, and mitigating risks in LLM-based agents, particularly for high-stakes scenarios where failures can have the most significant consequences for real-world users. This includes emphasis on both safety (e.g., preventing accidental agent harms emerging from typical benign inputs) and security (e.g., protecting agents from adversarial attack).
My current work focuses on the following high-level areas:
- Developing realistic and controlled evaluation frameworks for adversarial testing, enabling pre-deployment evaluations that capture safety and security risks under realistic real-world threat models.
- Automatically eliciting safety risks from computer-use agents, proactively and scalably surfacing long-tail unintended behaviors that emerge inadvertently within benign task and environment scenarios.
- Mitigating agent safety risks through continual learning, enabling agents to learn from prior safety-related experiences to avoid repeating harmful behaviors across future interactions.
Publications & Papers
RedTeamCUA: Towards Realistic Adversarial Testing of CUAs in Hybrid Web-OS Environments (Oral)
Paper · Website · Code
Zeyi Liao*, Jaylen Jones*, Linxi Jiang*, Eric Fosler-Lussier, Yu Su, Zhiqiang Lin, Huan Sun
(* denotes equal contribution)
The Fourteenth International Conference on Learning Representations
Adopted by QwenCUA for adversarial robustness evaluation.
(ICLR 2026)When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
Paper · Website · Code · Data
Jaylen Jones*, Zhehao Zhang*, Yuting Ning, Eric Fosler-Lussier, Pierre-Luc St-Charles, Yoshua Bengio, Dawn Song, Yu Su, Huan Sun
(* denotes equal contribution)
The Forty-Third International Conference on Machine Learning
(ICML 2026, ICLR AIWILD Workshop 2026 (Spotlight))When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
Paper · Website · Code · Data
Yuting Ning, Jaylen Jones, Zhehao Zhang, Chentao Ye, Weitong Ruan, Junyi Li, Rahul Gupta, Huan Sun The Forty-Third International Conference on Machine Learning
(ICML 2026)AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
Paper
Vishal Kumar, Zeyi Liao, Jaylen Jones, Huan Sun
(arXiv 2025)A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
Paper
Jaylen Jones, Lingbo Mo, Eric Fosler-Lussier, Huan Sun
2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics
(NAACL 2024)
Awards and Experience
- Oral Presentation for RedTeamCUA at the 2026 International Conference on Learning Representations, 2026
- Attended 2025 Conference on Language Modeling to present at COLM 2025 Workshop on AI Agents: Capabilities and Safety, 2025
- Lead Student Writer on accepted Open Philanthropy - Call for AI Safety Research grant, 2025
- Led as a Student Presenter at the Center for AI Policy’s Congressional Exhibition for Advanced AI, 2025
- Lead Student Writer on accepted Schmidt Sciences’ Safety Science Initiative grant, 2024
- Attended the 2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics, 2024
- Inaugural member of the L.I.F.E Foundation Fellowship at NKU, 2018
