About Me

Hi, I’m Jaylen Jones, a fifth-year PhD student in Computer Science & Engineering at The Ohio State University (OSU) advised by Prof. Huan Sun and Prof. Eric Fosler-Lussier. Prior to joining OSU, I received my B.S. in Computer Science from the College of Informatics at Northern Kentucky University (NKU).

Research Interests

My research centers on understanding, evaluating, and mitigating risks in LLM-based agents, particularly for high-stakes scenarios where failures can have the most significant consequences for real-world users. This includes emphasis on both safety (e.g., preventing accidental agent harms emerging from typical benign inputs) and security (e.g., protecting agents from adversarial attack).

My current work focuses on the following high-level areas:

  • Developing realistic and controlled evaluation frameworks for adversarial testing, enabling pre-deployment evaluations that capture safety and security risks under realistic real-world threat models.
  • Automatically eliciting safety risks from computer-use agents, proactively and scalably surfacing long-tail unintended behaviors that emerge inadvertently within benign task and environment scenarios.
  • Mitigating agent safety risks through continual learning, enabling agents to learn from prior safety-related experiences to avoid repeating harmful behaviors across future interactions.

Publications & Papers

  • RedTeamCUA: Towards Realistic Adversarial Testing of CUAs in Hybrid Web-OS Environments (Oral)
    Paper · Website · Code
    Zeyi Liao*, Jaylen Jones*, Linxi Jiang*, Eric Fosler-Lussier, Yu Su, Zhiqiang Lin, Huan Sun
    (* denotes equal contribution)
    The Fourteenth International Conference on Learning Representations
    Adopted by QwenCUA for adversarial robustness evaluation.
    (ICLR 2026)

  • When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
    Paper · Website · Code · Data
    Jaylen Jones*, Zhehao Zhang*, Yuting Ning, Eric Fosler-Lussier, Pierre-Luc St-Charles, Yoshua Bengio, Dawn Song, Yu Su, Huan Sun
    (* denotes equal contribution)
    The Forty-Third International Conference on Machine Learning
    (ICML 2026, ICLR AIWILD Workshop 2026 (Spotlight))

  • When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
    Paper · Website · Code · Data
    Yuting Ning, Jaylen Jones, Zhehao Zhang, Chentao Ye, Weitong Ruan, Junyi Li, Rahul Gupta, Huan Sun The Forty-Third International Conference on Machine Learning
    (ICML 2026)

  • AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
    Paper
    Vishal Kumar, Zeyi Liao, Jaylen Jones, Huan Sun
    (arXiv 2025)

  • A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
    Paper
    Jaylen Jones, Lingbo Mo, Eric Fosler-Lussier, Huan Sun
    2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics
    (NAACL 2024)

Awards and Experience