About me

I am a PhD student in Computer Science at Rutgers University and a member of RAISL. My research focuses on machine learning and reinforcement learning, with the goal of developing intelligent agents that can reason, learn, and adapt in complex environments.

I’m broadly interested in building AI systems that are sample-efficient and generalizable, drawing on deep reinforcement learning, multi-agent systems, and the theoretical foundations of learning algorithms.

I’m open to academic collaborations and ML / RL internship opportunities. Get in touch

Research interests

  • Deep reinforcement learning
  • Multi-agent systems
  • Learning theory
  • Efficient LLM training

News

Started my PhD at Rutgers University.

Selected publications

Google Scholar
Figure 1 from PR²: routing replay compared with predictive routing replay using an evolution predictor to reduce expert-selection mismatch.Figure 1Enlarge

2026SCALE Workshop at ICML 2026 · Poster

PR²: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning

Daize Dong, Junlin Chen, Haolong Jia, et al.

arXiv author list

Daize Dong, Junlin Chen, Haolong Jia, Jiang Liu, Jiawei Wu, Huanwei Di, Jialian Wu, Zhengzhong Liu, Zicheng Liu, Emad Barsoum, Dimitris N. Metaxas, Hongyi Wang.

Predictive routing replay reduces rollout–training mismatch and improves reinforcement-learning stability for mixture-of-experts language models.

Figure 1 from K2-V2: six benchmark comparisons for K2 Instruct and K2 Base against other language models.Figure 1Enlarge

2025Technical report · Revised 2026

K2-V2: A 360-Open, Reasoning-Enhanced LLM

K2 Team (including Haolong Jia)

Full author list

K2 Team: Zhengzhong Liu, Liping Tang, Linghao Jin, Haonan Li, Nikhil Ranjan, Desai Fan, Shaurya Rohatgi, Richard Fan, Omkar Pangarkar, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Seungwook Han, Bowen Tan, Gurpreet Gosal, Xudong Han, Varad Pimpalkhute, Shibo Hao, Ming Shan Hee, Joel Hestness, Haolong Jia, Liqun Ma, Aaryamonvikram Singh, Daria Soboleva, Natalia Vassilieva, Renxi Wang, Yingquan Wu, Yuekai Sun, Taylor Killian, Alexander Moreno, John Maggs, Hector Ren, Guowei He, Hongyi Wang, Xuezhe Ma, Yuqi Wang, Mikhail Yurochkin, Eric P. Xing.

A fully open 70B language model built from scratch, with transparency across data, pre-training, annealing, and SFT. Strong capabilities in math reasoning, long-context understanding, and tool use.

Figure 1 from EfficientLLM: an overview of efficiency techniques, datasets, model families, and six efficiency assessment metrics.Figure 1Enlarge

2025arXiv preprint

EfficientLLM: Efficiency in Large Language Models

Zhengqing Yuan, Weixiang Sun, …, Haolong Jia, et al.

Full author list

Zhengqing Yuan, Weixiang Sun, Yixin Liu, Huichi Zhou, Rong Zhou, Yiyang Li, Zheyuan Zhang, Wei Song, Yue Huang, Haolong Jia, Keerthiram Murugesan, Yu Wang, Lifang He, Jianfeng Gao, Lichao Sun, Yanfang Ye.

A benchmark of efficiency–performance trade-offs across LLM pretraining, fine-tuning, and inference, using six complementary efficiency metrics.

Figure 1 from Mora: a six-step multi-agent workflow for prompt enhancement, image generation and editing, video generation, extraction, and connection.Figure 1Enlarge

2024arXiv preprint

Mora: Enabling Generalist Video Generation via A Multi-Agent Framework

Zhengqing Yuan, Yixin Liu, …, Haolong Jia, et al.

Full author list

Zhengqing Yuan, Yixin Liu, Yihan Cao, Weixiang Sun, Haolong Jia, Ruoxi Chen, Zhaoxu Li, Bin Lin, Li Yuan, Lifang He, Chi Wang, Yanfang Ye, Lichao Sun.

A collaborative multi-agent framework for generalist video generation, with coordinated fine-tuning, synthetic training data, and human-in-the-loop data filtering.

Education

2025–Present

2022–2024

2018–2022

Internship

Sep 2024–Sep 2025

Contact

For research discussions, collaborations, or internship opportunities:

Figure 1