Figure 1Enlarge
PR²: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Predictive routing replay reduces rollout–training mismatch and improves reinforcement-learning stability for mixture-of-experts language models.
I am a PhD student in Computer Science at Rutgers University and a member of RAISL. My research focuses on machine learning and reinforcement learning, with the goal of developing intelligent agents that can reason, learn, and adapt in complex environments.
I’m broadly interested in building AI systems that are sample-efficient and generalizable, drawing on deep reinforcement learning, multi-agent systems, and the theoretical foundations of learning algorithms.
I’m open to academic collaborations and ML / RL internship opportunities. Get in touch
Started my PhD at Rutgers University.
Figure 1Enlarge
Predictive routing replay reduces rollout–training mismatch and improves reinforcement-learning stability for mixture-of-experts language models.
Figure 1Enlarge
A fully open 70B language model built from scratch, with transparency across data, pre-training, annealing, and SFT. Strong capabilities in math reasoning, long-context understanding, and tool use.
Figure 1Enlarge
A benchmark of efficiency–performance trade-offs across LLM pretraining, fine-tuning, and inference, using six complementary efficiency metrics.
Figure 1Enlarge
A collaborative multi-agent framework for generalist video generation, with coordinated fine-tuning, synthetic training data, and human-in-the-loop data filtering.
2025–Present
PhD student, Computer Science
2022–2024
M.S. in Natural Language Processing
2018–2022
B.S. in Information Security
Sep 2024–Sep 2025
Machine Learning Researcher · Internship
For research discussions, collaborations, or internship opportunities:
Haolong.Jia@rutgers.edu