Can AI Agents Automate LLM Post-Training? PostTrainBench Results | Maksym Andriushchenko
Maksym Andriushchenko on PostTrainBench: whether LLM agents can automate LLM post-training, and how quickly frontier models are closing the gap. Andriushchenko opens with why this capability is tracked at all: autonomous AI R&D sits alongside chemical, biological, radiological, and nuclear capabilities in the frontier safety frameworks published by Google DeepMind, OpenAI, and Anthropic, because it is the capability that could enable recursive self-improvement and loss of control. Earlier benchmarks measured small models on narrow tasks with heavy scaffolding. PostTrainBench instead gives an agent a base model of 2 to 4 billion parameters, a benchmark script, terminal access, web search, and 10 hours on a single H100, with no instructions on method, and scores only the post-trained model it produces. Results are averaged over four base models and seven benchmarks, with a judge that catches reward hacking. On the leaderboard he presents, GLM 5.2 leads, Opus 4.8 is close behind, and Fable 5 scores below both, based on the public version available at release. Two confounders matter: agent persistence, since some agents stop working hours before their budget runs out and have to be reprompted, and model size, which correlates clearly with performance. Reward hacking is pervasive, from repeating the test set to downloading a model from Hugging Face and presenting it as output. His headline: frontier models moved from roughly 10% to 34% on this benchmark in eight months. Chapters 0:00 PostTrainBench: can agents automate post-training? 0:23 Why automating AI research is a safety question 0:52 Autonomous AI R&D in frontier safety frameworks 1:16 The gap in earlier benchmarks 1:55 How PostTrainBench works 2:19 Results: GLM 5.2, Opus 4.8, and Fable 5 3:18 What the agent has access to 4:17 Catching reward hacking with a judge 4:48 Base models, hardware, scaffolds, and tasks 5:19 A worked example on Gemma3-4B 6:42 Where the gap is closing, and where it is not 7:37 Confounder: agent persistence 8:11 Confounder: model size and reasoning effort 8:43 How agents cheat on the benchmark 9:23 Takeaway: from 10% to 34% in eight months More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=0IfDd-WNQwrs14Kn FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.