AI
One of the most expensive bottlenecks in agent development can be made absurdly cheap through recursion, according to a new paper.
AI agents rely on good training exercises to produce consistent results. Humans have needed to build these complex exercises by hand and validate the agent's work, and humans are expensive.
But what if you could just magic up increasingly difficult tasks synthetically, through cheap recursion, to train agents to string together a whole lot of terminal commands into a workflow that solves a real-world problem?
That, said Tencent HY LLM Frontier researchers and collaborators in a new paper, is not only possible, but absurdly cheap, with a demonstrated case of 1,000 accepted tasks for $50, or $0.05 each.
The researchers said that fine-tuning open-weights models using the agent trajectories, the steps the agent took working on a task, saw meaningful improvements in the base model: "Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench."
The effect is to move the bottleneck for building terminal agents from "can we afford the data" to "how far do we want to push the recursion", said PwC Austria head of AI engineering Pascal Biese.
You can think of the approach described in the paper, Recursive Synthesis for Long-Horizon Terminal Tasks, as a reversal, said joint lead author Yucheng Shi.
Join peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events.
Already a member? Sign in