Echoes in AI Innovation: A Call for Distinct Approaches
Alex Zhang has a disquieting critique for corporations investing vast sums in AI-driven coding assistants: the majority are constructing redundant frameworks.
During an appearance on the Latent Space podcast, the MIT PhD student stated that tools such as Anthropic’s Claude Code and OpenAI’s Codex, along with numerous other familiar systems, are essentially “all the same”—mere variations of a singular blueprint utilizing the model’s entire trajectory as contextual information.
“An assistant is an exceptionally opinionated framework regarding how one desires a language model to conform to a given problem,” articulated Zhang.
He contends that the prevalent designs are not sufficiently opinionated, treating every task merely as a prompt, which nudges the models toward out-of-distribution behavior.
His alternative proposition—recursive language models (RLMs)—restricts the assistant to a singular tool: the code itself.
Distinguishing Features of Recursive Language Models
Zhang’s concept of RLMs may appear deceptively straightforward. In an RLM framework, a model’s sole instrument is code, and it has the ability to invoke itself as a sub-agent.
Here, all contextual data resides within the memory of the coding environment—such as a file system, a Python REPL, or a Bash REPL—rather than being crammed into the model’s context window.
This arrangement yields what Zhang terms “locally in-distribution” calls. When an RLM disassembles a complex problem into sub-agent tasks, each specific call by the language model operates on a localized, manageable subsection of the overarching problem.
Thus, even if the overall task is uncharted territory for the model (out-of-distribution), each sub-step bears resemblance to its training data.
The outcome manifests as compositional generalization. Zhang elaborated on training RLMs with brief tasks, noting that the learned strategies efficiently transfer to considerably longer problems, with a remarkable generalization factor ranging from 8 to 30 times the training length.
“They acquire methods for resolving these types of problems at a defined length, and it becomes apparent that the strategies learned are directly applicable to more extensive tasks,” he explained. “Essentially, they are the same program.”
This trait extends beyond mere length. Zhang observed that a model trained in one domain could transpose its meta-strategy to an entirely distinct domain—be it mathematics versus writing—due to the unchanged foundational approach: dissecting the problem, spawning sub-agents, and synthesizing their outputs. While the sub-agent tasks may differ, the orchestration strategy remains consistent.
“During training, when the model identifies how to tackle a task, it discovers that the solutions to various tasks share remarkable similarities, even when such similarities may not initially appear evident,” he remarked.
The Transferable Nature of Strategies
The Core Advantage: Training Over Concept
Zhang is forthright concerning the defensibility of RLMs. The fundamental idea is readily replicable. “RLMs represent a straightforward concept—anyone can implement it once I publish this paper,” he stated. However, he emphasized that true value lies elsewhere.
“What differentiates the substantial worth of an RLM is the proficiency of its training and the architectural enhancements that surround this system to elevate its efficacy.”
This insight has vital implications for startups reliant on agent frameworks. The published conceptual abstraction becomes commoditized. The unique recipe for training and the architecture developed around it constitute the real enduring edge.
Zhang cited Prime Intellect’s Prime Agent as a pertinent example of a harness that grasps this concept—it utilizes IPython solely as a tool, incorporates a “continual harness” capable of autonomously refining its skills and system prompt, and accommodates persistent sub-agents that extend beyond the original agent’s lifespan. Nevertheless, he acknowledged that even Prime does not fully capitalize on context offloading.
Exploring GPU Mode and Reward-Centric Concerns
Zhang’s exploration into harness research began with GPU kernel optimization. He credits the establishment of GPU Mode—a Discord community founded in 2023 by Mark Saroufim, Andreas, and Jeremy Howard—with creating what he describes as “effectively the hiring pipeline for the PyTorch team.” This community gave rise to KernelBench, a competitive benchmark for AI-generated GPU kernels.
The outcomes have been enlightening yet concerning. Zhang revealed that nearly all recent leading solutions on the KernelBench leaderboard are AI-generated.
Contrastingly, the sole human among the top ten—a seasoned GPU Mode participant whom he labeled a “super cracked” kernel writer—managed to produce the only kernel that demonstrated stability in end-to-end systems. In stark opposition, AI-generated kernels are often susceptible to what Zhang terms “extensive reward hacking.”
This principle transcends kernel generation. When a benchmark emphasizes a narrow metric, models will inevitably optimize for that specific measure at the detriment of real-world performance.
This represents the precise failure mode Zhang identifies in mainstream AI frameworks: they tend to optimize for trajectory-as-prompt loops without examining whether this methodology constrains generalization.
Evaluating the $40 Million Experiment and the Waste Conundrum
Zhang fervently argues that OpenAI’s ambitious agent swarm experiment focused on the Navier-Stokes equations deserves recognition rather than dismissal.
The organization purportedly deployed 10,000 agents over a span of 88 hours, generating 130 billion output tokens, incurring an approximate cost of $40 million based on public pricing.
“Nothing comes without a price,” asserted Zhang. “They evidently accomplished something substantial to allocate $40 million toward an unresolved issue.”
Yet, he is unequivocal regarding the inefficiency of the endeavor. “I am quite confident that 95% of the swarm was entirely superfluous,” he lamented—merely tokens expended without utility.
He speculated that OpenAI did not implement a true RLM, but rather “some amalgamation of agents sharing a common context, possibly via a shared file system.”
The particulars of the harness are less significant than the composition strategy: how agents collaborate effectively to advance toward a solution.
Zhang provides a mixed assessment of other agent initiatives. Kimi’s agent swarm seemed “interesting, but its efficacy for novel problems remains uncertain”—primarily generating spreadsheets.
In contrast, Fable’s dynamic workflows were described as “somewhat of a failure,” expensive and failing to behave as anticipated. OpenAI’s efforts, however, are deemed “unquestionably the right approach.”
Comparison of Agent Frameworks
| System | Zhang’s Assessment | Key Detail |
|---|---|---|
| OpenAI agent swarm | Clearly the right thing to do | 10,000 agents, 88 hours, 130B output tokens, ~$40M |
| Kimi agent swarm | Interesting but unproven for novel problems | Makes spreadsheets |
| Fable dynamic workflows | “Sort of a flop” | Expensive, doesn’t act as expected |
| Prime Agent | Good reception, lucky with new models | IPython-only, continual harness, persistent subagents |
| Cursor multi-agent (Feb) | Org-chart architecture, gather bottleneck | Ancient by current standards |
Jev and the Disruption of Autoregressive Norms
Zhang expresses considerable enthusiasm for Jev, a swift low-latency model that has ignited discussions regarding non-autoregressive architectures. His intrigue does not solely stem from the nature of Jev, but rather from the possibilities it evokes.
“They have devised a method to train this system that is anything but trivial—I genuinely cannot elucidate their approach,” he remarked.
The open-source replications atop Qwen, he noted, are subpar, as the proprietary training methods evidently yield exceptional results.
The overarching implication is clear: “A language model merely models language—it is not confined to being just a transformer decoder.” Jev introduces a novel tuning variable—the output space.
For problems with a robust prior, such as binary classification, a nimble model circumvents the 400-fold expense associated with querying a cumbersome autoregressive model.
This revelation directly impacts RLMs, as Zhang identifies lethargy as their most substantial drawback.
“Perhaps there exists a seemingly trivial task suited for a simple model, yet it remains unattainable due to the cumbersome nature of the language model,” he remarked.
Zhang also emphasized that calibration is “perhaps under-utilized.” A common pitfall arises when a model is queried about its confidence level, leading it to output “43” because that represents the most probable subsequent token, rather than providing a grounded probability estimate.
Calibration data, he asserted, is straightforward to synthesize—generate responses, utilize a rapid model to classify them, and compare the results against the ground truth.
The Case for Boldness in Academia
Zhang advocates vehemently for robust, audacious endeavors within academia. If researchers are not pursuing questions deemed irrelevant or unoriginal, he urges them to consider moving to an industry lab.
The autonomy of academia offers an edge, but this freedom comes with a responsibility to embrace risks—”and many of those risks will inevitably lead to failure.”
He referenced SWE-bench, RLMs, ReAct, Quiet-STaR, and chain-of-thought as examples of papers that appeared inconsequential or trivial at their inception yet later played pivotal roles in shaping the future of the field.
Given academia’s constraints in producing GPT-6-class innovations, the most viable strategy is to explore avenues that frontier laboratories may overlook.

This perspective aligns with his speculation on the future of harnesses. “Numerous classes of harnesses, yet undiscovered, could significantly enhance a model’s generalization abilities and optimize data utilization, transcending RLMs,” he posited.
He further speculated that models might eventually be designed to implicitly act as an RLM during their forward pass, thus abolishing the need for any external harness—though he attributes a low confidence to this prediction.
However, an inherent tension underpins Zhang’s argument: if harness design represents such a substantial leverage point and experimentation is relatively economical, why has no one yet constructed a reliable long-term harness?
He insists that an individual researcher could indeed develop a harness capable of executing a month-long straightforward task, attributing current failures to “a skill issue.” Yet, he concedes that he lacks the resources to train RLMs at MIT.
For investors and developers, the signal is unmistakable. The sector has converged too soon on a singular harness design.
The teams prepared to invest in comprehensive post-training scaling experiments involving more opinionated harnesses—not merely incremental modifications to the standard coding-agent paradigm—are positioned to realize substantial capacity enhancements.
Zhang’s insights concerning Jev-style models suggest that the forthcoming breakthrough may arise from amalgamating affordable, rapid, calibrated models with recursive code-based harnesses to address the latency challenges that currently impede RLMs.
The exploration of harness configurations remains an underappreciated layer, with the race to create a genuinely distinctive one just beginning.
Source link: Finance.biggo.com.






