Topic

GPT-2

All digests tagged GPT-2

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect thumbnail

· 19:39

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect

The presentation details an experiment where AI agents (Claude Code and Codex) competed in an 'Optimizer Speedrun' to achieve a new record for training a GPT-2-level model. While the agents successfully beat the human record, the speaker's key finding is that they achieved this by combining existing ideas rather than inventing novel optimizers or mechanisms. To advance AI research beyond mere evaluation, the speaker proposes an 'AlphaEvolve-style discovery loop' that integrates multi-agent interaction, quality feedback, and scaling elements.

Key takeaways

  1. AI Agents Beat Human Records in Speedrun 17:01

    In the Optimizer Speedrun, both Codex and Claude Code significantly outperformed the human record, achieving a new best record for training a GPT-2-level model. (10:21)

  2. Agents Exhibit Different Behaviors 13:56

    Codex was observed to write extensively in its 'scratchpad' (active memory), spawn more sub-agents, and burn more tokens compared to Claude Code, which frequently became idle and stated it could not improve the record. (7:46, 8:36)

  3. Lack of Novel Discovery

    Despite the impressive results, the models did not invent a new optimizer or mechanism. Instead, they combined existing ideas for small gains, suggesting that current methods are more geared toward evaluation than true discovery. (14:35)

  4. Proposed Discovery Loop

    The speaker proposes an AlphaEvolve-style multi-agent system that includes generators (LLMs), a reward mechanism (speedrun), a judge (quality feedback), and a scaling element to guide research toward novel breakthroughs. (15:40)

Watch on YouTube Full article