Topic

AI Research

All digests tagged AI Research

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect thumbnail

· 19:39

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect

The presentation details an experiment where AI agents (Claude Code and Codex) competed in an 'Optimizer Speedrun' to achieve a new record for training a GPT-2-level model. While the agents successfully beat the human record, the speaker's key finding is that they achieved this by combining existing ideas rather than inventing novel optimizers or mechanisms. To advance AI research beyond mere evaluation, the speaker proposes an 'AlphaEvolve-style discovery loop' that integrates multi-agent interaction, quality feedback, and scaling elements.

Key takeaways

  1. AI Agents Beat Human Records in Speedrun 17:01

    In the Optimizer Speedrun, both Codex and Claude Code significantly outperformed the human record, achieving a new best record for training a GPT-2-level model. (10:21)

  2. Agents Exhibit Different Behaviors 13:56

    Codex was observed to write extensively in its 'scratchpad' (active memory), spawn more sub-agents, and burn more tokens compared to Claude Code, which frequently became idle and stated it could not improve the record. (7:46, 8:36)

  3. Lack of Novel Discovery

    Despite the impressive results, the models did not invent a new optimizer or mechanism. Instead, they combined existing ideas for small gains, suggesting that current methods are more geared toward evaluation than true discovery. (14:35)

  4. Proposed Discovery Loop

    The speaker proposes an AlphaEvolve-style multi-agent system that includes generators (LLMs), a reward mechanism (speedrun), a judge (quality feedback), and a scaling element to guide research toward novel breakthroughs. (15:40)

Watch on YouTube Full article

Accelerate the self-improving AI loop with CoreWeave ARIA thumbnail

· 8:44

Accelerate the self-improving AI loop with CoreWeave ARIA

CoreWeave ARIA is an AI research and iteration agent integrated into Weights & Biases (W&B) designed to accelerate the self-improving AI loop. It addresses common challenges in AI development, such as stalled iteration cycles, massive data volume analysis, and manual dashboard creation. ARIA automates auto-research, analyzes training metrics and agent traces, generates comprehensive reports with suggested next steps, and assists in optimizing LLM prompts and agent performance.

Key takeaways

  1. Automated Auto-Research Loop

    ARIA can conduct auto-research by analyzing recorded training metrics and agent traces to uncover hidden insights. It generates visualization-packed W&B reports and automatically launches follow-up training experiments based on its findings, minimizing manual effort (5:51).

  2. Agent Performance Optimization

    ARIA supports agent development by analyzing production traces and suggesting improvements. It can specifically help refine system prompts and evaluate multiple prompt alternatives using defined datasets to achieve higher quality results at lower latency (7:07).

  3. Comprehensive Workflow Support 2:30

    Beyond research, ARIA handles time-consuming manual tasks like providing advice, generating code, and executing commands, all while supporting concurrent conversations that can continue running in the cloud (2:21).

Watch on YouTube Full article

Hugging Face Journal Club: Training AI Scientists to Replicate Research thumbnail

· 34:41

Hugging Face Journal Club: Training AI Scientists to Replicate Research

The discussion summarizes research on Faraday-27B, a model trained by Inherent designed for scientific replication—the ability to reproduce results from redacted ML/AI papers. The system uses Reinforcement Learning (RL) and integrates CodeX as a tool, allowing the agent to execute code within a simulated environment. Key methodological advances include using sophisticated rubric-based judges (generated via Claude) instead of simple verifiers, employing multi-rollout averaging to mitigate variance, and implementing weighted credit assignment across the agent's steps.

Key takeaways

  1. Scientific Replication Task

    The model is tasked with replicating missing figures from redacted ML/AI papers. This process requires the agent to use tools (like CodeX) and execute code in a simulated environment, moving toward full automation of AI R&D.

  2. Advanced Judging Mechanism 0:01

    Instead of simple verification, the system uses a rubric-based judge (generated by Claude) that assigns fine-grained points for correct reasoning, figure accuracy, and code writing. This process involves averaging judgments across multiple rollouts to prevent reward hacking.

  3. Performance & Scaling 0:02

    The trained Faraday model demonstrated strong performance, sometimes outperforming much larger models like Claude and GPT-5. Furthermore, the system showed generalization even when given increased compute resources (e.g., scaling up to 8 hours/8 B300s).

Watch on YouTube Full article