Latent Space

Inside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj

Published 2026-09-30 · Duration 40:24

Summary

OpenAI announced significant advancements in AI agent capabilities, particularly in 'Computer Use,' aiming for superhuman performance in software operation. Key announcements include **Dots**, a personal assistant product giving agents access to private Linux virtual machines in the cloud, and **GPT-6.1**, a new model offering substantial cost and speed improvements. For developers, the **Agents API** provides a unified platform for building on Computer Use, while the new **Decisions API** offers ultra-fast, low-latency classification and structured output, inspired by the recent 'Jev' model. The overall trend is moving toward an 'AI cloud' with higher-level primitives for memory, compute, and state management.

Download summary

Key takeaways

  1. Computer Use is Superhumaning 16:40

    Computer Use agents are dramatically improved, moving from human-level to potentially 'literally superhuman' performance in real-world software tasks. This is achieved by combining multiple modalities: screenshots, accessibility data, the DOM, Playwright, and generated JavaScript code, allowing agents to debug and recover from failures more effectively. (Transcript)

  2. Dots and Personal Computing 5:50

    The 'Dots' product gives agents access to their own dedicated Linux virtual computer in the cloud, enabling them to run full desktop applications and interact with complex, non-API-exposed software, expanding agent utility beyond simple web browsing. (Transcript)

  3. New Developer APIs for Scale 10:00

    New API features include **async tool calling**, **mid-turn steering**, and **WebSockets**, which enable bidirectional, real-time communication with the model. The **Agents API** centralizes Computer Use capabilities, allowing developers to build on a consistent, high-performance foundation. (Transcript)

  4. Decisions API for Speed 11:40

    The Decisions API is designed for extremely fast classification and structured output. It operates in parallel and is a smaller model than those used for complex Computer Use tasks, making it ideal for high-throughput, low-latency applications. (Transcript)

Technical details

  • Model & Performance 250s

    The new **GPT-6.1** model offers significant cost and speed advantages for Computer Use, reportedly being five times cheaper than Astra and seven times cheaper for specific Computer Use tasks. (Transcript)

  • Context & Input 1300s

    The **App Shots** feature enhances context by capturing not just a screenshot, but the raw metadata, accessibility representation, and DOM structure of an application, providing the Language Model (LM) with full context. (Transcript)

  • API Optimization 650s

    Performance gains are driven by techniques like **async tool calling** (allowing the model to continue reasoning while a tool runs), **mid-turn steering** (injecting messages during model reasoning), and **WebSockets** (enabling bidirectional communication). (Transcript)

  • Context Management 1800s

    Developers can manage large context windows using **compaction** techniques. Options include server-side compaction (automatic) or manual control via the `/compact` function, which is crucial for maintaining efficiency in long-running agent threads. (Transcript)

Mentioned resources

  • OpenAI DevDay (Event)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.