Topic

AI DevCon NYC 2026

All digests tagged AI DevCon NYC 2026

Uber Burned 6x Its AI Budget in Four Months thumbnail

· 10:18

Uber Burned 6x Its AI Budget in Four Months

The video provides a deep dive into the operational costs and optimization challenges of building agentic AI systems. Key themes include the critical need for cache-aware routing to manage computational costs, the alarming rate of AI budget expenditure (e.g., Uber's 6x increase in four months), and the finding that only a small fraction of AI spending translates into shipped, meaningful code. Speakers advocate for leveraging open-weight models, implementing smart evaluation gates, and optimizing knowledge base updates to prevent unnecessary human intervention.

Key takeaways

  1. AI Budget Overruns are Common 3:32

    Uber increased its AI budget by six times since 2024, spending the entire increase within four months, leaving them out of budget for the remainder of the year. (03:12)

  2. Low Dollar-to-Shipped-Code Ratio 3:32

    Only $18 of every $100 spent on AI actually reaches meaningful code that gets shipped to users. (03:12)

  3. Caching is Essential for Agent Workloads 0:23

    Properly implementing caching, especially for output/input tokens, is crucial for agent workloads, as it limits computation to only newly generated tokens. (00:00:23)

  4. Open-Weight Models Handle Significant Workload 5:23

    Open-weight models running on owned hardware can now handle approximately 80% of the required work, allowing organizations to avoid vendor lock-in. (05:23)

Watch on YouTube Full article

Meta's VR Codebase Nobody Wanted to Touch — Until This thumbnail

· 9:34

Meta's VR Codebase Nobody Wanted to Touch — Until This

This talk explores advanced strategies for modernizing legacy (brownfield) codebases using AI agents. Speakers argue that true modernization requires reverse-engineering the underlying business specifications and entity models, rather than simply performing 'lift-and-shift' migrations. Key architectural advice focuses on decoupling the system by ensuring the development process is not overly dependent on a single AI provider, model, or hardware environment, exemplified by the use of an MCP server to parallelize work across multiple on-demand environments.

Key takeaways

  1. Brownfield Code as a City, Not a Ball of Mud 0:21

    Legacy systems should be viewed as complex, functioning cities that have evolved over time, rather than a 'ball of mud.' The goal is to evolve the system into a more understandable, buildable structure, allowing for clear paths of development.

  2. Modernization Requires Specification Extraction 2:13

    True software modernization involves reverse-engineering the use case and entity model from existing code, tests, and documentation, and then generating new code based on that specification. Simply translating an old language (e.g., COBOL to Java) is considered 'lift and shift' and is insufficient.

  3. Decoupling the AI Stack (The Three Boxes) 4:23

    Architects must be wary of dependency on three uncontrolled components: the **harness** (interface), the **model host**, and the **model** itself. Building a digital product that can quickly switch between providers is critical for resilience.

  4. Parallelizing Refactoring with Agents 6:02

    AI agents can be taught specific refactoring patterns and then used to find similar patterns across a codebase, allowing multiple agents to work in parallel. This approach was used to move VR development off a single, powerful Windows machine.

Watch on YouTube Full article

Meta, Stanford & Odevo on Agentic Coding at Scale thumbnail

· 10:10

Meta, Stanford & Odevo on Agentic Coding at Scale

The session explores scaling agentic coding adoption from a single team to hundreds of engineers. Key findings highlight that while AI tooling can drive massive organic community growth (e.g., Meta reaching 80%+ weekly usage), success is highly dependent on organizational maturity. Speakers warn that deploying agents into an organization with weak software delivery practices will worsen outcomes, emphasizing that foundational improvements—such as robust CI/CD pipelines, dedicated platforms, comprehensive testing, and established coding standards—must precede advanced AI adoption.

Key takeaways

  1. Meta's Adoption Strategy 1:19

    Meta grew an organic community from ad hoc usage to over 40 times its original size. Weekly tool usage increased from under half to above 80%, demonstrating that sustained adoption can be achieved without mandatory enforcement. (00:01:39)

  2. Performance Spread and the 10x Engineer 2:47

    Studies across 150,000 engineers show the widest performance spread ever measured. Contrary to initial hypotheses, top performance is now being achieved by individuals skilled in creating and utilizing agents. (00:02:47)

  3. Prerequisites for Agentic Coding 5:11

    The 2025 DORA report warns that pointing agents at an organization already struggling with software delivery will make things worse. Successful adoption requires fixing fundamentals first: CI/CD, a platform, tests, and coding standards. (00:05:1)

Watch on YouTube Full article

NVIDIA, Docker & Hud on Agents in Production thumbnail

· 10:04

NVIDIA, Docker & Hud on Agents in Production

The discussion explores the operational challenges of deploying AI agents in a production environment (24/7 operation). Key insights emphasize that successful agent deployment requires shifting focus from root cause analysis to comprehensive context and observability. Speakers covered topics including using agents with combined data sources (Elastic logs + ServiceNow), redesigning automated fixes for human consumption, optimizing GPU utilization during tool calls, and leveraging AI-built tracing frameworks for debugging rare bugs.

Key takeaways

  1. Context over Root Cause Analysis 2:10

    When agents are running 24/7 in production, the most critical resource is context—understanding what changed yesterday and the relationships between services. This proactive data knowledge is more valuable than traditional root cause analysis.

  2. Automated Fixes Must Convince Humans 5:40

    Simply automating investigations and opening pull requests (PRs) for high-impact fixes is insufficient, as developers often ignore them. The output must be rebuilt to convince the human developer of its value and priority.

  3. GPU Idle Time During Tool Calls 7:10

    A counterintuitive finding is that when an agent makes a tool call, the GPU sits idle. Properly accounting for this CPU-intensive period allows users to serve roughly twice as many users compared to benchmark predictions that ignore tool calls.

  4. AI-Built Tracing Frameworks 9:00

    For debugging rare bugs, the most useful investment is getting AI to build a tracing framework. Providing traces from an overnight run allows the agent to pinpoint the exact problem rather than guessing or failing to reproduce the issue.

Watch on YouTube Full article

GitHub Next & Tessl on the Self-Merging Repo thumbnail

· 10:36

GitHub Next & Tessl on the Self-Merging Repo

The discussion outlines the evolution of software development from traditional CI/CD to a new paradigm: Continuous AI. Speakers presented models where automated agents handle code improvements, testing, and merging (Paul Stack). Key shifts include viewing continuous improvement as a system-level problem rather than an individual productivity issue (Don Syme), prioritizing fixing the build system over fixing the code itself (Patrick Debois), and leveraging advanced AI tools for knowledge retrieval and proactive information gathering (Robert Overweg).

Key takeaways

  1. Continuous AI is the Third Pillar 0:20

    The development process requires three pillars: Continuous Integration (CI), Continuous Deployment (CD), and continuous AI, which focuses on automated code improvement in the repository.

  2. Agent-Driven Merging Process 2:33

    Advanced pipelines allow agents to open a pull request, pass multiple reviews/gates, push changes, and auto-merge upon successful completion. The UAT (User Acceptance Testing) gate remains critical for preventing regressions before end-user release.

  3. Focus on System Improvement 8:07

    The primary mistake is fixing the code when an agent fails; the correct approach is improving the system that produced the faulty code. This shifts focus from 'fix the code' to 'fix the system.'

  4. Knowledge Retrieval and Briefing 9:20

    AI agents can transform company knowledge into a searchable resource, allowing users to query complex information in plain language or receive daily briefings rather than managing a backlog.

Watch on YouTube Full article

Cisco & Stanford on Why Skills Are the New Code thumbnail

· 9:44

Cisco & Stanford on Why Skills Are the New Code

The industry is shifting from viewing software development around explicit code and implementation toward one centered on high-level intent and 'skills.' This paradigm requires a layered agent stack (models, tools, context, harnesses) that must be managed rigorously. Experts highlight that skill sprawl leads to failure through overlap, drift, and lack of activation visibility. Crucially, the consensus is that achieving business value relies less on deploying increasingly powerful frontier models and more on sophisticated context engineering and centralized management of skills.

Key takeaways

  1. Skills as the New Code Paradigm 0:14

    Software development is transforming from revolving around code/implementation to revolving around intent and instructions. Skills must be treated as first-class citizens, not just configuration files (Guy Podjarny).

  2. Three Failure Modes of Skill Sprawl 3:29

    Skill sprawl negatively impacts teams through: 1) Overlap (multiple isolated implementations achieving the same outcome); 2) Drift (teams using outdated versions of skills); and 3) Lack of Activation (no visibility into whether a skill is actually being used by agents or humans).

  3. Context Engineering Beats Model Size 6:59

    For achieving business value, smarter context engineering is more critical than deploying the most advanced model. Mid-tier models (e.g., Sonnet, GPT medium reasoning) are often sufficient when provided with proper context and structured skills.

  4. Instruction Following Leakage 8:48

    Empirical testing involving 500 skills across 1,000 tasks revealed that over half (55%) of the time, models followed a skill's instructions even when the skill was not loaded. This suggests valuable information is already encoded in model weights.

Watch on YouTube Full article

Anthropic, OpenAI & Thoughtworks on Context Engineering thumbnail

· 10:08

Anthropic, OpenAI & Thoughtworks on Context Engineering

The core challenge in deploying AI agents is shifting from model intelligence to context engineering. Speakers from Anthropic, OpenAI, Thoughtworks, and Tessl argue that the surrounding context—including organizational knowledge, structured guides, and robust feedback loops—is the primary multiplier for agent capability. Key technical concepts include defining new constraints (human time, attention, context window), building specialized harnesses using computational tools like codemods and static analysis, and establishing a Context Development Lifecycle (CDLC) that runs parallel to the traditional Software Development Lifecycle (SDLC).

Key takeaways

  1. Context Engineering Multiplies Intelligence 0:24

    Model intelligence alone is insufficient for durable, scalable products. Context engineering provides the necessary domain-specific knowledge required for agents to succeed within an organization.

  2. Remaining Software Constraints 5:02

    Most traditional software engineering constraints are obsolete. The three remaining foundational limits when using human-agent teams are: human time (the most scarce resource), human/model attention, and the context window size.

  3. Agent Harness Architecture 8:41

    A coding agent harness requires two components: 'guides' that proactively point the agent forward, and 'sensors' that provide immediate feedback for self-correction (e.g., static analysis, logs).

  4. The Context Development Lifecycle (CDLC)

    Humans must own the CDLC while agents handle the SDLC. This involves generating context, evaluating agent performance via runtime observability, and optimizing skills in a continuous loop.

Watch on YouTube Full article