Topic

Build Engineering

All digests tagged Build Engineering

Your Coding Agent Deletes Its Memory After 30 Days thumbnail

· 9:25

Your Coding Agent Deletes Its Memory After 30 Days

The primary challenge in AI agent development is knowledge persistence: most learned information is lost when a session ends. The discussion outlines advanced strategies to build robust, long-term agent memory, moving beyond simple rule files like `AGENTS.md` or `CLAUDE.md`. Solutions include creating shared knowledge repositories, implementing a 'diary' to record the *rationale* behind decisions (not just who made them), and developing 'Super Agents' that process vast amounts of unstructured data to extract reusable skills and documentation.

Key takeaways

  1. Limitations of Simple Memory Files

    Storing agent memory solely in files (e.g., `CLAUDE.md`) is insufficient because the entire context must be loaded, potentially exposing sensitive information, and simple memory mechanisms can still be prone to duplication or forgetfulness.

  2. The Importance of Rationale (The Diary)

    Knowing *who* acted is merely attribution; the critical missing piece is knowing *why* the action made sense at the time. A dedicated 'diary' repository is needed to capture the reasoning and decision-making process, turning throwaway work into valuable lessons.

  3. Super Agents for Value Extraction

    To maximize the value of paid AI sessions, 'Super Agents' are proposed. These agents process codebases and multiple recorded sessions (VMs) to generate structured documentation, context, and reusable skills, preventing the loss of intellectual property.

Watch on YouTube Full article

How to Build a Company Knowledge Base (Full Tutorial) thumbnail

· 34:24

How to Build a Company Knowledge Base (Full Tutorial)

This tutorial provides a comprehensive guide on building a scalable company knowledge base using Fumadocs as the foundation and integrating Sanity as the Content Operations Platform (COP). The process moves from local Markdown files to a cloud-hosted, editable CMS, enabling advanced features like semantic search and an AI-powered chat widget. Key steps involve setting up the development environment, connecting Sanity via API keys, and deploying the CMS to a live, editable URL.

Key takeaways

  1. Initial Setup with Fumadocs

    Start by setting up the open-source Fumadocs framework using Node.js to create a foundational, beautiful documentation website structure. This establishes the basic content rendering pipeline.

  2. Integrating Sanity for CMS Functionality 30:27

    To enable non-technical content owners to edit content and provide robust search, the local content source must be switched to Sanity. This requires setting up the Sanity project, obtaining the API key and Project ID, and running `npm run Sanity setup` and `npm run import`.

  3. Advanced Search and AI Chat Widget

    Sanity automatically generates embeddings, enabling semantic search (searching by meaning, not just keywords). This foundation allows for the integration of an AI chat widget (using an LLM API key, e.g., OpenAI) that answers questions and links back to source pages.

  4. Deployment and Scalability

    The CMS can be deployed to a hosted URL using the Sanity CLI (`npx Sanity deploy`), providing a cloud-based interface for team members and AI agents to manage content.

Watch on YouTube Full article

Robot Demos Are Easy. Reliability Is Hard — Jason Ma, Dyna Robotics thumbnail

· 26:42

Robot Demos Are Easy. Reliability Is Hard — Jason Ma, Dyna Robotics

Dyna Robotics focuses on developing highly robust, generalist robotic policies for commercial deployment, arguing that high reliability is more critical than high performance in demos. The company utilizes a 'research and deployment flywheel' to build foundation models, achieving a 99.4% success rate in complex tasks like napkin folding over 24 hours. Key technical advancements include a 'pre-training data pyramid' (over 200,000 hours) and the use of reward models for scalable supervision, allowing the system to detect and recover from errors in long-horizon tasks.

Key takeaways

  1. The Reliability Gap in Robotics 10:10

    Achieving a high success rate (e.g., 99.4%) over extended periods (24 hours) is necessary for commercial viability, as standard models often stall at 80–90% success rates, making repeated tasks highly improbable.

  2. The Research and Deployment Flywheel 2:00

    Dyna Robotics combines frontier research with active commercial deployments to gather high-quality data, which informs and sharpens the focus of their model development, ensuring the product solves real-world problems.

  3. Scalable Error Recovery via Reward Models 18:59

    Instead of relying on manual oversight, the team developed reward models that score the robot's progress during complex tasks. Dips in this score signal a mistake, enabling targeted data collection and a human-in-the-loop active learning cycle for robust error recovery.

  4. Generalization Across Sites

    The model architecture is designed to generalize, allowing deployment at new customer sites (e.g., a laundromat, Red Bull events) without requiring site-specific fine-tuning or additional data.

Watch on YouTube Full article

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI thumbnail

· 28:18

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI

Deepak Pathak argues that robotics progress has stalled for approximately 70 years because the field has been treated as a hardware problem rather than a general intelligence problem. He introduces the concept of 'omni-bodied intelligence'—a single brain model applicable to any robot and any task, regardless of hardware. This approach leverages a 'data flywheel' that combines highly scalable data (simulation, human video) with high-quality, low-volume data (teleoperation) and, critically, real-world deployment data. Demonstrations include complex tasks like AirPods insertion, omelet cooking on $4,000 arms, and robust GPU assembly for NVIDIA's factory, showcasing the system's ability to handle real-world disturbances and zero-shot transfers.

Key takeaways

  1. Robotics Stagnation and the General Brain 5:57

    Robotics has historically been limited by approaching it as a hardware problem. The field is constrained by the lack of a general brain, leading to the 'Moravec's paradox' (what is easy for humans is hard for machines, and vice versa).

  2. The Data Bottleneck 9:02

    Collecting robot data via teleoperation is extremely slow and expensive. To reach the data scale of models like GPT-3, the entire US population would take over a century, highlighting the need for scalable data sources.

  3. Omni-bodied Intelligence and the Data Flywheel 12:02

    The proposed solution is an 'omni-bodied brain': one model for any robot and any task. This system utilizes a data flywheel, pre-training on scalable data (simulation, human video), post-training on teleoperation, and continuous improvement via deployment data.

  4. Real-World Deployment and Robustness

    The system demonstrates extreme robustness, performing tasks like GPU assembly in a randomized, noisy factory environment, and adapting to disturbances (e.g., recovering movement after disabling legs) without explicit mapping or planning.

Watch on YouTube Full article

What is LangSmith? thumbnail

· 5:33

What is LangSmith?

LangSmith is a comprehensive platform designed for the Agent Development Lifecycle (ADLC), enabling build engineers to build, test, deploy, and monitor LLM applications and agents. It functions as a tracing backend, providing crucial observability into complex agent behavior—which can involve dozens of model and tool calls—by tracking every step, diagnosing bugs, and facilitating continuous quality assurance through structured testing and production monitoring.

Key takeaways

  1. Agent Observability is Critical

    Agents are inherently 'black boxes'; LangSmith solves this by providing visibility into the sequence of model calls and tool decisions, which are not visible in the final output.

  2. Tracing Components 0:01

    LangSmith defines three components: a 'Run' (a single unit of work, e.g., one model call or tool call), a 'Trace' (a full pass through the agent, composed of multiple runs), and a 'Thread' (a conversation grouping multiple traces from one customer interaction).

  3. Testing and Validation Loop 0:02

    The platform uses Datasets (sets of examples), Evaluators (which score examples, potentially using an LLM-as-a-judge), and Experiments (running agents over datasets) to verify fixes and compare performance changes (regression testing).

  4. Production Monitoring 0:03

    In production, LangSmith allows online evaluators to score live traffic, generating dashboards that track scores, volume, latency, errors, and cost, and can trigger alerts or route traces to annotation queues.

Watch on YouTube Full article

Dexter Horthy: Why We Stopped Trusting AI to Write the Plan thumbnail

· 56:12

Dexter Horthy: Why We Stopped Trusting AI to Write the Plan

The discussion explores the shift in software development from writing code to managing 'software factories' powered by AI agents. The central thesis is that while AI agents can automate much of the implementation, the primary value shifts to defining and codifying *intent* (specs) and *preferences* (taste). The speaker argues that the process of continuous improvement—building the factory itself—is more critical than the act of reviewing individual code pull requests. Human review, therefore, evolves from checking syntax to verifying high-level architectural intent and system constraints.

Key takeaways

  1. The Spec is the New Code 3:44

    The industry trend is moving toward treating specifications (specs) as the primary, verifiable, and executable artifact. This approach aims to capture the full intent of a feature, which can then be compiled into code, rather than relying on the code itself as the source of truth.

  2. Context Engineering and the 'Dumb Zone' 10:03

    Context engineering is crucial for effective agentic development. Early models exhibited a 'dumb zone' where performance degraded significantly when the context window exceeded a certain token count (e.g., 100,000 tokens), emphasizing the need for intentional context management.

  3. The Value of the Software Factory 28:23

    A 'software factory' is a system that automates the entire development lifecycle (planning, building, reviewing, rolling out). The goal is to shift focus from fixing individual bugs to continuously improving the factory's processes and skills, thereby increasing overall velocity.

  4. The Persistence of Human Review 53:52

    While AI is powerful, the speaker asserts that there will always be 'alpha in reviewing something.' Human review will shift from checking code correctness to verifying high-level architectural decisions, business logic, and unique organizational 'taste' that models cannot inherently replicate.

Watch on YouTube Full article

Netlify's Dana Lawson: 'We Ain't Precious No More' thumbnail

· 10:06

Netlify's Dana Lawson: 'We Ain't Precious No More'

The landscape of software development is shifting from developer-centric to builder-centric, driven by AI agents. While agents enable non-technical users (Product Managers, designers) to open pull requests (PRs) and build applications, this transition introduces new challenges. Agents can fail by missing crucial product or design context, even when passing automated CI checks. Successful adoption requires platforms to be designed for these diverse 'builders' and necessitates that Product Managers evolve into 'agent orchestrators' who define the system's boundaries and ensure proper human control planes.

Key takeaways

  1. Agent Failure Due to Context Loss 2:24

    An agent, despite having access to skills, CI requirements, and passing automated checks, can fail by using a generic component (e.g., a generic React button) instead of a specific, context-aware component that holds critical requirements like accessibility patterns. (Marc Sloan, 00:00:24)

  2. Non-Technical PR Merge Metrics 3:26

    Across hundreds of organizations, 74% of PRs opened by non-technical individuals get merged, and 84% of those merge without any developer needing to push follow-up commits. (Tammuz Dubnov, 00:03:26)

  3. The Builder Persona Shift 5:21

    The platform is no longer built solely for developers. The rise of agents means that anyone—therapists, students, small business owners—can build, making the builder persona far broader. (Dana Lawson, 00:05:41)

Watch on YouTube Full article

Jev Explained for Python Developers thumbnail

· 16:51

Jev Explained for Python Developers

This video provides a deep dive into TypeSafe's Jev model, a novel classification model designed for building reliable, structured AI applications. Jev moves beyond traditional function calling by offering specialized methods—Choice, Score, and Null—to classify inputs. For build engineers, the key takeaways are the model's ability to facilitate complex decision-making (if/else logic) through structured API calls, coupled with significant performance advantages, being notably faster and cheaper than competitors like Claude Haiku.

Key takeaways

  1. Jev: A New Classification Paradigm

    Jev is presented as a new model category, optimized specifically for classification, which is a critical component for building reliable LLM-based systems. It is designed to be declarative, simplifying the need for complex system prompts.

Watch on YouTube Full article

The Hidden 50% Drop in AI Agents Following Your Rules thumbnail

· 29:18

The Hidden 50% Drop in AI Agents Following Your Rules

The increasing complexity of multi-agent AI coding systems has led to a critical loss of control, evidenced by a reported 50% degradation in agents' adherence to static instruction files like `AGENTS.md` and `CLAUDE.md` [00:14:28]. The talk argues that traditional agile rituals are being replaced by structured, technical controls: detailed specifications (specs), automated verification steps, and advanced merge tactics (like merge queues). To maintain control, developers must move beyond plain text instructions and adopt structured rule sets, such as those used by CodeRabbit, which force adherence across different models.

Key takeaways

  1. 50% Drop in Agent Adherence 10:28

    Baz's data shows a severe, almost overnight, degradation in the usage of static instruction files (`AGENTS.md`, `CLAUDE.md`) by coding agents, suggesting that model releases can break steering capabilities [00:14:28].

  2. The Shift from Rituals to Structure 19:12

    The bottleneck in software development has shifted from human capacity (PR bottleneck) to system consistency. The process is now being governed by three technical pillars: detailed specs (written in Markdown or linked to issues), automated verification, and advanced merge tactics [00:19:32].

  3. Structured Rules Outlast Plain Files 24:20

    Structured rule sets (e.g., CodeRabbit's JSON rule set) are significantly more effective at forcing agent adherence than plain instruction files (`CLAUDE.md`) because they provide a stronger, more consistent constraint across models [00:23:40].

  4. Long-Horizon Tasks are More Consistent 26:30

    While short tasks show high variability, long-horizon code sweeps demonstrate a larger likelihood of agents adhering to correct instructions due to the sheer number of turns and iterations, though users currently prefer faster, shorter loops [00:23:40].

Watch on YouTube Full article

How to use Jev to automate your business (Step-by-step w/ Treg) thumbnail

· 14:29

How to use Jev to automate your business (Step-by-step w/ Treg)

This talk introduces Jev, a specialized model designed for reliable, high-accuracy business automation rather than creative text generation. Unlike general-purpose LLMs, Jev is optimized for structured decision-making, providing probability distributions for a limited set of options. This makes it ideal for mission-critical workflows requiring near-100% accuracy, such as fraud detection, internal link mapping, and classifying user intent, while being significantly faster and cheaper than large general models.

Key takeaways

  1. Jev's Core Advantage

    Jev is designed for reliable, high-quality decision-making, outputting the probability of a list of given answers rather than predicting text token by token. This makes it extremely fast and cost-effective for high-volume business workflows.

  2. Confidence Scoring 2:00

    Every answer Jev provides comes with a probability distribution (confidence score). This allows developers to build sophisticated business logic (e.g., if confidence > 70%, auto-block; if 35% < confidence < 70%, request human review).

  3. Use Case: Browser Automation 3:40

    Jev can predict the next action (click, type) and the target UI element based on the DOM and interaction history, enabling fast and accurate browser and computer use for agent systems.

  4. Workflow Example: Fraud Detection 7:30

    By combining Jev with data services like Track, users can build automated pipelines to classify signups (e.g., fraud, upsell value, affiliate) using thousands of data points, making previously uneconomical automation possible.

Watch on YouTube Full article

Hugging Face's MCP Server: Only 62K of 10M Calls Matter thumbnail

· 9:44

Hugging Face's MCP Server: Only 62K of 10M Calls Matter

The video discusses the significant overhead and limitations inherent in current AI agent protocols, particularly the Model Call Protocol (MCP). Speakers highlight that complex agent interactions are often hampered by chatty, stateful handshakes and reliance on visual/pixel-based inference (the 'guessing game'). Solutions proposed include implementing Web MCP, which allows front ends to expose capabilities rather than just pixels, and building robust guardrails and validation logic directly into the protocol's plumbing (using lifecycle hooks) to prevent agents from reinventing existing components or making unauthorized calls.

Key takeaways

  1. Protocol Overhead is High 2:07

    A stateful MCP handshake is highly chatty. For every 10 million protocol messages, 1.2 million are 'initialize' events, but only 62,000 are actual tool calls, indicating significant protocol overhead (00:02:07).

  2. Web MCP Shifts Focus from Pixels to Capabilities 2:32

    Current web agents operate by observing screenshots, DOM, or accessibility trees, which is inefficient and consumes excessive tokens. Web MCP proposes letting the front end expose defined capabilities instead of relying on pixel-level guessing (00:02:32).

  3. Guardrails Must Live in the Plumbing 4:57

    Since developers cannot control what an LLM agent decides to call, guardrails must be implemented in the protocol's plumbing (e.g., using lifecycle hooks before or after a tool call) to validate outputs, such as ensuring an email is in a client's custom domain (00:04:57).

Watch on YouTube Full article

Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA thumbnail

· 18:22

Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA

The session addresses the massive engineering challenge of extracting structured, queryable data from enterprise agreements, which are often trapped in unstructured formats like PDFs. The scale is immense: Docusign processes approximately one million agreements daily, representing a potential $2 trillion in locked-up negotiated value. To solve the critical problem of complex tables (pricing tiers, SKUs, rate cards) that break generic extraction tools, DocuSign partnered with NVIDIA. They developed Neotron Parse, a purpose-built, small Vision Language Model (VLM) designed specifically as an extractor rather than a generator. This model provides single-shot processing for layout understanding, semantic structure, and accurate table preservation, achieving significantly higher speed and lower latency compared to general-purpose alternatives.

Key takeaways

  1. Scale of the Problem 3:40

    The enterprise agreement data corpus is massive, involving 1.9 million paying customers and a billion users, requiring the structuring of approximately one million agreements per day. The difficulty lies in the fact that agreements are hierarchical, not flat, and contain vital terms (pricing, SLAs) that are difficult to query.

  2. The Failure of Generic Extraction 5:59

    Traditional document extraction tools or generic VLMs fail when encountering complex tables, merged cells, or nested columns, leading to significant operational overhead for legal and procurement teams.

  3. Efficiency of Purpose-Built Models 9:40

    The Neotron Parse model, a small VLM (~850-900 million parameters), was designed for extraction, not generation. It demonstrated superior performance, running table extraction approximately 20 times faster than general alternatives, leading to lower latency and cost at scale.

Watch on YouTube Full article

If we want them to do Knowledge Work, design them as Knowledge Agents — Benjamin Clavié, Mixedbread thumbnail

· 17:55

If we want them to do Knowledge Work, design them as Knowledge Agents — Benjamin Clavié, Mixedbread

The talk distinguishes between 'Coding Agents' and 'Knowledge Agents,' arguing that most real-world tasks fall under the latter. While code provides durable cues (identifiers, file paths), knowledge work is inherently ambiguous, diffuse, and context-dependent (e.g., legal or medical research). The speaker posits that AI agents must be designed to mimic human knowledge work patterns—specifically, through advanced orchestration. This involves breaking down complex, open-ended problems, utilizing multiple specialized tools (primitives like BM25 and semantic search), and employing sub-agents (searchers) to synthesize findings into memos, thereby reducing the 'oracle gap' between perfect knowledge and the agent's output.

Key takeaways

  1. Knowledge Work vs. Coding Work

    Coding is a special, easy case because code has durable cues (identifiers, method definitions). Knowledge work, however, is defined by ambiguous information input and requires reconstructing intent and judgment, making it significantly harder for agents.

  2. The Knowledge Loop 9:15

    Human progress in knowledge work is driven by a single self-optimizing loop: better tools create new roles, new roles generate knowledge, and new knowledge demands better tools. This pattern should guide agent design.

  3. The Necessity of Orchestration

    Effective agent performance requires more than just a single tool. The most significant gains come from architectural improvements, such as having a main agent delegate tasks to specialized 'searcher agents' that return structured memos, which reduces the 'oracle gap' (the difference between perfect information and the system's output) by up to 40%.

  4. Tooling is not Neutral

    Tools are not merely incremental improvements; they are critical for overcoming performance ceilings. The ability to use a tool (e.g., a library catalog vs. physically searching archives) determines if a task is scalable and cheap enough to be practical.

Watch on YouTube Full article

Docker, Adobe & tldraw: Where Should Your Agent Run? thumbnail

· 10:08

Docker, Adobe & tldraw: Where Should Your Agent Run?

The discussion explores the critical architectural question of where AI coding agents should execute, presenting four distinct models: Docker advocates for secure microVM sandboxes; Helix ML proposes centralized, dedicated computing resources for each agent; Adobe demonstrates running the agent loop entirely within the browser tab; and tldraw visualizes agents collaborating as characters on an infinite canvas. The consensus highlights the trade-offs between isolation, centralized control, and environmental fidelity.

Key takeaways

  1. MicroVMs for Agent Sandboxes 0:09

    Docker recommends using micro VMs instead of traditional containers for agent sandboxes because enterprise security teams view shared kernels as an unacceptable isolation boundary.

  2. Centralized Agent Infrastructure 3:03

    Helix ML argues for giving every agent its own dedicated computer on centralized infrastructure (e.g., Kubernetes) to facilitate seamless handoffs of work across global time zones.

  3. Browser-Native Agent Loops 7:33

    Adobe demonstrated an agent that runs its entire loop and controls the browser from within the browser tab, showcasing the concept of the 'self-licking ice cream cone' (SLICC).

  4. Collaborative Canvas Agents 6:13

    tldraw presents agents as interactive characters on a canvas that can coordinate, plan, and execute tasks as a team, allowing for simultaneous, visible collaboration.

Watch on YouTube Full article

How Lyft Increased Its Agent Resolution Rate by 16% with LangSmith and LangGraph thumbnail

· 3:41

How Lyft Increased Its Agent Resolution Rate by 16% with LangSmith and LangGraph

Lyft addressed the challenge of scaling its customer support agent stack by replacing brittle, deterministic agents with a meta-agent architecture built on LangGraph and LangSmith. This new self-serve platform allows non-engineering personnel (PMs and ops) to deploy new agents via simple configuration and prompting, drastically reducing agent build time from six months to one to two weeks. This accelerated iteration cycle resulted in a 16% increase in the customer resolution rate.

Key takeaways

  1. Shift to Self-Service Agent Platform 2:25

    The team created a platform enabling PMs and ops to build and ship agents using domain knowledge and natural language prompting, minimizing the need for code changes (merely a config change).

  2. Architectural Improvement via Meta-Agent 3:35

    The system utilizes a meta-agent where all sub-agents are registered dynamically as nodes in the meta-agent, simplifying the composition and deployment of new agents.

  3. Significant Operational Gains

    The agent build time was reduced from six months to one to two weeks, allowing engineers to focus on complex, foundational improvements while increasing the overall resolution rate by 16%.

Watch on YouTube Full article

I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI thumbnail

· 16:04

I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI

The presentation compares the monitoring of decentralized, hyper-local crime data (Hamming's initial work) with the rapidly scaling, centralized risks of conversational voice agents. While voice AI is advancing rapidly, reliability remains the primary blocker for large-scale deployment. The speaker emphasizes that because voice agents are centralized, a single prompt or architectural change can have a massive 'blast radius.' He advocates for a continuous monitoring loop—including deep manual analysis, frequency/severity prioritization, and adversarial red teaming—to mitigate risks like unauthorized actions, incorrect information provision, and the leakage of PHI/PII.

Key takeaways

  1. Voice Agents vs. Crime Monitoring 7:12

    Crime incidents are generally hyper-local and decreasing, while voice agent usage is centralized and rapidly increasing, potentially handling a trillion calls annually. This centralization means a single failure point can impact millions of users.

  2. The Scale of Risk 8:43

    If a 1% error rate is assumed across annual calls, this equates to 10 billion potential bad interactions. In practice, monitoring 10,000 agents shows an error rate closer to 10%, manifesting as skipping eligibility checks or providing incorrect information.

  3. The Continuous Improvement Loop 11:44

    Fixing voice agent reliability requires a structured loop: Identify problems, prioritize by frequency and severity, understand the fix, execute the change, verify it hasn't caused regressions, and continue monitoring in production.

Watch on YouTube Full article

Tolan: Voice-First AI Companion — Paula Dozsa, Tolan thumbnail

· 15:08

Tolan: Voice-First AI Companion — Paula Dozsa, Tolan

Paula Dozsa, an engineer on the Tolan team, details the unique engineering challenges of building a voice-first AI companion compared to traditional text-based LLM applications. She emphasizes that voice introduces 'conversational volatility' (fast turns, interruptions) which requires fundamental shifts in pipeline design, including smart turn detection, tiered model routing based on emotional stakes, and rebuilding context per turn rather than relying on a continuous cache. The talk also covers using advanced AI agents (like Claude) to build the product itself, achieving significant improvements in stability and feature development.

Key takeaways

  1. Voice vs. Text LLM Assumptions 4:05

    Text chat assumes slow turns and stable context, while voice is characterized by fast turns and volatile context (interruptions, subject changes mid-sentence). This volatility requires building for messy, real-world speech patterns.

  2. Optimizing for Interruptions 5:38

    Instead of minimizing interruptions, the team focused on building smart turn-taking that reads speech patterns to prevent early, incorrect agent interventions. They paid an extra 60 milliseconds of latency to achieve this.

  3. Memory as Retrieval System 12:32

    To handle volatile context, memory is treated as a retrieval system, not a transcript. Facts and preferences are embedded, stored in a vector database (with sub-50ms lookups), and compressed nightly to resolve contradictions and merge duplicates.

  4. Context Reassembly

    Context must be reassembled from parts every single turn (summary, user persona, retrieved memories, tone guidance) because reusing old context is a 'trap' when the user pivots subjects.

Watch on YouTube Full article

Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind thumbnail

· 16:42

Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

Google DeepMind presented research on Speech-to-Speech models, positioning them as the foundation for the 'agentic future' of voice interaction. The core argument is that modern models must move beyond simple Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) chains. By leveraging natively multimodal pre-training (audio, video, text), these models achieve a 'trifecta' of conversational fluency, high intelligence (task completion/reasoning), and multimodality (handling video, screen shares, and documents). This enables complex, real-time applications like live multilingual translation and proactive, low-latency conversational agents.

Key takeaways

  1. The Shift from Cascaded to Unified Models 2:31

    Historically, speech processing required multiple hand-built components (feature extraction, acoustic modeling, language modeling, rescoring). Modern LLMs, trained on interleaved multimodal examples, collapse this chain, allowing a single model to understand and transition between audio, video, and text inputs.

  2. The Three Pillars of Speech-to-Speech Models 10:01

    A robust model must balance three vectors: 1) Conversational (low latency/snappy); 2) Intelligent (task completion, instruction following); and 3) Multimodal (accepting video, screen shares, and documents). Improving one vector often degrades the others (e.g., increasing intelligence can decrease time to first audio).

  3. Real-Time Multilingual Translation 5:27

    The model can perform streaming, real-time translation across multiple speakers and languages (e.g., English, Spanish, Italian, Chinese) with quality comparable to offline systems, a capability difficult for cascaded systems.

  4. Proactive Audio and Multimodal Output 13:57

    Advanced features include 'proactive audio,' where the model knows when to respond despite background noise or when another speaker is talking. Furthermore, the model can generate multimodal output, including customized real-time avatars with low-latency lip-syncing.

Watch on YouTube Full article

Event Recap: Build Smarter Voice Agents - New York Edition thumbnail

· 29:13

Event Recap: Build Smarter Voice Agents - New York Edition

This recap details the complexities of building and deploying production-grade voice AI agents across two distinct sectors: professional networking (Boardy) and regulated healthcare (Flagler Health). Key challenges discussed include maintaining conversational flow, establishing user trust, managing multi-party video meeting interactions, and ensuring subsecond latency for natural conversation. The discussion highlights the difference between highly structured, goal-oriented flows (healthcare) and highly conversational, relationship-driven interactions (networking).

Key takeaways

  1. Design Flow Differences 10:20

    Healthcare voice agents require highly structured, step-by-step flows with strict guardrails (e.g., collecting insurance info) to prevent medical advice or deviation. Conversely, networking agents are designed to handle highly conversational, open-ended interactions to facilitate connections.

  2. The Importance of Trust and Disclosure 21:20

    Building user trust is critical. Speakers emphasized that being upfront and immediately disclosing that the user is speaking to an AI (e.g., 'I'm Sarah and AI') is essential to prevent user frustration and loss of trust.

  3. Technical Challenge: Multi-Party Meetings 24:10

    Handling voice agents in multi-person video meetings (like Google Meet) is technically difficult. The primary challenge is determining when the agent should speak (turn-taking) to avoid false positives (randomly jumping in) or false negatives (failing to reply).

  4. Achieving Low Latency 25:00

    To feel like a natural conversation, the system must achieve subsecond latency. This requires advanced architecture, such as preemptively generating the entire voice pipeline while the user is speaking.

Watch on YouTube Full article

Agents Without Code: Skills, YAML, and Filesystems Replaced Python — Philipp Schmid, Google DeepMind thumbnail

· 18:28

Agents Without Code: Skills, YAML, and Filesystems Replaced Python — Philipp Schmid, Google DeepMind

The presentation details the evolution of LLM agents, demonstrating a shift from complex, brittle Python code loops to declarative, file-based definitions using system instructions and skills. The speaker shows that modern agent architectures, such as the Gemini API's anti-gravity agent, utilize a hosted sandbox and network proxy to manage state and credentials securely. This allows agents to operate using general-purpose tools (like GitHub CLI or Google Search) defined in files (e.g., `AGENTS.md`), drastically reducing the need for thousands of lines of custom orchestration code.

Key takeaways

  1. The Agent Evolution: Code to Files

    Agent development is moving away from writing explicit Python loops, JSON schemas, and tool routing logic. The core functionality is now expressed in files (Markdown/Skills) that define instructions, rules, and capabilities, allowing the model to use general tools.

  2. Server-Side State Management 14:02

    The new architecture handles complex tasks by moving loops, tool routing, session state, and context compaction to the server side, requiring only a single API call with new inputs.

  3. Security and Isolation 12:32

    A hosted sandbox and network proxy ensure that the agent never sees the actual credentials, injecting tokens only when outbound requests are made, and allowing domain restriction for enhanced security.

  4. Focus on Domain Logic 16:40

    The primary work for developers is now defining the domain instructions, rules, and evaluation criteria (the 'what'), rather than writing the infrastructure code (the 'how').

Watch on YouTube Full article