Topic

Scalability

All digests tagged Scalability

Dedicated Inference thumbnail

· 13:41

Dedicated Inference

This session outlines the AI inference landscape, detailing three production-grade paths for deploying custom models at scale: Serverless Inference, Dedicated Inference, and Inference on CKS. For build engineers, the key takeaway is selecting the right level of control—from rapid, managed iteration (Serverless) to explicit infrastructure control (CKS)—while maintaining a consistent, OpenAI-compatible endpoint for seamless integration into complex agentic frameworks.

Key takeaways

  1. The Inference Progression 0:15

    The deployment path moves from general-purpose, proprietary models (low control) to highly customized, controlled deployments that require explicit management of scale and cost.

  2. Serverless Inference 1:10

    Ideal for rapid model iteration and deployment, offering automatic scaling, observability, and native Weights & Biases tracing without requiring infrastructure management.

  3. Dedicated Inference 2:20

    Provides a controlled, production-ready path for custom models, allowing users to bring their own weights, select GPUs, and use open runtimes while CoreWeave handles operations and cost visibility.

  4. Inference on CKS 4:40

    Offers maximum control for ultra-high scale and latency-critical workloads. It is a fully managed service on CoreWeave Kubernetes Service (CKS), supporting runtimes like LLMD and Nvidia Dynamo for self-hosted, distributed inference.

Watch on YouTube Full article

Harness Engineering: Building the Production Cage for Powerful Domain Agents — Mike Chambers, AWS thumbnail

· 20:46

Harness Engineering: Building the Production Cage for Powerful Domain Agents — Mike Chambers, AWS

The presentation introduces 'Harness Engineering,' a critical concept for building production-grade AI agents at scale. Mike Chambers distinguishes between agents that are used (e.g., coding assistants) and agents that are built. For built agents, the harness encompasses all non-model components—such as memory, skills, tools, identity, and context management—that must scale independently. The core principle is that scaling these components separately, rather than deploying them in a single container, is essential for handling thousands of users and maintaining reliability.

Key takeaways

  1. Two Types of Agents 4:05

    Agents are categorized into 'agents we use' (productivity tools, coding assistants) and 'agents we build' (production-scale systems). The approach for built agents requires careful architectural planning.

  2. Defining the Harness 7:04

    A harness is defined by subtraction: take an agent and remove the model component; everything left over is the harness. This includes the infrastructure, skills, and tools.

  3. Scaling Built Agents 10:57

    For production agents, the harness must manage complex concerns like loop management, scaling, payments, identity, runtime, context management, and observability. Attempting to containerize everything together is incorrect for high scale.

  4. Avoiding 'Slop Ops' 10:07

    Build engineers must avoid 'slop ops' (clicking around a console to deploy resources). Instead, agents must build infrastructure using Infrastructure as Code (IaC) to maintain ownership and control over cloud deployments.

Watch on YouTube Full article

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO) thumbnail

· 56:30

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)

The discussion provides a deep dive into building highly scalable and resilient infrastructure, focusing heavily on state management challenges in large-scale distributed systems. Key engineering lessons include moving beyond simple benchmarks to model real-world failure modes (e.g., connection layer failures), optimizing for P99 latency when using object storage like S3, and adapting architecture to current cloud constraints, particularly the increasing demand for CPUs driven by AI/RL workloads.

Key takeaways

  1. Modeling Failure in CI 20:46

    To ensure system reliability, it is crucial to simulate low-level failures (like database connection loss) rather than just mocking components. The use of custom proxies or tools like `GDB` allows testing the application's failure handling at the connection layer, uncovering issues that are difficult to reproduce in production.

  2. The Importance of P99 Latency 30:27

    When designing large-scale systems, especially those involving multiple round trips (like navigating a tree structure on S3), optimization must focus on the P99 latency, not just the average (P50). This is critical for accurate performance prediction.

  3. CPU Scarcity in AI Workloads 47:25

    The demand curve for CPUs is shifting right due to AI and Reinforcement Learning (RL) workloads, which require significant CPU cycles for training and general-purpose agent execution. This scarcity is a major constraint that cloud providers are managing through power allocation.

  4. Architectural Simplicity Wins 51:27

    The principle of 'simplicity above everything' was key to the development philosophy, allowing for rapid iteration and focusing on core functionality rather than complex features. This approach helped achieve significant cost reductions (e.g., reducing a client's bill by 95%).

Watch on YouTube Full article