Weights & Biases

Building Blazing Fast AI-Native Apps: A Developer's Guide to the Database Underneath

Published 2026-10-08 · Duration 29:18

Summary

This guide outlines the architectural requirements for building AI-native applications, which place unique and demanding loads on the data layer, requiring capabilities beyond traditional databases. The session demonstrates how ClickHouse can serve as a unified core database, handling everything from structured OLTP transactions and high-volume event ingestion to complex vector search and ad-hoc analytics needed for agentic workflows. Key architectural considerations include managing high concurrency from unpredictable agents and ensuring low-latency data replication.

Download summary

Key takeaways

  1. AI-Native Data Demands 2:00

    AI applications require a data layer capable of handling vector search alongside structured queries, real-time feature lookups, high-cardinality event streams, and sub-second analytics at scale. As apps mature into agentic workflows, the database must absorb data from tool calls, retries, and reasoning traces.

  2. Columnar Storage for Analytics 7:40

    ClickHouse is an open-source columnar OLAP database. Its performance advantage stems from highly efficient data compression within columns, which reduces the amount of data read from disk into memory, thereby accelerating query execution and lowering costs.

  3. Agent Observability and Tracing 14:20

    For agentic systems, observability must focus on product quality—knowing *how* the agent is making decisions and calling tools—rather than just system uptime. Tracing is crucial for debugging complex agent workflows.

  4. Handling Agent Unpredictability 18:20

    Unlike predictable BI use cases, agents are highly iterative and exploratory. The architecture must support high concurrency and ad-hoc querying, requiring the database to handle unpredictable query patterns efficiently.

Technical details

  • Transactional Data Layer (OLTP) 320s

    For core data (e.g., food logs, activity updates) requiring updates and deletes, a transactional database is necessary. ClickHouse supports a managed PostgreSQL service, which can be optimized with NVMe storage for fast transactions and high concurrency.

  • Search and Retrieval 380s

    Semantic search requires capabilities like full-text search (using inverted indexes) and vector search (using Approximate Nearest Neighbor algorithms like HNSW). ClickHouse provides native support for these features, including AI embed functionality for indexing new data.

  • Real-Time Data Ingestion (CDC) 520s

    To move data from a transactional system (e.g., PostgreSQL) to an analytical store with low latency, Change Data Capture (CDC) is required. ClickHouse offers managed ingestion services (Click Pipes) and features like 'Wall Shadow,' which reads the write-ahead log (WAL) for exceptionally fast replication (benchmarked at 300,000 records/second).

  • Agentic Architecture Flow 1400s

    A typical AI-native flow involves a user query triggering an LLM, which uses tools (e.g., 'search foods' tool call) that execute SQL queries against the database, and the results are then used to draft a final response.

Mentioned resources

  • ClickHouse (Database)
  • Weights & Biases (Partner/Product)
  • Click Pipes (Managed Service)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.