# Building Blazing Fast AI-Native Apps: A Developer's Guide to the Database Underneath

## Executive summary

This guide outlines the architectural requirements for building AI-native applications, which place unique and demanding loads on the data layer, requiring capabilities beyond traditional databases. The session demonstrates how ClickHouse can serve as a unified core database, handling everything from structured OLTP transactions and high-volume event ingestion to complex vector search and ad-hoc analytics needed for agentic workflows. Key architectural considerations include managing high concurrency from unpredictable agents and ensuring low-latency data replication.

## Key takeaways

- AI-Native Data Demands: AI applications require a data layer capable of handling vector search alongside structured queries, real-time feature lookups, high-cardinality event streams, and sub-second analytics at scale. As apps mature into agentic workflows, the database must absorb data from tool calls, retries, and reasoning traces.
- Columnar Storage for Analytics: ClickHouse is an open-source columnar OLAP database. Its performance advantage stems from highly efficient data compression within columns, which reduces the amount of data read from disk into memory, thereby accelerating query execution and lowering costs.
- Agent Observability and Tracing: For agentic systems, observability must focus on product quality—knowing *how* the agent is making decisions and calling tools—rather than just system uptime. Tracing is crucial for debugging complex agent workflows.
- Handling Agent Unpredictability: Unlike predictable BI use cases, agents are highly iterative and exploratory. The architecture must support high concurrency and ad-hoc querying, requiring the database to handle unpredictable query patterns efficiently.

## Technical details

- Transactional Data Layer (OLTP): For core data (e.g., food logs, activity updates) requiring updates and deletes, a transactional database is necessary. ClickHouse supports a managed PostgreSQL service, which can be optimized with NVMe storage for fast transactions and high concurrency.
- Search and Retrieval: Semantic search requires capabilities like full-text search (using inverted indexes) and vector search (using Approximate Nearest Neighbor algorithms like HNSW). ClickHouse provides native support for these features, including AI embed functionality for indexing new data.
- Real-Time Data Ingestion (CDC): To move data from a transactional system (e.g., PostgreSQL) to an analytical store with low latency, Change Data Capture (CDC) is required. ClickHouse offers managed ingestion services (Click Pipes) and features like 'Wall Shadow,' which reads the write-ahead log (WAL) for exceptionally fast replication (benchmarked at 300,000 records/second).
- Agentic Architecture Flow: A typical AI-native flow involves a user query triggering an LLM, which uses tools (e.g., 'search foods' tool call) that execute SQL queries against the database, and the results are then used to draft a final response.

## Practical implications

- When designing for AI agents, assume highly unpredictable and exploratory query patterns, and architect the database to handle high concurrency and ad-hoc querying.
- Implement robust tracing and observability specifically for the agent's decision-making process (tool calls, reasoning steps), not just system uptime.
- Utilize CDC mechanisms (like ClickHouse's Wall Shadow) to ensure that transactional data is replicated to the analytical store with minimal latency for real-time insights.
- Leverage columnar databases for analytics to benefit from superior data compression, which improves query speed and reduces storage costs.

## Topics

AI Architecture, Data Warehousing, Vector Search, Change Data Capture, Agentic Workflows, ClickHouse, Weights & Biases, Click Pipes

Source: https://www.youtube.com/watch?v=p7yXYwEsEK8
