AI Engineer

Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava

Published 2026-10-04 · Duration 20:12

Summary

The core engineering challenge in enterprise AI is not the model itself, but how knowledge reaches inference. The speaker argues that teams often build AI systems by accident, treating prompt, memory, and weights as a single 'ladder.' Instead, they are three distinct tools for three different jobs: Prompt for small, stable behavior; Memory for current, large, and citable knowledge; and Weights (fine-tuning) for reflexes that have stopped changing. The key to robust AI architecture is recognizing these boundaries and building a circulating system (harness) that manages the flow of knowledge between these three components.

Download summary

Key takeaways

  1. Prompt: Behavior, not Facts 9:19

    The prompt should house small, stable, and editable instructions defining the agent's tone, persona, or behavior (e.g., offering human escalation after three failures). It should not be used to store facts, as this increases token cost and risks 'lost in the middle' problems.

  2. Memory: Current, Large, and Citable 10:59

    Memory (including external RAG and agent memory) is for knowledge that is current (changes faster than fine-tuning), large (cannot fit in the prompt), and citable (must point to a source). Access control (e.g., per-user scope) belongs here.

  3. Weights: What Has Stopped Changing 18:43

    Fine-tuning (weights) should only be used for reflexes or patterns that have stabilized and stopped changing (e.g., the format or structure of medical codes like ICD-10). Fine-tuning on facts (like runbooks or product catalogs) is often a retrieval problem.

  4. The Circulating Architecture

    A mature AI system is a loop: information moves from Memory to Prompt (context window) when a session starts, and patterns/formats move from Memory to Weights (fine-tuning) over time, allowing the agent to improve by doing its job.

Technical details

  • Knowledge Architecture Decision 0s

    The primary architectural decision is determining whether knowledge resides in the prompt, the retrieval/memory layer, or model adaptations like fine-tuning.

  • RAG Implementation for Code 1008s

    When building a RAG system for code, proper implementation requires code-aware chunking, using filtering, and denormalizing metadata (e.g., linking a code chunk to its specific repo and allowed users) to prevent 'RAG mush.'

  • Diagnostic Question for Memory 809s

    Knowledge belongs in memory if it is too large to fit in the prompt, changes faster than the model can be retrained, and requires access control (scope per user/repo).

  • Fine-Tuning Pitfall

    A common mistake is fine-tuning a model on documentation or runbooks when the issue is actually a retrieval problem (the model needs the right chunks, not a change in its weights).

Mentioned resources

  • ICD-10 codes (Technical Identifier)
  • Oracle (Company)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.