# Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava

## Executive summary

The core engineering challenge in enterprise AI is not the model itself, but how knowledge reaches inference. The speaker argues that teams often build AI systems by accident, treating prompt, memory, and weights as a single 'ladder.' Instead, they are three distinct tools for three different jobs: Prompt for small, stable behavior; Memory for current, large, and citable knowledge; and Weights (fine-tuning) for reflexes that have stopped changing. The key to robust AI architecture is recognizing these boundaries and building a circulating system (harness) that manages the flow of knowledge between these three components.

## Key takeaways

- Prompt: Behavior, not Facts: The prompt should house small, stable, and editable instructions defining the agent's tone, persona, or behavior (e.g., offering human escalation after three failures). It should not be used to store facts, as this increases token cost and risks 'lost in the middle' problems.
- Memory: Current, Large, and Citable: Memory (including external RAG and agent memory) is for knowledge that is current (changes faster than fine-tuning), large (cannot fit in the prompt), and citable (must point to a source). Access control (e.g., per-user scope) belongs here.
- Weights: What Has Stopped Changing: Fine-tuning (weights) should only be used for reflexes or patterns that have stabilized and stopped changing (e.g., the format or structure of medical codes like ICD-10). Fine-tuning on facts (like runbooks or product catalogs) is often a retrieval problem.
- The Circulating Architecture: A mature AI system is a loop: information moves from Memory to Prompt (context window) when a session starts, and patterns/formats move from Memory to Weights (fine-tuning) over time, allowing the agent to improve by doing its job.

## Technical details

- Knowledge Architecture Decision: The primary architectural decision is determining whether knowledge resides in the prompt, the retrieval/memory layer, or model adaptations like fine-tuning.
- RAG Implementation for Code: When building a RAG system for code, proper implementation requires code-aware chunking, using filtering, and denormalizing metadata (e.g., linking a code chunk to its specific repo and allowed users) to prevent 'RAG mush.'
- Diagnostic Question for Memory: Knowledge belongs in memory if it is too large to fit in the prompt, changes faster than the model can be retrained, and requires access control (scope per user/repo).
- Fine-Tuning Pitfall: A common mistake is fine-tuning a model on documentation or runbooks when the issue is actually a retrieval problem (the model needs the right chunks, not a change in its weights).

## Practical implications

- AI teams must adopt a diagnostic approach to determine if a knowledge gap requires a prompt update (behavior), a memory update (facts), or a model fine-tune (reflexes).
- Focusing on building the 'harness' (the system that circulates information) around the model is more critical than optimizing the model itself.
- The process of moving knowledge from Memory to Weights should only happen when human patterns have converged and the information is stable.

## Topics

AI Architecture, Retrieval Augmented Generation (RAG), Prompt Engineering, Model Fine-Tuning, Knowledge Management, ICD-10 codes, Oracle

Source: https://www.youtube.com/watch?v=qflLT3SoVbw
