# Build a Full-Stack GenAI Project in 4 Hours (FastAPI, React, Supabase)

## Executive summary

This tutorial provides an end-to-end guide for building a production-grade Retrieval Augmented Generation (RAG) application—a Document Copilot. The project uses SEC filings data and demonstrates the complete AI engineering workflow: from initial client brief analysis to setting up the full stack (FastAPI, React/TypeScript, Supabase Postgres with pgvector). Key phases covered include database schema design using SQLAlchemy/Alembic, implementing user authentication via Supabase Auth, building a front-end chat interface, and establishing a robust document ingestion pipeline that converts messy HTM files into structured Markdown chunks for vector embedding.

## Key takeaways

- Full Stack GenAI Architecture: The system is designed as a mono repo using FastAPI (backend) and React/TypeScript (frontend), connected via Supabase Postgres, which utilizes the pgvector extension for efficient vector storage and retrieval.
- Data Ingestion Pipeline: Raw SEC filings (HTM format) are processed using Dockling to convert them into clean Markdown. This structured data is then chunked, embedded via OpenAI, and stored in the database for RAG retrieval.
- Database Management: The project utilizes SQLAlchemy and Alembic for defining models (Users, Documents, Chunks, Messages) and managing schema migrations, ensuring a structured development process.

## Technical details

- Project Stack & Architecture: Uses FastAPI, React/TypeScript, Tailwind CSS, shadcn/ui, Supabase Postgres, pgvector, OpenAI for LLMs and embeddings, and Railway for deployment. The architecture is designed to handle complex data flow from user input -> backend API -> vector search.
- Database Modeling & Migration: Models include `users`, `source_documents`, `document_chunks`, `chats`, and `messages`. Database schema changes are managed using SQLAlchemy, with migrations tracked by Alembic. The use of pgvector enables vector search capabilities.
- Document Preprocessing: The ingestion pipeline converts raw HTM files into clean Markdown format using the Dockling library (a dev dependency). This step is critical for preserving table structures and ensuring high-quality input for chunking.
- Authentication Flow: User authentication is managed entirely by Supabase Auth, requiring configuration of environment variables (`SUPABASE_URL`, `SUPABASE_ANON_KEY`, etc.) and setting up callback URLs in the Supabase dashboard.

## Practical implications

- Teaches best practices for full-stack AI engineering by separating concerns (frontend/backend/database).
- Demonstrates how to build a production-ready RAG system that emphasizes source citation and fact-checking.
- Provides deep insight into managing complex dependencies, environment variables, and database migrations in real-world projects.

## Topics

AI Engineering, RAG Systems, Full Stack Development, Vector Databases (pgvector), Data Ingestion Pipelines, GitHub Repository, Supabase, Dockling Library

Source: https://www.youtube.com/watch?v=qF5il_9IwME
