Latent Space

Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI

Published 2026-08-21 · Duration 1:11:01

Summary

Simile AI aims to simulate human society by creating 'digital twins' of populations, moving beyond current Large Language Model (LLM) capabilities. The core thesis is that predicting human behavior requires modeling underlying 'social physics' and causal mechanisms, not just pattern recognition from web data. The company's approach integrates three primary data types—qualitative interviews, observational/transactional data, and Randomized Controlled Trials (RCTs)—to build highly accurate population and individual-level models. These simulations are intended to help solve 'wicked problems' like climate change and democratic instability by testing policies and interventions before real-world deployment.

Download summary

Key takeaways

  1. The Ambition: Simulating Society 2:00

    The ultimate goal is to simulate the world to answer complex societal questions (e.g., climate change, democratic instability) that are difficult to solve in reality. This is framed as a move from prediction to understanding the path to a desired outcome, similar to Thomas Schelling's work on agent-based modeling.

  2. Modeling Accuracy and Limitations 5:40

    Simile claims to have created digital twins that reproduce human behavior and attitudes 85% as accurately as people reproduce their own responses. They argue that frontier LLMs are optimized to be 'super rational,' while human behavior is often irrational, requiring bespoke training on behavioral data.

  3. The Necessity of Causal Data 7:28

    To model human decision-making, the most critical data is not just what people say (attitudinal) or what they do (observational), but the data describing the *cause and mechanism* of their decisions, best acquired through Randomized Controlled Trials (RCTs).

  4. Simulation vs. Prediction 10:00

    Simulation's highest form is not answering 'what will happen' (prediction), but defining the necessary steps to reach a specific goal (e.g., 'What path must we take to keep unrest to 1,000 years?'). This allows for counterintuitive, yet optimal, interventions.

Technical details

  • Generative Agents and Foundation Models 290s

    The work began with the Generative Agents paper (the 'Smallville paper') [300]. The initial interest in Foundation Models stemmed from their ability to process broad web data, which contains human behavioral data, allowing for emergent, realistic patterns.

  • Data Inputs for Behavioral Modeling 448s

    Simile structures data collection into three buckets: 1) Qualitative Interview Data (e.g., life stories, trauma); 2) Observational/Transactional Data (e.g., web scraping, purchase history); and 3) Causal Mechanism Data (e.g., RCTs, which test how changing one variable affects behavior).

  • Model Architecture and Training 780s

    The company trains two distinct model types: a Population-level model and an Individual-level model. This approach is necessary because modeling individual human behavior is a significantly harder task than modeling the general population.

  • Validation and Evals 840s

    A key validation paper involved bringing human participants to a virtual lab to complete surveys and behavioral studies. The resulting digital twins were able to replicate the source individuals' behaviors and attitudes with 85% accuracy.

Mentioned resources

  • Generative Agents paper (Research Paper)
  • Social Similac (Precursor Paper)
  • American Voices Project (Data Source)
  • Foundation Series (Thomas Pynchon) (Cultural Reference)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.