# Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI

## Executive summary

Simile AI aims to simulate human society by creating 'digital twins' of populations, moving beyond current Large Language Model (LLM) capabilities. The core thesis is that predicting human behavior requires modeling underlying 'social physics' and causal mechanisms, not just pattern recognition from web data. The company's approach integrates three primary data types—qualitative interviews, observational/transactional data, and Randomized Controlled Trials (RCTs)—to build highly accurate population and individual-level models. These simulations are intended to help solve 'wicked problems' like climate change and democratic instability by testing policies and interventions before real-world deployment.

## Key takeaways

- The Ambition: Simulating Society: The ultimate goal is to simulate the world to answer complex societal questions (e.g., climate change, democratic instability) that are difficult to solve in reality. This is framed as a move from prediction to understanding the path to a desired outcome, similar to Thomas Schelling's work on agent-based modeling.
- Modeling Accuracy and Limitations: Simile claims to have created digital twins that reproduce human behavior and attitudes 85% as accurately as people reproduce their own responses. They argue that frontier LLMs are optimized to be 'super rational,' while human behavior is often irrational, requiring bespoke training on behavioral data.
- The Necessity of Causal Data: To model human decision-making, the most critical data is not just what people say (attitudinal) or what they do (observational), but the data describing the *cause and mechanism* of their decisions, best acquired through Randomized Controlled Trials (RCTs).
- Simulation vs. Prediction: Simulation's highest form is not answering 'what will happen' (prediction), but defining the necessary steps to reach a specific goal (e.g., 'What path must we take to keep unrest to 1,000 years?'). This allows for counterintuitive, yet optimal, interventions.

## Technical details

- Generative Agents and Foundation Models: The work began with the Generative Agents paper (the 'Smallville paper') [300]. The initial interest in Foundation Models stemmed from their ability to process broad web data, which contains human behavioral data, allowing for emergent, realistic patterns.
- Data Inputs for Behavioral Modeling: Simile structures data collection into three buckets: 1) Qualitative Interview Data (e.g., life stories, trauma); 2) Observational/Transactional Data (e.g., web scraping, purchase history); and 3) Causal Mechanism Data (e.g., RCTs, which test how changing one variable affects behavior).
- Model Architecture and Training: The company trains two distinct model types: a Population-level model and an Individual-level model. This approach is necessary because modeling individual human behavior is a significantly harder task than modeling the general population.
- Validation and Evals: A key validation paper involved bringing human participants to a virtual lab to complete surveys and behavioral studies. The resulting digital twins were able to replicate the source individuals' behaviors and attitudes with 85% accuracy.

## Practical implications

- Businesses can use simulations for 'concept testing' (testing different messaging or products) and behavioral experiments, moving beyond simple market research.
- The technology allows companies to model specific, niche populations (e.g., 20s-30s in California) rather than just the entire US market.
- It enables the testing of complex policies and interventions (e.g., for CPG companies or financial services) by simulating the ripple effects across an entire society.
- The ability to reuse the same modeled individual across multiple, domain-agnostic studies saves significant time and cost compared to running physical human panels.

## Topics

Generative AI, Behavioral Modeling, Digital Twins, Simulation Theory, Agent-Based Modeling, Randomized Controlled Trials (RCTs), Foundation Models, Generative Agents paper, Social Similac, American Voices Project, Foundation Series (Thomas Pynchon)

Source: https://www.youtube.com/watch?v=KpOW9Pk4BUs
