Webinar: AI Agent Simulation of Human Behavior with Michael Bernstein
Stanford Online
Introducing Michael Bernstein 0:10
The session opens with an introduction of Michael Bernstein, a professor of computer science at Stanford University, where he holds fellowships and has won best paper awards at major human computer interaction conferences. His work has been covered by outlets like the New York Times and the Guardian, and he has received honors including an Alfred Sloan Fellowship and the Patrick J. McGovern Tech for Humanity Prize. Over five hundred attendees joined the webinar, which focuses on AI agents that simulate human behavior.
The problem of incomplete information 3:30
Bernstein explains that decisions in organizations, whether about customers, employees, or policies, are usually made with incomplete information about how people will actually react, since you can typically only launch once and cannot easily test outcomes in advance. He traces this challenge back to sociologist Robert Merton in 1906, who observed how hard it is to predict group behavior, using the example of everyone trying to escape crowded cities at the same time only to create new crowds elsewhere. This leads him to pose the idea of a 'what if machine' that could let you preview how people might respond to a decision before you actually make it, potentially leading to far better choices.
From old simulations to language models 9:01
Bernstein traces simulation of human behavior back to Thomas Schelling's 1978 agent-based models, which won a Nobel Prize and are still used today, including in pandemic modeling, and to entertainment examples like The Sims. He notes that past models were limited, either reducing a person to a handful of parameters or relying on rigid scripted rules, and academic literature described them as highly stylized with minimal real-world impact. He explains that his team recognized modern large language models like ChatGPT, Claude, and Llama had effectively trained on vast records of human behavior, including research literature and social media, allowing them to be prompted to role-play specific people with given backgrounds and traits.
Building Smallville and generative agents 12:00
This insight led to a project called generative agents, where the team built a simulated town called Smallville populated by twenty five AI agents, each autonomously living out a day, such as an artist painting or college students sleeping in and attending class. The work drew wide attention, including a recent article by Andreessen Horowitz suggesting this approach could become the next generation of market research tools. Bernstein describes how each agent is built from a persona, such as pharmacy shopkeeper John Lynn, who is given details about his family and relationships so the simulation behaves coherently from the start.
Interacting with the simulated town 15:31
Agents in Smallville act by generating natural language descriptions of what they are doing, which are then translated into movements and emoji in a video game style environment you can watch in near real time. Bernstein describes how you can intervene directly, for example posing as a reporter asking who is running for mayor, or acting as an agent's inner voice to nudge their decisions, or even setting a toaster on fire in the game world to see an agent react and put it out. These interventions show how the simulation can be used to explore how people might respond to unexpected events or changes in their environment.
Information Spreads Through the Town 16:30
In the simulation, agents talk to each other and pass along what they learn, the way John hears from his son Eddie that Eddie is composing a piece for his college music theory class, then relays that news later to his wife. To test something more complex, the researchers gave one agent, Isabella, who runs the town cafe, a single seed idea: plan a Valentine's Day party. Nothing in the system was built specifically for party planning, yet word spread naturally, Isabella recruited her friend Maria to help decorate, and by the next day 12 of the town's 25 agents had heard about the party, with 5 actually attending. A planted detail, that Maria had a crush on another agent named Klaus, led Maria to ask him to the party on her own. Later, other researchers reran the simulation and introduced a fake radio report about a communicable disease, and attendance collapsed except for Klaus, who hadn't heard the warning; a non-infectious disease report like diabetes complications left the party unaffected.
Three Building Blocks of the Agents 21:00
The architecture rests on three components. Memory comes first, a running log called a memory stream that records everything an agent observes, filtered using retrieval augmented generation, or RAG, which pulls up memories that are recent, important, and relevant to the current moment. Reflection comes second, where the system periodically asks an agent to draw higher-level conclusions from its raw memories, such as noticing that Klaus reading about gentrification and urban design means he spends a lot of time reading. Planning comes third, where agents sketch a full day, then break it into hours, then minutes, adjusting their plan whenever something unexpected happens in the environment.
Testing Whether Agents Are Accurate 26:32
Believability is not the same as accuracy, since cartoons are believable too. Two early approaches, demographic agents built from basic stats and persona agents built from narrative descriptions, both tended to produce stereotyped results, such as a model assuming a South Korean student would simply have rice for lunch.
Building Digital Twins From Interviews 29:00
A better approach used two-hour interviews with a representative sample of a thousand Americans, based on Stanford's American Voices Project script, covering life story, community, finances, health, and politics. Each interview became the memory for a digital twin agent, and both the real person and their agent then took the same battery of surveys, including the General Social Survey and the Big Five personality index, so their answers could be directly compared.
Measuring How Well Agents Replicate People 32:00
To test accuracy, everyone in the study took the same survey twice, two weeks apart, so the researchers could compare agent predictions against how much people naturally drift in their own answers. A score of 1.0 would mean an agent matches a person as well as that person matches their own later self, while random guessing scores about a third of that. Agents built from simple demographic labels reached about 70 percent, and agents built from full interviews reached about 85 percent, with similar results across other tools like the Big Five personality index and behavioral economics games. Interviews also reduced stereotyping, since knowing only that someone is a conservative pushes the model toward caricature, while a fuller interview lets it capture more nuance. Political affiliation was the hardest trait to model, with an 8 percent accuracy gap between best and worst groups, and far-right conservatives were the toughest to simulate, partly because the underlying language models resist voicing certain views and partly because those interviewees were guarded about their true opinions. Gender and race showed smaller gaps than expected, often under 1 percent once interviews were used. In a separate test, agents replicating five published studies from a top journal got four right and correctly flagged the fifth as bad science that failed to replicate even with real people, and a colleague's larger study of 100,000 people found simulation could predict effect sizes with a correlation between 0.85 and 0.9.
Building a Useful Agent Bank 37:00
For organizations wanting to build their own set of simulated people, defining agents by a single demographic trait like just being a Republican produces stereotyping and false certainty. Using five or six demographic variables is a workable middle ground, matching people about 70 percent as well as they match themselves. The strongest approach is gathering rich data such as full interviews, and surprisingly, deleting 80 percent of an interview at random only dropped accuracy from about 85 percent to about 79 percent, suggesting these interviews carry very dense information. The catch is relevance: interviewing someone about fashion won't help predict their views on climate change or retirement planning, so the gathered data needs to actually connect to the questions you plan to ask.
Risks and a Ladder of Trust 39:00
Simulations can fail in ways that matter. In one real case, ground-truth Pew data showed 13 percent of 18 to 35 year olds felt very familiar with their retirement plan fees, but the simulation guessed just 1.2 percent, an error that could mean wrongly dismissing a sizable group. Michael frames trust as a ladder: at the bottom, asking what might possibly happen works well today; a step up, estimating qualitative attitudes also works reasonably well with rich data; higher up, quantitative predictions like exact percentages are riskier since small numeric errors can mislead decisions; and at the top, full multi-agent simulations of markets or towns aren't yet reliable enough for real decision-making. His advice is to use simulation to narrow a hundred ideas down to five, then test those five on real people, and always keep agents' background data relevant, treat rough possibility questions as more forgiving than sharp quantitative ones, and validate key questions against a small real sample.
Practical Uses Already Working 44:30
One working application lets platform designers stress-test policies before launch by unleashing simulated trolls on a system, catching problems that would otherwise only surface as a live crisis, something Michael now has students do in his course on online platform design. Another use trains soft skills like conflict and negotiation, built with a Stanford business school expert. In an experiment, one group watched a lecture on conflict strategies while another watched the same lecture plus practiced in a simulated conflict, and although both groups scored the same on a written test, only the simulation group performed better in a real conflict afterward, cutting their use of antisocial strategies by two-thirds.
Commercial spinout and open research 47:01
Bernstein mentions a startup called Simile that spun out of the Stanford research he has been describing, while stressing that much of the underlying work remains open research rather than a product pitch. He frames the broader goal as building a what-if machine that lets people think through possible futures before they happen, and then opens the floor for audience questions.
Customer service and artificial companions 49:00
Asked whether AI can handle customer service, Bernstein points to the risks of chatbots making promises companies never intended, citing a Canadian airline whose chatbot promised a bereavement refund and was later held to that promise by the courts. He expects customer service bots to keep spreading despite this, but predicts the bigger coming issue will be entirely artificial characters, like those on character.ai, with people choosing to talk to fabricated personalities instead of other humans, a trend his colleague D. Young has already written a white paper about.
Mimicry, environment, and unpredictability 52:02
He describes the agents as mimics of behavior that must first interpret human data to reproduce it, not conscious or thinking beings. Replicating a full complex system remains an open frontier, since the environment itself, missing details as small as doors in early versions of the town, shapes roughly forty percent of behavior, and agents also need a backstory of what happened to them earlier in the day. He adds that even perfect simulation cannot yield one guaranteed outcome, illustrating this with a music experiment where over 14,000 people split into 16 parallel versions of a Spotify-like platform produced different hit songs each time, meaning simulations can only offer likely outcomes through repeated, Monte Carlo-style runs rather than single predictions.
Bias, risk prioritization, and future models 55:31
On applying this to design personas, he suggests ranking risks by how costly being wrong would be, similar to agile or lean-startup thinking, and even asking a reasoning model like GPT-5 what risks to test. He confirms LLM bias affects agent behavior, noting far-right conservatives were harder to model because models resist replicating racist or sexist views, and predicts a longer-term bias toward conflict avoidance, since models trained to be helpful and harmless make agents too agreeable, which will eventually require training separate models, like the Centaur model described in a recent Nature paper, specifically to reflect real human behavior. He closes by distinguishing his what-if machine from scenario analysis, calling it a tool meant to make scenario analysis faster and more effective.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

