Human simulation is a representation problem
Reading time: 2 minutesI believe social and human simulation is fundamentally a representation problem. The fundamental question is how little we need to know about a person to predict them reliably, and how cheaply we can keep that knowledge up to date.
Interviews, personas, demographics, and user modeling are all different attempts at acquiring information about people. But most approaches are still quite naive when it comes to representation itself. There has been surprisingly little thought in social simulation about what a good representation of a person should actually look like (a question cognitive science has studied for decades, and where I suspect many of the answers may already lie). Until that question is answered, running more surveys and more interviews may just be a slow and expensive way of piling up information.
To me, there are three fundamental questions: (1) how to compress and store information about a person, (2) how to structure it into layers that update at different frequencies to stay aligned with the real person, and (3) how to compose and extrapolate from these representations to unseen counterfactuals and unseen individuals, ultimately predicting entire populations from limited information.
All three are directly tied to capital efficiency. How much does it cost to represent a person, keep that representation up to date, and use it to make reliable predictions across new situations and populations?
Whether this efficiency can cross a certain threshold is what will ultimately determine whether human simulation can work at scale.