人脑气候预报:AI合成角色的技术原理、失效模式与前沿实践 AI Engineer 2026-07-29

范式转移:从角色提示到合成受访者

在这个大语言模型时代,大部分人都使用过角色提示词(Role Prompting)。这种通过设定特定身份来引导模型输出的原理,现已发展为合成角色(Synthetic Personas:利用大语言模型模拟特定群体或个人属性的AI智能体)这一全新的技术类别。企业正利用这些合成受访者来测试产品概念与文案,推动该领域从新奇尝试走向市场爆发。

作者将合成角色比作天气预报(Weather Forecasting:基于数据和算力对复杂系统未来状态进行的概率性预测)。正如天气预报得益于算力与数据的爆发,合成角色也是在相同背景下被解锁的。然而,正如天气预报只能预测未来数天一样,合成角色也运行在特定的边界和机制内,超出这一范围准确度就会急剧下降。因此,理解其局限性与理解其潜力同样重要。为了剥离市场上的宣传泡沫和盲目贬低,我们需要深入探索底层的技术细节。

Original English Source

Hello, I'm Ishan and welcome to can AI predict people like we predict the weather. A field guide to the nascent field of synthetic personas. Now, I'm sure most of you in this room at some point or another have prompted a large language model with a role prompt. You are a fill in the blank and then the task. Believe it or not, that core principle of steering a model's outputs as if it were a particular person or persona has turned into an entire category that companies are using to test product concepts and messaging against synthetic respondents. And it has moved from a novelty to market momentum, as you can see from both these headlines as well as the increase in funding for the last few years. And the running analogy I want to leave you with is that synthetic personas are like weather forecasting. Like weather forecasting, they were unlocked thanks to an increase in compute and data. And like weather forecasting, they operate within a particular regime and going past that sometimes can go outside of where they're accurate. So, for example, you can only predict the weather a certain number of days in advance. Similarly with synthetic personas, there's only so far you can go before you'll run into issues. And understanding those issues are as important as understanding their promise and their potential. So, I'm Ishan Nand. I'm the chief AI officer at Insight Sciences. We construct LLM synthetic personas for market research and market insights teams. And the reason for this talk is that most of the coverage in the space is very shallow, doesn't go into the technical details. It's either outright hype or outright dismissal. And it's really hard to separate the noise from what's real. So what I want to cover is that messy middle of the technical details. And you don't have to take my word for it, even though I'm a vendor in the space, because everything I'm going to talk about today is going to be based on published research. So we're going to cover why now for synthetic personas, how they fail, some techniques to inspire you, and then some metrics to judge whether your synthetic persona is accurate or not.

语言媒介:重塑人类行为预测的基石

在建立起天气预报这一认知框架后,我们不妨回顾历史。早在20世纪50和60年代,伴随早期计算机的诞生,Simulmatics公司就曾尝试利用粗糙的统计数据和当时的算力来模拟和预测选民行为,但最终宣告失败。然而,当今大语言模型的崛起为我们带来了全新的模拟媒介(Simulation Medium:将语言本身作为建模的基本原子单位,而非传统的数学公式)。过去,模拟必须将人类情感与决策转化为复杂的公式;而语言模型的本质是直接对人类的语言和逻辑进行建模,从而提供了一个前所未有的中间层。

通过实证研究证明,这种基于语言媒介的模拟过程是切实可行的。在一项极具代表性的研究中,研究人员对约一千名人类受访者进行了长达两个半小时的深度访谈、性格测试和问卷调查。随后,他们将这些访谈文本输入给AI代理,并让其参与相同的问卷和测试。结果表明,AI代理在预测人类行为和态度上的一致性(Alignment: 模拟输出与人类真实反应的契合度)达到了约83%。这有力地证明了利用语言建模进行人类群体预测的可行性。

Original English Source

Speaking of weather forecasting, another parallel is just like in the 1950s and '60s, we got computers that promised us, correctly, a future of accurate weather forecasts. We were also promised, believe it or not, people forecasts. This company, Simulmatics, were extensively covered by Jill Lepore, promised that they could simulate and predict the electorate using raw statistics and the computational power at the time. Fortunately, that turned out not to be the case. So you should approach claims like this with some humility. But we have something they did not have then. And that unlock is, again, more computational power, but also better modeling thanks to LLMs. And LLMs unlock a new kind of simulation. For the longest time, to simulate something meant to mathematize it in formulas or equations. But certain things, how we feel, how we act, what choices we make, aren't always succumbing to the equations. What LLMs offer us is a new medium, a new atomic unit of language itself that we can model against. Now granted, they are based on math under the hood, but it gives us this intermediary layer that we can construct and simulate against that we couldn't before. And the process can work. I want to share with you one of the most well-known demonstrations of the field. What they did is they took about a thousand humans. They put them through about two and a half hours of extensive interviews about their background and their views and their attitudes. And they put those people through a battery of personality tests and surveys. Then they took those transcripts and they passed it to an AI agent and they had the AI agent take the same set of surveys and personality tests. What they found was as the agents were basically about 83% aligned and predictive to the corresponding humans they were modeled against. Now one caveat is that number is normalized against the uncertainty and noise of the humans themselves. It's a theme we're going to come back to at the end of this talk.

即兴表演:隐性变量下的模型失控

虽然合成角色的预测潜力显著,但如果忽视其特有的失效模式,就极易被误导。失效模式一:隐性混淆变量(Latent Confounders:未在提示词中显式声明但被模型隐含关联的上下文变量)。在一项购买意愿测试中,人类的行为符合标准经济学规律,即随着价格上涨,购买概率下降。然而,大语言模型的预测曲线却呈现出反常的“倒U型”,即价格上涨时购买意愿反而上升。

通过后续实验,研究人员发现模型将“价格”自动当作了产品品质、保质期及竞品价格等潜在变量的代理指标。当提示词中缺失这些背景上下文时,模型就会开启即兴表演(Improv:模型在信息不全时根据概率联想自主补全背景的现象)。因此,为了防止模型胡乱推测,我们必须对合成角色进行深度锚定(Rich Grounding:在提示词中详尽描绘实验环境、心理动机以及研究设计本身的上下文),使模型拥有一个完全封闭且确定的运行宇宙,防止混淆变量干扰预测。

Original English Source

But don't get too excited because synthetic personas are different from regular experiments and they're liable to confuse and fool you if you don't know how they fail. So I'm going to cover three important failure modes that you need to know about when dealing with synthetic personas. To understand the first one, I want to consider this prompt these researchers gave. It's a very un It's a very ambiguous and very unsophisticated prompt. It basically says you're a customer, I'm going to show you a product, I'm going to tell you the category, I'm going to give you its price. Those are going to be the variables in the template and then I'm going to ask you to say whether you're going to purchase or not purchase. Willingness to pay, willingness to purchase is basically the test. What the researchers did is they recruited a panel of humans and put them through the same test and then they put the synthetic personas through the same test. What they found is very interesting. So the humans are here in red. They do exactly what you would expect from basic economic theory. As the purchase price increases, we see that the purchase probability goes down, slopes downward. But the LLMs did something different. They had this inverted U-shaped curve. And particularly problematic is this area right here, where as the price is increasing, the purchase probability is going up. That seems really bizarre. Through a series of additional experiments, what they discovered was that the LLM was using the price as a proxy for other properties about the product that the humans were considering were fixed. Things like the expiration date based on the price, what the price of competing products were also as the price changed. And those correlations, those latent confounders that weren't clear and immediate, were actually confusing the result. And the way to think about this is when an LLM is missing context, it has to potentially infer or invent confounders. Right? When we do a human experiment, if I put it like a gold watch on a table, I ask a human to walk in and estimate the price of it, everything about the environment is fairly fixed. The human and their decisions are the random variable. In a synthetic experiment, if you don't set it up properly, other parts of it actually become part of the random variable itself. I like to say if it's a poorly grounded persona, it's a little like the LLM is playing improv with you. It's like gold watch on a table? Oh, well, we must be in a jewelry store, right? It has to infer what's likely. And maybe this is a rich person, so they're more likely to purchase. And so the lesson is, we need to richly ground our personas in the personality, the context, and bizarrely, even the study's own construction. In a human subject experiment, you want to hide the study construction from the participant. But in the case of an LLM, they have no universe other than what's in the prompt, and you have to use the prompt to paint the world to prevent any type of confounders.

知行鸿沟:语言偏差与行动预测的错位

除了隐性变量的干扰,合成角色还面临着其他技术局限。失效模式二:提示词敏感性(Prompt Sensitivity:模型输出高度依赖于微小的提示词格式变化)。在选择题测试中,仅仅调换“是”与“否”的选项顺序,模型就会表现出极强的顺序偏见,最终导致结果均值被稀释为无意义的50/50分化。对此,我们必须进行鲁棒性测试(Durability Testing:通过改变顺序、措辞或引入对抗性挑战来评估模型态度一致性的测试)。

在解决格式敏感度的同时,我们还要直面失效模式三:知行鸿沟(Say-Do Gap:语言模型基于“人类如何说话”进行训练,而非“人类如何行动”)。这导致模型在预测表态性态度时效果较好,但在预测具体行为(如实际去健身房的频率)时准确率大幅下降。因为具体行动在训练数据中较少被直接记录,且无法直接转化为文本。为了解决这一痛点,在实际应用中,我们需要设计能够通过态度来三角定位(Triangulate:通过交叉比对多维度态度数据来推导真实行为概率的方法)出具体行为的间接问卷。

Original English Source

Another failure mode is prompt sensitivity. So, here's a researcher that took a question, they give the same question, same choices, they just swapped the order of the choices. Yes was the first one in the first question, yes was the second option in the second question. What they found was that the model had an extremely strong order bias. Basically, when they took the two results and they averaged them together, it washed out into noise, into 50/50. Now, humans do have a first order bias, but not to this extent. And so, the lesson here is that we need to durability test our personas to understand how they will change under reorderings, under rewordings, and even adversarial challenges to their opinions. The third and final area that I want to highlight is that LLMs are trained on what people say, and they're not trained on what people do. So, as a consequence, predicting stated attitudes tend to be easier than predicting actions or behaviors. Both because they're clear and likely to be in the text, but also because they are natively text themselves. So, this chart is from a bunch of researchers that used an LLM to try and predict known social science experiments. The original point of this chart is to show that the LLMs are about as good as the experts. LLM is in black in a circle, the experts are in blue, and you can see they're both doing about equally well in making the prediction. But, the point I want to draw you to is that there are two categories of experiments here. The top are surveys, those are natively language and text-based, and those reflect attitudes. And on the whole, the models tend to do better there. The bottom half is field experiments. Those are behaviors, and those are things that need to be transcribed into actions. They're less likely to be in the training data, and correspondingly, the LLM doesn't do as well. So, the lesson we often tell our clients is consider questions that triangulate to behavior from attitudes. As a hypothetical example, if you want to know about gym attendance, you might be better off, well, you can ask about both, but asking about attitudes towards working out rather than asking about attendance and see if that's a suitable proxy.

校准校准:从简单微调到语义映射

在掌握了失效模式后,我们可以采取相应的策略来优化合成角色。首先是基础的提示词校准(Prompt Calibration),如早期研究中使用的自述体文本补全。当简单的提示词无法满足要求时,第二种技术是微调(Fine-Tuning:在特定小规模数据集上继续训练模型以拟合特定分布)。研究表明,通过对特定群体的人类数据分布进行微调,不仅受训群体的拟合度大幅提升,甚至未参与微调的群体也同步改善。这表明微调的主要作用是帮助模型掌握“问卷作答”这一任务格式,释放其本就蕴含的泛化知识。

第三种前沿技术是语义分布映射(Semantic Distribution Mapping:通过对比模型生成的自由文本与人类撰写的锚定文本的语义相似度,来重构概率分布的方法)。该技术不再强求模型输出冰冷的“1到5”分值,而是让其输出自然的评语(例如“如果价格便宜我会试试”),随后通过计算该评语与人类极值文本(如“绝对不买”或“买20个”)的语义相似度,从而精准还原人类决策的离散分布,防止模型陷入“均值模糊”的陷阱,完整保留了群体决策的多样性与波动性。

Original English Source

Okay. Now, let's talk about three example techniques to kind of inspire your own synthetic personas. So, the first one is just prompting the model. Uh this right here is from the Argyle paper, which is really one of the seminal papers in this field. In fact, it's so early that the model they used was a text completion model. That's why this prompt isn't in the form of a chat. It's a statement of I am. So, they gave it a prompt that said, and for example, the middle column is basically where the context is. I am a strong liberal. I support progressive values, etc., etc. And at the end it says, "In 2016, I voted for." And they basically let the model sample its completions and it says, uh Hillary Clinton, Bernie Sanders, Hillary Clinton, and so forth. And you can see what happens with the conservative case on the top. Since this time, obviously, there've been a lot more prompting techniques. And a lot more models. And I can't tell you which prompting technique and which model is going to work best for your use case. What you are going to have to do is figure it out empirically by validating against some known human ground truth data. You'll have to do what these guys did. So, for example, here in this research, they're trying to figure out how well they can construct personas to represent voting patterns. What they found was they compared here on the left is reality and on the right is their four different types of persona constructions. And they didn't realize it at the time, but their persona construction was actually amplifying bias within the model as they got more and more detailed. And they found it was actually throwing it further and further astray from reality. So, you probably have a bunch of different ideas. The answer is you're going to have to test it and validate it against ground truth. The other natural thing you might expect is well, hey, we can fine-tune it, especially if it's missing data that isn't there, especially for example, if it's behaviors or something that wouldn't be in the training text. And this is a great paper to be inspired by for this. This is the Subpop paper. Basically, they construct a prompt template, which is the demographic information, then the survey question they want to ask, and then they compare the known human data distribution to the distribution that comes out of the model, and they do fine-tuning until the model and the human data align. Now, here's the interesting thing. When they did this, as you'd expect, the results that were from the populations they gave to the model, that's the ones in blue, improved. But very interestingly, the ones in white also improved by almost the same degree. Alignment improved even for the unseen groups. That seems almost magical. And some subsequent research has hinted that what might be really happening here is that the model itself has a latent understanding of these groups. It just didn't know how to express it in the format of surveys. And if you think about it, LLMs aren't used to doing surveys as a task. And so they aren't going to be as good as fitting it, especially to a prompt format they may not have seen on the first go-around. But fine-tuning actually is helping it learn the task or how to express itself. So, a lesson you can kind of take away is that your persona that you're looking for is in there. We just need to figure out the way to summon it or elicit it. And that lesson actually takes us to the third technique, which I want to highlight to show how sophisticated your techniques can get if you're just using so-called prompting alone, but using careful calibration and thinking. So, in this one, this team did something very clever. They set up a system prompt that was demographics. They showed a product concept. And then they asked how likely would you be to purchase this product? And they gave it the same scale from one to five, five being most likely, one being the least likely to purchase, like you'd expect, kind of your basic naive prompting pattern. And then they said, "Well, you know, hearkening back to that paper, although I don't know if they were inspired by it, they said, 'Well, large language models aren't used to doing surface, but they are more used to expressing themselves in text.' So, they said as instead of giving us a one to five rating, give us a set of text. So, the example here is, 'I'm somewhat interested. If it works well and isn't too expensive, I might give it a try.' And then to map that text to the one through five willingness to pay, they had humans write out corresponding text for what they would expect. So, if it's a one, "Hell no, I'll never buy that." Five, "Absolutely, I'll buy 20." Right? They had them write out examples of each one of the different options, and then they measured the semantic similarity between the text that came out of the model and those human examples. And that gave them a vector over which they can basically measure a probability distribution of where this text that came out of the model lands. So, what I like about this is it's actually a distribution. Kind of feels like, you know, human. Some days I might say four, some might Some days I might say five in this graph, but rarely would I say one, two, or three in this example. And what they were able to show is that they were not able to only reconstruct accurate values for willingness to pay. They were able to capture the distribution. Because one of the important failure modes we haven't talked about is that LLMs, even when they get the persona averages right, they very often lose the details. The variations get muddled together in the middle. This chart at the bottom, basically that horizontal axis is a measure of the entire shape similarity. And one means perfectly identical, and zero means not. And what you can see is the naive way, in the purple or I guess pink uh doesn't do as well as the yellow which is up near the top of the range. So that means it really did a good job not only understanding what the ultimate choice was but how well that choice varied.

分布对齐:确立人类数据的噪声底线

在构建出合成模型后,如何衡量其对齐度是至关重要的课题。需要明确的是,合成角色无法用于提高传统的统计显著性(Statistical Significance)。根据天气预报类比,在同一预报模型上重复运行一千次,并不会让明天的天气预测变得更准确,它只是降低了模型本身的方差,而无法消除基础偏差。因此,增加合成样本并不能凭空制造真实数据。

为了客观评估模型的可靠性,我们需要对比数据分布,并估算人类数据中的噪声底线(Noise Floor:人类在重复测试中由于自身不确定性而产生的天然偏差极限)。研究表明,人类受访者在两周后重新测试时,其自身前后的一致性仅为80%,这意味着模型的准确率上限受限于人类自身的随机性。在无法获取重复测试数据时,我们可以采用分半信度法(Split-Half Method:将真实数据随机拆分为两组以模拟合成与真实对比,从而计算基准噪声的方法)来人工确立噪声底线,从而理性评估AI的对齐表现。

Original English Source

Okay. Let's talk about how to measure alignment from a synthetic persona. One of the things that our traditional market research clients are sometimes surprised by and disappointed is that you cannot use statistical synthetic personas to boost statistical significance. You can take an underrepresented population and get more values out of it but you can't say it's statistically significant. And to understand this it helps to go back to that weather analogy. If I want to know how much it rains today in San Francisco and I used to live here so I know it rains a lot, I'd stick a weather gauge. If I want to know with more certainty, I'd stick a thousand weather gauges and those would increase the accuracy of my estimate. But if I want to know if it's going to rain tomorrow, if I take a forecast and I rerun it a thousand times without changing the input, that doesn't change my certainty of that forecast. It improves my estimate of what the model is telling me but it doesn't make the forecast itself more accurate. And that's what happens when you basically are rerunning a synthetic persona with no changes to input. So the lesson is more synthetic samples aren't actually going to improve your statistical significance for the most part. So what you need to do is you need to do what you do with weather forecast. You'd basically check against what actually happened or in our case what humans actually said. And that's where we're going to basically be measuring distributions of data. Unlike classic e-vals where there's clearly a right and wrong and you can score how many were right and how many wrong, now we need to measure the data as a comparison of distributions. And there are many ways for distributions to get wrong. They could be completely wildly off. They can as we mentioned get the average right but the shape of the distribution wrong. And so you're going to need multiple metrics to capture how well your model is reflecting different personas. Um I recommend using a correlation type metric along with one of these shape type metrics which capture what the underlying shape of the distribution is. The other thing you need to do is estimate the fundamental noise in your ground truth data. That experiment I talked about in the beginning where they got 83% accuracy, the key smart thing they did is they took those humans and they brought them back 2 weeks later and they redid the battery of surveys and personality tests and they found that the humans on average were only 80% consistent to themselves. So that sets a noise floor as how accurate our models could ever get because the humans themselves are fundamentally noisy. And so the 83% is actually normalized against that. If you can do this and bring your humans back, that's great. Very often you can't. So the way you can kind of artificially do this is take your ground truth human data, break it into two chunks, and then pretend one is synthetic and one is human, and then measure the correlation and repeat that hundreds and thousands of times and average it, and that'll set kind of a noise floor that your ground truth data where half of it's synthetic, half of it real, could be the level of accuracy you could hope to get. So hopefully by now you have an appreciation for why I think weather forecasts are the best lens to understand synthetic personas. They are not people, they are forecasts, and we should treat them accordingly. Both systems are bounded, both systems will be improving over time, and they're most trustworthy when they're validated against reality.

人机生态:将调研数据转化为可查询资产

展望未来,合成角色并非要取代人类研究,二者是互补关系。首先,我们正在步入一个人类决策日益被AI代理媒介化的时代,人类已不再是唯一的经济行为体。因此,仅针对人类的调研已不再是绝对的真理,我们需要研究的是由人和AI代理共同构成的人类+代理生态系统(Human-Agent Ecosystem)。

其次,合成角色能让数据变现为可查询的活资产(Living Queryable Asset:通过将静态调研数据转化为可动态交互、模拟的智能体系统,实现二次开发与随时追问的资产形态)。当企业在调研结束数月后产生新的研究假设时,无需重新进行高昂的线下问卷,而是可以直接向这群合成角色发起追问。通过生成式多智能体建模(Generative Agent-Based Modeling:利用多个合成智能体在模拟社会或市场环境中相互作用、交涉的技术),我们可以模拟出复杂的市场博弈与动态演化,开启全新的定量模拟范式。

Original English Source

Um Now synthetic personas are very often cast in the market against human research. And I think that's unfortunate because they're actually complementary to each other. Let me give you two reasons why. One is that we're entering an era where humans are no longer the sole economic actor. Every action your human customer is taking in terms of awareness, consideration, or a purchase decision to buy is being increasingly mediated by AI agents. So, a human-only study is actually not the gold truth. What we really need to understand is what does the human plus agent ecosystem look like? And then finally, the alternative to a synthetic persona is not human research. In most cases, it's no research or it's somebody's opinion. What really happens is you've done a survey of humans and you get a question and if it's in the survey, you can just answer it. That's very simple to do. But what typically happens is it's 2 months later and you're like, we need to answer this question which we didn't ask. Well, then somebody needs to be like, "Uh I think it would be this by extrapolation." Uh expert plus a synthetic persona is going to give you a better result to that. So, what we like to tell customers is synthetic extends your human data to more phases of your development process. It can go more places your existing research can't. One of the most exciting directions is to actually run simulations. We didn't get time for this, but it's called generative agent-based modeling where we can take each of these personas and simulate with the dynamics and how they'll interface interface and interact with each other. And ultimately, what this will let you do is turn your human data into a living queryable asset. If you're interested in doing that with your data, feel free to reach out to us. We help market research and insights teams generate and use synthetic personas in AI. You can find us on the web at insights sciences.ai and my contact information is on the slide. I hope you have a good conference. Thank you. >> [applause] [music]

📌 文中提及的人物和组织

公司/组织: Insight Sciences, Simulmatics

关键字: synthetic-persona market-research large-language-model simulation