Between Two Joels

On this episode of Between Two Joels, Dig Insights AI leaders Joel Armstrong and Joel Anderson dig into one of the most-cited papers in synthetic data research — and ask a simple question: what does 88% accurate actually mean? They unpack how digital twins are built, how accuracy gets measured, and why the headline number looks different once you understand what it's being compared to. Plus the latest AI news and a quick trivia game to close things out.

What is Between Two Joels?

Between Two Joels is your go-to source for making sense of AI. Hosted by Dig Insights AI experts, Joel Anderson our Chief Data Science Officer and Joel Armstrong, VP of AI, each episode explores the latest AI news and deep dives into specific topics, tools, trends, and ideas transforming business and insights.

Joel Armstrong (00:09.678)
Welcome to Between two Joels, I'm Joel Armstrong. Joel's favorite Mario game of all time is Super Mario World, and his second favorite is Super Mario Bros. three, whereas my favorite Mario game of all time is Super Mario Brothers three, and my second favorite is Super Mario World. Very different stuff.

Joel Anderson (00:11.788)
And I'm Joel Anderson.

Emma Sabry (00:24.598)
Let's start off with some news. One of the most interesting stories that we have didn't come from OpenAI, Google, or Meta. It came from the Vatican. Pope Leo is reportedly working on a major AI statement focused on human dignity, which really shows how AI is now colliding with religion, philosophy, and questions about what it means to be human.

Joel Armstrong (00:43.862)
This is the Pope in collaboration with one of the co founders from Anthropic, the collab we've all been waiting for. I mean, this is an interesting exercise. I'm curious to hear what the Pope has to say. I'm not sure what sort of general statements you might make. I do think it's interesting that Anthropic has done a very good job of positioning themselves as sort of the moral compass of the multi hundred billion dollar companies that want you to use AI for everything all the time. I think it's great. Yeah.

Joel Anderson (01:08.886)
Stereometry a lot more than symbol.

Joel Armstrong (01:11.01)
He does I I'm inclined to agree, though I'm not entirely sure why. But yes, they they seem to care a lot about model alignment, prioritize AI safety in a lot of the

Joel Anderson (01:20.168)
How anthropic was started in the first place. Yeah. I mean they kind of split. So Dari Amade I don't know if he was a co founder, just a very early entrant into open AI, but him and some of the other guys that left left on the on the idea that that they can build AI safer and more principled and and whatnot. and and even if you go back to some of the statements that Derry Amade will said while he was working for open AI, he was very vocal about the need for safety and sort of a measured approach without just unleashing everything.

Joel Armstrong (01:49.036)
Yeah. It's interesting too because they are finding that rather than sort of training specific rules in, the the alignment works better if you can instill ethics and moral guidelines into the models because they serve the same purpose in models that they do in humans, which is to provide guidelines to judge behavior against rather than specific sort of, you know, delineated here's what you can do and here's what you can't do.

Emma Sabry (02:11.374)
Okay, so the next news story that we have. For the past two years, the AI competition has mostly been framed as who was the smartest chatbot, but new data suggests that Anthropic has overtaken open AI and business adoption for the first time, largely driven by Cloud Code and Enterprise Developer Tools.

Joel Armstrong (02:31.084)
No, I mean I think Amazing Claude is very good. I have switched to Claude as my as my daily driver. I think I think like a distinction that's emerged that doesn't necessarily get captured in the sort of model horse race is that like model is just one component. And now something that we're seeing become really, really important is what's called the harness, which is the sort of like control system you have and it uses the model, is the engine, and the harness is sort of the car you build with that engine and you attach different tools to it and allows you to organize your file systems and to access different stuff and

To have sort of a central planning agent and all these sorts of things. And the effectiveness of your harness is like a primary determinant of the quality of the results that you get in a way that's not just dependent on the model anymore. And that's I think part of the thing that I've found is like, I don't really care. Like the models are all really, really good. And so jumping back and forth because one gets two percent better rather than having a like, you know, a single workflow that you can get good at and attach tools to and grow with is making a bigger difference over time.

Joel Anderson (03:23.852)
Yeah, exactly. And I think that's why Cloud's winning is because their own researchers are using it for their like all the work that they're doing. You know, they joke internally they're like, no one's writing code, it's all AI written code. And we're just finding better ways to like use the AI written code and using AI as part of our part of their processes when they're when they're building their systems. Totally. I think the reason Cloud is winning now. I mean you you could see the curve for a long time, you know, you know, you see these clouds skyrocketing and open AI still growing but not as fast. And then finally Cloud overtakes.

No big surprise because of the the trends. Tatchi PG kinda had a the huge moment at out the gate. But then since then a lot of people have moved and I think a lot of it is behind the sort of rhetoric from the different companies. You know, people look at open AI and they're sort of seen as the villain in some ways and anthropic, you know, maybe the sort of up and comer. But now that they're in the lead, I think we'll probably continue to see some oscillation.

Joel Armstrong (04:17.517)
So today we're gonna be getting right into it with digital twins. Specifically, we're gonna be talking about how you measure accuracy when it comes to digital twins, because we spend a lot of time thinking, talking about synthetic data. And computing. And computing. And so do all of you, as far as we can tell. I don't know who you are. People are all the time saying, Jules, when are you gonna tell us more about synthetic data? And we always have more to say about synthetic data. So we're gonna get started.

talking about accuracy in digital twins. So Joel, you've been doing a lot of work. We've been talking a lot the last couple of weeks about accuracy, particularly in regards to what's sort of like one of the major papers these days in synthetic data and digital twins. So you wanna tell us a little bit about that paper to get us started?

Joel Anderson (05:01.208)
Sure, yeah. Synthetic data means a lot of things to a lot of people. So we're specifically talking about replacing people with digital twins of themselves. I mean, not that people are doing it to themselves. We're talking about replacing respondents with digital twins of those same respondents. We collect some data about them and then use that data to extrapolate how they would answer on future surveys like quantitative.

Joel Armstrong (05:24.046)
surveys. Yeah. So this synthetic data is individual level, right? You take information about individual people and you simulate those individual responses. So you get similar structured data to how you would if you were running a survey or something like that, where you get rows of data that represent each of these synthetic individuals.

Joel Anderson (05:39.576)
Digital twins as a form of using synthetic data is one of the main ways that people are thinking about as like showing the most promise. So it's the area, you know, when these papers started coming out a couple of years ago or eighteen months ago or so, they were the areas that got us pretty excited about and we started diving into right away and looking at different ways that we could be doing this. People want to think about it as, you know, does it work or not? And it'd be great if the world was black and white, but it is not.

Joel Armstrong (06:02.432)
Yeah. And so that's why we've been doing a lot of work on this and a lot of talking about this is we want to be commenting meaningfully on that nuance, saying not just whether it works or not, but how well it works and under what conditions and what it means for it to work, and to try and help people understand all of those sort of nuances.

Joel Anderson (06:17.388)
Yeah, that's a big one is like what does it mean to work? You know, what does it mean if if digital twins are effective? You know, is there is there utility there? And that's kind of what we'll get into but when we talk about accuracy.

Joel Armstrong (06:28.448)
Absolutely. So the paper we're gonna be focused on today, what's the name of this paper? Twitter.

Joel Anderson (06:32.248)
Win two K five hundred. Yes. And then there's a long thing, colon. Yeah. This is a good paper because they they open sourced and it's by researchers from Columbia University in New York. And it's a great paper because they they open sourced the the data behind it, what their digital twin said, and what the humans said that they that they trained it on, and they did it with the best of intentions. I l I love it. The other major paper we won't be getting into is a paper by researchers at Stanford University.

Joel Armstrong (06:34.924)
Subtitle.

Joel Anderson (06:59.98)
And they built digital twins too. They seeded it with individual level data, but they did theirs with an interview instead of other quantitative research. But similar idea. But those researchers didn't open source their data, so we can't really talk about it.

Joel Armstrong (07:12.43)
So the paper, twin two K five hundred, why don't you tell us quickly what exactly it was at a high level that they did?

Joel Anderson (07:19.276)
The the two K stands for a twin obviously means they did digital twins. Two K means they have a sample of two thousand people. They started with I think twenty five hundred, but they did four waves of surveying. And then the five hundred refers to how many questions that they had the data for. So they they use the information from waves one to three when they create the persona behind these behind these people. And it's not just a general persona like a segment. It's like it's seated with all their individual

level information. So like asking them questions about their demographics, their attitudes, their behaviors, all this different type of stuff that you would get in a survey. They use that information then to ask it follow-up questions to then say, you know, what would you say for this?

Joel Armstrong (07:57.162)
Right. So you build out a profile of the individual with a bunch of information about them, demographics, all those sorts of things. You also build out the profile with a bunch of information about how they've answered real questions in the past based on these true individuals, right? And so you have sort of an understanding of this person that you can hand to the LLM who they are, what they're like, what they believe, what they think about a range of issues, and then you use that sort of seated persona to get more accurate and individual level responses from the digital twins.

Joel Anderson (08:26.008)
Yeah, that's right. And like part of their paper is to figure out, well, what's the best way to encapsulate all that prior information? Waves one to three, right? So do they do they just list all the questions and answers or do they summarize it into a a summary, like use AI to reduce all that information and s and summarize it? Or do they do it in some other formats, like they tested a bunch of those?

Joel Armstrong (08:44.344)
So that's our setup. So okay, so we digital twins. We have a bunch of information about these individuals. They're answering questions that we have, you know, real data from for decades. And so what are sort of the headline findings of the paper?

Joel Anderson (08:54.296)
The main headline finding of the paper was they had eighty eight percent relative accuracy. And obvious obviously when you hear that you go, Wow, it's like amazing.

Joel Armstrong (09:01.1)
Yeah, that's been a big thing, right? Like that people have heard about this paper. This has gained a lot of traction in the field when it comes to like how accurate do we actually think synthetic data can be because of this headline number. And so we're gonna get specifically into an aspect of this, which is what exactly does that 88% mean?

Joel Anderson (09:18.808)
There's many ways that you can calculate accuracy. The two sort of broad ways would be did they say the same thing or not? If you have a seven point scale, let's say it's a some sort of, you know, anchored scale on either end. So you strongly like something and strongly dislike something on the other end. If it's a if it's exact like an exact match, accuracy would be did they get a four on each? Did they get a seven on each? Et cetera. The downside of that is that if you get close, then it's penalized the same as if you got really far away.

Joel Armstrong (09:43.672)
Right. So the base version, this first simplest version would be is it exactly the right answer? You'd have a one if it's accurate and a zero if it's not, and then you'd tally up basically the percentage of what.

Joel Anderson (09:53.58)
Yeah, and if it's a seven point scale, then you'd have one in seven chance of getting that accurate if it's evenly distributed across the prob the prior probabilities across all of the options in the scale.

Joel Armstrong (10:02.636)
Right. But that doesn't account for the fact that answering close to right is better than answering way away from right. So if you have that seven point scale, what's what's next on the sort of list of things you

Joel Anderson (10:11.958)
So the other way of doing it would be to calculate accuracy. The the accuracy measure that they use the formula is one minus the distance between the predicted value and the actual value divided by the range of options in the scale. So if you're on the seven point scale, if you had a one versus a seven, then you're six apart. So you're the f maximum distance. And then you divide by the range of options. So seven minus one options in the scale is six options in the scale. So six over six is one. And then you'd have one minus one. So you'd have zero accuracy in that case.

But if you were a perfect hit, if your answer was seven and the prediction was seven, then you'd have seven minus seven, you know, is zero divided by the range of six is zero. So then you'd have equals a hundred because you w to one minus that value. You'd have a hundred percent accuracy for that that exact.

Joel Armstrong (10:56.834)
Match. Right. So if you if your correct answer was a seven and you got a six, you're off by one, which is one sixth of the possible gap. So you've got like an eighty four percent accuracy. Exactly. And that's basically the scale we're talking about. You get credit for being close and you get penalized more for being far away. Yeah, that's right. Now, one thing that's important to understand when it comes to this 88% headline metric, what does relative mean? And so the way that they did this, right, is that they assumed there was a ceiling on how accurate a result they could get. And they used what you would call test retest reliability.

Which is the fact that when you have someone take a test twice, they won't answer the questions exactly the same. If you got someone to rate a product right now and a week later, they would be roughly 80% similar in the way they would answer your large battery of questions, right? So it's really not reasonable to expect that a set of synthetic individuals would be more than 80% the same.

In their answers because humans aren't even 80% the same to themselves when you get them to answer multiple times. Now, equally interesting is the question of how you set a floor, because the same way that you kind of bring down the maximum reasonable expectation for what can be achieved, you also need some way of knowing what the like minimum bar that you can cross for it to be meaningful. And so we would call that like a naive baseline, right? If you have this one in six die, you know that there's a one in six chance of guessing correctly.

what face is going to be up, right? So it doesn't matter what your predictive method is, it needs to do better than random guessing. And random guessing means you get it right a sixth of the time.

Joel Anderson (12:22.37)
Yeah, that's exactly right. They assumed that there was an even chance for any of those options to be selected for any question. So it's called a uniform distribution. For calculating a random baseline, they used a random uniform distribution. It's kind of the the lowest possible bar that you could have for this. I would argue that it's not really a great random baseline.

Joel Armstrong (12:40.226)
And so what you mean by uniform distribution for this, what they did is they took, say, it was a seven point scale, they randomly selected from one to seven and they assumed any of those numbers were equally likely to have been the answer that was given.

Joel Anderson (12:51.042)
Yeah, that's right. But we just know from like doing many surveys over the years that like people don't pick the ends of the scale as often as they pick the middle of the scale. People sort of generally hedge and people are more likely to say somewhat agree than strongly agree.

Joel Armstrong (13:03.062)
And what this kind of points to, that's pretty interesting and maybe not as intuitive when we were talking about the die example, is that when you're doing predictive models, there's no single correct thing to choose as your naive baseline. You get to think through it and choose what you think would be the most reasonable way of saying we should perform to at least this level, and what you choose will impact the way you interpret your results.

Joel Anderson (13:22.764)
Yeah, exactly. So if you use their n version equal chance of all select all options on the scale, then you the naive baseline, if you do all the all the math and they they did it all, I verified it too. Then you get f fifty nine percent like that's your baseline accuracy, which is kind of an unintuitive number, but it is mathematically correct. But the choice is subjective about what to use for the random baseline. So they calculated fifty nine percent to be random chance. I did the math if you just instead of using a random baseline, if you use the midpoint every time.

Then you actually have a sixty eight percent chance.

Joel Armstrong (13:53.848)
So we know from a lot of historical market research and just surveys in general that people's answers are not randomly distributed across all seven points on a seven point scale or all five points on a five point scale as a naive baseline so that we can make sort of a more informed decision trying to figure out what do these numbers actually mean. So we were looking at a particular chart that shows sort of from zero to a hundred percent how much of an answer can be reasonably accounted for by including all the sort of sources of information that contribute to that.

And on that hundred percent bar, we brought the ceiling down from a hundred percent to eighty-two percent in line with what we were talking about. And we know that the floor was starting out at 50, 59% based on the uniform distribution method, but we didn't think that was very appropriate. So we were trying to figure out where we'd actually kind of squeeze those bars for us to sort of agree with the interpretation. One of the things that we came up with as a simple naive baseline was just the idea of the mode, right? The most common answer that you already got. So if you just apply the mode.

How accurate do you get is the question we were asking. And you ran the simulations on that. And what did we find with that?

Joel Anderson (14:57.548)
Yeah, so when we calculate that out, then it becomes seventy five percent accurate. Which is funny because their digital twins were seventy one point seven percent accurate. Right.

Joel Armstrong (15:06.542)
So this is kind of like a pretty big deal when it comes to interpreting the value of these digital twins.

Joel Anderson (15:11.758)
Yeah, it's pretty stark. We were we were trying to figure out well, where is this proper ceiling? You know, fifty-nine seems a little low. Sixty-eight was what we got if we just used the midpoint. So that was a pretty reasonable candidate. If we're showing that the digital twins only had a sliver of value, it helps you understand sort of the mechanics of the math behind this and what does it mean to calculate accuracy in this way. What we found is that using the mode actually had a higher accuracy at the individual level for predicting everybody. So what that means is most of the variation on on balance.

Because the moat had higher predictive accuracy at the individual level than the digital twins did, that suggests that the balance of the variation that we see in this data set was due to either noise or bias or or whatever. It's just saying that you do digital twins with all this individual information because that individual information helps us extrapolate and really know based on prior answers what they would say in the future. And what this is telling us is that you actually you throw all that away and you get better results.

If you just use the one number that everyone else that everyone says.

Joel Armstrong (16:14.478)
Right. And so the whole point of digital twins is to try and use as detailed of information as you can at the individual level to make sure that what's getting predicted meaningfully differs between individuals, right? That you're capturing individual level information. And what we found is that it is not capturing or representing individual level information, or if it is doing so, it is doing so in a way that does not meaningfully capture useful information. It turns out that if you're asking a question of these digital twins, the better answer to that question

is going to be just what was the most common answer for everybody from your training data for every single person.

Joel Anderson (16:48.258)
Yeah. And then th the obvious follow up to that is that that means that everyone has the same answer. So all subgroup analysis is the same. Right.

Joel Armstrong (16:56.054)
Can't do any subgroups.

Joel Anderson (16:58.56)
Really interpret anything because your results are one hundred percent on the most common thing. Right. The point is you would never actually do this. We're just highlighting that a very naive version of understanding what's the incremental value that this individual level information is providing to you, and we're saying that it's actually a net negative instead of a net positive versus another very naive way of doing this.

Joel Armstrong (17:20.053)
Right. And so this is kind of a big deal because we wanna be clear, we are not anti-synthetic data. We're very pro-synthetic data. We're super interested in it. We think that AI is a fundamentally fascinating mechanism for capturing large scale information, for compressing like just trillions of words of knowledge. Like there's something really important and interesting happening here. And we're trying to figure out what that is.

But this is one of the strongest pieces of evidence we'd seen so far. We kind of reached the conclusion that this will not help us achieve any sort of meaningful synthetic data or digital twin service or offering or anything like that. Like we cannot use this in a meaningful way.

Joel Anderson (17:56.898)
Yeah, this shows us at least for this methodology and this data set, is it not doing it?

Joel Armstrong (18:01.762)
Yeah, and to reiterate your point, kudos to the researchers. They did a good job. They you know, we're not saying that this is bad research, we're saying that definitely it's important to really dig into things because this stuff is complicated and nuanced.

Emma Sabry (18:16.11)
Now we're gonna move to our last little fun section. We're gonna do like kind of a trivia. So I've got a quote and then I've got a list of what the quote could be about. And you have to guess which one it is, okay? The first one is this technology may weaken human intelligence and memory. Is it about calculators, AI, television, telephones, or internet?

Joel Armstrong (18:42.446)
I'm gonna go with television.

Emma Sabry (18:45.812)
One, it's about calculators, this is from the nineteen seventies.

Joel Armstrong (18:49.196)
Yes. I love this subgenre of commentary, like the eighteen nineties like newspaper quotes about like ever no the like no one will talk to each other anymore 'cause we'll all be reading our own papers.

Joel Anderson (18:59.83)
Books come out and they're like, No one's gonna think for themselves 'cause they're all gonna have the same ideas. Yeah.

Emma Sabry (19:04.216)
People will lose the ability to think for themselves. Is this about calculators, AI, television, telephones, or

Joel Armstrong (19:10.914)
I'm doubling down on television.

Joel Anderson (19:13.112)
going with tele no. Internet. I think you're right. AI. AI. Nah. I mean that one did make sense. It's like it's like you know there's a trick question to it, so you don't want to go with the obvious answer.

Emma Sabry (19:28.216)
Children exposed to this too often may struggle socially.

Joel Armstrong (19:31.746)
Television.

Emma Sabry (19:34.658)
Social media is under like internet.

Joel Anderson (19:36.312)
Yeah, I'm going internet.

Joel Armstrong (19:37.558)
We have different strategies. Television. Yeah. Good old rock. Nothing beats rock.

Emma Sabry (19:41.752)
This technology could destroy entire industries and eliminate creative jobs.

Joel Armstrong (19:46.626)
I'll go with AI on this one. I don't know why I said it like that, but I'll go

Emma Sabry (19:50.614)
You guys are right. This invention will ruin meaningful conversation.

Joel Armstrong (19:51.854)
Nice.

Joel Armstrong (19:56.43)
Yeah.

Emma Sabry (19:57.659)
Yeah, that one's from the eighteen eighties. Workers fear they'll become obsolete because of automation.

Joel Armstrong (20:05.272)
Seems like an AI one as well.

Emma Sabry (20:07.818)
This one's a trick question, the industrial revolution.

Joel Armstrong (20:11.01)
That was a trick question. Yeah. I didn't know we were supposed to just be making up answers the whole time.

Joel Anderson (20:12.438)
Back to my social media off screen.

Joel Anderson (20:18.231)
Yeah.

Joel Armstrong (20:22.04)
Alright, so we got a pretty deep into the weeds with a little bit of digital twins today, but this is not the last time we'll be talking about synthetic data. We've got a couple more topics on the list we're excited to be talking about soon. But in the meantime, thank you for joining us and we'll talk to you again soon.

Joel Anderson (20:36.863)
Bye.