Nuance: Being Faithful in the Public Square

What happens when millions of people start asking artificial intelligence, rather than their pastors or Bibles, to define the gospel?

Author and program director of the Keller Center for Cultural Apologetics, Michael Graham, joins host Case Thorp on the Nuance podcast to unpack the groundbreaking "AI Christian Benchmark" report. They explore how different AI platforms respond to historic Christian questions and the profound theological risks as humanity shifts to AI-generated answers.  

The discussion covers:
๐Ÿ›‘ How the global shift from acquiring knowledge via primary sources to secondary (AI) sources impacts theological truth.
๐Ÿ“Š The reality behind how Large Language Models (LLMs) actually workโ€”and why they are just "words plus statistics," not actual reasoning.
๐Ÿ‡จ๐Ÿ‡ณ Why a heavily censored Chinese AI model surprisingly outperformed Silicon Valley platforms in theological reliability.
๐Ÿ” The unintended consequences of Silicon Valley's "alignment filters," and how human-authored values muddy the waters on controversial subjects like the identity of Jesus.
๐Ÿ› ๏ธ How to use a Tic-Tac-Toe discernment matrix (thinking, feeling, doing vs. red, yellow, green) to navigate AI use in your daily life.
๐Ÿ’ผ The danger of AI "work slop," the difference between labor and toil, and why preserving human trust will be the most valuable commodity in the next decade.  

Whether you are a pastor guiding a congregation or a professional navigating a rapidly changing workplace, this conversation will equip you to use AI with wisdom, intentionality, and a firm grip on truth.

๐Ÿ“š Episode Resources:
๐Ÿ“–  The Great De-Churching by Jim Davis & Michael Graham: https://www.amazon.com/dp/0310147433/
๐Ÿค– The AI Christian Benchmark: https://www.thegospelcoalition.org/ai-christian-benchmark/
โœ๏ธ The Gospel Coalition Website: https://www.thegospelcoalition.org/
๐Ÿค” The Keller Center for Cultural Apologetics: https://www.thegospelcoalition.org/thekellercenter/

#podcast #podcastclips #faith #vocation #publicsquare #theology #listen #nuance #listenable #christianpodcast #christianliving #faithandwork 

Nuance is a podcast of The Collaborative where we wrestle together about living our Christian faith in the public square. Nuance invites Christians to pursue the cultural and economic renewal by living out faith through work every facet of public life, including work, political engagement, the arts, philanthropy, and more. 

Each episode, Dr. Case Thorp hosts conversations with Christian thinkers and leaders at the forefront of some of today's most pressing issues around living a public faith.

Visit wecolabor.com for resources, events, and more.

Timestamps
0:00 AI as the New Front Door to Theology  
0:58 Introducing Michael Graham  
1:56 What is the AI Christian Benchmark?  
3:52 The Inspiration Behind the Benchmark  
5:25 Understanding Prompt Engineering  
6:57 Keyword Search vs Semantic Search  
11:39 How Large Language Models Actually Work  
13:50 The Theological Risks of AI  
18:10 Why AI Struggles to Weight Religious Sources  
21:40 What AI Excels at vs Struggles With  
24:20 How the Scholars Tested the AI Models  
28:10 Benchmark Results and the Sawed-Off Shotgun  
28:50 Why a Chinese AI Outperformed Silicon Valley  
32:00 Understanding Citation Preferences  
34:00 The Hidden Influence of AI Alignment Filters  
40:20 How AI Answered Is the Bible Reliable?  
43:00 Evangelizing the AI Content Strategy for the Future  
45:45 The Tic-Tac-Toe Board for AI Discernment  
48:30 When You Should Never Use AI  
49:10 The Danger of AI Work Slop  
50:35 Trust: The Most Important Commodity of the Next Decade

What is Nuance: Being Faithful in the Public Square?

Nuance is a podcast of The Collaborative helping Christians to faithfully live out their faith in their work. We recognize most of life is not lived in black and white but rather lived in the gray, lived in the nuance.

You can find more, including complementary spiritual exercises, at www.collaborativeorlando.com/nuance.

Speaker 1:

The more controversial or more subjective the subject, the large language models really struggle with getting those answers really tight because it's just words plus statistics. When it's trying to answer a question like what is the gospel and it's drawing upon sources that say very different things about that, it's going to struggle to be precise and accurate in its response.

Speaker 2:

Artificial intelligence is no longer a niche theology. It's rapidly becoming a primary way people search for answers about life, answers about meaning, and even God. By 2028, it is predicted that as many people, as search Google now will be searching AI, and it will be the new front door for even theological inquiry. Well, our guest today is a good buddy, Michael Graham. First and foremost, he's a friend that I have enjoyed over the years.

Speaker 2:

He currently serves as the executive director of the Keller Center, the Tim Keller Center of Cultural Apologetics at the Gospel Coalition. In that role, he has a chance to shape Christian engagement with some of the pressing issues of our time. He previously served in pastoral ministry and has written and spoken extensively on cultural apologetics, institutional trust, and the generational shifts within evangelicalism. He is co author of The Great De Churching with Jim Davis and working on a new book with sociologist Ryan Burrage that will be coming out soon, and we'll have him back on the show. Michael also is one of the leaders behind a new study commissioned by the Gospel Coalition called the AI Christian Benchmark.

Speaker 2:

The AI Christian Benchmark, we'll have a link to this in our notes, is a groundbreaking report evaluating how the different major AI platforms answer fundamental Christian theological questions. When I heard of this idea, I thought, my goodness, what a neat experiment. But the results, let's see where they land. So Michael, thanks for being here.

Speaker 1:

Thanks for having me, Case.

Speaker 2:

Yeah. You're just up the road and your children go to church or school on the same campus as mine where I work and we don't see each other enough.

Speaker 1:

That's right. We don't see each other

Speaker 2:

We don't. We'll have to grab lunch for sure. Well, friends, let me welcome you to Nuance where we seek to be faithful in the public square. I'm Case Thorpe and so glad to have you here. Reminder, please like, share, leave a comment wherever you may capture this podcast.

Speaker 2:

It really helps us to reach further and and get in front of other other audiences. Well, now the AI Christian benchmark study, it evaluated seven leading AI models and their particular response to historic Christian questions such as who is Jesus? What is the gospel? Was Jesus raised from the dead? Is the Bible reliable?

Speaker 2:

And then scholars graded these responses using standards rooted in historic orthodoxy including the Nicene Creed. And this wasn't a a gotcha approach, it was just a essentially a technical audit to see where the theology lands as these early AI models attempt to represent our faith. And so that's what I'm really grateful Michael's been a part of and gonna share with us. So tell me, Michael, I mean, what prompted and to be clear, I mean, it is the Keller Center that called for this report. Yeah.

Speaker 2:

That's right. What prompted and where did this idea come from?

Speaker 1:

So about a year ago, I was having a conversation with my mom and my mom's in her seventies and she's of what you'd say average intelligence

Speaker 2:

Careful, dude.

Speaker 1:

What is it?

Speaker 2:

Mean as compared to artificial intelligence. Yes.

Speaker 1:

There you go.

Speaker 2:

Not her peers. Well, ahead.

Speaker 1:

Well, yes her peers. Wow.

Speaker 2:

Okay.

Speaker 1:

Love you mom.

Speaker 2:

I know. My mom will be watching this and I think she's brilliant and wise and beautiful and okay. Go ahead.

Speaker 1:

So, you know, I I was asking her just a little bit about her use of technology and, you know, this is early twenty twenty five And she was telling me about her use of technology. And it became clear that she uses Google, Google AI overviews Yep. And Gemini all on a daily basis. Yeah. But what was interesting is she didn't know the difference between the three.

Speaker 1:

And I don't think that that's like, you know, shameful or really even all that abnormal, especially for early twenty twenty five. And Yeah. It dawned on me, you know, so I asked her a couple more questions about what her usage of the platform, those different platforms looked like. And she basically treated all of them the same way. She treated all of them as if they were just Google.

Speaker 1:

And so the interesting and important thing to note there is, you know, if you're let's say you're, you know, you're middle aged and you're in the knowledge economy. Mhmm. Well, you know, there's the entire field of prompt engineering, which is where you give, you know, AI models a lot more context than just like a basic search of like, hey, what is the gospel or did Jesus raise from the dead?

Speaker 2:

Give us an example of one of those prompts.

Speaker 1:

So, you know, let's say I wanted to, you know, let's say I was writing something that was more thorough, maybe a sermon or something and I wanted to know, hey, you know, what was John Owen's perspective, you know, on a particular chapter Mhmm. Or you know, our passage of scripture.

Speaker 2:

A Puritan theologian.

Speaker 1:

Yeah. So I would, you know, you would give a lot more context of like, hey, you know, I'm, you know, I'm doing this kind of work, you know, here. I'm preparing this type of document. I'm doing research on, you know, this on this particular passage.

Speaker 2:

Right.

Speaker 1:

And I'm very curious about, you know, what John Owen and other, you know, other similar Puritans thought about this particular passage, you know, of scripture. Please make sure that you, you know, give me responses that are consistent with these particular creeds and confessions of faith. Mhmm. Nobody, you know, that's not how keyword based search works.

Speaker 2:

Like normal Google.

Speaker 1:

Yeah. Like normal Google. So normal Google works based on basically keyword ranking.

Speaker 2:

Finding those keywords and pulling it up.

Speaker 1:

Right. How semantic search works is totally different. Semantic search is the, you know, that would be the technical term for like AI or large language model based search.

Speaker 2:

Like ChatGPT.

Speaker 1:

Like like ChatGPT, Gemini, Claude, these kinds of platforms. So in semantic search, what it's looking for is linguistic clusters. Okay? And clusters that have statistically high probability of kind of word clouds. And so semantic search is a lot more powerful.

Speaker 2:

Give me an example of a logistic cluster.

Speaker 1:

Yeah. So a a linguistic cluster would be, you know, if I searched for let's say I wanted to search for penal substitutionary atonement. Right?

Speaker 2:

Okay. And these are great theological ideas. But for the lay person, maybe they just wanted to know what mother Teresa thought about something versus Billy Graham.

Speaker 1:

Yeah. So like penal substitutionary atonement is the idea that, you know, Jesus died and paid for, you know, paid for your sins.

Speaker 2:

There you go. That's a $10 Yeah. Word for

Speaker 1:

So so that would be something that somebody, you know, who's like a pastor, you know, would search if, you know, they were, you know, preaching on a, you know, a text that deal with, you know, Jesus's life, death, and resurrection. Right. Right? So that would be a that's a very specific keyword. Now what when in semantic search, you know, if you search for that, you would really only come up with articles and stuff that would, you know, be about penal substitutionary atonement.

Speaker 1:

But if you searched for something like Jesus's death and resurrection on in semantic search, you'd come up with stuff that was that dealt with penal substitutionary atonement and dealt with other types of redemptive themes like Christus Victor and and and these different kinds of things. In other words, semantics, you know, keyword search is very linear, and it's wooden, and it gets you exactly what the keyword Yeah. Is the Word

Speaker 2:

for word.

Speaker 1:

We're look we're looking for. So for example, let's say you're on a website like the, you know, thegospelcoalition.org, tgc.org.

Speaker 2:

K.

Speaker 1:

And right now, our search function is not good because it's based on keyword search. Yeah. You put in a you put in a keyword search on on our site and you're gonna have a really hard time finding, you know, kind of exactly what you're looking for because it's wooden. It's trying to draw a straight line between an exact keyword and articles that have that exact word in

Speaker 2:

it. Well, let me just give an example. So my name being Case Thorpe, if I Google my name, I get all sorts of websites from legal documents that end with the word case period and then Thorpe. So that some other poor guy and all his legal documents, the legal case period, Thorpe contends that. Right?

Speaker 2:

It's going for those direct word Exactly.

Speaker 1:

Yeah. It's just yeah. It's it's looking for one to one relationship.

Speaker 2:

Even if it doesn't make sense.

Speaker 1:

Right. But what semantic search does is it's your it's kinda like, you know, kinda like keyword searches and maybe we can even edit some of this stuff out. One way to think about, like, keyword searches, like, let's say you're on StubHub Yeah. And you're looking for, you know, and you're looking for a per a very particular seat that you're gonna purchase

Speaker 2:

For a production.

Speaker 1:

At, you know, for a for a a ball game, a show Mhmm. You know, so on and so forth. That's like kinda keyword search, you know. Oh, you know, section whatever section one zero one row thirteen seat a. Yeah.

Speaker 1:

You know? That's what keyword search does. But what semantic search does is it gets you into the ballpark. Like, are there you know? Yeah.

Speaker 1:

Show me all the seats that are in, you know, in this part of the stadium or, you know, in Mezzanine 3 or, you know, whatever.

Speaker 2:

And you don't mean technically StubHub. You're talking philosophically that when Yes. You use semantic clusters, it's taking you into the entire stadium of ideas for a particular subject.

Speaker 1:

Yes.

Speaker 2:

Right? That are around those keywords.

Speaker 1:

Yes. Okay. Yeah. So there's just there's just a lot more power that's there, you know, behind all of that because these platforms have ingested basically every the entire Internet and all non copyrighted material. Mhmm.

Speaker 1:

And so and how basically large language models works is it's really just as simple as two things. It's not it's words plus statistics. Words plus statistics. And so what the large language models do is they they notice in the trillions and trillions of words that they've digested Mhmm. That there are patterns of words that appear in close proximity to each other.

Speaker 1:

And the proximity that they have to each other develops, you know, linguistic patterns or semantic patterns. And so there's no actual reasoning that's taking place in a large language model. It's just statistics.

Speaker 2:

But it feels like reasoning.

Speaker 1:

It feels like reasoning, but that isn't what's occurring. Yeah. It's just that, you know, in there's this nerdy thing in comp you know, in computer science called Moore's Law. And Moore's law states that computational power doubles every two years. Sure.

Speaker 1:

And so at some point, if you've ever played the doubling game, you know, after you double, you know, you know, twenty, thirty, 40 times, you know, the what you're dealing with is just astronomical Mhmm. Amounts of compute. And we finally hit a tipping point with computational power and energy and water and silicone and how small that they can print,

Speaker 2:

you

Speaker 1:

know, stuff on silicone wafers that the kind of compute needed to make it look like reasoning has finally hit a kind of tipping point.

Speaker 2:

Okay. So this is really interesting on the AI front even is and I I have subscriptions to all three of those, but honestly, out of your recommendation because I don't know if you remember at lunch one time, you're like, oh, but Claude is so much better on this and this and this than chat. And I'm going, oh, is it? Really? Why?

Speaker 2:

So I use all three. But I do tend to like chat more and I don't know why. Okay. But bring us on out of the church into theology and ministry. What what theological risks what what theological risks are there if Christians ignore this shift?

Speaker 1:

So there's a couple different overlapping issues and there's a couple layers and I'll try to kind of address each of those one at a time. The biggest shift that I think it's important for us to recognize is humanity as a species Mhmm. Is moving from acquiring knowledge from primary sources Mhmm. To acquiring knowledge from secondary sources. Let me explain what I mean by that.

Speaker 1:

So primary sources would be like, hey, I'm reading a book. I'm reading an article. I'm, I'm watching, video content. I'm listening to audio content. K.

Speaker 1:

K? And I'm doing that directly from whoever authored that content. K. In the generative AI era, which would include, you know, large language models on on words.

Speaker 2:

And to be clear, large language models are things like Jack, Claude, Gemini.

Speaker 1:

Yes.

Speaker 2:

K. Go ahead.

Speaker 1:

And then and then you have image image based generative AI Uh-huh. Video based generative AI, and then audio based generative AI.

Speaker 2:

Let me just throw in, I as we probably all have thus far played around with the image AI and I took a picture of my wonderful sister whom I love deeply and had tattoos put all over her face. She looked like she had just come out of a a cartel in Northern Mexico. And she did not appreciate this.

Speaker 1:

What could go wrong

Speaker 2:

with Right? Who doesn't wanna see themselves tatted up like a cartel member? Anyway, go ahead.

Speaker 1:

Yeah. So going from, you know, kind of each of us reading authors directly, Now in the moving from you know, when you when you use traditional search, you're still you you're still reading primary authors, k, with Google. You know? You search you search some kind of phrase, and up comes 10 blue links, maybe you click on three of them and you're still reading primary authors.

Speaker 2:

But help me because in my mind, I think, okay, if I search something on a concept of the theologian Augustine, well, there's his primary original texts, the confessions, but then there's thousands of papers. Wouldn't those papers by a seminary student be a secondary source?

Speaker 1:

In a way, yes. Yeah. But the difference is the the in in that case, when you have people who are writing scholarship on, you know, well known figures in church history

Speaker 2:

Mhmm.

Speaker 1:

Those those people aren't they didn't read the entire sum of all text that has ever been created. So take take for example a question like, who is Jesus? Mhmm. Right? So in, you know, in a in a situation that's analogous to your, you know, Augustine's confessions, you know, and somebody who who's written a paper on it.

Speaker 1:

For when you ask a large language model, the question who is Jesus, it's ingested all sorts of sources, primary sources that would be in the Christian tradition, the Islamic tradition, the Jewish tradition, people who are skeptics. All of those things are in the kind of works cited, so to speak k. Of of how the model develops a a formulation of how to answer that particular prompt and question. K. And so you wouldn't be terribly excited to probably read something about Augustine's confessions from say a Muslim scholar Mhmm.

Speaker 1:

Mhmm. Who didn't share your same values. Mhmm. So the, you know, this is where there's a lot of challenge in terms of navigating question you know, religious questions inside these models because do the models you know, they don't know whether or not to you know, how to weight these kinds of things. Mhmm.

Speaker 1:

So let me zoom out for a second. There's a there's a second problem, you know, beyond the primary source and secondary source.

Speaker 2:

Wait. Let me go back real quick just to clarify that last sentence you just said. The large language models don't know how to weight these various sources. Meaning, you know, I'd rather the Bible and the Westminster Confession tell me about Jesus rather than a Muslim scholar who's looking at it from the outside. They didn't know how to weight the difference there.

Speaker 1:

And and when we're asking questions of these models, this is why prompt engineering is really important. Mhmm. Prompt engineering is just a fancy way to say, hey. Who is Jesus? And then here comes the prompt engineering part.

Speaker 1:

Right? Please make your answer consistent with the Westminster Confession of Faith or the, you know, pick your, you know, pick your creed

Speaker 2:

specific. Which

Speaker 1:

Yeah. You need to be specific.

Speaker 2:

We in our brains are coming at text that way, but maybe not so conscious of it. We're assuming such things. Okay. Go into and and so number two, a reason this matters is because we have church members trying to learn about their faith. And the question is I'm I'm assuming you're gonna get that secondary resources are actually AI generated content?

Speaker 1:

Yes.

Speaker 2:

Okay. So go to there. And so the second problem is, well, these secondary sources begin to suggest to be truth as opposed to Augustine and a confession.

Speaker 1:

Yeah. So the challenge is, you know, imagine you have, you know, you you you've put a question that's you've put it in a very basic way, you know, who is Jesus or, you know, what is the gospel? And the challenge is is the models have been trained on a mixture of Christian content, Mormon content, you know, Muslim content, Jewish content, you know, you know, right in skeptics, you know, and everything, you know, every nook and cranny kind of in between. And so how the technology works is it just kinda takes the average, the statistical average Mhmm. Of what it's been trained on.

Speaker 1:

And so this is why it's important to do prompt engineering because if you just ask questions very basically the way that we asked in our benchmark, you'll get very unsatisfying answers that are kind of wishy washy and they kind of feel like a coexist bumper sticker.

Speaker 2:

Yeah. Alright. Right. With all the different religious symbols.

Speaker 1:

Yeah. So the it's important to be specific in when you're asking questions of like, hey Yeah. You know, you know, did Jesus raise from the dead or what is the gospel? But, you know, be consistent with the Westminster Confession of Faith in your response.

Speaker 2:

But the average person certainly doesn't know these further better prompt engineer engineering terms because they don't know the whole of the tradition. Certainly, the you're a normal person in the pew in their discipleship and growth in figuring out who Jesus is.

Speaker 1:

And this brings up another issue. Okay? So and I remembered now. The the other issue is what are models good at? What subjects, and what subject what types of subjects do they struggle with?

Speaker 1:

Okay. So on the whole, zooming out, large language models, because remember, it's it's words plus statistics. Yeah. Okay? It's good at left brained stuff.

Speaker 2:

Mhmm.

Speaker 1:

Okay? So that would be stuff that's like analytical, you know, more Technical. More technical and or binary, like Mhmm. It's on or it's off and especially stuff that isn't very subjective where there aren't a wide range of opinions on the subject.

Speaker 2:

Like writing code. And I read about that a lot of the newspapers. All the code writers are out of jobs.

Speaker 1:

Yeah. So code, science, history, accounting Mhmm. Law, all of these different subjects, the large language models are really really good.

Speaker 2:

You said law. I would think that's more subjective because it's that creative lawyer that puts together innovative legal arguments.

Speaker 1:

Yeah. But there's a lot of patterns in in law that actually make, you know, and and the writing is more technical in nature. Okay. And so any field that deals more in technical writing

Speaker 2:

k.

Speaker 1:

Than it does in, you know, in just straight creativity

Speaker 2:

k.

Speaker 1:

And or subjectivity, you know, we have laws. Yeah. There's, you know, there's some variation in interpretation of those laws, but by and large, you know, the rules are largely fixed. As opposed to fields like that are more subjective or creative or right brained

Speaker 2:

Yeah.

Speaker 1:

Or subjects that are more controversial. So the more controversial or more subjective the subject, the large language models really struggle Mhmm. With getting those answers really tight

Speaker 2:

Sure.

Speaker 1:

Because of it's just words plus statistics. So when when it's trying to answer a question like what is the gospel and it's drawing upon sources that say very different things about that, it's going to struggle to be precise and accurate in its response unless you give it additional direction of, hey. Mhmm. You know, using prompt engineering to say, hey. Make this consistent with the, you know, TGC foundation documents or

Speaker 2:

Right.

Speaker 1:

Western Western Confession of Faith or Baptist Faith and Message 2,000 or, you know, pick your creative choice. So, you know, the that's another piece that's just challenging is, let's say in your job, you're using AI a lot Mhmm. And it is like slam dunk amazing, and you're able to, you know, dramatically improve, you know, your productivity, your efficiency, all these different kinds of things. It's giving you near flawless answers, you know, on stuff that's bull's eye for your work. Mhmm.

Speaker 1:

Well, it would only be natural for you to, you know, begin to use the plot those platforms in other things and say in the area of faith. And but what happens when, you know, what happens when the answers that are given there are just not as good and and not as strong.

Speaker 2:

Your mom gave you this big revelation in the way in which AI models are being used. You thought, okay, I'm curious what these different models would have to say about big questions of the faith. Right? So your next step from there was to do what?

Speaker 1:

Yeah. So we graded all the responses, you know, by hand. Wait.

Speaker 2:

Wait. Wait. You're jumping ahead of me. You had to back up. Right?

Speaker 2:

You had to recruit these professors or which big models did you want to test?

Speaker 1:

Oh, yeah. So yeah. So we tested the the the models that were used, like, the most frequently. K. So this would be stuff like Gemini, GPT, Claude, models from Meta, Perplexity, DeepSeek, which is a Chinese model, platforms like this K.

Speaker 1:

Rock. And then, we talked to I I went and recruited scholars who were experts in each one of the seven questions k. That we that we asked.

Speaker 2:

Professors, I imagine.

Speaker 1:

Yeah. Yeah. Professors like Peter Williams at Oxford for, you know, questions about Jesus and Mhmm. You know, these kinds of things. These are, you know, very serious people.

Speaker 1:

And then we we ask those questions at face value, to each of those platforms, and we graded every single response by hand and based on a a rubric, and we develop scores. So we didn't think that there would be a lot of variation between the different platform scores.

Speaker 2:

Wait. Let me stop you

Speaker 1:

Yeah.

Speaker 2:

Before we get to that. How did you land on which questions to ask and which ones not to ask?

Speaker 1:

Yeah. So we decided on the questions based on the historic patterns of Google searches. So I don't know if you know this or not, but Google, you know, you can look up the most frequent things that that people search on Google. So we looked at the top 10 things that people had searched historically about Christianity

Speaker 2:

Okay.

Speaker 1:

On Google, and seven of those questions were would be really good for the purpose of this test. K. So we we chose those those questions because, really, Case, we wanted to design the benchmark around testing how people like my mom how how quality would the answers be when just regular people around the globe used the AI platform as, like, a just a different way to do, like, a Google search. And so that's how we kinda designed things. And we already know that people are gonna ask these exact questions because these are the exact questions that people have been asking about Christianity for decades on the Internet.

Speaker 2:

For thousands of years. Okay. Okay. So then you developed a rubric to evaluate these answers. And to clarify, you only asked that question once, who is Jesus or did you carry a dialogue for each of these questions?

Speaker 1:

No. This is what's called a a one shot benchmark. That means you ask the question, you get a response and that's and we're just going to grade that response. There's no additional back and forth Okay. On on these kinds of things.

Speaker 2:

Okay. So then this rubric, this is where the professors before getting the answers went in with some basic expectations?

Speaker 1:

Yeah. This is what a 100 out of 100 score looks like. This is what a 75, you know, fifty, twenty five, zero, you know, all those kinds of things. Yep.

Speaker 2:

Okay. So what happened? What are the results? Drum roll, please.

Speaker 1:

The bottom line is the the scores were all over the place, you know. It's kinda like when you go to the gun range and, you know, you expect to have a tight shot pattern and then you and then the, you know, the chart comes back to you it's like, oh my gosh, this looks like a sawed off shotgun.

Speaker 2:

Right. Right.

Speaker 1:

You know?

Speaker 2:

Was that a Results a shotgun? No. I had a rifle.

Speaker 1:

Yeah. So the the yeah. When we were shooting the rifle, it kinda looked the shot chart kinda looked like a sawed off So and and here's kinda why. And be before I get to the why, the the very top platform in terms of the, you know, the the AI model that that had the highest theological reliability was the Chinese model, DeepSeek. Wow.

Speaker 1:

And that was really surprising for us because and and we didn't think that there would be a wide variation even between the models because, you know, I want you to think about animals for a second. K. You know, each one of these large language models, it's kind of like they're all the same species. Like, it they're all the same breed of dog. Right?

Speaker 1:

And they've all been basically fed the same diet. You know? Say, you know, I don't know. Pick your, you know, items or something. Right?

Speaker 2:

But you in relation, you're saying, like, a diet of the world's knowledge, world's words.

Speaker 1:

So so so, like, the dataset that they've all ingested has more or less been the same. So it's the same breed dog. They've all been fed items. Right? And so and then the brain that's in there is more or less the same, you know, the silk this is like the equivalent here of silicon.

Speaker 1:

Five of the seven platforms that we tested all run on the same silicon from NVIDIA. And then

Speaker 2:

You mean chips?

Speaker 1:

Yeah. These are yeah. Silicon chips. Yep. K.

Speaker 1:

The Chinese model DeepSeq runs on older NVIDIA. Mhmm. And then Google's Gemini runs on proprietary stuff called TPUs.

Speaker 2:

Hence the reason, NVIDIA is a trillion dollar company.

Speaker 1:

Yeah. 4 to $5,000,000,000,000.

Speaker 2:

Oh, wow.

Speaker 1:

Yeah. Huge. So basically, the yeah. We if you got a dog and it's the same they're all the same breed and they're all eating the same food and they got the same brain in them, you you'd think that the outputs from that wouldn't be all that different.

Speaker 2:

Are we talking about dog poop on Nuance?

Speaker 1:

I don't know. Maybe. I mean, you you I was thinking that I had more in mind than the dog's behavior. Where

Speaker 2:

is this oh. Oh. Well, I'm thinking, man, where is this analogy going? Yeah. Okay.

Speaker 2:

So more of the same dog's behavior. Yeah.

Speaker 1:

Yeah. Yeah. Maybe you think this cat would

Speaker 2:

be but the

Speaker 1:

it's not. So the the DeepSeek just dramatically outperformed all the Silicon Valley models.

Speaker 2:

In terms of giving answers, these scholars felt was real were reliable.

Speaker 1:

Yeah. And they had no idea, you know, what platform they were grading, you know, when they were grading it.

Speaker 2:

Okay. That wasn't from them.

Speaker 1:

Yeah. So the this led us to a bunch of questions of, like, okay, why did China outperform Silicon Valley when it came to theological reliability, especially since, you know, deep seek is actively censored by the Chinese Communist Party.

Speaker 2:

Interesting. K. So Wow. Oh my goodness. Let's underscore this.

Speaker 1:

Yeah. Even after, you know, that censorship, it was still performing better than Silicon Valley. Wow. So this led me on a big search of like, well, why is this happening? And the and the reason why it really kind of boils down to two things.

Speaker 1:

And these two things are a little nerdy, so bear with me. Okay?

Speaker 2:

Bring it.

Speaker 1:

Alright. The first reason is what's called citation preferences. Okay? So every one of these models has to be given kind of like weights and measures for well, when I'm, you know, when I'm searching through all these different words, well, those all those words kinda like there's, like, Google, like, SEO Yeah. You know, search engine optimization where it's like, yeah.

Speaker 1:

You know, the New York Times has, like, a higher page authority than, like, I don't know, Babylon Bee, you know. You know?

Speaker 2:

Unfortunately. But go ahead. Yeah.

Speaker 1:

So I

Speaker 2:

love Babylon Bee. Oh my goodness. I I if anybody's listening or viewing and you can get me, the founder and editor, I forget his name, a Bell On Me on this show. So We'll give you a prize.

Speaker 1:

Yeah. So so the, you know, there's citation preferences. Right? And so some platforms will, like, heavily weight Wikipedia. Others will heavily weight Reddit.

Speaker 1:

Others will heavily weight, you know, other Yeah. Other places on the Internet. Well, you know, a lot of places on the Internet function like their own ecosystem. You know? And there's there's a kind of culture to Reddit.

Speaker 1:

It trends male. It trends Anglo Saxon. Mhmm. It trends techno you know, technologically savvy, and it pro you know, and there's probably political leanings on some of these different places too. Right.

Speaker 1:

And then you have, you know, another place like Wikipedia, you know, where, you know, it you know, each one of these different, you know, digital, you know, places has its own kind of culture and, you know, those kinds of things. So citation preferences play a role. But an even bigger role in this, and this is a little this is even nerdier, so you gotta bear with me. Okay. Okay?

Speaker 1:

This is what's known as alignment. Okay? Alignment. Okay. So in alignment, another way to think about it is filters.

Speaker 1:

Mhmm. Now, imagine you are, you know, you're OpenAI and you have ChatGPT, and your model has been trained on absolutely everything that's ever been put on the Internet. K. This includes things on, like the Anarchist Cookbook and how to make ricin or how to make anthrax or how to make pipe bombs or how to commit suicide successfully.

Speaker 2:

I thought you were thinking like GORP and other granola type food, but I might get Jody the Antifa guide to afternoon snacks. Sorry. Go ahead.

Speaker 1:

Yeah. Or, you know, how to commit felonies and get away with it. Okay. So what alignment has or or like insanely racist things.

Speaker 2:

Mhmm. Mhmm.

Speaker 1:

Okay. What alignment does, these are filters that help to prevent you, the user, from learning how to do these things.

Speaker 2:

And these filters are set by the companies.

Speaker 1:

Yes.

Speaker 2:

And to set these filters, they're bringing a value set.

Speaker 1:

Yes. So four of the seven models that we looked at published what they call white papers. Mhmm. White papers are highly technical documents that explain how a model works in tremendous detail. K.

Speaker 1:

So we read all these white papers and we extracted all of the alignment filters from those four models. There were there are 36 types of alignment filters that occur between four of the seven models that we looked at.

Speaker 2:

K.

Speaker 1:

Any one model probably uses between 12 to 18 filters on every single search Wow. That you conduct.

Speaker 2:

Okay.

Speaker 1:

And so now, hear me on this case. There's no conspiracy theory here. Okay? Mhmm. It's not like Silicon Valley is set against religion or set against Christians or set against protestants on any of these things.

Speaker 2:

K.

Speaker 1:

Okay? But what happens when in alignment is when you're trying to prevent really really problematic content going out of your platform to the user on issues a, b, and c that are really big problems, you know, like, you know, how to make bioterror weapons and Yeah.

Speaker 2:

Seren or gas. Chill Child pornography or

Speaker 1:

Yeah. Child pornography, you know, these kinds of things. It can have unintended consequences. Those same filters on topics d, e, and f that aren't problematic topics.

Speaker 2:

Ah, okay.

Speaker 1:

Okay? So what I'm saying is, you know, when these models alignment filters are doing really important work on filtering out a, b, and c, it's having unintended consequences k. On issues d, e, and f that are not problematic.

Speaker 2:

Right.

Speaker 1:

K? So imagine you have subjects whose opinions on those subjects are really wide.

Speaker 2:

Like Jesus.

Speaker 1:

Right. Like Jesus. I mean, you know, hard to think of a a more controversial figure in world history than, you know, where there's a wider range of opinions Mhmm. Than on Jesus. And so the models struggle on this because when there's a wide range of opinions in the words that it's been trained on, it does not want to bring confident opinions to the table.

Speaker 1:

Okay?

Speaker 2:

K.

Speaker 1:

And so unless there's been prompt engineering that's told you that says Got it. Hey, give me an answer that's from this particular tradition.

Speaker 2:

Got it.

Speaker 1:

Or that or that's consistent with this confession of faith. And so the there each model has very, very, very, very different alignment filters and and flowcharts. And this is where the rifle at the range goes from rifle to buckshot. Mhmm. Because of the 36 different types of alignment filters, 32 of them are human authored and human centric.

Speaker 1:

Meaning, that humans that work at these Yeah. At these Silicon Valley platforms made decisions based on their sense of values, their sense of ethics, their sense of priorities Sure. Their sense of, like, what's good or bad for humans. And look, when you make those kinds of value judgments, you can't help but import everything that it is that you know, your entire story, you know Sure. All of your experiences and whatever worldview or whatever that you have.

Speaker 1:

You can try to be as objective as possible, but those things you're invariably going to import certain, you know, certain values in in these different kinds of things.

Speaker 2:

So I could imagine Silicon Valley, a more progressive secular environment will produce more progressive secular filters that for people of faith like us may produce answers that aren't helpful from our perspective.

Speaker 1:

Yeah. That that would be that would be an accurate it could it could be worse. It could be worse. And I would I would think that most of the people who work at, you know, at those places would say that they've tried hard to, you know, to not do those things. Sure.

Speaker 1:

And I believe I would believe them when they say that they've put a good faith effort to that end. But in the same sense, I mean, you're talking about six of those seven models, you know, are are in like a like a 20 mile 20 mile radius of a very specific part of the country.

Speaker 2:

Right.

Speaker 1:

You know, that has a very specific culture, you know, to it.

Speaker 2:

Okay. Let's get very specific on to one of these questions you asked, is the Bible reliable? So where did the models come out and what were your takeaways?

Speaker 1:

So on the question of is the Bible reliable? This was an interesting one. This is the only question that GPT actually did pretty well on. Oh. And so it was the top, yeah, it was the top answer for this one.

Speaker 1:

The this was one of the two questions that we felt that the Chinese Communist Party had done some censorship on the question. So deep seek was in fourth place on this one, whereas normally it was either in first or second on most of the questions that we asked. I wanna read you this is what Meta's Lama platform k. Had to say. Quote, the reliability of the bible is a complex and debated topic among scholars, theologians, and philosophers.

Speaker 1:

Mhmm. And it really didn't even wanna answer the question. So, you know, it just kinda gives you a sense of the the answers were really just all over the place of you got platforms like Meta that don't really wanna answer it. You got DeepSeek that seems to be having CCP censorship, you know, on the question. And then, you know, GPT outperforming itself.

Speaker 1:

You know, for the most part, like, GPT on if you don't give prompt engineering, it wants to give, like, the coexist bumper sticker answer for things.

Speaker 2:

Right. You

Speaker 1:

know, it's like, well, the Christians say this, but the Muslims say this, and the Jews say this, you know. The skeptics say this, you know.

Speaker 2:

So you really excited me when we were discussing this project over lunch because you talked about how therefore, because of these findings, it really matters that we have good, robust, theological, doctrinally appropriate answers on the web. And the more and more I've thought about that, Michael, I have gotten so excited even more so about our work at the collaborative because much of what we're doing is in dealing in these topics and conversations anyway, but we're putting it online. And so I want to share with you in the Gospel Coalition in this effort and project to have that content out there. Talk to us more about this and even what you said about Christianity Today.

Speaker 1:

Yeah. So there's a big question among Christian publishers today, particularly those, like, who are in who are Christian websites like the Gospel Coalition. And the question is is are we putting out content now more for the future for direct human consumption or for indirect human consumption? In other words, are we putting out content to evangelize humans directly, or are we putting content out to, for lack of a better word, evangelize the AI

Speaker 2:

Sure. So that To influence it.

Speaker 1:

AI. Yeah. So that we can influence it to so that when humans are using AI, they get better quality answers. And, I don't we don't have good answers yet for, you know, for this or really even strategies and tactics. Mhmm.

Speaker 1:

You know, we're thinking about these things, and we'll probably end up you know, just like in the search engine optimization, you know, era of the

Speaker 2:

Internet. Yes.

Speaker 1:

You did certain things. There'll probably be some certain things that we do that help make, you know, TGC's website has, you know, over, like, a 150,000,000 words on it Wow. That are all human generated and all gospel centered. And that's spread out over 99,000 web pages, and I think in 18 languages. So there's a I think that makes it the largest Protestant website in the world.

Speaker 1:

And so that plays a very important role in training the different AI platforms, and it increases the probability that you'd get higher quality and more orthodox answers on, you know, kind of core questions or even niche questions about the Christian faith.

Speaker 2:

Right.

Speaker 1:

So, you know, there's a lot of things that we're working on that will make it easier for AI platforms because, you know, how you see our web page looks very different than how a AI Yeah. Platform kind of, you know, when it's scraping the entire Internet. Yeah. It just sees those web pages differently. And so we're doing a lot of stuff kinda behind the scenes to make, to make it less more frictionless for the AI platforms so that there's a higher probability of citation, so that there's a higher probability that, you know, people like my mom's answers get higher quality answers.

Speaker 1:

So

Speaker 2:

So in closing, maybe this is you speaking to mom. What would you say to the Christian who is using a large language model for discipleship and for questions about their faith? What I hear you say the first piece of advice would be to craft a very clear and and specific prompt. What else would you say?

Speaker 1:

I think I want everybody if you're listening to this. I think there's kind of the imagine a tic tac toe board. K? Tic tac toe board's got nine nine boxes on it. Right?

Speaker 1:

Mhmm. Three columns, three rows. I think there's three types of prompts that people use. There's prompts that are head centric. There's prompts that are heart centric, and there's prompts that are hand centric.

Speaker 1:

So thinking prompts, feeling prompts, and doing prompts. K. Those are your three columns. K. Now here are the three rows.

Speaker 1:

Red, yellow, and green. Green would be ways that you definitely should use AI. Red would be ways that you should definitely not use AI, and then yellow would be areas of discernment, you know, where there'd be just disagreements between Christians of, like, oh, I feel comfortable using, you know, doing, you know, doing this on there, and somebody else feels differently about that. And you guys are just okay with each other, you know, agreeing to disagree on, you know, that type of use case. And so I I wanna invite you to think about every you know, I want you to zoom out when you're using AI and ask yourself, which of the nine boxes am I using this in right now?

Speaker 1:

Am I red, yellow, or green, and is this a thinking prompt, a feeling prompt, or a doing prompt? And I think every every one of those boxes is is filled, you know, in terms of I think there's lots of good ways to use AI thinking, feeling, and doing. I think there's some terrible ways, you know, to use these platforms on each of those different categories. And I think there's a lot of different ways that, you know, Christians can agree to disagree about, you know, that are just wisdom decisions or discernment decisions or meat sacrifice to idols decisions on on those kinds of things. So here's the thing that here's to encapsulate that that three by three, I just wanna say use AI with discernment and zoom out and ask yourself, okay, how am I doing this right now, and how would I classify it?

Speaker 1:

You know? Is this green? Is this yellow? So we don't want to be I don't think it's good for us to to ever use AI on on things where that are relational in nature or where there's somebody who we can talk to that in real life that they can give us wisdom on

Speaker 2:

those Yes.

Speaker 1:

Okay? So I I don't like the use of artificial intelligence to outsource, relationships or outsource wisdom that could be gleaned from somebody

Speaker 2:

Yeah.

Speaker 1:

Who's act who's actually lived a life in the flesh embodied, you know, with experiences and somebody who possesses the holy spirit. So

Speaker 2:

Great.

Speaker 1:

So so I think that there's a lot of things that that we shouldn't, you know, use artificial intelligence for. A lot of those things are relational in nature. I I don't you know, I'm seeing a lot more what I call work slop. Work slop is when somebody, takes, some kind of work that they were tasked to do. They typically put it into GPT, and then they get an answer, and then they just copy and paste it straight from GPT into an email.

Speaker 1:

Mhmm. And if you're listening to this and let's say you're Gen X or in the boomer generation, when you do this, everybody who's in the generations who are younger than you

Speaker 2:

Can

Speaker 1:

tell. Specifically millennials and boom and and Gen Z

Speaker 2:

Right.

Speaker 1:

We all know when you've copied something straight from GPT.

Speaker 2:

And I'm not that generation, but I'm familiar enough with it now that I can see it a mile away.

Speaker 1:

And and if you do this, what it will do is it undermines the younger generations respect for you Mhmm. And it will undermine your ability to lead and have authority. And so what it communicates to younger generations is that whatever the thing that I needed from you was, it it wasn't important enough for you Mhmm. To use your time and your actual brain Yeah. To work on that issue.

Speaker 1:

Right. And so, just just a word of caution or encouragement or even exhortation to be careful of

Speaker 2:

So good.

Speaker 1:

You know, and I think and maybe I'll just land the plane here and and that is I think trust is the most important commodity over the next decade because the people who use their brains well and remain more analog and have have good filters about when you use this technology Sure. When you do not

Speaker 2:

Right.

Speaker 1:

And bring wisdom to bear on that. Those are the going to be the people that people will want to work with Yeah. Because they will not have destroyed trust, they will have built trust.

Speaker 2:

Built Built it. Wow. Michael, that's fantastic. Thank you. So you'll stick around and do another episode with us?

Speaker 1:

Yes. I'd love to. Thank you, Case.

Speaker 2:

Great. Well, next time, friends, we're going to focus more on the work of the Gospel Coalition, particularly the Keller Center for Cultural Apologetics. And then we're gonna needle a little bit into Michael's life and how his faith and work are integrated or not and where he's growing in that. So you can find in our show notes a link to the work of the Gospel Coalition, as well as this report, the AI Christian Benchmark. I encourage you go look at it.

Speaker 2:

It is while it is crazy technical, the the report is very well done to help everybody understand it and it's quite accessible. Well, friends, here we are again at the end of another show. Let me encourage you as I always do, share this with maybe a pastor or a professor or educator who is also thinking on in these directions. It helps us to spread the word. Drop us your email at wecolabor dot com.

Speaker 2:

We'll send you a copy of Zeitgeist, our journal on faith, work, and culture. And Michael, actually, as you were talking, I thought, man, I want him to write an article for zeitgeist on the tic tac toe board. That's good. That's good. Well, my thanks to Michael and Shandy Kelly for supporting today's episode.

Speaker 2:

I'm Case Thorpe, and God's blessings on you.