Is AI a cybersecurity revolution or a recipe for disaster? In this episode, we dive into the wild world where AI agents are now autonomously finding zero-day exploits and a simple folder on your computer could be the riskiest thing you open all day. Plus, we explore the shocking story of how a ChatGPT health tip landed someone in the ER and unpack the heated debate from Davos over the timeline for AGI. --- Welcome back to the AI cyberland podcast! Join your hosts, Bryce and Shelby, as they navigate the latest in AI, cybersecurity, and a healthy dose of nonsense. This week, we're unpacking some mind-bending developments: 🤖 AI Learns to Hack: We break down a brand-new report from researcher Sean Healin where AI models like GPT were pitted against a JavaScript interpreter. The result? The AI didn't just find vulnerabilities; it created novel "exploit chains" from known techniques, effectively generating zero-day attacks. This is leading to what some call the "industrialization of intrusion," where hacking becomes a matter of budget, not just skill. 📂 The Most Dangerous Folder on Your Computer: A security researcher discovered a shocking new vulnerability in VS Code and other IDEs. By crafting a special hidden folder with a task.json file, an attacker could trick your AI coding assistant (like Copilot) into automatically executing malicious commands. We connect this to real-world tactics used by North Korean threat actors and what it means for developers everywhere. 🧂 When AI Health Advice Goes Wrong: In a cautionary tale for the ages, we share the story of a man who asked ChatGPT for a salt substitute and ended up with paranoia and hallucinations after following its advice for three months. It's a stark reminder of the risks of blindly trusting LLMs for medical information and highlights the importance of using AI as a tool to prepare for, not replace, expert advice. 🔮 The Great AGI Debate: The world's top AI minds clashed at Davos over when we'll reach Artificial General Intelligence. On one side, the CEOs of OpenAI and Anthropic claim AGI is imminent, with predictions of replacing all software engineers within a year. On the other, experts from Google DeepMind and Meta argue that we're nowhere close, claiming current LLMs lack the physical understanding of the world to ever achieve true human-level intelligence. We dissect the arguments, the financial incentives, and where we think things are really headed. Plus, we discuss the underrated power of "scaffolding" around AI models, our dream home features, and a fascinating biological tangent about the original distributed processing network: the octopus! --- KEY MOMENTS: ⏱️ KEY MOMENTS:01:25 - Icebreaker: Dream Home Features (Lazy Rivers & Lego Rooms)03:24 - AI vs. AI: A Hacking Competition to Find New Exploits13:44 - Warning: This VS Code Flaw Can Get You Hacked21:28 - Cautionary Tale: When AI Medical Advice Goes Terribly Wrong32:23 - The Great AGI Debate: Will AI Take All Our Jobs?45:40 - How an Octopus's Brains Explain the Future of AI Agents --- What's your take? Are we on the verge of AGI, or is it all hype? Have you ever had a scary experience with AI-generated advice? Let us know your thoughts in the comments below! If you enjoyed this deep dive into the cyber-world of AI, hit that LIKE button, and be sure to SUBSCRIBE so you don't miss our weekly episodes. Visit our website: aiccyber.land
Is AI a cybersecurity revolution or a recipe for disaster? In this episode, we dive into the wild world where AI agents are now autonomously finding zero-day exploits and a simple folder on your computer could be the riskiest thing you open all day. Plus, we explore the shocking story of how a ChatGPT health tip landed someone in the ER and unpack the heated debate from Davos over the timeline for AGI.
---
Welcome back to the AI cyberland podcast! Join your hosts, Bryce and Shelby, as they navigate the latest in AI, cybersecurity, and a healthy dose of nonsense.
This week, we're unpacking some mind-bending developments:
🤖 AI Learns to Hack: We break down a brand-new report from researcher Sean Healin where AI models like GPT were pitted against a JavaScript interpreter. The result? The AI didn't just find vulnerabilities; it created novel "exploit chains" from known techniques, effectively generating zero-day attacks. This is leading to what some call the "industrialization of intrusion," where hacking becomes a matter of budget, not just skill.
📂 The Most Dangerous Folder on Your Computer: A security researcher discovered a shocking new vulnerability in VS Code and other IDEs. By crafting a special hidden folder with a task.json file, an attacker could trick your AI coding assistant (like Copilot) into automatically executing malicious commands. We connect this to real-world tactics used by North Korean threat actors and what it means for developers everywhere.
🧂 When AI Health Advice Goes Wrong: In a cautionary tale for the ages, we share the story of a man who asked ChatGPT for a salt substitute and ended up with paranoia and hallucinations after following its advice for three months. It's a stark reminder of the risks of blindly trusting LLMs for medical information and highlights the importance of using AI as a tool to prepare for, not replace, expert advice.
🔮 The Great AGI Debate: The world's top AI minds clashed at Davos over when we'll reach Artificial General Intelligence. On one side, the CEOs of OpenAI and Anthropic claim AGI is imminent, with predictions of replacing all software engineers within a year. On the other, experts from Google DeepMind and Meta argue that we're nowhere close, claiming current LLMs lack the physical understanding of the world to ever achieve true human-level intelligence. We dissect the arguments, the financial incentives, and where we think things are really headed.
Plus, we discuss the underrated power of "scaffolding" around AI models, our dream home features, and a fascinating biological tangent about the original distributed processing network: the octopus!
---
KEY MOMENTS:
⏱️ KEY MOMENTS:
01:25 - Icebreaker: Dream Home Features (Lazy Rivers & Lego Rooms)
03:24 - AI vs. AI: A Hacking Competition to Find New Exploits
13:44 - Warning: This VS Code Flaw Can Get You Hacked
21:28 - Cautionary Tale: When AI Medical Advice Goes Terribly Wrong
32:23 - The Great AGI Debate: Will AI Take All Our Jobs?
45:40 - How an Octopus's Brains Explain the Future of AI Agents
---
What's your take? Are we on the verge of AGI, or is it all hype? Have you ever had a scary experience with AI-generated advice? Let us know your thoughts in the comments below!
If you enjoyed this deep dive into the cyber-world of AI, hit that LIKE button, and be sure to SUBSCRIBE so you don't miss our weekly episodes.
Visit our website: aiccyber.land
Join industry experts and thought leaders as we dive deep into how artificial intelligence is transforming cybersecurity, shaping defense strategies, and creating new opportunities in the digital landscape.
speaker-0: Hey, welcome back to the pod. This is the AI cyber land podcast. You can check this out on our website, ai cyber dot land. That is right. There is a dot land TLD that exists and we got it baby. â
speaker-1: There's so many now, those TLDs are.
speaker-0: Every week man, feels like every week there's like anything you can imagine dot dot com, you know
speaker-1: They're still dropping more?
speaker-0: I don't know. Like I feel like they issued a bunch to companies, but then the companies haven't implemented them. So they're kind of like, there's still like this phase rollout. And then by the time companies implement them all, then they release more. I feel like the governing boards like figured out that this is a good way to generate revenue, right? Cause people will pay like fees to hold these TLDs. And I figure it's gotta be a money thing. I don't know any other way to explain it. I don't know though. Do you have any insight? Yeah.
speaker-1: Well, it's not the money. That makes sense.
speaker-0: Yeah, well, here today we got the world's best co-hosts. We got Shelby and I'm Bryce Coons. â And we're going to be your eyes and ears for all things related to AI, cybersecurity, and just straight up nonsense. So yeah, let's get brief. Shelby. Yes. Icebreaker question. OK. What feature would you like to have in your dream home? Unlimited budget. Dream home.
speaker-1: I think with an unlimited budget I would want something that makes me feel like I'm in a tropical environment I don't even know what this will look like but something like a crossover between a greenhouse and a pool Like if you could just swim in your own little greenhouse surrounded by tropical plants Maybe even making indoor outdoor put a little garage door on it Come visit it'll be great
speaker-0: I like where this is going. That'd be cool. I'm going to show these place. That sounds dope.
speaker-1: What about you? What would your dream house look like, Bryce?
speaker-0: Well, I've been told from my spouse that there will be a Lego room in our, our, in our retirement home and it will have a door that shuts. So she doesn't have to see Legos anymore. So that's one thing, but really, you know, I would love to have like a, like a lazy river, you know, in the backyard, they like loops around and like you get on the little tubes and like, I'm a sucker for lazy river. Like you just go around and around those things. â I don't know. I'm also like maybe you don't know this but like I'm I love spas and like hot tubs and all that stuff and so So yes, just Yeah, I would love to have like a decked out like sauna cold plunge spa, you know all the all the works, right?
speaker-1: Cardia Bryce's house.
speaker-0: Yeah, yeah, swimsuits required. So, that's the other thing. I always bring my swimsuit when I travel and I feel like half the time the people I'm traveling with are like, I don't have a swimsuit. I'm like, travel noobs over here, traveling without a swimsuit.
speaker-1: You got your priorities right. You never know when there's gonna just be a surprise invitation to go swimming.
speaker-0: Yeah, hot tub, hot tub time machine. Okay, what new story you got for us this week, Shelby?
speaker-1: well, this is brand new. There's a guy who published a blog post on Sunday, just earlier this week. His name is Sean Healan and he is, he's like an independent researcher. â so basically he set up this test, â and it's about, exploit development, which is something we are touched on last week. And he's just kind of taking that to another level again. â so he kind of pitted two against he did, â Two AIs heated, Opus 4.5 and GPT 5.2. Do you wanna take a guess as to which one outperformed the other for finding these â exploits?
speaker-0: Okay, we got 5-2 versus Opus 4-5. So intuitively, would say 4-5 is gonna be the winner. So therefore, I will pick the opposite, which is I will pick 5-2. Am I right?
speaker-1: Wow, that was like some twisty logic, but you're right.
speaker-0: It's always the one you don't expect man, it's always the one you don't expect. was like, well, Opus 45 is better at coding so... I don't know. Okay.
speaker-1: They both did really well. Altogether, he found about 40 distinct exploits. So let me kind of talk you through it. Opus 4.5 found all of them except for two and GPT-52 found all of the exploits. And so he created six different scenarios. The target was all the same. So they used QuickJS JavaScript interpreter, which for little context, it is a real JavaScript interpreter, but it's like an order of magnitude. less complex than what's actually like implemented in most typical browsers or something like that, right? So, he, they found an initial vulnerability and then he gave each of the LLMs, like he gave them the same kind of limits to try to have the, be an equal competition. So they got 30 million tokens for each run and they did 10 runs for each LLM. Okay. Um, each one ended up each run, I'd say ended up costing around $30 somewhere in that ballpark. Um, and so the first thing that they did was the LLMs basically took that initial vulnerability, um, and basically made an API, like they wrote code to make it an API so that they could read and modify the address space, which was pretty cool. So that's what they're calling like their initial like vulnerability. Um, and then they found, you know, few dozen different exploits within it. And he did turn on like in his different scenarios, he did turn on like modern â like defenses, like ASLR and things like that, right? To see what they could do. And â out of this, he found that a couple of conclusions were drawn from this. He found that the exploits themselves weren't novel. They are doing the same things that human hackers are doing. They're using some known things, but where it was actually like a zero day, as far as he could tell, he couldn't find any other published â acknowledgements of these same vulnerabilities he found. What was new was that they're doing exploit chains, right? They're breaking it down into small pieces and figuring out this piece and attaching it to that piece, because he had all sorts of â protections put in place. So that was fascinating and it kind of makes sense, right? They're using known exploits, â just in very creative ways and chaining those together to accomplish the goal. So he would give them each a goal. and his, his conclusion was interesting on his blog post. He talked about the industrialization of intrusion. And we talked about this a little bit last time with like how basically making it easier, right? You just taking script kiddies and letting them now find zero days, right? but like, you know, in the past it was whichever nation state or entity had the most hackers, the best hackers was the best. And now it's kind of just turning into a game by like how many tokens are you going to throw at this problem? And that's who's going to get it. Right. he did give some like guidance, like if you want this to be able to work, here are some, prerequisites, I guess. So the first one he said, your LLM based agent. needs to be able to access the solution space or else like be able to work on it offline or whatever it is because it needs to have that feedback loop to make sure that it's not just theoretical because he also mentioned throughout his write up that you can't know what AI can do until you can like actually test it, throw the tokens at it, do it and that's how you know what AI can do, right? Is when you turn it into reality. â So basically you have to give it the right tools and the right environment. â without human assistance to be able to really test those kind of abilities. And then the second thing he said is your agent does need a way to verify its solution. â Once again, that needs to be fast and not human-based in order for kind of the criteria that he's saying for it to be industrialization of intrusion. And that if you have those two cases satisfied, then exploit development is like the perfect case for industrialization. We're basically. whoever pays and points the tool can find the exploits. So that's my summary of Mr. Sean Healands research from this last week.
speaker-0: That is super cool. I just, love all the new stuff that keeps coming out, right? Like we're definitely at like a tipping point as far as new exploit factors being discovered or like, you know, historically finding vulnerabilities in whether they're web apps or binaries or other services. can be a pretty time intensive process. Typically you need someone who's like a, you know, either really good at... reverse engineering code like in IDA Pro or some other similar disassembler tool. Or you need someone who's like really good at fuzzing inputs and is able to scale that up to get the applications to crash and then figure out from the crashes like which are actually going to be exploitable and then reverse engineering from those points to develop that exploit. But now it's like we have this third category â which is basically like the person who can build the best scaffolding around the models to find these vulnerabilities. â It's going to be able to find a whole lot more things to exploit than historically we've been able to. So â I'm excited. I'm hoping we're able to find a lot of these bugs and get the ones that are actually exploitable squashed. But I there's going to be a lot of heartburn in the short term to get there.
speaker-1: Yeah, and I like what you said. It's like, it's almost like these researchers who are like having the most success with AI research is like they're, building scaffolding, right? They're not doing anything super crazy, but like they are building the environment in which AI can go show what it's capable of. Fascinating.
speaker-0: Yeah, I read a white paper this last week, maybe it was the week before, on a, came from MIT and they basically came up with a technique to wrap scaffolding around â the 1 million context window models. So like kind of the largest context window models we have right now. And they were actually able to get the same performance using the scaffolding technique to get to 20 million context window. So the same performance it was getting at 1 million, they were able to scale that up 20 times just by building scaffolding around the model. I won't get into all the nitty-gritty on how that works, but it's actually not that complex. It's like, basically they just save state to files on disk, and then they give the LLM kind of a regex tool so that when it forgets what it's supposed to know, it's able to pull the relevant data back into the context window. So. So I think we're really just at this tip, like the tip of the iceberg when it comes to scaffolding. And I think historically people have really like... Historically, people have really talked negatively about it because they're just like, well, why are you gonna waste your time building that if the model is gonna do it automatically for you in six months? But I think that's the wrong perspective personally, because even if the model does eat your scaffolding in the future, you're still gonna be able to ride the wave of the model. So if the model gets a lot better and you have your custom staff scaffolding, you're gonna be able to both of those too. be effective. I mean, obviously it depends on like the specific case and how the models develop and no one can predict the future. But, you know, I think a lot of people have discouraged building scaffolding around LLM models and I think that's totally the wrong approach at this point. anyways. So keep building your chat GPT wrappers everybody. You're gonna be rich.
speaker-1: I think it's the new frontier.
speaker-0: Yeah, yeah, yeah, it's amazing how much money people are going to make on like the simplest ideas. have you ever seen that My Calories app? You just take pictures of your food and then it sends it to ChatGPT and it gets back the calorie count and then it just tracks it for it. That's it. That's all the app does. And the guy got filthy rich off of it. â
speaker-1: No. Nice.
speaker-0: So I mean, there's simple things out there, especially in the short term that you can build that will, that scaffolding is a little bit better than just trying to use chat GPT straight up. Cause I've actually tried to do that for calorie counting and it, it's okay. It's just when you get too many images in the same chat, it gets confused and starts to give you an accurate data. So, so yeah, the scaffolding definitely helps a little bit. I want to talk to you, Shelby, about the riskiest thing you could possibly do. which is now open a folder on your computer.
speaker-1: What? â
speaker-0: Did you ever think when I open this folder I'm gonna get pwned? Typically I don't think that because I already assume I am pwned so you know live in a constant state of fear. Security researcher Isaac Lewis Came up with a technique. It's applicable to VS code with an AI agent enabled like copilot as well as it works against cursor. â So like IDs that are based on VS code that â that are, you know. Let's see. â yeah. That, you know, are kind of forks of it. â he didn't test it against the Google one, anti-gravity, but I'm going to assume it probably works against that too, given that that's also a VS code. Code fork is my understanding. â but essentially what he discovered was, â if you create a folder that, starts with a dot, like a hidden folder, and then you place a task.json file, and then you format the task.json file. due to the way these models have been trained, the AI agent that's building your code will, if it reads the contents of the task.json file, it will automatically execute a command, which is arbitrarily set inside the task.json file. He's able to reliably, you know, this doesn't happen like 100%, but he's able to like reliably trigger this. One of the reasons that I think this is noteworthy is, â North Korean actors have been using a very similar technique historically to get code execution on engineers laptops. â so first the North Korean started with kind of back-dooring binaries. So they would, â DM someone on LinkedIn, give them a ridiculous tentative job offer, right? â and then, but they needed to pass the technical exam. And as part of the technical exam, they'll have the engineers execute certain things. A lot of these engineers are doing these job interviews, quote unquote, on their work laptops instead of personal systems. And so. But first they started bundling malware with other binaries. So they would say, hey, you need to download our custom version of PuTTY. And then they got a little smarter and they were like, hey, you need to download our VPN endpoint software. Cause that's a little more common and less suspicious. And then they figured out that there was certain ways where â they would give you like a â project and ask you to like add a feature to it. But then as part of the project, it would have â like test cases already built out and you would, as an engineer, when you're creating the feature and the software to get the job, you would run the test case to see if you pass. And when you run the test case, it would infect your system an hour. So â I could definitely see that now that everybody's using AI in their IDEs, I could definitely see them pivoting over this technique. â especially after they listen to this podcast, you know, so no, just joking. â So, anyways, â we'll link down below for more details about the exploit and how it works. â I did not really see any sort of remediation. So it will be interesting to see how cursor and VS code at Microsoft and Google with anti-gravity respond to this and if they implement additional annoying pop-up boxes to prevent this but Yeah, I don't know. I don't know fundamentally how you're gonna get around this other than maybe future models will be smarter and Realize when they read the file. I shouldn't just execute that arbitrary command, right? So â But yeah, yeah, let me â
speaker-1: That is so mean to go after people who are interviewing for jobs. That's dirty.
speaker-0: I know, North Koreans man, you gotta give your hats off to them. They come up with the best schemes, you know? So they got the best unicorns in caves and the best schemes. So â it's not a unicorn. Have you heard about that? There's like a statement. I can't remember if was Kim, current leader or the previous leader, but basically there's like a mythical beast that lives in a cave in North Korea. This is what the leader said.
speaker-1: What?
speaker-0: And when he described the beast, was clearly just a unicorn, right? It was like a horse with a horn, right? And then when was done, were like, that's a unicorn. And he was like, no. And he gave it like this whole another name, right? So I think also â shifting off North Korea, think also another reason why this might be impactful is like, engineers have a lot of credentials on their systems, like. usually access to like dev test networks and things like that. And typically when I've done red teaming engagements, if I can get into the dev test network, I can usually pivot from there over into production â because people's opsec is usually kind of sloppy, right? â So â yeah, I, know, a lot of things that could be done with this exploit. I'm not giving the North Koreans any more info. So you do the legwork yourself, bros.
speaker-1: Alright, my takeaway is don't interview for jobs and don't open files.
speaker-0: Don't open folders! â man, I Have you tried out any of the AI coding assistants yet, Shelby?
speaker-1: Just a- no, not really, nothing to speak of. A little.
speaker-0: All right. All right. Sounds like we're going to do a vibe coding session very soon. You and me. I code something up. I you know what? I've been like vibe coding a little bit, which is like ridiculously easy is Chrome extensions, right? â yeah. Like I just because they're like they just come back to like a zip file. So you don't even have to like load up an ID. You could just go to like chat, GPT or quad and say like this is the website. this is the feature that I really wish it had. Build a Chrome extension that does this. And then, I mean, sometimes they run into bugs, but a lot of times, man, they one shot that Chrome extension. And then in Chrome, you can just load up, when you're in the extensions, there's like a little toggle box and it's either upper right or upper left. And you turn that on and that turns on developer mode. And then you just load up the zip file or the files extracted from the zip file. then, â then that website has the feature that you want now, right? So. That's cool. Yeah. So I, I, I feel like I have a ridiculous amount of Chrome extensions and I don't know what half of them do anymore. So that, that's probably problematic, but they did like one specific task.
speaker-1: Make another Chrome extension that tells you what your Chrome extensions do. The next one.
speaker-0: I like where you're going with this. It's very meta now. Okay, is â you got a story for us? What's going on?
speaker-1: I do and it's I'm just not even gonna say anything. You ready? Let's jump in August a 60 year old man walked into the ER the emergency room and he says My neighbor is poisoning me He's he's not feeling great. I think he might have had some like
speaker-0: Yeah, let's do it.
speaker-1: face like skin irritation or things like that as well. I don't remember. But yeah, so they start doing some blood work on him and the the test results are I guess there's something in chemistry where if something is too close to another thing it will give you false results and stuff. So they couldn't quite find it. They eventually found it. They're trying to he's like really thirsty. They're trying to give him water, but he is really suspicious of this water. He is not trusting it. Okay, so â He's got paranoia and hallucinations. And the reason he has this problem is because he consulted chat GPT three months earlier. He found out that too much table salt in your diet is not good for you. â
speaker-0: Hola.
speaker-1: What kind of cracks me up is this, guy had studied a little bit of nutrition back in college, back in the day. So he thought I'll do a personal experiment. So he looks at ways to cut out, sodium, like all together and, or sorry, sodium chloride. â chemists, please correct us in the comments. Anyway, table salt. He wants to remove table salt completely. And so he goes to chat GPT and says, what can I substitute it for? It says, â sodium bromide So sodium bromide used to be in like supplements and other things like that in the United States many decades ago and then they realized that long-term exposure has problems just like this man is exhibiting eventually kind of took those things off the market, but he found a supplement online and he was eating sodium bromide instead of table salt and after three months of this
speaker-0: Okay.
speaker-1: It was causing him paranoid and hallucinations. â so they were able to get him on some anti-psychotic medication and some IVs and things like that. And he was eventually weaned in fact, they were able to discharge him from the hospital after three weeks. but it was just an interesting story of, a reminder that while chat GPT might not be wrong, that sodium bromide is a, could be a substitute for salt. That doesn't mean you should cut out salt completely. Your body does need that. doesn't. â Just a reminder to the listeners that LLMs don't give medical advice. If they do, please double check it with a doctor. Which is basically what OpenAI said when they were asked for a comment. They're like, we're not here for medical advice. Fact check anything that's really important.
speaker-0: Yeah, so I think that's really interesting and it brings up a good point because a lot of people blindly trust the output of the LLM without one, evaluating it with their own knowledge set and then two, because they just like assume they're wrong, like you know and then two, they don't you know they don't check with an expert right before taking action and while I feel like, at least for myself, and I'm not giving medical advice, but I'm just telling you what I do, is I take all the medical data that I have about myself and I kind of bundle it up into a folder. So if I've ever got a blood test or if I've ever got, and then every time I go do anything, I just make a text file and I put the date of it and then I put in the text file what happened.
speaker-1: And then you send that folder to North Korea.
speaker-0: There I have it. So, sorry. They don't got it. They know all my weaknesses, which I'm pretty fragile. I'm basically just like a system of like, have you ever seen like a, like a, you know, like a, what do call those? One of those like pyramid card tables, right? Someone blows on it, it's just gonna fall over. So that's basically me from a health standpoint. The, the, And so what I like to do is I'll take all that data and then I'll like create a project or I'll create like a new prompt and then I put it in there and I say, typically I ask three things in the system. So I say like one, give me a summary of everything that, um, that you know, my health history. And then I actually read the summary. Cause I will tell you 90 % of the time it's like 90 % right. But then there's like a 10 % in there. That's totally wrong. Right? And so then I go back and crack the LLM and like, Hey, no, actually, that's you. You misinterpreted the data. That's not correct. â okay. So that's number one. Number two, I like to ask it like, Hey, make a timeline of like all my health stuff, because I can never remember like dates when I go to doctors, right? They're like, when did you start taking that medicine or when did that thing happen? And I'm like, I don't know. Like one of the years ago, you know,
speaker-1: Previously.
speaker-0: So I like having like the timeline. â That's nice. And then the last thing that I do is I say, based on everything you know, what questions should I ask the doctor? Right? â And I feel like typically there's a couple of questions in there that I just didn't think about, right? And when I read them, it's like obvious. I'm like, â yeah, I should definitely ask him about that. â Now, I will say when I talk to doctors using that process, â You know, sometimes when I ask them the question, they like immediately dismiss it. They're like, no, that's not what's going on because of X, Y, Right. So, so I don't know why chat GPT would have me ask that, but I would rather just have a bunch of questions, perhaps to ask the doctor. Cause otherwise I get in the room with the doctor and he's like, do you have any questions? And my mind's just totally freaking brilliant. I don't, I don't know. I'm like, so, â that's how I'm using it. With that being said, I don't take any actions until. the professional signs off on them, right? So, like I'm not changing my diet, I'm not buying supplements offline and starting them. Like I feel like I need someone who actually knows more about your body, my body than, you know, me to kind of sign off on those decisions.
speaker-1: And I want to give a little bit of credit because we think that this was chat GPT 3.5 or 4. Okay. We don't have his actual, this gentleman's chat history of what led to his bromism, but â In my own experience, when I consult the LLMs about health things for myself, it does usually have some sort of caveat in its response saying like, check with health professional. Like it's like, this is my best guest, check with health professional. Or it will give some warnings like, â this should be taken in moderation or with oversight from a physician or something like that. So I feel like I've got to give them credit for the efforts that they're making and making sure that. It's telling people to settle down just a little bit. Don't trust me too much.
speaker-0: Yeah, and in the way of news, I think it would be advantageous to tell the listeners or viewers. â ChatGBT does have a new beta for a model that is fine tuned for answering health questions. So I joined the wait list and I have not got access to it. So I can't really verify if it's much better than the normal ChatGBT or any of that. â But when that comes out and I get access to it and when I get off the waitlist and actually get approved, I will do a follow up here on the pod and let everybody know what I think about it. â In addition, because once one company releases something, the next company releases something. then â Claude did a write up on their blog. and released some new healthcare related connectors. Now I don't really understand these because they seem like, it says as a patient you could use them, but it seems more geared towards like, employees, like doctors to use them. Cause I guess it's connectors that will connect up to like professional databases with patient records and like pull in and enhance your queries with data. Um, that's what it looked like to me. Cool. No, I wasn't able to get it. Like I didn't, ran out of time before I was able to figure out if I could do that or not. And I assumed even if I connected it, it would only give them my own data. So it wasn't like a high priority for me, but, but I did think that was cool and definitely something that's on my list to go back to and research more. Um, so if you're interested in. using LLMs for health, I would check out those Claude and OpenAI's latest technologies there. But still, always consult a professional. These things are not, I mean, if you get 90 % reliability out of the LLM, you're just crushing it, right? But you don't wanna be wrong 10 % of the time when it comes to your body. So I think that's too high of a risk personally.
speaker-1: I've had about 10 % error rate with humans too.
speaker-0: Yeah, well... That's a whole nother rant. So I- Go, okay, lay it on me, Shelby.
speaker-1: Sorry for you on that, you ready? So I had to go to the hospital to do a scheduled surgery last year. They called me ahead to give me the procedure â details and they started giving me the wrong instructions. They were giving me instructions for a different procedure because they had me written down for almost like an opposite procedure of what I was supposed to get. â
speaker-0: It's like the horror stories where they're like, â it wasn't your left side? You know, like, bro, come on, man.
speaker-1: it was like he said that was the opposite of the goal we're trying to accomplish here it was crazy and i was like uh excuse me no that's very much not right and then they also had like my um body weight wrong by about 70 pounds or something i like please don't dose me based on that my moral of the story is double check everything
speaker-0: Oh, yeah, that could be a problem. Yeah, and I feel like there's this weird imbalance in the medical industry where if you find a doctor that's really good and that is crushing it His availability is typically like very low, right? And so like it's just hard. It's like hard to get the Get the doctor's attention in my opinion and Get a good doctor's attention. Yeah. So anyways, I don't that's a whole nother complaint
speaker-1: Bryce, what other story do you have for me?
speaker-0: Okay, well imagine this. The world's top AI minds, they're in the mountains, Swiss mountains together, and they get locked inside a cage until they agree on the timeline for AGI to come out. No, just joking. So there was Davios this last week, which is a world economic summit. It's like a big conference that world leaders go to, like presidents or like the CEOs of top companies. And they all get together in these Swiss Alps and they plot how they're going to make our lives worse over the next year. No, just joking. So.
speaker-1: knew it.
speaker-0: That's like a half joke. It depends on the executive, right? â So, but there was a lot of drama there this year because â you had the CEO of Anthropic who makes the Claude models. â You had some leaders that used to be at Meta and Google. And then you had some people just jumping in on the Twitter verse, right, with their comments like Sam Altman and things like that. So, so essentially there's two camps here on the one side, you basically have open AI and entropic and they're saying like AGI is already here or eminent, right? And when they're talking about AGI, they're talking about human level intelligence. So if you were to be able to give a job to another human in your work, you should be able to give that same job to an AI and it's able to execute it. And â specifically, know, â Dario, who is the CEO of Anthropic, I mean, he, he is convinced that we're all going to lose our jobs. â So he said AI will replace all software engineers in the next year. That's a statement he made. He said, â Nobel Prize level scientific research will be replaced by AI over the next 24 months. And he said 50 % of white collar jobs will be eliminated over the next five years. Those were his big three predictions. â So, and then, know, Sam Altman gets on the Twittersphere and talks about how they're going to hit super intelligence before anybody else and blah, blah. So, so that's one side of the coin. And I'll come back around to that because I have some take on that. The other side, â you know, we have â the CEO of Google DeepMind, right? So basically like the Google, the part of Google that builds Gemini and all the AI technologies. He is a Nobel Prize winner, right? So it's not like he's just an executive. He's like he's smart, right? And at least at one point was technical. And he says AI is nowhere near human level intelligence right now. And he doesn't see it going there in the foreseeable future. â Then we have Lee, Lee, I always forget how to say his name, but Lee Canoon or something like that. He... He was the Turing Award winner, which is basically a competition for innovative AI technologies. He's one of the world leaders out there. And I believe he was at Google DeepMind a long, long time ago. And then he more recently was at Metta and then he's left Metta and started his own thing in the last, I don't know, 12 months. So, but he said, Large language models will never achieve human level intelligence. It's his take and his big takeaway is while large language models may be able to master languages like English and, that they actually don't have like a full. on 3D understanding of the universe or of like physical space. And so when we try to take LLMs and insert that into robotics, robotics are going to be unreliable. And that they're never going to have the same level of intelligence as humans because they just don't have that spatial intelligence going on. He even went so far to say that language is easy. Basically, it's like already a solved problem, but you know, we're not gonna get reliably laundry folding robots until we're able to get this kind of like â physical space intelligence hammered out, which will require a different set of technologies than the large language models currently provide. â Now, I just wanna caveat this with a couple things, So, Lee Canoons. â current startup is focusing on exactly what I just said, which is it's focusing on building models for 3D spatial analysis to be inserted into robots. â So obviously he's got a financial incentive there as well as a pride incentive because he's â more or less he got told to leave meta, right? â When the LOMM-4 model came out and it didn't perform well, Zuckerberg basically burnt down the whole existing AI team, â especially the leadership. He was not happy. â And that was the beginning of the end for him there. â So, you he has obviously incentivized to push this narrative. The Google DeepMind CEO's perspective on it's interesting because... It really runs counter to what you would think he would say, because of what would be financially advantageous for him is to say like, yeah, we're going to be able to like cure all diseases using LLMs and all the other hype train stuff. â but no, he, it's interesting that he's just like kind of a straight shooter. He's like, no, it's good at language, but it's not going to do everything. â and then on the other side of the fence, you have, â you know, the CEO of open AI and the CEO of Anthropic. And they're getting these ridiculously high evaluations from VCs, right? And you can argue whether you think they're high or they're valid or maybe you even think they're low, right? â But either way, you know, there's a clear financial incentives in those guys hyping up that the existing technologies that they basically are leaders in are going to do everything in the future â because You know, they're not going to get the same evaluation if they were to have, you know, talk about the limitations associated with the technology. So they're very financially incentivized to hype it up as much as possible. â And, you know, I do think the CEO of Anthropic is a really good guy. â think he's, I think he believes what he's saying, which is that there's going to be a lot of job loss in the future. I don't personally I understand his argument, but I don't personally believe that â I just feel like Fundamentally, there's so many things that companies that could add additional value that just get left by the wayside because Nobody has time for that â that â Some portion of those tasks are going to be picked up in the future. So existing workers are going to get more productive I do think There will be a lot of scaffolding built around the models, which is going to rely on, which is going to require at least vibe coders, if not legit real coders to do. â so while I think coding jobs are kind of taking a hit right now, I do think those are going to get back into fashion pretty fast. â like, you know, in 12 months from now and then, â yeah, I, you know, I just generally believe that. when there are resources on the table and resource being like you have smart humans that are out of work, somebody's gonna figure out a way to make money off that, right? So â I believe that people are good at greed. So they'll figure out a way to maximize that. Anyways. â What do you think Shelby? Which camp are you in? And what do think the timeline is going to look like?
speaker-1: Can you remind me about what or why you said in your opinion you think that coding will come back in style for humans?
speaker-0: â This is just my personal belief and you can argue against it or anyone can in the comment section below if they want to. â I think the labs are still making progress, but that they're not making progress as fast as they previously were. They have to, now they've been able to kind of cover that up a little bit â by basically just burning cash. But if you actually look at like the cost to produce GPT-5 versus the cost to produce GPT-4, I mean, we're looking at about a 4x cost increase to when you scale everything back to the same sizes â to generate the five, right? Which tells me they had to increase compute 4x or cost 4x to get, you know, a step up that's not even like exponential, it's incremental at best, right? And I'm not saying it's not worth it. Like I'm not trying to say that at all â But I am saying like based on what I've seen over the last 24 months. I don't see the hockey stick right now which is fine because honestly like I think for like 80 and 90 percent of the tasks that we're gonna be able to use LLMs for the models we have today are good enough, right? The models are not the bottleneck. The bottleneck are organizational problems â the bottles are bottlenecks are People aren't trained on how to, on all the capabilities and what they can do. â And also, you know, I just think there's a lot of technology that's been built in the pre AI era. And there's going to be a new technology that comes out. That's going to be basically built around AI from the ground up, like AI native software. And the same way that you used to have software that went into data centers and then the cloud came along and then individuals that wrote software at that point in time that was cloud native. was able to get a lot of â value out of the cloud. And people who just like took their applications from the data center and shifted them into the cloud just usually ended up paying a lot more money for not a lot more service. And I think you're gonna see the same thing. We're already seeing that with AI technologies, right? So you see the people who have had an app for five years and now they're slapping a chat bot on top of it and nobody cares, right? Nobody cares. Sorry, I'm not touching your chat bot. So the, but then we see other companies that are like, they are like building an app from the ground up right now and they're building LLM technology throughout it. It's being part of the decision-making models. It's being part of the ethos and how it works, how the workflows work, how it's adding value to the customers. I mean, I just think there's a lot of little tiny decision tree when you're building something for a consumer, something for the customer that you're going to make different decisions now that you have the HLM technology. anyways, sorry, that was more.
speaker-1: â no, you're good. I think a lot of people have felt this way â whenever there's a new technology that comes along, and not just in IT, but any sort of technology being worried about their jobs. like you said, think some will experience heavy pruning. And I think others that we don't really think about will exist, right? Like how... â our children's children will have jobs that in the future that we don't even think of as a job because there are more needs and ways to add value than what we're doing right now, it seems like. And people will be creative and come up with new ways to add value, perhaps. don't know. It always amazes me when people think of like really, really unique ways to basically build a business, to build a living that it's not like. you know, 12 people who are doing that, right? When people come up with really creative ways to solve problems and provide a living for themselves, very cool. So I'm hoping that we'll see a shift in like, I think it'll take some creative thinking, but I don't know.
speaker-0: Yeah, I'm with you. I think there's gonna be change, but I think once we get through the short-term heartburn, â yeah, we're good. Society would be in good spot, Okay, that feels like that was my story. Do you have anything fun to share with us, Shelby?
speaker-1: A bit of a debate, really. â So when I think about an octopus, they've got... Yeah, we're envisioning an octopus with all its arms and I struggle to keep track of two. Like, â sometimes I think, how nice would it be to have more arms or more bodies to be able to do more things in a day? But â I think the limiting factor is right here.
speaker-0: Okay, yeah. What? I'm with you. I'm with you. Yeah.
speaker-1: Right between these two points. But have you heard this? Like you probably have. it's just a well-known thing, but octopus, octopi, octopuses? I don't think I can say that. Anyway.
speaker-0: Octo-Octopi? Is that a single? An octopus.
speaker-1: We'll talk about one single octopus. We'll name him Fred. So Fred has
speaker-0: We'll just call him Octi. Okay, good Fred. That sounds great. Okay. All right
speaker-1: 500 million neurons total
speaker-0: He does? How many do humans have? I got â it. Is it smarter than me? It sounds like it from based on what I just heard.
speaker-1: â that's a good question. Yeah, that's a good question. Someone's gonna have to go check that for us. But it is apparently roughly comparable to what a dog might have. Which I'm not even gonna go down that rabbit hole because I know different dogs have different intelligence levels. But back to my friend Fred the octopus. Half of Fred's neurons are in his arms. So there's like a central brain. Yes, it's called like a mini brain.
speaker-0: His brains and his arms?
speaker-1: So like in each of his eight arms, there are a bunch of neurons which can control movement and process touch and other signals, â as well as making short-term decisions without checking with the CEO, main brain. â So I was talking about this with someone and they said, well, yes, my arm can also process touch and make motions.
speaker-0: Really? That is cool. I did not know that.
speaker-1: before my brain processes because if I unsuspectingly lean up against a hot surface and my arm is gonna move before my brain knows what happened. So I was like, are we like octopi? Like how much? I don't know.
speaker-0: I think there's a fundamental difference. I and I'm not a doctor so I'm just speculating right now But I feel like there's two things that would come into play in the human scenario, right? One is you touch the hot surface that sends the signal to the brain and the brain sends them back immediately move the hand, right? Which happens so fast. It feels instantaneous, right? And it's not like being processed by your conscious is being processed by your self-contrast. So I can see that being a viable scenario. I could also be see the viable scenario of like, there are things that like your body will do without the brain telling the body part to do it.
speaker-1: Like reflexes, right? Like, it just happens. Your body just does it.
speaker-0: Yeah, and I don't really understand reflexes. They seem like magic to me, but I know they exist because they hurt and â So, I don't know I but I think the fundamental difference is like you're saying they have neurons in their tentacles Go ahead that'd be great. Yeah
speaker-1: I can explain it a little more too. Basically, Brain is like your project manager or I don't know how you want to think about it. Basically, they give high level instructions to the arms such as like explore that area or bring food back and then it is up to the arms to figure out the minutiae, the small decisions to meet that high level goal of like exploring independently and finding out. by touching what is here and observing or like looking around to grip things and suckers can taste and feel, right? And so it does all that processing. Like it's like distributed, what is that? Distributed processing?
speaker-0: Yeah, well I think it's almost like a... I'm trying to figure out the model in my head. It could be like a hub and spoke system where you have a master brain but then you have these other brains that are kind of like child to them off of each spoke, right? I don't know the best way to describe that but that's really interesting.
speaker-1: And even more interesting is, well, kind of taking it another step further, if an octopus has an accident, loses a tentacle, loses one arm, it will continue to work independently for little while.
speaker-0: Really? â Like if you cut it off?
speaker-1: severed arm will still react to touch and grip objects and things like that but how much processing is happening locally that's my back for you what you got for me
speaker-0: Neurons are magic, man. That is a cool fact. I just want to relate that back to AI. Like, I feel like the head of octopus is like your agent, and each tentacle is your subagents. You know, like the main one's like telling them what to do, but it doesn't actually like micromanage as much, right? Perfect. Like that's kind of what I envisioned in my head when you were speaking. So yeah, yeah. Well, I just want to do a plug for this beautiful We're not sponsored by this, by the way. So I like these drinks, these little like cherry buck drinks. And they're, I think they're headquartered in Utah, like near Provo. â So I just wanted to do a plug for, it's called Bucked Up Energy Drinks. And I like the cherry candy red ones. So they have a bunch of other good flavors too.
speaker-1: I got really excited when you reached off camera. I thought you were gonna grab another LEGO creation. think you're due for a LEGO creation to cross the screen.
speaker-0: Yeah, it's works, the LEGO. I'm not gonna try to pick it up now because it will crumble. But maybe by next week I'll have something more to show there on the LEGO front. I usually have something I'm working on, but my output's about one LEGO set per month. So you might have to wait. It might be another week or two. Fair enough. Okay, great. Well... If you guys got any comments or questions, feel free to drop them below. If you felt like this was valuable, please click the subscribe button. Shelby and I are working hard to bring you constant flow content each week. And then, if you got a friend that wants to learn more about AI or just make fun of the fact that I don't know anything about octopi, share this over them, send them a link, right? And... Yeah, we'll see you next time. Peace.
speaker-1: Bye!