the localhost :0001 Frankcx: [00:00:00] hey, everyone. Three months ago, we could not stop talking about price: cost per token, AI accountants. And then it changed. What changed? We're now talking about what intelligence costs and who gets to control it. From cost to control, that's the argument I've brought today . Frankcx: I wanna welcome you to our first show that we call The Local Host. There's a good joke in there from how we talk about local AI. I'm Frank, and this week I am your local host. This is episode :001, and if you don't get that joke, it means episode one for those that don't speak port. So there are five of us on this conversation. We're here to talk about local AI, have a good time. We're a group that got together normally anyway [00:01:00] just to talk about AI and what's going on, and we're just giving you the opportunity to listen in on our conversation. So hey, let's go around the horn. Frankcx: Tell us who writes your paycheck. Tell us a little bit about yourself your excitement about the show, and let's go ahead and get started. So Jacob, how about it? Jacob: how about it? Thanks Frank. Excited for this. My name's Jacob Rhodes, and I am passionate about technology and more recently very passionate about AI. Love getting to learn and study and grow and teach people, but mostly do a lot of listening and learning myself from the folks on this call or on this on this podcast. My paycheck is written by Microsoft, so that's an important disclosure. And my day-to-day is spent kinda learning and teaching about technology and computers and, Surface devices specifically. Frankcx: Awesome. Robert from Florida Robert: Yeah, coming to you from Northeast Florida. I my name is Robert. I am also funded and sponsored by Microsoft. I run I work in our engineering organization along with Neil, and we [00:02:00] run something called the AI Factory for Surface which is re- a really cool job that we get to go build a lot of different prototypes and demos and talk to customers about all the reasons why they should be thinking about running AI locally on their devices. Robert: Really good to be here. Awesome. Great to have you here, Robert. And then Chauncey Chauncey: Uh, yeah. Chauncey Larson. Uh, what does it say about me that when you say port, my mind doesn't go to technology, but instead wine? Um, I am the, uh, uh, I don't know, self-described, uh, technology nerd as well. Absolutely just, uh, e-enamored with all things technology, whether that's AI or hardware or software. Um, I too , am sponsored, as Robert mentioned, by Microsoft. Chauncey: Specifically, I'm in the, uh, Surface team alongside Jacob on the Surface marketing side. So, uh, I get to help tell all of our nerd stories. Uh, and so yeah, that's, that's what I do. Frankcx: Good thing we're recording this, , at midnight so that you're all off the clock, right? And last but not least, Mr. Neal Neil: Hello, everyone. My name is Neil [00:03:00] Misek, joining you today from Kirkland, Washington. A relatively new resident of the Pacific Northwest, born and raised in Chicagoland. As Robert mentioned, I'm a proud member of the Surface AI Factory sponsored by Microsoft, uh, helping customers and partners unlock the full potential of Surface. helping customers maximize not only the neural processing unit, but the CPU and GPU as well. And when back to the, the joke on port port to me as a distant sailor means the, the left side of the boat. Neil: I think we've got the full gamut here of IT, wine, boats you name it, we've probably got it covered here. Really looking forward to the session today, and I'll turn it back to you, Frank, our local host. Frankcx: Yeah. It's exciting that you moved to Kirkland 'cause Jacob and I both live in Kirkland. live in the annex part of Kirkland. Jacob's in the cool the sweet Frankcx: part Jacob: you gonna dox us, Frank? Frankcx: a coffee shop. Yeah, exactly. Chauncey: that mean that one of these upcoming podcasts you guys can all just go into one long, [00:04:00] like a table with a cool background? Frankcx: all wear Costco gear that says Kirkland on it, 'cause that's... Frankcx: and Chauncey: we're baking, breaking into other podcasts Frankcx: it. Costco could sponsor this. I wouldn't mind that, Jacob: They have a vested interest in local AI, I'm sure Frankcx: I'm sure they do. So that brings us to this, this show will have a regular format. We'll open up with our little bit of disclosure like we just did. Frankcx: And we'll also just tell a little bit of catching up and what's hot. But, the thing that I wanna talk about is also we- we're gonna have a section called The Rundown, and The Rundown is essentially where we just talk about what's big in the news, especially around open-weight models as well as local AI. And so the first two that I wanna bring personally is the fact that Qwen, which is a Chinese model, it's an open-weight model and it essentially there has been a really good model out there from Qwen called Qwen 3.6, and it had been out for a number of months. But Qwen recently released Qwen 3.8. The Rundown: News in AI Frankcx: Now, from my own [00:05:00] on my essentially rig that I have behind me called The Beast, which has two 3090 cards in it Qwen 3.8 is faster than Qwen 3.6, but it kinda doesn't think as well as it does. Qwen 3.6 essentially was a model that took its time to think about things and always got it right, where Qwen 3.8 is faster, but it is a little sloppier and it uses tools sometimes that it maybe doesn't need to or shouldn't in order to get some of its work done. Frankcx: Now, I'll follow up with that a little later, but it is interesting to see Qwen really charging forward. They have a model that can literally run on a 3090 card or I think could run on a Copilot+ PC if I'm not wrong. Maybe somebody could gut check me on that. But that's the cool thing about these models is the fact that they are essentially shrunken down quantized as... Frankcx: I don't think that's the Chauncey: Quantized? Quantized? Quantized? Robert: quantized. Yeah Frankcx: They are quantized down so that they could run on smaller footprints of hardware, whether that's on a Copilot+ PC or a [00:06:00] 3090 card or the new soon to be released RTX Spark Surface Laptop Ultra and the other devices that are in that classification. So this is all really relevant to how we think about it. And these models are super smart in how they work. They honestly stand head-to-head with some of the frontier models that are out there but they just have a smaller footprint and they maybe have a little bit smaller intelligence. Frankcx: But this new Qwen 3.8 model at 27 billion parameters can easily fit on a 3090 card and you too could benchmark this model just Frankcx: like I Jacob: I have been doing the same, Frank, as well. So I was there for the countdowns because the local like open weight community was like so passionate about like waiting for this 27 billion parameter variant. A lot of people are waiting for a mixture of experts variant as well that only loads a few billion parameters, but has 30 or 35 billion available. We haven't seen that yet, but I'd be very excited because you can get some really fast output that. But I've been loving it. I have a 3090 Robert: have you Jacob: too[00:07:00] Robert: parameter on a device yet, Jacob: have run on a few, yes. A couple I can talk about openly. Not all of them can I talk about openly. But, Robert: That's funny 'cause I went-- I loaded it up today to go run it on RTX Spark and it it threw a bunch of errors. So it w- it was having some problems. So I'm gonna have to continue to investigate what the problem is, but that was using LM Studio. Jacob: Yeah. Jacob: I will say I've had struggles with Elm Studio in running that model, Robert: I used Jacob: because... Robert: specifically Jacob: Ah, very interesting. I haven't tried that out yet, actually. I've been sticking to the standard Robert: j- Chauncey brought up a good point earlier that maybe it's a good time to ask around this whole open weight, Chauncey: Ja Robert: source, closed, m- maybe Chauncey: a general catch-up on some of the technology, 'cause like I'm pretty decent caught up, but there's barely-- there's so much to learn when it comes to AI. And yeah, help us understand, or maybe some other one else on here can... what's the difference between open source, open weight? Chauncey: Why should [00:08:00] that even matter to the folks that are listening or to anyone that's out there? Frankcx: Yeah, I'll give it a, I'll give it a go, and I'd love others' input here, but, this is actually leads to the tension of the week that we'll talk about. And that's the point of this show is we'll bring up a tension, essentially like a controversial or kind of something that has an opinion to it, and then I'll give a pitch, and then we'll have a discussion around it. Frankcx: But we'll jump into a little bit of that even tension right now, is that when you look at how models are built you have to think of it that there are frontier models out there like what OpenAI and Anthropic make with both ChatGPT as well as Claude. And those are essentially what we would call closed models. Frankcx: We don't know essentially how they are structured. We don't know how essentially some of the training happens. We don't know a lot of the kind of the data that goes in and out of them. We know the results that they have but we don't essentially know, how they're built. And there's some probably good things in there. Frankcx: Anthropic would probably argue that's out of safety because they wanna make sure that they can be running very safe models that aren't doing things that are [00:09:00] destructive and they are controlled, and therefore, having a closed-weight model like both Claude and OpenAI lead to those. Gemini is another model from Google that is a closed-weight model. Frankcx: Spark is a model right now is closed, but Meta is actually saying that they might make it open, meaning that they would release it to the public. So here's the difference is that an open-weight model is a model that you can actually, just like open source, you can see the structure of it. You can see the training. Frankcx: You can see essentially the weights and how it's structured. You can essentially quantize it down to be a smaller model or even distill it in a way that you could actually use it on a different leve- layers of hardware. I think for our discussion, when we talk about local AI, we have to be looking at right now anything that is a open-weight model because of the way that it has to be able to be shrunken down, quantized in order to run on local hardware. Right now there is, I think, don't know of too many closed-weight models that are running on local [00:10:00] hardware but maybe Microsoft is actually controlling some of those with some of its image Jacob: Th- But I Chauncey: hvad? Jacob: another step open weight that you touched on a little bit, Frank. I'm curious if Robert or Neil, you have any insights because like there's none of those open weight models are truly open source. I d- I Frankcx: They have licensing behind them Jacob: don't know if anyone has-- wants to go more into detail. I can take a stab at it, but there's... I think training is the biggest thing that the transparency in training is not there for most open weight models, but there are some that, that truly go the step further Neil: Yeah, I would say, from an industry perspective we're seeing a lot of customers request open weight models. Basically, instead of a universal model that's the right fit for healthcare, the right fit for financial services, right fit for state and local government, they want the ability to take these frontier models and bake them into their [00:11:00] workflows, and there's certainly an aspect of low-rank adaptive tuning, especially once we're talking small language models. But essentially, what are we trying to get out of that, right? When we're talking, trillions of parameters, right? The ability to take that off-the-shelf model, make it valuable for my business, have control over the inference, the responses back from that model that's the business priority. And a-agree with Frank's comments around the, Anthropic argument that we can't release these models as open weight from a safety perspective. But as you think about the future of customization, and instead of throwing every single prompt at the, at these frontier models being able to make them personalized to our business, to the specific use case, is the direction of travel I see the industry going. Neil: And I think the open weight approach helps companies get there faster. Chauncey: So basically Frankcx: yep Chauncey: comp-- open source is completely scrutinizable. You see the secret sauce, you can go [00:12:00] in and change and modify and take the code and do whatever you will with it. Open weight... yeah. Robert: Yeah Chauncey: weight is all the way up to the secret sauce basically, which is like you can modify and you can do a little bit of scrutiny, but for the most part, it's delivered as is, and the idea is that you might need to customize it a little bit, but beyond that it's still somewhat locked down. Chauncey: But not as locked down as a closed model Frankcx: think an example would be, I think Mistral is an interesting company. For those that aren't aware of Mistral, Mistral is a French AI company. They build their own model and it is an open weight model. But they actually sell the model more for essentially commercial customers to say, "We can give you this huge trillion-parameter model, but you could install it in your own data center, therefore it's sovereign," meaning that it is local to you. Frankcx: Now, local in this case isn't on a s- Copilot+ PC. It's literally like I own a data center and I can install this trillion-parameter model in my data center. It's still huge, but it [00:13:00] does give you essentially local control over it at a sovereign level, there's a lot of European customers that don't wanna share data onto US data centers for various geopolitical Chauncey: and is this quantizing something that we would expect customers to be doing? Is this something that the Mistral would be doing on behalf of customers? Is it a mixture of both? Like, when we go, especially with open-weight models being the main topic here what's the barrier to entry? Frankcx: that's a good question. I think that there-- you gotta have the right technology. You think of models, they have two sides to it. There's kind of the training and the quantization, and then there's the inferencing of just using it. And so I would think if you have the right hardware, you could do it, but it's certainly not a cheap or easy Jacob: You, you can distill-- the community can distill models, but they're-- it does require expertise. And so generally most of the AI model shops will put out specific weights and potentially a list [00:14:00] of quants as well. But the community then takes that and runs with it and can distill down some of these models. Jacob: I saw quant 3.8, 27 billion parameter model like running in one gig of RAM on a phone. Like Frankcx: Wow. Jacob: it's not gonna be accurate, it's not gonna be great, but Frankcx: you Jacob: y- Frankcx: run it on your Jacob: yeah, you c- you can do those sorts of things. I don't know, Robert, Neil, you got any more insight on that? Neil: I would say, a lot of the model makers will offer different tiers. You've got a version for your devs that own the beast like Frank does, and then you've got a version, that is designed to run on more mainstream user hardware. Neil: What is the quality output? Is it good enough, right? Is it good enough at that Robert: Yep. That's the balance Frankcx: Yeah. So I'm gonna keep things moving in the rundown 'cause I wanted to talk about one other piece of things that's in the news, and it's a little spicier. It's actually around these flock cameras that are Chauncey: Oh, man Frankcx: for those that don't know much about the flock cameras, and I am not an expert on them, but they do represent an interesting AI capability of [00:15:00] AI on-device as well as hybrid AI in the cloud. my understanding, what these flock cameras do is they essentially take pictures of license plates, and they're used by law enforcements to track down bad guys. But there's roughly 120,000 of these cameras out there throughout the United States right now. They have right now a seven-day retention, they can essentially be used to look at tracking of license plates into particular areas. But what I think is personally super cool about them, and this is controversial that I would say is super cool because no-- there are not a lot of people that are thinking these cameras are super cool. But I would think it's super cool from a technology standpoint just because it is doing inferencing locally on the device in order to analyze an object, like a license plate, collect that information, put it into a pixelized container, and then send it up so that it can be inspected at a cloud level where data analysis can be done at a much broader level. Frankcx: So to me, it's kinda playing this idea of, I think, where, many of our surface devices or any other on-device could be used in a way that you collect [00:16:00] and inference locally, and then that data is stored in the cloud that could be used for larger data analysis. I think there's a lot of different use cases out there that are like that. And unfortunately, the flock camera is getting a bit of flak because of the way people are misusing it at least in my opinion. But but I think that if there's better safeguards on it, it could honestly be a pretty interesting technology that-- to be replicated by many customers out there in the enterprise space for that. Frankcx: But I'm interested in others' opinion on that. Robert: Yeah, I think this whole thing started with a college student in Michigan, the police about who owned the data on those cameras. So it definitely brings up some controversial, Frankcx: privacy Robert: some controversial things. But to your point, Frank, I think the idea of the technology behind it absolutely makes sense in a lot of commercial scenarios. That obviously wouldn't have the same same controversial r- things that go along with it. thing too is there's a website, I think it's [00:17:00] called where you can actually go see a map of where all these cameras are. Just out of interest. Frankcx: can even look up Chauncey: Just out of interest Robert: Yeah. Frankcx: in the Frankcx: database too, right? Robert: If you wanna really geek out, it's it's kinda cool Jacob: yeah, at this point I've ac- I'm accepting that every time you go into a community that has this you're being tracked everywhere. Jacob: The technology of the processing of the data is super interesting to me and I agree, Robert, like I'm very interested in it. But the, the trust that puts in the data handlers, third-party organizations, and even, some government organizations is a whole nother topic that I'm not gonna get too into today, but Jacob: it's definitely interesting. Tension Pitch: Governance & Open Weights Frankcx: I think that brings us honestly to the tension pitch. And so here's the tension pitch. So this is essentially changing the piece, and we're gonna move out of our news of the week. But the tension pitch is talking about governance, right? It talks about this kinda concept where we're shifting from is the model good enough to being more like who's governing it and are we allowing it to be governed? [00:18:00] And the reason that this came up is because a lot of these open-weight models are jumping over safety barriers that are, that the closed-weight model manufacturers like Anthropic and OpenAI would say, "We take a lot more safeguards in how we produce and release our models than some of the open-weight models are doing out in the market space." And so there was a lot of tension brought up in the news that essentially said, "Hey, maybe we should ban these open-weight models." And then all of a sudden, Jensen Huang from NVIDIA posted his first post on X ever. He didn't post about the GPU, but he posted about this letter that he put out that essentially said how in favor he is of open-weight models. Now, you could call it like gold rush, where he builds the shovel that helps people find the gold, right? He's open and he wants people to use open-weight models because it benefits NVIDIA. But along that line, 270 people or co- organizations signed that letter to include Google and Microsoft. Frankcx: But there was [00:19:00] one company that was you know, interestingly enough, not signing the letter, and f- as of recording this, I don't think they've still li- signed the letter, and that's Anthropic. And Anthropic, essentially, you could make an argument that Claude, today is still the number one with fable AI in the market space. And I make this argument, when signing costs you nothing, everybody signs. But, when you're at the top of the heap of this model or of this essentially world, you would probably be like, "Hey, I'm gonna wait before I sign anything 'cause I'm not gonna give away the crown jewels if I own the crown." But what's interesting is that along the line Mr. Zuckerberg, Mark Zuckerberg, who has in the past, in my humble opinion, not really stood for open policy and, disclosure of how things work, he put out a manifesto of intelligence for everybody, and I highly recommend, everybody take a quick look at that. It's super cool, and he actually has some really good points in it. And it seems like it wasn't [00:20:00] written by AI. It's got Mark's kinda language and his speak inside of it. But he's essentially come out to say that intelligence is for everybody, and he's in big f- kinda proponent of open-weight models. Now, I would argue again that if he was at the top of the heap and he was running Anthropic, he probably wouldn't have come out with that letter. Robert: Yep Frankcx: I'm proponent of open-weight models, but it's interesting to see that, we are essentially having this policy that- When you sit and look at this idealistic world that we're living in of like openness and provisionals, I, I would love you guys to convince me that I'm wrong, that Mark and the o- the rest of the world are just trying to play catch up with Anthropic. Even Google in that space, and Microsoft, we're all trying to play catch up with Anthropic. And, it's interesting that OpenAI signed the letter 'cause you could honestly say that they week by week jump ahead of each other still in who's producing the top frontier model out there. But I'd love to open this up to discussion on, how we think about open weight models and, their value in society. But, you look [00:21:00] at Anthropic, and they take more of a safety look at it as well. Robert: I thi- I... Look, it's hard to argue. N- first of all, number one, anything that is in the form of a manifesto is never good. I don't think I've ever read a, a really positive manifesto. that's that's mistake number one, 6,500 words. And number two is it's interesting you mentioned about sovereignty, against the idea of open weight models, but this Open Weights and American AI Leadership Initiative, I think is what you were talking about, that everybody's been signing and touting. Initiative actually argues that downloadable models are critical for things like innovation, economic scalability, and sovereignty. So it's interesting to see kinda how the different sides are playing this. To your point, I think though, Frank, it's as positions on the stack change for each one of these companies, so will their their point of view Frankcx: Yeah. No I think that I'm interested in everybody [00:22:00] else's opinion here, as maybe we go around the horn or if anybody has opinions Jacob: I think it's all a chess game, to be honest. I think that it's literally like when someone's in the lead on a, a closed model they're trying to acquire customers, prove the value of these inflated market cap or inflated interest organizations. A- and then, everyone else in the industry is like, "Oh, we need to-- we need everything to be open." Now, I think the five of us here, at least for me personally, like I'm a huge proponent of people being able to tweak and play with and adjust, and I think for commercial customers, especially for organizations like that is critical. And there is trust that they can establish and great relationships with closed weight models and these companies and organizations. Jacob: But won't be the same. Like fine-tuning for your use case won't be the same if there's only a couple of gates, a couple of ways of entering that. Someone who's been playing around a lot with local [00:23:00] AI, personally and professionally, like more the merrier, we need this. But I'm a little fearful because there seems to be some stuff coming to a head right now. In the next few months, I think that like even licenses that some of these open weight models are releasing under might get a little bit more restrictive. We're seeing that a little bit with Qwen's biggest model is not under the Apache 2.0 license the same way that some of their smaller models are. And so I don't like the trend moving away from open weight. So if I could have signed that I would've, but I guess our-- my employer did. Chauncey: I love the altruistic vision of open, right? Which is this is the world that ultimately the technology should operate in. And there's arguments of the risks of open source in particular and there's definitely arguments for open weight and to what Jacob, you were just mentioning. Chauncey: I would also argue to a certain degree, some of this is just because they haven't figured out the cost model, right? Or like the, the revenue model, I should say. I, I-- Frankcx: is an IPO coming up for Anthropic, so Chauncey: Like at what point do they like just say, "Oh, [00:24:00] wait, like maybe we want to make money on open weight," or, "Maybe we want to make money on open source." Chauncey: Even though it's titled open source, like there's still plenty of open source technologies that still have some kind of funding, like revenue back to it. So I'm very curious to see how that ends up playing out for everyone, because ultimately, like that kind of dictates the usage of open weight models versus going to the cloud and the scalability of all of them. Chauncey: Because if you just shift costs from one place to the other, then you're creating more problems that way. Neil: So I think there's three layers to this. There's the argument for secu- security from a closed model perspective, right? There's also the economics of this, and then there's the advancements in local AI hardware. So I'll start with the latter, local AI hardware. The model... Frontier models in the cloud two years ago can now run locally on your device, right? Neil: So that's how fast this is moving. Chauncey: Mind-blowing Jacob: even one year ago [00:25:00] almost with Chauncey: Mind-blowing. Yeah Neil: It's amazing, right? And then furthermore, I think as a result of that, we're seeing these closed model providers offer N-1 free of charge, right? But yet Fable you have to use your usage credit. So you need a, a cloud subscription plus usage. And so I I agree with Anthropic's current stance of, "Hey, we're at the, we're at the top of the food chain here. Neil: You all have to pay us to access our closed model. Why would we share our secret sauce?" No different than an inventor wants to patent their invention, and they want an exclusive license to monetize that for a certain period of time. In the spirit of, ri- rising tides raises all boats I agree with the general consensus of, "Hey, this should be available to the community. Neil: We should have these models available with open weights." But from a, an IPO's coming I think Anthropic may be over-indexing on the [00:26:00] safety argument, and I agree Frankcx: Right Neil: really an, an economic one, right? They want people to pay for Fable. They want monthly active users. They want recurring revenue so that they can get the, the largest IPO possible. Neil: It'll be interesting post-IPO if that shifts from a regulatory perspective or if they really take more of the Jensen approach of now we're selling... Maybe the models become the, the shovels," right? In more of an open weight sense. Frankcx: Yeah. I was listening to this other podcast it told this interesting story of of Anthropic in that if you go back and look at why OpenAI was founded, it was founded by Elon Musk and some others because Gemma or Google and DeepMind had essentially, put a lock on the environment, and because Google was owning AI, there was a fear that Google would essentially commoditize AI. Frankcx: So OpenAI was formed in order to essentially take commoditization out of AI. That's changed, right? So then Dario Amodei left OpenAI to go form [00:27:00] Anthropic because he was fearful that OpenAI was gonna become a commodity and, make money off of AI. And then now Anthropic is, talking about doing a two trillion dollar IPO in the next, month or two. Robert: Good Lord Frankcx: You're looking at, like, where is the altruism in this, this story where you were like, "We started off against this fight against Google," and meanwhile, Google's probably just showing that they're just true to their word and trying to produce the best AI that's out there. But all the others are like, "Oh we like your model. Frankcx: Maybe we should make some money too." So that's just a, an interesting take. Chauncey: The reality is AI's expensive, right? It is expensive to maintain, it's expensive to train, it's expensive to run. So there-- that you've got to have some kind of funding and some kind of value, like at least in this capitalist market that we've got today. There's no getting around the way that we drive value towards those top contenders because they are the ones that are proving the most value out of what they're producing. Chauncey: So I'm not really surprised by [00:28:00] Anthropic going this direction. I'm-- I'd be very curious to see if it's a pattern that keeps on repeating and we get some random person at Anthropic starting, semi-closed AI or something like that. But it'd be curious to see how this all goes for sure. Chauncey: Yeah. Robert: It's a, it's an interesting point too, Chauncey, is are the vendors gonna start charging for some sort of license to run AI locally? Frankcx: Yeah. Robert: does that become, Frankcx: good question Robert: Does that become something? 'Cause to your point, running out of data center space, running out of power data sovereignty, all of those sort of things. Robert: It'll be interesting to see what happens Chauncey: And Frankcx: oh, go ahead Chauncey: I think Jacob, you and I were talking about at one point, does it make sense instead of running the model having a cost versus supporting the model having a cost? And so like you can get access to the model and you can run it all you want, but if you want like help from the company, then you probably need to pay for it. Chauncey: And so I think there's also creative ways to get people-- Yeah and to get people to like to be curious about and use, 'cause you still need [00:29:00] to generate like usage of things. You can't just put it behind a license 'cause then that ruins the whole model anyway, other model. But if I put some kind of support cost, if I put some kind of thing like that, that might he-help at least open up some other revenue streams. Chauncey: Sorry, Frank, you were gonna say something. Frankcx: No I was just gonna say The Vote Frankcx: let's put it to a vote. So my, my s- thesis is that openness is positional, not princled. Not principled, sorry. So the question would be is how many of you believe that or you believe more that, there's other reasons as to why open weight models exist? So let's go around the horn. Frankcx: Jacob Jacob: yeah you're it is positional. Like it it's for competition reasons I think Frankcx: Robert Robert: I'm gonna take the altruistic view and say the open part of this is the important part, right? It's to it's for the community. So it's, I don't think it's for greed or positioning. I think it's it's for the bigger community to drive drive innovation. Frankcx: Sure. Chauncey Chauncey: I'm too jaded. I'm going with positional[00:30:00] Frankcx: Positional. And then kneel Neil: It's positional. They want highest IPO valuation possible. Keep those GPUs coming, keep those engineers Frankcx: Yeah Neil: I, I think oh-- I think Anthropic would have a much different stance if they were third or fourth place right now Frankcx: It makes me wonder what is in it for the Chinese, in producing open weight models. Is it, I could be a conspiracy theorist and say it's to drive down the value of, AI at a frontier level by US AI. But I don't know. That's a pretty, that's a pretty spicy Robert: that's a whole different... now you get into national security concerns and Frankcx: Sure Robert: all kinds of other stuff, right? Like what ki- Jacob: But all, all of that matters, especially, maybe some stuff to talk about next time about how these models have been breaking out of their harnesses, have been breaking out of their sandboxes and like that, like the... We're at this point pretty soon where things are getting Frankcx: the other thing that was in the news this week was that OpenAI [00:31:00] put a two-week moratorium on the advancement of a particular model that they're working on because it was the same model that broke into Hugging Face and stole some cheat codes for an exam it was asked to Robert: It actually broke into my house while I was out the other night too, and stole some jewelry Chauncey: Drink some of your beer. Frankcx: there's an Robert: Yeah Frankcx: there. Yeah. All right, let's On My Device: Model Reviews Frankcx: let's move on to on my device. This is the, And I'll start off by talking about I, I'm, I was looking at all these different open weight models that were out because a bunch of them came out from essentially Meta produced this model called Glimmer, which they essentially said was faster and more powerful than any other open weight model out on the market. And on my thirty ninety that came out not to be true. But what's interesting is that if I thought of this as like a job interview, I'm gonna just give a quick rundown of five candidates that came for this job interview and the job that I would give them. So the first one is Gemma with-- from Google. Frankcx: This is the model I would hire. This is the fastest model that is out there locally [00:32:00] on the thirty ninety or even on a RTX Spark device. It is three times faster than anyone. It uses a twelfth of electricity on my device that any other model would have done. It is a daily driver. Google has excellent model in Gemma. Frankcx: It's an open weight model, and it probably would get the job nine times out of ten. But then there's Qwen 3.6, which is the previous model that was out there. This is a model that takes a ton of time making sure it gets the right answer. When I was giving it an exam, it took almost seven times longer. Frankcx: It took thirteen hours to answer an exam Chauncey: Oh Frankcx: essentially other models answered in under an hour. It really takes its time, and I think that was one of the reasons that Qwen came out with 3.8, because when 3.8 came out, 3.8 actually was eight times faster than Qwen 3.6 in my testing. But it actually it also used a ton of what's called scratch paper to do its work. Frankcx: It uses a big amount of memory and a lot of da-- or a lot of essentially tokens in order to do its work. And [00:33:00] while it's faster, it uses just a ton of space and tokens in order to get its answer done. And it actually used a tool that it wasn't supposed to use. It was kinda like the way I tested these models is I gave them an exam, but I also gave them kind of a real-world like scenario and it-- and I found that Qwen actually, 3.8 actually broke the rule one time that it shouldn't have. So I thought that was interesting. So it's its sibling. So I think the way they sped it up by, was by giving it, a little bit less accuracy, where Qwen 3.6 isn't gonna give you an answer till it's absolutely sure of it. 3.8, it will give you an answer a little faster. Then there's Glimmer from Meta. I... It's a careful reader. It's the only candidate that noticed two sources contradicting each other in the test and it was, but it was also last of the five on the exam. It just didn't, it didn't pass the exam as well as the other models did. Google and Qwen both did better at the exam than Glimmer did in that space. Frankcx: And then finally, there's what's called Mistral or Devastral. And this was a surprise. It actually [00:34:00] had the highest score on the exam but in real-world scenarios, it was the worst because it has the smallest memory in the field. And it's also it just doesn't have a l- a long memory towards understanding things, so I wouldn't give it a job. I'd give it a job as a chat Robert: May-maybe as an intern Frankcx: it... Yeah, as an intern or chat model type of thing, but I wouldn't give it... so it's interesting. I'll put all these results up on GitHub, but I thought it was interesting the results I even did on my own 3090. And that's the cool thing is that you can download these models, and you can put something like Cloud Code or Codex to work in order just to have it implement these models and test it on your own device. Frankcx: You don't have to take the word of other people out there. All right. What is everybody else working on? What are some cool things on your device? Robert: I, I'll go first. I 'cause it, I think this dovetails nicely into that. Forever I've been trying to figure out h-how do we bring coding, code generation to the device? And you mentioned Claude Code and Codex, and then I mentioned LM Studio [00:35:00] Bionic today, which is which is supposed to support some of that. Then you've got more traditional kinda IDEs. That's where I'm really struggling, is to figure out, what is that recipe is the model and the actual program or agent to do that. I haven't figured that out yet. That's been that's been the bane of my existence Frankcx: Sure. Like like a coding agent that runs locally on the device. I don't know, Chauncey, I know you gotta run. You wanna jump up next? Chauncey: What am I actually running locally? I'm not running a ton of stuff locally, shame on me. But I will say that I'm definitely digging into like Microsoft's answer to Claude, so Scout and just the unreal power that thing brings as a marketer. But even just the ability to have something that works alongside me. Chauncey: I, I actually use it less for like building material for me, but more as just like a, a admin to a certain degree, which is like part of that vision that folks have w- around agents is like it's supposed to be there to help you [00:36:00] day-to-day. I will have more to share about that later on. Chauncey: But yeah that's my thing I'm running right now. Jacob: I got Scout running locally this week. Chauncey: You, of course you did. You and I are gonna have to chat about that Jacob: and I also... Frankcx: you should tell us what Scout is Jacob: So s- so think of the, the ClawPod rage from, January-ish this year. Just local claws, local agents being able to do things. Scout is just Microsoft's current first one. We announced it, I think, at Build. There were some internal builds before that were pretty cool. But it's just, I, it's a very interesting tool that gives you access via MCP to stuff on your device, but also gives you access to, your cloud provider models that, in my case, my organization pays for. So we've got-- I've got access to all the latest models. Jacob: Chauncey, Chauncey: Sponsored by Microsoft. Jacob: when you're doing Scout? What's your favorite model right now? Chauncey: Oh I'm a big fan of the 5.6 Sol, but I was on Claude before that. But 5.6 Sol is pretty, pretty solid Jacob: Robert, Frankcx: I go back Jacob: what's [00:37:00] your best best cloud model right now that you're choosing? Neil: I'm having some fun with Groq 4.6. Chauncey: Ooh, all right Frankcx: It's like Neil: one. Frankcx: I always call Grok like your drunk uncle Neil: I'll go back to Scout for a moment. So you asked my favorite model. Choosing the right model for the right task, I think is super important. Neil: We're seeing some models really good at PowerPoint creation, right? Other models better suited for coding. To Robert's point I have been fixated on, all right, how do we start bringing some of this local? So I've actually built a harness in Scout. Every time I enter a prompt, I do a classification. Neil: "Hey, does this need to go to a frontier model, or can this inference Chauncey: Oh, cool Neil: I ran a report. So I've been an avid Scout user for a couple months now, just a couple months, and Scout told me that eighty-six percent of my inference is running locally. So clearly, I am not a frontier developer for most of the use cases I'm building. Eighty-six percent of that, of those inferences are running locally, [00:38:00] equivalating to sixteen hundred dollars in savings already. Frankcx: model that you're running local? Neil: Great question. Five four mini is my favorite multi- Chauncey: Wow, really? Neil: model. Also from a a prototype perspective, I really like Qwen anytime that math is involved. Imagine like an insurance agent claim, take a picture of a damaged vehicle. You can use a model like five three point five vision. But then if you're doing some calculations I'm getting great results with Qwen from purely mathematics. If I want a multi-model, a multimodal model, five four mini is my favorite right now. Robert: It's interesting 'cause I use Scout to build out, and you can get this on GitHub, my dream team, which Chauncey: Oh yeah, Robert: agents that Chauncey: set that up Robert: stuff for me on a daily. So I run it in Edge browser developer, it uses Ion 1.0 for a bunch of stuff, both the planner and the what's the other one? Robert: Execute or whatever it Jacob: Instructor, is it Robert: Thank you. Yeah. So it [00:39:00] uses it uses a bunch of those along with Phi-4 Mini when it needs to Jacob: gotta, I've gotta go down that path. I've had CLI build me those decision trees, Neil, at times where it's like, "Hey, use this local model some- sometimes," but I haven't been able to see benefits, and I've seen regression in some of my work, so I need to go and rebuild that. but I have run-- because Qwen 3.8 27 billion, which I've only been using for the last week and a half or whatever, but because that model is s- I think so smart for its size I've actually testing that as an orchestrator, and I know that it's not as smart as a frontier model, but it's fascinating how good it is at doing tool calls, at actually being able to make the decisions for, calling my work IQ MCP or doing something on device or en-enacting a skill to build a PowerPoint, and I've really been enjoying that. Jacob: [00:40:00] So I'm not an expert in that yet but I've had success in Scout and Copilot CLI as my harness Frankcx: it sounds like if there's something for everybody in the audience to do is go check out Scout, because it brings that governance and that understanding of what OpenClaw was doing, but at a Microsoft level. So yeah, very cool. Gentlemen, Wrap Up Frankcx: how about we do it a wrap here? I feel like we've drained the topic of open weight models and, but I'm excited. I love this podcast. I love you guys. I think we've covered the gamut from talking about all the different pieces. But hey, we'll have a new topic for next week. The plan is to put one of these shows out every week. We ran a little long today, but we'll try to come in right around h- half hour for everybody. Frankcx: Hopefully, you can start your week listening to this podcast. You can find us at thelocalhost.show, and Robert's art is coming, but the terminal that is on there is out there forever. There is no place like 127.0.0.1, gentlemen[00:41:00] Jacob: Thanks, Frank. Chauncey: Thanks, Frank. Frankcx: Bye for Chauncey: Bye, guys Robert: gang