Slightly Caffeinated

TJ Miller and Chris Gmyr compare the AI code review bots they are each building at work, from fixture repos of real PRs to requiring a line of code behind every finding. TJ also shares the eval-driven wizard he built to turn source material into learning content, sharing insights on building AI systems you can test and trust.

Links
References Mentioned
  • Hive - Chris's orchestration framework
  • Gorillaz - the show TJ saw in Detroit
  • Deltron 3030 - the opener and TJ's all-time favorite hip hop artist
  • Luma Brighter Learning - where TJ is building the learning content wizard and the PR review bot
  • Claude Opus 5.5 - the model TJ wants to move the content wizard to, then rerun the evals
  • Prism - TJ's Laravel AI package, which needs updates before the model upgrade
  • GitHub Copilot - the reviewer that caught real issues TJ's bot missed, and the inline comment style Chris wants to copy
  • CodeRabbit - one of the review services both teams passed on in favor of building their own
  • Amazon Bedrock - where Chris's reviewer runs to keep code inside their own infrastructure

Creators and Guests

Host
Chris Gmyr
Husband, dad, & grilling aficionado. Loves Laravel & coffee. Staff Engineer @ Rula | TrianglePHP Co-Organizer
Host
TJ Miller
Dreamer ⋅ ADHD advocate ⋅ Laravel astronaut ⋅ Building Prism ⋅ Principal at Geocodio ⋅ Thoughts are mine!

What is Slightly Caffeinated?

Join Chris Gmyr and TJ Miller as they dive into the world of PHP, Laravel, and all things programming, while also sharing insights on family life and other musings.

TJ (00:00)
Alright, welcome back to the Slightly Caffeinated Podcast. I'm TJ Miller.

Chris Gmyr (00:04)
And I'm Chris Gmyr

TJ (00:05)
So Chris, what's new in your world,

Chris Gmyr (00:07)
man, just cranking on the tokens. Cranking on Hive, all the projects, all the projects that work. I feel like that's my typical. I feel just like I don't know. I'm in my like AI bubble during the day and just like just cranking on stuff and then look up, it's like the end of the day and where where did everything go? And like I have meetings and stuff that throughout the day that breaks it up, but it's like, man, I like Hive just opens up like the

world of potential for working on too many things.

TJ (00:37)
Yeah.

Chris Gmyr (00:38)
Like it's it's been super helpful, but also it's like, is this too much during the day? But it's it's been pretty good. Got some scout stuff going on. he's going on an overnight this weekend and then we have our big fall camping trip next weekend that

TJ (00:52)
Yeah.

Chris Gmyr (00:53)
we're doing. So that'll be fun. Get outside, get outdoors, touch grass, all that. yeah.

TJ (00:59)
yeah. Yeah, have you

guys been hit with the leaves changing and big fall weather shifts yet?

Chris Gmyr (01:06)
not yet. It's been a little bit chillier in the morning, so that's been nice. It has been, you know, like eighty five, ninety degrees for a while. but yeah, it'll be coming. it's kinda weird down here because during the summer it gets so hot and we have like so much drought usually that like

TJ (01:24)
Mm.

Chris Gmyr (01:25)
leaves start changing and falling like in the middle of summer. 'Cause it's

TJ (01:28)
Ha ha ha!

Chris Gmyr (01:29)
just like it's just dead. so it's a little weird, but like we have like summer fall and then

actual fall when it gets

TJ (01:35)
Yeah.

Chris Gmyr (01:35)
chillier. yeah.

TJ (01:37)
Yeah, yeah, we're starting to see the shifts here. think the other morning I took my son to school and it was 45 outside. So it's we're definitely getting the shifts here. Yeah, like it

Chris Gmyr (01:45)
Risk.

TJ (01:48)
was it's only 60 now. So

We're definitely starting to see all the leaves change and everything. We've already done one cider mill trip for the fall. Got the donuts, the cider, gonna be doing that

Chris Gmyr (01:59)
nice.

TJ (01:59)
again pretty soon. Cause now we're out. I like, I just, live for that, those fall, fall vibes. Like gotta have the apple cider, gotta have the donuts, gotta have the caramel apples. So I gotta, gotta do a restock trip here pretty soon.

Chris Gmyr (02:13)
Yeah, yeah, totally. So yeah, what's new in your world?

TJ (02:17)
man, I feel like I've been in a weird funk the last couple weeks. I've not been working on much outside of Luma. Kind of just like doing chores and stuff around the house. You know, with the kiddo back at school, I'm trying to do a better job this year of like being on top of them academically.

So that's been a bit of an adjustment, just trying to make sure that like when he gets home, we check his portal for homework and like make sure he's doing all of that. And it's been a different shift in my routine coming out of the summer still. So I'm still kind of adapting to that. I'm such a creature of like routine and habit that once it gets shaken up, like it takes me a little while to settle back into things.

Chris Gmyr (02:59)
Mm-hmm.

TJ (02:59)
so yeah, we, we did like cider mill stuff. that was super fun. we went and saw gorillas the other night, which was

Chris Gmyr (03:07)
Okay.

TJ (03:08)
such a, such an amazing show. one of my favorite artists, deltron 30 30, they opened, this is my third time seeing him. Absolute hands down. Like I think my all time favorite hip hop artist. and,

Then the second opener was someone I had never heard of, but my wife's been listening to for years. It was a UK-based artist, so first time playing in Detroit and everything, and that was super cool. It was really cool to kind of see the crowd go from nobody there. It really seemed like their set started, everyone was kind of like, who is this? We don't know, so low hype.

But then by the end, you could tell everyone was all on board, everyone was super hype and the artists got a little choked up on stage about it, which was really cool. And then of course, like, Gorillaz just knocked it out of the park, which was a real treat. So like, I remember when Gorillaz came out and it was like super underground and like I was into them right off the rip. And so it's an artist I've always wanted to see. I could...

I could critique the show about some stuff, but I was entertained, I enjoyed it, I had a lot of fun. And

Chris Gmyr (04:18)
Yeah.

TJ (04:19)
it was cool that we purchased tickets separately from my sister and her husband, but they were like, they ended up being in the same section of us, a handful of rows up. So we did a double date, did dinner beforehand with them.

It was our first time doing a double day together. was a real treat to get to spend time with them. So all in all, that was good,

Chris Gmyr (04:41)
That's cool. Sounds awesome.

TJ (04:43)
So not outside of Luma, not burning as many tokens as I should be. And I feel like that's about to shift here pretty soon.

Chris Gmyr (04:50)
That's all right. Give the robots a break. They're tired.

TJ (04:52)
Yeah,

right? And they gotta breathe a little bit.

Chris Gmyr (04:54)
Yep, yep, totally.

TJ (04:56)
Yeah, so moving on from that coffee stuff, I don't know what we picked up. We finally got through the Costco stuff and we're on to another bean. It's not really my cup of tea. Like it's more on the fruity side of things. I like it, but I've not.

I've just never been like a fruity notes kind of coffee person. I'll drink it. It's good. It's caffeine. I need it. But I'm definitely kind of looking forward to getting out of these beans and into something else.

Chris Gmyr (05:24)
Yeah, yeah, totally. Cool.

TJ (05:26)
And then,

yeah, nothing, lots of hot cider. I've been doing that, you know. So I'll have my coffee and I'll have some cider in the morning. Yeah, just I'm all about it.

Chris Gmyr (05:31)
Yeah. Those fall pipes.

TJ (05:35)
I don't do the pumpkin spice, but.

Yeah, cool,

Chris Gmyr (05:40)
Yeah, for me. haven't had like anything too crazy. I did do like coffee and co working the other day with a coworker, which was fun. So I went to a different

TJ (05:48)
sick.

Chris Gmyr (05:49)
coffee shop and ended up getting like a harvest like chai. So it had like a whole bunch of different seasonings in it and like I like chai, like also.

It's

TJ (05:59)
yeah.

Chris Gmyr (05:59)
like a chai latte type of thing.

TJ (06:02)
Yeah.

Chris Gmyr (06:03)
So yeah, it was super good. so I got

one of those when I was there and two to go home for me and my wife for for later.

TJ (06:10)
Nice.

Chris Gmyr (06:11)
so it was definitely more of a dessert than like a coffee, but like there w

TJ (06:14)
Sure.

Chris Gmyr (06:14)
was at least some caffeine in there, so I'll take it.

TJ (06:17)
I do like myself a dirty chai every now and then, like for sure. And you know what, maybe I'm gonna make a trip today. You know, my son's got a half day, I'm gonna be out and about. Maybe I'm gonna go get myself an iced dirty chai. You know, that sounds really good actually. Yeah. Cool.

Chris Gmyr (06:30)
Yeah. Go do it. Treat yourself. Cool.

so last couple of times we talked about a whole bunch of my projects going on, but I wanted to circle back to you and some of the Luma projects that you've had going on recently.

TJ (06:46)
Yeah, so like on the the project front I haven't been doing a lot on like my personal projects But I've been cranking a ton on things over at Luma You know, I've built a ton of AI systems so far one of a couple projects recently We've been doing like we're We're like an education platform primarily for

the logistics industry around like OSHA and other regulations and we can do learning platform stuff for even like HR onboarding as well for like just onboarding into different things inside of companies. And I've been building a pretty robust like multi-step wizard for creating our learning content.

So trying to build a process that encodes our secret sauce of like what we put into like information design and all sorts of aspects of like learning, like our CEO has a PhD in learning. And so like I'm trying to build a system that applies a lot of those lessons that she's learned and like.

get a brain dump from her and integrate it into this. So you can start off with like providing a topic. Like this is what I want to build. We call them eNuggets. And so it's like, this is what you want. I want to build it about. And then you can like upload PowerPoints, Word docs, PDFs, big field to just like paste text even.

And then we like break that down and then put it through our information design process. And the interesting thing about this project is it's now going in two directions. We've got one for internal use that is way more in depth and way more hands on. like more of a hybrid AI assist rather than like really AI driven.

and this has like a lot of our internal glossary, like our internal phrases and words for things, like domain language heavy, but, you know, it, we were really getting like super strong results. So it's like, you, you do a step AI does a thing and you do another step AI does its thing and you kind of have these different approval gates as you go. But we're now also taking on.

part of the project is to empower our clients to do something similar. if like we may do your regulation content for you, like you have us, you request us to do it, we do it for you. And then that lessons available for your company. But maybe you don't want to pay us to do your HR onboarding, but you still want to as part of the platform. So

This other version that I'm kind of building is geared towards a little bit heavier AI automation. So it's like, give it a topic, upload your thing, and then we automate a whole bunch of steps that in our internal tool are actually like little bit more hands-on. So that's presented all sorts of really interesting challenges around like, what do we trust?

what do we trust the AI to do? What do we really need a person to be like hands on making decisions? Lots of, right now I'm like building out tons of like artists in command evaluations to like be able to kind of like, let's make some changes, run the evals, like go back and make more, use the informed things that come out of the evals back into like iterating on the

the prompts and the tool calls and all of this stuff to just try to squeeze out higher quality. But then also those serve as almost acceptance tests after the fact to make sure that like any changes we make moving forward, you know, that we are maintaining the same level of quality. Great example, I've built it using Sonnet 4.6, Opus 4.6,

or 4.8, and so I want to pretty soon here update to like Opus 5.5, Sonnet 5.5. I've got some prism implications that I need to work on in order to make that happen, but I want to, after updating those models, run the eval, see where we're at. Did we get a big quality gain?

out of that or do we maybe want to run a little bit older model that's maybe a little less expensive? we can play that game a little bit too.

Chris Gmyr (11:16)
Mm-hmm.

TJ (11:17)
So that's been, it's been really interesting building out this because it's...

I've learned so much more about building durable and testable AI systems. I've applied a lot of what I've learned, but I've learned a lot going through this too. So that's...

looking forward to next year and wanting to do some conference talks, I definitely see a lot of things that I can pull out of this to, you know, build a talk on like, you know, building durable AI systems and like, how do we test and evaluate those things? So that's been pretty cool, very in the weeds and a lot of back and forth with...

the learning team and like our CEO since she's got all of the knowledge on like the learning platform and the science behind learning that we're trying to encode. So that's been really fun too because I just haven't, I most of the time work on my own little island. And so this has been really fun getting to know and integrate a lot more with the rest of the company and teams. And then.

The other big, big project that I'm working on, and we kind of touched on this in an earlier episode too, is trying to lessen the review burden. We've got a Laravel application that is what we're calling like our back office Laravel app. And as part of that, whenever pull requests get opened or updated, it web hooks into this system. And then we have a bunch of what I'm calling lenses.

look at the PR through, has an agent look at the PR through these different lenses and evaluates that pull request. And then like post a comment feedback and then ultimately gives it like an approve, a request change or just like a comment status. So like it does some sort level of approval on the overall PR. And that's been.

that's been really challenging trying to figure out how to get it to leave quality reviews. Like it, in the past, it was leaving apps like factually wrong stuff. It would hallucinate things and that's maybe even worse. And so I kind of fixed all of that, but then we were getting these PRs where Co-Pilot was giving valid feedback on like here's like

some critical issues, some like warning issues, and they were very valid, but then our bot was coming back with like no feedback. So it's like kind of tuned it too far in the other direction. And so I'm finding it really challenging to figure out how to handle that in a nuanced way where like.

it's not giving feedback just for feedback sake, that it's actually finding things that are valuable because we would love to get to a place where that agent can call out X, Y, and Z things need to have human review on them, but everything else we can just rely on, you know, this review bot very confidently having feedback or not feedback on the PR. And then ultimately it's end status that like if it

we'd like to get to a place that if our code review agent approves the PR, we have the confidence to like maybe even auto merge that. That'd be a great place to be. the interim goal is to just try to lighten that review burden of like have humans put eyes on the things that humans need to have eyes on and then really lean on this thing to make it.

it's just a better experience because we've talked about, and more quantity of PRs are happening. PRs are getting a little bit bigger because the more agentic coding that's happening, that's just kind of a consequence of those things. it's been, it's not even like technically challenging. It's all of the prompting and like,

figuring out how to give it guidance of what's quality, what's not quality. And one of the things that I had it build is historically pull in PRs. And then myself or engineering manager and maybe one or two other engineers can go in and actually

give feedback on these things so that we can use that to feed into the system so that we have good example, actual examples from previous PRs of here's things that were called out. Ultimately, this would be an approval or a deny. And so we can have that data to improve the system and create this kind of like golden set of, you know, PR and result.

and then we can use that as part of the evaluation process as well. we're bringing, I'm building this like, yeah, like tagging in feedback system because I think the only way to get it is to maybe provide those examples of like, here's what's worth calling out, here's not, know, here's stuff that doesn't matter. Yeah, it's so nuanced and it's been.

really, really tricky to know. Like I said, technically, the code and the architecture and all that, not the challenging piece here at all.

Chris Gmyr (16:26)
Yeah. Yeah, totally. I know we talked about this a little bit before, but basically we're working on the same project in

TJ (16:33)
Yeah.

Chris Gmyr (16:34)
two different companies. But yeah, so I I don't remember how much in depth we went to into last time, but yeah, for the fixtures, I have a whole nother separate repo outside of the PR reviewer like GitHub action.

repo to pull in any of these fixtures. So I went through because I've gotten a bunch of feedback from other people around the company in different tech stacks, different teams. And it's like, hey, this PR is just, you know, the review is just not good. And I made comments and it didn't do anything, which it's expected because we weren't pulling in comments and stuff like that at the time. It's just like a

TJ (17:08)
you

Chris Gmyr (17:09)
fresh eye review, like knows nothing about anything really besides like the diff that we pushed in it.

so I've been able to capture, I don't know, probably twenty, twenty-five specific PRs from across different teams and tech stacks and and reads of like review and comments and all that stuff, and have a bunch of scripts to actually like pull that stuff in, see the whole diff of the changes and all the comments and the different rounds of reviews and all that stuff to take into account and especially from the humans.

commenting back to like the AI that like wasn't actually connected to the comments, but at least that was like some human feedback like in those loops. So very similar to what what you're doing with your engineering manager and like other people on the team to get that like kind of human the in the loop feedback of like, yes, this is right or this is wrong and getting all that calibration. So brought that all into a repo. And then I've been doing additional tests.

and

tracking with the AI reviewer. so kind of baseline of like your reviews can't get worse. So here's like the the diffs and reviews of like the baseline. So we did a lot of baseline testing to get kind of one-to-one match what was happening for those rounds and PRs and diffs and stuff like that. And then like as we iterate moving forward in this kind of like eval testing framework harness changes

Is like, okay, let's tweak this and let's make sure that we're you know, matching to the correct like if we're calling something out, point me to the line of code that you're actually like pointing out and having an issue because that fixed a lot of the hallucinations with like issues. So it's like, well, if you're gonna call something out, validate that you can actually point to the thing that you're calling out. And if you can't, then just drop it. So just that call out, like increased.

the correctness of the review probably like 15 or 20 points depending on the the PR you know on average. so

TJ (19:09)
Mm-hmm.

Chris Gmyr (19:10)
like things that. we're making sure that we're pulling in like the Claude or like agents MD, any sort of like nested claude and agent MDs, making sure rules are pulled in and aggregated into the reviewer as well. So

I've been trying to make context window changes and adjustments and making sure we don't, you know, blow out that window. trying to do it in like basically one shot, you know, as well. So everything stays fast and cheap. so lots of iterations happening. But what I want to do next, so probably into next week is mimic a little bit more of what Copilot is doing. So like individual comments on things instead of like one main comment on the PR.

And then we'll have emojis or thumbs up, thumbs down, along with comments that humans or agents can do back to the AI reviewer. And then that will funnel into like the next round. So we'll keep some sort of like aggregate file of changes in like a ledger of what was flagged as like incorrect or like what we fixed, and then pile all that into

round two's context and then round three's context and ideally getting smaller and smaller reviews of you know importance and not going too far into that like super critical adversarial round of reviews and just finding stuff to find stuff. We actually want like a green light that this is good and you can be okay to merge it. So

TJ (20:35)
Yeah, I

really, I really like, so what I had done before is I had it do rule matching on the glob patterns, just how like Claude has it. So it was only adding those to context as needed. And I had something in there before of like, your review has to be rooted in rules. So like, kind of like you have the lines, you have to be able to point this back to a rule, like, and it has to be like a violation of a rule. I really like that. I, I'm.

maybe going to go work with it to integrate this today of being able to concretely point to a line or lines in the code. This is validating this. It has to actually point back to real code that exists. The other piece that we're doing is we're all in on Anthropic and Cloud Code as an organization for doing our agentic development. So I'm now bringing in OpenAI.

as adversarial validation of things. I'm not doing a review with Claude and a review with OpenAI and then using the union of what they found. I'm basically having OpenAI come in and validate the findings that Claude had. So.

Yeah, so that's I think I want to go in add the the line attribution to the pull request because I think that's gonna be super valuable and Making it a little bit more concrete, but that was something that was funny that came out of it was Claude was like is it like asked me more or less if a Valid code review is like approved. I'm like, no, it's absolutely a valid state You know, like it

Chris Gmyr (22:08)
Mm.

TJ (22:09)
you might not find anything. It could be a good

good pull request. And so, like, yeah, no findings is valid, but like, I'd maybe be a little suspect of that, especially in the beginning. But the other, the bringing in the open AI, I think is gonna be really strong too. You know, is just finding additional blind spots to what Anthropix can have. So, yeah, we'll see how that goes. I'm in the weeds of it right now and...

I've spent a lot of time iterating on it, but we want to get to a place where we can heavily rely on it. you know, the time spent is worthwhile in getting it right. know, quality over quantity for sure.

Chris Gmyr (22:50)
Yeah. Yeah. Exactly. so yeah, the the agent side is moving. And then I've also tried to make some movements on the human side of reviews as well. So I've been making a bunch of changes to our engineering skills and there's one in there for like PR create. So if you do this, it'll like open it up as a draft, it'll assign it to you. it'll do like come up with its own title and description, link to a juror ticket, move it, all that stuff.

So previously, it was pretty verbose, as AIs, you know, do. no one was really like reading it or anything else like that. So I'm like, okay, how can we make these like PRs, you know, at the surface like better for humans to kind of crok? So I got rid of like as much of the pros as I could. And now I have the agent basically call out like super important changes that a human should be actually looking at.

in linking directly to that file or set of lines. So it could say like go to this you know security dot js file and like here's what you're looking for type of thing. And it'll limit that to like the top two or three items per PR. So we're trying to stay within like small to medium size, you know, PRs. So hopefully two to three will be you know good enough items in there. And then it'll

It'll say below that just in a sentence or two of what you can basically gloss over. It's like, yeah, here's a session log, you don't have to worry about that. Here's like a readme small change, you don't have to like

TJ (24:16)
Mm-hmm.

Chris Gmyr (24:17)
about that. So kind of like two human review, these, you know, two to three things. Here's what you can bypass, and then it kind of follows the rest the the team's PR template as per usual. And then at the very bottom.

We have I added agent instructions for the review. Yeah, all in a collapsed like details box. So it's

TJ (24:38)
Yeah.

Chris Gmyr (24:39)
collapsed by default. So like a human would have to click on it to actually like read all this stuff. And that's where like all the other pros is from the agent who's like implementing the the PR and branch. So ideally all of that will be put into the AI reviewer and workflow that it's a

okay, here's some really pointed things that we should look at in the review. And here's some like AI based notes for, you know, everything else. So hopefully all these changes, the AI feedback for the AI reviewer, having that iteration with humans in the future with like comments and emojis and and things like that will help. So I mean, we'll it's but it's a it's a lot to to put in.

TJ (25:16)
Yeah.

Yeah, we went and made a config file that lives in the root of the repo that has basically glob paths for must have human review for like auth and other like.

migrations and stuff like that where it's like that gets called out automatically in the PR like the agents comment on the PR of like these things have to be reviewed by a human and then if there is anything that needs human review that blocks it from auto approving the pull request because like then that means a human has to have the final

So yeah, it's so funny that we're working on the same project and like things that already exist, Copilot, CodeRabbit, like, know, the services are out there, but we want it to really reflect our process and the things that we care about. So even though we've like looked at other products, we've just, we keep coming back to, we really want to build this in-house because...

there's certain things that we really care about in certain ways, and we'd rather keep our workflow and all of that together. even though these things exist and they're pretty good, we're choosing to build our

Chris Gmyr (26:32)
Yeah, yeah. Same same reason here. it's just so much cheaper to build our own using our infrastructure stays within AWS bedrock for privacy, you know, reasons. And again, we can tune it and kinda do what we want with it. and it's so much cheaper than like paying for copilot or even code rabbit or anything else

TJ (26:49)
Mm-hmm. Yeah.

Chris Gmyr (26:51)
like that. And like I'm getting like fifty cent reviews per round. So I mean like a whole PR might be two to three bucks, but it's like that's

relatively inexpensive compared to some of the other options out there, especially for like a team of hundred, hundred twenty plus engineers doing

TJ (27:08)
Ours aren't that inexpensive yet. Ours, I think I just looked at one this morning that was $40 or so spend on it, but that was across a PR with like 12 pushes. So it's, you know, it's reviewed at least 12 times. So it adds up, but I'm trying to bring that cost down a little bit, being a little bit more efficient and then, you know, also bringing in open AI, like that's got

different cost profile too. we'll see how that works. But on that note, you want to wrap up? All

Chris Gmyr (27:35)
Yep, let's wrap up.

TJ (27:37)
right. Thank you all so much for listening to the Slightly Caffeinated podcast. Show notes, including all of the links and social channels, are down below and are also available at slightlycaffeinated.fm. If you have any questions for us or have content suggestions, go to the Ask a Question page on our site, and we will feature it on an upcoming episode. Thank you all so much for listening, and we'll catch you next week.