Agents and Engineers

Dan and Jonathan Bown open with the talk Jonathan gave at ODSC, "Practical Agent Ops: From POC to Prod with MLflow 3.0." MLflow 3.0 arrived last summer as the first stable release built for generative AI rather than traditional machine learning, and Jonathan's team used it to build an agent for pre-enrollment students. The centerpiece of that work was evaluation-driven development. Instead of jumping straight into a working prototype they aligning the business up front on what quality actually looks like before signing off on a model with inherently non-deterministic output.

The initial key to success was an Excel file. In it, the data science team had already assembled 150 ground truth examples, but left them untested and set aside while engineers focused on code. Jonathan's team paused the coding work and ran a simple foundation model against those examples first, landing at what amounted to a coin flip of useful versus hallucinated answers. From there they refined the examples with the business, loaded them into MLflow's evaluation datasets built from live traces, and iterated by versioning prompts and agent configurations.

Tooling came up repeatedly. MLflow's open source repo now ships a skill file that plugs into coding tools like Claude Code, which Jonathan called a game changer for keeping up with an API that changes at roughly a release a month. The Databricks AI Dev Kit, released around March, bundles skills for the Databricks SDK, CLI, data engineering, and analytics work, usable either inside Databricks' Genie Code pane or in outside tools such as Claude Code, AWS Kiro, or Google Antigravity. Jonathan said installing it produced a dramatic jump in output accuracy compared to coding assistants working from stale or incomplete context about Databricks and MLflow APIs.

Dan raised the idea that LLMs and agentic tools are becoming users of software in their own right, alongside humans, and Jonathan tied that to broader changes at WGU: more of the business, not just engineers, now writes system prompts and builds their own copilot-style agents. His own day to day has moved from core development toward AI enablement, meaning security review, best practices, and helping non-technical staff adopt evaluation-driven habits for the prompts and agents they build themselves.

Jonathan's path to WGU ran through Pentara, a biostatistics consultancy, and Zions Bancorporation, where he did quant finance work before a stint simulating financial products for WGU students. He became a founding member of WGU's MLOps team in 2023, when the university's machine learning was still traditional work like random forests and ensembles for predicting student outcomes, well before Databricks had built out MLOps tooling. Dan connected this to Hamel Husain's essay "The Revenge of the Data Scientist", and Jonathan agreed that evaluation-driven development brings the work full circle: checking evals and correctness is the generative AI analogue of checking a confusion matrix.

The pre-enrollment agent's rollout became the clearest illustration of the method. The first release, a bare foundation model with no WGU context, drew heavy negative feedback from the employees testing it, some of whom wanted to cancel the initiative. Jonathan's team treated that feedback as fuel, folding the failed questions into an evaluation dataset and iterating until they reached roughly 82 percent correctness and near-total relevance, at which point the same employees became enthusiastic supporters. He credited MLflow's architecture for building subject matter experts directly into the agent ops workflow rather than treating evaluation as a purely technical exercise.

Jonathan was candid about where his trust runs out. He does not trust a tool's first output even after a full planning session, citing a Kiro planning cycle from the day before that failed on the first try despite extensive back and forth. He is cautious about MLflow's fast release cadence outpacing its own skill files, and notably guarded about tools like OpenClaw and Claude Cowork that can reach into email or personal documents. Given how much effort WGU puts into protecting student data, he extends the same caution to his own personal information and limits what such agents can access.

On his team, Jonathan resists banning AI-generated code or stigmatizing it in review, and instead pushes everyone toward reviewing code outside their usual specialty, using AI review tools like Amazon Q or GitHub Copilot as a starting point rather than a final answer. He pushed back on the idea that tool usage equals productivity, warning about AI slop and noting that some of the heaviest users he knows are not the most productive. The thread ties back to evaluation-driven development's real thesis: start from value, not from the tool, a point he illustrated with WGU's Academic Virtual Assistant pilot, where a surprising result showed that students chatting with the assistant were more likely, not less, to still reach out to a human mentor afterward.

Full episode notes

Chapters

  • (00:00) - Introducing Jonathan Bown
  • (00:56) - ODSC talk: escaping POC prison with MLflow 3.0
  • (05:11) - The forgotten Excel file: rebuilding around evals
  • (09:36) - MLflow skills and the Databricks AI Dev Kit
  • (15:00) - When AI becomes the user of your software
  • (18:16) - How the day-to-day has changed in six months
  • (22:16) - Centralizing prompts, evals, and best practices
  • (25:38) - From quant finance to founding WGU's MLOps team
  • (31:02) - The Revenge of the Data Scientist
  • (35:01) - Has the job gotten easier or harder?
  • (37:33) - Thinking ten steps ahead with agentic coding tools
  • (42:15) - Leveling up junior engineers instead of gatekeeping review
  • (51:17) - Where trust breaks down: OpenClaw and personal data
  • (56:48) - The mental toll of managing agents versus writing code
  • (59:27) - How much detail agentic tools actually need in a prompt
  • (01:07:14) - Value over software: the Academic Virtual Assistant's surprise result

Links from the show

--------------------

Guests

-------

Jonathan Bown, Principal ML Engineer, WGU

Follow the podcast

-------------------

Follow Dan Gerlanc

-------------------

What is Agents and Engineers?

The podcast about Agentic AI and Software Engineering. Each episode is a conversation with people whose daily lives most intersect with AI and agentic systems. Join me as I follow the stories, the behind-the-scenes, and the real people behind the code.

I'm Dan Gerlanc, and welcome to Agents and Engineers. Today we're joined by Jonathan Bown. He's Principal ML engineer at Western Governors University, where previously he was an ML ops engineer and before that a data scientist. Prior to that, he worked as a statistician at Pentara. a quant analyst at Zion's Bancorp, and a software engineer and associate instructor at the University of Utah. Jonathan, welcome to the show. Great to have you on today.

Awesome to be here, Dan. Thanks. Thanks for having me. It's an honor to be on the show.

Yeah, so the first thing I wanted to talk to you a bit about is you gave a talk recently at ODSC Practical Agent Ops from POC to Prod with MLflow 3.0 So do you wanna tell us a little bit about that?

yeah, it was was very exciting to get accepted and have the opportunity to present there. I'm a big fan of ODSC and and what they do I I kinda heard this term like POC prison and I felt like you know, as we were evolving our agentics stack at WGU, there was an opportunity to share some of the learnings that we had as we Went through a lifecycle, an agent development lifecycle for kind of mid summer to end of last year. And MLflow had last summer come out with their you know version 3.0 stable release, and that was sort of their stepping into the generative AI stack and making MLflow work better with generative AI tooling and agent development because prior to that it was there was a heavy emphasis on traditional machine learning and for those of us that were trying to get generative AI workflows to work with MLflow there there was there was a challenge there. So at WGU I've kind of a I've I came from originally as a as a data scientist and kind of a consultant in that space and then went into MLOps. And we were building a platform for WGU to use with Databricks. And so we're using tools like you know Databricks, CLI, SDK, and then MLflow that have heavy integrations with each other to build a an ML op stack for WGU and for the data scientists to use. And The generative AI workflows just didn't really fit in with the traditional machine learning stack. And it just, you know, it was it was very rigid. It wasn't flexible for what we needed. And so they released 3.0 and we were really excited because now we could take a lot of these workflows that were kind of stalled and plug them into the ML ops platform. In my team, we were heavily focused on building a an agent for pre-enrollment students. And just found MLflow to be a d a joy to work with, MLflow 3.0. And it really unlocked and you know they they emphasize a lot this term called evaluation-driven development. With an agent development lifecycle, which is a deviation, I think, from traditional machine learning and even traditional software, where you're kind of jumping right into kind of a working prototype and then you're iterating over that and then ultimately getting kind of a sign-off from the business to say, hey, this this model that has inherently non-deterministic outputs is good to go, we're satisfied with the the quality of of results. And so we were using MLflow to kind of align the business around what what is a a achievable quality standard for this agent. And then we were focused on you know building the ground truth examples, validating those, running prototypes with the the the employees of WGU who actually answer questions for students during that that phase and just had a really great experience with that. We we achieved a really good quality standard and we also kind of we we also exceeded the expectations I think of the business in terms of what we were able to put out there. And so that the talk was okay, how did we how did we actually do that? What were the key components of MLflow that allowed that to happen? And how does the focus shift from just trying to get code working to actually achieving like a high quality agent result throughout this kind of new newer agent development lifecycle? So that was the premise

So would you say ML

Yes. Yes. So I mean, when we when we started the process last year, we we came into kind of a a data science a data science team that was already working on it. And they were working on it I would say in in the sense that it was kind of like a a traditional model, like that you would build maybe in like a a quant finance or like a data science kind of traditional machine learning. area where you're you're trying to get things working and you're trying to kind of plug different things together. and we came in and realized that there was this Excel file of like 150 ground truth examples and it was kind of left off to the side. It was it was not being tested against, it wasn't being validated. But the business had put a lot of time into generating that for us. And so we said, hey you guys you should pause on the on the like the actual coding And let's like actually shift over to the the this Excel file, which kind of sounds weird, but the the focus should be the evals. And what we did was kind of run an initial test and say how how well is just like a simple foundation model performing on these. and then move to, okay, we have this kind of baseline level of quality, which was originally kind of the flip of a coin, like it it half you know, half the time it it would generate something useful, half the time it would it would generate, you know, hallucinate or or generate garbage. And so we worked with the business to to like refine those examples and make those better, make those more achievable. And then we were able to plug that directly into MLflow because they have these now these evaluation data sets where you can actually take traces that come into the model. And you can just like you know checkbox them and say, I want these to be part of my evaluation data set. And then you can schedule runs or you can run manually on top of those. Ever every time you you iterate, you you can version your prompts, you can version you know the agent and and consistently measure progress as you're you know plugging in say something like rag or or you know, messing with the system prompt or messing with retrieval or something like that. And another thing that I I think is is happening too it in in general in the industries, I'm seeing I saw this a few times at ODSC actually where they they were saying, you know, you should move left to right in these in these agents, left being kind of just the most basic like just a completions call to a foundation model. And then maybe you upgrade that to a workflow and then if if that works well, then you move over to like an agent. Like you don't need to necessarily start with an an actual agent. and so that was kind of also the approach I took in the talk where I started, okay, let's just start with a basic completions call and kind of practice like some some basics with MLflow Let's get it set up. Let's let's get some initial evals plugged in and then okay, let's see where the foundation model is clearly clearly not working, which We were building a an MLflow assistant, which is very easy to show that it's not working because most foundation models aren't gonna have like the MLflow 3 That show, hey, this is clearly failing. It's it's picking up it's not picking up something that's clearly there. so and then kind of moved.

in the latest APIs for MLflow 3 things like that.

Exactly. And that's kind of even been my my eval as a developer is how well does is a chat experience or an agent coding tool doing with MLflow because you know, we use MLflow a lot and that's been one of the the hardest things and you know there's there's usually always gaps least in the documentation or you know things that the model the foundation models pick up on based on MLflow history that aren't necessarily accurate and they've even just recently put on a like a MLflow AI assistant onto the docs and that's been really useful because now I can like easily validate you know APIs and stuff like that. But but yeah it's it's very easy to generate ex

So is that something like an llms.txt that you can add to the context or a skill that you use with MLflow when you're working with it?

Yeah, so now what's great now is they a couple months ago they released MLflow repo actually has a a skill file that they that you can plug into different coding tools. And so that was that was a game changer for me when I was even developing the the material for the talk because you know, now I can just pull the skill into my local Claude code setup and I can yeah, you know, you know, generate some some code or some things that could you know, can enhance the material. Whereas before I was just having to like look everything up and and try and aggregate material that way. But but yeah, so they have skills now like that are In the MLflow open source repo. And then they also have Databricks has AI Dev Kit, which also has those skills as something you can integrate with your coding tool of choice. And so that's been a massive improvement for me. So AI Dev Kit is it's actually two different things.

What is the AI dev kit?

So we're we're in like I said, we're in Databricks a lot at at WGU that we're a Databricks customer. And they have a something called Genie Code, which is if you go into Databricks, there's a you know, a a pane you can open on the right hand side and you can chat and you can generate code and you can and that's been there for a while. but like I said, they just recently released MLflow skills. And they also just recently released AI Dev Kit, which lets you pull in skills that are very specific to Databricks SDK, Databricks CLI, Data Engineering, Data Analytics, all the Databricks products, you can pull skills in to either Genie Code or your coding tool of choice like Claude Code or you know, AWS Kiro, Antigravity, whatever. And You can actually develop Databricks, like very context-aware Databricks components, models, apps, different things, in in your tool of choice. And it's I I've I've noticed a dramatic improvement in the outputs just from walking through a very simple install process on that on my on my local machine, as well as you can you can plug those in, like I said, directly to Genie Code. So you're in Databricks. And you can actually enable specific skills, like I can say I want MLflow I want serving endpoints or whatever, and it's it's it's a dramatic improvement in outputs and the work that you like the back and forth that you have to give that coding assistant to actually get something that matches the API spec. so it's it's a really cool tool and I'm I'm kind of learning that it was only released I think in March. So we're all kind of learning that, figuring that out. It's really great tool.

So those are both the Databricks-specific SDKs APIs as well as specific to like data engineering or data science versus MLflow is Databricks has a hosted version, but there you could also do it open. There's the open source. Or I think MLflow is open source, so you can run it yourself, right?

Yep. Yep. yourself and it's kind of there's different paths you can take and I I'm I'm exploring all these paths because like for example you can set up your MLflow server. I didn't show this in the talk, but you can set up your MLflow server and you can connect the assistant in MLflow, which is you know, MLflow over time is looking more and more just kind of like the Dataverse UI, which is which is cool because it it just feels natural. but y you have your assistant just like you do on Databricks and you can actually connect your Claude code into that and then you can chat inside that MLflow server and it has access to say like agent traces or evaluation data sets or or system prompts, whatever. And you can you can get insights directly in that coding experience and they have tools behind the scenes that are making that more more relevant for as opposed to just using the terminal for Claude code with a skill installed. So there's there's a bunch of different paths you can take with it. And you know I find it all just kind of really it's all really enhancing the experience to to generate you know code and and and working things much much more quickly. Whereas before it was flip of a coin if it would even give you an API or or a set of you know code that was even accurate to the even, you know, the last year of versions. so that's one of the reasons I think I've been a little more hesitant to jump all into one tool 'cause I've lacked that context and I've had to provide that context. But the the tools like that and the alignment around skills I think has really been a been a huge boost for that.

Do you think this is a good example of I've seen people write about this that the next stage is AI or LLMs being the actual user of a lot of software versus historically you designed for a human user and now you also have to take into account that your user is agentic coding tool.

Yeah, that that's that's a really interesting point. And and I yeah, I I think that I mean all these tools, at least the the Databricks AI dev kit is coming out of the field engineering side of of Databricks. And so, you know, I'm not I'm not sure how much they're designing it. I mean, they're obviously designing it to have LLMs consume the skills and things like that, but they're also like you know, they're building things and they're finding you know, they're they're bumping up against walls with just the traditional tools, saying, I c I can't even use cod coding tools in my work, for for accelerating some of these things. and but it but it is it it's interesting 'cause it is kind of an internal product to them. Like they're it's it's technically open source because they they've opened it up, but they're not like

MLflow. Yeah.

Yeah, it's not like ex you can't just go to contribute to it. Like you have to meet like a pretty high standard to contribute to it. it is more of a like internal development tool, but yeah, it's kind of straddling that line of like, okay, there's customers that we have that are using LLMs to consume our libraries and our products and build products on our platform. And then, you know, we're also building products for customers. And so I think I think there's some good stuff coming out of that kind of know, joint application of those skills and things.

Yeah, for sure. I think it's definitely and get gets into the whole MCP or CLI, the interface, and or you just let the model use the SDK and write code. all all options.

Yeah. Yeah, I yeah, I I I'm I'm of the perspective now that I just kinda want it all. Like I I I'd I'd love you know, I you can when you walk through the setup of AI dev kit it it lets you kinda check the boxes of like which, you know, do you want data engineering, you want machine learning engineering, do you want data analyst and I j I just always choose them all 'cause I'm like I you know, we're as a machine learning engineering team, we're serving many different users at at any given time. data scientists, we're working with data analysts, we're working with data engineers, ML ops engineers. so we need the whole I need the whole stack. Like because if if I'm you know I could be building a Databricks app for just even a a consumer of data or a you know someone who's like in Databricks writing code and and actually building workflows on LLMs.

a lot has changed in the last six months. I mean how has your workflow day to day coding and working changed?

yeah, I I would say it's it's a product of a few things. It's a product of the tooling, obviously, like like we've talked about. It's it's changed in the sense that I have access to more context for these tools so I can be much more productive using, you know, just an off the shelf fork of VS code, for example, Cursor, Antigravity Kiro And as opposed to just kind of being a little hesitant and being like, is this gonna have the right context? Am I building something in Databricks? Am I building something locally just gonna have like a Python app? if I'm building a local Python app, I would six months ago I would have been much more open to jumping into like an AI tool, agentic tool, much more hesitant on Databricks, but now it's kind of a like let's just Let's just use it for both, make sure I have the right things plugged in. on the business side, I I I think there's been a more dramatic change where everyone else also has access to more AI tools. And the business is more open to people having that access. and so there's a lot more people that are not just consumers of like a chat experience, but actually, you know, writing system prompts, for example. Like We had an example where you know someone that interacts directly with students was building their own system prompt for s something like a copilot agent, right? as as opposed to just like chatting with a with a particular solution and and generating some output and and consuming the output. And so there's a lot more demand for in I guess AI enablement, like how do we enable people to be better users of AI? That could be from a security perspective, could be from an engineering perspective, could be from like, hey, you're using a system prompt. Maybe that belongs as like a skill in a broader ecosystem. That's like something I'm thinking about a lot more these days. but the day-to-day is a it it's a lot more interacting with the business and trying to understand what they need, as opposed to just we need a an AI chat experience that doesn't expose student data. That's kind of where things started out, I think, for the the AI engineering side or the the ML engineering that we've been involved with has started out and is now more of a okay, the business is gonna use AI. So how do we enable that? How do we make sure everything's secure? How do we make sure they're following best practices? How do we make sure they're incorporating, say, evaluation driven development in their process and their workflow? Like if you're just writing a system prompt, how do you how do you know that there's there's good quality behind that? How do you know that you're it's following your instructions every time? How do you know that the the outputs are high quality? So, you know, I'm I'm doing a lot more kind of evangelizing, I would say, of of these tools and best practices and as opposed to just like core development work where I'm, you know, just like looking at PRs, I'm I'm I'm writing PRs, you know, looking at our our AI stack and and trying to figure out what features we need. It's more like, okay, what does the business need? Is this is this stack even what they need or do they need something else? Do they just need an API key, for example, so That's that's been the more dramatic shift, but in addition, the you know, agentic tooling is becoming much more democratized and context aware and it's it's becoming also a different way of working. So learning that and and trying to apply that in a sophisticated way is is more of what I'm doing day to day.

And how are you centralizing that knowledge or training? Are you doing it as Databricks and MLflow provide a consistent way that people can store their prompts or evals? Is this still someplace where you have to come in and say, Okay, we're gonna start with what you have, but then we kind of take it over?

Yeah, and I I I think th I'd like it to be the latter. I and and not in the sense that I wanna take your you know, take the IP away from someone and say this is o we own this now. But it's more like, hey guys, there's a better way to do this. Like, you know, instead of writing a mile long system prompt, like let's let's cut let's cut down on that. Let's maybe we maybe we create a skill or we create you know, we plug in a Pydantic model for a structured output or or we we do a lot of things with the current, you know, open AI SDK for example. so I'd like it to be more where we're we're taking, you know, what they have or we're improving it and we're prov we're making it easier for them to to also make it better themselves. we're working closely with Databricks because they they have amazing solution architects. that are kind of like almost always like on call. Like I have a I have a chat with them and I just kinda message them and say, hey, we're having this issue or there's this use case or I have an idea for something and they're they're ready and willing to kind of jump on and and collaborate on something. So we have a lot of users of Databricks and a kind of a growing ecosystem of that platform at WGU 'Cause that's where our data is ultimately. it's not it's not necessarily like, we w we want to just evangelize Databricks. It's like, no, that's where our data is. And so let's let's work close to the data. Let's use MLflow for some of these ML ops type or you know, agent ops types workflows. Let's work with the business to understand what quality looks like and And then let's bring in the context from our data to then personalize the experience for students. Ultimately, I think that's that's one of the larger goals that we have as an organization is not just to use AI, but to personalize it for the student and and use use the context that we have about them, their progress in in their journey. to improve outcomes for their learning journey. So whether that's earlier graduation or earlier assessment passing, and then they can move on to the next one because it's it's a company competency based model. So they they progress kind of at their own pace. And so we want to make sure that we understand what their pace is. And there's a lot of different you know, people involved in that and and there's machine learning models and there's different predictive analytics that goes into that. And so how do we leverage that in our AI stack and make it, you know, a much more like unified experience? Those are those are kind of some of things that we're working on with with Databricks in particular.

And so, I mean, I think you'd said that you started more like many of us on the data science side, traditional ML side. Is that what prior to Gen AI world, a lot of the work you were doing around kind of traditional regression classification problems for this kind of this kind of problem?

Yeah. when I so my my first work with WGU was we were in I was working with I was consulting actually. So when I went to Pentara I was I was a consultant. It is a biostats company, but I I kinda went over there with a friend from from Zions. We were both in in the financial, you know, quant finance world. and went over, wanted to try something new. And then we ended up just like we we had another friend from from Zions that was at WGU. And so we were right back doing finance just with WGU because Pentara wanted to, you know, kind of break into some new territory. And so we were doing that. And we were working we were doing just l like basic simulation, I would say. We you know, some some small regression models, but more kind of like

Mm-hmm. Mm-hmm.

You know, we were looking at distributions of financial returns and you know it it was this kind of more of a startup environment. So we were looking at different financial products to again help improve student outcomes. And are there learners that you know we could reach with some different financial alternatives? So we were piloting some of those and we were doing a lot of simulation, a lot of projections and reporting type type stuff. and then that that group kind of got spun off into something else and I had the opportunity to kind of say, well, where do I want to go? And WG was very supportive of me kind of finding the right path. And I had always been interested in like machine learning was kind of my my side project ever since even I, you know, grad school and and undergrad, just you know, the machine learning class and and the computer science minor that I did was my my passion. And it was still kind of niche at the time, but I was always doing kind of side projects and and different things, Kaggle competitions, and and then so that that group got got spun off and they were founding a machine learning operations team and had the opportunity to jump in on that and be a founding member. And still very like the the the machine learning that WG was doing at the time was was very traditional, right? Like, you know, random forests and ensembles and a lot of different things to measure, you know, and predict outcomes for students. And so we were using like very specific Databricks tools to accomplish that. We didn't have an MLOps platform at the time. Databricks didn't really have a solution for that at the time back in 2023. And so we had to kind of like the AI tools that had come out, you know, in 2023, 2024 still were missing that context. And so we were still very much like a traditional like software engineering team. and then, you know, I I would I would work directly with data scientists. I would kind of put my ML engineer hat on and help them improve their workflow and get it in production. and then You know, when the the AI craze kind of caught on, it was definitely like, Okay, now we need solutions for AI usage at WGU and you know I I had had that experience as more of like, you know, traditional software engineering from MLOps and and was able to jump over into that. But I would yeah, so the workflow has has changed from okay, the the industry was kind of machine learning was kind of off to the side. It was maybe part of an overall product or solution. Like you would you would predict outputs and then those would be fed into something else. Like at Zion's we used, you know, random forest to like help with variable selection for just a regression model. but then yeah, generative AI transformed that into being just full on, you know, agent ops, LLMs, OpenAI, SDK, you know, completions, responses, agents and and things like that. So it's but but again, the evaluation driven development I think kind of brings it full circle because you're you're still kind of in the same way that you would be look checking your confusion matrix for a classification model, right? You're checking your evals, your correctness, your tool call correctness, your safety, all those things. You're you're checking those as you iterate. And and so I I think the it actually makes the machine learning workflow much more relevant. if if you're used to that sort of benchmarking and and evaluation type type workflow. So that hopefully that a that answers your question. It's it's definitely changed a lot even in the you know, s seven years since I graduated with my first masters, so

Yeah, that's one I think Hamel Husain has a article or blog post, The Revenge of the Data Scientist, which is the skills for working with LLMs and actually evaluating if you're getting what you want to be getting are what has been done in data science for a long time.

Yes. Hundred percent. Yeah. Yep. But it's it's i it's still this I I find that it's kind of this enigma to people and it's probably because data science was kind of an enigma, but you know, it was much more approachable from like the you know the measurements and the outputs. And now you're dealing with indeterministic outputs and and someone's actually interacting with the product. Like people weren't interacting with machine learning models before. Like maybe it would trigger something that you know, trigger recommendation for them or trigger a particular experience in in the tech that they were using, but they weren't actually interacting with the model, right? and now they're like you can take anyone a I can I can message any employee and say, Hey, can you try out this new agent that I've just deployed to our our, you know, AI gateway, our AI API that you can chat with and you can you can see for yourself, you know, and and tell me. And what was really interesting about that process is with our enrollment agent is you know the first the first go-around, like your first release, say you just you start on that left side and you release just a foundation model and they start asking it questions about WGU. I mean, we got so much negative feedback because it was like, this is not what I want, like this is this is hallucinating like crazy, it's not giving me WGU context. and we're like, it's okay. Like, you know, th th they're they're approaching it from the perspective of I I think we need to shut this down. This isn't working. Right? We're we're not gonna pursue this initiative. and we're like, no, this is this is part of the process. Like, trust me, I I love getting the negative feedback because then we can take the questions you asked and we can plug them into an evaluation data set, and then we can run the gamut again and we can we can make some improvements, we can make some some tweaks. And by the end of that process, when we were getting like, you know, 82% correctness and almost a hundred percent relevance, the feedback was completely the opposite. I mean everyone was so excited about it and they were like, this is awesome. This is exactly what we want. and so that just you know There there's the data science side of it where you are kind of like you're you're working with your stack and you're and you're checking the MLflow outputs, you're checking the the metrics and things like that, but you're also incorporating the SMEs into that that pipeline because you're also validating with them and MLflow actually has this on their kind of main architecture. They have the SME as part of the the workflow. for agent ops and so you kinda have to start with them and say what do you what do what do you think is good quality and then you come together and you build maybe you build an LLM-as-a-judge based on what they the feedback they provide or you you know build better evaluation data sets and then you test again and you refine that and so that process I think is is much newer to the equation because you know you could as a data scientist, kind of work in the background and s say, okay, we do have this objective, kind of more objective measure of success. And let's let's try and reach that. Let's try and get ninety-nine percent precision or or recall or whatever. so that the part of it where you're actually talking to the business and you're and you're talking to non technical users, that that's definitely a new a new thing that we're we're all kind of figuring out.

And do you think that's made the job easier, harder? A bit a mix of both.

A definitely a mix of both. the I w it it is a diff it's kind of it's leaning on different skill sets, of of mine that I you know, I I've used, but like, you know, when you have to go kind of workshop with people and you have to, you know, explain technical concepts to people that aren't working in AI. Maybe there's people that are hesitant about AI. There's people that have all sorts of opinions about AI. maybe s some of that are negative. and so you've gotta like do a lot of relationship building as part of the process. And that is something that I think sometimes engineers just don't really wanna do. And that's not a bad thing. It's it's just kind of something that we don't really think about as we're like going to school or we're like studying and we're we're learning things. you you've gotta do a lot of you know, back and forth and talking to people and and there's also a lot of solutions out there, right? There's a lot of third party tools that they can just go grab off the shelf. Maybe they can sign a contract with fairly easily. and so it it is a little bit more competitive in the sense that, you know, you have to communicate to them why going through this kind of painful process of evaluation driven development is worth it. as opposed to just saying, hey, there's this company that claims to have it all figured out for a particular use case and you know, you have you have to kind of sell things a little bit and be like, well, you know, that's great, but, you know, let's see what we can do. Let's see what we can do with student context. We have the data, we have Databricks, we have all these things. so there's the sales aspect that's kinda come into like me as the principal role that I have to I have to kinda sell things a little bit. I've I've become a little bit more of a salesman in the sense that, you know, I'm trying to leverage best practices, trying to leverage, you know, security and and and make sure we're we're always following like the WGU guidelines on AI. We're protecting student data, right? It's the law. So there's a lot of stuff like that that I'm having to do that wasn't traditionally part of part of the job of the engineer or the data scientist or things like that.

Yeah, I've talked with folks about how s in some ways, since the cost of generating code has come down, that the part that becomes more valuable is actually solving the problem. Or it lets a lot more engineers be in a position where you can say what's the problem and then you don't need a huge or as large of a team to build the solution as you might have used to. So it actually ends up pushing the work to the higher level of what's what's the problem we're solving versus being able to focus purely on the implementation.

Yeah, a hundred percent. so a couple a couple thoughts came to mind while you're you were describing that. the the the agentic coding tools place more emphasis and are actually better when you think ten steps ahead instead of one step ahead, right? Like if you are in Claude Code or AWS Kiro and you do slash plan. Which I'm I'm s you know fairly new to, but have, you know, seen some YouTube content and stuff recently where they're like, that's the and the creator of Claude Code even talking about it, where it's like that's that's how you get really good outputs and and you get something kind of on the first try that's like potentially works right off the bat with a higher probability of success. so thinking about a user story. as almost like its own product because you know if you if you're relying on like project management or or product to come in and and tell you what the entire solution should be without that technical expertise it's much harder for them. so you kind of have to as the engineer you have to kind of step up to a higher level and say, you know, I have I have a much quicker I have a much more enhanced ability to get something from start to finish. I can do it much quick much more quickly. what does that finished state look like? How do I plan this? Like a and and as you're planning with these tools, like it's asking you all these questions and you're like, I'm not used to thinking like this, just pulling a ticket out of the you know, my Jira board and saying, like, I'm gonna I'm gonna build this. I I usually have like a scope and I I have an acceptance criteria, but, you know, as you're now more capable, you can think further ahead. And so that's been an exercise for me to think just beyond like whatever the latest like this this core feature, this core implementation is. and it's it's everyone's definitely at a different level with that. It's not a it's it's hard to get, you know, an entire team to do that, I would say. But I I'm I'm definitely trying. I'm definitely trying to get my team to like use these tools more. I find that, you know, while there's like a lot of hype around them in in on YouTube and on, you know, social media and things like that, the adoption rate is actually I I feel like much lower than maybe we would expect. because it is a a big shift in how we work and that can be uncomfortable for people. And so like When when you see these low adoption rates, I was listening to a conversation with a professor from NYU who decided to create a course where he was showing students how to use these tools. he found that something like eighty percent of his class, two hundred students, eighty percent of two hundred students hadn't even used one, hadn't even tried one before. and I think there is a real hesitation to like shift that that way of working because it is it is a very different way of working. There's a lot more There's a lot of overhead when you're first kind of starting out. Maybe you want to plan with a tool, something like that. So but overall I think I think the shift is good. It it it at least in my little world, it it empowers me to be more proactive on the product lifecycle and think, you know. I I'm really only limited now by my imagination. What can I think up? What can I dream about to have as part of this these products and these these experiences for people inside outside WGU? So

And I th think another challenge is since these tools can accelerate the work so much that we almost at least I know from my end, need to step back and make sure that I understand areas that I have less experience in. And so I as someone who's working with a team of different skill levels. What I guess what do you as you have the team starting to use these tools, what do you tell junior developers about learning the fundamentals when in theory you could have an LLM, Claude or Kiro do the work for them?

That's a great question. so I definitely am not taking the approach where I'm saying throw throw everything out and and start fresh. I think that might be some people's approach. but what I am trying to do, I think 'cause it it you know, when you're reviewing a PR, it's not always clear if someone, say, generated it or they wrote it themselves. And I don't really want to make that a big deal. Like I don't want to put s a stigma around that and be like, this is AI generated. That's that's you know, I think there's some taboo around that a little bit in the engineering community. but what you want to do is is level them up on their capability to review and evaluate their own code. which was still relevant when they were just writing their own code, right? so As there's more PRs now than ever before, for sure. trying to level up the ability for everyone to review everyone else's code. it when we when I started on the team and we weren't using really any tools, it was okay, I'm gonna ask this specific person because this specific person worked on this code before. They they're a code owner or they are very knowledgeable about this feature. And me coming in not really like you know, I wasn't a founding member of this particular team. And so I wanted as a principal, I wanted to learn everything. So I was I was trying to review everyone's code and I'm like, well, why isn't everybody reviewing everyone's code? You know, we don't we don't have to require that. We don't have to require that every engineer is is hitting approve on every PR. But we want to give it make everyone feel like they have the opportunity to review, even if they don't really know that particular aspect of the code very well, even if it's a completely different language, if it's like a front end and you're asking a machine learning engineer to evaluate that versus a full stack. I'm trying to level up everyone's ability to review and feel empowered to review and leave comments. And using AI powered review tools, I think, is is one way of doing that. Not not necessarily saying I'm gonna have AI approve this PR and and merge it, but like I'm gonna have I think I think we have we have Amazon Q and we we had GitHub co-pilot it for the majority of last year where you would just say, Hey, I'm just gonna trigger the automatic review from this AI tool and and it would generate you know, s mostly some good stuff. Sometimes it was kind of out l out of left field. But like use that as a starting point. Like I I've I've become much more reliant on that in the sense that if I don't really know something as well I'm gonna jump into that AI review and I'm gonna get a perspective of like, okay, what are maybe some of the the initial kind of problems with this? And then where can I start my review from there? And so yeah, I would I would say building every feel making everyone feel empowered that they can review code. even if they they're not like fully aware of that. And ultimately that's also helping everyone level up their skill set, right? If if you can read code, if you can read jump from Python and jump into React and j you know jump into different like libraries and modules, open AI SDK versus Databricks SDK, things like that. I think that ultimately levels up the engineer anyway. even if they are, you know, generating or ri using agentic tools 'cause You know, the skills, the planning, the context of I know that this I know that OpenAI library or SDK does this particular thing. So I should be able to build, you know, this this tool that does this particular action. Like that knowledge I think is what's kind of coming out of and that that practice and that skill set is what's coming out of all of this. Is you're you're much more able you're much more quickly able to connect the dots with different tools and then, you know, let let Kiro or Claude Code go wild with it and then and then evaluate it. Don't just push it. Don't just create a PR off of it. Like definitely go in and read it and make sure you understand what's going on.

Yeah, I definitely agree. You can if you want to learn these things, you can in a way faster than you ever could before, because you can be like, Hey, explain this line of code to me which before you'd need, you know

Yes. Yep. Yep. Why is why is there a security vulnerability in this in this code? Like what what is this? Yeah, exactly.

Yeah, especially if you're working in an area where you might not have like a ton of front end experts, right? If you're depending on a data that a data science or ML engineering team, if you you're not in a large org where you have folks who really know React, then that getting that knowledge would be Hard without spending hours reading the docs, right? Or going and find that person. Yeah.

Yeah. Absolutely. Becoming a full stack engineer. Yeah, yeah. Exactly. Yeah, so it's just it's just all about getting everybody involved. We we have a really awesome UI kind of expert now that's d doing some of our front end stuff, in in conjunction with our full stack. And it's been really cool to get him more involved in the lower level details because now, like you know, you're even as someone who's not traditionally like an engineer or coder can come in and start to like even contribute, right, to the to the repo. Maybe they can, you know, you can teach a few th little things on GitHub and they can start contributing to it. Not that that's like what we what we envision for every single person on our team, but like they feel more empowered to actually jump into the code, right? And start like looking at what's actually going on, how are my designs influencing what's being done here? so it's it's making our team I I feel like much more collaborative and and leveling n there's not there's not so many silos, right? Where where people are getting s siphoned off into different parts of the of the repo, which I think is is problematic in in a lot of ways, but we're all kind of participating in this experience together and the the path to get something out is is shorter, or at least to prototype something.

Yeah, I think it lowers the barrier to entry in a way that before it just wouldn't be efficient because you can only remember or be an expert in so many things at a time. So that was that was always that's my that was always my excuse for not getting too involved in front end stuff that I'm just gonna forget all this.

Yeah. Yes. Well and that and that's where that's where the AI usage started for me, right? I like I I felt kind of an attachment to like writing my own Python. But like if I was if I was you know, in ML ops we were building a lot of GitHub actions. So, okay, now you're in YAML, you're in Bash, you're in you know, these these pre configured actions that are on GitHub and that was a new world to me. And when you're especially when you're designing your own actions and you're like, okay, I know what I want this to do, but like, you know, there's there's all these different ways in shell to make something happen. There's different ways to specify your variables and and, you know, pipe things around and, you know, rather than go and learn all that. by hand it was like, okay, I'm gonna rely on this, on ChatGPT or Grok or or whatever the chat tool of the day was. and just focus on things that I don't have the time to really go figure out all on my own. and that's where I kind of left it for a while and and then definitely evolved over time to be more of like a as as these tools mature, as they get better context and better knowledge bases, it's it's definitely a much better experience.

So there's a lot of positives, but where are the places that you still don't trust these tools?

Yeah, I that's that's a good question. I I mean I'm definitely trusting them more and more, but I I'm I'm not trusting the initial output. Like even if I go through plan mode, like it's tempting to be like, Okay, this works. But as I even experienced just yesterday, like after a a long kind of planning cycle, back and forth questions, first try didn't even work, right? So, The first output is where I'm like, okay, this is like let's be really careful here because I don't I don't know that this is gonna work and it probably won't work. I'm also if someone's just generating outputs and not using like a skill or a particular like context, aware type tool, and I I can usually I can usually spot that. and that's where I I'm Issues and then that's where again I'm trying to up level the the community here to be like, okay, there's there's better ways of going about this, right? There's better ways of you know generating a a notebook in Databricks or a tool calling agent in Databricks, things like that. so I see I I guess I I see a lot of the same patterns for things that are not all that sophisticated, and that's where I I kind of lose my trust a little bit because I'm like, okay, this probably wasn't iterated on very much. It was just kind of generated and stuck there and then they moved on to the next thing. I'm also I'm I'm still still fairly you know somewhat hesitant on the MLflow side because like it changes so fast. So like they're they're doing at least one release a month and you know, th they would have to update the skill really quickly too and and things like that. And there's also like, you know, the OpenClaw stuff is becoming much more popular. MLflow is now kind of piloting a an OpenClaw integration with you know a package on the MLflow repo that actually has a direct OpenClaw connection and you know giving it giving it act giving OpenClaw access to like personal information, I'm still very like guarded on that. I I I haven't really touched a lot of the, you know, say Claude Cowork plugins where you can connect your email and and some some of your more like personal information, your your documents, things like that. I'm still very hesitant on that. I don't trust it from more of a security side, I would say. just because I mean at at WGU we've put so much time into guarding students' data and making sure that that data doesn't go it and exist on someone else's server. and so I I feel the same kind of attachment to my personal information because of all that work. so not not plugging it into my email, not not plugging it into really personal documents or things like that. Strictly like, okay, I'm gonna create a project on Claude Cowork and I'm gonna I'm gonna choose which links or which files it gets as opposed to just like letting it letting it go wild or letting OpenClaw go wild and letting it text me and, you know, do these things. I'm still very hesitant on that. But, you know, it is a really cool technology. I like I'm I'm fascinated by it. There was some good content on it at ODSC and so I'm definitely fascinated by it, but still still hesitant to like go wild with it, I would say.

Yeah, I think once you're tying it up into any external data source or I mean even emails, right? Anyone can send you an email, so once you're injecting context that you don't control, then it's kind of the wild west, right? They can start telling the model to do something that

Yeah. Yes.

you don't want it to do, especially if it's on your own personal computer.

Yeah, and and even building building products even in Databricks or or outside that like, you know, an application that adheres to WGU best practices. Like maybe you can plug that into more of like a planning phase, which I think would be cool to to kind of try out, and maybe have that as a skill. Like you're a you're a WGU product builder and these are some of the the best practices we have. But trusting it to build something that is in line with our best practices, that's that's one angle where I definitely like I come at outputs for the first time and I'm like, okay, is this, you know, is this using any external APIs? Is this using any external models? Are we using, you know, internal models, internally hosted models and accounts, things like that? so the AI usage kind of governance side as well from a from a business. From the organization standpoint, that's that's still a concern.

This is a slightly different track, but managing agents I think is definitely a different cognitive mode than writing code yourself. How does that feel over a full day, a full week? Like are you do you feel more tired at the end of the day, less tired?

Absolutely.

Tired different.

I don't know. So I I'll draw on the experience from yesterday. I mean I I was going through a planning phase with with Kiro. I'm still kinda new to it. so it it required some constant attention, right? Because you it's asking you questions and you know, it's giving you options. So you can specify like A B or C or you can choose to provide more context. that kinda wore me out, honestly. even just that one kind of planning phase that I did yesterday. maybe it could be that I was just tired after a long day anyway. 'cause, you know, I'm I'm in meetings and, you know, d context switching all around. but but yeah, I mean, you've you've gotta be like and and these questions come out and you're just like, I hadn't even really thought of that. what is the right approach there? I don't know. And then and then you go through all that process and you're like, okay, yes, trust you know, trust this command, full permissions, things like that and and then and then it it says, Okay, I'm done, and then you've gotta go test it. And then you test it and it doesn't work. So then you've gotta like say, Hey, here's the error. Wha what's going on? and then it you kind of enter that loop again. it's not working again. Okay, it's a 429 error. It's a 500 error. you know what what what's what's the issue. but then you look at the code stack that it generated for you and you're like wow that that would have taken me a while. This is this is actually pretty cool. Maybe I have more ideas from looking at the code stack. So it is this yeah, it is an interesting shift of shifting from just like implementing to planning mode. And and even just prompting, like prompting separately, like if you're in the the AWS Kiro IDE or you're in Claude, you know, VS Code and you have Claude code in there, and you're just kind of prompt based code generation, it's it's still very it feels very different from that.

Do you find you still need to give detailed prompts beyond just here's a high-level set of requirements if you're building something? Or is it just simplified now, like here's kind of what I want to do, as much detail as possible, then kind of iterate or let it ask me questions.

Yeah, I mean I I'm I have a hard time writing like really long prompts. it's just I mean, I I like to write and I and I've definitely had to write a lot more as as the landscapes change and you know you're using these agen agentic tools and you have to write and kind of express your ideas more. so it's definitely a skill that I'm refining. And this was actually a topic that came up at ODSC as well. There was a talk from an MIT professor where he was like even if the agents get so capable that they can work autonomously, like you still need humans to have the skill to express those ideas in a meaningful way. Like it's not just a universal thing to express those ideas. So it's it's a skill that I'm refining, but I I do find that you know the It's a it's a balancing act 'cause you don't wanna give it too much and then and then constrain it so much that it doesn't have flexibility to kind of show you something better. so I kind of s I tend to stick more on the shorter side of of the prompting. especially especially in planning mode. I I mean my initial approach was just kinda to say, Hey, this is kind of my high level idea. here's here's like the environment I'm working in, here's kind of how I would like the experience to be. But let let's let's plan this out. and and and then I I I prefer to have it ask me questions as opposed to me having to like think of all those things. When I as a principal I'm doing so much context switching throughout the day. It's hard to really like get that focus time and and think through like a full kind of set of requirements that I'd want to give it initially.

Yeah, I I tend to agree that if you're working with like Claude Code or Codex they're pretty good at running the loop. Verse if you're building an AI workflow still, you might not want it to have to run like ten turns to get the answer. So then it the prompt still there is more valuable, but or you want it to be more detailed, but for work you're doing day to day, you're

Yes, absolutely. Yeah.

throwing a super powerful model at it that can run some arbitrary number of iterations. So it's it's gonna come up with something that probably will get you there.

Yeah, and that and that's that's probably another thing to call out with regard to the the trust aspect, right? Like trusting it to actually, you know, especially if you're working with a tool like OpenClaw, like trusting it to adhere to your like chosen token budget or like you know, your maybe your your subscription that you have or your API usage that you kind of have in mind. that's another area where trust is an issue. So yeah, absolutely with the AI workflows and you know, Databricks has some really cool ways of running these things in batch. and you know, I I definitely am like kind of like a zero trust with that kind of thing because I I've seen cost overruns from just some really kind of simple workflows. so definitely would you know, be much more cautious, like you said, with those AI workflows and and the token usage side. But but yeah, like if you're just building kind of an an app or something that, you know, someone could use AI with but it's not like part of the the the core workflow like it is definitely better to go through those step by step and kinda let it let it rip.

Yeah, you don't need to give everyone access to bash. Right. So they they might just need to be able to the look up what's what's the weather your weather app doesn't need bash access. It just needs to be able to make an API call to weather dot com or something like that.

Yeah. Yes. Yes, exactly. Yeah, yeah, yeah. Exactly. Yeah, web yeah, web browsing, stuff like that. Yeah, totally. It's it's kind of becoming like a s a staple in a lot of these things, but but yeah, absolutely. If it's using some more complex things it's it's worth like taking a step back. I like I kinda mentioned earlier, the adoption of these tools I think is among engineers I think is is probably lower than average. so my like kind of in just saying controversial opinion is like you should actually use them. I mean, you shouldn't necessarily outsource your thinking to them or like completely abandon all reason and just kind of let things work autonomously. Like you should you should definitely up level your ability to communicate, up-level your ability to review and read code and kind of quickly find insights or vulnerabilities. and then the other thing I would say is you know, I I think people need to be cautious about the thinking that a agentic tooling is one to one productivity, right? Like using this tool means more being more productive, right? Because, you know, we we still there's that term out there, right? The AI slop. And, you know, I I I do see you know, I tend to see that more often than I would like. And it it's y y I couldn't I couldn't really You know, the people that I I interact with day to day, it it's hard to predict who's using AI and who and who's not, right? Because people that are heavy users that like tell me they're heavy users, like I wouldn't have necessarily guessed that. Or people that don't use it at all, I also wouldn't necessarily have guessed that. There there's maybe they have high output, tons of PRs, tons of PR reviews, that sort of thing. so not falling into the trap of thinking that these tools are like a a perfect predictor of of and correlated to productivity because you could end up spending way more time with these things and, you know, refining outputs or refining the code as opposed to writing it yourself. I mean I've I've definitely had instances over the last year where, you know, I've been trying to use some of these tools and I'm like Whoa, like you're way off. Like I'm just gonna write this, you know, two cells on in my notebook that I need. as opposed to like five hundred lines of code. I think skills definitely help with that. So making sure you're kind of coming at it from like the planning, the spec driven development type type space. but yeah, th those would be kind of my maybe some opinions that people would would potentially find controversial or push back on is just, you know, I I think people should use them and learn them. I think it is moving in a really positive direction. And then also just being really careful about usage doesn't equal productivity gains. it's definitely helping people be more productive, but I think the people that are succeeding in the productivity space are the ones that are being much more intentional with their usage. Like much more confident going in that they're gonna get something high quality out as opposed to just kinda throwing something at the wall and and bringing something out.

Yeah, and I think this goes back as well to a bit of what we talked about earlier where actually solving problems that people have versus just building software are two different things.

Yes. Yes. And and I you know, I I'm hearing lately that, you know, there's now a AI AI generated code has made code much more of a commodity, right? There's a lot more code out there. And i it's not super clear whether that is a positive thing for engineering or a negative thing. But but you're right. It it it's sh it the emphasis is still on providing value. And I think that's the core message or the core kind of mission I think of evaluation driven development is you you're starting with value and what like what actually creates value. And then if you adhere to that system, then you're you're fairly likely to actually produce value in the end, as opposed to just throwing AI at something that where it doesn't really belong. So, you know, it's helpful to have someone come in and say, hey, this is what we would like you to do as a as an AI or an ML team. love that scenario. Love when there's someone creating all of our tickets for us and and saying, This is exactly what we need, this is this is the value that we want. but as things move more quickly, it's more incumbent upon us as engineers to actually focus more on that value and and maybe push back a little more and say, I think you're asking for this because you think AI is the value, but it's it's not. It's the it's the outcomes of that. Like, you know, if we're talking about at WGU, if we're talking about students, like students just using AI is not necessarily a valuable thing, but students using AI to we we published an article on LinkedIn, one of our senior executives, about one of the products that we built as a team, the Academic Virtual Assistant or AVA. it's it's on LinkedIn, it's on my LinkedIn. we actually one of the things we've been measuring while we've been doing this pilot is, you know, how is this AI usage, this chat experience they have with context, course context aware AI models, w how how is that influencing their journey as a student? And one of the surprising results was it it's actually increasing the likelihood that they call in to a a mentor. So it's actually improving their ability to the the the predictors I guess that would say this person's gonna reach out. Because maybe let's say they refined their question. Maybe they were nervous about asking and they, you know, went on in a session with Ava and they found, I actually do have a question about this. I'm gonna reach out to my mentor. So that was that was a surprising result. So it it can produce value. The value can be there, but you've gotta measure that, right? You've gotta have those metrics in place before you launch. Otherwise you're gonna miss stuff like that. And and then people are gonna come back to you and say, Where's the value in this? And you're gonna say, Well, it's it's it's right here. We we've been we've been measuring it. So you know, having a healthy level of of skepticism when someone comes and asks for an AI product, maybe it's not the right maybe they need to use AI differently. maybe they need to you know think about a a different tool or a simpler tool. Maybe you can offer the simpler solution. I often find that's the case. is we can, you know, provide like a lot of times simpler types of things and and get th the same or if not more value out. and then You know, it's kind of it's kind of on us as like AI experts and people that are working with this like day in, day out, to help drive the vision of it because I think there's a lot of like hype and you know you know there's a lot lot of companies are doing it, there's a lot of different kinds of products you can build. and so someone might see something and come to you and say, Hey, like this sounds like a really monumental thing for us to do and you're like, that's that's actually probably fairly straightforward given our tech stack. Let's let's see how quickly we can get there. And then you can you can much easier much more easily prototype it with the tools we have. I was in a meeting the other last week where they were kind of talking about where they would want to go with a product and they they were talking about it as if this if it was like, you know, months or weeks out. And I was like I was in the meeting with my camera off and I was like working with Kiro and I I'm like building it in real time. And I'm like like showed it to them after the meeting. I'm like, Hey, is this what you kinda had in mind? And they like, Wow, that's Yes, that you know, so there's a lot of opportunity for us to provide value that way in the sense that we're kind of showing what's possible with this technology to to the business, to the organization. And and that's you know, there's accelerating that that pipeline of like sh you know, showing value, getting feedback, showing value, getting feedback. That's the whole idea behind evaluation driven development. So it can work for agent ops, it can work for just kind of AI products in general or software in general. And so I think it's it's kind of a s a shift in skills, but it's it's gonna get us to I think it's heading in a in in in the right direction in in a good place. But we just can't it it's very easy to get distracted. It's very easy to like grab onto the next thing and say, we need this, we need this and be like, Hey, let's let's let's refocus on outcomes and and value and then you know, we can find the right tool for the job.

Yeah, I think just adding more features or doing more doesn't always solve someone's problem. I mean, look at Google Docs versus Microsoft Word. One has every feature known to word processing, another has eighty maybe eighty percent of those features.

Exactly. Right. Yeah. Ultimately you just need to like type something on print something out or or whatever. So yeah, a hundred that's a great example.

and I still also think that the part that takes just is hasn't been accelerated yet is actually verifying that the software is doing what you want it to do.

Yes. Yeah, mu much much

And you have to keep doing that over time as well.

Yeah, and it's actually the burden is higher, right? Because now now it's not just like a deterministic thing. It's it's indeterministic and you have to you have to add a you have to work on the harness around your AI workflow to actually make it, you know, more consistently produce what you want. There's still there's always gonna be that risk of something, you know, coming out that wasn't expected, right? but you you've gotta lean more into the checking that beforehand and testing it. And sometimes that means failing fast, right? Pushing something to production, even that may not be like this high quality standard, and having users test it. Maybe have a subset of users test it. We we have some really cool ways at WGU of having like students kinda come in and test things before we actually go to market. And that was something that definitely helped with the academic virtual assistant is we had a, you know, small group of students that were actually piloting it, giving feedback. And we've done that for multiple AI initiatives actually. And it's it's definitely helped us refine and and refocus on on what's most important. And so That's I think the if you look up a lot of these kind of MLflow 3.0 talks that they've given at, you know, Data + AI Summit last year or, you know, s some of these other ones they've been giving throughout the the last year, they definitely they always kind of draw that parallel where they're like, Okay, here's traditional software and here's my generative AI or like, you know, non-deterministic output workflow and there's there's more components to it because you're not you're not as easily able to just like write unit tests for for what you want. Your AI workflow in terms of like actual code could be really simple, right? It could be a a simple LangChain or LangGraph agent that's actually not that many lines of code. but like there's there's you know, you could cover that so easily, but to actually see whether that's working, that's that's gonna require a lot more tools.

Well Jonathan, thanks thanks again for joining us today. It was great to have you.

Yeah, it was it was a awesome conversation. thank you so much for for having it and congrats on the podcast and I I think it's it's a fantastic idea for you know, many, many more episodes to come, so I'm looking forward to to continue to follow it.

Thanks. enjoy enjoy doing it and having great folks and conversations on the show.

Thanks again, Dan.