Serious People Podcast


Rebekah Stevens
once spent months competing against an algorithm named Octopus, running weekly throwdowns to see who could staff a home care market better: her, with a spreadsheet and years of operating instinct, or a model built by her own data science team. She now leads Care Transformation at Honor, the senior home care platform, where that early contest became the blueprint for how the company decides what AI should touch, and what it shouldn't.

In this conversation, Noah and Rebekah get into how a business built entirely on trust and vulnerable moments decides where to let AI in. They cover why unstructured data doesn't automatically mean you need a language model, what it actually takes to turn a fuzzy client complaint into a precise, AI-ready metric, and why the hardest part of deploying AI is writing the goal correctly, not building the model.

You'll hear:
  • How Honor defined the outcome before automating staffing decisions
  • The real story behind Rebekah vs. Octopus, the human-versus-algorithm contest that built the system running Honor's staffing today
  • How her team decided a language model was the wrong tool for a problem everyone assumed needed one
  • Where Honor draws the line between AI working in the background and AI on the phone with a client
Timestamps: 
(00:00) Introduction
(02:05) Inside Honor: senior care built like a tech platform
(10:06) Turning client churn into a math problem
(16:38) Rebekah vs. Octopus: human vs. algorithm
(23:53) Where large language models actually showed up in the product
(26:59) When not to use an LLM
(30:10) What's blowing her mind about agentic AI right now
(33:01) Directed Work: letting the system decide who does what
(45:01) How AI will reshape product teams next year

Resources:
Valet — Episode Sponsor
Valet is the founding sponsor of the Serious People Podcast.

AI is getting very good at making useful things: dashboards, analyses, internal websites, and small tools. Making those things available to the rest of a team is a separate problem.

Valet lets you publish AI-created artifacts to a stable URL. New sites are private by default, and members of your organization can access them after signing in.

Try Valet at https://valet.dev 

Rebekah Stevens operator field guide: https://seriouspeople.valet.run/spp-rebekah-operator-field-guide/

Connect with Rebekah on LinkedIn: https://www.linkedin.com/in/rebekah-stevens-6751b41a/ 
Honor’s Website: https://www.honorcare.com/

Connect with Noah on LinkedIn: https://www.linkedin.com/in/noahlevin/ 
Noah’s Website: https://www.noahlevin.com/ 
Serious People Website: https://seriouspeople.ai/ 

  • (00:00) - Introduction
  • (02:05) - Inside Honor: senior care built like a tech platform
  • (10:06) - Turning client churn into a math problem
  • (16:38) - Rebekah vs. Octopus: human vs. algorithm
  • (23:53) - Where large language models actually showed up in the product
  • (26:59) - When not to use an LLM
  • (30:10) - What's blowing her mind about agentic AI right now
  • (33:01) - Directed Work: letting the system decide who does what
  • (45:01) - How AI will reshape product teams next year

What is Serious People Podcast?

Serious People Podcast: An Operator's Manual to AI is a podcast for business leaders and operators running complex, real-world businesses. The ones selling supplements, managing caregivers, or running service crews. Not software.

Host Noah Levin brings nearly two decades of experience at Amazon, Whole Foods, and healthcare tech to weekly conversations about what AI outcomes work inside a business, and what doesn't.

Equal parts practical and irreverent, but useful above all else.

[00:00:00] Rebekah Stevens:
The first inclination was like, great, we use AI elsewhere. Like, we use LLMs elsewhere to evaluate this. It's unstructured data. It makes perfect sense to put it in the LLM. But because LLMs are going to give you a probabilistic answer and not a deterministic one, they get it wrong a certain amount of the time.
And so you have to figure out, are you comfortable with that error rate, or are you gonna put an audit mechanism in place to ensure that you have correctly triaged and identified this feedback?

[00:00:25] Noah Levin:
Welcome to the Serious People podcast. Today, I'm talking to Rebekah Stevens, who leads care transformation at Honor, one of the world's largest home care providers for the elderly. Rebekah's job is to make sure that thousands of clients are receiving the care they need in their homes every day of the week, every hour of the day, and to build the tools necessary for this to happen at massive scale.
Rebekah is also my former boss and one of the sharpest operators I know. Today, she talks about how Honor picks the metrics that matter, uses human judgment to write software, and how they once made her fight with an octopus. What on earth am I talking about? And did the octopus win? You'll have to listen to find out. Here's Rebekah.

[00:01:08] Noah Levin:
Rebekah Stevens, my old boss, welcome to Serious People.

[00:01:12] Rebekah Stevens:
Thanks for having me. I will pretend to be a serious person for the next hour or so.

[00:01:16] Noah Levin:
That's the whole spirit of the Serious People brand, is pretending to be a serious person. I'm so grateful that you decided to say yes and come on this podcast and let me ask you to tell stories about our time together and your time well before our time together. You're one of these people who — you're like a born and bred operator who understands how to run a business on its own terms.
But then I think I was around for some of your coming of age as a product leader as well, and seeing you make the transition to doing both and using product as leverage for ops. And so I just, I admire the hell out of you, and I'm so excited to get you to tell some stories about your experience doing that, the trench warfare we've been in together at various points.

[00:01:56] Rebekah Stevens:
You need to hold me accountable. See, I can't come up with stories. You'll know the truth –

[00:01:58] Noah Levin:
—I have no leverage to hold you accountable now. But on the other hand, you also have no leverage to hold me accountable, so that works for both of us. Would you share just a little bit about your story at Honor? Talk about what Honor is, for people who have not seen Honor before or had a chance to use it as a customer, and your career journey, from when you started.

[00:02:15] Rebekah Stevens:
Honor, for those that aren't aware, is a platform for non-medical senior care.
So we're trying to provide care to seniors as they age so they can stay in their home, remain independent, age with dignity. Everything from companionship and transportation services to hands-on support so that people can get the support they need — bathing or toileting or medication reminders and meal prep — up to 24/7 support.

Honor as a kind of business model, the idea is that if we build a centralized technology and operations platform, we can support local agencies in spending more time with their local community, so current clients, potential clients, referral partners. And that is the business that we refer to as the Care Platform.
I led, for the majority of my first probably four years, our 24/7 client and Care Pro support, everything from your kind of traditional regional general management and account management, to the call-offs that happen inevitably at 2:00 AM and need to be restaffed, to the wildfires that require evacuations, everything else that comes with operations of a 24/7 home care business.

[00:03:21] Noah Levin:
Those are actual wildfires, like literal fires.

[00:03:23] Rebekah Stevens:
Literal fires. Yes. In this business, I mean literal fires, hurricanes, earthquakes, COVID, you name it. Also Care Pro recruiting and onboarding, which you of course are quite familiar with, and a few other operating teams. Probably about three years ago, we started doing a lot of work to understand what are what we refer to as the defects that lead to client churn. The number one defect is whether or not you have consistency for your clients. Do they see the same Care Pros in the home week after week? Which intuitively makes a ton of sense, but we could actually measure it, we could observe it, we could see where we were failing within our own client journey and even relative to the traditional network. And so I, along with a small team, took a market offline and we started to pilot a new way of staffing.

[00:04:07] Noah Levin:
What does that mean, take a market offline?

[00:04:09] Rebekah Stevens:
We took it outside of our kind of standard operating model. We said, “Okay, forget how we've done things previously. Let's operate this market differently.” And the reason I phrase it that way is what I actually did is I started to become a product manager, but I didn't know it at the time. I knew it as I was gonna take this market offline.

I luckily had great product counterparts, and the two of us and the broader team took this market offline. We sat in a room with PMs and designers and operators and data science and started to staff every visit and think about how could you optimize across the clients to get to a better client outcome, and how could you reduce this defect.

And so that was my entry into product. I'm fortunate at Honor opportunities are created, careers evolve. And so since that time have taken on much more of a product role. So I now lead a team called Care Transformation, which includes our product management team, as well as new revenue enablement.
So when we go to anything from a new state to opening a new business line, which you were obviously also critical in last year, testing out new kind of products or services, and bringing new owners onto the Care Platform, and have continued to support kind of various operating functions over the years.

[00:05:19] Noah Levin:
The thing that I find so fun about your story is that you were a tenured ops boss at the company. You were, I think, one of the most senior people at the company, both in terms of duration and, like, you were overseeing — was it hundreds of people?

[00:05:32] Rebekah Stevens:
Yeah, a couple of years.

[00:05:34] Noah Levin:
And I think we collectively, with the royal we, discovered this problem.
I say we, I wasn't even there. Discovered this problem that clients were churning because they weren't seeing the Care Pro that they expected on a regular basis. It was a staffing problem. We need to solve staffing. And so you got plucked off the top of this organization of 100 plus people to go, like, run a market and a spreadsheet yourself.

[00:05:59] Rebekah Stevens:
Pretty much.

[00:05:59] Noah Levin:
So, was it worth it?

[00:06:01] Rebekah Stevens:
Well, yeah. I mean, it was worth it for two reasons. It was worth it 'cause we did overhaul our staffing system, and I can objectively say that consistency for all of our clients has gotten better. But yeah, on a personal level, it's been incredibly rewarding. Like, I actually find it really fun to get to solve the same problems that I lived and breathed and know intimately so well, but with this, like, broader toolkit that is our engineering, data, and design teams, and being able to bring those two functions together.
I think one thing that's interesting about Honor and all kind of service-based products is at the end of the day, at least for the near term, probably for the foreseeable future, like, our product resolves to a human in another human's home. At some point, this still resolves to humans. The robots are coming.
We can save that for another podcast, but for now we don't have the robots and we still have humans.

[00:06:51] Noah Levin:
It's a very easy business to connect working your ass off day-to-day to having, like, a positive impact on the world, right? It's good for the people getting the service. It's good for the employees if we're doing our jobs right. It's a service we're all gonna need someday, God willing.
When I try and describe to people what Honor does, I find myself saying, it's Uber for, “I need someone to watch my mom in her old age.” And there's all kinds of reasons why that's not the right way to describe it. But the reason I like it, and then I'm—

[00:07:19] Rebekah Stevens:
Seth has spent years trying to overcome that analogy.

[00:07:22] Noah Levin:
Well, and there's — I'm saying it in part because I want you to take down, like, all the reasons why that's wrong. But, like, Uber is an example of a service that's all humans, or at least historically all humans, and it has existed on a small scale for a long time. We have a common boss/former boss who likes to say that humans are at best two-sigma machines. And if you want the scaled operation to be a six-sigma machine, you really have to have technology in the mix to make that happen. So that's why I like the analogy. But tell the fine people all the reasons that that's a terrible way of describing Honor.

[00:07:54] Rebekah Stevens:
First of all, I agree with that analogy. So I will give you that, and I think there's a ton we can learn—

[00:07:57] Noah Levin:
—that's good to know.

[00:07:58] Rebekah Stevens:
—how do you bring technology to a human services business, and obviously Uber would be an incredible company to replicate from that perspective. I think there are two big differences between Honor and Uber.
The biggest one, hands down for me, is Uber is a transactional business. You do not care who picks you up today versus who picked you up yesterday or who picks you up tomorrow. You care that they're reliable, you care that they're safe, they show up on time and get you to the right place, and we care about all those things too.

But in terms of our matching, that's the easy part of the match. If you then think about home care, you're talking about who's gonna walk into your home.
So all of a sudden you've got this, like, heterogeneous population you're trying to match with, and you're adding, like, personality as a dimension. And we haven't mastered that piece by any means, but it's a huge factor.

We also care about this concept of consistency. So it's not just who the match is, but it's this scheduling mix, that I want that same Care Pro who was here on Monday to be here on Wednesday and be here on Friday.

It's a far different dynamic than what Uber or more of the traditional marketplaces typically have to optimize for. The second one, which is it's really a three-sided marketplace, not just two.
So while day-to-day we provide Care Pros who go to clients' homes, you could think of that as, you know, drivers and riders. You've got this added party, which is the local franchise owners and their local teams, and they are critical in home care. We can't send a Care Pro to a client's home until we've gone into the home, done a consult, understand their care needs.

[00:09:22] Noah Levin:
Something that Honor has done an incredible job of is stepping back to understanding the objective function for the business. Meaning, like, what is the number that needs to go up or down in order for this whole business to work? I know there's a story behind sort of the invention of that metric and then the way it's been programmatized. Would you tell that story?

[00:09:40] Rebekah Stevens:
So the composite defect score is a composite of four individual defects that we identified were predictive of, is a client likely to churn? And the first version is actually specifically for starts of care. So is a client who has their first visit likely to churn in their first 30 days?
And it is, ballpark, they are four to five X more likely to churn if they have all four of these defects than if they have none of the defects.

[00:10:02] Noah Levin:
To even get to this point, like, what we've said as a company is, we wanna make money. The way you make money in home care is you get clients and keep them. And the measurement of how successful you are at keeping clients is you have a low churn rate.

[00:10:14] Rebekah Stevens:
That is correct.

[00:10:15] Noah Levin:
Right. So, already there, I think, by the way, Honor is winning, because a lot of companies are not even as precise as knowing that there is one output metric that needs to be sort of governing all else.
Before you get into the four metrics, can you talk about the definition of the consistency metric?
'Cause I think the nuance of how these metrics are defined, and how you have to actually be thoughtful about choosing one definition over another for the same concepts to matter, I think is worth getting into.

[00:10:41] Rebekah Stevens:
When we started down this journey, we met with a bunch of our owners that were on the Care Platform. Like, consistency was the word that everybody said mattered first. We all intuitively knew that in home care you care about seeing the same faces in your home.
Like, you know, we'll give ourselves enough credit for that one. The problem was no one had a standard way of measuring it. We had some of our original owners that were on the Care Platform with us, and we were talking about, “Well, how did you think about consistency before you came on the Care Platform?”
And every one of them had a different definition. There were 12 different ways to define consistency. We actually started with a different definition than the one that we landed on. But it counted up unique number of Care Pros over a certain number of visits, which intuitively made sense.
It was correlated with churn, but there were these edge cases that it kept falling apart on. And what we ended up coming up with is it is a rolling four-week metric defined on a per-week basis, based on the minimum number of Care Pros that is required for the number of hours that you have on your schedule, assuming a Care Pro can work up to 40 hours a week.
I know that's a mouthful. I don't expect anybody to remember that, but it was getting to that level of detail that finally had the highest predictability and the highest correlation with churn.

[00:11:53] Noah Levin:
Like, if someone has a four-hour a week, three times a week schedule—

[00:11:57] Rebekah Stevens:
Yeah.

[00:11:58] Noah Levin:
—better have one person delivering that care. But if they're 24/7 continuous coverage, then there's an allowance in that metric for having a larger number of people rotating through?

[00:12:09] Rebekah Stevens:
Exactly. The four 24/7 one rounds up to up to five Care Pros. And intuitively, if you think about breaking up a schedule, it makes sense that way. But it gives you a single metric that then allows you to compare the client experience between the 12-hour client and the 24/7 client, and it gives all of your staffing and optimization models one objective metric to optimize for across the market.

[00:12:31] Noah Levin:
Reflecting on our experience working on this metric together as humans, and also now the way that you would put an AI to work against the same metric, and how the specificity of the metric — it matters a lot to, like, running an ops team properly, right?

To be able to say this metric is so well specified that if you get the number to go down or up or whatever direction it needs to go, it almost doesn't matter how you did it. I mean, there needs to be guardrails obviously, but it means the outcome is what you wanted it to be.
Because the metric really is that precisely defined. And a few months ago, Claude and OpenAI both rolled out Goal. There's this function now in both these tools called /goal, and you can basically tell it to go run at a task, and it'll keep working until it hits the goal. There's things that are easy to keep within that loop that are all software, and then there are things that are harder, like running a skilled home care business. But it will not exit the loop until you've hit a very precisely specified goal or it feels like it's totally blocked.
And actually, the hard part of this now is writing the goal correctly.

[Sponsor — Valet] Noah Levin:
Today's episode is sponsored by Valet. I had a client tell me this week that they create tons of beautiful HTML artifacts with Claude, but they have no good way to share them with each other and with clients. Valet is built to share the things you make with Claude, with your team, your clients, or the whole world.
For this episode, I told Codex to take my conversation with Rebekah, pull out the best lessons for operators, add the source links and a few Serious People style visuals, and turn it into a companion page using Valet. Codex built the page, found Valet's agent instructions, and created a fully hosted page, which, by the way, contains some great takeaways and prompts from this episode.
What I like about Valet is that it handles all of the plumbing seamlessly, so I can focus on the content. Valet is the founding sponsor of Serious People. Check out this page I just made in the show notes, or go to valet.dev to try it yourself.

[00:14:25] Noah Levin:
So now Honor has these four metrics inside of this CDS uber metric. So, you were gonna go into what those four are.

[00:14:33] Rebekah Stevens:
So we have consistency that we just talked about. We have care quality that I touched on, so that's quality of the Care Pros in the home. We have canceled visits. This is another normalized metric, but it's a pretty liberal definition of canceled. So even if the client says, “You know what? Don't bother coming tomorrow. I don't want a new Care Pro in the home.” If it's anything that's kind of relatively short-term and we think it could be attributed to the staffing situation, we take the knock. And it's not just 'cause we're humble, it's because we saw that that actually is predictive of churn.

And then the last one is a metric called UHR, which is Unique Honor Reps, which is a proxy for kind of the quality of customer service that clients are getting when they call or text into our centralized support teams. Are they talking to one person? Are they talking to the same person? Are they getting their issues resolved quickly, or do they have a really disjointed experience?
So again, all of them kind of intuitively make sense, but now mathematically you can do exactly what you just said, which is set objectives, whether it's for AI, whether it's for your ops team, whether it's for a non-AI based optimization model, to solve for your primary objective. And what we now spend a lot of time on is understanding the dynamic between them.

When you have to make a trade-off between consistency and quality, where's that intersection? How do you think about optimizing the client experience in that situation?

[00:15:47] Noah Levin:
How do you think about optimizing? When there's a trade-off to be made — I mean, I imagine there's a lot we could talk about here, but at a high level, is there a tiebreaker?

[00:15:55] Rebekah Stevens:
The great news, it's all math.

[00:15:57] Noah Levin:
Great news.

[00:15:58] Rebekah Stevens:
The favorite buzzword and answer. It's all math. But what I mean by that is, it's a matrix, and you can look at all four defects and say, “Well, if I have good consistency and bad care quality, what's the impact? If I have good care quality and bad consistency, what's the impact? If I have both, what's the impact?” And use that to then rank your trade-off decisions.

But all of the systems, you know, our centralized staffing algorithms can use that type of a logic to build the rule set and build the optimization.
AI models are gonna optimize for whatever you tell it to optimize for. And so if you're not precise in how you define it, you're gonna end up with some really weird edge cases.

[00:16:38] Noah Levin:
You've got this well-defined goal, and you've been tasked with taking the market offline to figure out how best to do that. Can you talk through what you actually did when you took the market offline? Like, for the problem that you were trying to solve, what did it actually look like to go solve that?

[00:16:51] Rebekah Stevens:
So I think one thing that tripped me up is you're like, “Well, when do you bring in technology?” And it depends how you define bring in technology. We've been most successful when we've had the technologists, the people that understand the technology, sitting alongside the operators.
So one thing that worked really well when we took that market offline that I mentioned was we had myself, one of our most senior data scientists, one of our most senior PMs, one of our best designers all in the room together, even though the early days were truly whiteboard and notepad staffing, the way I would've just historically staffed as an operator.

The first iteration was exactly that. We had a spreadsheet. I said, “Here's how I think I would staff this market to try to get to better outcomes. Can I prove that I can get better outcomes than what we're doing today? Can I get the Care Pros to actually accept the visits? Can I, like, puzzle piece my way together?” It was intentionally a fairly small market so that you could look at all the puzzle pieces at once and say, “Is there a better way to do this?” And I now had the metric to check myself every week and say, “Yeah, old system or system-wide average is excess Care Pro of…” I'm gonna make up numbers here, “1.5, and I was able to get to 1.3 this week.”

Well, I had the data scientist sitting alongside me, and so he started to write scripts to do this. They would print to an Excel sheet. And these were all — it was an optimization algorithm based primarily on rule sets.
There was no LLM in this. This was more historical, just optimization. So phase one was just me staffing. Then phase two is we would run the — codename for the actual model project was Octopus — and we would run Rebekah versus Octopus. We'd have throwdown death matches of who could get to the better excess consistency score, week after week.

And, like, at first I was crushing it. I was just, I was dominating. But over time, the model started to catch up, because we would compete and then we would go audit the results. And I would sit next to our data scientist and we would look at and say, “Okay, well I did this.” And he would go, “Oh, well, we've constrained the model on that. Let's figure out how we can allow the model to flex in that way.” Or, “Oh yeah, that's an edge case the model failed out at and you were able to logically reason your way through. You need two Care Pros on a seven-day-a-week schedule. How do we integrate that into the model?” And so once we did enough rounds of that, we said, “Okay, let's take Rebekah out of the loop.”
We put a kind of standard operating team in place, and we said, “Every week you're gonna run off of the spreadsheet-generated recommendations.” So there's still no, like, hands-off-the-wheel automation happening, but spreadsheet, script would run, it would print your recommendations. Our Care Pro relationship managers, who are the operators who have the relationships with the Care Pros, they do most of our ongoing staffing, would take those recommendations and try to staff Care Pros against them.
Are the Care Pros accepting? Are those schedules sticking? Is there anything we're missing? What happens when a start of care comes in after the algorithm's already run for the day? Like, how do we reconcile those real-life operational challenges? And then once we were able to do that and we consistently were delivering outcomes that were better than the system average, we then said, “Okay, engineering team, it's your turn now to come see what we've been doing in, like, scrappy Python scripts and Jupyter Notebooks and Excel spreadsheets, and, like, turn this into production-grade code.”
And so they then came in, built the underlying infrastructure, built the UI so the rest of the team nationwide can use this, and that is what turned into Market Planner, which is now our centralized staffing system.

[00:20:16] Noah Levin:
So Rebekah versus Octopus is part of Honor lore at this point. It keeps coming back in different, more sort of meme-able formats. It was the stuff of legends when I got to Honor. And I think it's sort of looked at as, like, a best practice for how to approach a new problem, where, when we don't know how something is supposed to work, or we think we might have overcomplicated our own product and need to get out of our own way, you take a market offline, you go put a domain expert on it to do it by hand, you relearn the lessons of how to actually do it, and then you build the software incrementally around it.
But what's funny is, as I'm hearing you tell the story — I haven't thought through the story since my last six months immersed in AI — I really feel like there might be a different approach to it now. The loop that you and Patrick and the team were in, with, like, iterating on Octopus and then fighting it out with Rebekah, that's like a learning loop.

It sort of feels like reinforcement learning, and maybe it would take fewer cycles now.

[00:21:14] Rebekah Stevens:
I think it should take fewer cycles, because I think right now the benefit you get is that you should be able to get the learnings a lot faster. The iteration on the other end, frankly, the coding iterations to take those learnings and integrate them into the next, you know, version of the algorithm should be way faster.
I don't know if we're at the point that you could do it with just multiple agents without a human yet. Because part of what you would need is — you still need the arbitration of good, and you need someone to eyeball the guardrail metrics. And to your point, I have no doubt that the AI agents could get to a better consistency metric mathematically. What I don't know is how many rules they would figure out how to break along the way, because we hadn't thought to, like, appropriately constrain them in doing so.
There's still gonna be a role for the iteration with the agents. It might just be more like player coach rather than, like, two competitors. Like, I think you're right. I hadn't thought about it. I'll have to think a little bit more about how we'd actually, like, set that up.
So there's definitely something different there. It should be faster, but there was so much judgment in this one and there was so much just kind of defining what was allowed that I would be— I don't know that we could have just said, “Here's the rules of the game. Go play it,” and, like, stepped back and had a better outcome.

I think you could have wasted a lot of cycles as well.

[00:22:27] Noah Levin:
There was just so much work that happened before engineering got involved, or maybe they were involved in the offline prototyping, but the story of figuring out that churn is what matters, testing the hundreds of possible candidates for inputs into churn to resolve to the four that matter the most, then developing an offline prototype of something that could possibly improve that churn — and this is, like, years of work before the first engineer wrote the first line of code against Market Planner, which is the product solution to one of those four metrics.

[00:23:00] Rebekah Stevens:
I think the part that 100% would be faster, because we're dabbling with another iteration of, the—

[00:23:05] Noah Levin:
No, no. I thought we were done.

[00:23:07] Rebekah Stevens:
Never. You can always be better. No coding to do the prototypes, right? Like, that used to be a really strong position, is like engineering is a scarce resource.
Do not spend your engineering resources until we know what to build. Do whatever you can in spreadsheets. And the concept still applies, but I think with AI, the advantage is coding has become so cheap, you could probably get to a working prototype pretty quickly. Now, there's still a whole bunch of steps between that and actually wanting to put that into production code, and you would want your engineers involved for all of that.

But that's probably one, like, very explicit step that would be different today. You actually probably would not need to build the spreadsheet loop, because there's probably a faster, easier way to build the prototype of the actual staffing tool.

[0:23:53] Noah Levin:
When did LLMs first enter the picture in the product at Honor?

[0:23:57] Rebekah Stevens:
The first one that I think we used an LLM for is we collect care notes after every visit. Care Pros leave notes, they're visible to the family and kind of recording what happens. And in those notes, we had already had a machine learning algorithm scanning them for words that related to falls.
What we were trying to find out was, did the client potentially fall, which is a huge risk to our clients, and it may not have been reported. Inadvertently, maybe it mentioned that they'd fallen over the weekend outside of a visit, or they fell, but they got up and they were fine, and so that wasn't reported.
But we want to be able to understand that and investigate the situation. And so we'd had this model pre-LLMs, but it was an obvious use case once we had LLMs in place to update the models, to be able to scan and say, “Okay, well, you know, keyword searches are valuable, but you still miss a whole bunch of context. Can we train a model to be more nuanced in how they look for falls, and are there other things that they can identify? Are there other changes in conditions? Are there hospitalizations? Are there other risk factors that we can review these, you know, essentially audit logs that we have, these care notes that are submitted that there's some stat that I won't remember.”

We could never scale up a human team enough to, like, read all of them consistently. And so having an LLM-trained model — and it's frankly something we're still working on, you know, it'll be constantly evolving — identify risk factors.

So that was one. The second one was listening to all of our conversations with our clients and picking up on the nuance of when feedback is provided. We had created this feedback score, but one of the biggest issues we had is we don't get feedback on a lot of our Care Pros. Like, we ask for feedback at the end of every visit.

We send out surveys related to our Care Pros, but like any business, customer surveys are a relatively low response rate. And so we had a bunch of Care Pros that we just had no signal on. But we actually have a lot of conversations with our clients where they talk about their Care Pros, but it might not be the explicit point of the conversation.

So I call in, “Hey, Noah, you know, I wanted to schedule another visit for my dad on Thursday. Do you mind checking to see if Care Pro Joe's available first? Like, they had a great visit on Monday, and I thought they connected well.” Well, as an operator, what did you just hear? “Okay, I need to add a visit for Thursday. I need to reach out to Joe to see if Joe's available. I need to call this owner back and, like, confirm — or the family member back and confirm — if they can schedule the visit.” And the thing you probably didn't also pick up on and do is, like, “Oh, and I should add a note to Joe's profile that he received positive feedback from this client.”

Whereas now we listen to every call and there's an LLM scraping those calls that first identifies it as feedback and then can start to look and categorize, like, what was the type of feedback, document it, and even start to trigger, you know, a series of tasks that might be based on that feedback. So those are the two first ones I can think of where we truly integrated LLMs into our, like, production system, which I would differentiate from using LLMs in terms of, like, how we operate internally.

[00:26:59] Noah Levin:
When I was joining Honor, part of the pitch at the time was, “We're using AI to do this work.” It was early enough that that felt implausible to me and, like, a sales pitch. And I remember getting inside the company and two things happened. One is that I watched Ian, the president, talk about robots, which I still think he's probably right on a longer time scale. I watched some of our operators use the summarization feature where they could take this whole chronology of months or years of care and get the one-paragraph TL;DR on what is the state of this client right now. And it was this light bulb moment for me of, like, oh my God, our business just throws off unstructured data. It's just this, like, gobs and gobs of qualitative data that are unusable for so many important things because it's just all there. There's so much of it and we need to do something right now. And to be able to see that context collapsed for someone who has to answer the phone and interact with a client in real time about possibly an emergency — that was, like, my, oh, I get it now moment.

[00:28:07] Rebekah Stevens:
I forgot about that one. That one might have come before the feedback tool, so see, you can hold me—

[00:28:11] Noah Levin:
Did I win?

[00:28:12] Rebekah Stevens:
Yeah you might have. I think falls was still first, but, yeah, that's a great example. I think an interesting thing that we think about is, like, there's also the risk that you want to use LLMs for everything.
Now it's like, how are you using an LLM to solve that? It's kind of almost the first question, and one of the most important things is trying to figure out, like, is an LLM actually the right way to solve that?
So while I just gave the example of the totally unstructured call and an LLM being a great use case, because we actually need to identify that we're even receiving feedback before we can do anything with it. We have another way we can receive feedback through a form, but it currently comes in in an unstructured format, and it actually — there's a step where a operator needs to review it and determine was this positive or negative or neutral feedback.

And then the system kind of picks it up from there and continues on its way. And we were trying to figure out, like, why is this feedback not getting actioned on? And it's like, well, because there's a bottleneck in the system, that it actually has stops and a human has to review it, and we have to action on it.
And the first inclination was like, great, we use AI elsewhere. Like, we use LLMs elsewhere to evaluate this. It's unstructured data. It makes perfect sense to put it in the LLM. But because LLMs are going to give you a probabilistic answer and not a deterministic one, they get it wrong a certain amount of the time.
And so you have to figure out, are you comfortable with that error rate, or are you gonna put an audit mechanism in place to ensure that you have correctly triaged and identified this feedback? Well, the feedback's already coming in in a form. If you're getting it in a form anyways, you know what you can do?
You can ask one question in a structured format that says, “Is this positive or negative feedback?” And it is just one of those, like, when you look at it, you're like, “Oh, that's a far more elegant solution.” But everybody's first inclination was like, “Oh, it's unstructured data. It screams LLM. We should absolutely put an LLM on it. We already have one that does something kind of similar somewhere else in the system.”
And so, yes, when have we started to integrate LLMs? Like, the history is useful, but I think it's also almost as useful now to be like, when does just, you know, traditional programming or traditional solutions actually make more sense, or when might they be more appropriate for the problem?

[00:30:10] Noah Levin:
Doing simple things a hard way is, I think, a common pattern right now. So it's been a few months since I've been inside of Honor, and a lot has changed in those, yeah, few months in the world of, like, what's possible with AI. What's blowing your mind right now?

[00:30:26] Rebekah Stevens:
It's probably on two fronts. I mean, what's blowing my mind is that it's hard to keep up, right? I feel behind every day that I didn't, like, learn about the next tool or think about the next way to apply it. But within production, the ability to classify information and use that to trigger actions.
The kind of nuance that we can pick up through listening, and the amount of automation that that can then generate, is far more powerful than it was even a year ago, and I'm sure we'll say the same thing next year.
We've built on our agent, which enables our engineers to use agentic workflows to accelerate all of their at least first pass at coding. It's improving in terms of its ability to do code reviews, et cetera, though that's still a bottleneck.

But the reason I bring that up as blowing my mind is that's actually where it starts to change how do we build. It can be so much cheaper and so much faster to build that you think about the, like, product management life cycle and you're like, “Well, how much time do I spend on evaluating the value of this if it's only gonna cost me an hour to build it anyways?”

So I think the, like, velocity with which we can work is probably what's actually mind-blowing to me right now.

[00:31:35] Noah Levin:
It's amazing what these agents are capable of, and yet the gap between what they're capable of and what is necessary to get a feature over the line, especially in a business where there are life and death consequences to errors in some cases — it's like a Zeno's paradox of, like, you feel like you can never quite reach that finish line sometimes. Have you found these agents are, like, really limited in a specific area, or is there something where it's just, like, a totally unsolved problem at this point for an agent to go work on?

[00:32:01] Rebekah Stevens
If you have good patterns in place, if you have the data integrations in place that you need, they're incredibly good, right?

[00:32:08] Rebekah Stevens:
If you have the data inputs that you need and you have a pattern for automating something and you can say, “Do it like that, and I now need to automate these three other workflows,” go for it. Like, it's incredible. Where they struggled was either we, like, didn't have the infrastructure in place.
You had to make a ton of infrastructure level decisions that the agents weren't on their own gonna go design and develop, or you didn't want them to without quite a bit of oversight.

[00:32:30] Noah Levin:
What's sort of interesting about the variable work that we do as operators at Honor, with the tasks that are meted out to the operations team that have to be sort of worked in progression, is that we've been thinking in an AI mindset for actually a lot longer than AI's been around, with this program called Directed Work. The lens that I have on Directed Work in hindsight is that tasks kind of exist on this continuum from, like, really high judgment, requires expertise, and then it sort of continues down to this low judgment set of tasks.

[00:33:01] Noah Levin:
Can you tell the sort of the thesis of Directed Work and what that platform looks like and how it's being leveraged for automation?

[00:33:09] Rebekah Stevens:
Directed Work is a centralized work triage system, I guess, for lack of a better term. I, as an operator, would come in and I would think about, “What do I need to do today?”
And I'd go check in one system and see if there's anything there, and I'd check in another system and see if anything came in, and I would kind of use my best judgment to decide what I should work on at any given time.

And so the thesis of Directed Work is the system should be the brain behind the operation, and it should hold the memory and the context and the logic around what is most important to work on and who has the skills to do the work — and that could be an agent, that could be a human, that could be a really senior high judgment human — and allocate the work.
So in order for this to work, all of the work needs to be, like, broken down into states and then the tasks. So where are you at, at any given point, what is the condition of this task, and what is the task to be done?

[00:34:00] Noah Levin:
Just to make it concrete?

[00:34:01] Rebekah Stevens:
Yeah. A Care Pro could call off from their visit. That's gonna be an event, and that is going to trigger a series of tasks. When a Care Pro calls off, the Care Pro needs to be removed from the visit. The visit needs to be restaffed. In order to restaff the visit, you need to call the eligible Care Pros, or contact the eligible Care Pros. Once the visit is staffed or a Care Pro agrees, you need to assign them to the visit.
They need to be prepped for the visit. The family needs to be notified. And by breaking the workflows into individual tasks, each individual task can then carry with it the kind of context that it needs to complete that task, the logic for when it should be completed, when it's no longer relevant.
Like, you know, if the client calls in and says, “Never mind, cancel the visit,” the system needs to know, well, if the visit's canceled, you know, tasks five through eight can be canceled. I don't need to do them anymore. And who is qualified to do the task? And then to your question of, like, when can you start to automate things?

Well, then you can start to make those decisions and say, “Is this a purely administrative task?” Like, there's no value added by a human doing this task. They are just pushing a button, and I can automate that task. It needs to happen, but it can be fully automated. Is it a task that, you know, we have 100 people that are trained to do, and I should go find the first person who's available because it's an urgent task?
Or is it, like, a really high judgment or high context task where the relationship might matter more, and you're gonna assign that task out based on a different set of conditions, a different priority list, and hopefully — or maybe it has a different urgency associated with it.

So we continue to try to push Directed Work throughout the whole system. We're still working on it. And then once we have the rails in place, we can continue to iterate on what can be automated, what can be allocated to humans and which humans.

[00:35:54] Noah Levin:
Can we talk about culturally how this feels when you're sitting on the inside of the company? As you said, it's a human business, right? So the product that Honor provides is taking care of people in, like, the most, you know, fragile and vulnerable time and state in their lives. Usually the call that comes in, the initial customer service inquiry of what kind of care can I get from you, is, you know, during a hospital discharge or another acute event. It's this really, like, high emotional IQ, human-to-human business, and I think it attracts a certain type of human-oriented operator.

So as AI starts to be more useful in more use cases, how does it feel as an employee of Honor?

[00:36:37] Rebekah Stevens:
I don't have anywhere to compare it to because I've only been at Honor the time that we've gone through this journey, but I have to imagine, having spoken to friends and everywhere else, like, there's an element of uncertainty that everyone feels everywhere, and I don't think we are immune to that, right?
What does this mean? What does this mean for me? What does this mean for how we work? But I think part of how we've navigated it with Honor, and where I would say there's actually genuine enthusiasm, is if you can anchor on AI as a tool to accelerate your mission. One thing we've talked about, like, often you hear about, you know, AI as a cost-saving mechanism, and that is a possibility, but it also can be an accelerator to velocity of development.

It can enable you to have higher standards of care. It can enable you to do more personalization. It can improve your quality standards because you can run so many more tests, et cetera. And so part of what we've done is tried to focus on how can this help us accelerate the delivery of our mission? How do we use this new tool?

Because that's what it is, right? The computers were new tools. The internet was new tools. Excel at a time was a new tool. Like, AI is the newest of the tools. It's new to us. It feels different. It can feel threatening. But as with all other technology kind of changes, if you can harness it to drive your goals, to drive your mission, we may be able to hit our financial and business goals, but also provide more and better care to more people.

The other part from a culture perspective is where does it touch the client and consumer experience directly, and where is it behind the scenes? So I think we have been fairly conservative on if a client calls in, to your point, in a time of need, could we have a bot that they could speak to that could probably get most of their questions answered?

Do I actually think the idea of having a, like, personalized agent is actually incredibly compelling? Yes, but we're not there yet. The level of humanity that we would need to build into those agents is just not where we're at right now. And so I'm far more excited when it enables our really strong operators that you talked about, who are here because they care about seniors, because they care about our Care Pros, to spend less time doing administrative tasks and more time on the phone with clients, Care Pros, you know, partners, et cetera.

We're all here to provide really high-quality care. And how do you use it in the right use cases so that when that human touch is important, when there is real value in that human touch, you protect that and you use AI in the background to make that human touch better, to give the person on the phone way more context so it's a far better conversation, or to just free up their time so they can spend that time with the people they need to connect with.

[00:39:14] Noah Levin:
What are you gonna work on solving tomorrow?

[00:39:15] Rebekah Stevens:
There is still so much potential in improving our Care Pro client matching. And a lot of it is around, like, the nuance of personality matching, the nuance of schedules and schedule shaping.
Like, it's one thing to say I'm available, you know, from a Care Pro's perspective, Monday, Wednesday, Friday, 8:00 AM to 4:00 PM. It's a lot more complicated.

You know, like, there's just so much nuance in the real world of schedules and personalities and care needs, and I think we've got so much potential, and I think LLMs will be incredibly helpful in capturing what historically has been purely unstructured data, though we still have to figure out how do you leverage that unstructured data in a structured system to improve the outcomes.
But I think you end up with not just better matches on paper, but hopefully matches that Care Pros are more likely to accept, that are stickier with clients, and therefore lead to all the benefits we've talked about in terms of client retention.

Two is on our, like, engagement with our Care Pros and workforce management at scale. Workforce management at scale in a remote environment can feel transactional. And I think it's something we've tried to overcome with roles like the CPRMs, but the ability to think about how do we engage, support, train, you know, performance manage as appropriate a Care Pro workforce to build a community, I think is an area that we have a ton of opportunity to do.

Both leveraging LLMs and just because as we create the space to spend time there, I think will be one of the big differentiators in this space. And then my third one that I'm adding is as we start to solidify the systems that are core ADL support, we get to start to look beyond. We actually are finally at the, like, we have the time, we have the resources, we have the scale to start to dabble in how do you supplement core ADL support, core non-medical home care activities of daily living support, with whether it's technology.
We think about sensors or devices or wearables, whether it's additional services, how do you kind of create that full package for elder care? And again, it's been the goal for a long time, but I actually get to start to spend some time there.

[00:41:21] Noah Levin:
That's really exciting. Can you — I'm gonna make a unforced pun here — can you steel man the case for robots? And what I mean by this is, there are people who are looking a little bit farther ahead than the operators and who mean it, and will say it with a straight face, that robots are a part of the service offering in the not too distant future. What are they actually saying, and will I have a robot in my house taking care of me in my old age?

[00:41:50] Rebekah Stevens:
I think you will. You're not gonna be old for a long time, so I think you will. I feel confident—

[00:41:54] Noah Levin:
You don't know that.

[00:41:56] Rebekah Stevens:
I do fundamentally believe there's a role for robots. I think there's a couple of places that robots will come into play. One is, how do you supplement the humans that are in the home? For a lot of people, 24/7 care is incredibly expensive, and you may not actually need or want somebody in your home 24/7.
Like, I think about my parents, the last thing they're going to want is another person in their home. And so having some additional layer of security, if you think about it as, maybe it's monitoring just to start with, right?

If it's a robot, it's far more mobile. It can collect far more information than a static detection that's, you know, on a wall in your home. And when you leave that room, it loses sight of you. If you're out for a walk, it could capture it, things like that. So I think the mobility of monitoring and the ability to supplement what your in-home care solution looks like, I think is one of the entry points for robots.
I think probably the second place for robots is in more rote tasks, right? Like, you think about the other things that you need to stay safe in your home: cleaning, laundry, meal prep. Like, this is where the robots are actually starting to get, you know, at least in some of the demos, to look like they're fairly capable.
And again, it helps bring down the affordability of home care if you can supplement hands-on human care with a robot. I think the two longer-term, like, to-be-seen are how comfortable are people with robots as a companionship device? I know there are, like, the robot dogs and things, and for some people that may work.

I don't know. Like, that one, you know, to be seen if that replaces human-human contact for companionship in the home. And then I think the other one is hands-on physical care and support. The amount of dexterity and skill that will be required to think about transferring a human, like, from their bed into the shower to help them with showering and moving them again.

Nothing I've seen in the technology says we are close to that yet. But I do think as a supplement to in-home care, as an ability to bring down the overall cost of care, and in particular in that kind of monitoring space, there's huge potential for robots.

[00:44:03] Noah Levin:
I think about the LLM moment that I had when I joined Honor and the realization of, oh, this is how it can be useful and it's actually real enough that it can provide value. I think there's a moment like that coming for a lot of us in the next few years where we realize that the technology actually has gotten good enough, and that the things that feel like they should be weird about it either aren't as weird as they seem or are outweighed by the utility.
And for the business that Honor runs, there's just so much need for it, and affordability is still such a huge problem, that I think there's gonna be this natural shape to the product that comes where you get a robot with your human, and the robot stays a lot longer hours and does a lot more tasks, but the human is still a core part of it. And that there's just a natural complementarity there.

[00:44:46] Rebekah Stevens:
I agree. Yeah, and you may start to think about schedules differently, right? The human comes when the human's really needed, because you've got the robot to kind of help manage some of the other day-to-day.

[00:44:56] Noah Levin:
Always comes back to schedules with you.

[00:44:59] Rebekah Stevens:
Always schedules.

[00:45:01] Noah Levin:
On this note: when you think about the things you're working on now to solve at Honor and the direction that the technology's going, what do you think will be fundamentally different about either the product or the way the team works in a year?

[00:45:14] Rebekah Stevens:
I think the way the team works is the thing that's gonna change most fundamentally over the year, but I don't actually know what it looks like yet. I think everyone's defined roles are evolving, but I don't yet know in exactly what ways. I know, like, you talk about the kind of oscillation that everybody's on in this journey.
I don't think every PM is gonna become an engineer. I don't think every designer is gonna become a PM. I don't think every engineer is gonna become an operator or a PM. But I do think the kind of traditional agile process — a PM sitting at their desk and by themselves writing a PRD that they then hand to an engineer — like, I just don't think that kind of linear way of working is going to work the same way.
Iteration cycles are gonna be so much faster, it doesn't make sense. Agents are going to be part of the team, and so you have to figure out what's the agent's role versus what are the human's roles. You know, I know there's this concept of kind of builders in general. I don't know what that looks like at Honor yet.
But maybe that's what you'll have to have me back for, to figure out. I'll let you know in a year. So I think how we work will change a lot. I think the how it integrates in the product is a little bit more kind of continue down the path we're on. I mean, I could be surprised obviously, but I think we kind of see the path for that.
I think it is continue to build agentic workflows, continue to get work into Directed Work, continue to automate where automation makes sense, and probably start to dabble in more of the LLM-based, potentially consumer Care Pro-facing parts of the product. But again, that one feels like we kind of know where it's going a little bit more, and it's just a matter of executing.
I think the how we work is actually where we've got a lot more to figure out.

[00:46:49] Noah Levin:
Rebekah, if people want to be helpful to you, is there a place that they can go or something they can do to support your work?

[00:46:56] Rebekah Stevens:
I would love to connect with anyone that's interested in this space, and, you know, we're all trying to learn, so I'm sure there are people out there that I could learn a lot from as well, particularly those that have integrated AI into kind of human services businesses. I am on LinkedIn, so feel free to reach out and connect.

I am not a particularly active LinkedIn user, but I will be sure to—

[00:47:15] Noah Levin:
Why? What are you busy with?

[00:47:18] Rebekah Stevens:
You know, we got this whole care thing, those schedules, they keep having to be—

[00:47:21] Noah Levin:
—still fighting the Octopus.

[00:47:24] Rebekah Stevens:
I'm on a learning journey like everyone else here, so happy to continue to build the community, and we certainly welcome people that want to contribute to our mission at Honor as well.

[00:47:33] Noah Levin:
Well, I really appreciate you taking the time to chat and walking back through all the hard problems at Honor. I just, I so admire the way you approached them and the experience and the depth of caring that you bring to it. So thanks for sharing that with folks today. And, yeah, I appreciate you being on.

[00:47:48] Rebekah Stevens:
Thanks for having me. It was great to connect, and I am so excited for the work you're doing. I will continue to learn from you, so I appreciate it, Noah.

[00:47:55] Noah Levin:
Rebekah Stevens, thank you for joining Serious People.