High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Rebecca Hinds: evaluation

10 Jun 2026 The Cognitive Revolution Babysitting the Machine: Glean's Rebecca Hinds on the Hidden Human Labor of AI at Work

“You try to tweak one thing because the nature of the technology is probabilistic, it's not deterministic, you're not really sure what you tweaked that worked and didn't.”

— Rebecca Hinds

Source trail

Everything needed to verify it.

Speaker
Rebecca Hinds
Attribution
Verified speaker
Claim type
evaluation
Recorded
10 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, I used to tell people, When I did any sort of AI advisory consulting, I would say you could do a lot worse than as a leader just watching your token consumption, but definitely don't tell the team that's how you're gonna be measuring them. So it's really super easy to cheat on that. It's amazing that's actually, that's honestly been probably years ago that I was saying that, and it's funny to see that people are still shooting themselves in the foot that way. I guess in terms of understanding my fellow human. I struggle a little bit with the idea that this bot sitting work is so onerous. My attitude, which I don't expect everybody to share, but I just will give it to you for compare and contrast, is like, I used to have to do stuff and now I get to have AI largely in many cases do this stuff for me. I still have to make a contribution. But I definitely get a lot more done a lot faster. And I also get to learn a lot more about AI and what it can do, which I find to be always a very interesting question unto itself. And just a lot more because I'm able to take on so many more different things. I'm able to learn much more broadly and satisfy my curiosity in all kinds of ways that I never could before AI. So if you combine that with a And I would say mine is even more. I was at this AI event recursive last weekend and the question of people in the audience was, if your team had to replace you plus AI as it exists today, how many of you unaided by AI would they have to hire to get the same output? The median answer in that room was like basically two. In other words, people thought that they were twice as productive thanks to AI as they would be if they were unaided. And that's pretty much where I put myself as well. So even though, leaving that aside, okay, so people are reporting 13 hours gained. There's another stat in the report that says, let me make sure I get it exactly right. People are spending 6.4 hours, basically half of that time savings on this bot sitting activity. reviewing outputs, connecting things, different products that don't connect. And I think anybody listening to this show certainly had that experience of, okay, I got a clawed prompt, I got a clawed report or plan or whatever, but now I got to put that into copy and paste. So the phrase you just use is the human becomes the integration layer. And I felt that, and it is certainly tedious, although it's honestly also just a part of general computer work, right? We've all got kind of Slack and then this other task tracker and then there's e-mail. And so it's all a little bit disjointed anyway. All that to say, I don't really get it. Like, why is it so bad to be responsible for babysitting the bots or bot sitting? What is it that's really bothering people so much or alienating them so much about that? So it's the taking away from the meaningful work. And in the report, we look at three categories of interaction with AI. One is the bot sitting. One is the using the technology. So using the technology to do real work, you prompt it and It gives you the answer or it asks you a follow-up question in a way that you're moving work forward, iterating with the technology, as opposed to you're asking it a prompt. It doesn't have the context, so you're re-prompting it. We're seeing about 36% of all AI sessions fail, meaning a worker goes to use the technology and it's not successful. They either have to start completely from scratch or do significant rework. Imagine if it was right on the first time. Imagine if you could put those 6.4 or hours into either using the technology to drive work forward or the third category, which is learning or building agents, the time savings would be significantly higher. And so I think, again, there's a small component of bot sitting that I think is healthy. And for people who are curious, this is less of a problem. And I love to believe in human curiosity. I think the reality for many, and I think Nathan, you and I are probably an outlier here, the reality is employees are too to be curious right now. They're too overwhelmed with work to spend those one, two hours tinkering with the technology and prompting 4 different. tools and picking the right answer. And that's the problem. The fact that it's not meaningful work and it's not meaningful learning with the technology. We also look at along the different dimensions of bot sitting, what is most exhausting. In the report, we call it the exhaustion multiplier. And what we see is the highest exhaustion multiplier is associated with feeding AI context, right? Because that is in the best case, something your AI should know. know where the documents are, which documents are authoritative, and you should not be supplying that as a human in most cases. The other one is the debugging. The debugging. see an output, it's wrong, but because of the nature of LLMs, you're not quite sure why it's broken, what's wrong. You try to tweak one thing because the nature of the technology is probabilistic, it's not deterministic, you're not really sure what you tweaked that worked and didn't. That is the biggest contributor to this exhaustion multiplier. Yeah, that's interesting. The idea that I do have this experience sometimes. One thing I've been really enjoying doing lately is creating songs for each episode of the podcast. You can start thinking about if you want to request a genre. I try, but I can't always promise to be able to make something great in any given genre. It's really striking how sometimes I'll just run my kind of produce episode Claude code skill, which includes coming up with an idea for a song, writing lyrics, prompting Suno with that. Sometimes I show up and it's like banger immediately. And then other times I find myself sitting there and it is this kind of black box thing where you're like, you know, I tweak the style prompt, maybe I tweak the lyrics a little bit, go again. And for some reason it's just not working. It's just not landing. It's just not giving me what I want. And that can be definitely a sort of exhausting thing, especially because sometimes I get in this spot where I'm like, who even cares about these Like, am I doing this for anyone? Although actually I do get remarkably a lot of positive commentary on the songs. Anyway, I can relate to that sort of exhaustion point. I can definitely also relate to the shoveling context point. I used to have before my, you know, now much more integrated setup, I used to have just a single PDF that had a bunch of intro essays that I'd previously done for the podcast. And I found it was kind of exhausting, although this is like such a baby first world thing to say. Even just to go find that PDF every time and put it into the web UI so Claude would have it to use as examples. And it's like, man, that's really not much to complain about. And yet somehow it feels so much better now that those sort of moving of files and context around have been mostly automated away. Just on the time though, I mean, okay, so I'm empathizing with some of these problems, some of these pain points for sure, but it still seems like there's something, I guess one of the hypotheses we should always keep in mind is people maybe just in many cases don't like their jobs that much. I think this is something that the AI discourse broadly should remember much more than it does. And I've probably got on my soapbox enough times about that already, but that's for sure an ingredient in this overall recipe. But it was just striking still that like, okay, 13 hours saved, a little under half of that sort of reconsumed by doing this copying and pasting and double checking sort of work. But that still gives you almost a full workday back, right? Am I reading that right? Are people like, have we created a four day work week that we're just not ready to talk about?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence