Speakers in the public record
Claim mix
belief 28evaluation 3recommendation 3commitment 2preference 2disagreement 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
39 published records
“I think some way, like, right now, it just takes, I think we sort of force people to do the best practices of writing out sort of these full JSON schemas, but it would be really nice if you could just pass in a Python function as a tool.”
- Publisher
- Latent Space
“After the first go to Google, I think I tried to have it play Minecraft or something, and it actually installed and opened Minecraft.”
- Publisher
- Latent Space
“We can cut over to computer use if we're okay with moving on to topics on this, if anything else. I think we're good.”
- Publisher
- Latent Space
“I think there's, you know, the whole gamut from very simple to like very complex.”
- Publisher
- Latent Space
“I would say on the positives, like, there is really sort of incredible progress that's happened in the last five years that I think will be a big unlock for robotics.”
- Publisher
- Latent Space
“I will say SWE-Bench just released SWE-Bench Multimodal, which I believe is either entirely JavaScript or largely JavaScript.”
- Publisher
- Latent Space
“I think the computer use demo that we released is an extension of that. It has the same bash and edit tools, but it also has the computer tool that lets it get screenshots and move the mouse and keyboard.”
- Publisher
- Latent Space
“We're not going to talk about the training for that because that's still confidential. But I think Anthropic's done a really good job, like applying the model to different things.”
- Publisher
- Latent Space
“I think that for robotics, the limiting factor is going to be reliability, that these models are really good at doing these demos of doing laundry or doing dishes.”
- Publisher
- Latent Space
“I think there's still need for sort of many different varied evals. Like sometimes you do really care about just sort of greenfield code generation.”
- Publisher
- Latent Space
“XML is good for everyone, not just Cloud. Cloud was just the first one to popularize it, I think.”
- Publisher
- Latent Space
“Now, I would say that, you know, this is definitely an area of future research, especially if we talk about these problems that are going to take a human more than four hours.”
- Publisher
- Latent Space
“I think there's been a lot more progress in the software in the last few years. And I think a lot of the humanoid robot companies now are really trying to build amazing hardware.”
- Publisher
- Latent Space
“I think the downside is just that it adds a lot of complexity and it adds a lot of extra tokens.”
- Publisher
- Latent Space
“I think examples are really good in things like descriptions, like so many, like just using the Linux command line, like how many times I do like dash dash help or look at the man page or something.”
- Publisher
- Latent Space
“There's certain things that it's not very good at yet. But I'm really excited, I think, most broadly, not just for new things that weren't possible before, but as a much lower friction way to implement tool use.”
- Publisher
- Latent Space
“We really care about providing a nice sort of... Making the default safe, I think, is the best way for us to do it.”
- Publisher
- Latent Space
“I mean, this, so I'd say that a lot of previous agent work built sort of these very hard coded and rigid workflows where the model is sort of pushed through certain flows of steps. And I think to some extent, you know, that's needed with smaller models and models that are less smart.”
- Publisher
- Latent Space
“Like, you're kind of skeptical about self-driving as a business. So I want to double click on this a little bit, because I mean, I think that shouldn't be taken away.”
- Publisher
- Latent Space
“E2B is a close friend of ours that Alessio has led around in, but also I think there's others where they're focusing on snapshotting memory so that it can do time travel for debugging.”
- Publisher
- Latent Space
“How did you get, you joined Anthropic, did you already know you were going to work on of the stuff you publish or you kind of join and then you figure out where you land? I think people are always curious to learn more.”
- Publisher
- Latent Space
“One is like the, I'd say maybe like doing the easier solution rather than the hard solution. And I'd say the second one, I think what you're talking about is like the lazy model is like when the model says like dot, dot, dot, code remains the same.”
- Publisher
- Latent Space
“From the day-to-day job. But I think one of the most interesting things about SWE-Bench is that all these other benchmarks are usually just isolated puzzles, and you're starting from scratch.”
- Publisher
- Latent Space
“I think Cursor might have said that they actually have a separate model for file editing.”
- Publisher
- Latent Space
“I think Eider, they have a really good blog where they explore some of these different methods for editing files, and they post results about them, which I think is interesting.”
- Publisher
- Latent Space
“The way I think about this is that humans, even like very smart humans still use sort of checklists and use sort of scaffolding for themselves.”
- Publisher
- Latent Space
“For people who don't know, who maybe haven't dived into SWE-Bench, I think the general perception is they're like tasks that a software engineer could do.”
- Publisher
- Latent Space
“I think initially people were like, oh, maybe you're just getting lucky with XML.”
- Publisher
- Latent Space
“I think the really interesting question for me, for all the startups out there, is this kind of divergence between the benchmarks and what real customers will want.”
- Publisher
- Latent Space
“The question is, are those fairly inaccessible or are they just impossible because of the descriptions? But I think certainly some of the tasks, especially the ones that the human graders reviewed as like taking longer than four hours are extremely difficult.”
- Publisher
- Latent Space
“The other thing I was going to say is that SWE-Bench is certainly hard to implement and expensive to run because each task, you have to parse, you know, a lot of the repo to understand where to put your code.”
- Publisher
- Latent Space
“I think for agents in general, like having a planning step at the beginning, one, just having that plan will improve performance on the downstream task just because it's kind of like a bigger chain of thought, but also it's just such a better UX.”
- Publisher
- Latent Space
“You know, the classic idea is you can have a plug that can fit either way, and that's dangerous, or you can make it asymmetric so that it can't fit this way, it has to go like this, and that's a better tool because you can't use it the wrong way.”
- Publisher
- Latent Space
“I think another one is like, if you have a real coding agent, you don't want to have it start on a task and like spin its wheels for hours because you gave it a bad prompt.”
- Publisher
- Latent Space
“I think another area, at least the agent we created, didn't have any multimodal abilities, even though our models are very good at vision.”
- Publisher
- Latent Space
“I think we'll go into suite agent in a little bit, but I kind of reject the fact that, you know, you need to choose one prompt and like have your whole performance be predicated on that one prompt.”
- Publisher
- Latent Space
“I would say that as people are building systems around agents, I think the more you can separate out the different kinds of work the agent needs to do, the better you can tailor a prompt for that task. And I think that also creates a lot of like, for instance, if you were trying to make an agent that could both solve hard programming tasks, and it could just write quick test files for something that someone else had already made, the best way to do those two tasks might be very different prompts.”
- Publisher
- Latent Space
“We actually got acquired about six months ago, but I had left Cobalt about a year ago now, because I was starting to get a lot more excited about AI.”
- Publisher
- Latent Space
“I use XML in other models as well, and it's just a really nice way to make sure that the thing that ends is tied to the thing that starts. That's the only way to do code fences where you're pretty sure example one start, example one end, that is one cohesive unit.”
- Publisher
- Latent Space