tool / uses
Zo
“I think there’s been one thing, I use another thing called zo, which is kinda like a cloud computer plus agent.”
Public evidence record
Partner · Decibel
Books, apps, and tools
tool / uses
“I think there’s been one thing, I use another thing called zo, which is kinda like a cloud computer plus agent.”
tool / likes
“First of all, you have very good support for mocking in unit tests, which is something that a lot of other frameworks don't do. So, you know, my favorite Ruby library is VCR because it just, you know, it just lets me store the HTTP requests and replay them.”
other / likes
“Yeah, I'm a big fan of Simulative AI. We had a summer of Simulative AI. Another term we're trying to coin.”
Claim ledger
43 transcript-backed records
01 / evaluation
“I was gonna say, to me, driving feels like a great next token prediction thing because you’re kinda like on a path and like, it doesn’t really matter what you’ve done before.”
02 / evaluation
“The problem is that you need twice the amount of ram, twice the amount of, you know, it’s like, it’s kind of taxing on the machine.”
03 / evaluation
“From a workload perspective, you’re thinking this is gonna be like a read heavy thing because they’re doing recommend.”
04 / evaluation
“I think there's the other side of MCPs that people don't talk as much about because it doesn't go viral, which is building the servers.”
05 / evaluation
“And then we got computer use. Which I think Operator was obviously one of the hot releases of the year.”
06 / evaluation
“Yeah. And the big thing was like the mixed trial price fights, you know, and I think now it's almost like there's nowhere to go, like, you know, Gemini Flash is like basically giving it away for free.”
07 / evaluation
“Um, and I think today the, the problem is that, Yeah, the agents are, that most people are building are good at following instruction, but are not as good as like extracting them from you.”
08 / evaluation
“I think to me that the most interesting is like rest and GraphQL is almost more interesting in the world of agents because agents could come up with so many different things to query versus like before I always thought GraphQL was kind of like not really necessary because like, you know what you need, just build the rest end point for it.”
09 / evaluation
“I think inference time compute is bad for open source just because, you know, Doc can donate the flops at training time, but he cannot donate the flops at inference time.”
10 / evaluation
“Was there kind of like a threshold of employees and team size where you felt like, okay, maybe that worked. Now it doesn't work anymore.”
11 / evaluation
“Because there was a manifesto for responsible AI that hundreds of VCs and people signed and I don't think anybody actually thinks about it anymore.”
12 / evaluation
“And eventually, you know, I started doing venture six, five years ago. And I think just like so many people in Europe reach out and ask, hey, can you like talk to our team and they just cannot comprehend like the risk appetite that people have here.”
13 / evaluation
“When I tried to set up Slack, it was like, hey, give me access to all channels and everything, which for the average person probably makes sense because you don't want to re-prompt them every time you add new channels.”
14 / evaluation
“I've written a post called Maximum Enterprise Utilization, kind of like you have MFU for GPUs, but it's basically like so many people are focused on, oh, it's going to like displace jobs and whatnot. But I'm like, there's so much work that people don't do because they don't have the people.”
15 / evaluation
“You know, I think even today people will tell you, oh, models are not really good at X because they were not good 12 months ago, but they're good today.”
16 / evaluation
“You have a lot of companies that sound the same, but like none of them are really working. So obviously the problem is not solved.”
17 / evaluation
“I think today it's almost like, hey, if I can use Dash to like access my Google Drive file, why would I pay Google for like their AI feature?”
18 / evaluation
“It's not, it's not what gives us our edge, but it certainly means that then we don't have to build it and maintain it afterwards. So, it's a really good first step, I think, in, like, the overall maturity of the fine tuning product and API in terms of where they're going to see those early products.”
19 / evaluation
“So obviously there's a lot of interest. And I think some of the initial jailbreaks, I got fine-tuned back into the model, obviously they don't work anymore.”
20 / evaluation
“I think there was obviously the Kepler, and then there was Chinchilla, and then people kind of got the Llama scaling law, like the 100 to 200x parameter to token ratio.”
21 / evaluation
“Are people just finding out recently about these problems because now the scores are getting so high that you're actually inspecting the benchmarks and maybe in the past you were scoring so badly that maybe you weren't as worried about the overall quality?”
22 / evaluation
“I think the other thing to talk about here is whether or not humans are good at judging and evaluating these models.”
23 / evaluation
“For example, with one system, there are like 780 endpoints. And if you're actually trying to do vector similarity, it's not that good because the people that wrote the specs didn't have in mind making them like semantically apart.”
24 / evaluation
“People were like, okay, this is better than this on this benchmark, blah, blah, blah, because maybe they did not have a lot of use cases that they did frequently.”
25 / evaluation
“Cloud is better. It's very good, you know, it's much better, it seems to me, it's much better than GPT 4 at doing writing that is more, you know, I don't know, it just got good vibes, you know, like the GPT 4 text, you can tell it's like GPT 4, you know, it's like, it always uses certain types of words and phrases and, you know, maybe it's just me because I've now done it for, you know, So, I've read like 75, 80 generations of these things next to each other.”
26 / evaluation
“I think before we wrap, you have written a blog post that can show about good hearts law impact in ML, which is, you know, when you measure something, then the thing that you measure is not a good metric anymore because people optimize for it.”
27 / evaluation
“I think if anything, RAG's complexity goes up and up the more you use it, you know, because you have more data sources, more things you want to put in there.”
28 / evaluation
“I think in Europe, I walked through a lot of the posters and whatnot, there seems to be mode collapse in a way in the research, a lot of people working on the same things.”
29 / evaluation
“I think people come on your website today and they say, you raised a hundred million dollars Series A.”
30 / evaluation
“I know we kind of binned the lightning round in the last few episodes, but I think for you two, one of the questions we used to ask is like, what's the most interesting unsolved question in AI?”
31 / evaluation
“I was mostly helping people write their own code, you know, so even if you have the best inline completion, it doesn't help me do my job.”
32 / evaluation
“I think the cool things about these models is like people that are not traditionally technical can do a lot of very advanced things.”
33 / evaluation
“I think, to me, the biggest takeaway was like and I was talking with Mike Conover, another friend of the podcast, about this is they're kind of staying in the single threaded, like, synchronous use cases lane, you know?”
34 / evaluation
“So I think the easiest thing for people to grasp so far has been Mojo, which is a superset of Python. And I think everybody talks about that because it's easier to grasp, but Modular's goal is to build a unified AI engine.”
35 / evaluation
“I think in one of your previous podcasts, you mentioned leaving people behind, you know, that are like not experts in certain things and they can't contribute.”
36 / evaluation
“There's a bunch of vector databases that are killing each other out there to get people to embed data in them, and you're like, I love you all.”
37 / evaluation
“If you have a eight bits model quantized down, you need one byte per parameter. So for example, in an H100, which is 80 gigabyte of memory, you could fit a 70 billion parameters in eight, you cannot fit a FP32 because you will need like 280 gigabytes of memory.”
38 / evaluation
“I think I've talked about this on the podcast, but this idea of like just-in-time UIs, you know, like each type of user wants to interact in a different way.”
39 / evaluation
“I think I talked about it on the podcast before, but like the switch from syntax to like semantics, like developers used to be focused on the syntax and not the meaning of what they're writing.”
40 / evaluation
“I'm building an agent internally for us. And Guardrails are obviously very exciting because once you set the initial prompt, like the model creates its own prompts.”
41 / evaluation
“Yeah. This is super interesting because right now a lot of products are kind of the same because all I do is they call it the model and some are prompted a little differently, but you can only guess so much delta between them in the future.”
42 / evaluation
“I think that's gonna look really different in opensource models because just hosting a model doesn't have a lot of value.”
43 / evaluation
“The question is, if you wanna compete against these companies, maybe the model is not what you're gonna do it with because the open source kind of commoditizes it.”