High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Aman Sanger

Published podcast speaker

Claims
15
Episodes
1
Shows
1
Named items
0

Claim ledger

What Aman said.

2 transcript-backed records

01 / evaluation

We found that it was not just good at creating net new things, but refactoring code, editing code, helping you debug kind of every single aspect of software development felt so different with these models.

“We found that it was not just good at creating net new things, but refactoring code, editing code, helping you debug kind of every single aspect of software development felt so different with these models.”
Speaker
Aman Sanger
Publisher
Latent Space

02 / evaluation

I've been meaning to do it at some point, but there's this paper called Babel code and they have a library which I think literally translates human eval into all other languages. And I think that would be a really good test because the other issues, a lot of the models that perform really well on human eval are pure Python, right?

“I've been meaning to do it at some point, but there's this paper called Babel code and they have a library which I think literally translates human eval into all other languages. And I think that would be a really good test because the other issues, a lot of the models that perform really well on human eval are pure Python, right?”
Speaker
Aman Sanger
Publisher
Latent Space
Search evidence