High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Dries Smit

Published podcast speaker

Claims
10
Episodes
1
Shows
1
Named items
0

Claim ledger

What Dries said.

4 transcript-backed records

01 / evaluation

Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet?

“Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet?”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

02 / evaluation

So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit.

“So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

03 / evaluation

We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position.

“We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

04 / evaluation

Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up.

“Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk
Search evidence