01 / observation
It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great.
“It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great.”
- Speaker
- Bronson Schoen
- Episode
- RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
- Publisher
- The Cognitive Revolution