High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Andreas Stuhlmüller

Published podcast speaker

Claims
28
Episodes
2
Shows
2
Named items
0

Claim ledger

What Andreas said.

28 transcript-backed records

01 / recommendation

Mentions personal use of systematic literature view API.

“When I run systematic literature views, often I kind of use the systematic literature view API and I iterate using like various other models on the protocol for a while, and then I run it in the background and retrieve it.”
Publisher
The Cognitive Revolution

03 / belief

People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.

“People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.”
Publisher
The Cognitive Revolution

04 / belief

I still find, especially as a company that is trying to be extremely mission focused and how do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality, I think by default, you're probably just not going to accomplish that.

“I still find, especially as a company that is trying to be extremely mission focused and how do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality, I think by default, you're probably just not going to accomplish that.”
Publisher
The Cognitive Revolution

07 / belief

The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data.

“The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data.”
Publisher
The Cognitive Revolution

08 / disagreement

I think sometimes people like come to us and are like, well, who's going to win in AI for science? And I think that's just an absurd thing to say because science is such a big space.

“I think sometimes people like come to us and are like, well, who's going to win in AI for science? And I think that's just an absurd thing to say because science is such a big space.”
Publisher
The Cognitive Revolution

09 / preference

I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc.

“I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc.”
Publisher
The Cognitive Revolution

10 / evaluation

I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time.

“I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time.”
Publisher
The Cognitive Revolution

11 / evaluation

I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal.

“I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal.”
Publisher
The Cognitive Revolution

12 / evaluation

These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning.

“These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning.”
Publisher
The Cognitive Revolution

13 / evaluation

Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process.

“Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process.”
Publisher
The Cognitive Revolution

15 / belief

In terms of like cost and compute, I think the closed models make up more of the budget since the main cases where you want to use closed models are cases where they're just smarter, where no existing open source models are quite smart enough.

“In terms of like cost and compute, I think the closed models make up more of the budget since the main cases where you want to use closed models are cases where they're just smarter, where no existing open source models are quite smart enough.”
Publisher
Latent Space

16 / belief

I think templates are a specific case of this where you're like, okay, well, there's just particular sequences of actions that you often want to chunk and have available as primitives, just like in normal programming.

“I think templates are a specific case of this where you're like, okay, well, there's just particular sequences of actions that you often want to chunk and have available as primitives, just like in normal programming.”
Publisher
Latent Space

17 / belief

I think if you can substantially improve how quickly people find new discoveries or avoid controlled trials that don't go anywhere, I think that's just huge amounts of money.

“I think if you can substantially improve how quickly people find new discoveries or avoid controlled trials that don't go anywhere, I think that's just huge amounts of money.”
Publisher
Latent Space

19 / belief

I guess to be clear, at the very beginning, we had humans do the work. And then I think the first models that kind of make sense were TPT-2 and TNLG and like Yeah, early generative models.

“I guess to be clear, at the very beginning, we had humans do the work. And then I think the first models that kind of make sense were TPT-2 and TNLG and like Yeah, early generative models.”
Publisher
Latent Space

21 / belief

I think we'll probably want to think about more semantic pieces like a building block is more like a paper search or an extraction or a list of concepts.

“I think we'll probably want to think about more semantic pieces like a building block is more like a paper search or an extraction or a list of concepts.”
Publisher
Latent Space

23 / belief

I think the answers are actually a little bit clearer on the just kind of basic robustness side of where you can import ideas from normal software engineering and normal kind of DevOps.

“I think the answers are actually a little bit clearer on the just kind of basic robustness side of where you can import ideas from normal software engineering and normal kind of DevOps.”
Publisher
Latent Space

24 / evaluation

In one sense, I think you're right that throw everything into the context window thing is easier to maintain because you just can swap out a model.

“In one sense, I think you're right that throw everything into the context window thing is easier to maintain because you just can swap out a model.”
Publisher
Latent Space

25 / recommendation

So specifically in the past, I think a lot of ranking was kind of per item ranking where you would score each individual item, maybe using increasingly expensive scoring methods and then rank based on the scores. But I think list-wise re-ranking where you have a model that can see all the elements is a lot more powerful because often you can only really tell how good a thing is in comparison to other things and what things should come first.

“So specifically in the past, I think a lot of ranking was kind of per item ranking where you would score each individual item, maybe using increasingly expensive scoring methods and then rank based on the scores. But I think list-wise re-ranking where you have a model that can see all the elements is a lot more powerful because often you can only really tell how good a thing is in comparison to other things and what things should come first.”
Publisher
Latent Space

27 / prediction

I think in many ways, the approach is still the same because the way we are building illicit is not let's train a foundation model to do more stuff.

“I think in many ways, the approach is still the same because the way we are building illicit is not let's train a foundation model to do more stuff.”
Publisher
Latent Space
Search evidence