High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Latent Space / episode intelligence

Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit

11 Apr 2024 29 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 17evaluation 7uncertainty 2recommendation 2prediction 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

29 published records

03 / belief

In terms of like cost and compute, I think the closed models make up more of the budget since the main cases where you want to use closed models are cases where they're just smarter, where no existing open source models are quite smart enough.

“In terms of like cost and compute, I think the closed models make up more of the budget since the main cases where you want to use closed models are cases where they're just smarter, where no existing open source models are quite smart enough.”
Publisher
Latent Space

05 / belief

I think templates are a specific case of this where you're like, okay, well, there's just particular sequences of actions that you often want to chunk and have available as primitives, just like in normal programming.

“I think templates are a specific case of this where you're like, okay, well, there's just particular sequences of actions that you often want to chunk and have available as primitives, just like in normal programming.”
Publisher
Latent Space

06 / belief

I think if you can substantially improve how quickly people find new discoveries or avoid controlled trials that don't go anywhere, I think that's just huge amounts of money.

“I think if you can substantially improve how quickly people find new discoveries or avoid controlled trials that don't go anywhere, I think that's just huge amounts of money.”
Publisher
Latent Space

11 / uncertainty

Yep. And then just to recap as well, like the models you were using back then were like, I don't know, would they like BERT type stuff or T5 or I don't know what timeframe we're talking about here.

“Yep. And then just to recap as well, like the models you were using back then were like, I don't know, would they like BERT type stuff or T5 or I don't know what timeframe we're talking about here.”
Speaker
Shawn Wang
Publisher
Latent Space

12 / belief

I guess to be clear, at the very beginning, we had humans do the work. And then I think the first models that kind of make sense were TPT-2 and TNLG and like Yeah, early generative models.

“I guess to be clear, at the very beginning, we had humans do the work. And then I think the first models that kind of make sense were TPT-2 and TNLG and like Yeah, early generative models.”
Publisher
Latent Space

15 / belief

I think we'll probably want to think about more semantic pieces like a building block is more like a paper search or an extraction or a list of concepts.

“I think we'll probably want to think about more semantic pieces like a building block is more like a paper search or an extraction or a list of concepts.”
Publisher
Latent Space

17 / belief

I think the answers are actually a little bit clearer on the just kind of basic robustness side of where you can import ideas from normal software engineering and normal kind of DevOps.

“I think the answers are actually a little bit clearer on the just kind of basic robustness side of where you can import ideas from normal software engineering and normal kind of DevOps.”
Publisher
Latent Space

19 / uncertainty

You know, the fun thing you can do with a credit system, which is data for data, basically you can give people more credits if they give data back to you. I don't know if you've already done that.

“You know, the fun thing you can do with a credit system, which is data for data, basically you can give people more credits if they give data back to you. I don't know if you've already done that.”
Speaker
Shawn Wang
Publisher
Latent Space

20 / evaluation

He was saying that at a large, well-resourced hospital, like a city hospital, there might be a team of infectious disease specialists who can help interpret these results. But at under-resourced hospitals or more rural hospitals, the primary care physician can't interpret the test results, so then they can't order it, they can't use it, they can't help their patients with it.

“He was saying that at a large, well-resourced hospital, like a city hospital, there might be a team of infectious disease specialists who can help interpret these results. But at under-resourced hospitals or more rural hospitals, the primary care physician can't interpret the test results, so then they can't order it, they can't use it, they can't help their patients with it.”
Speaker
Jungwon Byun
Publisher
Latent Space

21 / evaluation

In one sense, I think you're right that throw everything into the context window thing is easier to maintain because you just can swap out a model.

“In one sense, I think you're right that throw everything into the context window thing is easier to maintain because you just can swap out a model.”
Publisher
Latent Space

22 / recommendation

So specifically in the past, I think a lot of ranking was kind of per item ranking where you would score each individual item, maybe using increasingly expensive scoring methods and then rank based on the scores. But I think list-wise re-ranking where you have a model that can see all the elements is a lot more powerful because often you can only really tell how good a thing is in comparison to other things and what things should come first.

“So specifically in the past, I think a lot of ranking was kind of per item ranking where you would score each individual item, maybe using increasingly expensive scoring methods and then rank based on the scores. But I think list-wise re-ranking where you have a model that can see all the elements is a lot more powerful because often you can only really tell how good a thing is in comparison to other things and what things should come first.”
Publisher
Latent Space

24 / prediction

I think in many ways, the approach is still the same because the way we are building illicit is not let's train a foundation model to do more stuff.

“I think in many ways, the approach is still the same because the way we are building illicit is not let's train a foundation model to do more stuff.”
Publisher
Latent Space

26 / recommendation

That's why we're launching this new set of features called Notebooks. It's very much inspired by computational notebooks, like Jupyter Notebooks, you know, DeepNode or Colab, because they're so powerful and so flexible.

“That's why we're launching this new set of features called Notebooks. It's very much inspired by computational notebooks, like Jupyter Notebooks, you know, DeepNode or Colab, because they're so powerful and so flexible.”
Speaker
Jungwon Byun
Publisher
Latent Space

29 / evaluation

Basically, I think I'm very just impressed by how first principles, your ideas around what the workflow is. And I think that's why you're not as reliant on like the LLM improving, because it's actually just about improving the workflow that you would recommend to people.

“Basically, I think I'm very just impressed by how first principles, your ideas around what the workflow is. And I think that's why you're not as reliant on like the LLM improving, because it's actually just about improving the workflow that you would recommend to people.”
Speaker
Shawn Wang
Publisher
Latent Space
Search evidence