01 / belief
I think at the time, the models hadn't been trained at all to be faithful to a text.
“I think at the time, the models hadn't been trained at all to be faithful to a text.”
- Speaker
- Jungwon Byun
- Publisher
- Latent Space
Latent Space / episode intelligence
Speakers in the public record
Claim mix
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
29 published records
01 / belief
“I think at the time, the models hadn't been trained at all to be faithful to a text.”
02 / belief
“I think we have some specific requirements around being able to stream work back very quickly to our users.”
03 / belief
“In terms of like cost and compute, I think the closed models make up more of the budget since the main cases where you want to use closed models are cases where they're just smarter, where no existing open source models are quite smart enough.”
04 / belief
“In fact, I think you're doing an offsite and we're actually organizing our biggest AI UX meetup around whenever she's in town in San Francisco.”
05 / belief
“I think templates are a specific case of this where you're like, okay, well, there's just particular sequences of actions that you often want to chunk and have available as primitives, just like in normal programming.”
06 / belief
“I think if you can substantially improve how quickly people find new discoveries or avoid controlled trials that don't go anywhere, I think that's just huge amounts of money.”
07 / belief
“I think there may still be like some particular tools like for genomics, like particular types of modules of genes and proteins and whatnot.”
08 / belief
“I think you're about to start us on like GPT-3 and how that changed things for you.”
09 / belief
“I think one thing the longer context models changed for us is maybe a focus from breaking down tasks to breaking down the evaluation.”
10 / belief
“I think we also end up effectively monitoring by trying to evaluate new models as they come out.”
11 / uncertainty
“Yep. And then just to recap as well, like the models you were using back then were like, I don't know, would they like BERT type stuff or T5 or I don't know what timeframe we're talking about here.”
12 / belief
“I guess to be clear, at the very beginning, we had humans do the work. And then I think the first models that kind of make sense were TPT-2 and TNLG and like Yeah, early generative models.”
13 / belief
“I think in some cases, we also use two different models for the uncertainty estimates than for the question answering.”
14 / belief
“I think GPT-4 unlocked tables for us, processing data from tables, which was huge.”
15 / belief
“I think we'll probably want to think about more semantic pieces like a building block is more like a paper search or an extraction or a list of concepts.”
16 / belief
“I think the very first version probably used the Semantic Scholar or API or something similar.”
17 / belief
“I think the answers are actually a little bit clearer on the just kind of basic robustness side of where you can import ideas from normal software engineering and normal kind of DevOps.”
18 / belief
“I think the interesting insight that we got from talking to David Luan, who is CEO of multimodality has effectively two different flavors.”
19 / uncertainty
“You know, the fun thing you can do with a credit system, which is data for data, basically you can give people more credits if they give data back to you. I don't know if you've already done that.”
20 / evaluation
“He was saying that at a large, well-resourced hospital, like a city hospital, there might be a team of infectious disease specialists who can help interpret these results. But at under-resourced hospitals or more rural hospitals, the primary care physician can't interpret the test results, so then they can't order it, they can't use it, they can't help their patients with it.”
21 / evaluation
“In one sense, I think you're right that throw everything into the context window thing is easier to maintain because you just can swap out a model.”
22 / recommendation
“So specifically in the past, I think a lot of ranking was kind of per item ranking where you would score each individual item, maybe using increasingly expensive scoring methods and then rank based on the scores. But I think list-wise re-ranking where you have a model that can see all the elements is a lot more powerful because often you can only really tell how good a thing is in comparison to other things and what things should come first.”
23 / evaluation
“For all intents and purposes, the Elicit and also the Ought that existed before then were very different from what I started.”
24 / prediction
“I think in many ways, the approach is still the same because the way we are building illicit is not let's train a foundation model to do more stuff.”
25 / evaluation
“I think GPT-3 was a big change because it kind of said, oh, now is the time that we can use AI to build these tools.”
26 / recommendation
“That's why we're launching this new set of features called Notebooks. It's very much inspired by computational notebooks, like Jupyter Notebooks, you know, DeepNode or Colab, because they're so powerful and so flexible.”
27 / evaluation
“Because I just didn't find them reliable because they just hallucinated their own uncertainty.”
28 / evaluation
“I think if illicit hallucinates something about the abstract, that's really not good.”
29 / evaluation
“Basically, I think I'm very just impressed by how first principles, your ideas around what the workflow is. And I think that's why you're not as reliant on like the LLM improving, because it's actually just about improving the workflow that you would recommend to people.”