app / uses
Claude
“I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case.”
Public evidence record
Published podcast speaker
Books, apps, and tools
app / uses
“I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case.”
Claim ledger
11 transcript-backed records
01 / belief
“I feel like there's always an explore-exploit trade-off, and you want to navigate those two thoughtfully, and sometimes you want to do one, sometimes you want to do the other. And maybe the cost of exploring has gone down a little bit, but I think there are still some things where you know you have to get it right the first time, or it's still, it's not all software engineering has literally gone to zero.”
02 / recommendation
“But now let me think about how, for example, like a sensitivity analysis, like how sensitive are my findings to different changes in input parameters, logical consistency checks, things like that. And so that's where I think we need a lot more investment in infrastructures and building kind of independent checks that don't don't rely just on checking chain of thought monitoring.”
03 / evaluation
“I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature.”
04 / recommendation
“I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case.”
05 / belief
“I think at the time, the models hadn't been trained at all to be faithful to a text.”
06 / belief
“I think one thing the longer context models changed for us is maybe a focus from breaking down tasks to breaking down the evaluation.”
07 / belief
“I think we also end up effectively monitoring by trying to evaluate new models as they come out.”
08 / belief
“I think GPT-4 unlocked tables for us, processing data from tables, which was huge.”
09 / evaluation
“He was saying that at a large, well-resourced hospital, like a city hospital, there might be a team of infectious disease specialists who can help interpret these results. But at under-resourced hospitals or more rural hospitals, the primary care physician can't interpret the test results, so then they can't order it, they can't use it, they can't help their patients with it.”
10 / evaluation
“I think GPT-3 was a big change because it kind of said, oh, now is the time that we can use AI to build these tools.”
11 / recommendation
“That's why we're launching this new set of features called Notebooks. It's very much inspired by computational notebooks, like Jupyter Notebooks, you know, DeepNode or Colab, because they're so powerful and so flexible.”