Evidence receipt / evaluation
Published · transcript-backedBaris Gultekin: evaluation
14 Jan 2026 The Cognitive Revolution Snowflake VP of AI Baris Gultekin on Bringing AI to Data, Agent Design, Text-2-SQL, RAG & More
“I don't think that adoption of this technology requires human-like intelligence because Even for the specific things that these models and these applications do well, that is such high value that we're seeing huge adoption, as you all know, we're seeing huge adoption of AI already.”
Source trail
Everything needed to verify it.
- Speaker
- Baris Gultekin
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 14 Jan 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, interesting. On the sort of performance reliability side, my experience has often been And sometimes it's for good reason. Certainly in like the self-driving car realm, there's a certain logic to saying, we don't just want these things to be like roughly human level. We want them to be like clearly a step up before we're going to adopt them society wide. Good news. It seems like we're getting there. What do you, what do people have in mind as the intuitive standard of performance? Is it like they want these agents to be perfect? Is it that they want them to be like at the level of the human that used to do the job? Is there some like heuristic in the middle that you think people often land on? Yeah. Super interesting concept, right? The more natural the interface is, the more human-like intelligence we expect intuitively. If I'm talking to the agent versus typing, I think talking has much higher expectations, for instance, versus I'm just typing. I know I'm typing to a computer, so the expectations become a little less high. I don't think that adoption of this technology requires human-like intelligence because Even for the specific things that these models and these applications do well, that is such high value that we're seeing huge adoption, as you all know, we're seeing huge adoption of AI already. And then it keeps getting better and it keeps getting better at a super rapid phase, rapid pace. Yeah, I'm very excited about where the technology is. Before going into your expectations for the year ahead, what are you seeing in terms of guardrails. And obviously one big pattern that I think is like very natural to you guys is sourcing answers back to the document or the sort of authoritative place from which it came. Beyond that though, we've got this whole constellation of different patterns, right? In terms of you can filter inputs for appropriateness, you can filter outputs, you can log things and post-process logs. You can, AWS has, I think, a really interesting new service called automated reasoning checks. where you can put a policy in, they convert with a language model, your natural language policy into a set of rules and values. And then they do like literal formal methods to ensure that at runtime, like the agent or whatever the system, whatever it gave you back, that it actually passes those like formal reasoning checks that were derived originally from a natural language policy. That's like pretty interesting and pretty cutting edge from what I've seen. But I think in most places, my sense is like the frontier model companies are doing a ton of this stuff. Anthropic has pushed this to probably farther than anyone when it comes to preventing you from using Claude to do certain things in the biosphere. But are people at the enterprise level actually doing much of it? Or are they just saying, this thing seems to work, we've got an eval set, it passes, and we'll go with that?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.