High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Baris Gultekin: evaluation

14 Jan 2026 The Cognitive Revolution Snowflake VP of AI Baris Gultekin on Bringing AI to Data, Agent Design, Text-2-SQL, RAG & More

“80% to 90% of all data is unstructured data. And because there weren't a lot of easy ways to process this, it was not necessarily seen as the most usable data, and it is now very usable.”

— Baris Gultekin

Source trail

Everything needed to verify it.

Speaker
Baris Gultekin
Attribution
Verified speaker
Claim type
evaluation
Recorded
14 Jan 2026
Publisher
The Cognitive Revolution

Transcript context

…Could you give a sense for sort of the balance of structured, unstructured, or then I guess you're also saying that like structured data is getting structured through the process of basically AI retro annotation. My sense is that I think of, oh gosh, there's so many vector databases. The exact name of the one where the founder told me, this is slipping my mind. But one of the interesting things that I've understood to be happening in general with business data is that structured data was kind of like the tip of the iceberg in many organizations where it was the most usable kind, but it was actually like a relatively small amount of the data. It was Anton from ChromaDB who said that most of the data that was going into ChromaDB had never been in a database before at all. It was just lying around in various places. So have you seen a sort of great unlocking of people dumping more and more data into Snowflake because now they have ways to make it useful where it just previously wasn't even worth it? Yeah, we're absolutely seeing this. 80% to 90% of all data is unstructured data. And because there weren't a lot of easy ways to process this, it was not necessarily seen as the most usable data, and it is now very usable. both from a kind of extract and then bring structure to it perspective, as well as just talk to all of that data, find the right information using vector DBs, for instance, and then build agents and chat experiences off of it. So we're seeing this, and it is, again, playing out in both ways. More and more data is getting structured so that you could run analytics on it down the road. as well as you can just use all of this data in conjunction with the structured data that you have. So for instance, if you want to build anything, let's say a wealth management agent, so you still need to be able to look up what the stocks are doing in a structured way, but you also have all of the equities research that's in PDFs that you'd like to be able to use. So being able to combine both structured and unstructured is incredibly important for real-world use cases. Where would you say we are on text-to-sequel today? It's been a while since I've done a show on text-to-sequel. There have been probably two episodes on this theme, historically. And... I guess last I checked, there was a range of opinions where some people were like, Yeah, it's just not really there. Other people is there, but you have to do a lot of work to make sure that you have a good semantic understanding because a lot of the SQL databases that come in are like, there's multiple columns and there's tribal knowledge on teams that we don't use that column anymore. We haven't deleted it, but we don't use it anymore and it's superseded by this. There's all these little nuances that live in people's heads. if they have enough of that kind of context. What does a process look like today and how good does it get if new customers, I want to start to enable these, talk to my data with a sort of text-to-SQL kind of strategy. What does that look like now?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence