Speakers in the public record
Claim mix
evaluation 5prediction 3belief 2
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
10 published records
“So we take the last couple of frames from our video. And we take the last couple of frames from our video attend that, along with the set of prompts that we provided, they could come from the future, [00:13:00] they could come from anywhere in the video, as well as reference object pointers, saying, by the way, here's what we've found so far attending to the last few frames has the interesting benefit of allowing it to model complex object motion without actually By limiting the amount of frames that you attend to, you manage to keep the model running in real time.”
- Publisher
- Latent Space
“This is so interesting because the the original diffusion transformer paper from Facebook actually showed that, in fact, the specific hyperparameters of the transformer didn't really matter that much.”
- Publisher
- Latent Space
“The best source of data we have is like image all text pairs on the internet and that's pretty low quality. So yeah, I, I think our solution here is really just we need to teach them how to operate on individual tasks and figure out how to scale that out.”
- Publisher
- Latent Space
“When we were thinking of ways to add value to our academic conference coverage, we realized that there was a lack of good talks, just recapping the best of 2024, going domain by domain.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“See these details. So it hypothesizes that models that have been initialized with, with Clip as their vision encoder, they don't have fine grained details and the, the features extracted using Clip because Clip sort of doesn't need to find these fine grained [00:22:00] details to do its job correctly, which is just to match captions and images, right?”
- Publisher
- Latent Space
“The, the, the, the bigger question, like, why isn't it transferring to object detection, especially like real time object detection. I think, in my mind, there are two answers.”
- Publisher
- Latent Space
“By relying on this long training cycle. And then LWdebtor also shows superior performance to our favorite data set, Roboflow 100 which means that they do better on the real world, not just on Cocoa.”
- Publisher
- Latent Space
“You know, if we just keep bumping the parameter count and increasing the example scene, which is the, the, the line of thinking for language models, then it'll keep getting better. So how does it actually do at finding, oh, it also improves with resolution, which you would expect for a model that This is the ImageNet classification accuracy, but yeah, it does better if you increase the resolution, which means that it's actually leveraging and finding fine grained visual features.”
- Publisher
- Latent Space
“Lava is, interestingly, extremely negatively correlated with this dataset. It does much, much, much, much worse [00:24:00] than random guessing, which means that this process has done a very good job of identifying hard images for, for Lava, specifically.”
- Publisher
- Latent Space
“I think the thing that's really exciting about this is it makes it possible for for developers to build using the 2B param [00:44:00] model and just explore, build their application, and then once they're ready to deploy figure out what exactly they need out of the model and prune those capabilities into a smaller form factor that makes sense for their deployment target.”
- Publisher
- Latent Space