Andrej Karpathy predicts that future AI architectures will converge on cognitive tricks similar to those evolution developed, such as sparse attention and modified MLPs.
But we're going to converge on a similar architecture cognitively. In 10 years, do you think it'll still be something like a transformer, but with much more modified attention and more sparse MLPs and so forth?
Named things
DeepSeek v3.2 · tool · mentions