Andrej Karpathy observes that large language models lack a distillation phase equivalent to human sleep, where experiences are analyzed and distilled into weights.
We don't have an equivalent of that in large language models. That's to me more adjacent to when you talk about continual learning and so on as absent.