LLMs are trained on approximately 10^13 tokens, equivalent to 2x10^13 bytes of training data, which would take a human 170,000 years to read at eight hours a day.
those LLMs are trained on enormous amounts of text. Basically the entirety of all publicly available text on the internet, right? That's typically on the order of 10 to the 13 tokens. Each token is typically two bytes. So that's two 10 to the 13 bytes as training data. It would take you or me 170,000 years to just read through this at eight hours a day.