Pretraining
This video presents the same text shown beside it, spoken and on screen. It adds nothing the text does not say.
State
A large language model is a transformer pretrained on next-word prediction over vast text collections — billions of knobs descending against prediction error.
Show
More training text than a person reads in ten thousand lifetimes; grammar, style, idiom, and fact-shaped regularity fall out of one objective.
Watch for
"Fact-shaped" is the honest term — descent rewards producing the text that typically follows, which overlaps truth imperfectly.