Skip to main content

Pretraining

This video presents the same text shown beside it, spoken and on screen. It adds nothing the text does not say.

State

A large language model is a transformer pretrained on next-word prediction over vast text collections — billions of knobs descending against prediction error.

Show

More training text than a person reads in ten thousand lifetimes; grammar, style, idiom, and fact-shaped regularity fall out of one objective.

Watch for

"Fact-shaped" is the honest term — descent rewards producing the text that typically follows, which overlaps truth imperfectly.