2020
Language Models are Few-Shot Learners
GPT-3: the moment scale alone became capability, and prompting replaced fine-tuning.
“We train GPT-3, an autoregressive language model with 175 billion parameters, and test its performance in the few-shot setting.”
Scale as capability
Brown et al. trained GPT-3 at 175B parameters and found that many tasks improve smoothly with scale. The model does not need fine-tuning for every new benchmark. A short prompt with a few examples is often enough.
In-context learning
The surprising behavior is few-shot learning: the model reads a task description and a handful of input-output pairs in its context window, then completes the next item. No weight update occurs. The task is inferred at inference time.
What it changed
GPT-3 made scale feel like a product strategy. It also exposed the limits of prompt-only learning: strong on many benchmarks, brittle on others, and expensive to run. Every later frontier model inherits this paper's question: what does scale buy you, and what still requires training?
- collection
The idea lineage of the model you talked to this morning, in reading order.
← previous · 2019
The Bitter Lesson
next · 2020 →
Scaling Laws for Neural Language Models