2023
Sparks of Artificial General Intelligence: Early experiments with GPT-4
The field looking hard at GPT-4 and daring to ask whether it had built something general.
“We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more.”
Beyond benchmarks
Bubeck et al. did not rely on a single score. They probed GPT-4 with tasks requiring planning, abstraction, and tool use. The paper documents behaviors that look like reasoning even when the mechanism is opaque.
The AGI framing
The title provoked debate. Critics said the evaluation was anecdotal; supporters said standard benchmarks had already saturated. The paper's value is descriptive: it catalogues what a frontier model could do in early 2023 before the public had access.
What remains open
"Sparks" is the careful word. The model still hallucinates, forgets, and fails on adversarial prompts. The paper asks whether capability clusters into general intelligence or a mosaic of narrow tricks. The question outlived the hype.
- collection
The idea lineage of the model you talked to this morning, in reading order.
← previous · 2022
Training Language Models to Follow Instructions with Human Feedback
next · 2024 →
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?