#fundamentals #ai-history

A Short History of AI: What Changed, What Didn't

Mathieu Hadjimegrian Mathieu Hadjimegrian 7 min

The wish for thinking machines is older than modern computing: it appears in stories, philosophy, and mechanical calculation. But those cultural precedents are not a scientific lineage; the modern history of AI begins when computation became formal enough to describe, test, and engineer.

That distinction matters. AI is often narrated as a rush of breakthroughs leading naturally to a machine that thinks like a person. The record is less tidy. Each important system made a particular capability practical under particular conditions. Each also left a boundary behind: a task it could not transfer, a context it could not represent, a claim it could not justify. The history is not a march from “not intelligent” to “intelligent.” It is a history of changing engineering bargains.

Selected milestones in AI development from 1943 to 2022, a visual preview.

A neuron becomes a formal component

In 1943, Warren McCulloch and Walter Pitts published A Logical Calculus of the Ideas Immanent in Nervous Activity. Their model treated a simplified neuron as a threshold-like logical unit: given inputs, it could produce an output according to a formal rule. That was a consequential abstraction. It made networks of neuron-like components something that could be reasoned about mathematically rather than merely imagined biologically.

What it could do was clarify a possible computational language for networks. What it could not do was learn from experience, model a whole brain, or understand an environment. The connectionist tradition begins here as an abstraction, not as a proof that intelligence had been reconstructed.

Seven years later, Alan Turing changed the terms of a related debate. In Computing Machinery and Intelligence, he did not settle the meaning of “thinking.” Instead, he proposed the imitation game: an operational way to ask what a machine’s behaviour might lead an observer to conclude.

The move was practical and still useful. It redirected an unbounded philosophical question towards observable interaction. But an interaction test is not a complete account of intelligence, and Turing did not call the field “artificial intelligence.” A system can produce convincing behaviour in a setting without that behaviour establishing how broadly it can reason, what it knows, or how reliably it acts outside the setting.

The name arrived with a research programme. In a proposal dated 31 August 1955, John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon proposed the Dartmouth Summer Research Project on Artificial Intelligence; the workshop took place in the summer of 1956. The phrase gave related work a shared field and a shared ambition.

That was a real institutional capability: researchers could now gather questions about reasoning, learning, language, and machines beneath a recognisable label. It was not the arrival of a generally intelligent machine. Early symbolic programs illustrated both the promise and the boundary of the approach: encode a formal problem, search within it, and obtain a result. A program can be powerful inside a formal representation without possessing broad, transferable reasoning.

When conversation looks like understanding

In 1966, Joseph Weizenbaum published ELIZA, a program designed to study natural-language communication between people and machines. Its DOCTOR script simulated a Rogerian psychotherapist through pattern matching and transformations of a user’s words.

ELIZA could sustain a recognisably conversational exchange with remarkably little machinery. That mattered because it exposed a durable human tendency: fluent surface interaction can invite an inference of depth. Yet the program did not ground language in a world model, maintain a clinical understanding of the person speaking, or provide a reason to infer a therapist—or a mind—behind the exchange. It matched and transformed patterns. The gap between a compelling interface and a warranted conclusion about capability would recur throughout AI’s history.

The following decades also made another lesson hard to avoid: demonstrations and dependable systems are not the same thing. Historiography commonly identifies periods described as “AI winters,” after waves of expectation outpaced the capabilities, resources, and reliability available to programmes at the time. The dates and causes should not be treated as a single rigid story. The larger point is more useful: enthusiasm does not remove engineering constraints.

Expert systems showed both sides of that lesson. They encoded explicit IF–THEN rules obtained from human experts and could be useful in narrow, carefully bounded domains. Their strength was legibility. A rule could be inspected, discussed, and adjusted. Their weakness was also structural: knowledge acquisition, exceptions, maintenance, and transfer beyond the domain were difficult. Symbolic reasoning did not disappear; it demonstrated that explicit structure is valuable when a problem can be represented clearly. It did not, however, turn a collection of rules into general competence.

A spectacular win, still a narrow one

Deep Blue made the distinction between achievement and generality visible to a global audience. In the 1997 New York rematch, IBM’s chess system defeated Garry Kasparov by 3½–2½. The result was striking, and it should not be minimised. Deep Blue solved a demanding competitive task through specialised chess search and engineered evaluation.

But its accomplishment was precisely bounded by that design. A win at chess did not show that the system could reason generally, move knowledge into an unrelated domain, or understand the world in the human sense. This is the discipline a history of AI requires: a result can be both extraordinary and specialised. Treating every high-profile benchmark as evidence of a general mind turns an engineering fact into a philosophical overclaim.

Scale makes neural methods practical

By the 2000s, the practical question had increasingly become one of learning from examples at scale. Statistical learning offered ways to extract regularities from data rather than specify every rule in advance. That did not displace symbolic methods everywhere; different representations remained useful for different tasks. But more available data, more computation, and architectures suited to the task made some previously difficult approaches operationally competitive.

AlexNet provided a clear, measurable marker in 2012. Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton entered a deep convolutional network in the ImageNet Large Scale Visual Recognition Challenge. Their paper reports a winning top-5 test error of 15.3%, compared with 26.2% for the next-best entry, using two NVIDIA GTX 580 GPUs.

The capability was concrete: on a defined visual-classification benchmark, a deep network combined with suitable data and GPU computation produced a measured improvement on that benchmark. This was an important moment for the connectionist tradition because it showed what deep networks could do under material conditions that earlier decades often lacked.

Its limit is equally concrete. A benchmark score is not general perception. It is bounded by its dataset, its labels, its evaluation metric, and the conditions represented in training and testing. AlexNet did not solve vision, erase competing methods overnight, or establish that a system could interpret every visual situation robustly. It made a particular kind of progress visible—and measurable.

Learning, search, and attention

AlphaGo made a different combination of methods visible in 2016. Google DeepMind’s system defeated Lee Sedol 4–1 in Seoul. The technical account describes a system that combined supervised learning from human games, reinforcement learning through self-play, and Monte Carlo Tree Search.

The system could play Go at an exceptional level by joining learned representations with structured search. That combination matters: learned systems need not operate without explicit search, and statistical learning can work alongside other computational techniques. But AlphaGo was not a general player, a general scientist, or an autonomous agent outside the problem it was engineered to solve. Its strength was an example of focused system design, not evidence that task boundaries had ceased to matter.

In 2017, Attention Is All You Need introduced the Transformer, an architecture based on attention rather than recurrence or convolution. The paper evaluated the model first on machine translation. This is an important boundary to preserve: it was not presented as a chatbot paper, and later developments should not be projected backwards as if they were already inevitable.

Still, the Transformer became a foundational architecture for later large language models. It is best understood as a bridge in the history: an architectural building block that later systems could use at larger scales and in different products. The three traditions in this story did not collapse into a single winner. Symbolic approaches continued to offer explicit structure; neural methods learned useful representations; statistical methods organised learning from data. Modern systems often reflect tensions and combinations among them rather than a clean replacement of one by another.

Three AI approaches shown as symbolic reasoning, neural methods and statistical learning; visual preview.

Modern AI systems combine methods; they do not erase constraints.

2022: a long history becomes widely visible

On 30 November 2022, OpenAI released ChatGPT as a free research preview. Its announcement described training using Reinforcement Learning from Human Feedback (RLHF), using the same methods as InstructGPT. The release mattered because a conversational interface made an accumulated technical lineage accessible to a far wider public.

What changed in public view was access and format, not the birth of language models. Conversational ease did not establish factual reliability, general intelligence, or autonomous agency. It made the old ELIZA lesson newly urgent: a system’s fluent interface is evidence that it can produce fluent interaction; it is not, by itself, evidence for every capability people may infer from that interaction.

Public bibliography

Related Articles