An LLM does not come from nowhere. The first ChatGPT gave us a way to converse with a model from the GPT-3.5 family. But GPT was already doing something remarkable before it became a chat interface: learning to continue a piece of text. How does predicting what comes next turn into answering a question?
The answer takes us through deep learning, its wider family of machine learning, and natural language processing (NLP), the work of making computers handle human language. Let's start inside the model and follow those ideas towards GPT.

Deep learning: patterns through layers
Compare “I charged my phone” with “The café charged my card”. The same word means something different in each sentence, and you recognise which meaning fits without consulting a list of rules. A deep-learning model learns from many examples instead. It processes text through layers of calculations. During training, it adjusts numerical values called weights, so it can respond differently to “phone” and “card” without a hand-written rule for every use of charged.
That is what deep describes: layers of computation, not depth of understanding. It is a way to learn patterns in data. It can help with language, but it could just as well be used for images or sound. Why call this learning at all?
Machine learning: the larger family
Machine learning (ML) means fitting a model to data so it can make predictions on new examples. Imagine teaching an email filter with messages marked “spam” or “not spam”. It learns from those examples, then you see how it handles messages it has never seen. It need not contain a deep neural network to be useful.
Deep learning is one approach within ML: the model has many layers, and training adjusts their weights. So deep learning is not a step that happened before machine learning. We have simply started close to GPT and stepped back to see its family. Both belong to artificial intelligence (AI), but neither tells us why GPT works with language in particular.
NLP: the work being done with language
That is where NLP comes in. It is the field concerned with computational work on human language: recognising names, classifying a message, translating a sentence or generating one. Those jobs are different. They do not all require the same kind of model.
Stanford CoreNLP shows one way to tackle them: it can split a sentence into pieces, spot names and dates, and analyse its grammar. Some language tasks use rules, some use conventional ML, and some use deep learning. CoreNLP is not GPT. It helps make the distinction visible: NLP is the problem space; a model is one way of working in it.
One NLP question brings us closer to ChatGPT: given a piece of text, what might come next? A language model learns to assign probabilities to possible continuations. Older models could do this with counts of short word sequences. Neural language models learn from a much wider range of examples. GPT belongs to this latter family.
Where they meet: the language model
Now the pieces fit. GPT is a deep-learning model (within ML) built for language work, the subject of NLP. The letters stand for Generative Pre-trained Transformer. Generative means it can produce text; pre-trained means it has learnt from text before you use it; Transformer names the neural-network architecture it uses. Its attention mechanism helps relate pieces of a passage to one another as it processes them.
GPT is a family, not a synonym for NLP or for every LLM. An LLM is a language model trained at large scale, especially in terms of its parameters and data; there is no single size at which a model suddenly becomes “large”. ChatGPT is the conversational application built around a model. Its first release used a model from the GPT-3.5 family, further trained to respond in dialogue. That extra training matters: being good at continuing text is not quite the same as being helpful to someone asking a question.

How it learns
Training starts with text turned into tokens: pieces that might be whole words, parts of words or punctuation. The exact split depends on the tokenizer. Each piece is represented numerically so the model can process it. A technical name or a line of code may break into different pieces from an ordinary word.
For GPT-style pre-training, the exercise is deceptively simple. Show the model “My phone battery was almost empty, so I plugged it in to ...” and have it predict the next token. The training text provides what actually followed. The training process compares prediction with that token and changes the model's weights to reduce the error. Repeated across many passages, this teaches the model relationships between words, code, explanations and other patterns in its training material. It does not give the model a fact-checker.
That first stage is self-supervised: the text itself supplies the prediction target. The original ChatGPT then underwent further training with examples of conversations and human preferences about responses. That helps explain why it usually answers a question instead of merely carrying on a sentence. The exact training recipe varies by model, and a model is not a dependable, searchable copy of all the documents it saw.
How it infers
When you ask ChatGPT a question, the conversation is split into tokens and passed to a trained model. It generates a reply one token at a time, using each new token as context for the next. That is inference. Your question does not normally change the model's learnt weights.

Some ChatGPT models offer a reasoning mode. Instead of moving straight to the reply, the model generates additional internal tokens to work through intermediate steps and consider alternatives before answering. You do not normally see those steps. It is still generating tokens during inference, not training itself again. The extra work can help with a problem that takes several steps, but thinking longer does not make a wrong assumption true.
Try giving it a short, public technical note. Ask for a summary, then ask a question that requires two details from the note to be combined. Check each claim in the reply against the note: which sentences support it, and what has the model added on its own? A fluent answer, even after more reasoning, is not a substitute for evidence you can check.
Sources and further reading
- Lior Gazit and Meysam Ghaffari, Mastering NLP from Foundations to LLMs (Packt, 2024): https://www.packtpub.com/en-us/product/mastering-nlp-from-foundations-to-llms-9781804619186
- OpenAI, Improving Language Understanding by Generative Pre-Training (2018): https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
- Vaswani and colleagues, Attention Is All You Need (2017): https://arxiv.org/abs/1706.03762
- Stanford NLP Group, Stanford CoreNLP: https://stanfordnlp.github.io/CoreNLP/index.html
- OpenAI, Introducing ChatGPT (2022): https://openai.com/index/chatgpt/