Skip to content
LLM EngineeringIntermediate

From GPT-2 to GPT-4 and Beyond: How Large Language Models Evolved

Technical walkthrough of LLM evolution from GPT-1 to o3: scaling laws, RLHF training pipeline, chain-of-thought reasoning, and inference-time compute.

TUTAIMarch 20, 202618 min read

The GPT family traces one of the most consequential scaling experiments in machine learning. If you are new to generative AI fundamentals, start there first. This article covers the architectural and training innovations from GPT-1 (2018) through the reasoning models of 2025, focusing on the technical decisions that drove each capability jump.

The Scaling Trajectory: Parameters, Data, Compute

The GPT family of models provides the clearest empirical record of what happens when you systematically increase three variables: parameter count, training data volume, and compute budget. Each generation tested a specific hypothesis about scale, and each produced qualitatively different capabilities.

GPT-1 (2018) introduced the decoder-only Transformer architecture for generative pre-training. At 117 million parameters, it was roughly 10x larger than prior language models. It trained on BookCorpus (approximately 7,000 unpublished books). GPT-1 established the pre-training plus fine-tuning paradigm as standard procedure for NLP. Its limitations were significant: repetitive or nonsensical text generation, inability to track long-term dependencies, and failure to reason across multiple dialogue turns.

GPT-2 (2019) scaled to 1.5 billion parameters, more than 10x GPT-1 and 100x larger than previous state-of-the-art language models. Training data expanded to 40 GB of filtered web text (WebText, curated from Reddit outbound links and filtered from Common Crawl). The embedding dimension grew to 1,600. GPT-2 tested the scale hypothesis: the observation that training with larger data and larger models could develop new capabilities automatically, without explicit supervision. Emergent abilities appeared at this scale, including zero-shot task performance that GPT-1 could not achieve. The model still struggled with coherence over longer passages, failed to capture subtleties such as humor or irony, and was prone to generating misleading content in response to ambiguous or malicious prompts.

GPT-3 (2020) pushed to 175 billion parameters (100x GPT-2) while retaining the same fundamental architecture. Embedding size reached 12,288 dimensions. GPT-3 was the first general-purpose system capable of writing code, performing language translation, solving arithmetic, and designing websites through in-context learning alone. The GPT-3 paper marked a paradigm shift from "fine-tune a model for each task" to "prompt a large model for any task." OpenAI released it through a private beta program and launched the OpenAI API for developer access.

GPT-3.5 / InstructGPT (2022) was a fine-tuned version of GPT-3 that marked the first major application of RLHF in production LLMs. The goal was to make GPT models follow user instructions more accurately and helpfully. InstructGPT (January 2022) led to text-davinci-003 (September 2022, adding RLHF alignment) and then to gpt-3.5-turbo (March 2023, powering ChatGPT). This release brought guardrails into focus: preventing harmful, biased, or unsafe outputs became a central concern alongside capability improvements.

GPT-4 (March 2023) is rumored to use approximately 1.8 trillion parameters across a mixture-of-experts architecture. The key architectural advancement was native multimodality: GPT-4 processes both textual and visual inputs. It demonstrated substantially improved safety and reliability through extended training with human feedback. OpenAI stopped publishing architectural details starting with GPT-4, but the model's performance on standardized exams (passing the bar exam, scoring in the 90th percentile on SAT) confirmed a qualitative jump in reasoning capability. GPT-4 also introduced custom GPTs, allowing users to create specialized versions of the model for specific tasks or domains without coding skills.

The progression follows a clear pattern: each generation increased parameters by roughly two orders of magnitude while expanding training data proportionally.

ModelYearParametersKey Innovation
GPT-12018117MDecoder-only pre-training + fine-tuning
GPT-220191.5BScale hypothesis, zero-shot transfer
GPT-32020175BIn-context learning, few-shot prompting
GPT-3.52022175B (fine-tuned)RLHF alignment, instruction following
GPT-42023~1.8TMultimodality, advanced reasoning

Architecture Evolution: Decoder-Only Transformers at Scale

Every GPT model shares the same foundational architecture: a decoder-only Transformer. Unlike BERT-style encoder models that process bidirectional context, GPT models use masked self-attention where each token can only attend to previous tokens in the sequence. This autoregressive design makes them natural text generators.

GPT-1 replaced RNNs and LSTMs with this decoder-only Transformer, gaining two advantages: improved parallelization during training and better handling of long-range dependencies. The architecture stack consists of multi-head self-attention layers followed by position-wise feed-forward networks, with layer normalization and residual connections.

Subsequent models scaled this architecture without fundamental structural changes. GPT-2 increased the number of layers, attention heads, and embedding dimensions. GPT-3 scaled the same design further. The core transformer block remained consistent across generations. What changed was the number of blocks stacked, the width of each layer, and the training procedure applied after pre-training.

GPT-4o (May 2024) introduced a natively multimodal architecture. Rather than a text model with vision bolted on, GPT-4o processes text, images, audio, and video through a unified model. It achieves lower latency (0.37 seconds vs 0.55 seconds for GPT-4 Turbo) and higher throughput (109 tokens/second vs 20 tokens/second). The model was likely trained with a combination of RLHF and AI feedback (RLAIF) for safety alignment.

GPT-4o mini (2024) applied knowledge distillation to compress the larger model's capabilities into a smaller, faster variant estimated at around 8 billion parameters. In distillation, the full GPT-4o serves as the "teacher" and GPT-4o mini as the "student," trained to replicate the teacher's output distributions and internal patterns. The student model is trained to match both output distributions and behavioral patterns (a mix of behavioral and optional layer-wise distillation). GPT-4o mini also inherited safety features from GPT-4o, including instruction-hierarchy guardrails that prioritize developer instructions over conflicting user inputs. The result preserves much of the original performance while being significantly cheaper to run.

GPT-4.1 and o4-mini (April 2025) continued this trajectory. GPT-4.1 improved performance and lowered cost compared to GPT-4o, while o4-mini provided a smaller, faster version of GPT-4o optimized for cost-efficient reasoning. These models reflect OpenAI's strategy of maintaining both a frontier capability track and a cost-optimized track for production deployments. For practical guidance on working with these models through APIs, including function calling and cost optimization, see the companion article.

The Training Pipeline: Pre-training, Supervised Fine-Tuning, RLHF

Modern LLMs are trained in three distinct stages, each with a different objective. Understanding this pipeline is essential for anyone working with or deploying these models.

Stage 1: Pre-training

Pre-training uses unsupervised language modeling on massive text corpora. The objective is straightforward: predict the next token given all previous tokens. The model processes trillions of tokens from web crawls, books, code repositories, and curated datasets. This stage consumes the vast majority of the compute budget, often millions of GPU-hours.

What pre-training produces is a base model with broad knowledge of language structure, world facts, reasoning patterns, and code syntax. However, a base model is not instruction-following. If you prompt it with a question, it may complete the text plausibly rather than answering the question. It has no concept of helpfulness, safety, or conversational norms.

# Conceptual pre-training objective (next-token prediction)
# For a sequence [x1, x2, ..., xn], maximize:
# P(x_t | x_1, x_2, ..., x_{t-1}) for all t

import torch
import torch.nn.functional as F

def compute_pretraining_loss(model, input_ids):
    # Shift inputs to create targets
    logits = model(input_ids[:, :-1])
    targets = input_ids[:, 1:]

    # Cross-entropy loss over vocabulary
    loss = F.cross_entropy(
        logits.reshape(-1, logits.size(-1)),
        targets.reshape(-1)
    )
    return loss

Stage 2: Supervised Fine-Tuning (SFT)

Supervised fine-tuning transforms the base model into an instruction-following assistant. Human annotators write thousands of prompt-response pairs demonstrating desired behavior: answering questions accurately, refusing harmful requests, formatting responses clearly, and following complex multi-step instructions.

The model is then fine-tuned on this curated dataset using standard supervised learning. The loss function is identical to pre-training (next-token prediction), but the training data now consists of high-quality demonstrations rather than raw web text.

SFT is what turns a text completion engine into something that behaves like a useful assistant. InstructGPT (2022), the precursor to GPT-3.5 and ChatGPT, was the first major model to apply this approach at scale, and it represented the return of fine-tuning after GPT-3 had emphasized in-context learning.

Stage 3: Reinforcement Learning from Human Feedback (RLHF)

RLHF aligns the model's outputs with human preferences on dimensions that are difficult to capture through supervised examples alone: tone, helpfulness, nuance, and safety.

The RLHF pipeline has two sub-stages:

Reward model training. Human annotators rank multiple model outputs for the same prompt from best to worst. These preference rankings train a separate reward model that learns to score any given output for quality. Given a prompt and a response, the reward model outputs a scalar score indicating how well the response matches human preferences.

Policy optimization. The LLM is then optimized using Proximal Policy Optimization (PPO) to maximize the reward model's score while staying close to the SFT model's distribution (to avoid reward hacking). A KL-divergence penalty prevents the model from finding degenerate outputs that exploit the reward model without being genuinely helpful.

# Simplified RLHF objective
# maximize: E[R(prompt, response)] - beta * KL(policy || sft_model)

# Where:
# R = learned reward model score
# beta = KL penalty coefficient
# policy = current model being optimized
# sft_model = reference model from Stage 2

The three-stage pipeline (pre-training, SFT, RLHF) is now the standard procedure for training production LLMs. Each stage addresses a different aspect: knowledge acquisition, instruction following, and preference alignment.

Emergent Capabilities: Chain-of-Thought and Few-Shot Learning

As models scaled, capabilities appeared that were absent in smaller models and were not explicitly trained. These emergent abilities fundamentally changed how practitioners interact with LLMs.

In-Context Learning and Few-Shot Prompting

GPT-3 demonstrated that a sufficiently large model can learn new tasks at inference time by conditioning on examples provided in the prompt. This is in-context learning (ICL): the model discovers patterns in the input-output pairs within its context window and applies those patterns to new inputs. Critically, no gradient updates occur. The model weights are frozen. The "learning" happens entirely through attention over the prompt.

Few-shot prompting provides multiple examples before the target query:

Classify the sentiment:

Text: "This movie is great, I had a great time watching it."
Sentiment: Positive

Text: "I've never seen a worse movie, it was a waste of time."
Sentiment: Negative

Text: "The food is bad and the service should be improved"
Sentiment:

The model generalizes the classification pattern from the examples and applies it to the new input. Zero-shot prompting (no examples) and one-shot prompting (a single example) work similarly but with decreasing reliability.

Chain-of-Thought (CoT) Prompting

Chain-of-thought prompting, introduced by Wei et al. (2022), instructs the model to produce intermediate reasoning steps before arriving at a final answer. This technique dramatically improves performance on arithmetic, commonsense reasoning, and multi-step logic problems.

The simplest form is zero-shot CoT, triggered by appending "Let's think step by step" to the prompt. Few-shot CoT provides worked examples that include the reasoning trace:

Q: A juggler can juggle 16 balls. Half of the balls are golf balls,
   and half of the golf balls are blue. How many blue golf balls are there?
A: There are 16 balls in total. Half of the balls are golf balls.
   So there are 16 / 2 = 8 golf balls. Half of the golf balls are blue.
   So there are 8 / 2 = 4 blue golf balls. The answer is 4.

Without CoT, the same model often outputs incorrect answers because it attempts to jump directly to the solution without decomposing the problem.

Self-Consistency

Self-consistency (Wang et al., 2022) extends CoT by sampling multiple reasoning paths and selecting the most frequent final answer through majority voting. Rather than relying on a single greedy decode, the model generates several independent chain-of-thought traces. If three out of five traces produce the same answer, that answer is selected. This approach reduces variance and improves accuracy on problems where a single reasoning chain might take a wrong turn.

Self-Refine

Self-refine creates an iterative improvement loop: a generator LLM produces an initial output, a critic LLM reviews the output and provides feedback, and then the generator revises its response based on that feedback. The feedback can be extrinsic (from an external verifier or tool) or intrinsic (from the LLM evaluating its own generation). This pattern is foundational to modern agent architectures where models iterate toward better solutions rather than producing a single attempt.

Self-Ask

Self-ask prompting teaches the model to decompose complex questions into sub-questions, answer each sub-question independently, and then combine the intermediate answers. For example, given the question "Who was president of the U.S. when superconductivity was discovered?", a standard CoT approach might hallucinate the answer. Self-ask decomposes it: (1) When was superconductivity discovered? (2) Who was president at that time? Each sub-question is easier to answer correctly, and the composed answer is more reliable.

Reasoning Models: o1, o3, and the Shift Toward Deliberative AI

GPT-4 demonstrated strong general capabilities but still failed at certain hard reasoning tasks, particularly in mathematics and competitive programming. OpenAI's o1 (September 2024) and o3 (April 2025) represent a fundamentally different approach to this problem.

What Makes Reasoning Models Different

Standard LLMs allocate roughly the same amount of computation per token regardless of problem difficulty. A question about the capital of France receives the same processing depth as a graduate-level physics derivation. Reasoning models break this pattern by spending variable amounts of computation at inference time based on problem complexity.

The o1 model uses long chain-of-thought reasoning: before producing an answer, the model generates an extended internal reasoning trace. This trace includes problem decomposition, hypothesis generation, self-critique of partial solutions, and exploration of alternative approaches. OpenAI reported that o1's performance improves consistently with both more reinforcement learning during training (train-time compute) and more time spent thinking during inference (test-time compute).

Benchmark Performance

The o1 model achieved substantial improvements over GPT-4o on reasoning-heavy benchmarks:

  • MATH benchmark: 94.8% (o1) vs 60.3% (GPT-4o)
  • MMLU: 92.3% (o1) vs 88.0% (GPT-4o)
  • GPQA Diamond (PhD-level science): Chemistry 64.7% (o1) vs 40.2% (GPT-4o), Physics 92.8% (o1) vs 50.5% (GPT-4o)

On standardized exams, o1 scored 83.3% on AP Calculus (vs 71.3% for GPT-4o) and 95.6% on the LSAT (vs 69.6% for GPT-4o).

o3: Scaling Test-Time Compute

The o3 model (2025) is the refinement of o1, incorporating all learnings from the experimental o1 deployment. On the ARC-AGI semi-private evaluation, o3 (high-tuned configuration) achieved 88% accuracy, compared to roughly 32% for o1 at the high-compute setting. This came at significant per-task cost (in the hundreds of dollars range), illustrating the compute-accuracy tradeoff inherent in test-time scaling.

The key shift is architectural: while traditional LLMs scale primarily through larger models and more training data, reasoning models scale through inference-time computation. The model can allocate more "thinking" to harder problems, producing a cost-performance curve rather than a single fixed capability level.

Do Reasoning Models Truly Reason?

Research from Shojaee et al. (2025) investigated this question using controllable puzzle environments (such as Tower of Hanoi variants) that allow precise manipulation of compositional complexity. Their findings identify three performance regimes:

  1. Low-complexity tasks: Standard models surprisingly outperform reasoning models, which waste tokens on unnecessary deliberation.
  2. Medium-complexity tasks: Reasoning models demonstrate clear advantage, as additional thinking time helps decompose the problem.
  3. High-complexity tasks: Both model types experience complete collapse, suggesting current architectures hit a ceiling regardless of compute allocation.

For correctly solved problems, the reasoning model tends to find answers early at low complexity but at higher complexity often fixates on an incorrect early answer and wastes the remaining token budget. This suggests that current reasoning models use pattern-matching and heuristic search rather than true logical deduction.

LLM Scaling Laws and What They Tell Us About the Future

The Kaplan et al. (2020) scaling laws paper established that language model performance (measured by cross-entropy loss) improves as a smooth power law with three factors: model size (N), dataset size (D), and compute budget (C). The relationship follows:

L(N) ~ (N_c / N)^alpha_N
L(D) ~ (D_c / D)^alpha_D
L(C) ~ (C_c / C)^alpha_C

where L is loss and the alpha exponents are empirically determined constants. The critical finding: for optimal performance, all three factors must be scaled in tandem. Scaling model size alone while holding data constant hits diminishing returns. Scaling data without increasing model capacity wastes compute.

Chinchilla (Hoffmann et al., 2022) refined these laws and showed that many large models were significantly undertrained relative to their parameter count. The Chinchilla-optimal ratio suggests approximately 20 tokens of training data per parameter. By this metric, a 70B parameter model should train on roughly 1.4 trillion tokens.

Two Scaling Frontiers

The field now operates on two parallel scaling axes:

Train-time scaling continues to increase model size and training data. GPT-4 trained on an estimated 13 trillion tokens. Each generation requires roughly 3-10x more compute than the previous one.

Test-time scaling (inference compute) is the newer frontier. Reasoning models like o1 and o3 demonstrate that spending more computation per query can substitute for some training-time investment. This creates a spectrum: for a fixed total compute budget, the optimal split between training and inference depends on the task distribution.

The practical implication for engineers is that the cost-performance frontier is no longer one-dimensional. A smaller model with a reasoning loop may outperform a larger model on structured tasks, while the larger model wins on broad knowledge retrieval. Model selection now requires understanding both the capability profile and the inference cost structure.

The Distillation Path

Knowledge distillation offers a third scaling strategy. Rather than deploying the largest model, organizations can train a smaller student model to replicate the teacher's behavior. GPT-4o mini demonstrates this: at an estimated 8 billion parameters, it inherits much of GPT-4o's capability while being suitable for latency-sensitive and cost-constrained deployments. The student model is trained on both ground truth labels and the soft probability distributions from the teacher, capturing richer information than hard labels alone.

These three strategies (bigger pre-training, inference-time computation, and distillation) are complementary. The most effective production deployments combine all three: a large model for difficult queries, a distilled model for routine requests, and reasoning loops for tasks requiring multi-step logic.

Safety, Guardrails, and Data Considerations

As LLMs became more capable, safety and responsible deployment became a parallel engineering challenge spanning guardrails, data quality, privacy, and accountability.

Guardrails and Instruction Hierarchy

With ChatGPT's release, preventing harmful, biased, or inappropriate outputs became critical. Guardrails address three main concerns: safety and ethical use (blocking harmful content generation), factual reliability (reducing hallucinations), and user trust (compliance with legal and organizational policies in sensitive domains like healthcare, education, or finance).

GPT-4o mini introduced an instruction-hierarchy safety strategy. The model assigns different privilege levels to message types: system messages receive the highest priority, user messages medium priority, and tool outputs the lowest. If a user input conflicts with the developer's original system instructions, the model defaults to the developer's guidance. This layered approach helps defend against prompt injection attacks embedded in external content.

Training Data Limitations

LLMs depend on massive text corpora, but not all training data is reliable, relevant, or representative. Some data may be outdated, inaccurate, or biased. Certain languages and domains have significantly less data available than others. For example, English dominates NLP solutions at roughly 67% of available data, while low-resource languages represent only about 6%, despite speakers of those languages making up 68% of the global population. This imbalance leads to lower performance for underrepresented languages and domains.

Privacy and Accountability

LLMs can process and generate sensitive personal information including health records, financial data, and identity details. They can also inadvertently leak information from their training data. Regulatory frameworks such as the EU AI Act and GDPR impose requirements on AI systems that handle personal data, including consent for processing, data minimization, and security guarantees. Organizations deploying LLMs must evaluate and monitor these systems regularly, as they can produce wrong answers, misleading advice, or inaccurate predictions with real-world consequences.

Summary

The evolution from GPT-1 to o3 traces a path from simple next-token prediction at 117 million parameters to deliberative reasoning systems with trillions of parameters and variable inference compute. The key transitions were:

  1. GPT-1 to GPT-2: Validation of the scale hypothesis. Larger models develop capabilities without explicit supervision.
  2. GPT-2 to GPT-3: In-context learning emerges. The paradigm shifts from fine-tuning to prompting.
  3. GPT-3 to GPT-3.5/ChatGPT: The three-stage training pipeline (pre-training, SFT, RLHF) produces models that follow instructions and align with human preferences.
  4. GPT-3.5 to GPT-4: Multimodality and substantially stronger reasoning. OpenAI closes the research publication loop.
  5. GPT-4 to o1/o3: Test-time compute scaling. Models that allocate variable computation based on problem difficulty.

Each transition expanded both the technical capabilities and the design space for practitioners building on these systems. The training pipeline, scaling laws, and emergent prompting techniques covered here form the foundation for engineering work with any modern LLM, not just the GPT family.

Share

Ready to go beyond theory?

Explore TUTAI's hands-on AI courses for practitioners and build real-world AI projects with expert guidance.

Explore the courses

Related Articles