HOW AI ACTUALLY WORKS ← LEARNSTART

How AI
actually works

From pure math to chips to agents that do your work — the whole magic trick in one coffee break, explained like you're five. No equations, promise.

← → arrow keys · hover the dotted words for pro-level detail
The big secret

AI doesn't know.
It guesses

Underneath every clever answer is one simple game: guess the next token. The model plays it absurdly well — and those percentages come from a step called softmax.

FIG.01 — The only game it plays THE CAT SAT ON THE ___ MAT 94% SOFA 5% MOON 1% ← picks this
Step 1 — The math

Words become numbers

The model can't read. Every word becomes an embedding — a list of numbers — pushed through a neural network: billions of dials called weights, each followed by a tiny gate called ReLU.

FIG.02 — A word, entering the machine CAT 1.8 -0.3 2.2 0.7 -1.1 0.4 … ITS "MEANING", AS NUMBERS THE FORMULA × BILLIONS OF DIALS
Step 2 — Inside a neuron

One neuron, one gate

Rosenblatt's 1958 blueprint still holds: multiply inputs by dials, add them up, squeeze through a gate — the activation function. The gates got smoother over the decades; stack millions of gated neurons and you're doing deep learning.

FIG.03 — The perceptron, and its gates INPUTS ×W Σ ADD UP GATE OUT STEP · 1958 HARD YES / NO SIGMOID · 1986 SMOOTH MAYBE RELU · 2011 IF POSITIVE GELU · NOW SMOOTH RELU THE GATE'S SHAPE, THROUGH THE DECADES →
Step 3 — Training

Learning = turning dials

Show it a sentence, hide the ending, let it guess. Wrong? Nudge every dial slightly downhill — that's gradient descent — until the wrongness score, the loss, stops shrinking. Repeat trillions of times.

FIG.04 — The training loop WATCH IT LEARN — SAME BLANK, THREE ROUNDS → THE CAT SAT ON THE DOG BED MAT ✕ WRONG — NUDGE THE DIALS ~ CLOSER — NUDGE AGAIN ✓ RIGHT — DIALS LOCKED IN ONE OF 1,000,000,000,000 DIALS LOSS LOWER = SMARTER ATTEMPTS SO FAR: 0
Step 4 — The chips

One genius vs ten thousand helpers

A CPU is one brilliant chef. A GPU is a stadium of kids each doing one tiny sum, all at once. Dial-math is really one enormous matrix multiply — exactly what GPUs were built for.

FIG.05 — CPU vs GPU CPU — 1 BIG BRAIN GREAT AT HARD, ONE-AT-A-TIME JOBS GPU — 10,000 TINY CALCULATORS ALL WORKING AT THE SAME TIME
Step 5 — Transformers

Every word watches every other word

The 2017 trick that changed everything is called attention: to understand "IT", the model glances back at every word and finds "THE CAT". Stack that trick dozens of layers deep and you've built a transformer — the T in Chat-GP-T.

FIG.06 — Attention, watching THE CAT SAT BECAUSE IT WAS TIRED "IT" = THE CAT · 91% SURE EVERY WORD DOES THIS TO EVERY OTHER WORD, IN PARALLEL
Step 6 — The family tree

Many brains came before

AI tried many brain shapes first: the one-neuron perceptron, stacked ANNs, loops that remember (RNNs), grids that see (CNNs), duelling twins that paint (GANs). Then 2017 happened.

FIG.07 — 60 years of brain shapes 1958 PERCEPTRON ONE FAKE NEURON 1986 ANN NEURONS, STACKED 1997 RNN · LSTM LOOPS THAT REMEMBER 2012 CNN GRIDS THAT SEE 2014 GAN FORGER VS DETECTIVE 2017 TRANSFORMER ATE THEM ALL EACH SHAPE SOLVED ONE THING · THE TRANSFORMER SOLVED LANGUAGE — AND SCALED
Step 7 — Why the transformer won

It killed the bottleneck

Older brains — autoencoders & seq2seq — squeezed a whole sentence through one tiny summary, like retelling a movie from a single sticky note. The transformer let every word talk to every word, all at once. No bottleneck, fully parallel — so it could scale to the whole internet.

FIG.08 — Before and after 2017 BEFORE — SEQ2SEQ / AUTOENCODER ONCE UPON A TIME… ONE TINY SUMMARY — THE BOTTLENECK LONG STORY IN → START FORGOTTEN · ONE WORD AT A TIME AFTER — TRANSFORMER · 2017 ONCE UPON A TIME EVERY WORD ↔ EVERY WORD · ALL AT ONCE · NOTHING FORGOTTEN
Step 8 — New models

Bigger brains that now think first

From the single-neuron perceptron of 1958 to today's deep neural networks: more dials + more examples + more chips = smarter. That's scaling. The newest models also think before answeringreasoning.

FIG.09 — Scaling, then reasoning BIGGER… TRILLIONS 1958 2020 NOW 1 DIAL 175 BILLION DIALS SIZE ≈ SMARTS — THE SCALING LAW …AND THEY THINK FIRST EXPLAIN MY ELECTRIC BILL → HERE'S THE BREAKDOWN ✓ DRAFT IN PRIVATE → ANSWER IN PUBLIC THAT PAUSE IS THE REASONING
Step 9 — Your request, then vs now

Same question, more machinery

Ask GPT-3.5 (2022) and your words ran straight through — one pass, one answer. GPT-4 (2023) added vision, longer memory, safety checks. Ask GPT-5.6 today and a router picks a brain, it thinks in private, calls tools and checks itself — all before you see a word.

FIG.10 — The pipeline, by generation GPT-3.5 · 2022 YOU ASK MODEL ANSWER ONE PASS · FAST · OFTEN WRONG GPT-4 · 2023 YOU ASK BIGGER MODEL SAFETY CHECK ANSWER + VISION · LONG MEMORY · GUARDRAILS GPT-5.6 · 2026 YOU ASK ROUTER THINKS IN PRIVATE… CALLS TOOLS CHECKS · ANSWERS A WHOLE PIPELINE PER QUESTION SAME GAME UNDERNEATH — GUESS THE NEXT TOKEN — WITH EVER MORE MACHINERY AROUND IT
Step 10 — Fine-tuning

School, then job training

Pre-training = reading half the internet. School. Fine-tuning = a short course on top: manners, honesty, or a specialty like medicine or code. Job training.

FIG.11 — Base model, then the polish BASE MODEL KNOWS EVERYTHING, HELPS WITH NOTHING TRAINED ON THE INTERNET + A FEW GOOD EXAMPLES FINE-TUNE: MANNERS · SAFETY · SPECIALTY ASSISTANT SAME BRAIN, NOW ACTUALLY USEFUL
Step 11 — Agents

Give the brain hands

A chatbot can only talk. An agent can act: think → use a tool → look at what happened → try again. Around the loop until the job is done.

FIG.12 — The agent loop 1 · THINK 2 · USE A TOOL SEARCH · CODE · FILES · EMAIL 3 · CHECK RESULT 4 · TRY AGAIN OR FINISH ✓ ROUND AND ROUND, UNTIL DONE
Step 12 — Harnesses

The cockpit around the brain

Raw models never ship alone. A harness wraps them — instructions, tools, memory, guardrails. It's the difference between an engine and a car you can drive.

FIG.13 — Model + harness = product WATCH THE COCKPIT ASSEMBLE → MODEL THE BRAIN RULES — ITS ORDERS TOOLS — ITS HANDS MEMORY — ITS NOTES GUARDRAILS — LIMITS = A PRODUCT YOU CAN TRUST CHAT · COPILOT · AGENT
The whole story

Sand to systems

That's the stack: math on chips, transformers trained into models, polished by fine-tuning, given hands as agents, wrapped in harnesses. Now you know how AI works.

HOW AI WORKS THE WHOLE STACK · IN ONE PICTURE 1 · MATH — GUESS THE NEXT WORD 2 · CHIPS — 10,000 TINY CALCULATORS 3 · TRANSFORMER — WORDS WATCH WORDS 4 · MODEL — BIGGER + THINKS FIRST 5 · FINE-TUNE — SCHOOL → JOB 6 · AGENT — GIVE IT HANDS 7 · HARNESS — THE COCKPIT BENJAMINSANGWA.COM/LEARN
More at /learn
The poster keeps animating — open it in any browser