Understanding Artificial Intelligence
"AI" is one of the most overloaded words in technology — used for self-driving cars, photo filters, search engines, and chatbots, often in the same sentence. The diagram below is the simplest accurate picture of how the pieces fit. Click any level to read more.
Every level is a narrower slice of the one above. The reverse is not true — most AI is not Generative AI, and most Machine Learning is not Deep Learning.
Any system that makes a judgment a person would otherwise make — by rule or by data.
Instead of following rules, the system finds patterns from labelled examples.
Many-layered networks that build up understanding from raw pixels, audio, or text.
Produces new content — sentences, images, code — that wasn't in its training data.
Keep reading
See where these systems are actually running in AI in the World, or jump to the trade-offs that matter most in Risks & Responsibility.
Artificial Intelligence (AI)
Artificial Intelligence is a research field, not a product. It started in the 1950s with one question: can a machine do the things that, when humans do them, we call "thinking"? That question has been answered piece by piece ever since.
When someone says "AI did X," it is usually worth asking which kind. The answer often clarifies whether the achievement is impressive, routine, or somewhere in between.
Most AI systems run on a three-step feedback loop: sense the environment through some input (sensors, text, images), process that input through a model (rules, statistics, a neural network), and act on the result (output a decision, send a command, generate a response).
The interesting question is what goes inside the "process" step. Old-style AI used logic programmed by humans — if the patient has these symptoms and not those, suggest this diagnosis. Modern AI replaces the rules with a model trained on data: show it ten thousand labelled cases, and it learns the pattern itself.
Both approaches are still in use, often together. A self-driving car uses learned models to identify pedestrians, but uses hand-written rules to decide what to do at a stop sign. Treating "AI" as one technique misses the fact that real systems are usually a quilt.
- Healthcare: medical imaging, drug discovery, scheduling and triage.
- Finance: fraud detection, credit scoring, algorithmic trading.
- Logistics: route optimisation, demand forecasting, warehouse robotics.
- Communication: spam filters, translation, captioning, voice assistants.
- Search and recommendation: what you see on social media and streaming.
- Public infrastructure: traffic signals, energy grids, weather prediction.
Most of these are not branded as "AI" in the marketing sense. They are quiet AI — running in the background of systems you already use. The flashy generative AI you have heard about is a small fraction of the AI economy by revenue or impact.
The interesting frontier is no longer raw capability. It is interpretability (can we understand why a model made a decision?), reliability (can we trust it on cases unlike anything in its training data?), and alignment (does it actually do what we asked, or what we said we asked?).
These are unsolved problems, and they are the bottleneck on most ambitious deployments — not the technology itself.
Machine Learning (ML)
Machine Learning is the part of AI where the system improves from experience rather than from a programmer rewriting its rules. Instead of telling the computer "emails with the word 'lottery' are spam," you show it ten thousand emails labelled spam or not-spam, and the system finds patterns you would never have thought to write down.
This shift — from rule-writing to example-collecting — is the most important change in software in the last 30 years. It also explains why "data is the new oil" became a cliché: the quality of the data largely determines the quality of the result.
- Supervised learning. Train on labelled examples — emails marked spam, X-rays marked tumour or not. Most practical ML in industry today.
- Unsupervised learning. No labels — the model finds structure on its own. Useful for grouping similar customers or spotting anomalies.
- Reinforcement learning. The model acts, observes the result, and adjusts. Used for game-playing AIs and robotic control.
- Self-supervised learning. The newer category that powers most large models. The system invents its own labels by hiding parts of the data and learning to predict what was hidden — which is how language models learn from raw text without anyone labelling anything.
- Recommendation systems on streaming, shopping, and social platforms.
- Predictive maintenance — flagging a factory machine before it breaks.
- Credit scoring and fraud detection.
- Email filtering and search ranking.
- Speech recognition and voice transcription.
- Demand forecasting in supply chains.
Most of these have been quietly improving for a decade. Modern ML usually does not arrive as a dramatic launch — it arrives as a feature getting slightly better every quarter, in ways nobody announces.
The interesting movement in ML is toward systems that learn continuously from the data they encounter in production, rather than being trained once and frozen. This is harder than it sounds — models that learn from new data can also learn the wrong patterns or be manipulated by bad actors feeding them poisoned input.
The other movement is toward smaller, cheaper, more efficient models that run on a phone or a low-power chip rather than in a giant data centre. The largest models grab the attention; the smallest ones grab the deployments.
Deep Learning (DL)
Deep Learning is the technique that made AI suddenly useful in the 2010s. The "deep" refers to how many layers of mathematical transformation the model puts the data through. Each layer learns slightly more abstract features than the one before it — in an image model, edges first, then shapes, then objects, then scenes.
Worth knowing: the "neurons" in a neural network are loosely inspired by biological neurons but are not actually similar to them. The metaphor is helpful for intuition and deeply misleading for everything else.
- CNNs. Convolutional networks. The architecture behind image recognition — looks for spatial patterns and combines them into higher-level features.
- RNNs / LSTMs. Earlier architectures for sequential data. Largely replaced by transformers for new work, but still common in legacy systems.
- Transformers. The architecture behind nearly every modern language model. Its key idea — attention — is a way for the model to weigh which parts of the input are most relevant when producing each part of the output.
- Diffusion models. The architecture behind most image generators. Trained to gradually transform pure noise into a coherent image, conditioned on a prompt.
- Image and video recognition — from medical imaging to self-driving cars.
- Speech recognition, synthesis, and translation.
- Language understanding — search ranking, autocomplete, content moderation.
- Protein structure prediction (AlphaFold), drug discovery, materials science.
- Robotics perception and control.
Most of these were thinkable but not workable before deep learning. The shift from "barely working in a lab" to "running in production at scale" happened in roughly a decade — unusually fast for any technology.
The next wave is multimodal — single models that handle text, images, audio, and video together rather than as separate problems. The practical effect is that a model can read a chart, listen to a recording, watch a clip, and answer questions across all three sources in one conversation.
The deeper open question is whether scaling up keeps yielding gains. So far, more compute and more data have produced more capable models — but researchers disagree on whether that curve continues, plateaus, or reverses. The honest answer is that nobody knows yet.
Generative AI
Generative AI is the part of deep learning that creates new content rather than classifying or predicting existing content. A spam filter sorts emails into two buckets — that's not generative. ChatGPT writes a new sentence that wasn't in its training data — that is.
These systems do not "know" anything in the way a person knows things. They model statistical patterns in their training data and produce outputs that match those patterns. That sounds limiting — and yet, on enough data, those patterns include nearly everything humans have written down.
For language models, the core idea is surprisingly simple: predict the next word. Given the start of a sentence, what is the most likely word to come next? Train a large transformer on billions of sentences from the internet, books, and code, and you get a system that can continue almost any text plausibly.
Image generators work differently. A diffusion model is trained on billions of image-caption pairs and learns to gradually transform random noise into an image matching the caption. At generation time, it does that reverse process from pure noise all the way to a finished image, conditioned on your prompt.
One important consequence: these systems can be confidently wrong. A model that predicts what looks plausible is not the same as a model that knows what is true. Treating their output as a starting point rather than a final answer is the single most important skill in using them well.
- Conversational AI. Chatbots, virtual assistants — including the chat in the bottom-right corner of this site.
- Content creation. Drafting articles, scripts, emails. Useful as a first draft to edit, less useful as a finished product.
- Code generation. Writing, completing, debugging code. Increasingly the way most professional developers work.
- Design and media. Generating images, animations, voice tracks, short videos from text prompts.
- Scientific work. Generating candidate molecules, protein designs, and material structures faster than humans can hand-design them.
The same capability that enables all of these also enables the misuse cases on the Risks page — phishing, fakes, impersonation. The technology does not care which it is being used for. The choice is human.
The near-future direction is multimodal and agentic. Multimodal means the same model handles text, images, video, and audio in one conversation. Agentic means the model does not just answer your question — it takes a sequence of actions on your behalf, like booking a meeting or finishing a multi-step task across tools.
Both directions raise harder questions than today's chat-only models do. A model that can take actions can take the wrong actions. A model that handles your voice, photos, and documents knows much more about you than a text-only model can. The capabilities will arrive faster than the safeguards.