This isn't a course that hands you everything. It's a map. The teaching is done by the best free resources on the internet; I've put them in order and explained the ideas in plain language first.
No jargon for the sake of it. And one thing matters more than any video: you have to do it yourself.
Due to my busy work schedule, a detailed intuitive explanation with images will roll out slowly over time. In the meantime, try to get a headstart here — a beginner-friendly e-book is coming :3
The real thing a degree gives you isn't the facts. It's the ability to teach yourself anything, for the rest of your life. This field never stops moving, so learning how to learn on your own is the one skill that never goes out of date.
This is free because I learned everything for free. Every resource on this page is one I actually used. No paywall, no catch.
The order is deliberate. Maths first, with code running alongside it. Then data science and machine learning, where those ideas turn into working predictions. Then deep learning, the same ideas stacked higher. Then LLMs, last on purpose — skipping ahead feels faster but it's how people get stuck copying code they can't debug.
You don't need equal amounts of each. The maths that earns its keep is roughly half statistics and probability, a third linear algebra, and a smaller slice of calculus. Two things also live somewhere people don't expect: reinforcement learning sits at the end of deep learning rather than with the other ML methods, because anything you'd build with it uses a neural network; and deployment sits inside data science, because a model nobody can use isn't finished.
The language underneath everything. Start with statistics and basic algebra, build up gradually.
Python, properly, once. Run this alongside the maths, not after it.
Where the maths meets real data: clean it, explore it, let a model find the patterns, then ship the result.
The same idea, stacked into layers: neurons, then vision, sequences, transformers, and reinforcement learning last.
Prompting, the API, context and cost, then agents that use tools on your behalf.
Not a phase you reach, a habit you start on day one. Small finished projects teach more than the next course.
An autodidact is someone who teaches themselves. It's a habit, not a talent: read or watch a little, then immediately try it with your own hands before moving on.
Watching someone solve a problem feels like learning, but understanding doesn't stick until you get stuck and work your way out. Every resource below is something to use actively — pause the video and code along, redo the maths on paper, break the example and fix it. If a section feels too hard, drop back one level rather than pushing through confused.
Don't do these one at a time, front to back. Run them in parallel and build up together. Khan Academy is the spine here because it starts gently and never assumes you remember the last thing.
Read: Why you need to understand the maths behind ML (intuitive understanding)
Run both at once. Algebra 2 is the toolkit; stats teaches you to reason about data and chance early.
The bridge — functions, graphs and the ideas calculus is about to lean on.
Calc 1 for how things change; College Algebra to firm up the foundations; Linear Algebra because data lives in matrices.
Deepens the calculus you'll meet again inside how models learn.
The most useful maths for this whole field, and the one beginners skip. It's how you describe data, measure uncertainty, and tell a real pattern from random noise. Begin it alongside Algebra 2.
Statistics is just making honest statements about data when you can't see everything. You have a sample, not the whole world, so you learn how confident you're allowed to be.
Bring this in at step 03, next to Calculus 1. It's the maths of vectors and matrices, and since data is stored as grids of numbers, this is how data physically moves through every model.
A spreadsheet is a matrix. Linear algebra is the rules for doing arithmetic on whole tables of numbers at once instead of one cell at a time.
The smallest slice of the three, so don't let it eat your schedule. Climb Pre-calc → Calc 1 → Calc 2. You mainly need the core idea behind how a model improves itself, not every technique.
Calculus is the maths of change and slopes. Later, "the model is learning" really means "follow the slope downhill to make the error smaller." That's it.
Run this next to the maths, not after it. Maths without code is theory you can't test; code without maths is copying things you don't understand.
You are not training to be a software engineer. Write a script from scratch, read someone else's and follow it, work out why yours is broken without panicking. Python is the only language you need, learn it properly here, once.
Read: Guide to learning how to code in the sea of vibecoders
Every program ever written is made of about six ideas, and you can learn all of them in a fortnight. Pick one resource below and finish it end to end rather than sampling all of them. Typing the examples out yourself is not optional.
A program is a list of instructions the computer follows in order. Variables are labelled boxes for values, conditionals choose between paths, loops repeat work, functions are recipes you write once and reuse.
Variables & types — named boxes holding numbers, text, or true/false; most beginner bugs are a value being a different type than you assumed. Control flow — if this, do that; conditionals pick a branch, loops repeat until a condition ends them. Functions — a block of work with a name, inputs and an output; small functions are easier to fix than one long script. Data structures — lists, dictionaries, sets; choosing the right container makes half your problems disappear.
Code being broken is the normal state of code. The gap between a beginner and a professional is mostly how calmly and quickly they find the mistake.
An error message is not noise, it's the computer telling you where it got confused. Read it bottom-up, then print the values around it until reality stops matching what you assumed.
Read the error slowly — most errors say exactly what's wrong. Print everything — narrow down where your belief and the truth split apart. Shrink the problem — a ten-line failing example is easy to fix; a 300-line one isn't. Rubber ducking — explain the code line by line to an object or a chat window; you'll usually catch it mid-sentence.
The stuff no course teaches and every job assumes. The terminal is how you run things; Git is how you save your work in a way that lets you undo mistakes and show your history.
Git is a save system for a whole project. You take snapshots as you go, so you can always get back to a version that worked. GitHub is where those snapshots live online, and it doubles as your CV in this field.
An LLM will happily hand you working code for anything in this section. That's exactly the problem — the understanding you're building comes from getting stuck and unstuck, and accepting a finished answer skips the part that does the teaching.
There's a difference between a tutor and a ghostwriter. Asking it to explain an error, or review code you already wrote, is a tutor. Asking it to write the function is a ghostwriter, and you'll feel the hole later. Once you can genuinely follow every line it produces, use it freely — phase 05 is all about that.
This is where most of the real work happens. Surprisingly little of the job is inventing clever algorithms — most of it is getting data, cleaning it, looking at it properly, applying tools that already exist, and getting the result somewhere people can use.
You know Python from phase 02. This is the data dialect: NumPy for fast arrays, Pandas for tables, SQL for getting data out of a database. Beginners skip SQL and regret it.
Pandas is a spreadsheet you drive with code. SQL is how you ask a database "give me just these rows." NumPy is the fast maths underneath both.
The unglamorous skill that quietly decides everything. Real data is messy: missing values, typos, wrong formats. A model fed bad data gives bad answers, no matter how fancy it is.
Garbage in, garbage out. Cleaning is the boring-but-essential work of making messy data trustworthy before you ask any questions of it.
EDA is the habit of properly looking at a dataset before you trust it with anything. You count things, plot things, hunt for whatever is weird. It's the cheapest step in the whole process and catches mistakes that would otherwise quietly ruin a model three weeks later. Expect to loop with cleaning several times before you train anything.
EDA is getting to know your data before you trust it. Plot everything, count everything, stay suspicious. A five-minute histogram has saved more projects than any clever algorithm.
Shape first — rows, columns, types, missingness; always look at raw rows too. Distributions — skew, double humps, impossible values. Relationships — correlation is not causation. Leakage — the killer bug: a column that secretly contains the answer makes your model look perfect in testing and useless in reality.
The secret that makes it all click, worth reading twice: a machine learning model is just a maths function. Numbers go in, a prediction comes out. "Training" is searching for the version of that function that fits your data best.
Machine learning is finding the best simple formula that connects your data. You have inputs and known answers; the machine tweaks a formula until its outputs match the answers as closely as possible. Then you feed it new inputs and trust its prediction.
Weights — how important each input is. Bias — a baseline that shifts the whole result up or down. Most of what you'll meet is supervised learning, where every example comes with the right answer attached, plus some unsupervised methods that find structure with no answers at all. Reinforcement learning, the third branch, waits until phase 04.
Models affect real people. They can inherit bias from their data and make unfair decisions at scale. Knowing how to spot and reduce that harm is part of doing this work responsibly, not an optional extra.
A model trained on biased data will repeat that bias confidently. Ethics is learning to ask "who could this harm, and how would I know?" before you ship anything.
A model sitting in a notebook on your laptop isn't finished, it's a sketch. Deployment turns it into something other people can use: packaging it, putting it behind an address your app can call, and watching it afterwards so you notice when it starts going wrong.
MLOps is DevOps for models. This is everything around the model file: how it gets built the same way twice, how it gets served, and how you find out when the world moved on without it.
Serving — wrap the model in a small API; FastAPI plus Docker is the standard first stack. Reproducibility — same data, code and settings, versioned. Monitoring & drift — the world keeps changing, your model doesn't; watch the inputs and predictions, not just whether the server is up. Pipelines — collect → clean → train → evaluate → deploy, on a schedule, without you.
The piece of deployment infrastructure that arrived with modern AI. An embedding is a list of numbers a model produces to stand for a thing, arranged so similar things end up close together. A vector database stores millions of them and answers "what's nearest to this?" in milliseconds — powering semantic search, recommendations, and the retrieval half of RAG.
An embedding is meaning turned into coordinates. A vector database is the map you look things up on: instead of matching words, you find the nearest points.
Similarity search — cosine similarity is the usual ruler. Chunking — how you split a document into passages is the biggest quality lever in a retrieval system. ANN indexes — approximate nearest-neighbour search trades a sliver of accuracy for enormous speed.
You already know a machine learning model is just a maths function. Deep learning is what happens when you stack many of those functions on top of each other, so the output of one becomes the input of the next, layer after layer. That lets the model learn patterns far too complicated for a single formula. Everything famous, from vision models to language models, is this same idea arranged in different shapes.
Learn these before any specific model. They're the building blocks every deep learning system shares.
Neuron — one tiny maths function: takes inputs, multiplies each by a weight, adds a bias, passes on a single number. Activation function — lets the network bend and curve instead of only drawing straight lines. Layers — neurons lined up side by side, then stacked; functions feeding functions. Backpropagation — how the network learns from mistakes: traces the error backwards and nudges every weight a little in the direction that reduces it, millions of times.
The models that see. Built for images, where what matters is the pattern in a patch of nearby pixels: an edge, a corner, eventually a face.
A CNN scans an image in small tiles looking for little patterns, then combines them into bigger ones. Early layers spot edges; later layers spot whole objects.
Models for data where order matters: text, speech, time series. They read one step at a time and carry a little memory of what came before.
These models remember the previous steps as they read, the way you hold the start of a sentence in mind to understand its end.
The architecture behind modern language models and most of today's AI. Instead of reading step by step, a transformer looks at all the words at once and learns which ones to pay attention to.
A transformer reads everything at once and asks, for each word, "which other words should I pay attention to here?" That's what makes it so powerful, and it's still just stacked maths functions.
Everything so far learned from an answer key. Reinforcement learning throws the key away: an agent acts in an environment, receives a reward or penalty, and gradually works out which behaviours pay off. It's here rather than in phase 03 because anything you'd build today uses a neural network as the agent's brain — and it sets up the next section, since chat models are tuned with RL from human feedback.
Supervised learning is learning from an answer key. Reinforcement learning is learning from consequences: try something, see whether it went well, do more of what worked.
Agent & environment — the decision-maker and the world it acts on. Reward — one number saying "good" or "bad"; design it badly and the agent optimises exactly what you measured, not what you meant. Policy — the strategy; in deep RL, the policy is a neural network. Explore vs exploit — take the reward you know, or gamble on a better one.
Large language models like ChatGPT and Claude are the tools you'll reach for daily, both to learn faster and to build things. Used well, an LLM is a patient tutor who never gets tired of your questions. Used badly, it quietly teaches you wrong things. Most apps with a built-in chatbot are just LLM wrappers with a prompt telling it to act as the platform's help bot.
The skill isn't just typing a question. It's knowing how to ask, how to check the answer, and how to plug a model into your own code. Learn it last, on purpose, because it's most useful once you understand enough to spot when the model is wrong.
The fastest way to get unstuck. Paste an error, ask it to explain a concept five ways, or have it quiz you. But it can sound confident and still be wrong, so treat every answer as a smart friend's guess, not gospel.
An LLM predicts the next likely words, it doesn't look things up or truly "know." That's why it can invent facts. Use it to understand and explore, then verify anything that matters against a real source.
A fancy name for asking clearly. The model can only work with what you give it, so the more specific you are about the goal, format, and examples, the better the answer.
A vague question gets a vague answer. Tell it who to be, what you want, and show one example. "Fix my code" is weak; "You're a Python tutor, here's my code and the error, explain the bug then show the fix" is strong.
How you go from typing in a browser to building your own apps. An API is your code sending a message to the model and getting a reply back — same conversation, done in Python instead of a chat box.
An API call is one message in, one message out. Your program keeps the conversation; the model just answers the latest version of it, from scratch, every time — the API is stateless.
Messages & roles — a system instruction, then alternating user/assistant turns. Tokens — roughly three-quarters of a word each; you're billed per token in and out. Temperature — how much randomness in the wording; not an intelligence dial. Streaming — the reply piece by piece, far better to sit in front of. Ask for structured output (JSON) when your code needs to use the answer directly, and wrap calls in retries for rate limits and timeouts. Keep API keys in an environment variable, never in the file.
The context window is the model's desk — everything it can see at once has to fit on it: system prompt, conversation so far, any documents, plus room for the reply. Fill it with junk and answers get worse, not better, so deciding what earns a place is the real skill (context engineering).
Caching: if the start of your prompt is identical on every call, you're paying to re-read it every time. Prompt caching keeps that processed opening warm, so repeat calls are cheaper and faster. The rule: stable content first, changing content last — caching matches on the prefix, and one edited character near the top throws away everything after it.
The context window is how much the model can hold in its head at once, and you pay for all of it, every turn. Caching is paying properly once for the part that never changes, then almost nothing to reuse it.
Batching — async jobs at a steep discount when latency doesn't matter: bulk classification, backfills, evals. Cost control — measure tokens per request, then route: a small cheap model for easy steps, the expensive one only where it earns its place.
A chatbot answers. An agent does. Give the model a goal and a set of tools it can call — search the web, read a file, run code, query a vector database — and let it run. It asks for a tool, your code executes it, hands back the result, it decides what's next. Repeat until the job's done.
The model never runs anything itself; it only ever asks. Everything around it — tool definitions, execution code, memory, permission checks, the loop itself — is the harness. That's the part you build. Loop engineering is the design work inside that cycle: what goes back into context, when to stop, what happens when a tool errors, how much you're willing to spend.
An agent is a model in a loop with tools. It decides, your code acts, the result goes back in, round again. The cleverness lives in the loop, not the model.
Tools — describe them as carefully as your prompt, since that's all the model has to go on. Evals — keep a set of known-good tasks and re-run them every time you change the prompt or the loop, or you're guessing. Most problems don't need an agent — one good prompt, or a fixed chain of two or three calls, is cheaper and more predictable. When you do build one, give it the fewest tools that can do the job, plus a hard cap on steps and spend.
This is, and always will be, free. There's no paywall coming, and you never have to give a thing to use any of it.
But if it helped you, and only if you can comfortably spare it, a small tip helps me keep making free things and keeps learning accessible for the next person. Completely optional. No pressure, no hard feelings.
Tip on Ko-fiYou don't need to finish every resource before you start. Learn a little, build a small project, get stuck, look it up, keep going. That loop is what turns you into someone who can do this.