10 lessons · 10–16 h (Parts 1–2)
Stage 14 · Advanced Robotics / Reinforcement Learning
Parts 1–2: RL foundations for robots (MDP, policy/value, exploration, reward shaping) plus value-based deep RL / DQN intuition (why tables break, function approximation, Q-net + replay + target, stability traps). Soft-recommend Stages 11–13; hard locks cleared; no payment gating. Browser educational toys only — NOT MuJoCo/Isaac/stable-baselines. Parts 3–4 (policy gradients, sim-to-real caveats) planned later.