Stage 14 · Advanced Robotics / Reinforcement Learning · Lesson 2 · 12–18 min
States, Actions & Rewards
An MDP is the robot’s decision loop written as math: state, action, reward, next state — the learning twin of Stage 13’s world model + act cycle.
Stage 14 · Advanced Robotics / Reinforcement Learning · Lesson 2 · 12–18 min
An MDP is the robot’s decision loop written as math: state, action, reward, next state — the learning twin of Stage 13’s world model + act cycle.