AIHUB Robotics Lab

Stage 14 · Advanced Robotics / Reinforcement Learning · Lesson 2 · 12–18 min

States, Actions & Rewards

An MDP is the robot’s decision loop written as math: state, action, reward, next state — the learning twin of Stage 13’s world model + act cycle.