Reinforcement Learning (Q-Learning)

Watch an agent learn to navigate a grid world through trial and error. Rewards and penalties guide the agent to discover optimal paths (policies).

Speed200ms
Learning Rate (α)0.50
Discount Factor (γ)0.95

Training Metrics

Episodes
0
Epsilon
1.000
Success (last 10)
0%
Best steps
Avg steps (10)
Steps-per-episode chart appears after training…

The agent uses Q-learning to navigate to the Goal (green) while avoiding the Pit (red) and Obstacles (black). Arrows show the learned policy; brighter arrows = higher confidence. Once the greedy path reaches the goal and the last 5 episodes succeed, a solid line marks the discovered shortest route.

S
X
G
StartGoal (+100)Pit (−50)ObstacleAgentCurrent best guess
Go Deeper

Learn Reinforcement Learning on DataCamp

Curated courses and career tracks to take your understanding from this demo to real-world mastery. All links open directly on DataCamp.

DataCamp
Cirby AI

Cirby AI

ML / AI Mastery Learning Assistant

Powered by Gemini AI • ML / AI Mastery