SyncAI.news, a Varaisys broadcasting
Deep Q-Learning with Space Invaders
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

Deep Q-Learning with Space Invaders

Unit 3, of the Deep Reinforcement Learning Class with Hugging Face 🤗

⚠️ A new updated version of this article is available here 👉 https://huggingface.co/deep-rl-course/unit1/introduction

This article is part of the Deep Reinforcement Learning Class. A free course from beginner to expert. Check the syllabus here.

⚠️ A new updated version of this article is available here 👉 https://huggingface.co/deep-rl-course/unit1/introduction

This article is part of the Deep Reinforcement Learning Class. A free course from beginner to expert. Check the syllabus here.

In the last unit, we learned our first reinforcement learning algorithm: Q-Learning, implemented it from scratch, and trained it in two environments, FrozenLake-v1 ☃️ and Taxi-v3 🚕.

We got excellent results with this simple algorithm. But these environments were relatively simple because the State Space was discrete and small (14 different states for FrozenLake-v1 and 500 for Taxi-v3).

But as we'll see, producing and updating a Q-table can become ineffective in large state space environments.

So today, we'll study our first Deep Reinforcement Learning agent: Deep Q-Learning. Instead of using a Q-table, Deep Q-Learning uses a Neural Network that takes a state and approximates Q-values for each action based on that state.

And we'll train it to play Space Invaders and other Atari environments using RL-Zoo, a training framework for RL using Stable-Baselines that provides scripts for training, evaluating agents, tuning hyperparameters, plotting results, and recording videos.

So let’s get started! 🚀

To be able to understand this unit, you need to understand Q-Learning first.

  • From Q-Learning to Deep Q-Learning
  • The Deep Q Network
    • Preprocessing the input and temporal limitation
  • The Deep Q-Learning Algorithm
    • Experience Replay to make more efficient use of experiences
    • Fixed Q-Target to stabilize the training
    • Double DQN

From Q-Learning to Deep Q-Learning

The Q comes from "the Quality" of that action at that state.

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News