本课程主要介绍强化学习的基础知识,其目标是帮助同学们快速、顺利地进入强化学习及其应用领域的研究工作。课程主要内容包含有限马尔可夫决策过程,动态规划,无模型预测与控制(SASA,Q-Learning),价值函数逼近(DQN),策略梯度方法(REINFORCE),执行者/评论者方法(AC,TRPO,PPO),连续动作空间的确定性策略(DDPG)。
☆18Oct 17, 2022Updated 3 years ago
Alternatives and similar repositories for A05_rl
Users that are interested in A05_rl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 利用遗传算法做基于客流需求的列车时刻表的优化☆15Apr 25, 2021Updated 5 years ago
- 画出列车运行图,给出列车运行的最佳调度☆15Mar 9, 2020Updated 6 years ago
- ☆10Jul 13, 2019Updated 7 years ago
- ☆10Jun 13, 2023Updated 3 years ago
- Markdown 语法文档 整理与修缮☆13Jun 25, 2019Updated 7 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Implementation of the TD3 algorithm written in Pytorch☆12Dec 8, 2022Updated 3 years ago
- Some notes about reinforce learning, self-driving cars and leetcode☆22Mar 26, 2022Updated 4 years ago
- 这是高华的部分。列车运行图综合运用系统☆21Dec 8, 2022Updated 3 years ago