针对最经典的表格型Q learning算法进行了复现,能够支持gym中大多数的离散动作和状态空间的环境,譬如CliffWalking-v0。
☆10Jan 2, 2021Updated 5 years ago
Alternatives and similar repositories for Q-learning
Users that are interested in Q-learning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementations of Influential Recommender System☆12Oct 29, 2024Updated last year
- The code of paper LMC: Fast Training of GNNs via Subgraph Sampling with Provable Convergence. Zhihao Shi, Xize Liang, Jie Wang. ICLR 2023…☆48Feb 15, 2023Updated 3 years ago
- 基于 LeRobot 和 MuJoCo 的机器人学习教程,包含 ACT、pi0、SmolVLA 模型的完整复现:数据采集、训练与部署。☆16Apr 26, 2026Updated 2 months ago
- Python 高级编程☆15Dec 18, 2019Updated 6 years ago
- Langchain Agent finetuning using 7B - LLAMA 2 , on hotpotQA (Retroformer framework)☆16Sep 5, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Analyse Social Network of co-authors in DBLP website (https://dblp.uni-trier.de) using NetworkX.☆13May 27, 2020Updated 6 years ago
- An analyzer for PE files☆16Jan 14, 2026Updated 6 months ago
- An ASCII Header Generator for Network Protocols☆14Dec 12, 2024Updated last year
- A solutions manual for Set Theory by Thomas Jech☆14Aug 12, 2018Updated 7 years ago
- ☆17Nov 3, 2024Updated last year
- A novel template-free retrosynthesizer that can generate diverse sets of reactants for a desired product via discrete conditional variati…☆15Aug 7, 2022Updated 3 years ago
- ☆11Jan 6, 2024Updated 2 years ago
- Executive control code for STRANDS robots.☆11Feb 13, 2020Updated 6 years ago
- The code of paper *Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy Optimization*.☆18Mar 26, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Must-read papers on Knowledge Graph Reasoning (KGR)☆21Mar 16, 2020Updated 6 years ago
- ☆13Oct 5, 2021Updated 4 years ago
- Implementation of the paper Unsupervised Domain Adaptation by Backpropagation☆11Dec 1, 2018Updated 7 years ago
- Code and data for paper named: Large language models for automatic equation discovery of nonlinear dynamics☆14Mar 6, 2025Updated last year
- Neural theorem proving evaluation via the Lean REPL☆24Jul 12, 2025Updated last year
- ☆20Jan 26, 2024Updated 2 years ago
- ☆11Apr 23, 2025Updated last year
- OpenAI gym environments for goal-conditioned and language-conditioned reinforcement learning☆14Jan 27, 2026Updated 5 months ago
- ☆10Jun 7, 2021Updated 5 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [T-RO] Python implementation of PRobabilistically-Informed Motion Primitives (PRIMP)☆14Apr 19, 2024Updated 2 years ago
- ☆15Jul 7, 2024Updated 2 years ago
- Deep Q learning algorithm written on PyTorch for solving 2D robot arm reacher☆12Feb 19, 2020Updated 6 years ago
- SSGCN☆11Jul 23, 2020Updated 6 years ago
- 最基本的基于蒙特卡洛搜索树(MCTS)的五子棋。☆13Apr 8, 2021Updated 5 years ago
- ☆11Jul 1, 2024Updated 2 years ago
- Official Code Repository for the POLICEd-RL Paper: https://www.roboticsproceedings.org/rss20/p104.html☆14Mar 4, 2025Updated last year
- Code for Policy Bifurcation in Safe Reinforcement Learning☆10Jul 4, 2025Updated last year
- ☆38Dec 26, 2022Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Must-read papers on Knowledge Graph Embedding☆29Oct 15, 2020Updated 5 years ago
- ☆30Dec 27, 2024Updated last year
- [NeurIPS 2023] Large Language Models Are Semi-Parametric Reinforcement Learning Agents☆40May 2, 2024Updated 2 years ago
- A molecule generative model used interaction fingerprint (docking pose) as constraints.☆15Feb 13, 2022Updated 4 years ago
- Android一些我看到的开源项目☆20Sep 11, 2017Updated 8 years ago
- ☆11Jul 29, 2021Updated 4 years ago
- Code repository for the CoRL 2021 paper "RoCUS: Robot Controller Understanding via Sampling"☆11Mar 24, 2022Updated 4 years ago