Trains an agent with (stochastic) Policy Gradients(actor-critic) on Pong. Uses OpenAI Gym.
☆18Jan 10, 2025Updated last year
Alternatives and similar repositories for pong_actor-critic
Users that are interested in pong_actor-critic are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Published by Packt☆11Jan 18, 2021Updated 5 years ago
- This is a caffe implementation to visualize the learnt model (circa 2014)☆60Aug 4, 2021Updated 5 years ago
- Custom ROS messages for Kobuki☆11Oct 28, 2022Updated 3 years ago
- Code for abstracting, evaluating, and visualizing Markov Decision Processes.☆11Jan 12, 2017Updated 9 years ago
- ROBEL: Robotics Benchmarks for Learning with low-cost robots (dev fork)☆13Jun 16, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ROS package for robot learning☆17Oct 16, 2019Updated 6 years ago
- ☆14Dec 10, 2017Updated 8 years ago
- Experimentation with Streamlit for personal LLM tool☆15Jun 19, 2023Updated 3 years ago
- ☆12Dec 8, 2016Updated 9 years ago
- [DEPRECATED] Advantage Actor Critic model in PyTorch inspired by OpenAI baselines TensorFlow implementation☆52Feb 4, 2020Updated 6 years ago
- Robust Multi-Agent Reinforcement Learning with State Uncertainty☆12May 30, 2023Updated 3 years ago
- Code for "Predictive-Corrective Networks for Action Detection"☆16Nov 29, 2017Updated 8 years ago
- Retrieve metrics from the InCites web service☆12Jan 20, 2021Updated 5 years ago
- In this repository I'll be programming the cool exercises of the Book Reinforcement-Learning: An introduction by Sutton☆14Apr 15, 2018Updated 8 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Temporal Difference Learning based Backgammon game using Neural Network based model☆11Mar 13, 2018Updated 8 years ago
- This is an implimentation of Value Iteration Networks (NIPS2016 best paper) in keras☆17Jan 6, 2018Updated 8 years ago
- An opensource implementation of kanerva coding for use in reinforcement learning research☆11Mar 28, 2026Updated 5 months ago
- 𝔸𝕄𝔹ℝ𝕆𝕊𝕀𝔸: A Benchmark for Parsing Ambiguous Questions into Database Queries☆16Oct 31, 2024Updated last year
- A CUDA implementation of the ZeroOut tensorflow custom op, just for fun☆11Feb 1, 2017Updated 9 years ago
- Use pytorch the right way http://pytorch.org/docs/☆22Nov 15, 2017Updated 8 years ago
- Audio Masking Methods☆12Nov 15, 2019Updated 6 years ago
- A set of core libraries for useful DSP related classes that are used by multiple White Elephant Audio VSTs and Audio Units☆15Apr 18, 2026Updated 5 months ago
- A summary of my recently surveyed papers. Some papers on Arxiv with unimpressive results are not included.☆25Apr 18, 2018Updated 8 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- self implementation of DPPO, Distributed Proximal Policy Optimization, by using tensorflow☆12Sep 1, 2017Updated 9 years ago
- Voice Music Separation competing for 6th Huawei Cup in ZJU☆11Jun 2, 2015Updated 11 years ago
- ☆13Nov 8, 2022Updated 3 years ago
- Predictive State Recurrent Neural Networks☆18May 18, 2020Updated 6 years ago
- Matlab implementation for the paper "Efficient Belief Propagation for Early Vision"☆11Jun 9, 2016Updated 10 years ago
- Mancs: A multi-task attentional network with curriculum sampling for person re-identification☆13Aug 5, 2019Updated 7 years ago
- Repo for PyData 2018 tuorial☆12Oct 18, 2018Updated 7 years ago
- Extract images and annotation files from the Caltech Pedestrian Dataset.☆14Nov 29, 2022Updated 3 years ago
- A Python library for parsing OSM streams.☆15May 8, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Trust Region Policy Optimization with Generalized Advantage Estimator☆16Nov 15, 2018Updated 7 years ago
- Vue app for https://github.com/bearpelican/musicautobot☆17Dec 10, 2022Updated 3 years ago
- DF2Net☆13Aug 18, 2018Updated 8 years ago
- Using Deep Learning to predict audio quality.☆18Jan 31, 2020Updated 6 years ago
- Using Multiple GPU with tensorflow☆13Dec 28, 2018Updated 7 years ago
- Working Theano implementation of Pixel RNN on MNIST.☆76Jul 12, 2016Updated 10 years ago
- Python command line tool to search and download guitar tabs from ultimate-guitar.com☆16Mar 3, 2019Updated 7 years ago