☆19Oct 30, 2025Updated 9 months ago
Alternatives and similar repositories for KnapsackRL
Users that are interested in KnapsackRL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for ICLR 2022 Paper (HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning)☆12Nov 28, 2023Updated 2 years ago
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)☆59May 12, 2025Updated last year
- Code for Paper (Policy Optimization in RLHF: The Impact of Out-of-preference Data)☆29Dec 19, 2023Updated 2 years ago
- Irene is a python package that aims to be a toolkit for global optimization problems that can be realized algebraically. It generalizes L…☆15Updated this week
- Code for Paper (ReMax: A Simple, Efficient and Effective Reinforcement Learning Method for Aligning Large Language Models)☆202Dec 16, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆10Jul 13, 2024Updated 2 years ago
- Code for Adam-mini: Use Fewer Learning Rates To Gain More https://arxiv.org/abs/2406.16793☆457May 13, 2025Updated last year
- Code for the paper: Why Transformers Need Adam: A Hessian Perspective☆65Mar 11, 2025Updated last year
- Implementation of Action Matching for the Schrödinger equation☆25Jun 18, 2023Updated 3 years ago
- ☆23Feb 4, 2025Updated last year
- Official implementation of the paper `Augmenting GAIL with BC for sample efficient imitation learning` in PyTorch☆35Jan 3, 2021Updated 5 years ago
- ☆14May 30, 2019Updated 7 years ago
- This project applies Monte Carlo Tree Search (MCTS) to a simple grid world.☆10May 30, 2018Updated 8 years ago
- Official codebase for CuGRO: Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay☆34Apr 14, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- sc14 matlab application☆14Nov 24, 2014Updated 11 years ago
- TD-VAE in PyTorch☆10May 28, 2019Updated 7 years ago
- Revisiting Peng's Q(lambda) for Modern Reinforcement Learning☆15Jul 23, 2021Updated 5 years ago
- ☆15Jan 27, 2025Updated last year
- ☆15Feb 22, 2018Updated 8 years ago
- AISTATS 2019: Reference-based Adversarial Sampling & Its applications to Soft Q-learning☆15Jan 21, 2019Updated 7 years ago
- ☆10Oct 18, 2021Updated 4 years ago
- Welcome to the land of C Neuro.☆51Jun 28, 2026Updated last month
- Ant Gather and Ant Maze envs, separated from RLLab☆11Aug 2, 2018Updated 8 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A fun review of spectral clustering with MATLAB demos I made for the NU machine learning meetiup in 2014☆12Mar 4, 2016Updated 10 years ago
- 16824 homework: weakly supervised object detection with PyTorch☆13Sep 5, 2018Updated 7 years ago
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search in ACL'25☆103Jun 16, 2025Updated last year
- No.3 solution of Tianchi ImageNet Adversarial Attack Challenge.☆12Apr 22, 2020Updated 6 years ago
- Reproduction of the complete process of DeepSeek-R1 on small-scale models, including Pre-training, SFT, and RL.☆31Mar 11, 2025Updated last year
- This is the CUDA GPU implementation + Python interface (using PyTorch) of DCI. The paper can be found at https://arxiv.org/abs/1512.00442…☆13Dec 20, 2023Updated 2 years ago
- ☆15Jan 9, 2026Updated 7 months ago
- Levin tree search guided by both a policy and a heuristic function☆19Jul 13, 2023Updated 3 years ago
- Exploration Strategies for Deep Reinforcement Learning☆39Oct 31, 2018Updated 7 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- python algorithms to solve sparse linear programming problems☆34Jul 6, 2023Updated 3 years ago
- 我们是第一个完全可商用的角色大模型。☆39Aug 11, 2024Updated 2 years ago
- Using PyTorch autograd to compute Hessian of Perplexity for Large Language Models☆29Apr 17, 2025Updated last year
- Reinforcement learning tutorials using the rlberry library.☆18Jan 9, 2023Updated 3 years ago
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆20Jan 19, 2025Updated last year
- Code and Data for "Characterizing Multi-Domain False News on Weibo and the Underlying User Effects"☆19Aug 24, 2022Updated 3 years ago
- This repository contains the replication of the iGSM dataset generation process from the Physics of LLM paper by Zeyuan Zhu.☆17Sep 13, 2024Updated last year