Verifiers for LLM Reinforcement Learning
☆84Sep 11, 2025Updated 11 months ago
Alternatives and similar repositories for verifiers-deepresearch
Users that are interested in verifiers-deepresearch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Updated this week
- ☆69May 23, 2025Updated last year
- aka "Bayesian Methods for Hackers": An introduction to Bayesian methods + probabilistic programming with a computation/understanding-firs…☆11May 9, 2015Updated 11 years ago
- Waffer-thin FlaskGPT on Vercel.☆12Jun 1, 2023Updated 3 years ago
- Exploring Applications of GRPO☆252Aug 25, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆21Mar 25, 2025Updated last year
- Paper Implementation of Self-Rewarding Language Models☆13Feb 1, 2024Updated 2 years ago
- Official repo for "Binary Retrieval-augmented Reward Mitigates Hallucinations"☆16Nov 13, 2025Updated 9 months ago
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆30Nov 20, 2024Updated last year
- RAG Tool using Haystack, Mistral, and Chainlit. All open source stack on CPU.☆23Oct 14, 2023Updated 2 years ago
- Seamless Voice Interactions with LLMs☆12Oct 28, 2023Updated 2 years ago
- Codebase exploration with AI research agents☆21Feb 25, 2025Updated last year
- [ICML 2025] Logits are All We Need to Adapt Closed Models☆23May 2, 2025Updated last year
- Project of ACL 2025 "UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models"☆15Mar 25, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Apps that run on modal.com☆13Sep 14, 2025Updated 11 months ago
- Lego for GRPO☆30May 27, 2025Updated last year
- ☆22Nov 28, 2024Updated last year
- Using open source LLMs to build synthetic datasets for direct preference optimization☆72Feb 29, 2024Updated 2 years ago
- A proxy for minimax-m2, enabling interleaved thinking, and tool calls.☆39Nov 21, 2025Updated 9 months ago
- ☆26Nov 26, 2024Updated last year
- ☆28May 19, 2025Updated last year
- Testing paligemma2 finetuning on reasoning dataset☆18Dec 28, 2024Updated last year
- j1-micro (1.7B) & j1-nano (600M) are absurdly tiny but mighty reward models.☆108Jul 19, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- macOS computer use CLI — screenshots, input simulation, app management, session orchestration☆25Mar 24, 2026Updated 5 months ago
- Visual demo of DSPy's prompt optimization on Gradio☆17Apr 14, 2025Updated last year
- Demo of knowledge graph creation and Graph RAG with Dspy and Kuzu☆22Jun 30, 2025Updated last year
- ☆10Jul 13, 2024Updated 2 years ago
- Our library for RL environments + evals☆4,593Updated this week
- Embedding models from Jina AI☆66Jan 18, 2024Updated 2 years ago
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated 2 years ago
- ☆30Oct 7, 2024Updated last year
- aws lambda bash template, lambda bash shell script wrapped in nodejs☆10Sep 4, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Fused Qwen3 MoE layer for faster training, compatible with Transformers, LoRA, bnb 4-bit quant, Unsloth. Also possible to train LoRA over…☆258Jul 24, 2026Updated last month
- Implementing cognitive architecture and psychological memory concepts into Agentic LLM Systems☆551Dec 12, 2024Updated last year
- Implementation of Recursive Language Model paper from scratch☆47Feb 10, 2026Updated 6 months ago
- Mixtral finetuning☆19Feb 2, 2024Updated 2 years ago
- Train an agent to generate high quality summaries☆45Jan 29, 2026Updated 7 months ago
- Luth is a state-of-the-art series of fine-tuned LLMs for French☆47Oct 12, 2025Updated 10 months ago
- A command-line utility to manage MLX models between your Hugging Face cache and LM Studio.☆89Nov 11, 2025Updated 9 months ago