☆56Jun 25, 2026Updated 3 months ago
Alternatives and similar repositories for futuresim
Users that are interested in futuresim are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codebase from our first release.☆63Sep 13, 2026Updated 2 weeks ago
- ☆21Apr 3, 2026Updated 6 months ago
- ☆57Mar 18, 2026Updated 6 months ago
- Official code and dataset for our paper: RefineBench: Evaluating Refinement Capability of Language Models via Checklists☆16Dec 1, 2025Updated 10 months ago
- ☆35Sep 12, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆21Oct 31, 2024Updated last year
- ☆28Jun 22, 2026Updated 3 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆80Updated this week
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆14Jun 28, 2025Updated last year
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated last year
- Repo for Paper: Discovering Interpretable Algorithms by Decompiling Transformers to RASP☆16May 25, 2026Updated 4 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆20Jul 4, 2025Updated last year
- Official code repository for the paper "Internal Activation as the Polar Star for Steering Unsafe LLM Behavior"☆16May 31, 2026Updated 4 months ago
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆75Apr 8, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆33Apr 29, 2026Updated 5 months ago
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.☆20Updated this week
- Storing the LongCoT-mini results for RLM(GPT-5.2)☆20Apr 26, 2026Updated 5 months ago
- ☆120Updated this week
- ☆19Feb 12, 2026Updated 7 months ago
- Repository for the paper: "TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining" ACL Oral 2025☆24Sep 11, 2026Updated 3 weeks ago
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- Head Vis Public Release☆41May 4, 2026Updated 4 months ago
- Source code for the collaborative reasoner research project at Meta FAIR.☆116Mar 26, 2026Updated 6 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Landing page for MIB: A Mechanistic Interpretability Benchmark☆26Aug 15, 2025Updated last year
- The code base for the article "Pivot Based Language Modeling for Improved Neural Domain Adaptation", NAACL 2018☆16Jul 14, 2019Updated 7 years ago
- A drop-in replacement for the standard Categorical Cross-Entropy (CCE) loss that significantly improves OOD and Calibration performance w…☆53Apr 6, 2026Updated 5 months ago
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025☆18Jan 12, 2026Updated 8 months ago
- CS194-196 Course Project☆15Feb 20, 2025Updated last year
- ☆79Mar 6, 2025Updated last year
- [WWW 2026 Oral] MoE-CL:Self-Evolving LLMs via Continual Instruction Tuning☆22Dec 1, 2025Updated 10 months ago
- ☆21Apr 21, 2026Updated 5 months ago
- ☆83Apr 26, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official implementation of the ΔBelief-RL method.☆31Feb 28, 2026Updated 7 months ago
- SHARPIE: Shared Human-AI Reinforcement Learning Platform for Interactive Experiments☆29Aug 3, 2026Updated 2 months ago
- EdgeBench: Unveiling scaling laws of learning from real-world environments☆457Sep 20, 2026Updated last week
- DASH: Detection and Assessment of Systematic Hallucinations of VLMs☆16Jul 2, 2025Updated last year
- Engine for collecting, uploading, and downloading model activations☆31Apr 2, 2025Updated last year
- The code for "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by VIdeo SpatioTemporal Augmentation" [CVPR2025]☆20Feb 27, 2025Updated last year
- ADAG: Transluce's MLP neuron-level circuit tracing library☆39Apr 10, 2026Updated 5 months ago