Inverse Scaling in Test-Time Compute
☆26Dec 3, 2025Updated 9 months ago
Alternatives and similar repositories for inverse-scaling-ttc
Users that are interested in inverse-scaling-ttc are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspace☆19Oct 21, 2024Updated last year
- Official implementation for "Mixture of In-Context Experts Enhance LLMs’ Awareness of Long Contexts" (Accepted by Neurips2024)☆14Jan 7, 2025Updated last year
- The official baseline implementations for Chronocept. Published at EACL 2026.☆10Aug 4, 2026Updated last month
- ☆20Nov 4, 2025Updated 10 months ago
- [EMNLP 25] An effective and interpretable weight-editing method for mitigating overly short reasoning in LLMs, and a mechanistic study un…☆20Aug 24, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Code for the EMNLP24 paper "A simple and effective L2 norm based method for KV Cache compression."☆19Dec 13, 2024Updated last year
- 📄 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay☆25Jul 17, 2026Updated 2 months ago
- ☆57Jul 7, 2025Updated last year
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆30Apr 17, 2026Updated 5 months ago
- ☆565Updated this week
- ☆10Jan 28, 2019Updated 7 years ago
- Trace origins, shared sources, and contamination risk☆29May 27, 2026Updated 4 months ago
- Restore safety in fine-tuned language models through task arithmetic☆33Mar 28, 2024Updated 2 years ago
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2024] Code for the paper "MoE-RBench: Towards Building Reliable Language Models with Sparse Mixture-of-Experts"☆11Jul 1, 2024Updated 2 years ago
- This repo contains the source code for: Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs☆44Aug 14, 2024Updated 2 years ago
- [ICLR 2022] "Bayesian Modeling and Uncertainty Quantification for Learning to Optimize: What, Why, and How" by Yuning You, Yue Cao, Tianl…☆14Aug 19, 2022Updated 4 years ago
- this is based on the paper Chain-of-Retrieval Augmented Generation☆15Mar 29, 2025Updated last year
- Explanation Optimization☆13Oct 16, 2020Updated 5 years ago
- GOPHI: an AMR-to-English Verbalizer☆12Feb 5, 2020Updated 6 years ago
- Source code to accompany research paper on training multi token prediction language models using self-distillation.☆42Feb 21, 2026Updated 7 months ago
- Source code and data of our paper "Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation" (https://arxiv.org/…☆11Jun 21, 2023Updated 3 years ago
- ☆13Jan 14, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- OpenVLThinker [NeurIPS 2025] & OpenVLThinkerV2 [COLM 2026]☆157May 25, 2026Updated 4 months ago
- [COLM'25] A Controlled Study on Long Context Extension and Generalization in LLMs☆65Mar 9, 2026Updated 6 months ago
- ☆29Apr 19, 2026Updated 5 months ago
- A neural network for sarcasm detection I trained on the reddit sarcasm database☆13Jul 27, 2017Updated 9 years ago
- Reproduction of "Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization" for the Reproducibility challenge@NeurIPS…☆11Jan 14, 2020Updated 6 years ago
- An algorithm that lets a robot track and follow a person.☆11Jun 1, 2015Updated 11 years ago
- Implementing -- Histogram of oriented gradients / Support Vector Machine / TensorFlow☆11Mar 15, 2017Updated 9 years ago
- Official code for the paper "HEXA-MoE: Efficient and Heterogeneous-Aware MoE Acceleration with Zero Computation Redundancy"☆15Mar 6, 2025Updated last year
- A simple Docker sandbox example and a ready-to-use autograder API. Based on asynchronous FastAPI and disposable Docker containers. Three …☆15Jan 10, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆16Sep 3, 2026Updated last month
- ☆10Apr 26, 2023Updated 3 years ago
- [ACL 2025] An inference-time decoding strategy with adaptive foresight sampling☆108May 18, 2025Updated last year
- ☆15Jan 14, 2026Updated 8 months ago
- [ICML 2023] "Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?" by Ruisi Cai, Zhenyu Zhang, Zhangyang Wang☆16May 4, 2023Updated 3 years ago
- Experiments for efforts to train a new and improved t5☆76Apr 15, 2024Updated 2 years ago
- This repository contains the source code for the article "Towards Feature Selection for Ranking and Classification Exploiting Quantum Ann…☆10Jul 27, 2022Updated 4 years ago