☆70Aug 6, 2025Updated last year
Alternatives and similar repositories for gpt-oss-reverse-engineering
Users that are interested in gpt-oss-reverse-engineering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution☆105Sep 24, 2025Updated last year
- [NeurIPS 2022] Your Transformer May Not be as Powerful as You Expect (official implementation)☆35Aug 6, 2023Updated 3 years ago
- 👌[ICLR 2025] TFG-Flow: Training-free Guidance in Multimodal Generative Flow☆20Mar 4, 2025Updated last year
- ☆32Jul 2, 2025Updated last year
- ☆39Dec 14, 2025Updated 9 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆11Apr 3, 2023Updated 3 years ago
- Efficient Scaling laws and collaborative pretraining.☆24Jul 19, 2026Updated 2 months ago
- Simple (fast) transformer inference in PyTorch with torch.compile + lit-llama code☆11Aug 29, 2023Updated 3 years ago
- Accelerate LLM preference tuning via prefix sharing with a single line of code☆52Jul 4, 2025Updated last year
- Code for "Towards Revealing the Mystery behind Chain of Thought: a Theoretical Perspective"☆21Jul 16, 2023Updated 3 years ago
- [ICML 2021] This is the official github repo for training L_inf dist nets with high certified accuracy.☆41Mar 16, 2022Updated 4 years ago
- Homepage for ProLong (Princeton long-context language models) and paper "How to Train Long-Context Language Models (Effectively)"☆266Sep 12, 2025Updated last year
- ☆17Aug 23, 2025Updated last year
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Adversarially Robust Generalization Just Requires More Unlabeled Data☆11Aug 8, 2019Updated 7 years ago
- NVIDIA cuTile learn☆168Dec 9, 2025Updated 9 months ago
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding☆281Aug 31, 2024Updated 2 years ago
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- ☆15Jan 27, 2025Updated last year
- Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation, ICML 2024☆25Jun 26, 2024Updated 2 years ago
- A sumary of MoE experimental setups across a number of different papers.☆16Feb 16, 2023Updated 3 years ago
- ☆53May 20, 2025Updated last year
- Quantized Attention on GPU☆45Nov 22, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Quartet II Official Code☆82May 1, 2026Updated 4 months ago
- ☆20Sep 28, 2024Updated last year
- Unified tracing profiler and visualizer, one timeline from CPU to GPU/HES to in-kernel zones☆147Updated this week
- [ICLR26] AI-based scaling law discovery☆36Jan 30, 2026Updated 7 months ago
- A survey of manufacturer-provided DRAM operating parameters and timings as specified by DRAM chip datasheets from between 1970 and 2021. …☆11May 4, 2022Updated 4 years ago
- ☆158Jun 22, 2023Updated 3 years ago
- prompt tuning, management, playground lab(w/ Python Streamlit, FastAPI)☆13Mar 14, 2025Updated last year
- Code of ICML paper arxiv.org/abs/2302.08105☆14May 4, 2023Updated 3 years ago
- ☆117May 29, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The unofficial CLI of Amazon S3 Vectors (Preview) in Rust☆17Jul 19, 2025Updated last year
- Code for the paper "The Journey, Not the Destination: How Data Guides Diffusion Models"☆27Dec 12, 2023Updated 2 years ago
- ☆15Jul 17, 2024Updated 2 years ago
- ☆42May 26, 2026Updated 3 months ago
- ☆26May 20, 2025Updated last year
- ☆892Sep 15, 2025Updated last year
- Tile-based language built for AI computation across all scales☆191Aug 25, 2026Updated 3 weeks ago