☆114Sep 3, 2026Updated this week
Alternatives and similar repositories for sotaku
Users that are interested in sotaku are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 4 months ago
- ☆16Aug 25, 2026Updated last week
- Minimal and highly hackable implementation of Looped Transformers with GPT☆25Mar 8, 2026Updated 5 months ago
- A sample pattern for running CI tests on Modal☆19Apr 12, 2025Updated last year
- The best ChatGPT that $100 can buy.☆58Aug 26, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Training code for Sparse Autoencoders on Embedding models☆40Jul 11, 2026Updated last month
- wandb-compatible single-process server, and a bad pun☆18Mar 24, 2026Updated 5 months ago
- ☆49Updated this week
- Well documented examples of running distributed training jobs on Modal☆33Aug 11, 2026Updated 3 weeks ago
- Compile programs directly into transformer weights. Includes a 2D convex-hull KV cache with O(log n) inference.☆217Jun 1, 2026Updated 3 months ago
- Official implementation of Stackelberg PPO for morphology–control co-design.☆17Mar 17, 2026Updated 5 months ago
- ANE accelerated embedding models!☆20Dec 11, 2024Updated last year
- 100M tokens. Infinite compute. Lowest val loss wins.☆530Jul 3, 2026Updated 2 months ago
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.☆121Aug 11, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Modal LLM LLama.cpp based model deployment as part of series of Model as a Service (MaaS)☆17Mar 23, 2026Updated 5 months ago
- [ICML 2026] Code for Equilibrium Reasoners: learning attractor dynamics for scalable reasoning☆48Aug 4, 2026Updated last month
- Pytorch implementation of Deepmind's WaveRNN model☆13Apr 5, 2020Updated 6 years ago
- Storing the LongCoT-mini results for RLM(GPT-5.2)☆20Apr 26, 2026Updated 4 months ago
- ☆88Mar 12, 2026Updated 5 months ago
- ☆12Dec 28, 2021Updated 4 years ago
- A webhook that integrates the W&B model registry with Modal Labs☆15Dec 24, 2023Updated 2 years ago
- Load and run Llama from safetensors files in C☆16Oct 24, 2024Updated last year
- Using modal.com to process FineWeb-edu data☆20Aug 25, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆12Jul 8, 2024Updated 2 years ago
- Because it's there.☆16Sep 22, 2024Updated last year
- A varitation graph tool☆10Dec 23, 2019Updated 6 years ago
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆11Oct 31, 2024Updated last year
- Orchestrate Modal and OpenAI workloads with Dagster☆13Dec 11, 2024Updated last year
- Code for "Evidence of Learned Look-Ahead in a Chess-Playing Neural Network"☆31Jun 4, 2024Updated 2 years ago
- Official implementation of "Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought" (NeurIPS 2025)☆44Oct 8, 2025Updated 10 months ago
- An introduction to LLM Sampling☆80Dec 15, 2024Updated last year
- BPE modification that implements removing of the intermediate tokens during tokenizer training.☆27Nov 25, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Rusty CUDA wrapper☆37Dec 19, 2021Updated 4 years ago
- Official Pytorch Implementation of "Zero-Shot Off-Policy Learning" (ICML 2026)☆25Feb 16, 2026Updated 6 months ago
- Privacy-preserving k-means clustering on data owned by multiple parties☆14May 10, 2016Updated 10 years ago
- Messing around with delimited continuations, fibers, and algebraic effects☆16Oct 2, 2021Updated 4 years ago
- [ICML 2022] Official implementation of "Score-Guided Intermediate Layer Optimization: Fast Langevin Mixing for Inverse Problems".☆12Jul 19, 2022Updated 4 years ago
- various experiments for scaling inference time compute with small reasoning models☆17Jan 16, 2025Updated last year
- Frontend for managing Hack as a Service apps☆14Mar 6, 2023Updated 3 years ago