gzip Predicts Data-dependent Scaling Laws
☆35May 28, 2024Updated 2 years ago
Alternatives and similar repositories for complexity-scaling
Users that are interested in complexity-scaling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Mar 31, 2024Updated 2 years ago
- ☆27Jul 9, 2024Updated 2 years ago
- 0-Shot Tokenizer Transplant☆14May 16, 2025Updated last year
- Multi-Word Probabilistic based supertokenizer☆15May 15, 2025Updated last year
- Adversarial Training and SFT for Bot Safety Models☆41Apr 18, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆44Jun 19, 2024Updated 2 years ago
- High-performance tokenized language data-loader for Python C++ extension☆15Jul 22, 2024Updated 2 years ago
- An implementation of the Llama architecture, to instruct and delight☆21May 31, 2025Updated last year
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 4 months ago
- ☆54May 20, 2024Updated 2 years ago
- LayerNorm(SmallInit(Embedding)) in a Transformer to improve convergence☆61Feb 21, 2022Updated 4 years ago
- ☆16Aug 7, 2026Updated last month
- Tensor-Slayer : Manipulate weights and tensors of LLMs to achieve performance upgrades and introduce a novel inferenceless mechanistic in…☆28May 27, 2025Updated last year
- Tools for formatting large language model prompts.☆13Dec 19, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- PyTorch implementation for MRL☆23Feb 22, 2024Updated 2 years ago
- Simple Transformer in Jax☆143Jun 22, 2024Updated 2 years ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆22Aug 27, 2026Updated last week
- Why Do We Need Weight Decay in Modern Deep Learning? [NeurIPS 2024]☆73Sep 25, 2024Updated last year
- sketch-rnn demo for seoul mediacity biennale 2018☆13Sep 4, 2018Updated 8 years ago
- Your favourite classical machine learning algos on the GPU/TPU☆23Dec 14, 2025Updated 8 months ago
- A light tensor library in zig.☆77Feb 9, 2025Updated last year
- A spinning Rubik's Cube with functionality all made with CSS only☆13Jan 8, 2024Updated 2 years ago
- PDF Reader but with a chat bar...☆11Jan 21, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Direct Preference Optimization for RWKV, aiming for RWKV-5 and 6.☆11Mar 1, 2024Updated 2 years ago
- Easily turn large sets of audio urls to an audio dataset.☆21Dec 27, 2022Updated 3 years ago
- Code for ACL 2023 paper titled "Lifting the Curse of Capacity Gap in Distilling Language Models"☆29Jul 14, 2023Updated 3 years ago
- A toolkit for scaling law research ⚖☆69Jan 27, 2025Updated last year
- [TMLR'25] Official implementation for "Large-Scale Targeted Cause Discovery via Learning from Simulated Data"☆28Sep 30, 2025Updated 11 months ago
- Website for the MIT/Harvard Computational Neuroscience Journal Club☆11Apr 7, 2025Updated last year
- [ICML 2025] From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories and Applications☆52Oct 30, 2025Updated 10 months ago
- a fast implementation of BM25☆10Sep 15, 2022Updated 3 years ago
- ☆18Jun 12, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- TLDR; Leetcode and CSS Battles into one platform☆17Sep 25, 2022Updated 3 years ago
- ☆93Aug 18, 2024Updated 2 years ago
- I modified some code of K-BERT so that it can be fit to English datasets Topics Resources☆11Dec 15, 2022Updated 3 years ago
- An llm wrapper for OpenAI☆13Dec 14, 2024Updated last year
- ☆15Oct 24, 2023Updated 2 years ago
- seqax = sequence modeling + JAX☆199Jul 23, 2025Updated last year
- Repository for the code of the paper "Neural Networks Regularization Through Class-wise Invariant Representation Learning".☆12Oct 1, 2017Updated 8 years ago