Unofficial Implementation of Selective Attention Transformer
☆20Oct 31, 2024Updated last year
Alternatives and similar repositories for selective-attention-transformer
Users that are interested in selective-attention-transformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Updated this week
- ☆19Jun 3, 2024Updated 2 years ago
- ☆17Sep 11, 2026Updated 2 weeks ago
- Code for "What really matters in matrix-whitening optimizers?"☆25Oct 31, 2025Updated 10 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Unofficial Implementation of Chain-of-Thought Reasoning Without Prompting☆35Mar 19, 2024Updated 2 years ago
- Example code of Sparse Gaussian Process Attention (ICLR 2023)☆26Sep 15, 2025Updated last year
- Code for Augment & Reduce, a scalable stochastic algorithm for large categorical distributions☆10May 16, 2018Updated 8 years ago
- A no-string API framework for deploying schema-based reasoning into third-party apps☆23Updated this week
- pytorch☆10Apr 13, 2022Updated 4 years ago
- ☆26Jun 29, 2025Updated last year
- Mixture of Lora Experts☆11Apr 7, 2024Updated 2 years ago
- ☆10Oct 12, 2021Updated 4 years ago
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for "Exponential Family Estimation via Adversarial Dynamics Embedding" (NeurIPS 2019)☆14Nov 26, 2019Updated 6 years ago
- ☆19Mar 25, 2025Updated last year
- The Compositionality article class.☆14Mar 16, 2026Updated 6 months ago
- ☆13Jun 15, 2021Updated 5 years ago
- ☆13Feb 2, 2023Updated 3 years ago
- Create string diagrams with LaTeX!☆14Jan 3, 2025Updated last year
- Evals meant to evaluate language models' ability to reason over long contexts.☆10Sep 12, 2024Updated 2 years ago
- Self-Teaching Notes on Gradient Leakage Attacks against GPT-2 models.☆14Mar 18, 2024Updated 2 years ago
- Code for Semi-crowdsourced Clustering with Deep Generative Models☆12Dec 9, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Implementation of SelfExtend from the paper "LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning" from Pytorch and Zeta☆13Nov 11, 2024Updated last year
- Code accompanying VarGrad: A Low-Variance Gradient Estimator for Variational Inference☆12Oct 12, 2020Updated 5 years ago
- personal info☆11Mar 23, 2024Updated 2 years ago
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.☆16May 25, 2025Updated last year
- G. Peyré, L. Chizat, F-X. Vialard, J. Solomon, Quantum Optimal Transport for Tensor Field Processing, Arxiv, 2016☆10Apr 13, 2017Updated 9 years ago
- [NeurIPS 2021] "Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models" by Boxin Wang*, Chejian Xu*, Shuoh…☆13Apr 3, 2023Updated 3 years ago
- ☆11Jul 25, 2021Updated 5 years ago
- Stock analysis and prediction - fundamental, quantitative, technical analysis and machine learning.☆13May 1, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆13Mar 10, 2026Updated 6 months ago
- Code for paper 'Zero-Shot Scene Graph Generation via Triplet Calibration and Reduction' (TOMM 2023)☆10Sep 6, 2025Updated last year
- [NeurIPS 2024 Datasets and Benchmarks Track] Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime☆26Mar 27, 2025Updated last year
- PyTorch utilities for ML, specifically speech☆13Jan 30, 2024Updated 2 years ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆20Jul 4, 2025Updated last year
- ☆15Oct 19, 2024Updated last year
- Category Theory for Quantum Natural Language Processing☆11Feb 22, 2023Updated 3 years ago