[ICLR 2025 Spotlight] Official Implementation for ToST (Token Statistics Transformer)
☆135Feb 25, 2025Updated last year
Alternatives and similar repositories for ToST
Users that are interested in ToST are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository of Polarity-aware Linear Attention for Vision Transformers (ICLR 2025)☆91Aug 12, 2026Updated 2 weeks ago
- Flash-Linear-Attention models beyond language☆21Aug 28, 2025Updated last year
- ☆17Feb 23, 2025Updated last year
- This repository includes the official implementation our paper "Scaling White-Box Transformers for Vision"☆48Jun 3, 2024Updated 2 years ago
- Source code of our ICML 2025 paper "Flowing Datasets with Wasserstein over Wasserstein Gradient Flows"☆21May 21, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official PyTorch implementation of The Linear Attention Resurrection in Vision Transformer☆15Sep 7, 2024Updated last year
- [𝗜𝗖𝗠𝗟 𝟮𝟬𝟮𝟲] Dispersion loss counteracts embedding condensation and improves generalization in small language models☆19May 21, 2026Updated 3 months ago
- [ICML 2025] Fourier Position Embedding: Enhancing Attention’s Periodic Extension for Length Generalization☆120Jun 2, 2025Updated last year
- [ECCV2022] Gumbel Optimised Loss for Long Tailed Instance Segmentation.☆18Nov 24, 2022Updated 3 years ago
- Offical implementation of "Temporal-wise Attention Spiking Neural Networks for Event Streams Classification" (ICCV2021)☆23Oct 16, 2023Updated 2 years ago
- Official implementation of the paper "Scalable Image Coding for Humans and Machines Using Feature Fusion Network".☆16Jan 8, 2025Updated last year
- Code for APLA: A Simple Adaptation Method for Vision Transformers☆16Apr 3, 2025Updated last year
- The code for the paper "LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling" (NeurIPS'24).☆15Dec 25, 2024Updated last year
- Code for the paper "Cottention: Linear Transformers With Cosine Attention"☆21Nov 15, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Here is the resources and code for the LotteryCodec.☆28Nov 3, 2025Updated 9 months ago
- Source codes of Learning Causal Representations for Robust Domain Adaptation (IEEE TKDE)☆12Feb 14, 2022Updated 4 years ago
- 复现Drone-YOLOv8s,论文三明治结构中DW卷积核存在疑点,均改为3*3.☆31Jul 15, 2024Updated 2 years ago
- [NeurIPS 2024] Image Understanding Makes for A Good Tokenizer for Image Generation☆21Dec 17, 2024Updated last year
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs☆71Mar 22, 2026Updated 5 months ago
- [ICLR 2025 & COLM 2025] Official PyTorch implementation of the Forgetting Transformer and Adaptive Computation Pruning☆152Feb 25, 2026Updated 6 months ago
- [NeurIPS25 Spotlight] Official Implementation for CBSA (Contract-and-Broadcast Self-Attention)☆36Apr 3, 2026Updated 4 months ago
- All Points Matter: Entropy-Regularized Distribution Alignment for Weakly-supervised 3D Segmentation (NeurIPS 2023)☆32Nov 3, 2023Updated 2 years ago
- Code for the paper "Interpreting and Improving Diffusion Models from an Optimization Perspective", appearing in ICML 2024☆15Sep 30, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICCV2025 highlight]Rectifying Magnitude Neglect in Linear Attention☆64Jul 24, 2025Updated last year
- Contains scripts to allow reproducibility of all results shown in "TeLU Activation Function for Fast and Stable Deep Learning"☆47Jan 6, 2025Updated last year
- A Tight-fisted Optimizer (Tiger), implemented in PyTorch.☆12Jun 26, 2024Updated 2 years ago
- ☆15Apr 6, 2023Updated 3 years ago
- UniGSC is a highly modular and extensible framework for compressing static and dynamic Gaussian Splats, supporting both video and point c…☆20Dec 18, 2025Updated 8 months ago
- Stick-breaking attention☆63Jul 1, 2025Updated last year
- ☆47Jun 16, 2025Updated last year
- Official implementation of "ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot Learning" [CVPR 2…☆30Sep 1, 2025Updated last year
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMs☆17Jul 18, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated 2 years ago
- [CVPR2025] Breaking the Low-Rank Dilemma of Linear Attention☆45Mar 11, 2025Updated last year
- Unofficial implementation of Tensorial Radiance Fields (Chen & Xu ‘22)☆38Feb 17, 2026Updated 6 months ago
- [ ICLR 2025 ] CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compression☆19Apr 4, 2025Updated last year
- A high-efficiency text embedding and reranking model based on RWKV architecture.☆20Jan 10, 2026Updated 7 months ago
- ☆39Oct 16, 2024Updated last year
- Official pytorch implementation for ESCNet:Edge-Semantic Collaborative Network for Camouflaged Object Detection☆23Jul 3, 2026Updated last month