Unofficial PyTorch implementation of Fastformer based on paper "Fastformer: Additive Attention Can Be All You Need"."
☆131Sep 6, 2021Updated 4 years ago
Alternatives and similar repositories for Fastformer-PyTorch
Users that are interested in Fastformer-PyTorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A pytorch &keras implementation and demo of Fastformer.☆192Sep 22, 2022Updated 3 years ago
- Implementation of Fast Transformer in Pytorch☆176Aug 26, 2021Updated 4 years ago
- ☆13Aug 13, 2020Updated 6 years ago
- FastFormers - highly efficient transformer models for NLU☆705Mar 21, 2025Updated last year
- Code for experiments done for EMNLP2020.☆11Dec 8, 2022Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Optimizing bit-level Jaccard Index and Population Counts for large-scale quantized Vector Search via Harley-Seal CSA and Lookup Tables☆22May 18, 2025Updated last year
- ☆19Oct 10, 2020Updated 5 years ago
- FairSeq repo with Apollo optimizer☆113Dec 20, 2023Updated 2 years ago
- Trains Transformer model variants. Data isn't shuffled between batches.☆147Oct 5, 2022Updated 3 years ago
- transformers go brrr...☆148Feb 15, 2022Updated 4 years ago
- Interactive tree-maps with SBERT & Hierarchical Clustering (HAC)☆30Dec 31, 2024Updated last year
- ☆28Oct 6, 2020Updated 5 years ago
- [NAACL 2021] This is the code for our paper `Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self…☆205Aug 17, 2022Updated 3 years ago
- Locally Enhanced Self-Attention: Rethinking Self-Attention as Local and Context Terms☆20Nov 29, 2021Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An open-source AutoML Library based on PyTorch☆306Jul 6, 2026Updated last month
- Tagger for explicit cause-and-effect relationships in text☆11Jan 8, 2020Updated 6 years ago
- On Generating Extended Summaries of Long Documents☆78Jan 26, 2021Updated 5 years ago
- WaveGlow vocoder with VQVAE☆61Jun 18, 2019Updated 7 years ago
- ☆31Jan 16, 2021Updated 5 years ago
- Implementation of Perceiver, General Perception with Iterative Attention, in Pytorch☆1,218Jun 8, 2026Updated 2 months ago
- A series of BERT and Albert model checkpoints trained to reduce gendered correlations in pre-training☆11Oct 22, 2020Updated 5 years ago
- ☆74Jul 2, 2021Updated 5 years ago
- ☆17Sep 22, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- (ACL-IJCNLP 2021) Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models.☆21Jul 13, 2022Updated 4 years ago
- Pytorch implementation of Compressive Transformers, from Deepmind☆165Oct 4, 2021Updated 4 years ago
- SummVis is an interactive visualization tool for text summarization.☆253Jun 17, 2022Updated 4 years ago
- Meta learning for generative models.☆16Jul 24, 2019Updated 7 years ago
- Sessa: Selective State Space Attention☆18Apr 28, 2026Updated 3 months ago
- VirusMVP is an interactive heatmap-centric app that integrates viral genomic mutations, lineage information and curated functional impact…☆15Jul 21, 2026Updated 3 weeks ago
- Some demos using Nvidia RAPIDS for Cheminformatics☆14Aug 17, 2020Updated 5 years ago
- An intelligent, flexible grammar of machine learning.☆82Jul 29, 2021Updated 5 years ago
- ☆75Nov 19, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- An implementation of Performer, a linear attention-based transformer, in Pytorch☆1,181Feb 2, 2022Updated 4 years ago
- Locality Preserving Dense Graph Convolutional Networks with Graph Context-Aware Node Representations☆11Jan 21, 2021Updated 5 years ago
- ☆184May 26, 2023Updated 3 years ago
- Binary Passage Retriever (BPR) - an efficient passage retriever for open-domain question answering☆175Jun 6, 2021Updated 5 years ago
- Demo for DART, Audio Imagination workshop submission in NeurIPS 2024☆16Apr 22, 2026Updated 3 months ago
- The code related to the paper☆14Mar 1, 2023Updated 3 years ago
- Implementation of H-Transformer-1D, Hierarchical Attention for Sequence Learning☆167Feb 12, 2024Updated 2 years ago