Sparse Attention with Linear Units
☆20Apr 21, 2021Updated 5 years ago
Alternatives and similar repositories for rectified-linear-attention
Users that are interested in rectified-linear-attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Contrastive Object-level Pre-training with Spatial Noise Curriculum Learning☆20Feb 4, 2022Updated 4 years ago
- Code for ICML 2020 paper: Do RNN and LSTM have Long Memory?☆17Jan 6, 2021Updated 5 years ago
- ☆13Mar 30, 2022Updated 4 years ago
- [ICLR 2025] Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception☆16Jul 4, 2025Updated last year
- MDRDC dataset and used baselines☆11Feb 20, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This is the public github for our paper "Transformer with a Mixture of Gaussian Keys"☆29Aug 13, 2022Updated 4 years ago
- Multivariate time-series forecasting with LSTNET and soft-DTW loss☆29Jun 3, 2020Updated 6 years ago
- Custom Keras layers for implementing multi-dimensional recurrent neural networks (MDRNNs) described in Alex Graves's paper https://arxiv.…☆10Apr 27, 2020Updated 6 years ago
- A PyTorch implement of Dilated RNN☆11Dec 31, 2017Updated 8 years ago
- EMNLP 2021: A Label-Aware BERT Attention Network for Zero-Shot Multi-Intent Detection in Spoken Language Understanding☆10Apr 8, 2022Updated 4 years ago
- 软微新圣经----大兴究竟有什么可以输?☆14Sep 18, 2022Updated 4 years ago
- ☆11Mar 8, 2022Updated 4 years ago
- datetime模块的C语言实现,《奔跑吧,Python君》系列相关代码☆10Apr 30, 2023Updated 3 years ago
- ☆14Apr 1, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The repo for ACL2021 findings paper - Don't Miss the Labels: Label-semantic Argumented Meta-Learner for Few-Shot Text Classification☆15Mar 24, 2022Updated 4 years ago
- 面向 Codex、Claude Code、OpenAI Agents 等 AI Agent 的专业个股投研 Skill。☆35May 14, 2026Updated 4 months ago
- 基于预训练BERT和GAT的剧本角色情绪识别研究☆13Dec 15, 2023Updated 2 years ago
- Draw 3D bounding box for objects on image. Based on Tensorflow☆12Apr 10, 2019Updated 7 years ago
- NLP 相关岗位 笔试面试资源汇总☆16Jun 17, 2021Updated 5 years ago
- 中国人民大学 YOJ 题库☆15Jun 9, 2022Updated 4 years ago
- An implementation of Transformer with Expire-Span, a circuit for learning which memories to retain☆34Oct 30, 2020Updated 5 years ago
- Official code for the paper "Why Do Self-Supervised Models Transfer? Investigating the Impact of Invariance on Downstream Tasks".☆16Dec 7, 2021Updated 4 years ago
- Code for the paper "Adaptive Transformers for Learning Multimodal Representations" (ACL SRW 2020)☆43Oct 20, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- AAAI 2024, "Working Memory Capacity of ChatGPT: An Empirical Study".☆15Feb 10, 2025Updated last year
- Code of our IJCAI2021 paper: "Learning Class-Transductive Intent Representations for Zero-shot Intent Detection"☆15Sep 10, 2021Updated 5 years ago
- ☆11Jun 28, 2020Updated 6 years ago
- ☆33Apr 12, 2021Updated 5 years ago
- [ACL‘20] Highway Transformer: A Gated Transformer.☆33Dec 5, 2021Updated 4 years ago
- [ECCV 2020] Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes