PyTorch implementation of the Differential-Transformer architecture for sequence modeling, specifically tailored as a decoder-only model similar to large language models (LLMs). The architecture incorporates a novel Differential Attention mechanism, Multi-Head structure, RMSNorm, and SwiGLU.
☆86Oct 27, 2024Updated last year
Alternatives and similar repositories for Differential-Transformer-PyTorch
Users that are interested in Differential-Transformer-PyTorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An open source community implementation of the model from "DIFFERENTIAL TRANSFORMER" paper by Microsoft.☆41Updated this week
- ☆13Oct 14, 2024Updated last year
- Implementation of BitNet1.58b☆15Jul 9, 2024Updated 2 years ago
- Linear Algbra lib in MoonBit☆14Apr 26, 2026Updated 3 months ago
- Pandas lib in Moonbit☆10Apr 26, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆20Updated this week
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated last year
- A toolkit for researchers in the multimodal sound separation.☆16Oct 20, 2023Updated 2 years ago
- Exquisite video generation☆15Feb 18, 2024Updated 2 years ago
- GoldFinch and other hybrid transformer components☆16Dec 9, 2025Updated 7 months ago
- noise reduction☆17Jul 3, 2024Updated 2 years ago
- A PyTorch implementation of Determinantal Point Process Likelihoods for Sequential Recommendation☆12Dec 9, 2024Updated last year
- Official repo of Future-aware Diverse Trends Framework for Recommendation☆11Jul 22, 2022Updated 4 years ago
- Spatial Spectral Machine Learning☆14Oct 15, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation: "AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation"☆19Oct 9, 2025Updated 9 months ago
- Automatic speech annotator processing speech with voice activaty detection, overlapping speech detection, speaker diarization and automat…☆33Jun 14, 2024Updated 2 years ago
- code for "Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification"☆10Mar 19, 2022Updated 4 years ago
- Weakly-Supervised Residual Evidential Learning for Multi-Instance Uncertainty Estimation (ICML 2024)☆15Jul 19, 2024Updated 2 years ago
- [CVPR'24] Solving the Catastrophic Forgetting Problem in Generalized Category Discovery https://arxiv.org/pdf/2501.05272☆16Dec 24, 2024Updated last year
- ☆13Apr 26, 2026Updated 3 months ago
- ☆24Sep 25, 2024Updated last year
- This is the our implementation for the paper: Exploring Mixed Information Flow for Cross-domain Sequential Recommendations☆12Aug 17, 2020Updated 5 years ago
- Entity-Aware and Motion-Aware Transformers for Language-driven Action Localization(IJCAI-22)☆12Oct 11, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Streaming Audiotransformers for online Audio tagging☆57Jun 14, 2024Updated 2 years ago
- ☆12Mar 30, 2021Updated 5 years ago
- HGRN2: Gated Linear RNNs with State Expansion☆58Aug 20, 2024Updated last year
- Med-DANet Series (ECCV 2022 & WACV 2024)☆13Jan 2, 2024Updated 2 years ago
- ☆27Apr 2, 2024Updated 2 years ago
- ☆14Jul 11, 2023Updated 3 years ago
- Sequence alignement methods with helpers for PyTorch.☆24Nov 30, 2022Updated 3 years ago
- [ICMR 2025] Official Repository for The Paper, Let Network Decide What to Learn: Symbolic Music Understanding Model Based on Large-scale …☆19Aug 17, 2025Updated 11 months ago
- NanoGPT (124M) in 5 minutes☆16Feb 14, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Implementation of: Hydra Attention: Efficient Attention with Many Heads (https://arxiv.org/abs/2209.07484)☆14Jan 8, 2023Updated 3 years ago
- 【ICML2026】Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning☆27May 18, 2026Updated 2 months ago
- This repo holds the official code and data for "Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with H…☆15May 21, 2024Updated 2 years ago
- [ICML 2025] Fourier Position Embedding: Enhancing Attention’s Periodic Extension for Length Generalization☆118Jun 2, 2025Updated last year
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention (NeurIPS'25 Spotlight)☆26Feb 22, 2026Updated 5 months ago
- ☆25Nov 3, 2025Updated 8 months ago
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 8 months ago