π Sliding Window Attention Training for Efficient Large Language Models
β20Jun 7, 2026Updated 2 months ago
Alternatives and similar repositories for swat-attention
Users that are interested in swat-attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding.β13Nov 19, 2024Updated last year
- This respository is used for time reasoning task for mult-session dialogue system.β18Feb 7, 2026Updated 6 months ago
- Codebase for Instruction Following without Instruction Tuningβ36Sep 24, 2024Updated last year
- We introduce EMMET and unify model editing with popular algorithms ROME and MEMIT.β29Dec 16, 2024Updated last year
- β17Nov 20, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of "MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model". Our coβ¦β26Dec 20, 2024Updated last year
- Official resource for paper Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models (ACL 20β¦β18Aug 12, 2024Updated 2 years ago
- Tensor Parallelism with JAX + Shard Mapβ11Sep 29, 2023Updated 2 years ago
- β18Mar 16, 2026Updated 5 months ago
- β18Oct 12, 2025Updated 10 months ago
- The code implementation of Skill-MoEβ47May 22, 2026Updated 3 months ago
- β12Sep 18, 2024Updated last year
- Accurate and fast KV cache compression with a gating mechanismβ29Jul 27, 2026Updated last month
- Collection of useful kernelsβ20Oct 3, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- π₯This is a repository of paper list for streaming LLMs/MLLMs.β27Apr 19, 2026Updated 4 months ago
- Machine Learning based toxicity prediction tool for small molecules.β11Feb 13, 2024Updated 2 years ago
- [ACL 2024 Findings] Learning Fine-Grained Grounded Citations for Attributed Large Language Modelsβ20Oct 24, 2024Updated last year
- β16Mar 17, 2025Updated last year
- [EMNLP 2025 Oral] IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agentsβ22Sep 16, 2025Updated 11 months ago
- Tools for exploring Transformer neuron behaviour, including input pruning and diversification.β11Jun 6, 2023Updated 3 years ago
- JAX for Graphcore IPU (experimental)β21Mar 12, 2024Updated 2 years ago
- [ECCV 2024] The official PyTorch implementation of the "Plain-Det: A Plain Multi-Dataset Object Detector".β30Dec 8, 2024Updated last year
- Official implementation of paper VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interactβ¦β45Feb 5, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Reinforcement Learning Projectβ12Jan 16, 2017Updated 9 years ago
- a transformer implemented primarily using einops and trained on the tinystories datasetβ14Jun 21, 2024Updated 2 years ago
- A cycle-accurate RISC-V CPU simulator + RTL modeling library in pure Python.β19Aug 14, 2026Updated 2 weeks ago
- Benchmarking LLMs in Real-World Memory-Driven Interactionβ47Apr 7, 2026Updated 4 months ago
- Turbo Lossless - 1.33x Smaller, 2.93x Faster, Decode with 1 ADD operationβ15Apr 4, 2026Updated 4 months ago
- Contains from-scratch implementation of the MobileNetV1, V2 and V3 paper with PyTorch. Each model architecture is contained in a single fβ¦β19Aug 10, 2022Updated 4 years ago
- (Verilog) A simple convolution layer implementation with systolic array structureβ14May 9, 2022Updated 4 years ago
- Python library for analyzing data quality and its impact on model performance across classification and object-detection tasks.β18Updated this week
- β53Apr 9, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of paper 'Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference' [NeurIPS'24β¦β26Jun 14, 2024Updated 2 years ago
- Text2Mem: A Unified Memory Operation Language for Memory Operating Systemβ58Jan 7, 2026Updated 7 months ago
- Source code of the IPDPS '21 paper: "TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs" by Yuyao Niu, Zhengyangβ¦β13Aug 12, 2022Updated 4 years ago
- β10Jul 4, 2022Updated 4 years ago
- Official InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Showsβ20Nov 4, 2025Updated 9 months ago
- [ACL 2025] Knowledge Unlearning for Large Language Modelsβ49Sep 18, 2025Updated 11 months ago
- β12Nov 24, 2023Updated 2 years ago