Long Context Pre-Training with Lighthouse Attention
☆68Jul 31, 2026Updated 3 weeks ago
Alternatives and similar repositories for lighthouse-attention
Users that are interested in lighthouse-attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 3 months ago
- Data and code for ACL 2026 Paper "Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems…☆19Apr 30, 2026Updated 3 months ago
- Website using FinBERT + live financial news scraping to assess short-term investment potential.☆13Mar 7, 2025Updated last year
- Official PyTorch Implementation of Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention☆269May 25, 2026Updated 2 months ago
- ArxivDaily☆13Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- SKT A.X LLM 3.1☆13Jul 24, 2025Updated last year
- Conformer block with Rotary Position Embedding, modified from lucidrains' implement☆19Sep 13, 2024Updated last year
- Code & Data for our Paper "RobustGEC: Robust Grammatical Error Correction Against Subtle Context Perturbation" (EMNLP 2023)☆17Jan 23, 2024Updated 2 years ago
- Data pipeline for HRM-Text pretraining☆70May 21, 2026Updated 3 months ago
- ☆14Nov 18, 2025Updated 9 months ago
- Implementation of 'Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis', in MLX☆24Oct 30, 2024Updated last year
- Official Implementation for NorMuon paper☆86Apr 30, 2026Updated 3 months ago
- Analyzing partial dimensional collapse in non-contrastive self-supervised learning. "Understanding Collapse in Non-Contrastive Siamese Re…☆16Nov 12, 2023Updated 2 years ago
- Don't just regulate gradients like in Muon, regulate the weights too☆32Jul 30, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation☆298Feb 18, 2026Updated 6 months ago
- Compositional Muon release☆25Jun 5, 2026Updated 2 months ago
- Delta Attention Residuals - supplementary code and pretrained models☆43May 20, 2026Updated 3 months ago
- TACS: Taxonomy Adaptive Cross-Domain Semantic Segmentation☆12Jul 14, 2022Updated 4 years ago
- Fork of github.com/UCSBarchlab/OpenTPU for the TGPTPU project☆15Jun 1, 2025Updated last year
- ☆15Aug 26, 2024Updated last year
- ☆21Aug 26, 2025Updated 11 months ago
- ☆14Apr 7, 2025Updated last year
- Mixture of Experts from scratch☆14Apr 12, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- FORMED: Foundation Model Repuposed for Medical Time Series Classification☆17Nov 14, 2025Updated 9 months ago
- Go implementation of the Gun distributed graph database☆11Feb 26, 2019Updated 7 years ago
- Voice synthesis library for Text-to-Speech applications (HTS Engine rewrite in Rust language)☆13Updated this week
- CUDA_C编程权威指南示例代码☆13Mar 22, 2023Updated 3 years ago
- HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness☆130May 9, 2026Updated 3 months ago
- Explorations into the proposed SDFT, Self-Distillation Enables Continual Learning, from Shenfeld et al. of MIT☆32Feb 6, 2026Updated 6 months ago
- Code for NeurIPS 2024 Paper - Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass☆21Aug 22, 2024Updated 2 years ago
- JAX support for tvm-ffi abi☆26May 14, 2026Updated 3 months ago
- COS-PLAY: Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play☆30Jul 11, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.☆257Jun 29, 2026Updated last month
- ☆25May 23, 2026Updated 3 months ago
- The open source implementation of the multi grouped query attention by the paper "GQA: Training Generalized Multi-Query Transformer Model…☆17Dec 11, 2023Updated 2 years ago
- A fast, memory-efficient exact MaxSim kernel for late-interaction retrieval and reranking.☆26May 27, 2026Updated 2 months ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆40Jul 20, 2026Updated last month
- parallel LSTM from Were RNNs All We Needed?☆30Oct 11, 2024Updated last year
- Collection of resources for RL and Reasoning☆27Feb 3, 2025Updated last year