☆16Apr 7, 2024Updated 2 years ago
Alternatives and similar repositories for GemmaLongText
Users that are interested in GemmaLongText are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Jun 24, 2024Updated 2 years ago
- LLM checkpointing for DeepSpeed/Megatron☆27Nov 30, 2025Updated 9 months ago
- ☆11Jun 1, 2023Updated 3 years ago
- ☆18Sep 22, 2024Updated 2 years ago
- 持续追踪ChatGPT相关的技术资料和行业进展。☆12Apr 24, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ppo+action mask for atari tennis agent☆12Mar 2, 2023Updated 3 years ago
- [ACL'24 Oral] Analysing The Impact of Sequence Composition on Language Model Pre-Training☆24Aug 18, 2024Updated 2 years ago
- ☆14Jul 11, 2024Updated 2 years ago
- [ICCV 2025] Dynamic-VLM☆28Dec 16, 2024Updated last year
- ProxyExplainer for Graph Neural Networks☆16Oct 24, 2024Updated last year
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- Create your map such as (Google map or Here Map or Mapbox or .....)☆20Mar 24, 2021Updated 5 years ago
- ☆12Updated this week
- ☆13May 9, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆17Jul 10, 2023Updated 3 years ago
- Official code for the paper "HEXA-MoE: Efficient and Heterogeneous-Aware MoE Acceleration with Zero Computation Redundancy"☆15Mar 6, 2025Updated last year
- XGEN-MM(BLIP3) Autocaptioning Tools☆17Jun 20, 2024Updated 2 years ago
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination.☆22Jul 18, 2025Updated last year
- code for ACL24 "MELoRA: Mini-Ensemble Low-Rank Adapter for Parameter-Efficient Fine-Tuning"☆33Feb 19, 2025Updated last year
- Quack is a free and open-source chat application designed for private use. Although it doesn't have any unique features, it combines the …☆47Jul 30, 2026Updated last month
- ☆12Feb 12, 2026Updated 7 months ago
- Unofficial implementation of paper "InstructionNER: A Multi-Task Instruction-Based Generative Framework for Few-shot NER" (https://arxiv.…☆38Feb 14, 2024Updated 2 years ago
- 石蒜摇摇乐vscode插件☆13Aug 31, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs☆16Apr 18, 2025Updated last year
- Reproduction of the complete process of DeepSeek-R1 on small-scale models, including Pre-training, SFT, and RL.☆33Mar 11, 2025Updated last year
- B站API的Golang版本,提供视频源解析,排行获取等常用接口☆12Jan 19, 2016Updated 10 years ago
- [EMNLP 2023] TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding☆50Jan 9, 2024Updated 2 years ago
- Sogou RPC benchmark base on Sogou C++ Workflow☆10Sep 4, 2020Updated 6 years ago
- Implementation of PALI3 from the paper PALI-3 VISION LANGUAGE MODELS: SMALLER, FASTER, STRONGER"☆146Aug 28, 2026Updated 3 weeks ago
- Measuring memory usage in C and C++☆29Nov 3, 2016Updated 9 years ago
- create verify code using canvas☆17Sep 20, 2021Updated 5 years ago
- Implementation of Alexander A. Stepanov inverted Index Compression algorithms☆21Nov 10, 2015Updated 10 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆24Aug 17, 2024Updated 2 years ago
- ☆20May 5, 2024Updated 2 years ago
- Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models☆68Nov 1, 2024Updated last year
- resizable hashing strategy for large-scale storage☆25Oct 6, 2019Updated 6 years ago
- [ICME2024, Official Code] for paper "Bringing Textual Prompt to AI-Generated Image Quality Assessment"☆21Jul 9, 2024Updated 2 years ago
- ☆42Nov 20, 2025Updated 10 months ago
- USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference☆695Sep 16, 2026Updated last week