Softened ROSA QKV Operators for Training Next-Generation LLM Models
☆39Aug 5, 2026Updated last week
Alternatives and similar repositories for rosa_soft
Users that are interested in rosa_soft are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 10 months ago
- Reinforcement Learning Toolkit for RWKV.(v6,v7,ARWKV) Distillation,SFT,RLHF(DPO,ORPO), infinite context training, Aligning. Exploring the…☆64Sep 19, 2025Updated 10 months ago
- A large-scale RWKV v7(World, PRWKV, Hybrid-RWKV) inference. Capable of inference by combining multiple states(Pseudo MoE). Easy to deploy…☆51Oct 21, 2025Updated 9 months ago
- Official Chinese documentation for RWKV | RWKV官方中文文档☆15Jun 10, 2026Updated 2 months ago
- ☆20Aug 1, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- State tuning tunes the state☆35Feb 12, 2025Updated last year
- 用户友好、开箱即用的 RWKV Prompts 示例,适用于所有用户。Awesome RWKV Prompts for general users, more user-friendly, ready-to-use prompt examples.☆34Apr 13, 2026Updated 4 months ago
- A high-efficiency text embedding and reranking model based on RWKV architecture.☆20Jan 10, 2026Updated 7 months ago
- ☆183Jan 13, 2026Updated 7 months ago
- The WorldRWKV project aims to implement training and inference across various modalities using the RWKV7 architecture. By leveraging diff…☆70Mar 18, 2026Updated 4 months ago
- Solving puzzles with RWKV locally in your browser.☆13Mar 31, 2026Updated 4 months ago
- Mini_RWKV_V7_LM Only 34.2M params (also have RWKV7s architecture [deep embedding]/[deep embedding attention) with Full Training code & da…☆90Jan 26, 2026Updated 6 months ago
- Efficient implementations of state-of-the-art linear attention models in Pytorch and Triton☆50Apr 2, 2026Updated 4 months ago
- ☆50Jul 3, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A specialized RWKV-7 model for Othello(a.k.a. Reversi) that predicts legal moves, evaluates positions, and performs in-context search. It…☆44Jan 25, 2025Updated last year
- [EMNLP 24] Source code for paper 'AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tu…☆13Dec 15, 2024Updated last year
- ☆12Dec 14, 2024Updated last year
- ☆12Aug 2, 2026Updated last week
- Course Project for COMP4471 on RWKV☆17Feb 11, 2024Updated 2 years ago
- ☆12Dec 21, 2024Updated last year
- RADLADS training code☆46May 7, 2025Updated last year
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated last year
- Accepted to ICLR 2025. MetaMetrics is a calibrated meta-metric designed to evaluate generation tasks across different modalities aligned …☆15Dec 30, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Efficient PScan implementation in PyTorch☆17Jan 2, 2024Updated 2 years ago
- ☆81May 15, 2024Updated 2 years ago
- This repo is an exploratory experiment to enable frozen pretrained RWKV language models to accept speech modality input. We followed the …☆54Dec 23, 2024Updated last year
- ☆151Nov 22, 2024Updated last year
- ☆20Mar 11, 2025Updated last year
- Flutter adapter and Dart FFI runtime bridge for on-device inference with rwkv-mobile.☆18Jul 23, 2026Updated 3 weeks ago
- Inference RWKV with multiple supported backends.☆95Jul 17, 2026Updated 3 weeks ago
- RWKV infctx trainer, for training arbitary context sizes, to 10k and beyond!☆149Aug 13, 2024Updated 2 years ago
- 更简单的微调,提供便捷脚本,微调说明☆35May 30, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Finetune Nemo parakeet ASR model with new language (support 8 bit optimizer). Experimental birwkv-fastconformer TDT for long-form ASR(8.5…☆26Nov 27, 2025Updated 8 months ago
- Adaptation of titans-pytorch to llama models on HF☆24Mar 6, 2025Updated last year
- This project is to train an RWKV LLM for TTS generation which compatible to other TTS engine(like fish/cosy/chattts).☆101Oct 8, 2025Updated 10 months ago
- ☆27Feb 26, 2026Updated 5 months ago
- Model Doctor: A Simple Gradient Aggregation Strategy for Diagnosing and Treating CNN Classifiers [https://arxiv.org/pdf/2112.04934.pdf]☆15May 13, 2023Updated 3 years ago
- Helper Tool for Card Preview Modding in Clash Royale☆17Nov 3, 2025Updated 9 months ago
- Tools to isolate speaker and transcribe unstructured audio clips☆11Dec 4, 2022Updated 3 years ago