FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)
☆18Jul 1, 2026Updated 2 months ago
Alternatives and similar repositories for flashinfer-windows
Users that are interested in flashinfer-windows are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pynini for Windows☆26Apr 11, 2025Updated last year
- Windows-native adaptation fork of Hermes Agent, based on upstream Hermes Agent 0.13.0. Improves local runtime environment, path handling,…☆20May 15, 2026Updated 4 months ago
- Memo's Blog☆27Jan 21, 2026Updated 7 months ago
- Controllable Language Model Interactions in TypeScript☆10May 17, 2024Updated 2 years ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code☆10Mar 8, 2022Updated 4 years ago
- Official source code repository for paper BubbleRAG.☆17Jun 1, 2026Updated 3 months ago
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 5 months ago
- ☆18Jun 6, 2026Updated 3 months ago
- [ICLR'26] Official code for "Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Trai…☆16Mar 31, 2026Updated 5 months ago
- A single dynamic MCP server that turns JSON configs into working API tools☆23Jul 30, 2026Updated last month
- ☆18Dec 1, 2025Updated 9 months ago
- An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io☆16Apr 18, 2024Updated 2 years ago
- FastWave is a lightweight diffusion model for general audio super-resolution (any -> 48 kHz). SOTA quality reconstruction metrics with ju…☆20May 16, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- tf-openpose and unity IK☆10Jul 2, 2020Updated 6 years ago
- LatentMAS with kNN kv cache pruning | up to 40% more memory efficient and 30% faster☆19Dec 10, 2025Updated 9 months ago
- ☆14Nov 22, 2022Updated 3 years ago
- 基于有赞云开放平台sdk源码迭代的api库☆11Feb 17, 2024Updated 2 years ago
- ☆27Jul 13, 2026Updated 2 months ago
- Reorganizes Booru Datasets from Gwern to be valid for DeepDanbooru☆12Aug 5, 2021Updated 5 years ago
- [TrustNLP@NAACL 2025] BiasEdit: Debiasing Stereotyped Language Models via Model Editing☆18Sep 30, 2025Updated 11 months ago
- RAG based agent with chDB(ClickHouse)☆23May 14, 2025Updated last year
- Semantic Web Memory for Intelligent Agents☆19Jul 12, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is a fork of the https://dockbox.dev project I made☆172Updated this week
- Business Data Benchmark (BDB) is a set of real-world questions to evaluate AI systems connected to business data.☆25Dec 3, 2024Updated last year
- Kimi K2 Thinking Agentic Search Unofficial Implementation☆15Nov 9, 2025Updated 10 months ago
- Surgically de-slop LLMs☆15Jun 1, 2025Updated last year
- ☆14Feb 18, 2025Updated last year
- 微信加群活码☆12Aug 12, 2019Updated 7 years ago
- SR²AM: Efficient Agentic Reasoning Through Self-Regulated Simulative Planning☆21May 22, 2026Updated 3 months ago
- The source code for the paper CrossSinger (asru2023)☆18Oct 12, 2023Updated 2 years ago
- Run Docker in Android (via Termux)☆18Jan 21, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code corresponding to Generative Semantic Workspaces - Long term Structured Memory for Large Language Models - AAAI 26 (Oral), ICML 26☆26Aug 1, 2026Updated last month
- The official implementation of LIFT: Improving Long Context Understanding of Large Language Models through Long Input Fine-Tuning☆15Mar 14, 2025Updated last year
- Automatic1111 to InvokeAI prompt resolver☆17Jun 29, 2024Updated 2 years ago
- Helper for FFmpeg allowing multi-threaded downloading of M3U8 video streams☆15Feb 20, 2022Updated 4 years ago
- [NeurIPS 2025] A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings☆28Sep 9, 2026Updated last week
- An official source code of Choi et al., High-quality Frame Interpolation via Tridirectional Inference, WACV 2021 paper.☆16Mar 26, 2021Updated 5 years ago
- Official Code Repositiry for "RaDeR: Reasoning-aware Dense Retrieval Models" accepted at Main Conference EMNLP 2025☆18Jun 23, 2025Updated last year