gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。
☆253May 9, 2026Updated 2 months ago
Alternatives and similar repositories for gpt_server
Users that are interested in gpt_server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- OpenAI Router 轻量级、持久化、零配置的 OpenAI API 统一网关☆25Updated this week
- Evaluation for AI apps and agent☆46Jan 18, 2024Updated 2 years ago
- Openai style api for open large language models, using LLMs just as chatgpt! Support for LLaMA, LLaMA-2, BLOOM, Falcon, Baichuan, Qwen, X…☆2,459Sep 26, 2024Updated last year
- code for piccolo embedding model from SenseTime☆143May 21, 2024Updated 2 years ago
- Unify Efficient Fine-tuning of RAG Retrieval, including Embedding, ColBERT, ReRanker.☆1,126Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- 通义千问VLLM推理部署DEMO☆643Mar 28, 2024Updated 2 years ago
- Accelerating GOT-OCRv2 with VLLM☆10Nov 15, 2024Updated last year
- Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-p…☆9,442Updated this week
- TextEmbed is a REST API crafted for high-throughput and low-latency embedding inference. It accommodates a wide variety of embedding mode…☆28Sep 5, 2024Updated last year
- aigc evals☆10Dec 2, 2023Updated 2 years ago
- Streaming ASR and TTS based on FastAPI+ sherpa-onnx☆221Nov 2, 2025Updated 8 months ago
- LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA☆520Dec 31, 2024Updated last year
- 基于Funasr的[实时]AI语音助手☆25Dec 18, 2025Updated 7 months ago
- pure go for rwkv☆18Dec 31, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- unified embedding model☆876Sep 1, 2023Updated 2 years ago
- Netease Youdao's open-source embedding and reranker models for RAG products.☆1,881Sep 9, 2025Updated 10 months ago
- 用于大模型 RLHF 进行人工数据标注排序的工具。A tool for manual response data annotation sorting in RLHF stage.☆254Aug 1, 2023Updated 2 years ago
- this repo is mnbvc text quality classification using fastText☆16Oct 2, 2023Updated 2 years ago
- 360LayoutAnaylsis, a series Document Analysis Models and Datasets deleveped by 360 AI Research Institute☆305Sep 10, 2024Updated last year
- A simple wrapper around "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching" that provides an OpenAI-compatibl…☆14Feb 7, 2025Updated last year
- accelerate generating vector by using onnx model☆18Jan 23, 2024Updated 2 years ago
- Interpretable Word Sense Representations via Definition Generation☆10Mar 6, 2025Updated last year
- LMDeploy is a toolkit for compressing, deploying, and serving LLMs.☆7,970Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.☆3,123Updated this week
- A Scheduler for Batched LLM Inference☆19Oct 5, 2025Updated 9 months ago
- Imitate OpenAI with Local Models☆91Aug 27, 2024Updated last year
- ☆27Jul 18, 2023Updated 3 years ago
- ☆14Mar 7, 2025Updated last year
- Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpali☆2,889Mar 24, 2026Updated 3 months ago
- ☆61Jan 3, 2025Updated last year
- Retrieval and Retrieval-augmented LLMs☆11,968Apr 22, 2026Updated 3 months ago
- [ACL 2025] AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark☆167Mar 29, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- FastAPI Server Implementation for Bilibili Index TTS☆25Apr 13, 2025Updated last year
- 文本去重☆77May 23, 2024Updated 2 years ago
- ☆16Jul 29, 2025Updated 11 months ago
- Firefly: 大模型训练工具,支持训练Qwen2.5、Qwen2、Yi1.5、Phi-3、Llama3、Gemma、MiniCPM、Yi、Deepseek、Orion、Xverse、Mixtral-8x7B、Zephyr、Mistral、Baichuan2、Llma2、…☆6,646Oct 24, 2024Updated last year
- 面向金融领域的小样本跨类迁移事件抽取 第三名 方案及代码☆16Dec 23, 2020Updated 5 years ago
- KnowFlowRAG☆516May 27, 2026Updated last month
- 通用向量搜索服务☆32Mar 21, 2022Updated 4 years ago