This code sets up a simple yet robust server using FastAPI for handling asynchronous requests for embedding generation and reranking tasks using the BAAI M3 multilingual model.
☆72May 8, 2024Updated 2 years ago
Alternatives and similar repositories for baai_m3_simple_server
Users that are interested in baai_m3_simple_server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- *high-load* benchmarking tool☆18Jul 17, 2026Updated last week
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆33Jan 23, 2025Updated last year
- allowing R users to work with dlib through Rcpp☆13Apr 11, 2018Updated 8 years ago
- CPython 파헤치기 스터디☆16Jul 13, 2024Updated 2 years ago
- Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpali☆2,892Mar 24, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- ☆13Nov 26, 2021Updated 4 years ago
- Hugging Face RoBERTa with Flash Attention 2☆24Sep 14, 2025Updated 10 months ago
- A blazing fast inference solution for text embeddings models☆4,959Updated this week
- A collection of models for TensorFlow Go☆12May 29, 2022Updated 4 years ago
- ☆48Jun 20, 2024Updated 2 years ago
- The Batched API provides a flexible and efficient way to process multiple requests in a batch, with a primary focus on dynamic batching o…☆161Jul 14, 2025Updated last year
- A context-aware embedding similarity score☆11Aug 23, 2023Updated 2 years ago
- TextEmbed is a REST API crafted for high-throughput and low-latency embedding inference. It accommodates a wide variety of embedding mode…☆28Sep 5, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- See how HTTPX, Requests, and AIOHTTP libraries compare for sending network requests and find out which one may fit your case better.☆22Sep 25, 2025Updated 10 months ago
- ☆10Feb 12, 2024Updated 2 years ago
- Few-shot text classification with meta learning and BERT☆11Jun 14, 2021Updated 5 years ago
- Deep & Cross Network in Tensorflow☆15May 30, 2018Updated 8 years ago
- Paystack payment gateway for Frappe/ERPNext☆22Oct 13, 2025Updated 9 months ago
- A curated list of outstanding, actively maintained vector search frameworks and engines, libraries, cloud services, and research papers f…☆18Mar 21, 2026Updated 4 months ago
- rerank library for easy reranking of results☆56Sep 17, 2024Updated last year
- Experimentation on google's gemma model☆15Mar 6, 2024Updated 2 years ago
- Using multiple LLMs for ensemble Forecasting☆16Jan 17, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Open-source Human Feedback Library☆11Oct 25, 2023Updated 2 years ago
- GVC decomposition in R☆20Jun 9, 2026Updated last month
- Train a SmolLM-style llm on fineweb-edu in JAX/Flax with an assortment of optimizers.☆19Jul 24, 2025Updated last year
- How to Build an AI Children’s Book Service☆30Nov 2, 2023Updated 2 years ago
- 텍스트 전처리 강의☆13Nov 7, 2019Updated 6 years ago
- Trigram tokenizer module for SQLite FTS5☆14Feb 22, 2021Updated 5 years ago
- The fastest FlashText library for Python☆26Jul 4, 2024Updated 2 years ago
- ☆11Mar 17, 2026Updated 4 months ago
- ☆21Mar 12, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- An experimental htmx extension for streaming contents using http streaming☆12Feb 23, 2024Updated 2 years ago
- Examples for QinYan GLMs☆13Sep 3, 2024Updated last year
- code for piccolo embedding model from SenseTime☆143May 21, 2024Updated 2 years ago
- Postgres-GPT combines pgvector and OpenAI to make a SQL-based knowledge repository from Markdown files that allows easy access to relevan…☆22Oct 26, 2023Updated 2 years ago
- Code for AUC Mu☆14Oct 17, 2019Updated 6 years ago
- Build Text Rerankers with Deep Language Models☆265Feb 20, 2024Updated 2 years ago
- Pure Go MPEG-1 Audio library☆11Nov 20, 2020Updated 5 years ago