A high-throughput and memory-efficient inference and serving engine for LLMs
☆15Jan 22, 2025Updated last year
Alternatives and similar repositories for vllm
Users that are interested in vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A fully cuda implementation of DCNv2(deformable convolution) forward. Without dependent of cuTorch(THC).☆10Dec 9, 2019Updated 6 years ago
- A multimodal large-scale model, which performs close to the closed-source Qwen-VL-PLUS on many datasets and significantly surpasses the p…☆14Feb 5, 2024Updated 2 years ago
- (My internship selection task at LearnOpenCV | Big Vision LLC) OpenCV based dimensional measurement of a book cover using Homography and …☆13Mar 2, 2018Updated 8 years ago
- Official Implementation of APB (ACL 2025 main Oral) and Spava (ACL 2026 main).☆37Apr 6, 2026Updated 5 months ago
- HCCR中文手写汉字识别 (网站在线实时推理)☆13Feb 2, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆14May 6, 2024Updated 2 years ago
- 生成用于训练CRNN的图片数据☆20Apr 13, 2018Updated 8 years ago
- 链家网深圳所有租房信息爬取☆12Feb 7, 2017Updated 9 years ago
- darknet2onnx2tensorrt for yolov3-tiny on nvidia tx2☆15Apr 29, 2019Updated 7 years ago
- The code for reproducing "Frame Difference-Based Temporal Loss"☆11Sep 11, 2021Updated 5 years ago
- ☆19Updated this week
- Table Structure Recognition☆28Jul 25, 2024Updated 2 years ago
- ☆11Oct 25, 2022Updated 3 years ago
- A hobby project that dewarps book pages in images☆19Jan 5, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICME 2021 Oral] Official implementation for "FGF-GAN: A Lightweight Generative Adversarial Network for Pansharpening via Fast Guided Fil…☆11Mar 29, 2022Updated 4 years ago
- 【ICCV 2023】Towards Instance-adaptive Inference for Federated Learning☆12Mar 31, 2025Updated last year
- apollo r3.0 感知移植(参考百度开发者套件)☆19Oct 29, 2019Updated 6 years ago
- Implementation of the paper : Not all attention is needed - Gated Attention Network for Sequence Data (GA-Net) [https://arxiv.org/abs/191…☆13Aug 20, 2020Updated 6 years ago
- Are Intermediate Layers and Labels Really Necessary? A General Language Model Distillation Method ; GKD: A General Knowledge Distillation…☆34Aug 4, 2023Updated 3 years ago
- Scan your documents with this simple Python script.☆13May 14, 2017Updated 9 years ago
- Implementation of the paper "Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network"☆16Nov 1, 2019Updated 6 years ago
- 利用 tesseract 解析简单数字验证码图片☆22May 25, 2018Updated 8 years ago
- Clustering with Orthogonal AutoEncoder☆15May 6, 2019Updated 7 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- SRD: A Tree Structure Based Decoder for Online Handwritten Mathematical Expression Recognition☆21Jul 20, 2020Updated 6 years ago
- Implementation of 'Paint With Words' method in eDiff-I☆27Jan 31, 2023Updated 3 years ago
- A polygon detector based on obb-yolov3 (WIP)☆17Jul 21, 2021Updated 5 years ago
- Official code for the paper "FairerCLIP: Debiasing CLIP’s Zero-Shot Predictions using Functions in RKHSs".☆16Oct 14, 2025Updated 11 months ago
- [EMNLP'23] Code for Generating Data for Symbolic Language with Large Language Models☆18Oct 21, 2023Updated 2 years ago
- ☆22Feb 14, 2019Updated 7 years ago
- ☆19Jun 25, 2019Updated 7 years ago
- [TIM 2025] Towards Accurate Readings of Water Meters by Eliminating Transition Error: New Dataset and Effective Solution☆20Mar 5, 2025Updated last year
- Source code of our MM24 paper "Harmfully Manipulated Images Matter in Multimodal Misinformation Detection"☆19Aug 10, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- 根据全能扫描王的文档扫描实现简单的图片扫描功能☆20Nov 4, 2019Updated 6 years ago
- Code for "Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations"☆22Nov 16, 2022Updated 3 years ago
- ☆48Feb 7, 2025Updated last year
- An implementation for MLLM oversensitivity evaluation☆18Nov 16, 2024Updated last year
- (ACL 2025) MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale☆50Jun 4, 2025Updated last year
- 这是一个用于计算ViT及其变种模型的GradCAM自动脚本,可以自动处理批量的图像 A GradCAM automatic script to visualize the model result☆18Dec 16, 2024Updated last year
- A Survey on Interpretable Cross-modal Reasoning☆14Oct 12, 2023Updated 2 years ago