Cortex.Tensorrt-LLM is a C++ inference library that can be loaded by any server at runtime. It submodules NVIDIA’s TensorRT-LLM for GPU accelerated inference on NVIDIA's GPUs.
☆41Sep 26, 2024Updated last year
Alternatives and similar repositories for cortex.tensorrt-llm
Users that are interested in cortex.tensorrt-llm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- cortex.llamacpp is a high-efficiency C++ inference engine for edge computing. It is a dynamic library that can be loaded by any server a…☆44Jul 4, 2025Updated last year
- ☆22Mar 25, 2025Updated last year
- A small utility library for parsing GGUF file info☆31Jan 27, 2025Updated last year
- Local AI API Platform☆2,751Jul 4, 2025Updated last year
- Attempt at cog wrapper for nightmareai/real-esrgan for larger images☆16Sep 28, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Inference Llama/Llama2/Llama3 Modes in NumPy☆21Nov 22, 2023Updated 2 years ago
- ☆12Nov 8, 2023Updated 2 years ago
- A Deeplearn Model to rec table in photo with ncnn. 一个深度学习模型用于检测图片中的表格 画像内のテーブルを検出するためのディープラーニング モデル☆20Mar 2, 2025Updated last year
- ☆140Apr 23, 2024Updated 2 years ago
- 🌐 OpenCrawl: An ethical, high-performance web crawler built for scale A powerful web crawling library that respects robots.txt and rate…☆28Apr 3, 2025Updated last year
- Attempt at cog wrapper for segmind/SSD-1B☆10Dec 11, 2023Updated 2 years ago
- Run AuraFlow on Replicate☆14Jul 12, 2024Updated 2 years ago
- An open-source framework designed to extend and maintain the ESCO taxonomy☆18Mar 12, 2026Updated 6 months ago
- Examples using MLX Swift☆13Apr 9, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Scaling is a distributed training library and installable dependency designed to scale up neural networks, with a dedicated module for tr…☆66Nov 18, 2025Updated 10 months ago
- fast-embeddings-api☆16Nov 23, 2023Updated 2 years ago
- Generative AI powered Note taking mobile app converts your voice recording or other audio file into short notes and customized questions …☆18Dec 8, 2023Updated 2 years ago
- Real-world Conversational AI personas.☆23Oct 23, 2023Updated 2 years ago
- A Discord bot that answers questions about Replicate.☆16Jan 5, 2024Updated 2 years ago
- SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution☆14Jan 12, 2024Updated 2 years ago
- ☆29May 27, 2026Updated 3 months ago
- Web accessible Inkscape inside an Alpine Container☆19Updated this week
- Moved to here: https://github.com/lyogavin/airllm☆37Aug 1, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Safely push a Cog model version by making sure it works and is backwards-compatible with previous versions.☆17Dec 4, 2025Updated 9 months ago
- Converting Chinese sentences into pinyin sequences, implemented in C++, very fast and easy to deploy.☆23Jan 5, 2026Updated 8 months ago
- A sensible approach to writing common code for react and react-native☆16Jan 8, 2019Updated 7 years ago
- An Android app for real-time facial emotion recognition, designed to improve accuracy for Middle Eastern faces and women wearing hijabs. …☆21Sep 11, 2023Updated 3 years ago
- Trying to build an all in one speech-text language model - a bit like GPT-4o☆22Jun 1, 2024Updated 2 years ago
- Cog wrapper for FalconsAi / nsfw_image_detection☆19Aug 6, 2025Updated last year
- 33B Chinese LLM, DPO QLORA, 100K context, AirLLM 70B inference with single 4GB GPU☆14May 5, 2024Updated 2 years ago
- Browser extensions for the Knowledge application☆33Jul 16, 2022Updated 4 years ago
- Cog implementation of the ByteDance/Hyper-SD Flux.1-Dev 8-step LoRA☆16Mar 18, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for paper: "QuIP: 2-Bit Quantization of Large Language Models With Guarantees" adapted for Llama models☆40Aug 4, 2023Updated 3 years ago
- Improve Devcontainer Creation☆17Sep 3, 2026Updated 2 weeks ago
- ☆21Mar 3, 2025Updated last year
- ☆62Nov 21, 2024Updated last year
- Cog wrapper for collabora/WhisperSpeech☆25Mar 5, 2024Updated 2 years ago
- ☆18Mar 19, 2023Updated 3 years ago
- Algorithm study using python day by day☆13Apr 9, 2017Updated 9 years ago