Lab 5 project of MIT-6.5940, deploying LLaMA2-7B-chat on one's laptop with TinyChatEngine.
☆18Dec 1, 2023Updated 2 years ago
Alternatives and similar repositories for LLaMA2-7B-on-laptop
Users that are interested in LLaMA2-7B-on-laptop are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 模型加速/模型压缩(已完成所有Lab)☆11Dec 24, 2023Updated 2 years ago
- Research in compressing convolutional layers of CNN using low-rank Tucker tensor decomposition☆12Nov 1, 2023Updated 2 years ago
- The official implementation of the paper "Affective Faces for Goal-Driven Dyadic Communication."☆15Jan 27, 2023Updated 3 years ago
- Structured Pruning Adapters in PyTorch☆19Aug 30, 2023Updated 2 years ago
- MOFY: MOsaic For You 실시간 불특정 인물 비식별화☆14Jun 22, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- channel pruning for accelerating very deep neural networks☆13Mar 8, 2021Updated 5 years ago
- ☆178Aug 9, 2023Updated 2 years ago
- ☆10Feb 7, 2022Updated 4 years ago
- In this project, we provide a strong template for a PyTorch project. The purpose of this repository is to provide an example (and strong …☆15Apr 21, 2021Updated 5 years ago
- ☆13Jun 12, 2025Updated last year
- An implementation of Distortion-Free Wide-Angle Portraits on Camera Phones☆10Dec 24, 2019Updated 6 years ago
- Source code of our TNNLS paper "Boosting Convolutional Neural Networks with Middle Spectrum Grouped Convolution"☆12Apr 14, 2023Updated 3 years ago
- Rust bindings for SPDK☆12Mar 5, 2020Updated 6 years ago
- ☆79Nov 5, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [NeurIPS 2024] Search for Efficient LLMs☆16Jan 16, 2025Updated last year
- a port forwarding tool similar to lcx☆10Mar 14, 2019Updated 7 years ago
- TinyML and Efficient Deep Learning Computing☆20Apr 26, 2024Updated 2 years ago
- Minimal implementation of Denoised Smoothing (https://arxiv.org/abs/2003.01908) in TensorFlow.☆20Aug 4, 2021Updated 4 years ago
- A Valgrind extension for CUDA, unofficial mirror for https://www.hlrs.de/organization/av/spmt/research/cudagrind/☆10Aug 5, 2015Updated 10 years ago
- ☆15Feb 1, 2016Updated 10 years ago
- This is a repository of coursework project for the Stanford Compilers MOOC course. The result is a fully-working compiler for the COOL Pr…☆18Sep 11, 2023Updated 2 years ago
- Implementation of the paper : Not all attention is needed - Gated Attention Network for Sequence Data (GA-Net) [https://arxiv.org/abs/191…☆13Aug 20, 2020Updated 5 years ago
- [ ICCV'2025 Poster ] SecDOOD is a secure cloud-device collaboration framework for efficient on-device OOD detection without requiring de…☆15Oct 10, 2025Updated 9 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [부스트캠프 AI Tech 3기 / CV-10] 인물 기반 예능 숏폼 영상 생성기, #눈#사람 ⛄️ (22.05.16 - 22.06.15)☆16Apr 14, 2023Updated 3 years ago
- 南京大学软件分析作业☆14Jul 30, 2022Updated 3 years ago
- ☆19Apr 11, 2024Updated 2 years ago
- ☆10May 25, 2017Updated 9 years ago
- [SP 2024] A Novel Recursive Least-Squares Adaptive Method For Streaming Tensor-Train Decomposition With Incomplete Observations. In Elsev…☆15Jan 2, 2024Updated 2 years ago
- ☆10Feb 17, 2022Updated 4 years ago
- This repo contains the Assignments from Cornell Tech's ECE 5545 - Machine Learning Hardware and Systems offered in Spring 2023☆44May 31, 2023Updated 3 years ago
- Computer Vision Fall 2019 by Chiu-San Fu@ CSIE NTU Taiwan☆10Jan 11, 2020Updated 6 years ago
- [ICLR 2023] PyTorch code for DFPC: Data flow driven pruning of coupled channels without data.☆15Aug 25, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆10Nov 14, 2023Updated 2 years ago
- Asynchronous Rust bindings for SPDK.☆18Nov 1, 2022Updated 3 years ago
- code for the paper "A Statistical Framework for Low-bitwidth Training of Deep Neural Networks"☆29Oct 31, 2020Updated 5 years ago
- CUDA_C编程权威指南示例代码☆13Mar 22, 2023Updated 3 years ago
- A docker image for One Student One Chip's debug exam☆10Sep 22, 2023Updated 2 years ago
- ☆48Nov 1, 2025Updated 8 months ago
- Sirius, an efficient correction mechanism, which significantly boosts Contextual Sparsity models on reasoning tasks while maintaining its…☆21Sep 10, 2024Updated last year