This repository contains an implementation of the LLaMA 2 (Large Language Model Meta AI) model, a Generative Pretrained Transformer (GPT) variant. The implementation focuses on the model architecture and the inference process. The code is restructured and heavily commented to facilitate easy understanding of the key parts of the architecture.
☆75Oct 1, 2023Updated 2 years ago
Alternatives and similar repositories for LLaMA2
Users that are interested in LLaMA2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Inference Llama 2 in one file of pure Haskell (A port of llama2.c from Andrej Karpathy)☆14Oct 17, 2025Updated 9 months ago
- PyTorch Quantization Framework For OCP MX Datatypes.☆16May 30, 2025Updated last year
- [ACL'26 Workshop] KoViDoRe: Korean Visual Document Retrieval Benchmark☆24Jul 2, 2026Updated 3 weeks ago
- ☆11Oct 11, 2023Updated 2 years ago
- My defense presentation☆10Mar 7, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆11Feb 3, 2025Updated last year
- An implementation of the base GPT-3 Model architecture from the paper by OPENAI "Language Models are Few-Shot Learners"☆22Jun 29, 2024Updated 2 years ago
- Wrapper to easily generate the chat template for Llama2☆65Mar 10, 2024Updated 2 years ago
- Scaling Sparse Fine-Tuning to Large Language Models☆19Jan 31, 2024Updated 2 years ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- Code for data reduction and analysis of Galaxy Zoo 2☆14May 20, 2016Updated 10 years ago
- LLaMA 2 implemented from scratch in PyTorch☆375Sep 25, 2023Updated 2 years ago
- The official code and dataset for EMNLP 2022 paper "COPEN: Probing Conceptual Knowledge in Pre-trained Language Models".☆21Mar 9, 2023Updated 3 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for "Automatic Circuit Finding and Faithfulness"☆19Jul 11, 2024Updated 2 years ago
- Ultra-minimal autoregressive diffusion model for image generation☆21Jul 20, 2026Updated last week
- Norwegian question answering dataset☆15Feb 3, 2024Updated 2 years ago
- D.Com 학우들을 위한 커리어 조언 Repo☆12May 17, 2023Updated 3 years ago
- ☆38Jun 2, 2026Updated last month
- Kanban board made with TailwindCSS☆11Jun 10, 2021Updated 5 years ago
- Deepseek-CoT☆10Oct 6, 2024Updated last year
- 🚀 [ICLR '25] RocketEval: Efficient Automated LLM Evaluation via Grading Checklist☆17Aug 21, 2025Updated 11 months ago
- ☆16Jan 14, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Unofficial reimplementation of ViR: Vision Retention Networks by Hatamizadeh et. al. (https://arxiv.org/abs/2310.19731)☆18Jul 26, 2024Updated 2 years ago
- Creates CMM script that can directly executed on Kaggle from easy merge script☆14Mar 6, 2026Updated 4 months ago
- Template repo for Python projects, especially those focusing on machine learning and/or deep learning.☆15Jan 14, 2026Updated 6 months ago
- Official Implementation of paper: [Nav-R2:Dual‑Relation Reasoning for Generalizable Open‑Vocabulary Object‑Goal Navigation]☆20Dec 10, 2025Updated 7 months ago
- [ICLR 2024] Unveiling the Pitfalls of Knowledge Editing for Large Language Models☆22Jun 13, 2024Updated 2 years ago
- Code for the AAAI 2024 Oral paper "OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Model…☆72Mar 7, 2024Updated 2 years ago
- Multi-LexSum is an abstractive summarization dataset for US Civil Rights Lawsuits☆23Dec 15, 2022Updated 3 years ago
- Code for Personalized Large Language Models via Selective Prompt Tuning☆10Jun 26, 2024Updated 2 years ago
- Follow-up Studies on STEGO☆13Sep 27, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Repo accompanying our paper "Do Llamas Work in English? On the Latent Language of Multilingual Transformers".☆87Mar 11, 2024Updated 2 years ago
- a simplified version of Meta's Llama 3 model to be used for learning☆44May 21, 2024Updated 2 years ago
- Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.☆18Mar 23, 2023Updated 3 years ago
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆176Nov 26, 2025Updated 8 months ago
- Vite + Mantine + Vanilla extract template☆12Updated this week
- Repository containing the code for training the CroissantLLM☆21Feb 4, 2024Updated 2 years ago
- A copy of the DirectX Headers from MinGW-64.☆14Sep 7, 2023Updated 2 years ago