☆131Jul 28, 2026Updated last month
Alternatives and similar repositories for public-tasks
Users that are interested in public-tasks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Vivaria is METR's tool for running evaluations and conducting agent elicitation research.☆142May 18, 2026Updated 4 months ago
- METR Task Standard☆197Feb 3, 2025Updated last year
- ☆164Oct 16, 2025Updated 11 months ago
- ☆36Jun 4, 2025Updated last year
- ☆15Jul 12, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆75Jun 16, 2026Updated 3 months ago
- Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons☆15Feb 13, 2023Updated 3 years ago
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆17Feb 23, 2026Updated 7 months ago
- ☆155Jul 23, 2025Updated last year
- 📚📚📚📚📚📚📚📚📚 Reading everything☆17Mar 11, 2026Updated 6 months ago
- ☆24Jul 28, 2026Updated last month
- Pin files for contextual, codebase-level AI assistance.☆16Jul 11, 2024Updated 2 years ago
- 🔥 A repository for collecting cyberdefense thoughts, books, and documents about AI cyberdefense☆13Jul 2, 2023Updated 3 years ago
- smolLM with Entropix sampler on pytorch☆148Oct 31, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Work in progress! I don't recommend looking at the code right now.☆25Updated this week
- Repo for the paper on Escalation Risks of AI systems☆45Apr 12, 2024Updated 2 years ago
- Training GPTs to solve interaction nets☆18Aug 14, 2024Updated 2 years ago
- Inspect: A framework for large language model evaluations☆2,850Updated this week
- Situational Awareness Dataset☆55Dec 14, 2024Updated last year
- ☆22Sep 9, 2021Updated 5 years ago
- Public repository containing METR's DVC pipeline for eval data analysis☆331Mar 6, 2026Updated 6 months ago
- Approximating the joint distribution of language models via MCTS☆22Nov 3, 2024Updated last year
- Inverse Constitutional AI [ICLR 2025]: compressing pairwise preference data into a short constitution of principles.☆42May 6, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🧠 Inspecting complexity and goal-directedness of imagination in an fNIRS BCI system.☆11Aug 26, 2023Updated 3 years ago
- epsilon machines and transformers!☆40Apr 16, 2026Updated 5 months ago
- Mechanistic Interpretability Visualizations using React☆366Apr 30, 2026Updated 4 months ago
- Karpathy's llama2.c transpiled to MLX for Apple Silicon☆14Dec 28, 2023Updated 2 years ago
- Collection of evals for Inspect AI☆680Updated this week
- A collection of projects designed to help developers quickly get started with building deployable applications using the Anthropic API☆31Dec 27, 2024Updated last year
- Mamba support for transformer lens☆20Sep 17, 2024Updated 2 years ago
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Resources for skilling up in AI alignment research engineering. Covers basics of deep learning, mechanistic interpretability, and RL.☆248Aug 11, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Dark Patterns in Chatbot Design☆21Jun 15, 2024Updated 2 years ago
- Code for creating a 3D lidar map of Leslie St and Highway 7.☆10Jan 4, 2021Updated 5 years ago
- Experimental LLM interface exploring new ways to use AI to improve human thinking☆21Apr 13, 2026Updated 5 months ago
- (Model-written) LLM evals library☆19Dec 13, 2024Updated last year
- A Primer for Decentralized Identifiers☆10Nov 11, 2021Updated 4 years ago
- TPU pod commander is a package for managing and launching jobs on Google Cloud TPU pods.☆20Sep 24, 2025Updated last year
- Aidan Bench attempts to measure <big_model_smell> in LLMs.☆321Jun 26, 2025Updated last year