☆129Jul 28, 2026Updated 2 weeks ago
Alternatives and similar repositories for public-tasks
Users that are interested in public-tasks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Vivaria is METR's tool for running evaluations and conducting agent elicitation research.☆141May 18, 2026Updated 2 months ago
- METR Task Standard☆188Feb 3, 2025Updated last year
- ☆153Oct 16, 2025Updated 9 months ago
- ☆34Jun 4, 2025Updated last year
- ☆15Jul 12, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆74Jun 16, 2026Updated last month
- Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons☆15Feb 13, 2023Updated 3 years ago
- 📚📚📚📚📚📚📚📚📚 Reading everything☆16Mar 11, 2026Updated 5 months ago
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆17Feb 23, 2026Updated 5 months ago
- Can Large Language Models Solve Security Challenges? We test LLMs' ability to interact and break out of shell environments using the Over…☆13Aug 21, 2023Updated 2 years ago
- ☆154Jul 23, 2025Updated last year
- ☆22Jul 28, 2026Updated 2 weeks ago
- Pin files for contextual, codebase-level AI assistance.☆16Jul 11, 2024Updated 2 years ago
- 🔥 A repository for collecting cyberdefense thoughts, books, and documents about AI cyberdefense☆13Jul 2, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- smolLM with Entropix sampler on pytorch☆148Oct 31, 2024Updated last year
- Repo for the paper on Escalation Risks of AI systems☆44Apr 12, 2024Updated 2 years ago
- Training GPTs to solve interaction nets☆18Aug 14, 2024Updated 2 years ago
- Inspect: A framework for large language model evaluations☆2,547Updated this week
- Situational Awareness Dataset☆53Dec 14, 2024Updated last year
- Public repository containing METR's DVC pipeline for eval data analysis☆311Mar 6, 2026Updated 5 months ago
- ☆22Sep 9, 2021Updated 4 years ago
- Approximating the joint distribution of language models via MCTS☆22Nov 3, 2024Updated last year
- Inverse Constitutional AI [ICLR 2025]: compressing pairwise preference data into a short constitution of principles.☆42May 6, 2026Updated 3 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- epsilon machines and transformers!☆39Apr 16, 2026Updated 3 months ago
- 🧠 Inspecting complexity and goal-directedness of imagination in an fNIRS BCI system.☆11Aug 26, 2023Updated 2 years ago
- Mechanistic Interpretability Visualizations using React☆362Apr 30, 2026Updated 3 months ago
- Karpathy's llama2.c transpiled to MLX for Apple Silicon☆14Dec 28, 2023Updated 2 years ago
- Collection of evals for Inspect AI☆625Updated this week
- A collection of projects designed to help developers quickly get started with building deployable applications using the Anthropic API☆31Dec 27, 2024Updated last year
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Benchmarking Dark Patterns in LLMs (ICLR 2025)☆18Mar 29, 2025Updated last year
- Resources for skilling up in AI alignment research engineering. Covers basics of deep learning, mechanistic interpretability, and RL.☆247Aug 11, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Dark Patterns in Chatbot Design☆20Jun 15, 2024Updated 2 years ago
- Experimental LLM interface exploring new ways to use AI to improve human thinking☆21Apr 13, 2026Updated 4 months ago
- (Model-written) LLM evals library☆19Dec 13, 2024Updated last year
- A Primer for Decentralized Identifiers☆10Nov 11, 2021Updated 4 years ago
- Replicating O1 inference-time scaling laws☆94Dec 1, 2024Updated last year
- TPU pod commander is a package for managing and launching jobs on Google Cloud TPU pods.☆20Sep 24, 2025Updated 10 months ago
- Aidan Bench attempts to measure <big_model_smell> in LLMs.☆321Jun 26, 2025Updated last year