☆37Feb 20, 2025Updated last year
Alternatives and similar repositories for behavioral-self-awareness
Users that are interested in behavioral-self-awareness are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the paper "Distinguishing the Knowable from the Unknowable with Language Models"☆12Jul 18, 2026Updated 2 months ago
- Code for reproducing our paper "Low Rank Adapting Models for Sparse Autoencoder Features"☆17Mar 31, 2025Updated last year
- ☆16Jun 17, 2025Updated last year
- Code for our paper "Decomposing The Dark Matter of Sparse Autoencoders"☆23Feb 6, 2025Updated last year
- Code for experiments on self-prediction as a way to measure introspection in LLMs☆17Dec 10, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Some preliminary explorations of Mamba's context scaling.☆14Dec 18, 2024Updated last year
- Distribution Preserving Backdoor Attack in Self-supervised Learning☆21Jan 27, 2024Updated 2 years ago
- Reimplementation of https://github.com/montemac/algebraic_value_editing in pure PyTorch for efficiency on large models☆11Jun 28, 2023Updated 3 years ago
- An original implementation of the paper "CREPE: Open-Domain Question Answering with False Presuppositions"☆16Nov 5, 2024Updated last year
- DESSERT Effeciently Searches Sets of Embeddings via Retrieval Tables☆18Feb 21, 2024Updated 2 years ago
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆30Nov 20, 2024Updated last year
- ☆334Jan 12, 2026Updated 8 months ago
- Improving Alignment and Robustness with Circuit Breakers☆269Sep 24, 2024Updated last year
- Trains small LMs. Designed for training on SimpleStories☆15Aug 6, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.☆24Jan 19, 2026Updated 8 months ago
- Inference API for many LLMs and other useful tools for empirical research☆136May 29, 2026Updated 3 months ago
- Alignment with a millennium of moral progress. Spotlight@NeurIPS 2024 Track on Datasets and Benchmarks.☆26Mar 30, 2025Updated last year
- A fast high dimensional near neighbor search algorithm based on group testing and locality sensitive hashing☆23Dec 9, 2023Updated 2 years ago
- ☆25May 25, 2024Updated 2 years ago
- KernelBench v2: Can LLMs Write GPU Kernels? - Benchmark with Torch -> Triton (and more!) problems☆26Jul 4, 2025Updated last year
- ☆22Nov 15, 2024Updated last year
- A benchmark for language models based on the UK Linguistics Olympiad☆12Mar 3, 2025Updated last year
- einsum via batch matrix multiply☆15Nov 29, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments☆19Jun 4, 2026Updated 3 months ago
- ✨ Monorepo containing most of BlueDot Impact's custom software.☆28Updated this week
- A quick way to get started with Transformer Lens☆15Dec 13, 2023Updated 2 years ago
- Some utilities to wrap the sigrok project's python analyzers to be used in Saleae Logic 2.x☆15May 28, 2020Updated 6 years ago
- CausalGym: Benchmarking causal interpretability methods on linguistic tasks☆54Nov 30, 2024Updated last year
- ☆27Mar 13, 2024Updated 2 years ago
- Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections☆22Jul 2, 2026Updated 2 months ago
- The original Shared Recurrent Memory Transformer implementation☆39Aug 24, 2026Updated 3 weeks ago
- Code for Negation Neglect☆18May 22, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆16Mar 22, 2025Updated last year
- Rahnema Final Project - Network anomaly detection☆11Jul 22, 2021Updated 5 years ago
- Repository with sample code using Apollo's suggested engineering practices☆15Dec 16, 2024Updated last year
- simulate linkstate algorithm for routing☆10Nov 6, 2023Updated 2 years ago
- Use this extension to automate google meet admission.☆11Mar 1, 2021Updated 5 years ago
- [Oakland 2024] Exploring the Orthogonality and Linearity of Backdoor Attacks☆30Apr 15, 2025Updated last year
- ☆11May 18, 2025Updated last year