☆34Nov 7, 2024Updated last year
Alternatives and similar repositories for Truth_is_Universal
Users that are interested in Truth_is_Universal are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆44Feb 11, 2025Updated last year
- Math evaluations of llama models.☆10Jan 3, 2024Updated 2 years ago
- Datasets used in the paper "Reward hacking behavior can generalize across tasks"☆16Aug 17, 2025Updated 11 months ago
- ☆34Nov 16, 2025Updated 8 months ago
- Code for the ICLR 2024 paper "How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions"☆74Jun 19, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Creating a game to play Figgie & Train an agent to play against☆15Dec 3, 2022Updated 3 years ago
- Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals; ACL 2024☆13May 24, 2024Updated 2 years ago
- Official Code for our paper: "Language Models Learn to Mislead Humans via RLHF""☆20Oct 11, 2024Updated last year
- ☆16May 7, 2026Updated 3 months ago
- Template for Python-based data science projects in the Alexandra Institute.☆12Jul 25, 2026Updated 2 weeks ago
- Website for the MIT/Harvard Computational Neuroscience Journal Club☆11Apr 7, 2025Updated last year
- Scalable DBSCAN and OPTICS for clustering high-dimensional datasets using random projections☆14Nov 1, 2024Updated last year
- Code for "Astraea: Grammar-based Fairness Testing"☆10Jan 7, 2022Updated 4 years ago
- This repository contains the code and data for the paper "SelfIE: Self-Interpretation of Large Language Model Embeddings" by Haozhe Chen,…☆58Dec 9, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The project page of paper: Aha! Adaptive History-driven Attack for Decision-based Black-box Models [ICCV 2021]☆10Feb 23, 2022Updated 4 years ago
- Compositional Muon release☆23Jun 5, 2026Updated 2 months ago
- ☆18Mar 16, 2026Updated 4 months ago
- ☆15Apr 27, 2024Updated 2 years ago
- Official Implementation of "The Graph Database Interface: Scaling Online Transactional and Analytical Graph Workloads to Hundreds of Thou…☆14Jul 2, 2025Updated last year
- ☆95Jan 22, 2025Updated last year
- Can We Trust Large Language Models?: A Benchmark for Responsible Large Language Models via Toxicity, Bias, and Value-alignment Evaluation☆25Oct 12, 2023Updated 2 years ago
- This is official project in our paper: Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers☆31Jan 13, 2024Updated 2 years ago
- Chess agent specification gaming☆25Aug 3, 2026Updated last week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [NeurIPS 2025] The official implementation of the paper "DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agen…☆59Jul 16, 2026Updated 3 weeks ago
- LoFiT: Localized Fine-tuning on LLM Representations☆45Jan 15, 2025Updated last year
- ☆14Mar 23, 2023Updated 3 years ago
- Data set for LREC 2020 paper "I Feel Offended, Don't Be Abusive!"☆19Sep 23, 2023Updated 2 years ago
- Sparse probing paper full code.☆68Dec 17, 2023Updated 2 years ago
- BAD: BiAs Detection for Large Language Models in the context of candidate screening (EECS 692)☆12Feb 14, 2024Updated 2 years ago
- Official repository for the paper "Gradient-based Jailbreak Images for Multimodal Fusion Models" (https//arxiv.org/abs/2410.03489)☆20Oct 22, 2024Updated last year
- Discriminative Feature Selection via A Structured Sparse Subspace Learning Module☆12Apr 15, 2022Updated 4 years ago
- Open source replication of Anthropic's Crosscoders for Model Diffing☆68Oct 27, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons☆15Feb 13, 2023Updated 3 years ago
- Dcraw native builds for Windows OS☆12Dec 29, 2017Updated 8 years ago
- Learning Certified Individually Fair Representations☆25Nov 7, 2020Updated 5 years ago
- Implementation of the paper "Hallucination Detection in LLMs Using Spectral Features of Attention Maps"☆16Oct 18, 2025Updated 9 months ago
- Interpretable unified language safety checking with large language models☆32Apr 15, 2023Updated 3 years ago
- Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval☆52Jan 6, 2026Updated 7 months ago
- ☆195Mar 8, 2026Updated 5 months ago