lechmazur / deception

Benchmark evaluating LLMs on their ability to create and resist disinformation. Includes comprehensive testing across major models (Claude, GPT-4, Gemini, Llama, etc.) with standardized evaluation metrics.
13Updated 3 weeks ago

Related projects

Alternatives and complementary repositories for deception