yance521 / super_pe_evaluationView on GitHub
A TDD-style prompt evaluation skill — build reusable eval datasets, run prompt versions against locked test cases, score outputs with evidence, and compare results without overwriting history. Local-first, platform-agnostic, append-only.
21Jul 9, 2026Updated last week

Alternatives and similar repositories for super_pe_evaluation

Users that are interested in super_pe_evaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?