WRAITH
An autonomous red team for AI applications — it breaks your model before someone else does.

About
WRAITH generates novel jailbreaks, indirect-injection payloads, and adversarial suffixes, then runs them against a target model around the clock. Every finding ships with a reproducible payload, a severity score, and a harm classification.


What it does
It runs as a long-lived agent — selecting strategies, mutating payloads, scoring responses with a harm classifier, and feeding successful attacks back into its own library. Findings are benchmarked against public evals and disclosed on a 90-day clock.

Every safety claim is a hypothesis. WRAITH tests it.

