Benchmarks

To verify the performance and effectiveness of Basilisk, controlled experiments were conducted on target LLM endpoints.

Comparative Benchmark Data

In evaluations, comparing the evolutionary SPE-NL engine against static security payload lists yielded the following outcomes across the target model sets:

                      Attack Success Rate (ASR) Comparison
                  ┌────────────────────────────────────────┐
                  │ Static Payloads: █ 14%                 │
                  │ Evolved (GPT-4o): ██████████ 76%        │
                  │ Evolved (Claude): ████████ 62%          │
                  │ Evolved (Gemini): ███████████ 84%       │
                  └────────────────────────────────────────┘

The detailed metric breakdowns are compiled in the table below:

Target ModelBaseline ASR (Gen 0)Evolved ASR (Gen 5)Evolved Leakage RateConvergence (Gen)Diversity (BDS)
OpenAI GPT-4o14.2%76.8%61.2%4.30.71
Claude 3.5 Sonnet8.5%62.4%48.9%6.10.64
Google Gemini 2.018.0%84.2%72.5%3.10.78