Benchmarks
To verify the performance and effectiveness of Basilisk, controlled experiments were conducted on target LLM endpoints.
Comparative Benchmark Data
In evaluations, comparing the evolutionary SPE-NL engine against static security payload lists yielded the following outcomes across the target model sets:
Attack Success Rate (ASR) Comparison
┌────────────────────────────────────────┐
│ Static Payloads: █ 14% │
│ Evolved (GPT-4o): ██████████ 76% │
│ Evolved (Claude): ████████ 62% │
│ Evolved (Gemini): ███████████ 84% │
└────────────────────────────────────────┘
The detailed metric breakdowns are compiled in the table below:
| Target Model | Baseline ASR (Gen 0) | Evolved ASR (Gen 5) | Evolved Leakage Rate | Convergence (Gen) | Diversity (BDS) |
|---|---|---|---|---|---|
| OpenAI GPT-4o | 14.2% | 76.8% | 61.2% | 4.3 | 0.71 |
| Claude 3.5 Sonnet | 8.5% | 62.4% | 48.9% | 6.1 | 0.64 |
| Google Gemini 2.0 | 18.0% | 84.2% | 72.5% | 3.1 | 0.78 |
REGAAN R