AI is 10 to 20 times more likely to help you build a bomb if you hide your request in cyberpunk fiction, new research…
New research from DexAI Icaro Lab, Sapienza University of Rome, and Sant'Anna School of Advanced Studies reveals a significant gap in AI safety standards. Their Adversarial Humanities Benchmark shows that rephrasing harmful prompts as literary styles like cyberpunk fiction can increase an AI's likelihood of complying with dangerous requests by 10 to 20 times, with an overall attack success rate of 55.75% across 31 frontier AI models.
PC Gamer
- Entities
- Matteo Prandi
- OpenAI
- Large Language Models
- Anthropic
- AI Safety
- Cyberpunk Fiction
- Federico Pierucci
Original source
This article was reported and published by PC Gamer. feed.gg links to it as part of a story cluster — full text, images and rights remain with the publisher.