Adversarial Prompts Degrade LLM Instruction-Following Performance
Key point
Adversarial user prompts consistently lowered IFEval performance across 14 LLMs.
Details
Evaluating 14 instruct models with IFEval, hostile user prompt reduced instruction-following performance across all model families.
Comparing performance against a length-matched neutral control, three independent training recipes in the 7~8B FP16 range all showed significant declines.
- Llama 3.1 8B Instruct: 76.3 → 66.9, residual -9.8pp
- Mistral 7B Instruct: 60.2 → 55.8, residual -6.2pp
- Qwen3 8B Instruct: 78.8 → 72.4, residual -6.1pp
- Average residual: -7.4pp, relative decline about 10.2%
Reproducibility was confirmed broadly.
- Llama 3.1 8B FP16 -9.8pp, Q4 MLX -9.5pp, 70B Q4 MLX -6.4pp
- Mistral 7B FP16 -6.2pp, Q4 MLX -7.7pp, Mistral Large 123B Q4 MLX -5.6pp
- Qwen3 0.6B Q4 MLX -9.6pp, 8B FP16 -6.1pp, 8B Q4 MLX -7.6pp
- Qwen3 30B-A3B Q4 MLX -8.1pp, 32B Q4 MLX -7.2pp
All comparisons were p < .001 under paired bootstrap N=10,000. The decline diminished as scale increased, but the effect did not disappear even at 123B.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.