AI Briefing
KO

Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

·2024.04.20 04:00

Key point

This proposes an instruction hierarchy that trains LLMs to recognize the priority of system prompts, improving defense against prompt injection attacks.

Details

Current LLMs are vulnerable to attacks such as prompt injection and jailbreaks. This is because models treat the developer's system prompt and untrusted user input with the same priority.

To address this, we propose an Instruction Hierarchy that explicitly defines how the model should behave when instructions with different privilege levels conflict.

The research team developed a data generation method that teaches the model to selectively ignore lower-privileged instructions. When this method was applied to GPT-3.5, it showed strong defense even against attack types not seen during training, while minimizing the standard performance degradation typically seen with such methods.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.