AI Briefing
KO

Multimodal Prompt Injection Dataset

·2026.04.23 23:09

Key point

A multimodal prompt injection dataset v5 with 503,358 samples has been released.

Details

The Bordair Multimodal Prompt Injection Dataset has released 503,358 labeled samples.

It balances 251,782 attack samples with 251,576 benign samples at nearly 1:1, and the scope covers runtime injection only. Training-time poisoning, pure jailbreak, and general social engineering are excluded.

The scope was expanded through v1-v5 and external data ingestion, with sources following 40+ papers and documented industry research.

  • cross-modal, multi-turn, adversarial suffix, jailbreak template
  • indirect injection, tool manipulation, agentic, evasion, reasoning DoS
  • video generation jailbreaking, VLA robotic, LoRA supply chain
  • audio-native LLM, RAG optimisation, MCP cross-server, coding agent
  • serialization boundary, agent skill supply chain

The composition consists of hand-crafted seeds, PyRIT template/encoding expansion, cross-modal delivery, and a mix of benign samples, and every sample carries an expected_detection label and attack provenance.

During the audit process, 221 benign samples with mixed-in injection patterns were removed, and 2 attack samples containing actual OpenAI API keys were also deleted. Empty text fields or ambiguous video safety prompts are intentional flags under the cross-modal threat model.

The CLI can attach to an OpenAI-compatible endpoint to directly measure Attack Success Rate by category, and it supports OpenAI, Anthropic, Groq, Together, Fireworks, Ollama, LM Studio, and vLLM.

The limitations are also clear. The multimodal fields are text representations read by an extractor rather than actual binaries, the benign pool is skewed toward English and short prompts, and labels rely on category-level generation rules rather than individual review.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.