Indirect Prompt Injection in Municipal Document-Processing Copilots: Attacks, Defences and Harm
Jorge Cisneros-González, José Antonio Ondiviela García, Javier Sánchez-Soriano
Source record
Source: Crossref
Published: Sep 5, 2026
DOI: 10.62161/sauc.v12.6399
Open original source ↗Source abstract
Public administrations are deploying large language model (LLM) assistants that process, summarise, classify and validate citizen-submitted documents. These copilots are exposed to indirect prompt injection: instructions hidden in manipulated documents that reach the model as if they were data. We develop a municipal aid-procedure copilot and evaluate its robustness across six attack objectives, five delivery vectors, five defences and four LLMs, with 3,000 attack evaluations and 1,000 legitimate evaluations. We combine deterministic detection with a human-validated LLM judge. Defences reduce attack success, but unevenly across objectives. They largely neutralise imperative attacks, such as request misrouting, while barely affecting summary falsification, revealing a provenance gap: the inability to determine whether an output value originates from an authoritative field or attacker-controlled content. Moreover, attack success and potential harm are decoupled: payment fraud and rule-exfiltration attacks have the greatest potential for harm despite intermediate success rates. We frame the risk within a human-in-the-loop model, where automation bias may amplify harm.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.