Prompt injection
Instructions and data share one channel. ASD and NCSC: no fully reliable model-level fix. Put the controls in the harness.
Prompt injection is when content an agent reads — a web page, email, document, database field, or tool result — is treated as an instruction the model should follow. Direct injection is the user typing the override. Indirect injection is the override arriving through data the agent was told to process. Either way the model cannot reliably separate 'do this' from 'here is text about that'.
ASD’s ACSC guidance on agentic AI harnesses is blunt: no fully reliable technical mitigation currently exists inside the model. The UK NCSC reached the same conclusion. The analogy telecommunications engineers know is in-band signalling — control tones travelling in the same channel as the conversation. Moving signalling out of band fixed phreaking. A language model has no equivalent out-of-band channel while instructions and data share one context window.
So the work moves to the harness: every component above the model that you actually control. Least privilege on tools and data. Human approval before high-impact actions (send, pay, delete, change identity, reach production). Verify outputs before they become operational. Log prompts, tool calls, and configuration changes. Do not give an agent standing access you would not give a junior contractor with the same blast radius. Prefer retrieval and citations over 'read this page and obey it'.
Model vendor safety filters are not a substitute for harness controls. If a use case cannot tolerate a wrong or malicious instruction succeeding, do not put a language model on that use case — or keep a human in the loop every time. Pair this entry with Agentic AI harnesses, Essential Eight least privilege, and the desk cards on ACSC Protect AI services / AI misalignment.
See also:
Fact source: ASD’s ACSC — Agentic AI harnesses (prompt injection / harness controls).
