Prompt injection
Prompt injection occurs when instructions hidden in content such as email, webpages, documents, CRM notes, or retrieved knowledge alter an agent's behaviour.
Start Here07 / CONTROL / NEXIUS CONCEPT GUIDE
Turn AI risk and governance requirements into visible ownership, permissions, approvals, evidence, escalation, and accountability.
Start with the answerDIRECT ANSWER
Human control means named people remain accountable for agent outcomes and retain informed authority at consequential points. AI risk analysis explains what can go wrong, governance defines the required responsibilities and controls, and Mission Control operationalises them through identities, permissions, evidence, approvals, escalation, monitoring, and incident response.
Prompt injection occurs when instructions hidden in content such as email, webpages, documents, CRM notes, or retrieved knowledge alter an agent's behaviour.
Tool poisoning occurs when malicious instructions or deceptive information are placed in tool descriptions, schemas, metadata, errors, or outputs that an agent trusts.
THE LETHAL TRIFECTA
Security researcher Simon Willison calls this combination the lethal trifecta. Each capability can be useful on its own. When one agent holds all three in the same session, malicious content can steer it from reading a secret to sending that secret outside the organisation.
Customer records, internal documents, credentials, email, financial information, or any other data the attacker cannot reach directly.
Email, webpages, uploaded files, search results, retrieved knowledge, tool output, or other material an attacker may be able to shape.
Email, messages, web requests, shared documents, API calls, or any channel that can transmit information beyond the trusted boundary.
Untrusted instructions enter through content. The hijacked agent retrieves private data. Its outbound capability delivers that data to the attacker.
Break the chain by removing at least one capability from any unapproved session. If the workflow genuinely needs all three, separate reading from execution, minimise the data exposed, constrain destinations, and require independent human approval before information can leave.
WHAT THE CONCEPT CHANGES
The organisation should know the agent's purpose, accountable owner, approved skills, data boundary, tools, credentials, and operating status.
Event logs should retain the sources, tool calls, decisions, outputs, exceptions, approvals, and hand-offs needed to explain what happened.
High-impact actions require stronger review, while low-risk repeatable work can operate with monitoring and exception-based intervention. The control model should be explicit.
Content an agent reads may contain instructions designed to override the user's intent. The system must preserve the difference between trusted policy and untrusted material being analysed.
A poisoned tool description, schema, output, or dependency can redirect an agent even when the user's prompt is safe. Tools need provenance, review, scoped permissions, and continuing monitoring.
No model-layer defence is perfect. Hard limits on data, tools, identity, network access, approvals, and irreversible actions reduce the damage when detection fails.
Control depth should reflect who may be affected, the sensitivity of the data, the authority delegated to AI, the reversibility of actions, and the financial, legal, safety, operational, and reputational consequences of failure.
A named person or committee owns the outcome, residual risk, approval boundaries, and response to incidents. Human oversight must occur at a point where intervention is informed, timely, and capable of changing the result.
A launch review is not enough. Models, prompts, knowledge, tools, providers, data, users, and threats change. Testing, monitoring, audit trails, incident learning, and change control must continue across the complete lifecycle.
DECISION GUIDE
Every production agent needs a named business purpose, accountable human owner, risk tier, approved skills and tools, data boundary, credentials, and current operating status.
Prompts can guide behaviour but cannot provide final protection. The application and system of record must enforce identity, permissions, validation, approvals, transaction rules, and prohibited actions.
Mission control should connect intent, source context, tool calls, policy decisions, approvals, outputs, exceptions, cost, and hand-offs so leaders can inspect both an individual action and the operating pattern.
Classify webpages, emails, documents, tickets, retrieved knowledge, and tool output as untrusted content. Preserve source labels and prevent external content from silently becoming operating instructions.
Maintain a governed tool registry. Review tool descriptions, schemas, dependencies, permissions, errors, and runtime outputs before approval and whenever the tool or manifest changes.
Separate reading, drafting, reviewing, and execution. Give each agent only the tools and data required for its role, then require human approval before sending, deleting, spending, deploying, or modifying important records.
Record the purpose, affected people, data, model, provider, tools, decisions, actions, and consequence of failure. A low-impact internal drafting aid and an agent that can approve credit, send money, or alter customer records should not pass through the same governance path.
Use IMDA guidance as the Singapore-wide operating foundation, add MAS FEAT, Veritas, model-risk, and other supervisory expectations for financial services, and map the control lifecycle to NIST AI RMF or an appropriate management standard. Also identify applicable law, regulation, contract, and internal policy; a voluntary framework does not replace them.
Define acceptance tests, independent review, monitoring thresholds, human checkpoints, incident response, change approval, and retirement criteria. Keep enough documentation to reconstruct what the system was allowed to do, which version ran, what evidence was reviewed, and who accepted the residual risk.
THE NEXIUS OPERATING INTERPRETATION
AI risks explain what can go wrong, governance defines the responsibilities and controls, and Mission Control operationalises both through named ownership, bounded authority, approvals, evidence, escalation, and continuing oversight.
AUTHORITATIVE SOURCES
These external sources provide the original research, engineering guidance, standards, or platform material informing this guide. Inclusion does not imply endorsement or partnership.
NEXIUS FIELD NOTES
These practitioner notes by Darryl Wong show how Nexius applies and adapts the concepts. They are implementation guidance—not claims that Nexius originated the underlying concepts.
COMMON QUESTIONS
No. A dashboard displays information. Mission control also applies operating rules, permissions, approvals, escalation, evidence retention, and accountable ownership.
No. Review should be proportional to consequence, ambiguity, novelty, and risk. Well-designed systems concentrate human judgement where it matters.
It is the combination of access to private data, exposure to untrusted content, and the ability to communicate externally. Together, these capabilities can let a prompt-injection attacker steer an agent to retrieve sensitive information and send it outside the organisation.
No. Instructions and model safeguards can reduce risk, but they are not a complete security boundary. Safe deployment also requires untrusted-content handling, least privilege, isolation, approval gates, monitoring, and evidence.
No. The attack can sit in a tool description, schema field, parameter guidance, error message, runtime output, plugin document, or compromised dependency that the agent interprets as trusted context.
No. SMEs should start with bounded use cases, limited permissions, observable workflows, human ownership, and controls proportionate to the consequence of failure.
The IMDA Model AI Governance Framework is practical governance guidance, not a substitute for applicable law or sector regulation. Organisations must separately identify obligations such as data-protection, consumer, employment, intellectual-property, cybersecurity, financial-services, and contractual requirements.
MAS guidance is especially relevant to financial institutions and AI or data-analytics use in financial products, services, and regulated operations. Other organisations can still learn from FEAT, Veritas, and model-risk-management practices, but should not imply MAS compliance or approval without a proper assessment.
A governance framework describes the organisational principles, responsibilities, and practices expected. AI Verify adds a testing framework and toolkit that can help assess selected governance principles through technical tests and process checks. Testing supports governance; it does not replace accountable decisions or continuous oversight.
No single checklist proves that an AI system is safe, lawful, fair, or suitable. A defensible decision depends on the use case, applicable obligations, effective controls, test evidence, known limitations, independent review where appropriate, and continuing monitoring.
Begin with a bounded use case, a named human owner, explicit success and stop conditions, and the minimum data and tool access needed. Connect the concept to a real workflow before expanding it.
Turn AI risk and governance requirements into daily, visible, enforceable human control. Human ownership, proportionate permissions, observable evidence, and clear escalation should remain part of the operating design.
Define measurable outcomes before implementation, then review quality, time, cost, exceptions, human acceptance, evidence completeness, and unintended consequences. Improve or stop the workflow when the evidence does not support expansion.
COMPLETE RISK LANDSCAPE
Prompt injection and tool poisoning are gateway threats. Review the wider business-process risk landscape, including excessive permissions, data leakage, unsafe actions, poisoned memory, weak evidence, runaway loops, brittle integrations, and compliance drift.