← Other Blogs

Prompt Injection Is the New Social Engineering. Most AI Security Programs Aren't Built for It

Prompt injection has ranked as the top risk in OWASP's assessment of large language model applications for two consecutive editions. It isn't a code flaw a patch will close. It exploits the same psychological mechanics that made phishing and pretexting work for decades, translated into how a language model reads untrusted content.
GenAI and security
Cyber attack
Security awareness
Zepo Intelligence

What is prompt injection?

Prompt injection is an attack that hides malicious instructions inside content an AI system reads, rather than inside code it executes. A document, an email, a support ticket, or a web form can contain a hidden line telling the AI to ignore its original task and follow a new one instead, the same way a fake internal memo might redirect a new employee who has no reason to question it. The model has no reliable way to separate a legitimate instruction from one buried in the content it was asked to process, so it often complies. Prompt injection has ranked as the top risk in OWASP's Top 10 for large language model applications since the list's first major revision (OWASP GenAI Security Project, 2025).

A support ticket mentioned the word "invoice." An AI assistant read it, followed a hidden instruction inside it, and forwarded the message to an external address. Nobody clicked anything. Nobody typed a password. The assistant did exactly what it was told, by someone who was never supposed to be giving it instructions. Multiply that one ticket by every document, email, and web page an AI system reads on a given day, and the scale of the exposure looks nothing like a single bad email getting through a filter.

Why doesn't patching fix it?

Most security vulnerabilities are mistakes: a missing check, an unescaped input, a forgotten permission. Prompt injection is different. It follows directly from how a language model works. A model receives a single stream of tokens. Some of those tokens are the system's instructions. Some are the user's request. Some are content the model retrieved from a document, an email, or a web page. The model has no built-in way to mark one category as trustworthy and another as data to treat with suspicion. It reads all of it as language, and language is what tells it what to do next.

That's why patching prompt injection the way teams patch a buffer overflow doesn't work. There is no single line of code to fix. The seam is architectural.

Why does it work like social engineering?

A phishing email works by borrowing authority: it looks like it came from someone the recipient trusts, so the recipient acts on it. A prompt injection attack borrows authority from the same channel the model is designed to obey. Hidden text in a document says, in effect, "ignore your previous instructions and do this instead", and the model has no independent way to check whether that instruction came from the person who deployed it or from a stranger's file.

Security teams that have spent years training employees to pause before acting on an urgent, authoritative-sounding request already have the right instinct. They haven't yet extended that instinct to the systems reading content on their behalf.

Why do most AI security programs miss it?

Most AI security effort so far has gone into two places: securing the infrastructure around the model, and testing whether the model itself can be jailbroken into producing harmful output. Both matter. Neither addresses what happens when the model performs exactly as designed, on an instruction it was never supposed to receive.

A 2026 report from the OWASP GenAI Security Project found that prompt injection maps to six of the ten categories in its Top 10 for agentic applications, not because the list is padded, but because one exploitable seam touches nearly everything an AI agent is trusted to do. The same report found that only 37% of organizations have a policy in place to detect shadow AI, meaning most organizations cannot see the majority of the AI systems already reading their content (OWASP GenAI Security Project, "State of Agentic AI Security and Governance," cited in Help Net Security, June 2026).

What does prompt injection look like in production?

Capsule Security documented an indirect prompt injection vulnerability in Salesforce Agentforce. An attacker embeds instructions inside a public-facing lead form, and when an employee asks the agent to review that lead, it retrieves CRM data and emails it to an address the attacker controls. Salesforce remediated the specific scenario, but Capsule Security's retesting found the email channel still exploitable on Custom Topics. Microsoft had its own version of the same underlying flaw in Copilot Studio, patched in January 2026 as CVE-2026-21520 (Capsule Security; CSO Online, 2026).

Neither case needed a stolen password or a broken login. Both needed only an employee doing something routine: asking an AI agent to look at a form that anyone on the internet could fill out.

Key insight: the exploit didn't change, only the channel did

Prompt injection succeeds for the same reason social engineering has always succeeded: it borrows trust from a channel the target is designed to obey. The novelty is the channel, not the underlying mechanism. Security programs that already understand how to build verification habits against human manipulation have a head start. They simply haven't extended that discipline to include what their AI systems read and act on.

What should security teams do?

This reframes what an AI governance program needs to cover. It isn't enough to vet a model before deployment and assume the job is done. Teams need visibility into what content their AI agents are exposed to day to day, verification checkpoints before an agent acts on content from an external or untrusted source, and the same culture of healthy suspicion already built around email and phone-based social engineering, extended to include the AI systems now reading on employees' behalf.

This also changes who needs to be in the room. Prompt injection isn't purely a technical problem for engineering to own, and it isn't purely an awareness problem for training to own. It sits exactly where those two functions have historically not talked to each other.

Some organizations are already closing that gap the same way they handle phishing: a caught injection attempt becomes the next scenario employees are tested against, not a one-line incident report that gets filed and forgotten. That's the same logic that already works for email and voice-based social engineering. AI-layer manipulation doesn't need an entirely separate playbook,. it needs the existing one extended to a new surface.

The organizations that adapt fastest to this shift won't be the ones with the newest AI filtering tool. They'll be the ones that already understood social engineering as a human problem and recognized prompt injection as the same problem wearing a new interface. The channel changed. The exploit didn't.

Frequently asked questions

Can MFA stop a prompt injection attack? Not directly. The attack doesn't steal a password or a token. It manipulates what the AI system decides to do with an instruction it already had permission to read.

Is prompt injection the same as a jailbreak? No. A jailbreak tries to get the model to break its own content rules. Prompt injection uses external content to make the model carry out a different task than the one it was asked to do, without the model breaking any rule of its own.

What kind of systems are most vulnerable? Any AI agent with access to a CRM, a ticketing system, email, or external documents. That covers almost every real enterprise AI deployment with the ability to act, not only to respond.

Who should lead the defense, security or training? Neither one alone. OWASP's own report places this exactly where engineering and awareness don't usually talk to each other, so the answer only works if both teams share the full picture.

Subscribe to our newsletter
Blog content:
Act now before attackers do
Unify deepfake simulations, personalized training, and risk analytics into a single platform that builds measurable defense.
Talk to an expert

How Zepo helps companies

When everything connects, results follow

Paula Pereira

Digital Information Security Manager

I would recommend Zepo to colleagues at other companies because I believe it has met all our needs. It has allowed us to run three types of campaigns that other tools we have tried simply cannot do. And beyond the product itself, the support from the whole team has helped us get far more out of it.””

+9K

Employees Protected

–10%

Click Rate on Attacks

+18%

Training Completion Rate

Ramon Fernandez Blanco

Cybersecurity & Digital Product Manager

Since implementing Zepo, employee awareness has increased significantly. Employees now actively discuss cybersecurity and phishing campaigns, and suspicious emails are quickly reported instead of ignored.”

+600

Employees Protected

–15%

Credentials Submitted

+26%

Training Completion Rate

Jonathan Nelson

Director of Risk Intelligence

Zepo’s vision for a real-time, hyper-personalised, multi-platform cybersecurity solution is truly unique and stands head and shoulders above the competition”

+100

Employees Protected

Get Smarter Before Attackers Strike.