View on GitHub

Prompt injection

Methodologies planned

One sentence hidden in a web page can take over your AI agent.

The idea

An agent that reads email, documents or web pages treats all of it as text in its context. If that text contains instructions, the model may follow them. That’s prompt injection, and it’s the main security risk for agents with tools.

There’s no single fix, so the defense is layered: separate trusted and untrusted content, give tools the least access they need, and require confirmation for anything risky.

What the lesson will build

Key ideas

The video

When it’s published, the code will live in methodologies/ and this page will link to it.


All topics · Suggest a topic