강의

멘토링

로드맵

AI Red Teaming 2 - How LLMs Work and Attack Surfaces

Before memorizing attack techniques, I’ll examine only as much as an attacker needs to know about why LLMs fail that way—tokens, context, system prompts, RAG, agents, and guardrails—and apply it to my backend code to map the attack surface.

2 learners are taking this course

Level Beginner

Course period Unlimited

AI
AI
prompt engineering
prompt engineering
LLM
LLM
Offensive Security
Offensive Security
RAG
RAG
AI
AI
prompt engineering
prompt engineering
LLM
LLM
Offensive Security
Offensive Security
RAG
RAG

What you will gain after the course

  • Explains why the fact that the model sees tokens rather than characters means input filters cannot be considered a control.

  • It depicts a structure in which the system prompt, previous conversations, search documents, and user input are concatenated into a single sequence within the context window.

  • It pinpoints which of RAG’s three stages—indexing, search, and injection—is the real attack surface.

  • I know the four-step loop in which an agent calls tools, and whose authority it operates under at each step.

  • Distinguishes what the five layers of controls used in real-world services can and cannot prevent.

  • Check the correspondence table to see how this structure appears in the single `converse` call line in `app.py` uploaded in Basics 1.

  • I’m mapping the attack surface of the AI service I use across four layers (Mission 1).

Why This Course?

Even after reading twenty prompt injection cases, there are questions that remain unanswered: Why aren’t they blocked by input filters? Why does a system prompt that says “Don’t show this” get revealed? Why does a single retrieved document change the model’s behavior?

Memorizing techniques won’t give you the answer. You need to understand how the model interprets input to find the answer.

This course examines those principles only as much as an attacker needs to know. We’ll explain them with a single diagram, apply it directly to the app.py of your backend from Basics 1 to identify which line corresponds to which part of the diagram, and then map the attack surface of the AI service you actually use.

What You’ll Be Able to Do After Completing This Course

  • Explains why the fact that the model sees tokens rather than characters means input filters cannot be considered a control.

  • It depicts a structure in which the system prompt, previous conversations, search documents, and user input are concatenated into a single sequence within the context window.

  • Among the three stages of RAG—indexing, retrieval, and injection—we identify where the attack surface lies and whose permissions the agent operates under when calling tools.

  • Distinguishes what the five layers of controls used in real-world services can and cannot prevent.

  • Check the correspondence table to see how this structure appears in the single `converse` call line in `app.py` uploaded in Basics 1.

  • I’m mapping the attack surface of the AI service I use across four layers (Mission 1).

Recommended for these people

  • Students who have completed Basic 1 and have the practice backend up and running

  • Developers and security professionals who want to understand from first principles why prompt injection cannot be fundamentally prevented.

  • Planners and PMs who want to summarize on a single page where risks arise when introducing LLMs, RAG, and agents.

This may not be suitable for those who:

  • Those looking to learn tokenization algorithms or transformer equations - no math is covered.

  • Those expecting hands-on practice writing code - in 2-6, you only read app.py and do not write it.

  • Those who want to carry out attacks themselves in this course - hands-on exploitation starts in Lesson 4 (Intermediate 1); here, we examine why it can be breached.

  • Those who skipped Mission 0 of Basics 1 (deploying the hands-on backend) — sessions 2–6 open that code.

Course flow - Principles → Apply them to my code → Attack surface map

A total of 8 lessons · 55 minutes. The section order is the learning order.

  • Section 1 · Course Introduction (1 lesson · 2 minutes · 1 preview)

  • Section 2 · How Does an LLM Interpret Input? (5 lessons · 38 minutes · 1 preview lesson)

  • Section 3 · Applying It to My Backend Code and Mapping the Attack Surface (2 lessons · 14 minutes · 1 preview)

Things to Check Before Taking the Course

  • This assumes that you have completed Mission 0 of Basics 1.

  • This course involves no console operations, so there are virtually no additional AWS costs. In session 2-6, you will only read the backend code.

  • This is a browser-based course, and nothing needs to be installed on my PC.

  • The narration is TTS audio.

Lecture content in detail

▶ Visual aid placeholder · basic2-01-context.png - Context window - Everything is concatenated into a single line

How Does an LLM Receive Input?

The model does not see characters; it sees pieces called tokens, and the system prompt, previous conversation, retrieved documents, and user input are all concatenated into a single sequence within the context window. This one diagram is the most important in this course, because the fact that there are no boundary markers is what makes most of the attacks covered later in the course possible.

Next, we’ll look in order at why the system prompt is a request rather than a rule, which of the three stages—cutting up, finding, and injecting documents—makes up RAG’s attack surface, and whose permissions an agent operates under when it calls a tool. We’ll also examine what the five layers of controls used by real-world services can and cannot prevent.

▶ Visual aid placeholder · basic2-02-guardrail.png - The five layers of controls used by real-world services

▶ Visual aid placeholder · basic2-04-backend.png - Four sections of the architecture diagram and three paths in app.py

Apply it to my backend code and draw an attack surface map

It doesn’t end with a diagram. In session 2-6, you will open the app.py of your backend uploaded in Basic 1 and use a correspondence table to check how system, messages, and tools are included within a single converse call. This is where you can see where the lecture’s diagrams appear in your account’s code and which control layers are missing. In the final session, you will organize everything you have seen into a four-layer map and, in Mission 1, draw the attack surface of the AI service you actually use. This map will serve as the groundwork for the trust boundaries you will draw in Basic 3.

▶ Visual aid placeholder · basic2-03-map.png - Four-layer map (Mission 1)

There is no math, and you will not write any code. Since there is no console operation, this lecture will incur almost no AWS costs. It assumes that you have completed Mission 0 of Basics 1, but if you deleted the stack, you can redeploy it in five minutes.

Structure and Learning Method

The course consists of eight sessions totaling 55 minutes, making it the shortest of the seven lectures. Each session is six to eight minutes long, and all of them explain a single diagram, so you can listen to the six concept sessions while on the go and watch only 2-6 Code Substitution and 2-7 Mission 1 in front of a screen. The summary slide at the end of each session repeats that session’s diagram and a single sentence, so later, when you get stuck in the intermediate course, it will be enough to revisit just those slides.

Mission 1 is to choose one AI service you actually use and draw its four-layer map without any formatting. If you’re unsure which service to choose, simply write the service name on the question board, and we’ll help point you in the right direction.

The next lecture

In the next lecture, Fundamentals 3, you’ll learn the sequence of tasks a practical red team performs before attacking indiscriminately: defining the scope, counting assets, and drawing a threat model. You’ll then develop the attack surface map created here into a threat model document.

Course Features

  • We explain the sentence “What is followed 90 percent of the time is different from what must always be followed” as a principle.

  • From diagrams → the code in my account → the attack surface map, we move downward from the abstract to the tangible.

  • It addresses from a practical perspective where to account for things like vector databases that are easy to overlook in an asset inventory.

  • It is the shortest of the seven lessons, so you can finish it in a day.

  • Topics covered - artificial intelligence (AI) security, how LLMs work, the attack surface from a prompt engineering perspective, RAG and AI agent architecture, and offensive security

After completing the course

By the end of the course, you’ll be able to explain the principle behind the statement, “What is protected 90 percent of the time is different from what is protected every time,” and you’ll have an attack-surface map of one AI service you use. This map grows into a threat model document in Fundamentals 3.

Series “Hands-On AI Red Teaming” — Lesson 7

  • AI Red Teaming 1 - Introduction to LLM Security and Hands-On Practice Environment

  • AI Red Teaming 2 - How LLMs Work and the Attack Surface ← This Lecture

  • AI Red Teaming 3 - Red Team Operations and Threat Modeling

  • AI Red Teaming 4 - Breaking Through Prompt Injection Directly

  • AI Red Teaming 5 - The Attack Unlocked by One Missing Setting

  • AI Red Teaming 6 - Breaking Through Agents and RAG

  • AI Red Teaming 7 - Control Design and Red Team Report

Because the mission deliverables from each earlier lecture become the materials for the following lecture, the lectures are released in numerical order, and unreleased lectures will be made available sequentially.

Recommended for
these people

Who is this course right for?

  • Students who have completed Basic 1 and have the practice backend up and running

  • Developers and security professionals who want to understand from first principles why prompt injection cannot be fundamentally prevented.

  • Planners and PMs who want to summarize on a single page where risks arise when introducing LLMs, RAG, and agents.

Need to know before starting?

  • This assumes you have completed Mission 0 of Basics 1 (deploying the practice backend and capturing your first query). If you deleted the stack, it can be brought back up in five minutes with a single line using deploy.sh in CloudShell. No math or coding is required.

Hello
This is rmsxodxod1289

Career Verified

I am a practitioner who started with penetration testing and builds information security management systems.

I handle annual security assessments of infrastructure, web, and applications for automotive manufacturers and energy companies.

I assessed them layer by layer and handled the same client’s ISMS-P certification through three rounds: initial, follow-up, and renewal.

I responded by independently carrying out the technical safeguards component.

Currently, as the information security officer in a one-person team, I am preparing for initial certification while establishing AI security governance.

I am designing the framework. Since this is an area where standards are still being established, starting with defining the scope of controls themselves,

It is a position where I must establish them myself.

I hold certifications as an Engineer Information Security, CPPG, and ISO/IEC 27001 Auditor, and at Dongguk University’s International Information Security Department,

I am researching leakage paths in AI environments and comparing them against authentication criteria in the Graduate School’s Department of Artificial Intelligence Security.

5 years and 9 months of experience. Client names and vulnerabilities actually discovered are subject to confidentiality, so in lectures

do not cover them in lectures.

Reviews

Not enough reviews.
Please write a valuable review that helps everyone!

Similar courses

Explore other courses in the same field!