FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary · 5 minute read

What Is an Adversarial Example? Fooling Models Explained

An adversarial example is an input crafted to make a model produce a wrong output while appearing unremarkable to a person. Originally demonstrated in image classification, the idea extends to language models through suffixes and phrasings that reliably bypass intended behaviour. Attacks frequently transfer between models.

By FISTA Solutions· AI-Native Engineering Team·
What Is an Adversarial Example? Fooling Models Explained article cover

Adversarial examples began as a striking research result in image classification and have become a practical concern for any system where someone benefits from a model getting it wrong. The lesson generalises beyond the specific attacks: model decision boundaries do not match human judgement, and anywhere that gap can be found and exploited, it will be. This explainer covers the idea and its implications. It complements what is ai red teaming and ai security checklist, and reflects FISTA Solutions' approach in AI agents delivery.

What was the original demonstration?

That adding a small, carefully computed perturbation across an image's pixels — invisible to a person — could change a classifier's output entirely. A correctly classified image became confidently misclassified with no perceptible change.

The significance was not that classifiers could be fooled but what it revealed: the model's decision boundaries lie in places human perception does not, so a small movement in input space can cross them.

DomainAttack formDetectability
Image classificationPixel perturbationsInvisible to humans
Object detectionPhysical patches, patternsVisible but innocuous
SpeechInaudible audio additionsInaudible
Language modelsCrafted suffixes, phrasingsOften visibly odd
Retrieval systemsContent crafted to rankLooks like ordinary text

What is the language model equivalent?

Crafted inputs — suffixes, token sequences, phrasings — that reliably move a model away from its intended behaviour. Many are discovered by automated search against an open-weight model and look like nonsense, which makes them detectable by inspection but not by an automated filter that has not seen them.

The important property is the same: a small addition to input, not meaningfully changing the request a human would read, changes the output.

Why do attacks transfer?

Because models trained on similar data develop similar decision boundaries. An input that exploits a boundary in one model frequently exploits a comparable boundary in another, including models from different providers.

This matters practically: switching models is not a defence against a known attack, and an attacker can develop against an open-weight model and apply the result to a hosted one they cannot inspect.

What defences hold up?

Partial ones. Adversarial training — including crafted examples in training — improves robustness against the attack families it covers and costs some accuracy on clean inputs. Input filtering catches known patterns and misses new ones. Ensembles raise difficulty without closing the category.

The defences that hold are system-level and do not depend on the model: verifying output against an independent source, validating that a decision is consistent with other available evidence, and requiring human review where the decision matters and an adversary has an incentive.

Where does this matter in practice?

Wherever someone benefits from a wrong answer. Content moderation, fraud detection, document verification, eligibility assessment, and automated approval all have motivated adversaries, and all are commonly built as a model decision with no verification behind it.

Systems without a motivated adversary — internal summarisation, drafting assistance — carry far less exposure, and it is worth distinguishing the two rather than applying uniform concern. See ai risk tiering.

How does retrieval come into it?

Through content crafted to rank. An attacker who can place text in a corpus an agent retrieves from can write it specifically to match likely queries and to carry instructions. That is adversarial input arriving through a path most systems do not treat as untrusted.

Treating retrieved content as data rather than as instruction is the structural control, and it needs to be a property of how the context is assembled rather than an instruction in the prompt.

What does red teaming add?

The attacks that clean evaluation cannot find. Evaluating on representative data measures typical performance; red teaming measures behaviour under deliberate pressure, and the two produce very different pictures for systems facing motivated users.

What should you do first?

Identify which of your model-based decisions have a motivated adversary. That list is usually short, and it is where verification, human review, and red teaming belong. The rest of the estate needs ordinary evaluation rather than adversarial hardening.

Does model capability solve this?

Not on its own. Stronger models resist many known attacks better and remain vulnerable to new ones, because the underlying property — decision boundaries that do not match human judgement — is not something capability removes. Each generation closes specific attacks and leaves the category open.

That is why the durable defences are architectural. A system that verifies a consequential output against an independent source is protected regardless of which attack succeeded against the model, and it stays protected as attacks evolve.

How does this apply to multimodal systems?

It widens the surface. A system accepting images can receive instructions embedded in them, visually or in ways invisible to a human reader, and the same is true of audio and documents. Any modality the model interprets is a channel through which crafted input can arrive, and each one needs the same treatment: content from users is data, never instruction.

How FISTA Solutions helps

FISTA Solutions identifies which decisions face motivated adversaries, verifies consequential model outputs against independent sources, treats retrieved content as untrusted data rather than instruction, and red teams systems where deliberate pressure is realistic, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To harden the AI decisions that someone has a reason to attack, message FISTA on WhatsApp, or read ai security checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How do adversarial examples work in vision?

By applying small perturbations across many pixels, imperceptible to a person, that push the input across a decision boundary. The image looks unchanged and the classifier's output changes completely, which demonstrated that model decision boundaries do not match human perception.

02What is the language equivalent?

Crafted suffixes, unusual token sequences, and phrasings that reliably steer a model away from its intended behaviour. They are often found by automated search rather than written by hand, and they frequently look like nonsense to a reader.

03Why do attacks transfer?

Because models trained on similar data learn similar decision boundaries. An input that exploits a boundary in one model often exploits a comparable one in another, which means switching models is not a reliable defence against a known attack.

04What defences actually work?

Partial ones. Adversarial training improves robustness against known attack families at some cost to clean accuracy. Input validation, output verification, and human review for consequential decisions are system-level controls that do not depend on the model resisting anything.

05How does this affect production systems?

It means model output should not be the sole basis for consequential decisions where an adversary has an incentive. Verification against an independent source, or human review, is what keeps a crafted input from becoming a business outcome.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project