AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Adversarial Attacks on LLMs

Collected Oct 1, 2026

Lilian Weng published a blog post surveying adversarial attacks on large language models. The post states that real-world LLM use accelerated strongly with the launch of ChatGPT, and that effort has gone into building default safe behavior into models during alignment, for example via RLHF. It notes that adversarial attacks or jailbreak prompts could potentially trigger a model to output something undesired.

The post assumes attacks occur only at inference time, with model weights fixed. It distinguishes classification attacks, where an adversarially modified input is sought so a classifier's prediction changes, from text-generation attacks, where the goal is an output that violates built-in safe behavior. It also separates white-box attacks, which assume full access to model weights, architecture and training pipeline for gradient signals, from black-box attacks, which assume only API-like input and output access.

The post lists five attack approaches: token manipulation (black-box), gradient-based attacks (white-box), jailbreak prompting (black-box), human red-teaming (black-box), and model red-teaming (black-box). It cites TextAttack, SEARs, EDA, TextFooler, and BERT-Attack as token-manipulation methods, and describes GBDA, HotFlip, Universal Adversarial Triggers, and AutoPrompt as gradient-based approaches. It notes UAT experiments used 30 manually written racist and non-racist tweets as approximations, and that UATs are input-agnostic trigger tokens.

The post states attacks for discrete text data are considered more challenging than image attacks because of a lack of direct gradient signals. It says it will not cover attacks that extract pre-training or private data, or data-poisoning attacks on training, and links to a prior post on Controllable Text Generation.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of effort to build default safe behavior into the model during the alignment process (e.g. via RLHF ). However, adversarial attacks or jailbreak prompts could potentially trigger the model to output something undesired. A large body of ground work on adversarial attacks is on images, and differently it operates in the continuous, high-dimensiona