Computer Science > Computation and Language
arXiv:2607.21619 (cs)
Submitted on 1 Jun 2026
Title
Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-Based Stylistic Triggers Optimization
Authors
Bingjun Luo, Jialin Guo, Yue Yao, Xinpeng Ding
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance across various tasks, yet their safety alignment remains susceptible to jailbreak attacks. While existing content-based jailbreak methods often yield inconsistent results against rapidly evolving MLLMs, they fail to exploit non-content-based vulnerabilities. In contrast to prior work, we empirically identify a “Stylistic Inconsistency” in MLLMs: these models robustly understand content regardless of visual style, yet their defense mechanisms can be easily bypassed by specific stylistic triggers. Based on this finding, we propose Adversarial Style Optimization (ASO), a plug-and-play enhancement module designed to amplify existing visual jailbreak attacks. ASO fine-tunes an image-editing model to superimpose an optimized stylistic modification onto a given adversarial image. This optimization is driven by a Group Relative Policy Optimization (GRPO) agent guided by a Structurally-Tiered Reward Function, which combines a logit-based signal for detecting explicit refusals with a high-fidelity semantic evaluation from a powerful judge model. Extensive experiments demonstrate that ASO significantly improves the Attack Success Rate (ASR) of state-of-the-art attacks, highlighting stylistic biases as a scalable vector for red-teaming MLLMs. Our code is available at: https://github.com/bingjunluo/ASO.
Comments
Accepted by CVPR 2026 (Oral).
Subjects
- Computation and Language (cs.CL) – Primary
- Artificial Intelligence (cs.AI)
How to Cite
arXiv:2607.21619 [cs.CL]
(or arXiv:2607.21619v1 [cs.CL] for this version)
DOI: https://doi.org/10.48550/arXiv.2607.21619
via ArXiv CL
