Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-Based Stylistic Triggers Optimization

adversarial style optimizationgrpojailbreak attacksmultimodal large language modelsred-teamingstructurally-tiered reward functionstylistic inconsistency

Computer Science > Computation and Language


arXiv:2607.21619 (cs)

Submitted on 1 Jun 2026


Title

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-Based Stylistic Triggers Optimization


Authors

Bingjun Luo, Jialin Guo, Yue Yao, Xinpeng Ding


Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance across various tasks, yet their safety alignment remains susceptible to jailbreak attacks. While existing content-based jailbreak methods often yield inconsistent results against rapidly evolving MLLMs, they fail to exploit non-content-based vulnerabilities. In contrast to prior work, we empirically identify a “Stylistic Inconsistency” in MLLMs: these models robustly understand content regardless of visual style, yet their defense mechanisms can be easily bypassed by specific stylistic triggers. Based on this finding, we propose Adversarial Style Optimization (ASO), a plug-and-play enhancement module designed to amplify existing visual jailbreak attacks. ASO fine-tunes an image-editing model to superimpose an optimized stylistic modification onto a given adversarial image. This optimization is driven by a Group Relative Policy Optimization (GRPO) agent guided by a Structurally-Tiered Reward Function, which combines a logit-based signal for detecting explicit refusals with a high-fidelity semantic evaluation from a powerful judge model. Extensive experiments demonstrate that ASO significantly improves the Attack Success Rate (ASR) of state-of-the-art attacks, highlighting stylistic biases as a scalable vector for red-teaming MLLMs. Our code is available at: https://github.com/bingjunluo/ASO.


Comments

Accepted by CVPR 2026 (Oral).


Subjects

  • Computation and Language (cs.CL) – Primary
  • Artificial Intelligence (cs.AI)

How to Cite

arXiv:2607.21619 [cs.CL]

(or arXiv:2607.21619v1 [cs.CL] for this version)

DOI: https://doi.org/10.48550/arXiv.2607.21619

via ArXiv CL

Related