Preprint

Memory Makes the Poison:
Over-Memorization Drives Visual Data Poisoning in LVLMs

Sy-Tuyen Ho1 Yaseema R. A. Epa2 Yasoda L. A. Epa2 Andrew Mendez1 Huy Nghiem1 Quan H. Nguyen1 Xudong Jiang3 Alex Kot3 Furong Huang1 Ngai-Man Cheung2

1University of Maryland, College Park    2Singapore University of Technology and Design    3Nanyang Technological University

TL;DR. Over-memorization during fine-tuning, rather than visual perturbations, drives LVLM poisoning. RejectShield rejects poisoned samples before fine-tuning, reducing attack success while preserving model utility.
RejectShield defense results vs. existing defenses across LLaVA 1.5 and MiniGPT4-v2

Figure 1. (1) Controlled experiments (red) isolate the effect of data memorization during fine-tuning, showing that LVLMs over-memorize injected concepts, leading to hallucinations even without adversarial perturbations. (2) Our analysis explains why existing purification-based defenses (DiffPure, DiffJPEG) fail: they address the wrong root cause. (3) RejectShield (pink) directly disrupts memorization by rejecting poisoned samples, reducing attack success rates by up to 99% while preserving model utility.

Abstract

Large Vision-Language Models (LVLMs) excel across tasks, yet their safety and security remain underexplored. Among emerging threats, LVLM poisoning attacks pose a serious risk by inducing targeted hallucinations in fine-tuned models. Although effective, the root cause of these attacks remains poorly understood. The SOTA attack, ShadowCast, originally attributes its success to carefully injected visual perturbations — leading existing defenses to focus on purification methods, which have proven largely ineffective.

In this work, we argue this gap stems from a limited understanding of LVLM vulnerabilities during fine-tuning. We systematically study the fine-tuning process and, for the first time, identify over-memorization as the key vulnerability: LVLMs tend to over-memorize fine-tuning concepts, directly leading to hallucinations. Our finding overturns the original justification — the dominant driver is memorization of injected concepts, not the visual perturbation. Guided by this insight, we introduce RejectShield, a simple rejection-based defense that explicitly disrupts memorization. Across eight settings spanning attack variants, attack goals, model families, and access regimes, RejectShield reduces attack success by up to 99% while largely preserving normal performance.

Multimodal input exacerbates memorization in LVLMs compared to counterpart unimodal LLMs

We design controlled experiments that isolate data memorization from visual perturbations by replacing poisoned images with their benign counterparts while keeping all other inputs identical. This isolates the memorization effect.

Memorization comparison between LVLMs and unimodal LLMs

Figure 2. Left: Fine-tuned LVLMs exhibit a sharp jump in attack success (from ~0% to >90%) once the injection ratio exceeds 1%, confirming rapid concept memorization. Right: Unimodal LLMs with the same language backbone (Vicuna 1.5 7B) remain robust even at 5% injection. The only variable is the presence of multimodal visual input — confirming that multimodality is the amplifying factor.

RejectShield: Reject, Don't Purify

Inspired by our findings, RejectShield takes a fundamentally different approach from existing defenses. Rather than attempting to purify or reconstruct poisoned images, it employs an adversarial detector to reject them outright — eliminating the memorization opportunity entirely.

RejectShield vs. purification-based defense pipeline comparison

Figure 3. Existing purification-based defenses apply image purifiers to poisoned fine-tuning data, but leave the caption intact — the very signal the LVLM memorizes. RejectShield instead uses an adversarial detector to filter out poisoned samples entirely, removing the memorization trigger at its source. No ShadowCast poison data is needed to train the detector.

Model Utility Preservation

RejectShield preserves model utility across both GQA and VizWiz benchmarks under all poison ratios, matching the undefended (No Defense) and clean model baselines. Results follow ShadowCast experimental setups on LLaVA 1.5.

Task Defense Benchmark 0% 1.4% 2.9% 4.3% 5.7%
Trump→Biden No Defense GQA 59.88 59.57 59.53 59.09 59.37
Trump→Biden RejectShield GQA 59.20 59.44 59.32 59.21 59.49
Trump→Biden No Defense VizWiz 56.42 56.22 56.31 55.98 56.43
Trump→Biden RejectShield VizWiz 55.78 55.77 55.83 56.15 55.82
Engine→Fuel No Defense GQA 59.88 59.50 59.74 59.39 59.59
Engine→Fuel RejectShield GQA 59.26 59.19 59.15 59.13 59.17

Data Memorization as a General LVLM Vulnerability

Beyond ShadowCast, our findings reveal data memorization as a fundamental, general vulnerability of LVLMs exposing new attack surfaces that do not require visual perturbations at all.

Memorization-Based Attacks (Earlier Attack Stage)

Adversaries can exploit memorization directly using only standard fine-tuning procedures and benign destination data — no suspicious visual perturbations needed. Such attacks are stealthier and harder to detect without memorization awareness.

🛡

MemDefense: LLM-Powered Monitoring

We propose MemDefense, an LLM-based monitoring tool that analyzes the textual content of fine-tuning datasets for overrepresented concepts (e.g., flagging "the current U.S. president Joe Biden" as anomalous). Together with RejectShield, it provides a more comprehensive safeguard for LVLM fine-tuning pipelines.

Citation

@misc{rejectshield2026,
  title  = {Memory Makes The Poison: Over Memorization Drives Visual Poisoning in {LVLM}s},
  author = {Sy-Tuyen Ho, Yaseema Rusiru Ariyarathna Epa, Yasoda Lasiru Ariyarathna Epa, Andrew Mendez, Huy Nghiem, Quan H. Nguyen, Xudong Jiang, Alex Kot, Furong Huang, Ngai-Man Cheung},
  year   = {2026},
  url    = {https://openreview.net/forum?id=bnWb25IRxx}
}