
MMAligner Raises Refusal of Unsafe Multimodal Inputs to 99 Percent
MMAligner studies why multimodal inputs can bypass safeguards and proposes representation calibration that moves unsafe inputs back into the model's refusal region.
MMAligner Raises Refusal of Unsafe Multimodal Inputs to 99 Percent
A multimodal model may refuse an unsafe text request yet answer a semantically equivalent request delivered through an image or a text-image combination. The MMAligner paper, submitted to arXiv on August 6, 2026, investigates why this safety gap occurs.
The problem is representation alignment
The researchers found that safety mechanisms learned from text remain present, but unsafe multimodal inputs can shift into internal representations outside the model's refusal boundary. MMAligner calibrates those representations back toward the refusal region while attempting to preserve benign inputs.
According to the paper, the method raised the average refusal rate for unsafe multimodal inputs to 99 percent across multiple open-source models while keeping utility degradation below 2 percent. The paper is scheduled to appear at ACM CCS 2026 and offers a notable direction for repairing multimodal safety inside model representations rather than relying only on external filters.