From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models Through the lens of Social Relationship
Abstract
Ensuring the trustworthiness of Multimodal Large Language Models (MLLMs) requires rigorous evaluation of gender bias, particularly in socially sensitive domains. Existing benchmarks primarily assess bias in isolated contexts, overlooking its subtle emergence through interpersonal interactions. We introduce Genres, a benchmark for evaluating Gender bias in MLLMs through the lens of social Relationships in narrative generation. Genres features dual-character tasks that capture rich interpersonal dynamics and support fine-grained, multidimensional bias assessment. Experiments on both open- and closed-source MLLMs reveal persistent, context-dependent biases that remain undetected in single-character settings. Beyond diagnosis, we further explore two complementary generation-stage mitigation strategies. Genres underscore the importance of relationship-aware benchmarks for diagnosing subtle, interaction-driven gender bias and provide actionable insights for mitigation.