JMedEthicBench: A Multi-Turn Adversarial Benchmark for Japanese Medical Ethics Alignment in LLMs
Junyu Liu ⋅ Zirui Li ⋅ Qian Niu ⋅ Zequn Zhang ⋅ Yue Xun ⋅ Wenlong Hou ⋅ Shujun Wang ⋅ Yusuke Iwasawa ⋅ Yutaka Matsuo ⋅ Kan Hatakeyama-Sato
Abstract
As Large Language Models (LLMs) are increasingly deployed in healthcare worldwide, robustness against adversarial jailbreaking becomes essential for maintaining medical ethics compliance, calling for comprehensive adversarial benchmarks. However, existing benchmarks remain predominantly English-centric and limited to single-turn attacks. To address these gaps, we introduce JMedEthicBench, the first multi-turn adversarial benchmark for evaluating medical ethics alignment of LLMs in Japanese. Grounded in 67 guidelines from the Japan Medical Association, our benchmark comprises over 50,000 adversarial conversations generated using seven automatically discovered jailbreak strategies. Through a dual-LLM scoring protocol, we evaluate 22 models and find that commercial models maintain strong resistance to attacks while medical-specialized models exhibit increased vulnerability. Furthermore, safety scores decline significantly across conversation turns (median: 9.5 to 5.5, $p < 0.001$), demonstrating that multi-turn escalation systematically circumvents defense mechanisms. Cross-lingual evaluation reveals that medical model vulnerabilities persist across Japanese and English, indicating inherent alignment limitations rather than language-specific factors. Our findings suggest that domain-specific fine-tuning may inadvertently weaken safety alignment and that multi-turn jailbreak attacks represent a distinct threat surface requiring dedicated defense strategies.
Successful Page Load