CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
Abstract
We introduce CreativityNeuro (CN), a data-free method for enhancing creative behavior in LLMs via contrastive weight steering. We evaluate CN on various creativity assessments and report three main findings: (1) On the Divergent Association Task (DAT), CN improves performance by up to 14 human percentile points, with gains that are consistent across semantically diverse prompts. (2) In a large-scale human evaluation (N=720) on the Alternative Uses Test (AUT) and the Task Task (TT), CN-enhanced models achieve significant improvements in originality, surprise, and creativity, demonstrating transfer to more complex, open-ended settings. Compared to contrastive activation steering (CAA), we find that while activation steering can match CN on the DAT, it does not achieve consistent improvements on held-out tasks (AUT, TT). (3) CN-enhanced models reduce factual reasoning scores on MMLU by 1--11\%, and attempts to explicitly preserve such capabilities by incorporating MMLU prompts into prompt contrast sets reverse nearly all divergent thinking (DT) gains, providing evidence that DT and factual reasoning rely on functionally entangled weights.