Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions
Paper โข 2606.24082 โข Published
How to use Lab-MSP/comparative-reasoning-sft with PEFT:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Omni-3B")
model = PeftModel.from_pretrained(base_model, "Lab-MSP/comparative-reasoning-sft")The SFT model from Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions (Interspeech 2026). Given two utterances, it answers which one has higher arousal, valence, or dominance.
This repository contains a LoRA adapter (rank 64, alpha 128, all linear layers) for Qwen/Qwen2.5-Omni-3B.
| Training | SFT, lr 1e-4, total batch size 32 |
| Training data | 10k MSP-Podcast v2.0 preference pairs per attribute (30k total), target: answer only |
| MSP-Podcast test (A / V / D / Avg) | 88.1 / 87.8 / 86.7 / 87.5 |
| BIIC-Podcast / WHiSER (Avg) | 76.0 / 89.8 |
Merge the adapter and evaluate with the code repository:
swift export --adapters Lab-MSP/comparative-reasoning-sft --merge_lora true \
--model Qwen/Qwen2.5-Omni-3B --output_dir outputs/experiments/sft_merged
bash src/eval.sh msp_test outputs/experiments/sft_merged
Prompt format (two audios followed by the question):
system: You are a helpful assistant for emotion comparative reasoning.
user: <audio><audio>You will hear two audio clips. Clip 1 is the first audio clip. Clip 2 is the second audio clip. <attribute definition> Which clip has higher <attribute>?
The model answers <answer> Clip 1 </answer> or ... Clip 2 ....
@inproceedings{naini2026comparative,
title = {Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions},
author = {Naini, Abinay Reddy and Kim, Jaeyeon and Yang, Chao-Han Huck and Watanabe, Shinji and Busso, Carlos},
booktitle = {Interspeech},
year = {2026}
}
Base model
Qwen/Qwen2.5-Omni-3B