| 英文摘要 |
Speech Emotion Recognition (SER) is a key technology within the myriad of speech solutions. A unique fairness issue in SER stems from the inherent emotional perception bias present in the data labels provided by raters. To enhance both recognition performance and fairness in SER, addressing rater bias is paramount. In this study, we propose a two-stage framework. In the first stage, we generate debiased representations using a fairness-constrained adversarial framework. Subsequently, in the second stage, following gender-wise perceptual learning, we empower users with the ability to toggle freely between specific gender-wise perceptions as needed. We utilize two significant fairness metrics to evaluate our results, demonstrating that our distributions and predictions across genders are fair. Further analysis indicates that our model effectively mitigates the influence of gender perspectives in the feature learning space. |