ISSN 0439-755X
CN 11-1911/B
主办:中国心理学会
   中国科学院心理研究所
出版:科学出版社

心理学报 ›› 2026, Vol. 58 ›› Issue (11): 2355-2370.doi: 10.3724/SP.J.1041.2026.2355 cstr: 32110.14.2026.2355

• 研究报告 • 上一篇    下一篇

视听类别学习中的诊断性信息整合:“视”频增益与“听”界壁垒

林蕴旖, 黄丹阳, 黄晖, 唐雯勤, 刘志雅   

  1. 华南师范大学心理应用研究中心/心理学院, 广州 510631
  • 收稿日期:2025-09-09 发布日期:2026-09-10 出版日期:2026-11-25
  • 通讯作者: 刘志雅, E-mail: zhiyaliu@scnu.edu.cn
  • 基金资助:
    广东省脑认知与人的素质发展基础学科研究中心项目(2024B0303390003); 广东省哲学社会科学规划一般项目(GD25CXL03)

Asymmetry of diagnostic information integration in audiovisual category learning

LIN Yunyi, HUANG Danyang, HUANG Hui, TANG Wenqin, LIU Zhiya   

  1. Center for Studies of Psychological Application / School of Psychology, South China Normal University, Guangzhou 510631, China
  • Received:2025-09-09 Online:2026-09-10 Published:2026-11-25

摘要: 本研究采用家族相似性类别学习范式, 探究视听双通道诊断性信息在多维度类别学习中的整合机制及其不对称性。两个实验共招募153名大学生, 结合计算模型分析学习后的类别表征。实验1考察听觉信息诊断性对视觉类别学习的影响, 操纵听觉诊断性为高(0.9)、一致(0.75)和低(0.6)三种水平。结果发现, 当视听诊断性一致时, 学习者对双通道特征维度的加工最为充分, 表现出跨通道整合的增益效应。实验2则考察视觉信息诊断性对听觉类别学习的影响, 发现诊断性一致条件下的整合增益没有出现, 视觉诊断性对听觉特征维度的学习没有显著影响, 体现视听信息加工的相对独立性。综上, 多维度类别学习中的跨通道信息整合存在不对称性:视觉任务可以整合听觉信息, 而听觉任务对视觉信息的整合深度有限。本研究首次揭示了多维度听觉类别学习的学习效率与表征策略, 并讨论了视听类别学习在表征机制上的差异。

关键词: 类别学习, 家族相似性, 特征诊断性, 视听整合, 类别表征

Abstract: Examining multidimensional visual and auditory category learning, as well as the interactions involved in this process, is important for understanding how individuals integrate information in complex environments and for revealing the cognitive mechanisms underlying natural category learning. Feature diagnosticity, defined as the probability that a feature predicts category membership, is a core construct in category structure and plays a unique role in categorical representation. The present study investigated how diagnostic information from one modality affects multidimensional category learning in another modality and explored the representational strategies underlying this process. We hypothesized that, due to inherent processing differences between the visual and auditory modalities, cross-modal audiovisual integration was asymmetric. Specifically, visual tasks were more likely to show cross-modal integration, whereas auditory tasks were less likely to exhibit such effects.
Two experiments employed a family-resemblance category learning paradigm. In Experiment 1, participants completed a visual classification task with fixed visual feature diagnosticity (0.75), while auditory feature diagnosticity was manipulated across three levels (high: 0.9; congruent: 0.75; low: 0.6) to examine its influence on visual category learning. Experiment 2 adopted the reverse task structure to investigate the influence of visual feature diagnosticity on auditory category learning. A total of 87 and 66 participants were randomly assigned to Experiments 1 and 2, respectively. Dependent measures included accuracy and reaction time during the learning phase, accuracy and the number of acquired dimensions in the dimension test phase, and accuracy in the final test phase. Computational modeling was also conducted to identify participants’ categorization strategies.
In Experiment 1, diagnosticity conditions significantly influenced learners’ processing of cross-modal features. When auditory and visual diagnosticity were congruent, participants showed the highest accuracy on both visual and auditory dimensions, acquired more diagnostic dimensions, and generalized their learning to low-similarity exemplars. These findings demonstrated a “co-frequency enhancement” effect, namely, an audiovisual integration gain under congruent diagnosticity. In contrast, Experiment 2 showed that visual diagnosticity had no significant effect on auditory category learning. Auditory feature acquisition did not differ across diagnosticity conditions, whereas visual feature accuracy increased with diagnosticity. This pattern suggested an “auditory boundary barrier” effect, namely, the relative independence of audiovisual information processing. Cross-experimental comparisons confirmed that auditory category learning was less accurate overall than visual category learning. However, no significant differences were found between the two modalities in attention span or learning outcomes. Both tasks were primarily characterized by prototype-based strategies, although visual tasks included more users of hybrid strategies, while auditory tasks included more random responders.
The present study reveals an asymmetric mechanism of audiovisual interaction in category learning. Visual tasks can integrate auditory information relatively thoroughly, whereas auditory tasks, despite being able to effectively attend to visual information, show limited cross-modal integration. This study provides the first evidence regarding the learning efficiency and representation strategies of multidimensional auditory category learning and discusses differences in representational mechanisms between visual and auditory category learning. Future research could use larger sample sizes and optimized experimental designs to further examine the robustness of these effects.

Key words: category learning, family resemblance, feature diagnosticity, audiovisual integration, categorical representations

中图分类号: