ISSN 0439-755X
CN 11-1911/B

Acta Psychologica Sinica ›› 2026, Vol. 58 ›› Issue (11): 2355-2370.doi: 10.3724/SP.J.1041.2026.2355

• Reports of Empirical Studies • Previous Articles     Next Articles

Asymmetry of diagnostic information integration in audiovisual category learning

LIN Yunyi, HUANG Danyang, HUANG Hui, TANG Wenqin, LIU Zhiya   

  1. Center for Studies of Psychological Application / School of Psychology, South China Normal University, Guangzhou 510631, China
  • Received:2025-09-09 Published:2026-11-25 Online:2026-09-11

Abstract: Examining multidimensional visual and auditory category learning, as well as the interactions involved in this process, is important for understanding how individuals integrate information in complex environments and for revealing the cognitive mechanisms underlying natural category learning. Feature diagnosticity, defined as the probability that a feature predicts category membership, is a core construct in category structure and plays a unique role in categorical representation. The present study investigated how diagnostic information from one modality affects multidimensional category learning in another modality and explored the representational strategies underlying this process. We hypothesized that, due to inherent processing differences between the visual and auditory modalities, cross-modal audiovisual integration was asymmetric. Specifically, visual tasks were more likely to show cross-modal integration, whereas auditory tasks were less likely to exhibit such effects.
Two experiments employed a family-resemblance category learning paradigm. In Experiment 1, participants completed a visual classification task with fixed visual feature diagnosticity (0.75), while auditory feature diagnosticity was manipulated across three levels (high: 0.9; congruent: 0.75; low: 0.6) to examine its influence on visual category learning. Experiment 2 adopted the reverse task structure to investigate the influence of visual feature diagnosticity on auditory category learning. A total of 87 and 66 participants were randomly assigned to Experiments 1 and 2, respectively. Dependent measures included accuracy and reaction time during the learning phase, accuracy and the number of acquired dimensions in the dimension test phase, and accuracy in the final test phase. Computational modeling was also conducted to identify participants’ categorization strategies.
In Experiment 1, diagnosticity conditions significantly influenced learners’ processing of cross-modal features. When auditory and visual diagnosticity were congruent, participants showed the highest accuracy on both visual and auditory dimensions, acquired more diagnostic dimensions, and generalized their learning to low-similarity exemplars. These findings demonstrated a “co-frequency enhancement” effect, namely, an audiovisual integration gain under congruent diagnosticity. In contrast, Experiment 2 showed that visual diagnosticity had no significant effect on auditory category learning. Auditory feature acquisition did not differ across diagnosticity conditions, whereas visual feature accuracy increased with diagnosticity. This pattern suggested an “auditory boundary barrier” effect, namely, the relative independence of audiovisual information processing. Cross-experimental comparisons confirmed that auditory category learning was less accurate overall than visual category learning. However, no significant differences were found between the two modalities in attention span or learning outcomes. Both tasks were primarily characterized by prototype-based strategies, although visual tasks included more users of hybrid strategies, while auditory tasks included more random responders.
The present study reveals an asymmetric mechanism of audiovisual interaction in category learning. Visual tasks can integrate auditory information relatively thoroughly, whereas auditory tasks, despite being able to effectively attend to visual information, show limited cross-modal integration. This study provides the first evidence regarding the learning efficiency and representation strategies of multidimensional auditory category learning and discusses differences in representational mechanisms between visual and auditory category learning. Future research could use larger sample sizes and optimized experimental designs to further examine the robustness of these effects.

Key words: category learning, family resemblance, feature diagnosticity, audiovisual integration, categorical representations