ISSN 1671-3710
CN 11-4766/R
主办:中国科学院心理研究所
出版:科学出版社

心理科学进展 ›› 2026, Vol. 34 ›› Issue (12): 2309-2320.doi: 10.3724/SP.J.1042.2026.2309 cstr: 32111.14.2026.2309

• 研究前沿 • 上一篇    下一篇

从胎儿到新生儿:语音感知发展的连续性视角

王梦寰, 许秋阳, 梁丹丹   

  1. 南京师范大学文学院, 南京 210097
  • 收稿日期:2025-12-12 出版日期:2026-12-15 发布日期:2026-09-30
  • 通讯作者: 梁丹丹, E-mail: 03275@njnu.edu.cn
  • 基金资助:
    2022年国家社会科学基金重点项目(22AYY013)资助

From fetus to newborn: A continuity perspective on the development of speech perception

WANG Menghuan, XU Qiuyang, LIANG Dandan   

  1. School of Chinese Language and Culture, Nanjing Normal University, Nanjing 210097, China
  • Received:2025-12-12 Online:2026-12-15 Published:2026-09-30

摘要: 人类新生儿在出生数小时内即可表现出对母语、母亲声音及语音细节的敏感, 其发展基础可能在宫内已经形成。本文从发展神经科学视角出发, 梳理宫内声学环境、胎儿听觉系统成熟及产前听觉经验对新生儿语音感知的影响。现有研究表明, 宫内低通滤波使胎儿主要接触并记忆母语相关的韵律模式, 且这些经验可延续至出生后早期; 出生后, 全频段语音输入与皮层成熟共同促进音段加工, 其神经机制涉及双侧颞-额叶网络及母语相关神经振荡活动的变化。基于此, 本文在既有围产期连续性视角的基础上构建语音感知发展的产前-产后连续性框架, 认为新生儿语音感知是在听觉系统成熟、产前韵律经验与产后语音输入共同作用下连续发展的。未来研究需加强多模态与纵向研究, 拓展早产儿、声调语言及多语环境样本, 并探索其在语言发育障碍早期识别与干预中的价值。

关键词: 产前听觉经验, 宫内声学环境, 新生儿语音感知, 韵律加工, 神经机制

Abstract: Newborns show remarkable sensitivity to speech within the first hours and days of life. They can distinguish their native language from unfamiliar languages, prefer their mother's voice, respond to prosodic and emotional features of speech, and show early sensitivity to vowel and consonant contrasts. These findings have often been interpreted as evidence for broad biological preparedness or as the beginning of rapid postnatal language learning. This review adopts a prenatal-postnatal continuity perspective and asks how prenatal auditory experience, especially low-frequency prosodic experience in the womb, may contribute to the emergence of newborn speech perception.
The review first considers the physiological and acoustic conditions under which prenatal speech experience occurs. The intrauterine environment functions as a natural low-pass filter. High-frequency components of speech, which are crucial for fine-grained consonantal distinctions, are strongly attenuated by maternal tissues, the uterine wall, and amniotic fluid. By contrast, low-frequency information such as rhythm, stress, intonation, pitch contour, and the temporal envelope of speech is relatively well preserved. Thus, the fetus does not encounter speech as a full-spectrum acoustic signal. Prenatal speech input is mainly composed of prosodic and rhythmic information, with only limited access to segmental cues, especially low-frequency vowel-related information. This acoustic constraint defines what can reasonably be learned before birth and also helps explain why prosody occupies a privileged position in fetal and neonatal speech perception.
The review then synthesizes behavioral, physiological, and neuroimaging evidence suggesting that prenatal auditory learning is not limited to general sound exposure. Studies using fetal heart rate, movement responses, fetal magnetoencephalography, electroencephalography, and functional near-infrared spectroscopy indicate that late-gestation fetuses and newborns can respond to frequency changes, rhythm, familiar voices, and repeatedly presented speech materials. After birth, newborns show preferences for the maternal voice, the native language, and speech passages heard prenatally. These findings suggest that prenatal auditory experience can generate memory traces that carry over into the neonatal period. Because prosodic information is the most stable and accessible component of speech in the womb, such learning is most plausibly organized around rhythm, stress, intonation, and other slow temporal properties of speech.
On this basis, the review develops a speech-perception-specific account of prenatal-postnatal continuity. This account does not assume that fetuses fully acquire phonemes before birth. Instead, it distinguishes several possible roles of prenatal experience. First, prenatal experience may support familiarity with global prosodic patterns of the maternal language, including rhythm, stress, and intonation. Second, low-frequency vowel-related information, especially cues associated with the first formant, may provide a limited prenatal interface for later segmental processing. Third, prenatal prosodic experience may serve as a slow temporal scaffold that helps the newborn brain organize the full-spectrum speech signal after birth. Once the infant is exposed to air-conducted speech, richer acoustic information becomes available, including consonantal details, spectral contrasts, and fine phonemic cues. Under the joint influence of this new input and rapid cortical maturation, the newborn brain may quickly enter a phase of segmental learning and neural tuning.
This framework also highlights the role of neural maturation. The development of speech perception cannot be reduced either to prenatal experience or to postnatal input alone. The fetal auditory system becomes functional during gestation, but cortical and fronto-temporal networks continue to mature rapidly around birth and in early infancy. Prenatal experience provides structured input within the limits of the womb, whereas birth introduces a major acoustic transition from a low-pass-filtered, prosody-dominant environment to a full-spectrum speech environment. The newborn brain must therefore recalibrate its auditory and speech-processing system. In this sense, birth is better understood not as an absolute starting point, but as a transition event that reorganizes the relationship among available input, neural readiness, and learning demands.
By integrating findings on intrauterine acoustics, fetal auditory maturation, prenatal prosodic learning, and newborn segmental development, this review attempts to clarify how prenatal and postnatal factors may jointly shape early speech perception. Compared with discussions that focus mainly on postnatal neonatal or infant speech perception, this perspective extends the developmental window backward into fetal life. Compared with broader accounts of prenatal experience or perinatal development, it focuses specifically on speech perception and on the transition from prenatal prosodic predominance to postnatal segmental development.
Future research should examine this continuity account using longitudinal and multimodal designs. Combining fetal and neonatal measures such as fetal magnetoencephalography, electroencephalography, functional near-infrared spectroscopy, ultrasound-based behavioral indices, and postnatal behavioral paradigms may help distinguish the respective contributions of prenatal exposure, birth-related recalibration, and postnatal speech input. Studies of preterm infants, tone-language environments, bilingual or multilingual exposure, and high-risk populations will be especially informative. Such work may also help identify early neural markers of atypical speech perception and inform early screening and intervention for developmental language disorders.

Key words: prenatal auditory experience, intrauterine acoustic environment, newborn speech perception, prosodic processing, neural mechanisms