Mowa widziana wzmacnia rozróżnianie fonemów w górnym zwoju skroniowym człowieka
Visual speech enhances phoneme separability in human superior temporal gyrus
W skrócie
[Preprint - wstępne wyniki] Badacze analizowali aktywność mózgu pacjentów z epilepsją, którzy słuchali lub widzieli słowa wymawiane (np. czytali z ruchu warg). Odkryli, że widzenie ruchu warg znacznie ułatwia mózgowi rozróżnianie poszczególnych dźwięków mowy (fonemów), szczególnie na początku wyrazów. Wynika z tego, że połączenie słuchu i wzroku powoduje szybsze i dokładniejsze rozpoznawanie słów poprzez poprawę przetwarzania fonemów w mózgu.
Oryginalny abstract (angielski)
Visual speech, such as lipreading, facilitates spoken word recognition, but the neural mechanisms underlying audiovisual speech perception remain poorly understood. Visual cues may disambiguate fine-grained articulatory features during early perceptual stages or instead integrate with speech at more categorical, phoneme-level stages. To test how speech representations are modulated by visual input, we analyzed intracranial electroencephalography (iEEG) signals recorded from 12 epilepsy patients performing an audiovisual speech perception task. Participants perceived 16 monosyllabic words presented in auditory-only, visual-only, or congruent audiovisual formats. Words were constructed from four onset consonants (/b/, /g/, /m/, /n/) and four rimes (vowel nucleus and any coda consonants). We examined event-related potentials (ERP) in superior temporal gyrus (STG) and trained support vector machine (SVM) classifiers to decode word identity from neural activity at individual electrodes. Discrete and continuous confusion matrices captured complementary changes in classification accuracy and normalized inverse classification loss, a continuous proxy for classifier confidence. Decoding performance was hierarchically evaluated at the word, phoneme, and phonetic feature levels to determine the representations affected by visual speech. Congruent audiovisual speech increased classifier confidence for phoneme-level representations and improved decoding accuracy at both the word and phoneme level, without corresponding effects on phonetic features. Time-resolved analyses further revealed earlier successful decoding for audiovisual than auditory-only speech, with audiovisual enhancement primarily observed for onset consonants rather than rimes. Together, these findings suggest that visual speech sharpens primarily categorical phoneme representations in STG, with accelerated speech processing and improved word recognition emerging as downstream consequences of phoneme-level enhancement.