Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

MEG Decoder Maps Perceived Speech

MEG Decoder Maps Perceived Speech

HuggingFace·Sunday, August 9, 2026
  • •MEG-to-audio model retrieves perceived speech with 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates
  • •Researchers reduced subject-specific representation from 270 to 25 branches and used about 20 times fewer decoder parameters
  • •Paired MEG occlusions found 15 of 19 stimulus features contributed, led by silence and sound intensity
  • •MEG-to-audio model retrieves perceived speech with 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates
  • •Researchers reduced subject-specific representation from 270 to 25 branches and used about 20 times fewer decoder parameters
  • •Paired MEG occlusions found 15 of 19 stimulus features contributed, led by silence and sound intensity
  • •MEG-to-audio model retrieves perceived speech with 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates
  • •Researchers reduced subject-specific representation from 270 to 25 branches and used about 20 times fewer decoder parameters
  • •Paired MEG occlusions found 15 of 19 stimulus features contributed, led by silence and sound intensity
  • •MEG-to-audio model retrieves perceived speech with 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates
  • •Researchers reduced subject-specific representation from 270 to 25 branches and used about 20 times fewer decoder parameters
  • •Paired MEG occlusions found 15 of 19 stimulus features contributed, led by silence and sound intensity

Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova and Alex Ossadtchi reported a MEG-to-audio retrieval model for perceived speech in a paper published on Aug 2 and submitted to Hugging Face Papers on Aug 7. The model retrieves short speech segments from non-invasive magnetoencephalographic recordings, or MEG (brain magnetic-field recording), by training deep networks with a CLIP-style objective against wav2vec2.0 audio embeddings. The paper says earlier high-performing decoders could retrieve perceived speech, but their weights did not map clearly onto electrophysiological quantities, leaving unclear which speech properties drove retrieval.

The researchers redesigned both the frontend and decoder of an existing high-performing MEG-to-audio retrieval architecture. They replaced spatial attention on a flattened sensor layout with spherical harmonics defined on the three-dimensional MEG helmet geometry, reduced the subject-specific representation from 270 to 25 branches, added a temporal filter to each branch to match a neuronal source in space and time, and made the convolutional decoder shallower. Ocular and cardiac components were removed before training to reduce the risk of stimulus-locked shortcuts.

On MEG-MASC, the model reached 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions while using about 20 times fewer decoder parameters. The authors said the weights map to source space and recover generators consistent with the speech-perception network. They also reported that left-lateralized branches carry higher-frequency rhythmic components not evident on the right.

Paired MEG occlusions showed that 15 of 19 stimulus features contributed to retrieval, with the largest effects for silence, sound intensity, vowels and acoustic onsets. Random word lists behaved differently: substituting narrative MEG into them improved retrieval, which the paper says indicates that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target could be reduced to about twelve learned feature dimensions without loss of accuracy, while strong temporal compression caused a clear loss. A paper author summarized the result as a model that is approximately 20x smaller while remaining in the same performance range as prior state-of-the-art systems.

Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova and Alex Ossadtchi reported a MEG-to-audio retrieval model for perceived speech in a paper published on Aug 2 and submitted to Hugging Face Papers on Aug 7. The model retrieves short speech segments from non-invasive magnetoencephalographic recordings, or MEG (brain magnetic-field recording), by training deep networks with a CLIP-style objective against wav2vec2.0 audio embeddings. The paper says earlier high-performing decoders could retrieve perceived speech, but their weights did not map clearly onto electrophysiological quantities, leaving unclear which speech properties drove retrieval.

The researchers redesigned both the frontend and decoder of an existing high-performing MEG-to-audio retrieval architecture. They replaced spatial attention on a flattened sensor layout with spherical harmonics defined on the three-dimensional MEG helmet geometry, reduced the subject-specific representation from 270 to 25 branches, added a temporal filter to each branch to match a neuronal source in space and time, and made the convolutional decoder shallower. Ocular and cardiac components were removed before training to reduce the risk of stimulus-locked shortcuts.

On MEG-MASC, the model reached 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions while using about 20 times fewer decoder parameters. The authors said the weights map to source space and recover generators consistent with the speech-perception network. They also reported that left-lateralized branches carry higher-frequency rhythmic components not evident on the right.

Paired MEG occlusions showed that 15 of 19 stimulus features contributed to retrieval, with the largest effects for silence, sound intensity, vowels and acoustic onsets. Random word lists behaved differently: substituting narrative MEG into them improved retrieval, which the paper says indicates that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target could be reduced to about twelve learned feature dimensions without loss of accuracy, while strong temporal compression caused a clear loss. A paper author summarized the result as a model that is approximately 20x smaller while remaining in the same performance range as prior state-of-the-art systems.

Read original (English)·Aug 9, 2026
#meg#speech decoding#wav2vec2#clip style objective#cortical sources#perceived speech#meg masc#occlusion