跳到主要内容

研究与发表

语音端的误差沿这条路回到解码器训练因此不需要成对的神经信号与语音标注

横向拖动查看完整流程 →

点任意一站看这一步在做什么

01皮层信号ECoG

电极阵列贴在皮层表面,记录说话时的高伽马活动。每位受试者的电极位置都不同,这是这类工作首先要处理的麻烦。

图示按A neural speech decoding framework leveraging deep learning and speech synthesis(Nature Machine Intelligence, 2024)的框架重画,参数名与数值取自原文。

陈旭鹏

Xupeng ChenCTO

Ph.D., Electrical Engineering, NYU

NYU Tandon School of Engineering · Advisors: Yao Wang, Adeen Flinker

Google Scholar

Top 5% of all research outputs scored by Altmetric

论文发表

研究图谱

点图上任意一点跳转至对应文字

横向拖动查看完整图谱 →

2024神经信号与语音

A neural speech decoding framework leveraging deep learning and speech synthesis

Nature Machine Intelligence · Vol. 6, No. 4, pp. 467–480

  • 期刊与会议 5 篇
  • 预印本 7 篇
  • 获奖或高关注

12 篇 · 2020 到 2026 · 三条线

  1. 2024Top 5% of all research outputs (Altmetric)

    A neural speech decoding framework leveraging deep learning and speech synthesis

    Xupeng Chen, Ran Wang, Amirhossein Khalilian-Gourtani, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, Adeen Flinker

    Nature Machine IntelligenceVol. 6, No. 4, pp. 467–480doi:10.1038/s42256-024-00824-8

    An ECoG decoder that translates cortical signals into interpretable speech parameters, paired with a differentiable speech synthesizer. Reproducible across 48 participants; the 3D ResNet decoder reaches PCC 0.804 against the original spectrogram, and stays high under causal-only operation as required for real-time neural prostheses.

    ECoG 解码器输出语音参数,经可微语音合成器还原为频谱,与原始语音频谱比对后反向传播的框架图

    横向拖动查看完整图版 →

    Fig. 1整个框架的闭环:ECoG 解码器把皮层高伽马信号译成一组可解释的语音参数,可微语音合成器再把这组参数还原成频谱。合成器可微,于是语音端的损失能一路回传到解码器 —— 训练不需要成对的"神经信号—语音"标注去硬对齐。原图 ↗·CC BY 4.0
    原始语音频谱与解码得到的频谱逐词对照,八个英文单词的谐波结构一一对应

    横向拖动查看完整图版 →

    Fig. 2c–d上行是受试者真实说出的八个词的频谱,下行是仅凭皮层信号解码出来的频谱。谐波结构、起止时刻、清浊变化都对得上 —— 48 名受试者上,3D ResNet 解码器与原始频谱的相关系数达到 0.804。原图 ↗·CC BY 4.0
    左侧大脑皮层热力图,显示遮蔽各区域后解码相关系数的下降幅度,热区集中于感觉运动皮层与颞上回

    横向拖动查看完整图版 →

    Fig. 4遮住某块皮层,解码质量掉多少 —— 掉得越多,颜色越深。热区集中在感觉运动皮层与颞上回,两种解码器、因果与非因果设定下都稳定复现。这条结论不依赖模型选择,说明模型学到的是生理上说得通的东西。原图 ↗·CC BY 4.0
  2. 2026

    Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

    Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li

    arXiv preprintarXiv:2604.27720

    Asks why telling a vision-language model where to look often makes its answer worse. Benchmarks frontier VLMs on medical grounding, separates the two failure modes — the box lands in the wrong place, or the box is right but cropping to it throws away the context the answer needed — and shows what fine-tuning does and does not fix.

  3. 2026

    Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering

    Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang

    arXiv preprintarXiv:2604.27724

    MedVRAG retrieves over ~350K page images from the literature rather than over plain text, then decides for itself whether the evidence it pulled is enough — and goes back for more if it is not. Most questions settle in one round; the hard ones are exactly the ones that need a second and third.

  4. 2026

    Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

    Chenqian Le, Ruisi Li, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Amirhossein Khalilian-Gourtani, Tianyu He, Adeen Flinker, Yao Wang

    arXiv preprintarXiv:2604.18920

    Surface EMG from the face and neck, compared across spoken, mimed and silently imagined speech. Asks how much of the articulatory signal survives when no sound is produced — the question any silent-speech interface has to answer before it can work.

  5. 2025

    VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

    Chenqian Le, Yilin Zhao, Nikasadat Emami, Kushagra Yadav, Xujin "Chris" Liu, Xupeng Chen, Yao Wang

    arXiv preprintarXiv:2509.09015

    Decodes what a person is looking at from fMRI, with one model shared across subjects instead of one model per brain. A token-merging encoder maps voxels into a CLIP embedding space, so adding a new subject costs a fraction of the parameters a per-subject model would.

  6. 2025

    Machine Learning-Based Prediction of Speech Arrest During Direct Cortical Stimulation Mapping

    Nikasadat Emami, Amirhossein Khalilian-Gourtani, Jianghao Qian, Antoine Ratouchniak, Xupeng Chen, Yao Wang, Adeen Flinker

    arXiv preprintarXiv:2509.08703

    Predicts which cortical sites will arrest speech when stimulated — the mapping a surgeon currently has to establish by stimulating the awake patient site by site. Same signal, read ahead of time instead of found the hard way.

  7. 2025

    When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

    Xiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen, Ziji Zhang, Yingying Zhuang, Narayanan Sadagopan, Anurag Beniwal

    arXiv preprintarXiv:2505.11423

    Chain-of-thought prompting is supposed to make models better. On instruction following it frequently makes them worse: the model reasons its way past a constraint it would have obeyed if asked directly. Measures where the trade-off bites and what mitigates it.

  8. 2025

    Transformer-based neural speech decoding from surface and depth electrode signals

    Junbo Chen, Xupeng Chen, Ran Wang, Chenqian Le, Amirhossein Khalilian-Gourtani, Erika Jensen, Patricia Dugan, Werner Doyle, Orrin Devinsky, Daniel Friedman, Adeen Flinker, Yao Wang

    Journal of Neural EngineeringVol. 22, No. 1, 016017doi:10.1088/1741-2552/adab21

    SwinTW, a transformer that works with arbitrarily positioned electrodes by using their 3D cortical coordinates rather than a 2D grid. Subject-specific models on low-density 8×8 ECoG reach PCC 0.817 across 43 participants, freeing future speech prostheses from grid-only electrode layouts.

  9. 2024

    R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest

    Xupeng Chen, Zhixin Lai, Kangrui Ruan, Shichu Chen, Jiaxiang Liu, Zuozhu Liu

    arXiv preprintarXiv:2410.20327

    Injects the region a clinician would actually look at straight into the visual token stream, so the model reasons over the lesion instead of over the whole slide. The line of work that the 2026 grounding benchmark above grew out of, and argued with.

  10. 2024

    A corollary discharge circuit in human speech

    Amirhossein Khalilian-Gourtani, Ran Wang, Xupeng Chen, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, Adeen Flinker

    Proceedings of the National Academy of Sciences (PNAS)Vol. 121, No. 50, e2404121121doi:10.1073/pnas.2404121121

    Identifies a reproducible corollary discharge source in ventral speech motor cortex that fires before articulation and predicts how strongly auditory cortex is suppressed during speech — how the brain tells its own voice apart from the world.

  11. 2023

    Distributed feedforward and feedback cortical processing supports human speech production

    Ran Wang, Xupeng Chen, Amirhossein Khalilian-Gourtani, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, Adeen Flinker

    Proceedings of the National Academy of Sciences (PNAS)Vol. 120, No. 42, e2300255120doi:10.1073/pnas.2300255120

    Uses causal, anticausal and noncausal convolutional architectures to disentangle motor control from sensory processing during speech, revealing a mixed feedforward/feedback recruitment across perisylvian cortex.

  12. 2020Best Paper Award Finalist

    Stimulus Speech Decoding From Human Cortex with Generative Adversarial Network Transfer Learning

    Ran Wang, Xupeng Chen, Amirhossein Khalilian-Gourtani, Zhaoxi Chen, Leyao Yu, Adeen Flinker, Yao Wang

    IEEE International Symposium on Biomedical Imaging (ISBI)

    Applies GAN transfer learning to decode heard speech from cortical recordings, addressing the scarcity of paired neural-speech data.

页面上的 3 张原图取自上列论文,均为开放获取、以 CC BY 4.0 授权,署名与许可链接见各图图注。图版按版面宽度缩放,个别图只取了其中的子面板,像素颜色未作改动。其余论文的图版权归出版方,未获授权不转载。

从论文到产品

学术研究与落地

01

留下没说出的话

研究

CTO 和首席科学家陈旭鹏在博士期间的研究聚焦于从神经信号到语音解码,解码器输出可解释的语音输出并且稳定复现。我们也会聚焦于还原人本身未发声时,脑内已经形成的想法。这也融入了维度之门的技术基因。

落到产品上

这正是我们那句「每秒八比特」的出处。今天的产品还够不到皮层,但方向是同一个:MinuteX 先把说出口的那部分完整留住,再往说得含糊、说了一半、当时没来得及说的地方推。

对应论文Nature Machine Intelligence 2024Journal of Neural Engineering 2025PNAS 2024arXiv:2604.18920arXiv:2509.08703

02

让大模型耳聪目明

研究

近两年这条线在做多模态:模型答对了,不代表它看对了地方。我们把「看哪里」和「答什么」拆开分别评估,发现前沿模型在医学影像上经常框全错却答得很像样;检索也从纯文本换成对页面图像检索,让证据能被指着看。

落到产品上

任何一条结论,只要是从图像里读出来的,就绕不开同一个问题:它必须能指回自己究竟看了哪一块。这就是产品里 evidence 溯源那套东西的来历。不能溯源的结论,我们不让它出现在界面上。

对应论文arXiv:2604.27720arXiv:2604.27724arXiv:2410.20327arXiv:2509.09015

03

会推理,不等于听得懂话

研究

我们希望让模型在回答之前再想一想,但这种情况会让指令遵循变得更差。这篇文章揭示了模型的推理时间过长会推翻本来的约束,以及如何缓解模型过度推理带来的认知负荷。

落到产品上

多智能体系统中,某个 agent 会自作主张,错误顺着链路一路传下去。所以我们把校正做成显式动作(CONFIRM / CORRECT / DISMISS / UPSERT)。这是我们把前沿研究用在产品实践的结果。

对应论文arXiv:2505.11423