الانتقال إلى المحتوى الرئيسي

الأبحاث والمنشورات

يعود الخطأ من جانب الكلام إلى وحدة فك التشفير على طول هذا المساروبالتالي فإن التدريب لا يتطلب أزواجًا من الإشارات العصبية وشروح الكلام

اسحب أفقيًا لعرض العملية الكاملة →

انقر على أي توقف لمعرفة ما تفعله هذه الخطوة

01إشارة القشريةECoG

يتم توصيل مجموعة من الأقطاب الكهربائية على سطح القشرة وتسجل نشاط غاما العالي أثناء الكلام. يختلف موقع الأقطاب الكهربائية لكل موضوع، وهي إحدى المشكلات الأولى التي يجب التعامل معها في هذا النوع من العمل.

زر الأيقونةA neural speech decoding framework leveraging deep learning and speech synthesisتمت إعادة رسم إطار عمل (Nature Machine Intelligence, 2024)، وأخذت أسماء المعلمات وقيمها من النص الأصلي.

تشن شيوبنغ

Xupeng ChenCTO

Ph.D., Electrical Engineering, NYU

NYU Tandon School of Engineering · Advisors: Yao Wang, Adeen Flinker

Google Scholar

Top 5% of all research outputs scored by Altmetric

المنشورات العلمية

خريطة الأبحاث

انقر على أي نقطة في الصورة للانتقال إلى النص المقابل

اسحب أفقيًا لعرض الخريطة الكاملة →

2024الإشارات العصبية والكلام

A neural speech decoding framework leveraging deep learning and speech synthesis

Nature Machine Intelligence · Vol. 6, No. 4, pp. 467–480

  • المجلات والمؤتمرات 5 مقالات
  • 7 مطبوعات أولية
  • الحائز على جائزة أو رفيعة المستوى

12 مقالة · 2020 إلى 2026 · ثلاثة أسطر

  1. 2024Top 5% of all research outputs (Altmetric)

    A neural speech decoding framework leveraging deep learning and speech synthesis

    Xupeng Chen, Ran Wang, Amirhossein Khalilian-Gourtani, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, Adeen Flinker

    Nature Machine IntelligenceVol. 6, No. 4, pp. 467–480doi:10.1038/s42256-024-00824-8

    An ECoG decoder that translates cortical signals into interpretable speech parameters, paired with a differentiable speech synthesizer. Reproducible across 48 participants; the 3D ResNet decoder reaches PCC 0.804 against the original spectrogram, and stays high under causal-only operation as required for real-time neural prostheses.

    يقوم جهاز فك ترميز ECoG بإخراج معلمات الكلام، والتي يتم استعادتها إلى الطيف بواسطة مركب الكلام القابل للتمييز، ويتم مقارنتها مع طيف الكلام الأصلي ثم يتم نشرها بشكل عكسي.

    اسحب أفقيًا لعرض الصورة كاملة →

    Fig. 1الحلقة المغلقة للإطار بأكمله: يقوم جهاز فك ترميز ECoG بترجمة إشارة جاما القشرية العالية إلى مجموعة من معلمات الكلام القابلة للتفسير، ويقوم مُركِّب الكلام القابل للتمييز باستعادة هذه المجموعة من المعلمات إلى نطاق. إن جهاز النطق قابل للتمييز، لذا يمكن نشر الخسارة في جانب الكلام وصولاً إلى وحدة فك التشفير - لا يتطلب التدريب محاذاة صارمة لتعليقات "الإشارة العصبية للكلام" المقترنة.الصورة الأصلية ↗·CC BY 4.0
    تتم مقارنة طيف الكلام الأصلي والطيف الذي تم فك تشفيره كلمة بكلمة، وتتوافق الهياكل التوافقية للكلمات الإنجليزية الثماني مع كلمة واحدة.

    اسحب أفقيًا لعرض الصورة كاملة →

    Fig. 2c–dالصف العلوي هو طيف الكلمات الثماني التي يتحدث بها الشخص بالفعل، والصف السفلي هو الطيف الذي تم فك شفرته من الإشارة القشرية وحدها. البنية التوافقية، ولحظات البداية والنهاية، والتغيرات في الوضوح والتعكر كلها متسقة - في 48 موضوعًا، وصل معامل الارتباط بين وحدة فك ترميز ResNet ثلاثية الأبعاد والطيف الأصلي إلى 0.804.الصورة الأصلية ↗·CC BY 4.0
    تُظهر الخريطة الحرارية للقشرة الدماغية اليسرى الانخفاض في معامل ارتباط فك التشفير بعد إخفاء كل منطقة. وتتركز المناطق الساخنة في القشرة الحسية الحركية والتلفيف الصدغي العلوي.

    اسحب أفقيًا لعرض الصورة كاملة →

    Fig. 4إن تغطية منطقة معينة من القشرة ستؤدي إلى تقليل جودة فك التشفير - فكلما زاد فقدانها، أصبح اللون أغمق. تركزت المناطق الساخنة في القشرة الحسية الحركية والتلفيف الصدغي العلوي، وتم تكاثرها بشكل ثابت في كل من أجهزة فك التشفير، والإعدادات السببية وغير السببية. لا يعتمد هذا الاستنتاج على اختيار النموذج، مما يشير إلى أن النموذج تعلم شيئًا منطقيًا من الناحية الفسيولوجية.الصورة الأصلية ↗·CC BY 4.0
  2. 2026

    Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

    Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li

    arXiv preprintarXiv:2604.27720

    Asks why telling a vision-language model where to look often makes its answer worse. Benchmarks frontier VLMs on medical grounding, separates the two failure modes — the box lands in the wrong place, or the box is right but cropping to it throws away the context the answer needed — and shows what fine-tuning does and does not fix.

  3. 2026

    Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering

    Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang

    arXiv preprintarXiv:2604.27724

    MedVRAG retrieves over ~350K page images from the literature rather than over plain text, then decides for itself whether the evidence it pulled is enough — and goes back for more if it is not. Most questions settle in one round; the hard ones are exactly the ones that need a second and third.

  4. 2026

    Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

    Chenqian Le, Ruisi Li, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Amirhossein Khalilian-Gourtani, Tianyu He, Adeen Flinker, Yao Wang

    arXiv preprintarXiv:2604.18920

    Surface EMG from the face and neck, compared across spoken, mimed and silently imagined speech. Asks how much of the articulatory signal survives when no sound is produced — the question any silent-speech interface has to answer before it can work.

  5. 2025

    VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

    Chenqian Le, Yilin Zhao, Nikasadat Emami, Kushagra Yadav, Xujin "Chris" Liu, Xupeng Chen, Yao Wang

    arXiv preprintarXiv:2509.09015

    Decodes what a person is looking at from fMRI, with one model shared across subjects instead of one model per brain. A token-merging encoder maps voxels into a CLIP embedding space, so adding a new subject costs a fraction of the parameters a per-subject model would.

  6. 2025

    Machine Learning-Based Prediction of Speech Arrest During Direct Cortical Stimulation Mapping

    Nikasadat Emami, Amirhossein Khalilian-Gourtani, Jianghao Qian, Antoine Ratouchniak, Xupeng Chen, Yao Wang, Adeen Flinker

    arXiv preprintarXiv:2509.08703

    Predicts which cortical sites will arrest speech when stimulated — the mapping a surgeon currently has to establish by stimulating the awake patient site by site. Same signal, read ahead of time instead of found the hard way.

  7. 2025

    When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

    Xiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen, Ziji Zhang, Yingying Zhuang, Narayanan Sadagopan, Anurag Beniwal

    arXiv preprintarXiv:2505.11423

    Chain-of-thought prompting is supposed to make models better. On instruction following it frequently makes them worse: the model reasons its way past a constraint it would have obeyed if asked directly. Measures where the trade-off bites and what mitigates it.

  8. 2025

    Transformer-based neural speech decoding from surface and depth electrode signals

    Junbo Chen, Xupeng Chen, Ran Wang, Chenqian Le, Amirhossein Khalilian-Gourtani, Erika Jensen, Patricia Dugan, Werner Doyle, Orrin Devinsky, Daniel Friedman, Adeen Flinker, Yao Wang

    Journal of Neural EngineeringVol. 22, No. 1, 016017doi:10.1088/1741-2552/adab21

    SwinTW, a transformer that works with arbitrarily positioned electrodes by using their 3D cortical coordinates rather than a 2D grid. Subject-specific models on low-density 8×8 ECoG reach PCC 0.817 across 43 participants, freeing future speech prostheses from grid-only electrode layouts.

  9. 2024

    R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest

    Xupeng Chen, Zhixin Lai, Kangrui Ruan, Shichu Chen, Jiaxiang Liu, Zuozhu Liu

    arXiv preprintarXiv:2410.20327

    Injects the region a clinician would actually look at straight into the visual token stream, so the model reasons over the lesion instead of over the whole slide. The line of work that the 2026 grounding benchmark above grew out of, and argued with.

  10. 2024

    A corollary discharge circuit in human speech

    Amirhossein Khalilian-Gourtani, Ran Wang, Xupeng Chen, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, Adeen Flinker

    Proceedings of the National Academy of Sciences (PNAS)Vol. 121, No. 50, e2404121121doi:10.1073/pnas.2404121121

    Identifies a reproducible corollary discharge source in ventral speech motor cortex that fires before articulation and predicts how strongly auditory cortex is suppressed during speech — how the brain tells its own voice apart from the world.

  11. 2023

    Distributed feedforward and feedback cortical processing supports human speech production

    Ran Wang, Xupeng Chen, Amirhossein Khalilian-Gourtani, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, Adeen Flinker

    Proceedings of the National Academy of Sciences (PNAS)Vol. 120, No. 42, e2300255120doi:10.1073/pnas.2300255120

    Uses causal, anticausal and noncausal convolutional architectures to disentangle motor control from sensory processing during speech, revealing a mixed feedforward/feedback recruitment across perisylvian cortex.

  12. 2020Best Paper Award Finalist

    Stimulus Speech Decoding From Human Cortex with Generative Adversarial Network Transfer Learning

    Ran Wang, Xupeng Chen, Amirhossein Khalilian-Gourtani, Zhaoxi Chen, Leyao Yu, Adeen Flinker, Yao Wang

    IEEE International Symposium on Biomedical Imaging (ISBI)

    Applies GAN transfer learning to decode heard speech from cortical recordings, addressing the scarcity of paired neural-speech data.

الصور الثلاث الأصلية الموجودة على الصفحة مأخوذة من الأوراق المذكورة أعلاه. جميعها مفتوحة الوصول ومرخصة بموجب CC BY 4.0. يمكن العثور على رابط التوقيع والترخيص في وسيلة إيضاح كل شخصية. يتم تغيير حجم لوحة الصورة وفقًا لعرض التخطيط. بالنسبة لبعض الصور، يتم التقاط اللوحات الفرعية فقط، ولا يتغير لون البكسل. حقوق الطبع والنشر للأوراق المتبقية مملوكة للناشر ولا يجوز إعادة إنتاجها دون إذن.

من الأوراق إلى المنتجات

البحث الأكاديمي والتنفيذ

01

اترك الكلمات التي لم تقال

الأبحاث

ركز CTO والعالم الرئيسي Xupeng Chen خلال الدكتوراه على فك ترميز الكلام من الإشارات العصبية، مع مخرجات كلامية قابلة للتفسير ومستقرة التكرار. ونواصل الاتجاه نفسه: استعادة الأفكار التي تتشكل قبل أن ينطق الإنسان. هذا جزء من الحمض التقني لبوابة الأبعاد.

تقع على المنتج

ومن هنا تأتي عبارة "ثمانية بتات في الثانية". لا يمكن لمنتجات اليوم أن تصل إلى القشرة، لكن الاتجاه هو نفسه: يحتفظ MinuteX أولاً بالجزء الذي يقال بالكامل، ثم يدفعه إلى المكان الغامض، ونصف القول، ولم يكن لديه الوقت لقوله في ذلك الوقت.

الورق المقابلNature Machine Intelligence 2024Journal of Neural Engineering 2025PNAS 2024arXiv:2604.18920arXiv:2509.08703

02

اجعل آذان وعين النموذج الكبير واضحة

الأبحاث

في العامين الماضيين، كان هذا الخط يقوم بعمل طرق متعددة: لمجرد أن النموذج أجاب على السؤال بشكل صحيح، فهذا لا يعني أنه نظر إلى المكان الصحيح. قمنا بفصل وتقييم "أين ننظر" و"ما يجب الإجابة عليه" بشكل منفصل، ووجدنا أن النماذج المتطورة غالبًا ما تجيب على أسئلة لائقة على الرغم من أن جميع المربعات كانت خاطئة في الصور الطبية؛ تم تغيير البحث أيضًا من استرجاع النص العادي إلى استرجاع صور الصفحة، بحيث يمكن الإشارة إلى الأدلة.

تقع على المنتج

أي استنتاج، طالما أنه يُقرأ من الصورة، لا يمكنه تجنب نفس المشكلة: يجب أن يكون قادرًا على الإشارة إلى الجزء الذي نظر إليه المرء. هذا هو أصل شيء تتبع الأدلة في المنتج. نحن لا نسمح بظهور الاستنتاجات التي لا يمكن تتبعها على الواجهة.

الورق المقابلarXiv:2604.27720arXiv:2604.27724arXiv:2410.20327arXiv:2509.09015

03

القدرة على التفكير لا تعني القدرة على فهم الكلمات

الأبحاث

نود أن نجعل النموذج يفكر مرة أخرى قبل الإجابة، ولكن هذا الوضع سيجعل متابعة التعليمات أسوأ. تكشف هذه المقالة أنه إذا كان وقت الاستدلال للنموذج طويلاً جدًا، فسوف يقلب القيود الأصلية، وكيفية تخفيف العبء المعرفي الناجم عن الاستدلال الزائد للنموذج.

تقع على المنتج

في نظام متعدد الوكلاء، قد يتصرف وكيل واحد من تلقاء نفسه ويمرر الخطأ عبر السلسلة. لذلك جعلنا التصحيح فعلا صريحا: CONFIRM وCORRECT وDISMISS وUPSERT. هذه هي الطريقة التي تتحول بها الأبحاث المتقدمة إلى ممارسة منتجية.

الورق المقابلarXiv:2505.11423