{"id":"yu-zhang","hidden":false,"ipfs":"QmaZsGskZs43gbbpkZw89L4cWFJsPUDyQcHLrWsdPm3ymL","language":"en","transactionHash":"0x982b35a51b19512110f145df019e241ea0ef481fe98559aeaeadabee61a546f2","created":"2025-08-21T06:18:26.446Z","updated":"2025-08-21T06:27:23.739Z","title":"Yu Zhang","summary":"Yu Zhang is a Research Scientist at Meta’s Superintelligence team focusing on deep learning for speech recognition and synthesis and large-scale self-supervised and multilingual speech models.","content":"**Yu Zhang** is a research scientist specializing in deep learning for speech, including automatic speech recognition, speech synthesis, and self-supervised and multimodal speech models.[[4]](#cite-id-BAo3S9fuVG) He is a Research Scientist on Meta's [Superintelligence](https://iq.wiki/wiki/meta-superintelligence-team) team and has previously held research roles at Google DeepMind, OpenAI, and Microsoft.[[3]](#cite-id-7LPjfBfXXn)\n\n## Education\nYu Zhang earned a Ph.D. in Computer Science from the Massachusetts Institute of Technology (MIT), where he studied from 2012 to 2017, and a B.S. in Computer Science from Shanghai Jiao Tong University, where he studied from 2005 to 2009.[[3]](#cite-id-7LPjfBfXXn) At MIT he was a member of the Computer Science and Artificial Intelligence Laboratory (CSAIL), conducting research as part of the Spoken Language Systems Group under the supervision of Dr. James Glass. His academic work centered on applying machine learning models to challenges in speech and language processing. In fall 2009, prior to his doctoral studies, he also served as a teaching assistant for a course on Statistical Learning.[[1]](#cite-id-k9y5FmSYyl)[[3]](#cite-id-7LPjfBfXXn)\n\n## Career\nZhang began his career in academic research at MIT's CSAIL, where his work primarily focused on machine learning applications for speech recognition, speaker verification, and language identification. He was an active participant in the IARPA Babel Program, a research initiative aimed at advancing multi-lingual speech recognition capabilities, particularly for low-resource languages.[[1]](#cite-id-k9y5FmSYyl) His research during this period explored the use of advanced deep learning architectures, such as deep neural networks and Recurrent Neural Networks (RNNs), to solve complex problems in speech processing. Specifically, his work investigated techniques like Long Short-Term Memory (LSTM) for distant speech recognition, the extraction of deep neural network bottleneck features for improved acoustic modeling, and the use of i-vector based approaches for normalizing speaker and environmental variability in audio signals.[[1]](#cite-id-k9y5FmSYyl)\n\nAfter his tenure in academia, Zhang transitioned to research roles in the technology industry. He worked as an intern at Microsoft Research Asia from 2010 to 2012 and later as a Research Intern at Microsoft in 2014, contributing to speech and language technology projects.[[3]](#cite-id-7LPjfBfXXn) From 2017 to 2023 he was a Staff Research Scientist at Google DeepMind, where his work included large-scale automatic speech recognition, text-to-speech synthesis, self-supervised and semi-supervised speech learning, and multilingual and multimodal speech–text systems. This work is reflected in publications such as SpecAugment, Conformer, LibriTTS, WaveGrad, w2v-BERT, ContextNet, Google USM, and FLEURS.[[4]](#cite-id-BAo3S9fuVG) He then served as a Member of Technical Staff at OpenAI from October 2023 to July 2025, before joining Meta in July 2025 as a Research Scientist on the [Superintelligence](https://iq.wiki/wiki/meta-superintelligence-team) team, focusing on advancing foundational models for speech and multimodal understanding.[[3]](#cite-id-7LPjfBfXXn)[[4]](#cite-id-BAo3S9fuVG)[[2]](#cite-id-wQ09dXK2H1)\n\n### Major Works and Publications\nThroughout his career, Yu Zhang has co-authored numerous research papers that have been presented at major machine learning and signal processing conferences, including the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) and Interspeech. His publications reflect work on deep learning for speech recognition, feature extraction, and acoustic model training.\n\nA selection of his published works includes:\n\n* **\"Highway Long Short-Term Memory RNNs for Distant Speech Recognition\" (2015):** This paper investigated the application of Highway LSTM networks, a variant of recurrent neural networks, to enhance the accuracy of speech recognition systems when the audio is captured from a distance.\n* **\"Prediction-adaptation-correction Recurrent Neural Networks for Low-resource Language Speech Recognition\" (2015):** This work introduced a specialized RNN architecture designed to improve speech recognition performance for languages with limited training data.\n* **\"Language ID-based Training of Multilingual Stacked Bottleneck Features\" (2014):** This research explored a method for training multilingual acoustic models by using language identification to inform the extraction of stacked bottleneck features from a deep neural network.\n* **\"Extracting deep neural network bottleneck features using low-rank matrix factorization\" (2014):** This publication proposed an efficient method for extracting compact, information-rich bottleneck features from deep neural networks by applying low-rank matrix factorization.\n* **\"Joint Learning of Phonetic Units and Word Pronunciations for ASR\" (2013):** This paper focused on a method for simultaneously learning phonetic units and their corresponding word pronunciations to improve the performance of automatic speech recognition (ASR) systems.\n* **\"A new i-vector approach and its application to irrelevant variability normalization based acoustic model training\" (2011):** This work introduced a novel approach using i-vectors, a low-dimensional representation of audio segments, to normalize for variabilities such as speaker characteristics and channel noise in acoustic model training.\n* **\"An evidence framework for Bayesian learning of continuous-density hidden Markov models\" (2009):** This early work presented a Bayesian framework for learning the parameters of Hidden Markov Models (HMMs), a foundational statistical model used in speech recognition.\n\nHis later research includes contributions to methods and datasets for speech recognition and synthesis, such as SpecAugment for data augmentation, convolution-augmented and context-aware architectures like Conformer and ContextNet, and the LibriTTS text-to-speech corpus. It also covers transfer learning from speaker verification to multi-speaker TTS, WaveNet- and diffusion-based speech generation with works like WaveGrad, self-supervised and semi-supervised speech models including w2v-BERT and large-scale semi-supervised ASR, and multilingual speech systems and benchmarks such as Google USM and FLEURS.[[4]](#cite-id-BAo3S9fuVG) The full list of his publications spans topics in automatic speech recognition, speech synthesis, self-supervised learning, and multilingual and multimodal speech processing.[[1]](#cite-id-k9y5FmSYyl)[[2]](#cite-id-wQ09dXK2H1)[[3]](#cite-id-7LPjfBfXXn)[[4]](#cite-id-BAo3S9fuVG)[[5]](#cite-id-SeIAELJX0p)[[6]](#cite-id-4glueoI0bX)\n\n## Interviews\n\n\n### LTI Colloquium at Carnegie Mellon #01\nOn November 15, 2024, Yu Zhang was a featured speaker at the LTI Colloquium organized by the *Language Technologies Institute at Carnegie Mellon University (LTI at CMU)*. His presentation, titled *“Hearing the AGI: from GMM-HMM to GPT-4o”*, examined the historical development and current directions of speech recognition research.\n $$widget0 [YOUTUBE@VID](https://youtube.com/watch?v=pRUrO0x637A)$$ \nIn his talk, Zhang outlined the progression from early Gaussian Mixture Model–Hidden Markov Model (GMM-HMM) systems to large-scale, multimodal architectures based on self-supervised transformer models. He noted that advances in the field have been driven not only by the expansion of datasets and model size but also by the scaling of computational resources and by overcoming system-level engineering challenges.\n\nAccording to Zhang, self-supervised learning has played a central role in enabling models to utilize large amounts of unlabeled audio, which has expanded the capacity and performance of speech systems. He also observed that speech processing requires substantially more computational power than text, as it must address additional factors such as background noise, silence, and diverse acoustic conditions.\n\nZhang further discussed the shift from automatic speech recognition toward multimodal systems that combine speech, text, and vision. He emphasized that next-token prediction approaches, similar to those used in GPT-style language models, are central to this transition. He also pointed out that traditional metrics such as *Word Error Rate (WER)* do not always reflect human judgments of quality, highlighting the importance of developing more representative evaluation methods.\n\nIn addressing safety and reliability, Zhang remarked that speech models may present unique risks, as their outputs can appear more persuasive when incorrect. He identified alignment, benchmarking, and efficient handling of long-context inputs as ongoing research needs. He concluded by noting that the integration of speech with text and vision is likely to play a major role in the advancement of multimodal systems and their potential contribution to artificial general intelligence, but emphasized that progress depends on both scientific research and practical engineering solutions.[[7]](#cite-id-dfIIibO8Pq)","recentActivity":null,"operator":{"id":"0x1E23b34d3106F0C1c74D17f2Cd0F65cdb039b138"},"categories":[{"id":"people","title":"People in crypto"}],"tags":[{"id":"AI Token"},{"id":"Researcher"}],"images":[{"id":"QmfHCybM56c7PP36mNifysZctStp3VkQymCLF7CnsLAi1B","type":"image/jpeg, image/png"}],"media":[{"name":"1702171103571.jpeg","id":"QmazPC4DJtzUf4jxJbn7YfjvbqmkojBmfhXnin1Tv2e8Xn","size":null,"type":null,"source":"IPFS_IMG"},{"name":"pRUrO0x637A","id":"https://www.youtube.com/watch?v=pRUrO0x637A","size":null,"type":null,"source":"YOUTUBE"}],"linkedWikis":{"founders":[],"blockchains":[]},"events":[{"title":"Intern at Microsoft Research Asia","date":"2010-01-01","type":"DEFAULT","description":"Started an internship at Microsoft Research Asia, working on speech and language technology.","link":"","country":"","continent":"","multiDateStart":null,"multiDateEnd":null,"id":"502cd01d-ebcb-4a25-81ae-60416546cdca"},{"title":"Research Intern at Microsoft","date":"2014-06-01","type":"DEFAULT","description":"Joined Microsoft as a Research Intern, contributing to projects in speech and language processing.","link":"","country":"","continent":"","multiDateStart":null,"multiDateEnd":null,"id":"dea4604c-948b-4e6e-806e-893ae9b33ecd"},{"title":"Staff Research Scientist at Google DeepMind","date":"2017-09-01","type":"DEFAULT","description":"Became a Staff Research Scientist at Google DeepMind, working on large-scale speech recognition and speech synthesis.","link":"","country":"","continent":"","multiDateStart":null,"multiDateEnd":null,"id":"ca6201e2-f769-45d9-8e69-e44e1537a603"},{"id":"a1711070-68a6-4ae6-9fac-853936f0ba91","type":null,"date":"2018-01-01","multiDateStart":null,"multiDateEnd":null,"country":null,"continent":null,"title":"Staff Researcher at DeepMind","description":"Worked as a Staff Researcher at DeepMind, contributing to artificial intelligence research and development.","link":null},{"id":"accb9fdb-da18-41c7-bba9-0d7c6d7b31a4","type":"DEFAULT","date":"2023-10-01","multiDateStart":null,"multiDateEnd":null,"country":"","continent":"","title":"Member of Technical Staff at OpenAI","description":"Joined OpenAI as a Member of Technical Staff (MTS), contributing to speech and multimodal AI systems.","link":""},{"title":"LTI Colloquium talk at CMU: \"Hearing the AGI\"","date":"2024-11-01","type":"DEFAULT","description":"Delivered an invited LTI Colloquium talk at Carnegie Mellon University on the evolution of speech recognition and multimodal models.","link":"https://www.youtube.com/watch?v=pRUrO0x637A","country":"","continent":"","multiDateStart":null,"multiDateEnd":null,"id":"f6fa4863-2770-4f49-ac3e-6893aa6269bb"},{"id":"fc1e6bc7-5b65-4435-a925-3d8440fab80f","type":null,"date":"2025-07-01","multiDateStart":null,"multiDateEnd":null,"country":null,"continent":null,"title":"Joined Meta's Superintelligence Team","description":"Joined Meta as a Software Engineer on the newly formed \"Superintelligence\" team, focusing on backend systems for advanced AI development.","link":null}],"founderWikis":[],"blockchainWikis":[],"metadata":[{"id":"website","value":"https://people.csail.mit.edu/yzhang87/"},{"id":"linkedin_profile","value":"https://www.linkedin.com/in/yu-zhang-07a43016/"},{"id":"email_url","value":"mailto:sjtuzy@gmail.com"},{"id":"words-changed","value":"323"},{"id":"percent-changed","value":"43.74"},{"id":"blocks-changed","value":"content, tags, summary"},{"id":"wiki-score","value":"70"},{"id":"references","value":"[{\"id\":\"k9y5FmSYyl\",\"url\":\"https://people.csail.mit.edu/yzhang87/\",\"description\":\"Yu Zhang's academic profile at MIT CSAIL\",\"timestamp\":1755756984414},{\"id\":\"wQ09dXK2H1\",\"url\":\"https://medium.com/g-able/unpacking-metas-superintelligence-team-a-deep-dive-into-their-ambitious-ai-play-99cc821cffb7\",\"description\":\"Analysis of Meta's Superintelligence team\",\"timestamp\":1755756984414},{\"id\":\"7LPjfBfXXn\",\"description\":\"LinkedIn: Yu Zhang\",\"timestamp\":1755757125809,\"url\":\"https://www.linkedin.com/in/yu-zhang-07a43016/\"},{\"id\":\"BAo3S9fuVG\",\"description\":\"Google Scholar: Yu Zhang\",\"timestamp\":1755757232327,\"url\":\"https://scholar.google.com/citations?user=EilVnKwAAAAJ&hl=en\"},{\"id\":\"SeIAELJX0p\",\"description\":\"It’s Known as ‘The List’—and It’s a Secret File of AI Geniuses\\n\",\"timestamp\":1755757305427,\"url\":\"https://www.wsj.com/tech/meta-ai-recruiting-mark-zuckerberg-openai-018ed7fc\"},{\"id\":\"4glueoI0bX\",\"description\":\"Meet the AI Superstars\",\"timestamp\":1755757340779,\"url\":\"https://timesofindia.indiatimes.com/education/news/meet-the-ai-superstars-who-are-making-more-than-nba-stars/articleshow/122162700.cms\"},{\"id\":\"dfIIibO8Pq\",\"description\":\"November 15th LTI Colloquium Speaker - Yu Zhang\\n\",\"timestamp\":1755757505059,\"url\":\"https://www.youtube.com/watch?v=pRUrO0x637A\"}]"},{"id":"previous_cid","value":"QmaZsGskZs43gbbpkZw89L4cWFJsPUDyQcHLrWsdPm3ymL"},{"id":"commit-message","value":"Expand Yu Zhang summary and timeline (+110w)"},{"id":"previous_cid","value":"QmaZsGskZs43gbbpkZw89L4cWFJsPUDyQcHLrWsdPm3ymL"}],"user":{"id":"0x8af7a19a26d8fbc48defb35aefb15ec8c407f889"},"author":{"id":"0x1E23b34d3106F0C1c74D17f2Cd0F65cdb039b138"},"views":1962}