Speech Technology
Статистика- Последний пост
- 14 авг.
- Последнее чтение
- 15 авг.
- Постов за неделю
- 4
- Всего постов
- 21
- Тип
- открытый
- Язык
- und
- Категория
- Познавательное (по похожим)
- В каталоге с
- 12 авг.
- 1/24сутки в ленте
- 419
- 1/48двое суток
- 480
- 1/72трое суток
- 517
Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.
Посты
One more agentic thing https://github.com/InteractiveASR/AgenticASR https://interactiveasr.github.io/ https://arxiv.org/abs/2604.09121 Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Peng Wang, Yanqiao Zhu, Zixuan Jiang, Qinyuan Chen, Xingjian Zhao, Xipeng Qiu, Wupeng Wang, Zhifu Gao, Xiangang Li, Kai Yu, Xie Chen Recent years have witnessed remarkable progress in automatic speech recognition (ASR), driven by advances in model architectures and large-scale training data. However, two important aspects remain underexplored. First, Word Error Rate (WER), the dominant evaluation metric for decades, treats all words equally and often fails to reflect the semantic correctness of an utterance at the sentence level. Second, interactive correction-an essential component of human communication-has rarely been systematically studied in ASR research. In this paper, we integrate these two perspectives under an agentic framework for interactive ASR. We propose leveraging LLM-as-a-Judge as a semantic-aware evaluation metric to assess recognition quality beyond token-level accuracy. Furthermore, we design an LLM-driven agent framework to simulate human-like multi-turn interaction, enabling iterative refinement of recognition outputs through semantic feedback. Extensive experiments are conducted on standard benchmarks, including GigaSpeech (English), WenetSpeech (Chinese), the ASRU 2019 code-switching test set. Both objective and subjective evaluations demonstrate the effectiveness of the proposed framework in improving semantic fidelity and interactive correction capability. We will release the code to facilitate future research in interactive and agentic ASR.
Big paper on state of the art in speech translation (87 pages) Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign https://aclanthology.org/2026.iwslt-1.39.pdf This paper reports on the outcomes of the shared tasks organized as part of the 23rd International Workshop on Spoken Language Translation (IWSLT). The workshop covered ten major challenges in spoken language translation, including speech-to-text translation for both high-resource and low-resource language pairs, customized speech translation, speech generation, instruction-following speech processing, and the evaluation of speech translation systems. The shared tasks received strong participation, with more than 30 teams submitting runs. This year’s edition broadened the range of tasks, placing particular emphasis on speech generation and evaluation metrics.
Duplex paper from today https://github.com/duplexgen/duplexgen-code https://arxiv.org/abs/2607.26178 DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-Tür Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. https://huggingface.co/datasets/sarvamai/indic-diarbench https://www.sarvam.ai/blogs/indic-diarbench
Topic of today - autoregression vs diffusion in TTS First of all the presentation from Meta. Claims diffusion/autoregression decision could be dynamic https://www.youtube.com/watch?v=kMimQxIJLos Second paper from today on the similar topic https://arxiv.org/abs/2607.04140 DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Junwon Moon, Yejin Lee, Seungbeom Kim, Hoseong Ahn, Sewoong Park, Heeseung Kim, Kyuhong Shim Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, since the model receives the full input text before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a 1/t-weighted training objective and a time-shifted inference schedule that together defer low-confidence positions to later steps. Trained on only 585 hours of LibriTTS, DELTA-TTS achieves a 1.75% WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens 3.3x faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates the hallucinations observed in AR generation.
Deepmind also releases something https://github.com/google-deepmind/phasecoder https://arxiv.org/abs/2601.21124 PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only spatial audio encoder that is agnostic to microphone geometry. PhaseCoder takes raw multichannel audio and microphone coordinates as inputs to perform localization and produces robust spatial embeddings. We demonstrate that Gemma 3n LLM can be fine-tuned to reason over "Spatial Audio Tokens" produced by PhaseCoder. We show our encoder achieves state-of-the-art results on microphone-invariant localization benchmarks and, for the first time, enables an LLM to perform complex spatial reasoning and targeted transcription tasks from an arbitrary microphone array.
Piotr is a reincarnation of Lennart https://x.com/PiotrZelasko/status/2085709891604254788
https://www.reddit.com/r/homeassistant/comments/1vhofep/hacked_and_debloated_an_echo_dot_2_local_llm/
There is a big interest in full duplex as I see, here is a nice collection of papers https://github.com/Ruiqi-Yan/Awesome-Full-Duplex-SDM
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B NVIDIA NemotronLabs VoiceChat is a 11B end-to-end, real-time speech full duplex (FD) model for conversational AI that jointly performs streaming speech understanding and speech generation [1, 2]. Unlike traditional cascaded stacks (ASR → LLM → TTS), this model achieves full duplex, real-time, seamless voice interaction in one unified architecture, eliminating the need for multiple models or API handoffs, thus reducing end-to-end latency. It sets new benchmarks by bringing open, robust, and highly natural conversation capabilities. Moreover, NVIDIA NemotronLabs VoiceChat is the first open full-duplex model to support tool calling while maintaining a natural conversation flow during tool execution. For each tool, a specific “on-hold” message can be defined that will be spoken by the agent as soon as the LLM generates the text that will trigger the tool call and response.
Things move on in OpenAI as well. Interesting that voice model is separate. And no turn detector anymore. https://x.com/OpenAI/status/2084378415818579975 Lots of interesting technical details, from realtime inference, to dynamic compaction, to WebRTC optimization. https://openai.com/index/continuous-voice-interaction-with-gpt-live/
One more https://github.com/AnXMuy/AgenticASR
https://huckiyang.github.io/voice-memory/ from NVIDIA https://arxiv.org/abs/2607.26410 Voice Memory for Agentic Speech Recognition Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain this http URL and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The me
https://huggingface.co/nyralabs/CrisperWhisper2.0_large Most speech-to-text systems never actually decide whether to write down what was said or what was meant. They inherit that choice from their training data and apply it inconsistently. CrisperWhisper 2.0 makes it an explicit, controllable choice. One recording, two transcripts: Verbatim, exactly what was said, in one consistent format: [um] so we we need to, to reschedule the th- thursday meeting to [uh] march third at nine thirty [laughter] Intended, the clean version the speaker meant, with numbers, dates, and emails formatted the way you'd write them: So we need to reschedule the Thursday meeting to March 3 at 9:30. On top of that: Word-level timings. Around 30 ms mean boundary error on read speech and 41 ms on conversational speech, the most precise word timing of any system we benchmarked, on both. Verbatimize. Upgrade transcripts you already have: given audio plus a trusted clean transcript, the model reproduces your content word-for-word and inserts only the disfluencies and vocal events actually present in the audio (rare-word recall jumps from 6.8% to 96.1% vs. re-transcribing). This turns the world's abundant clean corpora into verbatim ones, ready for TTS data, clinical speech analysis, and dataset construction. Multilingual. Verbatim and intended modes work across most languages Whisper supports. CrisperWhisper 2.0 tops the Nyra Verbatim Speech Benchmark leaderboard for disfluency F1 across ten languages, ahead of every closed-source alternative we tested. Seamless longform. Audio of any length, transcribed without the usual chunk-boundary artifacts: each window continues from the words already transcribed (conditional continuation), so there are no duplicated or dropped words at the seams and no fragile timestamp-token bookkeeping. Production inference. A CTranslate2 runtime with speculative decoding and built-in mitigation of Whisper's looping-hallucination failure mode.
Some recent Uzbek things https://huggingface.co/datasets/k2speech/FeruzaSpeech - single speaker 40 hours TTS dataset https://huggingface.co/collections/navai-uz/navai-whisper-collection - recently trained Whisper models from Navai https://navai.pro https://huggingface.co/instinct-org/collections - some loosely organized data https://huggingface.co/datasets/OvozifyLabs/asr_evaluate_set - evaluation dataset with Telegram messages https://huggingface.co/datasets/openbank-uz/youtube_transcriptions - large autotranscribed dataset (300k rows, gemini transcribed) https://huggingface.co/uzinfocom-edu-ai/asr-uz-fastconformer-large - recently trained fastconformer (cv + issai + uzvoice + islomov) https://huggingface.co/Abduqayum/whisper-uzbek-medium-callcenter - recently trained whisper https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h - dataset from the above, gemini transribed
Everyone builds self-improvement loops in LLMs, I wonder how they could look like in ASR/TTS. Not many publications on that yet.
We compared three LALM judges against a calibrated human panel across 15 dimensions of speech quality. The LALMs tracked humans closely on relevance, answer quality, and instruction following—what was said—but were much less reliable on naturalness, emotion, pronunciation, and overall feel—how it was said. https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl
Interesting math on speech LLM https://arxiv.org/abs/2604.08003v1 Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Yuan Xie, Jiaqi Song, Guang Qiu, Xianliang Wang, Ming Lei, Jie Gao, Jie Wu Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three metrics to characterize how training paradigms allocate entropy reduction between the speech encoder and the LLM. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a principled multi-stage training strategy grounded in capability-boundary awareness, optimizing parameter efficiency and hallucination robustness. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on Mandarin and English benchmarks show that our method achieves competitive performance with state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.
Interesting project https://github.com/Xiaobin-Rong/unipase UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Xiaobin Rong, Zheng Wang, Yushi Wang, Jun Gao, Jing Lu Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tailored for USE. At its core is DeWavLM-Omni, a unified representation-level enhancement module fine-tuned from WavLM via knowledge distillation on a large-scale supervised multi-distortion dataset. This module directly converts degraded waveforms into clean and linguistically faithful phonetic representations, ensuring robust enhancement with minimal linguistic hallucination. Based on these enhanced phonetic representations, an Adapter generates enhanced acoustic representations containing rich acoustic details, which a neural Vocoder uses to reconstruct corresponding high-fidelity 16-kHz waveforms. A PostNet then converts the waveforms to 48~kHz before resampling them to their original rates, enabling seamless handling of inputs and outputs at multiple sampling rates. Experimental results on several evaluation datasets, covering sub-tasks and full tasks, demonstrate that UniPASE achieves superior or competitive performance compared with existing state-of-the-art models. The proposed model also serves as the backbone of our submission to the URGENT 2026 Challenge, which achieved 1st place in the objective evaluation. The source code and audio demos are available at this https URL.
SGLang-Omni does serious job on optimizing speech models (Higgs TTS too) https://x.com/YichiZ03/status/2078588932191895976