Охват к подписчикам
49,4%
ERR
Реакции к просмотрам
0,00%
0 на 21 постов
Пересылки к просмотрам
1,41%
255
Постов в день
0,6
всего 21
Где отзываются чаще
доля реакций к просмотрам- 14 авг.One more agentic thing https://github.com/InteractiveASR/AgenticASR https://interactiveasr.github.io/ https://arxiv.org/abs/2604.09121 Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Peng Wang, Yanqiao Zhu, Zixuan Jiang, Qinyuan Chen, Xingjian Zhao, Xipeng Qiu, Wupeng Wang, Zhifu Gao, Xiangang Li, Kai Yu, Xie Chen Recent years have witnessed remarkable progress in automatic speech recognition (ASR), driven by advances in model architectures and large-scale training data. However, two important aspects remain underexplored. First, Word Error Rate (WER), the dominant evaluation metric for decades, treats all words equally and often fails to reflect the semantic correctness of an utterance at the sentence level. Second, interactive correction-an essential component of human communication-has rarely been systematically studied in ASR research. In this paper, we integrate these two perspectives under an agentic framework for interactive ASR. We propose leveraging LLM-as-a-Judge as a semantic-aware evaluation metric to assess recognition quality beyond token-level accuracy. Furthermore, we design an LLM-driven agent framework to simulate human-like multi-turn interaction, enabling iterative refinement of recognition outputs through semantic feedback. Extensive experiments are conducted on standard benchmarks, including GigaSpeech (English), WenetSpeech (Chinese), the ASRU 2019 code-switching test set. Both objective and subjective evaluations demonstrate the effectiveness of the proposed framework in improving semantic fidelity and interactive correction capability. We will release the code to facilitate future research in interactive and agentic ASR.0,00%
- 11 авг.Big paper on state of the art in speech translation (87 pages) Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign https://aclanthology.org/2026.iwslt-1.39.pdf This paper reports on the outcomes of the shared tasks organized as part of the 23rd International Workshop on Spoken Language Translation (IWSLT). The workshop covered ten major challenges in spoken language translation, including speech-to-text translation for both high-resource and low-resource language pairs, customized speech translation, speech generation, instruction-following speech processing, and the evaluation of speech translation systems. The shared tasks received strong participation, with more than 30 teams submitting runs. This year’s edition broadened the range of tasks, placing particular emphasis on speech generation and evaluation metrics.0,00%
- 11 авг.Duplex paper from today https://github.com/duplexgen/duplexgen-code https://arxiv.org/abs/2607.26178 DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-Tür Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.0,00%
- 11 авг.A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. https://huggingface.co/datasets/sarvamai/indic-diarbench https://www.sarvam.ai/blogs/indic-diarbench0,00%
- 7 авг.Topic of today - autoregression vs diffusion in TTS First of all the presentation from Meta. Claims diffusion/autoregression decision could be dynamic https://www.youtube.com/watch?v=kMimQxIJLos Second paper from today on the similar topic https://arxiv.org/abs/2607.04140 DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Junwon Moon, Yejin Lee, Seungbeom Kim, Hoseong Ahn, Sewoong Park, Heeseung Kim, Kyuhong Shim Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, since the model receives the full input text before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a 1/t-weighted training objective and a time-shifted inference schedule that together defer low-confidence positions to later steps. Trained on only 585 hours of LibriTTS, DELTA-TTS achieves a 1.75% WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens 3.3x faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates the hallucinations observed in AR generation.0,00%
- 7 авг.Deepmind also releases something https://github.com/google-deepmind/phasecoder https://arxiv.org/abs/2601.21124 PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only spatial audio encoder that is agnostic to microphone geometry. PhaseCoder takes raw multichannel audio and microphone coordinates as inputs to perform localization and produces robust spatial embeddings. We demonstrate that Gemma 3n LLM can be fine-tuned to reason over "Spatial Audio Tokens" produced by PhaseCoder. We show our encoder achieves state-of-the-art results on microphone-invariant localization benchmarks and, for the first time, enables an LLM to perform complex spatial reasoning and targeted transcription tasks from an arbitrary microphone array.0,00%
- 7 авг.Piotr is a reincarnation of Lennart https://x.com/PiotrZelasko/status/20857098916042547880,00%
- 7 авг.https://www.reddit.com/r/homeassistant/comments/1vhofep/hacked_and_debloated_an_echo_dot_2_local_llm/0,00%
- 6 авг.There is a big interest in full duplex as I see, here is a nice collection of papers https://github.com/Ruiqi-Yan/Awesome-Full-Duplex-SDM0,00%
- 5 авг.https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B NVIDIA NemotronLabs VoiceChat is a 11B end-to-end, real-time speech full duplex (FD) model for conversational AI that jointly performs streaming speech understanding and speech generation [1, 2]. Unlike traditional cascaded stacks (ASR → LLM → TTS), this model achieves full duplex, real-time, seamless voice interaction in one unified architecture, eliminating the need for multiple models or API handoffs, thus reducing end-to-end latency. It sets new benchmarks by bringing open, robust, and highly natural conversation capabilities. Moreover, NVIDIA NemotronLabs VoiceChat is the first open full-duplex model to support tool calling while maintaining a natural conversation flow during tool execution. For each tool, a specific “on-hold” message can be defined that will be spoken by the agent as soon as the LLM generates the text that will trigger the tool call and response.0,00%
- 4 авг.Things move on in OpenAI as well. Interesting that voice model is separate. And no turn detector anymore. https://x.com/OpenAI/status/2084378415818579975 Lots of interesting technical details, from realtime inference, to dynamic compaction, to WebRTC optimization. https://openai.com/index/continuous-voice-interaction-with-gpt-live/0,00%
- 1 авг.One more https://github.com/AnXMuy/AgenticASR0,00%