Recent posts
Ubranch Conformer、このエンコーダーモデルは知りませんでした https://ieeexplore.ieee.org/document/10700716
こちらはWhisper EncoderをSparse Autoencoder(SAE)で解析した研究みたいです。内部表現を分解して、どのような情報が表現されているのかを調べるだけでなく、その特徴を操作して認識や翻訳への影響も解析しているのが面白そうです。 https://arxiv.org/html/2605.12225v2
ASRの内部表現をMechanistic Interpretabilityの観点から解析した研究みたいです。Linear ProbingやLogit Lensだけでなく、Activation Patchingを用いてEncoderからDecoderへの情報伝搬やhallucinationの原因を因果的に解析しているのが面白そうです。

Neta Glazer, Yael Segal-Feldman, Hilit Segev, Aviv Shamsian, Asaf Buchnick, Gill Hetz, Ethan Fetaya, Joseph Keshet, Aviv Navon, "Beyond Transcription: Mechanistic Interpretability in ASR," https://arxiv.org/abs/2508.15882
めっちゃ強そうに見えますね…

Introducing Qwen-Audio-3.0-ASR-Flash: More context-aware. Stronger domain-term recognition. 🚀Our latest ASR model upgrades: • Context consistency • Domain-term recognition • Custom hotwords • Speech polishing into structured transcripts ⚡️In internal tests: • Medical term recall: 95.36% • Industrial term recall: 93.24% Qwen-Audio-3.0-ASR-Flash-Streaming: https://t.co/QqTp4HkknJ Qwen-Audio-3.0-ASR-Flash-Filetrans: https://t.co/7llmle4deF Qwen-Audio-3.0-ASR-Flash: https://t.co/VxruY0Ruhi

Cross-Attentionに単語境界を教師あり学習させることで、タイムスタンプ精度を改善しているみたいです。非流暢音声でも高精度なタイムスタンプが得られているのが面白そうです。

[CL] Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing L Wagner, M Zusag, B Thallinger [nyra labs] (2026) https://arxiv.org/abs/2607.18934




単語レベルのタイムスタンプの精度高いみたいです

CrisperWhisper 2.0: production-grade verbatim speech recognition - 29ms word-timing accuracy - controllable toggle between verbatim/clean - 10 languages - seamless longform inference https://huggingface.co/nyralabs/CrisperWhisper2.0_large


秋ASJのプログラム公開されてますね https://acoustics.jp/annualmeeting/program/
正常聴力者と難聴者を対象に、駅などの仮想音響環境で語音了解度や聴取努力、補聴器の効果を多面的に測定した研究みたいです。 https://www.tandfonline.com/doi/full/10.1080/14992027.2026.2615093
MossFormerGAN_SE_16K日本語ASRの前処理でどうなのか試してみたいです https://github.com/modelscope/ClearerVoice-Studio/tree/main/clearvoice

認識の前処理としての強調はMossFormerが良さそうなんですかね
言語に依存しない聴覚評価について整理したレビュー論文みたいです 読んでみたいです https://www.frontiersin.org/journals/neuroscience/articles/10.3389/fnins.2026.1887850/full
ZipformerとCR-CTCを使ったTransducerモデルが公開されていますね CR-CTCなしバージョンもあるみたいですがそれより精度高いみたいですね https://huggingface.co/soundsgoodai/Zipformer-cr-ctc-transducer-XL-290M


