音声処理 × 安全性・評価

HF、実世界向けASRベンチFFASRを公開

HF、実世界向けASRベンチFFASRを公開

✎ ストーリー本文

Hugging Face が実世界での ASR(自動音声認識)性能を測る FFASR Leaderboard を公開。従来の clean 音源ベースのベンチマークでなく、現実的な雑音・アクセント環境での認識精度を評価する軸を提示。arxiv では音声用の U-Net 派生アーキテクチャや Latent Representation 系の研究が続き、ASR コミュニティの評価軸が「静かな部屋の精度」から「現場で使える精度」へシフトしつつある。ハイプ抑制的な指標だが、業務適用の判断には こうした「real world 前提」のベンチが必須で、静かな地味な良い動き。

▲ 公式・報道
公式

Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World

Hugging Face Blog ・ 2026-06-24 ・ 📌

Hugging Face、実環境ASR評価「FFASR Leaderboard」公開

公式

Automating fork maintenance with AI agents

Cohere Blog ・ 2026-06-25

Cohere、AIエージェントでvLLM forkの保守を自動化

学術(arxiv ほか) 25本 ▾
学術

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-25

学術

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

arXiv cs.CL (Computation and Language) ・ 2026-06-25

学術

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

arXiv cs.CL (Computation and Language) ・ 2026-06-25

学術

Heterogeneous Neural Predictivity from Language Models During Naturalistic Comprehension

arXiv cs.LG (Machine Learning) ・ 2026-06-25

学術

FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

arXiv cs.CL (Computation and Language) ・ 2026-06-25

学術

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

arXiv cs.LG (Machine Learning) ・ 2026-06-25

学術

Real-Time Voice AI Hears but Does Not Listen

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-24

学術

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-24

学術

Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

Fully Differentiable Neural Forced Alignment via Soft Dynamic Programming

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS

arXiv cs.CL (Computation and Language) ・ 2026-06-24

学術

L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

arXiv cs.CL (Computation and Language) ・ 2026-06-23

学術

Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-23

学術

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation

arXiv cs.CL (Computation and Language) ・ 2026-06-23

学術

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

arXiv cs.CL (Computation and Language) ・ 2026-06-23

学術

Measuring User's Mental Models of Speech Translation in Human-AI Collaboration

arXiv cs.CL (Computation and Language) ・ 2026-06-23

学術

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-23

学術

Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet

arXiv cs.CL (Computation and Language) ・ 2026-06-23

学術

NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction

arXiv cs.LG (Machine Learning) ・ 2026-06-23

← ストーリー アーカイブ