AI 语音与语音技术新闻
发布时间:2026-07-27 | 浏览:1
了解 AI 语音技术、语音合成和不断发展的监管环境的最新动态
Microsoft Unveils MAI-Voice-1: Hyper-Realistic Speech Generation from Just One Minute of Audio
Microsoft launches three new foundational AI models including MAI-Voice-1, which delivers hyper-realistic voice synthesis and custom brand voice creation, marking a major leap in enterprise TTS capabilities.
ElevenLabs Reaches $11 Billion Valuation, Eyes IPO as Voice AI Becomes Enterprise Standard
AI voice startup ElevenLabs raises $500 million at an $11 billion valuation, tripling its worth in just over a year while forging major partnerships with IBM and planning a potential IPO.
Global AI Voice Regulation Tightens: EU AI Act Deepfake Rules Take Effect as Voice Cloning Crosses 'Indistinguishable Threshold'
As voice cloning technology reaches human-level quality, regulators worldwide respond with new laws — the EU AI Act's deepfake labeling rules, the US ELVIS Act, and emerging biometric voice data protections reshape the industry landscape.
OpenAI Launches Voice Engine to the Public: Real-Time Conversational TTS Now Available to All Developers
After over a year of limited preview, OpenAI opens Voice Engine to all API developers, introducing real-time streaming TTS with emotional awareness and 40+ language support at significantly reduced pricing.
Google DeepMind Brings Studio-Quality TTS to Smartphones with SoundStorm 2 Edge — No Internet Required
Google DeepMind announces SoundStorm 2 Edge, a compact on-device TTS model that runs entirely on mobile hardware, delivering studio-quality voice synthesis without cloud connectivity and opening new possibilities for offline accessibility.
AI Dubbing Market Surges Past $2 Billion as Hollywood, Streaming Giants, and Game Studios Embrace Automated Localization
The AI-powered dubbing and localization market crosses the $2 billion mark in Q1 2026, driven by adoption from Netflix, Disney+, and major game publishers seeking to reach global audiences at a fraction of traditional costs.
Apple Unveils 'Personal Voice 2.0' in iOS 20: On-Device Voice Cloning Creates Your Digital Twin in 3 Minutes
Apple announces Personal Voice 2.0 at its spring event, allowing users to create a highly realistic clone of their own voice in just 3 minutes of recording — all processed entirely on-device with Apple Silicon, positioning it as the privacy-first alternative to cloud-based voice AI.
Spotify Rolls Out AI Voice Translation for Podcasts Globally: Your Favorite Hosts Now Speak 40 Languages in Their Own Voice
Spotify launches its AI-powered podcast translation feature worldwide, using voice cloning technology to automatically dub podcasts into 40 languages while preserving each host's unique voice characteristics — opening 100,000+ shows to global audiences overnight.
FDA Clears First AI Voice Assistant for Clinical Use: Voice-Based Patient Screening Enters the Hospital
The FDA grants its first clearance for an AI voice assistant designed for clinical patient interaction, allowing automated voice-based symptom screening and triage in emergency departments — marking a historic milestone for voice AI in healthcare.
Meta Releases Llama-Voice: First Fully Open-Source TTS Model to Match Commercial Giants in 50+ Languages
Meta drops Llama-Voice under an Apache 2.0 license, delivering near state-of-the-art voice synthesis, zero-shot voice cloning from 10 seconds of audio, and 52-language coverage — all runnable on a single consumer GPU.
NVIDIA Launches Voice Foundry NIM: Blackwell-Optimized Microservices Cut Real-Time TTS Costs by 70%
NVIDIA unveils Voice Foundry, a dedicated suite of NIM inference microservices for TTS and STT optimized for Blackwell GB200 hardware, promising sub-80ms first-token latency and 70% lower per-character costs for enterprise voice applications.
Audible Opens AI-Narrated Audiobook Catalog to 400,000 Backlist Titles — Narrators Split on Landmark Royalty Model
Amazon's Audible launches the industry's largest AI-narrated audiobook catalog, adding 400,000 previously unnarrated titles using voice clones of consenting narrators, with a first-of-its-kind per-listen residual model that splits the narration community.
Google Launches Gemini 3.1 Flash TTS: 70+ Languages, Multi-Speaker Dialogue, and a Top Spot on the Artificial Analysis Leaderboard
Google introduces Gemini 3.1 Flash TTS, a new text-to-speech model with audio tags for fine-grained vocal control, native multi-speaker dialogue, and 70+ language support — landing in the 'most attractive quadrant' of the Artificial Analysis TTS leaderboard with an Elo of 1,211.
OpenAI Launches GPT-Realtime-2: Voice Models with GPT-5-Class Reasoning, Live Translation, and Streaming Transcription
OpenAI introduces three new Realtime API voice models — GPT-Realtime-2 with GPT-5-class reasoning, GPT-Realtime-Translate covering 70+ input languages, and GPT-Realtime-Whisper for live transcription — quadrupling the context window to 128K tokens and bringing voice agents closer to production-ready workflows.
Microsoft Launches MAI-Voice-2 at Build 2026: Expressive Speech and Zero-Shot Voice Cloning Across 15 Languages
Microsoft unveils MAI-Voice-2, calling it the most expressive and natural-sounding text-to-speech model it has built, expanding from English-only to 15 languages with granular emotion control, code-switching, and zero-shot voice prompting from a few seconds of audio.
Wispr Hits ~$2 Billion Valuation as AI Voice Dictation Becomes a Workplace Standard
Wispr, the startup behind the AI dictation tool Wispr Flow, is raising roughly $260 million at a near-$2 billion valuation led by Menlo Ventures — nearly tripling its worth in six months as voice-to-text moves from novelty to everyday workplace productivity tool.
FTC Begins Enforcing the TAKE IT DOWN Act: Platforms Face $53,088-Per-Violation Penalties for AI Deepfakes
The FTC's civil enforcement of the TAKE IT DOWN Act took effect on May 19, 2026, requiring platforms to remove nonconsensual intimate imagery — including AI-generated deepfakes — within 48 hours, with penalties of $53,088 per violation. The agency promptly sent warning letters to major platforms and 'nudify' websites.
Poland Government Takes Stake in ElevenLabs, Launches AI Lab to Build Voice AI from Europe
The Government of Poland invests in ElevenLabs through its Vinci/BGK Group, joining Andreessen Horowitz and Sequoia as a strategic backer, while launching AI Lab Poland to nurture the next generation of voice AI companies with global ambition.
ElevenLabs Launches Dubbing v2: Emotion-Preserving AI Dubbing Across 90+ Languages
ElevenLabs releases Dubbing v2, a breakthrough AI dubbing model that preserves the original speaker's emotion, tone, and pacing across 90+ languages by conditioning directly on the performance rather than just transcripts.
ElevenLabs Partners with UK Government to Bring Voice AI to Public Services, Doubles London Headquarters
ElevenLabs signs a Memorandum of Understanding with the UK's Department for Science, Innovation and Technology to deploy voice AI in public services, focusing on accessibility for the visually impaired, elderly, and linguistically diverse communities.
Rumik Launches Silk Mulberry 1.5: 'Describe a Voice Into Existence' with Plain-Language Prompts, Matching Commercial TTS Giants at 95% Lower Cost
Indian AI startup Rumik releases Silk Mulberry 1.5, a text-to-speech model that replaces preset voice menus with plain-language voice descriptions, achieving MOS scores competitive with ElevenLabs and Google at roughly $0.0046 per minute.
Michael Caine's AI Voice Narrates 13-Hour 'The Odyssey' Audiobook — 20 AI Characters, Original Score, Built by 4 Producers in 6 Weeks
ElevenLabs releases a cinematic audiobook of Homer's The Odyssey narrated by an authorized AI replica of Sir Michael Caine's voice, featuring ~20 AI-generated character voices, original music, and sound design — all produced by a four-person team in six weeks.
Five9 Launches Voice AI Agents with ElevenLabs, Deepgram, and OpenAI Under the Hood — Targeting Legacy IVR Replacement
Five9 unveils Voice AI Agents at Customer Contact Week 2026, combining ElevenLabs TTS, Deepgram ASR, and OpenAI reasoning in a proprietary three-model architecture built to replace scripted IVR systems with natural, human-like voice self-service.
xAI Launches Voice Agent Builder: No-Code Platform Harnesses Grok Voice to Beat GPT and Gemini in Telephony Benchmarks
Elon Musk's xAI enters the voice AI market with Voice Agent Builder, a no-code platform powered by Grok Voice Think Fast 1.0 that scores 67.3% on the τ-voice Bench — far outpacing Google Gemini 3.1 Flash Live (43.8%) and OpenAI GPT Realtime 1.5 (35.3%) — with pricing starting at $0.05 per minute.
Bland.ai Raises $50M Series C After 180 Investor Rejections, Now Powers 3.5 Million Voice Calls Per Week
San Francisco voice AI startup Bland.ai closes a $50 million Series C led by Dell Technologies Capital, bringing total funding past $100 million — after founders were rejected by 180 investors who told them 'phone calls won't exist in a year.'
NetEase Youdao Releases Confucius4-TTS: Open-Source 14-Language Voice Cloning from Just 3 Seconds of Audio
Chinese edtech giant NetEase Youdao open-sources Confucius4-TTS under Apache 2.0, a 1.3B-parameter voice cloning model achieving 85%+ voice similarity from 3 seconds of audio across 14 languages — with no reference text needed for cross-lingual cloning.
NO FAKES Act Unanimously Passes Senate Judiciary Committee, Creating Federal Voice and Likeness Protection
The bipartisan NO FAKES Act clears the Senate Judiciary Committee by unanimous voice vote, creating a federal intellectual property right over AI-generated digital replicas of voice and visual likeness — with platform liability, 70-year post-mortem protections, and DMCA-style takedown provisions.
Kotoba Technologies Raises $10 Million to Bring Real-Time Voice AI to East Asian Languages
San Francisco and Tokyo-based Kotoba Technologies raises an additional $10 million in seed funding led by Kindred Ventures, with Salesforce Ventures and Sony Innovation Fund participating, to expand its Koto voice AI model optimized for Japanese, Korean, and Chinese — languages spoken by roughly 1.6 billion people.
ViiTorVoice-NAR Goes Open Source: First TTS Model That Edits Single Words Inside Finished Audio
Chinese startup Yunshang Qulv releases ViiTorVoice-NAR under Apache 2.0, introducing word-level audio editing that replaces individual words without regenerating surrounding content — alongside sub-60ms latency and benchmark-leading accuracy on both English and Chinese.
OpenAI Launches GPT-Live: Full-Duplex Voice Model Lets ChatGPT Listen and Speak Simultaneously
OpenAI rolls out GPT-Live-1 and GPT-Live-1 mini globally, introducing full-duplex architecture that enables ChatGPT to listen and speak at the same time — with background task delegation to GPT-5.5 for complex reasoning, marking voice AI's shift from turn-based chat to continuous conversation.
Gradium Raises $100M Seed Round Backed by Nvidia to Build Ultra-Low-Latency Voice AI
Paris-based voice AI startup Gradium, spun out of French research lab Kyutai, extends its seed round to over $100 million with Nvidia joining as a strategic investor — signaling that the race to eliminate latency in AI voice conversations is attracting infrastructure-level capital.
Tencent Cloud Partners with Inworld AI to Deliver One-Stop Real-Time Voice AI with Sub-130ms Latency Across 100+ Languages
Tencent Cloud and Inworld AI announce a strategic partnership integrating Inworld's top-ranked TTS models into Tencent RTC's global infrastructure, creating a production-grade voice AI solution with sub-130ms first-chunk latency, 100+ language support, and voice cloning — backed by 3,200+ global edge nodes.
Rime Raises $24 Million Series A to Build Enterprise-Ready Speech-to-Speech Voice AI
San Francisco-based Rime raises $24 million led by M13 to scale its linguistics-first speech-to-speech voice AI platform, already handling nearly 100 million calls monthly for Mayo Clinic and Dialpad — and hires ex-Meta audio research lead Rafael Valle as Chief Science Officer.
Omilia Launches Lexis: First Native Generative TTS Built Into an Enterprise Contact Center Platform
Omilia releases Lexis, a generative text-to-speech model built natively into its Cloud Platform — not via third-party API — delivering sub-45ms latency, 25+ languages, instant voice cloning, and PCI-DSS, HIPAA, and GDPR compliance for regulated enterprise contact centers.
Ex-Waymo Engineer Raises $28 Million to Bring Autonomous-Vehicle-Grade Testing to Voice AI Agents
Coval, founded by a former Waymo evaluation infrastructure lead, raises $28 million from Norwest, Base10 Partners, and Twilio Ventures to build simulation-based testing for voice AI agents — applying the same safety discipline that made self-driving cars possible to autonomous voice systems.
Alibaba's Qwen-Audio-3.0-TTS Tops Global Leaderboard: 16 Languages, Voice Cloning, and Sub-200ms Latency
Alibaba releases Qwen-Audio-3.0-TTS in Flash and Plus variants, claiming the #1 spot on the Artificial Analysis Speech Arena with an Elo score of 1,234 — surpassing Google, ElevenLabs, and Cartesia in blind listening tests while adding support for 16 languages and 20 Chinese dialects.
ByteDance Unveils Seed Audio 1.0: A Single Prompt Generates Dialogue, Sound Effects, and Ambient Audio in One Pass
ByteDance's Seed team launches Seed Audio 1.0, an 'audio creation model' that jointly generates speech, sound effects, music, and environmental audio within a unified framework — eliminating the multi-tool pipeline that traditional audio production requires and delivering 90%+ usable audio in a single inference step.
Japan Moves to Legally Protect Voices from Unauthorized AI Use — First-of-Its-Kind Framework in Asia
Japan's Ministry of Justice unveils draft guidelines extending right-of-publicity protections to voices, proposing that unauthorized AI-generated voice content causing harm be treated as a civil violation — responding to over 43,000 suspected unauthorized uploads and an estimated ¥4.5 billion in economic losses to rights holders.
Speechify's Simba 3.2 Tops Global TTS Leaderboard: Consumer-First Voice Model Beats OpenAI, Google, and ElevenLabs in Blind Tests
Speechify's Simba 3.2 reaches #1 on the Artificial Analysis TTS leaderboard, beating ElevenLabs, OpenAI, Google DeepMind, and Cartesia in independent blind listening tests — while priced at just $6–10 per million characters, roughly 15x cheaper than rivals in the top ten.
Verbatik Launches First MCP Text-to-Speech Server: 2,700+ Neural Voices Now a Native Capability for Claude, Codex, and Other AI Assistants
London-based Verbatik Technologies releases the first text-to-speech server for the Model Context Protocol (MCP), connecting 2,700+ neural voices across 50+ languages directly to AI assistants — turning voice generation into a native agent capability rather than an external API call.
AMD Lemonade 11.0 Brings Text-to-Speech and Voice Cloning to Open-Source Local AI Server
AMD releases Lemonade 11.0, a major update to its open-source local AI server, adding text-to-speech with voice cloning and voice design via the OpenMOSS backend — enabling entirely offline, privacy-preserving speech synthesis on consumer AMD hardware.
Deepgram Brings Enterprise Voice AI to Snapdragon PCs: Nova-3 Speech Recognition Runs Fully On-Device with 6.89% Word Error Rate
Deepgram partners with Qualcomm to optimize its Nova-3 speech-to-text model for Snapdragon X Series processors, enabling real-time enterprise voice recognition that runs entirely on-device via the Hexagon NPU — no cloud required — targeting AI PCs, automotive, XR, and industrial edge deployments.
China's National Security Ministry Warns AI Voice Cloning Now Takes Just 3 Seconds — Outlines Three Major Risk Categories
China's Ministry of State Security issues a landmark public advisory warning that voice cloning platforms now require as little as 3 seconds of audio to create convincing fakes, outlining risks of impersonation fraud, voice rights infringement, and public opinion manipulation — while detailing an existing three-pillar legal framework for enforcement.
ElevenLabs Explores $22 Billion Tender Offer, Doubling Valuation in Six Months as Voice AI Revenue Surges Past $5 Billion ARR
ElevenLabs enters early-stage negotiations for a secondary share sale at a $22 billion valuation — doubling its February 2026 price tag — backed by surging enterprise adoption, $5 billion-plus ARR, and 33 million AI-powered conversations handled by its agents platform in the first half of 2026.
Cartesia Launches Sonic 3.5 and Ink 2: SSM Architecture Makes It the First Provider to Top Both TTS and STT Leaderboards
Cartesia releases Sonic 3.5 (TTS) and Ink 2 (STT) built on State Space Model architecture rather than Transformers, achieving sub-90ms latency and 42-language support — becoming the only provider to simultaneously hold the #1 ranking for both speaking and listening on Artificial Analysis benchmarks.
Vapi Raises $50M at $500M Valuation After Amazon Ring Picks Its Voice AI Platform Over 40 Rivals
Voice AI infrastructure startup Vapi closes a $50 million Series B led by Peak XV Partners at a $500 million valuation, after Amazon Ring selected its platform from over 40 vendors to handle 100% of inbound customer support calls — with the company now processing over 1 billion total calls.
Lingraphica Launches AI-Powered 'Conversations' AAC Tool: Voice AI Enables Spontaneous Dialogue for People With Speech Disabilities
Lingraphica releases Conversations, an AI-powered communication tool for its AAC devices that transcribes a conversation partner's speech and suggests contextually relevant responses — cutting conversation time by 80% and achieving a 100% user test completion rate in clinical trials.