Recent posts
Every business is unique. Unique customers, unique products, unique markets you serve and a unique culture. Features like keyterm prompting are a nice and easy way to inject context to make your agents smarter at understanding your users & customers. Really cool feature

New to Ink-2: keyterm prompting Pass rare terms like product names, industry jargon, and even French words for accurate transcription. 20% higher keyword recall, with no added latency. Learn more → http://cartesia.ai/blog/keyterm-prompting/?utm_source=x
AI factory goes brrrrrrrrrr

Excited that @cartesia is in @cursor_ai's India campaign with some of the most ambitious founders of Indian origin (including in Delhi where I grew up) featuring amazing folks @amanrsanger, @vipulved, @tankots, @mukundjha, @ManishaRaisingh, @regards_rishi, @_sankyy

Founder mode is great, I can feel the restless energy through the page (no pun intended)

JUST IN: Sergey Brin to reportedly take direct oversight of Gemini as Google restructures its AI leadership.
Our London office is growing quickly, reach out to @_albertgu or me if you’d like to work on a very different, new and exciting research agenda to advance multimodal models

Good night DeepMind. Wow. FT: "Google is shifting control of its AI effort from London back to Silicon Valley" "executives and board members remained concerned by its weaker position in coding models and enterprise AI, where Anthropic and OpenAI have established an early lead." "Several current and former DeepMind employees said the reorganisation had sent shockwaves through the London-based lab, with some fearing it marked the end of the research culture that Hassabis has long protected. One former DeepMind executive at a rival company said they had already received calls from multiple staff who said they were ready to jump ship." "Over the past year, Hassabis has devoted more time to Isomorphic Labs, the AI drug discovery company he founded"
It's so exciting that there are genuinely amazing researchers starting new companies outside of big tech, taking swings at the biggest problems in AI The reason I ended up in AI was the deep infectious passion and enthusiasm in the community to solve the seemingly impossible
Then you melt all the stages and learn at test time — and the distinction between training and inference disappears in a lot of interesting scenarios

@aaronburnett 99% of compute long-term will be for inference
Nobody has really figured out how to connect the world of atoms to the world of bits Connecting knowledge & reasoning to the physical world really feels like the next big intelligence frontier A few new ideas are all that’s needed
Synthesizing speech is a pretty hard problem, our blog post lays out the challenges in building and evaluating these systems for production use

"Is this TTS model good?" gets harder to answer as models improve. "Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first. We wrote up the failure modes that make TTS eval hard: https://www.cartesia.ai/blog/is-this-tts-model-good
.@j0nathanj is working on an incredibly important problem — making robots intuitive and intelligent the team is humble, creative and extremely brilliant this is their first release, really pumped about what they’re building

We put 100 real AI-powered robots online. Anyone in the world can control them right now, from a browser. Go make one do something:
Super cool collaboration with @paraga, @georgepickett and team. Access the web and bring in the freshest information to your voice agent while it talks and listens. All at un-parallel-ed speeds.

Today, @cartesia and @p0 are making it possible for your voice agents to search the web at conversational speed. Cartesia builds the fastest voice models available, and Turbo mode for Parallel Search extends that same low latency to web search. An agent built on Cartesia can query the live web without interrupting the natural flow of conversation. Learn more 👇
.@_albertgu and I are organizing a research happy hour in London on Monday July 13th. Come hang out! https://luma.com/3p1u9pgf
The @cartesia team continues to kill it This eval measures quality when models mimic the same voices. Stronger models generalize better.

Announcing the Controlled Voice Arena Leaderboard comparing Text to Speech models on the same set of 8 cloned voices The Controlled Voice Arena standardizes, through voice cloning, the set of voices that each model’s performance is evaluated on - separating specific voice preference from broader aspects of model quality. It complements our Provider Voice Arena, where each model uses a select set of its own available voices. We have generated speech samples on models that offer voice cloning abilities using the same voice categories as our existing Provider Voice Arena, namely: 2 US Male voices, 2 US Female voices, 2 UK Male voices, 2 UK Female voices. Each model has been cloned on the same 1-2 minute audio recordings for each voice. Key results ➤ Overall: @cartesia Sonic 3.5 leads (1,122 Elo), followed by @ElevenLabs Eleven v3 (1,088) and @inworld_ai Realtime TTS-2 - Research Preview (1,070) ➤ US accent: Cartesia Sonic 3.5 leads (1,139 Elo), followed by ElevenLabs Eleven v3 (1,104) and Inworld Realtime TTS-2 - Research Preview (1,059) ➤ UK accent: Cartesia Sonic 3.5 also leads (1,103 Elo), with Inworld Realtime TTS-2 - Research Preview (1,075) moving ahead of ElevenLabs Eleven v3 (1,067) into 2nd ➤ Open weights: @FishAudio S2 Pro leads (1,034 Elo), followed by @MistralAI Voxtral TTS (1,024) and @resembleai Chatterbox (930) See more details below ⬇️

Shipped one of the most common ask from users: how does a voice change when you play it on different channels like telephony and the web? Try it out at http://play.cartesia.ai

Memory is a huge challenge for frontier AI systems, and true personalization remains out of reach until progress is made on it. Excited to see close friends and former lab mates go after these hard problems with new ideas @EyubogluSabri @MayeeChen @dan_biderman @scott_linderman

@elise_aiみたいな信頼できるパートナーとの密なフィードバックループのおかげで、業界の最高水準の企業が必要とするカパビリティに向かって、めちゃ速いペースでモデルを改善できてる。 一緒に、可能性のフロンティアを押し広げてる。
原文を表示 (en)
Thanks to the tight feedback loop with close partners like @elise_ai, we've been improving our models very quickly towards the capabilities that the best companies in the space need. Together, we're pushing the frontier of what's possible.

We didn't switch to @cartesia’s Sonic 3.5 because it was incrementally better, we switched because nothing else came close. Since deploying Sonic 3.5, @elise_ai has seen a 2.9% lift in conversion and a 12.2% increase in customer engagement. 🚀
また1位だ!
原文を表示 (en)
#1 again!

Cartesia Sonic 3.5 is the #1 streaming TTS model on Voice Arena US English Leaderboard. In the overall leaderboard (streaming + non-streaming) it jumped from rank #9 → #2, moving ahead of Grok TTS, ElevenLabs v3 & OpenAI's gpt-4o-mini-tts. Sonic-3.5 is the latest TTS model from @cartesia . It supports 42 languages, with 500+ voices available out of the box. The model has been highly preferred among raters on @voicearena_ai . Results backed by 11,110 blind, head-to-head listener votes

新しい音声テキスト変換モデルInk-2をリリースしました。Artificial Analysisで第1位です。 ストリーミング向けに構築—低レイテンシー、高速eagerモードと、ユーザーが話し終わったかを検出するセマンティックエンドポイント組み込み 新しいアーキテクチャとアルゴリズムによってこのPareto支配が実現しました
原文を表示 (en)
Our new speech-to-text model Ink-2 is out and #1 on Artificial Analysis. It’s built for streaming — low latency, fast eager mode and built in semantic endpoints to detect when users are done talking New architectures & algorithms made this Pareto-dominance possible

Cartesia Ink-2 debuts as #1 for accuracy on the brand-new streaming speech-to-text leaderboard from @ArtificialAnlys! We designed Ink-2 from the ground up for voice agents - with low latency, eager transcripts, and semantic endpointing.




