Back to people
@krandiash
K

Karan Goel

音声
@krandiash

founder ceo @cartesia, derp learned @stanfordailab @mldcmu @iitdelhi

20KFollowers1.0KFollowing2.1KPostsView on X

Recent posts

Every business is unique. Unique customers, unique products, unique markets you serve and a unique culture. Features like keyterm prompting are a nice and easy way to inject context to make your agents smarter at understanding your users & customers. Really cool feature

@cartesia
C
Cartesia@cartesia

New to Ink-2: keyterm prompting Pass rare terms like product names, industry jargon, and even French words for accurate transcription. 20% higher keyword recall, with no added latency. Learn more → http://cartesia.ai/blog/keyterm-prompting/?utm_source=x

Excited that @cartesia is in @cursor_ai's India campaign with some of the most ambitious founders of Indian origin (including in Delhi where I grew up) featuring amazing folks @amanrsanger, @vipulved, @tankots, @mukundjha, @ManishaRaisingh, @regards_rishi, @_sankyy

Photo 1

It's so exciting that there are genuinely amazing researchers starting new companies outside of big tech, taking swings at the biggest problems in AI The reason I ended up in AI was the deep infectious passion and enthusiasm in the community to solve the seemingly impossible

Then you melt all the stages and learn at test time — and the distinction between training and inference disappears in a lot of interesting scenarios

@elonmusk
E
Elon Musk@elonmusk

@aaronburnett 99% of compute long-term will be for inference

Nobody has really figured out how to connect the world of atoms to the world of bits Connecting knowledge & reasoning to the physical world really feels like the next big intelligence frontier A few new ideas are all that’s needed

.@j0nathanj is working on an incredibly important problem — making robots intuitive and intelligent the team is humble, creative and extremely brilliant this is their first release, really pumped about what they’re building

@Enigma_AI
E
Enigma@Enigma_AI

We put 100 real AI-powered robots online. Anyone in the world can control them right now, from a browser. Go make one do something:

Super cool collaboration with @paraga, @georgepickett and team. Access the web and bring in the freshest information to your voice agent while it talks and listens. All at un-parallel-ed speeds.

@cartesia
C
Cartesia@cartesia

Today, @cartesia and @p0 are making it possible for your voice agents to search the web at conversational speed. Cartesia builds the fastest voice models available, and Turbo mode for Parallel Search extends that same low latency to web search. An agent built on Cartesia can query the live web without interrupting the natural flow of conversation. Learn more 👇

.@_albertgu and I are organizing a research happy hour in London on Monday July 13th. Come hang out! https://luma.com/3p1u9pgf

The @cartesia team continues to kill it This eval measures quality when models mimic the same voices. Stronger models generalize better.

@ArtificialAnlys
A
Artificial Analysis@ArtificialAnlys

Announcing the Controlled Voice Arena Leaderboard comparing Text to Speech models on the same set of 8 cloned voices The Controlled Voice Arena standardizes, through voice cloning, the set of voices that each model’s performance is evaluated on - separating specific voice preference from broader aspects of model quality. It complements our Provider Voice Arena, where each model uses a select set of its own available voices. We have generated speech samples on models that offer voice cloning abilities using the same voice categories as our existing Provider Voice Arena, namely: 2 US Male voices, 2 US Female voices, 2 UK Male voices, 2 UK Female voices. Each model has been cloned on the same 1-2 minute audio recordings for each voice. Key results ➤ Overall: @cartesia Sonic 3.5 leads (1,122 Elo), followed by @ElevenLabs Eleven v3 (1,088) and @inworld_ai Realtime TTS-2 - Research Preview (1,070) ➤ US accent: Cartesia Sonic 3.5 leads (1,139 Elo), followed by ElevenLabs Eleven v3 (1,104) and Inworld Realtime TTS-2 - Research Preview (1,059) ➤ UK accent: Cartesia Sonic 3.5 also leads (1,103 Elo), with Inworld Realtime TTS-2 - Research Preview (1,075) moving ahead of ElevenLabs Eleven v3 (1,067) into 2nd ➤ Open weights: @FishAudio S2 Pro leads (1,034 Elo), followed by @MistralAI Voxtral TTS (1,024) and @resembleai Chatterbox (930) See more details below ⬇️

Photo 1

Shipped one of the most common ask from users: how does a voice change when you play it on different channels like telephony and the web? Try it out at http://play.cartesia.ai

Photo 1

Memory is a huge challenge for frontier AI systems, and true personalization remains out of reach until progress is made on it. Excited to see close friends and former lab mates go after these hard problems with new ideas @EyubogluSabri @MayeeChen @dan_biderman @scott_linderman

@EngramLab
E
Engram@EngramLab

http://x.com/i/article/2069463677733142528

@elise_aiみたいな信頼できるパートナーとの密なフィードバックループのおかげで、業界の最高水準の企業が必要とするカパビリティに向かって、めちゃ速いペースでモデルを改善できてる。 一緒に、可能性のフロンティアを押し広げてる。

原文を表示 (en)

Thanks to the tight feedback loop with close partners like @elise_ai, we've been improving our models very quickly towards the capabilities that the best companies in the space need. Together, we're pushing the frontier of what's possible.

@ShrutiGoli
S
Shruti Goli@ShrutiGoli

We didn't switch to @cartesia’s Sonic 3.5 because it was incrementally better, we switched because nothing else came close. Since deploying Sonic 3.5, @elise_ai has seen a 2.9% lift in conversion and a 12.2% increase in customer engagement. 🚀

また1位だ!

原文を表示 (en)

#1 again!

@voicearena_ai
V
Voice Arena@voicearena_ai

Cartesia Sonic 3.5 is the #1 streaming TTS model on Voice Arena US English Leaderboard. In the overall leaderboard (streaming + non-streaming) it jumped from rank #9#2, moving ahead of Grok TTS, ElevenLabs v3 & OpenAI's gpt-4o-mini-tts. Sonic-3.5 is the latest TTS model from @cartesia . It supports 42 languages, with 500+ voices available out of the box. The model has been highly preferred among raters on @voicearena_ai . Results backed by 11,110 blind, head-to-head listener votes

Photo 1

新しいビデオが近々公開予定です 👀

原文を表示 (en)

A new video is coming soon 👀

Photo 1

新しい音声テキスト変換モデルInk-2をリリースしました。Artificial Analysisで第1位です。 ストリーミング向けに構築—低レイテンシー、高速eagerモードと、ユーザーが話し終わったかを検出するセマンティックエンドポイント組み込み 新しいアーキテクチャとアルゴリズムによってこのPareto支配が実現しました

原文を表示 (en)

Our new speech-to-text model Ink-2 is out and #1 on Artificial Analysis. It’s built for streaming — low latency, fast eager mode and built in semantic endpoints to detect when users are done talking New architectures & algorithms made this Pareto-dominance possible

@cartesia
C
Cartesia@cartesia

Cartesia Ink-2 debuts as #1 for accuracy on the brand-new streaming speech-to-text leaderboard from @ArtificialAnlys! We designed Ink-2 from the ground up for voice agents - with low latency, eager transcripts, and semantic endpointing.

Photo 1