Back to people
@krandiash
K

Karan Goel

音声
@krandiash

founder ceo @cartesia, derp learned @stanfordailab @mldcmu @iitdelhi

21KFollowers1.0KFollowing2.1KPostsView on X

Recent posts

At Cartesia, our #1 value is "Get the fundamentals right". This value has been with us since Day 0. It shaped the research @_albertgu led during our PhDs: rather than using the same algorithms & architectures for every problem, he asked whether they were the right tool for the job. New algorithms will help us process the raw data in domains like audio, video, health, and robotics, and SSMs were born from that observation. At @cartesia, the same value is core to how we build models. We've obsessed over every detail in building our 10-layer cake (parfait?): data, systems, algorithms, post-training & RL, evals, inference, serving, platform, product surfaces, and customer feedback loops. These foundations enable us not only to produce world-class models, but to rapidly turn research advances into production-ready systems that scale reliably across millions of interactions. As a researcher myself, it's exciting to see that we're #1 on AA. From my perspective, the most exciting result is our leading position on the Controlled Voices benchmark. It's hard to game that benchmark because @ArtificialAnlys clones the same 8 voices on every model. It truly tests the model's ability to generalize. We're both #1 and #2 on that benchmark now with Sonic-3.6 and Sonic-3.5. We're just scratching the surface. Our latest algorithms haven't fully made their way into these models yet. The work is advancing faster than ever before, and we have exciting new results in model intelligence & data efficiency coming soon. Audio is the first step - it’s one of the simplest physical signals of the world. We’re tackling multimodal models next - models that will bridge the gap between knowledge & reasoning (text), and the physical world (other domains). Our core belief is that a fundamentally sound approach that uses the right algorithms will generalize to all the data arising from the dance of atoms in our universe. From audio to video, robotics, biology, and everything else. Expect many more exciting releases from us in the coming months.

@cartesia
C
Cartesia@cartesia

Introducing Sonic-3.6: our most lifelike TTS yet, and another step change in naturalness across 44 languages. Just three months after our Sonic-3.5 launch, we've made fundamental model improvements based on feedback from the teams building on Sonic. We're #1 on @ArtificialAnlys across both the provider and controlled voice streaming leaderboards — best-in-class voices + the best underlying model. Available in beta today. Hear it for yourself 👇

Photo 1Photo 2

good afternoon

Photo 1

We're partnering with leading AI natives, and it's great to see our models directly impact their business outcomes. @elise_ai and @minnasong are no exception, with an amazing mission to improve how we live using AI and voice systems, including right here in San Francisco!

@elise_ai
E
Elise Labs@elise_ai

We partnered with @a16z to look at data from the millions of renter conversations we handle every month. We found that AI-handled calls now last as long as human ones. Part of how we got there is by partnering with teams like @cartesia to improve latency and voice quality. We switched to Sonic 3.5, Cartesia's newest text-to-speech model, and saw nearly immediate results: a 2.9% lift in conversion and 12.2% increase in customer engagement.

Photo 1

A lot of custom post-training work being done today will be deprecated in favor of smarter models and continual learning 10-100x intelligence with continual learning is around the corner Compute, time and effort to heavily customize models is not a great use of resources while these capabilities are emerging. They will disrupt the assumptions people are building around today. It analogizes to adapting open source BERT variants to build models for language in 2019. LLMs disrupted this by solving the same problem through in-context learning with much less data and effort. 1. The base model will become fundamentally smarter, there are 10-100x improvements to be made in pre-training alone. 2. Continual learning algorithms fundamentally change the interface by which models are updated, how easy they are to maintain and use, and how much effort is required to adapt them to your domain and context. 3. The context and data sources from the domain, and the judgment of what good looks like will of course continue to matter. But even there, the level of data curation and feedback needed will shrink dramatically. Of course, it's not possible for folks to sit around and do nothing, but it's good to understand the challenges the current stack might be up against in the near future.

Every business is unique. Unique customers, unique products, unique markets you serve and a unique culture. Features like keyterm prompting are a nice and easy way to inject context to make your agents smarter at understanding your users & customers. Really cool feature

@cartesia
C
Cartesia@cartesia

New to Ink-2: keyterm prompting Pass rare terms like product names, industry jargon, and even French words for accurate transcription. 20% higher keyword recall, with no added latency. Learn more → http://cartesia.ai/blog/keyterm-prompting/?utm_source=x

Excited that @cartesia is in @cursor_ai's India campaign with some of the most ambitious founders of Indian origin (including in Delhi where I grew up) featuring amazing folks @amanrsanger, @vipulved, @tankots, @mukundjha, @ManishaRaisingh, @regards_rishi, @_sankyy

Photo 1

It's so exciting that there are genuinely amazing researchers starting new companies outside of big tech, taking swings at the biggest problems in AI The reason I ended up in AI was the deep infectious passion and enthusiasm in the community to solve the seemingly impossible

Then you melt all the stages and learn at test time — and the distinction between training and inference disappears in a lot of interesting scenarios

@elonmusk
E
Elon Musk@elonmusk

@aaronburnett 99% of compute long-term will be for inference

Nobody has really figured out how to connect the world of atoms to the world of bits Connecting knowledge & reasoning to the physical world really feels like the next big intelligence frontier A few new ideas are all that’s needed

.@j0nathanj is working on an incredibly important problem — making robots intuitive and intelligent the team is humble, creative and extremely brilliant this is their first release, really pumped about what they’re building

@Enigma_AI
E
Enigma@Enigma_AI

We put 100 real AI-powered robots online. Anyone in the world can control them right now, from a browser. Go make one do something:

Super cool collaboration with @paraga, @georgepickett and team. Access the web and bring in the freshest information to your voice agent while it talks and listens. All at un-parallel-ed speeds.

@cartesia
C
Cartesia@cartesia

Today, @cartesia and @p0 are making it possible for your voice agents to search the web at conversational speed. Cartesia builds the fastest voice models available, and Turbo mode for Parallel Search extends that same low latency to web search. An agent built on Cartesia can query the live web without interrupting the natural flow of conversation. Learn more 👇

.@_albertgu and I are organizing a research happy hour in London on Monday July 13th. Come hang out! https://luma.com/3p1u9pgf

The @cartesia team continues to kill it This eval measures quality when models mimic the same voices. Stronger models generalize better.

@ArtificialAnlys
A
Artificial Analysis@ArtificialAnlys

Announcing the Controlled Voice Arena Leaderboard comparing Text to Speech models on the same set of 8 cloned voices The Controlled Voice Arena standardizes, through voice cloning, the set of voices that each model’s performance is evaluated on - separating specific voice preference from broader aspects of model quality. It complements our Provider Voice Arena, where each model uses a select set of its own available voices. We have generated speech samples on models that offer voice cloning abilities using the same voice categories as our existing Provider Voice Arena, namely: 2 US Male voices, 2 US Female voices, 2 UK Male voices, 2 UK Female voices. Each model has been cloned on the same 1-2 minute audio recordings for each voice. Key results ➤ Overall: @cartesia Sonic 3.5 leads (1,122 Elo), followed by @ElevenLabs Eleven v3 (1,088) and @inworld_ai Realtime TTS-2 - Research Preview (1,070) ➤ US accent: Cartesia Sonic 3.5 leads (1,139 Elo), followed by ElevenLabs Eleven v3 (1,104) and Inworld Realtime TTS-2 - Research Preview (1,059) ➤ UK accent: Cartesia Sonic 3.5 also leads (1,103 Elo), with Inworld Realtime TTS-2 - Research Preview (1,075) moving ahead of ElevenLabs Eleven v3 (1,067) into 2nd ➤ Open weights: @FishAudio S2 Pro leads (1,034 Elo), followed by @MistralAI Voxtral TTS (1,024) and @resembleai Chatterbox (930) See more details below ⬇️

Photo 1

Shipped one of the most common ask from users: how does a voice change when you play it on different channels like telephony and the web? Try it out at http://play.cartesia.ai

Photo 1