Back to people
@robrombach
R

Robin Rombach

画像生成
@robrombach

Krawallkrümel. Generative Models at https://t.co/1xqMb617gc, made with ❤️

15KFollowers572Following798PostsView on X

Recent posts

Man I am so proud of our team. I love doing this work with our amazing group of people around the globe, while operating out of a headquarter in🇪🇺. More models are in the pipeline.

@arena
A
Arena.ai@arena

Last week, @bfl_ai launched their first video model, FLUX 3 Video. They have been testing an update on @arena, and the new version is ranked #2 in the Text-to-Video Arena! With 1496 pts, the updated FLUX 3 Video is just 16 pts behind the #1 spot, Gemini Omni Flash (1512 pts). This model is coming soon. Congrats to the @bfl_ai team on the release!

Man I am so proud of our team. I love doing this work with this amazing group of people, and out of a headquarter in🇪🇺. More models are in the pipeline.

@arena
A
Arena.ai@arena

Last week, @bfl_ai launched their first video model, FLUX 3 Video. They have been testing an update on @arena, and the new version is ranked #2 in the Text-to-Video Arena! With 1496 pts, the updated FLUX 3 Video is just 16 pts behind the #1 spot, Gemini Omni Flash (1512 pts). This model is coming soon. Congrats to the @bfl_ai team on the release!

🔥

@ostrisai
O
Ostris@ostrisai

I have playing with @bfl_ai Black Forest Labs FLUX 3 for a few days now. (Thank you for the early access!) And I played with it... a lot. Some may say, too much. 😅 Here are some of my favorite gens. Needless to say, I love this model. Audio on! More in 🧵

🤖🌲🐐

@elvisnavah
E
Elvis Nava@elvisnavah

Four years ago, in the early days of my PhD, I was obsessed with Stable Diffusion, and with the work of @robrombach, @pess_r, @andi_blatt and their team in image generation. You could say I had one north star: take these generative techniques and bring them to robotics. Today, it feels unreal to announce @mimicrobotics' official collaboration with @bfl_ai and to reveal FLUX-mimic, our next-generation Video-Action Model for general purpose dexterity. The bet behind it: control reduces to visual prediction. A model that can predict how a scene will unfold has already learned the physics of manipulation. So we built our Video-Action Model architecture on top of FLUX 3, the strongest video backbone available today. We're already testing and deploying it with manufacturing leaders like @AudiOfficial, on complex manipulation long considered impossible for conventional automation. FLUX-mimic is the culmination of years of work: the promise of true multimodality extended beyond the visual, into the domain of action. A new bet on general-purpose manipulation. And we're just getting started.

Yesterday, I had the insane experience of joining the G7 summit in Evian and present our perspective on open innovation in AI to the G7+ world leaders. With openness under pressure around the world, it is vital that we preserve a culture that makes open and responsible development the norm, not the exception My full remarks below. > Thank you, President Macron, for organizing this important conversation, and for bringing together this excellent group of people. I think I represent the youngest company in this room. I am Robin Rombach, co-founder and CEO of Black Forest Labs: a 90-person frontier AI lab, built on both sides of the Atlantic in Germany and the United States. We develop “world models”. These are visual AI models that understand and simulate our complex physical world. Many of you have probably never heard of us. But if you’ve ever generated an AI image or AI video—I’m sure some of you have done so—you’ve probably encountered our research. Not far from here, at universities in Germany, our team pioneered and published many of the techniques that gave rise to visual generative AI. Our team has co-developed 3 of the top 5 most popular open generative AI models of all time, including Stable Diffusion, and the only open models as popular as DeepSeek. We’ve talked a lot about language models today. They are remarkable. But language is a compressed, and ultimately limited, representation of the world. We are building models that learn from thousands of years of video, and can reason in complex visual environments—the rich, messy, complicated world we actually live in. These are still early days, but this technology is going to transform every corner of the economy. For that reason, we believe that visual models, alongside language models, will soon be critical economic infrastructure. That’s why we are passionately committed to open innovation. Sharing our technology openly means that businesses around the world can build their own AI systems, rather than renting them from a handful of companies. Open technology is vital for transparency, competition, and strategic independence in AI. A culture of open innovation made AI possible in the first place, including the transformer paper that gave birth to today’s language models, and we need to preserve that culture. But we are fully aware that open innovation poses unique challenges, especially in image, video, and audio. Open AI models can be misused or modified to produce unlawful and deeply harmful content, and it can be difficult to withdraw these models once they're released. Our commitment is to show that innovation in visual AI can be both open and responsible, not one or the other. This is partly a technical challenge, and we are making good progress. For example, our latest open models demonstrated over 10 times fewer vulnerabilities for sexual deepfakes and child abuse material than the open models released by other Big Tech firms. But there is a role for governments too, from safety tools to evaluation standards to targeted regulation. For example, we have welcomed the leadership of the Trump Administration, as well as the European Union and the UK Government, in developing new legal strategies to combat sexual deepfakes. It’s crucial that we strike the right balance. The future of our societies depends on getting this technology out into the world safely. But a climate of fear around open technology, or a focus on suppression over diffusion, will leave the world reliant on a handful of firms for critical infrastructure. We are ready to work with you all to make open and responsible innovation the norm, not the exception.

Photo 1

フォレストバイブスな Diffusion Circle! 🌲 元々の Diffusion Circle オーガナイザー @sedielem が今年 CVPR に直接参加できないので、彼と調整して、今回は僕たちが Diffusion Circle をホストすることになりました。 フロー マップ、マルチモーダル拡散モデル、変分フロー マッチング、アクション予測、その他ホットなデノイジングの話題に興味がある人は寄ってみてください。 @sumith1896@dustin_podell がコーディネートします。集合場所:Upper Lobby F / Exhibit Hall F Entrance(画像を参照)。 6月6日(土)午後3時30分開催!

原文を表示 (en)

Diffusion Circle with forest vibes! 🌲 Since the original Diffusion Circle organizer @sedielem cannot make it to CVPR in person this year, we coordinated with him and are happy to host the Diffusion Circle this time. Stop by if you are interested in flow maps, multimodal diffusion models, variational flow matching, action prediction, and all the other hot denoising stuff. @sumith1896 and @dustin_podell will coordinate. Meeting point: Upper Lobby F / Exhibit Hall F Entrance; see attached pics. Happening Saturday, June 6, 3:30pm!

Photo 1Photo 2

New paper out! We present a training method for multimodal generative models, called Self-Flow, which combines classic flow matching and representation learning. Why? Unlike most representation alignment methods, our new approach does not require external, pretrained models and thus scales gracefully to joint multimodal training on images, videos and audio. How? It combines per-timestep flow matching with dual-timestep representation learning, improving the models' internal representations. This approach outperforms prior methods and shows promising scaling behavior in multimodal pretraining. It also enables downstream applications such as action prediction for embodied AI. webpage+paper: https://t.co/qzGQGj8JYk code: https://t.co/edhfdVEqSf Credit to @hila_chefer, @pess_r, Dominik, @dustin_podell, Vikash, @Vinh_Suhi and Antonio. If you enjoy doing open research like this, come and join BFL! We are actively hiring🌲

Photo 1