Back to people
@drfeifei
F

Fei-Fei Li

リサーチ
@drfeifei

Cofounder/CEO @theworldlabs, Prof (CS @Stanford), Co-Director @StanfordHAI, #AI #SpatialIntelligence #GenAI #computervision #robotics #AI-healthcare

957KFollowers1.2KFollowing3.0KPostsView on X

Recent posts

I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between @StanfordSVL and @NVIDIARobotics !

@DrJimFan
J
Jim Fan@DrJimFan

We scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot policies used to live their lives a few frames at a time (< 0.1 sec), instantly forgetting what just happened. We pushed to 3 orders of magnitude beyond SOTA. Introducing RoboTTT. Test-Time Training (“TTT”) carries a tiny model *inside* the model. Every incoming sensor reading triggers one gradient step on that tiny core, so the history keeps getting compressed into its weights. The hidden state has a fixed size (literally a small neural net), so the robot can “grok” arbitrarily long experience with little overhead. Learning continues indefinitely after deployment. We can then put an entire video in context as prompt! RoboTTT enables one-shot in-context learning from human video: in circuit board assembly, a human demonstrates a never-seen configuration once, and the robot imitates it faithfully. Humans drop things all the time, but we pick them up so fast that we don’t even notice. That reflex to fix is half of our physical competence. RoboTTT shows self-improvement on the fly: the robot is skilled at recovering from its own errors mid-episode, and each fix enters its context to inform the next move. The TTT core distills a general-purpose, failure-to-correction mapping from the training data. One more thing. What excites me the most is a new Context Scaling Curve: from 128 to 8K timesteps, closed-loop performance hill-climbs steadily with no sign of saturation. 8K-context pretraining beats 1K by 62%. What LLM enjoys, robotics should too. Soon, even 1M context is not a fantasy. Deep dive in thread:

58672165112KXで開く

AI’s next chapter will be defined not only by technical progress, but by how responsibly and thoughtfully we bring it into the world. I’m looking forward to be speaking on stage at Ai4 2026, Aug 4–6 in Vegas. Register here at https://ai4.io/register/ #Ai42026

Your thoughtful reflection is so inspiring and encouraging @smallfly ! As everyone talks about AI and automation, human creativity, story telling and productivity are even more important and essential to our society. @theworldlabs is founded on the premise of empowering human ingenuity and productivity. We are very grateful to be able to work with people like you! 🙏🌐

@smallfly
H
Hugues Bruyère@smallfly

@FastCompany just published a great piece on @theworldlabs , @drfeifei , Marble, and the idea that spatial intelligence / world models may be one of the next big shifts in AI. I was happy to be quoted in the article, but I also wanted to share more context about my own experience with World Labs and Marble, and why this direction is especially interesting to me. https://t.co/mdWBmSuNBe My starting point: volumetric capture — For the past few years I’ve been exploring and using volumetric capture and reconstruction (photogrammetry, NeRFs, 3D Gaussian Splats) mostly capturing locations around Montreal. Alleys, museums, urban interiors. I love every step of it: the capture itself, the pipeline, and what can be done with the output. Turning real spaces into real-time explorable systems. I do this personally, sharing explorations here, and professionally as chief technologist, and co-founder of Dpt. Physical reality + generative manipulation — In my work I’m especially drawn to mixing physical reality with generative and digital manipulation: using physical interfaces (light, clay, ink, ... ) to drive generative AI pipelines, building mixed reality prototypes that reshape your surroundings, or starting from real captured spaces and transforming them using tools like Marble. Like many people, I saw the World Labs announcement on Twitter in September 2024, and Marble when it surfaced in early December. But by then, I already had a sense something was coming. The first conversation — As someone deep into volumetric capture and radiance fields, I obviously knew about @BenMildenhall and his pioneering work on NeRF. To my surprise, Ben reached out to me in late June 2024. He’d been following some of my experiments and wanted to chat about my process and workflows and how I was using this “stuff” creatively. At that point he didn’t share what he was building, but we had a genuinely great conversation about radiance fields, AI, and my work. He was curious about the creative perspective, not just the technical one. When the World Labs announcement dropped a few months later, it all made sense. I understood what Ben had been working on, and why the creative angle mattered to them. Then in August 2025, he invited me to try the Marble beta, and I’ve been experimenting with it since. Experimenting with Marble — The first thing I used Marble for was materializing scene and world concepts during ideation at the studio, and seeing if and how it could fit into our production pipeline. In parallel, I dove into a series of experiments focused on world manipulation: starting from real captured spaces and transforming them using Marble. I’d already been exploring that idea using img2img diffusion with ControlNet on NeRF renders, real-time video streams, and even mixed reality using headset camera feeds. But Marble brings something different. It generates persistent, spatially cohesive 3D worlds that can be rendered in real time across a wide range of devices. That’s a real shift. Experiment 01: Parallel Realities — The first experiment, Parallel Realities, starts from a volumetric capture of a real location, reconstructed as 3D Gaussian Splats. Using Marble, I generate an alternate version of that same space, something informed by the original architecture: abandoned, nature-reclaimed, alternate era. Then, using Spark (World Labs’ 3D Gaussian Splatting renderer for THREE.js) I make both realities coexist in the same spatial coordinate system. From there, I use a portal UX mechanic to let the user step between the real reconstruction and the Marble-generated version. Experiment 02: Hidden Depth The second experiment, Hidden Depth, does not transform a space as much as expand it. A captured location has a visual boundary (a mural, a doorway, a dark corridor) and Marble generates what exists beyond it. For example: a Montreal alley has a painted mural; step through it and you’re inside a world informed by what is actually depicted there. World Labs showcased part of this work here: https://t.co/0RQTDWsgs2 And in their Spark 2.0 post: https://t.co/X34yzkLBOm The project page is here: https://t.co/T6Qxuuq9RJ Why this matters to me — Being able to start from a real 3D Gaussian Splat scene and manipulate it with Marble opens up a lot of ideas. The 3DGS pipeline is becoming an increasingly compelling foundation for exploration, experimentation, and storytelling. What matters most to me right now is more control. The more I can steer the generated scene or world, the more useful the tool becomes. I want more features like the already existing multiple input images and Chisel, the blockout-based approach. I would like better local control, the ability to expand a generated world more and more while preserving coherence, and the ability to directly import 3D Gaussian Splat scenes to be used as a starting point. I want more ways to shape the result, not just a “prompt and hope” approach. — It is exciting to see this field moving from research and demos toward actual creative workflows.

科学研究は文明を進歩させ、医学から材料科学、脳科学から物理学まで、世界中の人々が最も深刻な課題を解決するための基盤です。これが可能なのは、科学者がAIベースのツールを含む、最先端の研究ツールにアクセスできるときだけです。

原文を表示 (en)

Scientific research is fundamental to advancing civilization and helping people globally to solve the most critical problems, from medicine to materials, from brain science to physics, and much beyond. This is only possible when scientists have access to the best tools of the time to conduct scientific research, including having access to AI-based tools.

3.1K467378200KXで開く

創意工夫と想像力が本当に素晴らしい!@theworldlabs が @withloreco の才能あふれるチームとパートナーシップを組んで、彼らの素晴らしいアイデアをユーザーが楽しめるインタラクティブな体験に変えることができたこと本当に感謝です!🤩

原文を表示 (en)

The creativity and imagination is out of the world! So grateful that @theworldlabs got to partner with the amazing talents @withloreco to translate their incredible ideas into an interactive experiences for users to enjoy!🤩

@theworldlabs
W
World Labs@theworldlabs

We turned dreams into worlds. Then filled them with history's greatest minds. Not a video. A world, running directly in your browser. Step inside ↓

大規模生成モデルの現代に対応したビジュアル生成用のベンチマークデータセットに、とても興奮してます!🤩

原文を表示 (en)

I’m very excited by this new benchmark dataset for visual generation that is suitable for the modern era of large scale generative models!🤩

@keshigeyan
K
Keshigeyan Chandrasegaran@keshigeyan

1/ Introducing GPIC: a Giant Permissive Image Corpus and benchmark for visual generation! 🚀100M VLM-captioned image-text pairs for training 📊1M image-text pairs for benchmarking 🖼️~28 trillion pixels 🤗Centrally Hosted ✅Fully permissive for research + commercial use Dataset, benchmark and models🧵👇 Co-led with @KyleSargentAI

Photo 1

本当に誇りに思ってる @_amirabs @sadeghian_ali ; 一緒に仕事をしてくれて、@PlayAstrocade の進捗を見守れて本当に素晴らしいです!🚀

原文を表示 (en)

So proud of you guys @_amirabs @sadeghian_ali ; it's been wonderful working with you guys, and seeing the progress of @PlayAstrocade !🚀

@PlayAstrocade
A
Astrocade@PlayAstrocade

We raised $56M to help build the next era of interactive entertainment. Series B led by @sequoia, Series A led by Sea. Astrocade lets anyone create games with AI, play them with friends, and share them with millions. But this isn’t about replacing creativity. It’s about giving more people the tool to bring their taste, humor, stories, and craft to life. Today, the fun goes public.

It’s 11th year and counting! Teaching the first lecture of @cs231n every year has been a highlight of my spring seasons. As usual, I asked students which departments or schools they come from @Stanford . Increasingly, students raise their hands to indicate that they come from all seven schools on campus, from @StanfordEng to @StanfordMed @StanfordHumSci @StanfordGSB @StanfordLaw @StanfordEd @stanforddoerr . AI is truly a horizontal technology that excites students across all backgrounds and disciplines!🤩

Photo 1Photo 2