Back to people
@leerob
L

Lee Robinson

コーディング
@leerob

Model behavior @SpaceXAI. Helping train useful models.

281KFollowers872Following17KPostsView on X

Recent posts

How can you trust @Bot​s to operate autonomously? This is something we've carefully designed. First, we use an LLM to automatically review every action your bot takes. If anything is unexpected, it will ask for your approval. But this isn't enough. We also have customizable rules to allow or deny actions. For example, I have bots that are connected to my X, Gmail, etc. I explicitly block any write/destructive actions and have it ask for my approval first. This is a good balance between YOLO mode and reviewing everything by hand. Although, we do support YOLO if you like to live dangerously!

Photo 1
92444560102KXで開く

I thought the growth of Cursor was insane, but Grok @Bot is blowing me away. This feels like something special. Will be training Grok 4.7 to make sure it’s exceptional in the Bot harness!

3.4K119183292KXで開く

Grok @Bot has made a few simple yet powerful technical decisions that I believe make it easy and enjoyable to use. 1. The best UI is none at all. The product interface is dramatically simpler than alternatives without sacrificing functionality. How is this possible? It's one of the first products designed for current frontier model capabilities and has a UI restrained enough to remain easy to use as models improve exponentially. Everyone knows how to text. 2. A thin harness for the client, a thick harness for the server. You might have noticed the app feels very fluid to use, even for a beta product. This is primarily because of everything we didn't have to build. The app harness is essentially a single tool to send messages between the client and server. The complexity moves to the server, where you can still use the coding agent harness with specialized tools as needed. This helps make the UI fast and responsive on desktop and mobile. 3. An always-on computer. Most coding agents and assistants today start fresh with every question you ask. Some of these sessions are on your local machine and others happen in the cloud. We believe strongly that cloud is the future, which is why it's the only option. Further, rather than spinning up virtual machines for every conversation, your bots connect to their own computer. This means you can still run agents on the bot's persistent filesystem. It's closer to what programmers have been doing by using Tailscale from their phones to connect to a remote computer and run an agent TUI. You get those capabilities without the hassle. 4. Your bots can use the browser. Coding agents have shown that most work on a computer can be expressed and run as code. You can ask for a task in natural language and the agent will decide to write a script to complete it. This is amazing, but there's still many tasks which can't be completed without logging into a website and clicking around the browser. Models and harnesses are now good enough to reliably handle this. The combination of writing code and using browsers means you can automate almost any task on a computer. Further, you can ask Grok Bot to record you doing the task, and then turn it into something repeatable.

1.7K102898221KXで開く

Big day! Cursor has officially joined SpaceX. I’m proud of our team and what we’ve accomplished so far. Our momentum is accelerating with Grok 4.6 and Grok Bot as well. I’ll be working on making Grok useful, tasteful, and safe. Onward!

@cursor_ai
C
Cursor@cursor_ai

Cursor is now part of @SpaceX. Today, we have officially closed our acquisition. We will join the @SpaceXAI team to help make Grok the world's most useful AI and improve Grok Build, Grok Bot, Grok API, Cursor, and more. SpaceX has built some of the most inspiring and impressive technology in the world, and we’re grateful for the opportunity to become part of such a special company. Onwards.

4.4K114118209KXで開く

Gemini 3.7 Flash is now available in Cursor! Here are some of the results on our evals ↓

Photo 1

We've released the model card for Grok 4.6! It goes in-depth on the capabilities of the model across many different evals for coding, engineering, knowledge work, and more. We also discuss pre-deployment safety testing and our safeguard stack. https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf

Photo 1
1.5K100386346KXで開く

We’re making great progress on improving the writing quality and design taste of Grok. Still a lot more to do, but I’m encouraged by how quickly we’re iterating. Please let me know any feedback. Grok 4.6 soon!

4.8K169244289KXで開く

There's something missing from the open vs. closed models debate that has been bothering me. The better analogy, to me, is managed vs. self-hosted infrastructure for AI. Maybe a company wants to use an open weight model because they want to do additional training on top, bringing their data and domain knowledge. This requires software and services to do additional training (e.g. Tinker and friends) as well as to serve the model (e.g. SGLang). Not all of this software is open source today! Once you have successfully trained a model with your enterprise data and expertise, you now need to deploy and serve it for customers. You can partner with an inference company to run the software and hardware. But if ownership was your primary concern, you still want to control the hardware and storage, and you now also need to run infra and secure GPU capacity. There are other valid reasons to be open. In particular, the entire industry benefits when companies training models release data or research about their work. It also allows capitalism and free markets to do their thing, increasing competition and ultimately providing better options for customers. So we should all encourage openness. The reason I prefer the managed vs. self-hosted infra framing is that we can learn from the past decade of cloud infrastructure. It's important and healthy to have both, and a great self-hosted alternative ultimately pushes the managed versions to innovate. The decision to run infra then comes down to more standard business reasons: attracting talent, the cost and maintenance of the hardware, and the importance of uptime and reliability to the business. Many businesses will say, actually, I don't want to staff and run a training and inference team, and I'm happy to pay API pricing for intelligence. And others will do the opposite and invest heavily here. We need both! As an aside, the capability of open models will reach a point where we need to be very intentional about how they are deployed. But I think this problem is solvable, whether it is sharing research early and weights later, or also open sourcing the safety stack to properly serve the model. I don't have a perfect answer here but I think the ecosystem should figure it out together. Full disclaimer, I work at a company which has released both open and closed models. There are probably people more knowledgable than myself of the open weights ecosystem. If that's you, curious if you disagree with any of this.

Humans try hard things, fail, learn, and get better through repetition. AI models aren't as different as you might think. Here's how models "learn" explained in simple terms. http://x.com/i/article/2080455318128009217

2.0K2122.8K216KXで開く

Grok 4.5 is really good at React. It's also very affordable and token efficient!

Photo 1Photo 2
3.7K526279720KXで開く

My talk from AI Engineer is now live! It covers how we're automating parts of AI research and building systems to rapidly improve our models. I cover some of the work our team did to train Grok 4.5 together with SpaceXAI.

@willcb
W
will brown@willcb

incredibly chill fun laid-back talk from @leerob describing how cursor has fully solved RSI

We just doubled the included usage of Cursor models on all plans. Enjoy more access to Grok 4.5 and Composer 2.5!

7.0K366477459KXで開く

New open model from Thinking Machines! 1T MoE, 1M context, multimodal, some solid evals. Really nice blog post 👏

@thinkymachines
T
Thinking Machines@thinkymachines

Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵

I've been thinking about this post a lot. It represents a bigger shift in startup marketing to me. I don't want to see your copycat funding launch video. Cool, you raised a bazillion dollars. It's overplayed. Show me the product! Show me the customers!

@Etched
E
Etched@Etched

We're coming out of stealth. We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised. Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads. Our first racks ship this summer.

Photo 1
99422322163KXで開く

She’s found the ancient scrolls.

Photo 1

Are current LLMs incompatible with great creative writing? I can't tell if it's cope or not, but it seems like even with the best models, I still can't get them to write like humans would. For coding, there is a verifiable reward like it compiling or tests passing. But for creative work like writing, it's much more subjective. I have struggled to prompt / harness the models to write truly amazing work. They are fantastic for spell checking, grammar suggestions, and taking on different personas to read and critique work. Maybe it's because I'm only doing nonfiction, and to write something top 0.1% means that you need to think over a long horizon and develop an interesting insight about the world. Great writing is clear thinking. I've even asked models to try 10 different versions of a blog post, then have a council of models grade and critique the results and pick the best parts... and still I end up with this lowest common denominator slop. Skill issue? Someone show me the way.

You can now try Kimi K2.7 in Cursor! Results from our evals ↓ Interesting to see the comparison with GLM 5.2.

Photo 1

高品質なevalsを作ることはますます重要なスキルになっています。 特に転職や AI業界への進出を目指しているなら、あなたが興味を持つタスク/ドメインでモデルをベンチマークしてみることをお勧めします。 うまくやれば、モデルを学習している企業の注目を集めることができます。

原文を表示 (en)

Building high-quality evals is an increasingly important skill. Especially if you're trying to land a job or get into AI, I'd recommend trying to benchmark models on a task/domain you care about. If done well, you'll get the attention of any company training models.

@cursor_ai
C
Cursor@cursor_ai

We're sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.

Photo 1

CursorでGLM 5.2を試すことができるようになりました! より有用なオープンモデルが増えることに期待しています。Fireworksとのパートナーシップをありがとうございます。evals の結果 ↓

原文を表示 (en)

You can now try GLM 5.2 in Cursor! Excited to see more useful open models, thank you to Fireworks for partnering here. Results from our evals ↓

Photo 1
2.8K148330314KXで開く