Recent posts
The conjecture is wrong, here's an AI-generated counter example

Regular reminder -- the set of public ARC 3 games is called "demonstration set", not "eval set" nor "training set". It is not meant to be used as training data, and it is not meant to be used as an eval. Scores on the public demonstration set are not indicative of scores on the actual benchmark. The demonstration set is intended to demonstrate the format and drive human engagement. The private eval set is substantially more difficult and more novel. The top score on the Kaggle leaderboard today is 2.70% (that's the semi-private set -- at the end of the competition, the submissions are scored on the fully private set).

Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks off.
Builders respect builders. The loudest, most toxic haters are almost always the ones who have never built a thing -- the Nobody McPoasters.
The reports of the demise of Google are greatly exaggerated. I wouldn't underestimate them

JUST IN: Sergey Brin to reportedly take direct oversight of Gemini as Google restructures its AI leadership.
The Keras community call is starting now -- link to join in tweet below.

The Keras community meeting will take place this Friday at 10am PT -- the team will present the latest developments in the Keras ecosystem, in particular the new vLLM integration. Anyone can join the call. Please use this link https://t.co/tllqzfieXS to join when the meeting starts (10am Friday).

The Keras community call starts in 15 minutes!

The Keras community meeting will take place this Friday at 10am PT -- the team will present the latest developments in the Keras ecosystem, in particular the new vLLM integration. Anyone can join the call. Please use this link https://t.co/tllqzfieXS to join when the meeting starts (10am Friday).

With agentic AI, workflows are increasingly CPU hungry. The share of cognition moving to the CPU keeps increasing.

Scoop: AWS engineers have been told to conserve CPU compute to make sure the cloud giant has enough for its customers, with some having to wait days to have the compute they need. GPUs have long been scarce — now high demand for CPUs and memory is leading to new constraints
The Keras community meeting will take place this Friday at 10am PT -- the team will present the latest developments in the Keras ecosystem, in particular the new vLLM integration. Anyone can join the call. Please use this link https://t.co/tllqzfieXS to join when the meeting starts (10am Friday).

We need open frameworks to evaluate model behavior. Discussions need to be grounded in auditable measurements rather than "us vs them" vibes. @cyrilgorlla and the team at CTGT are doing important work in this space

@ReedAlbergotti broke it at @semafor this morning, and his question is the one that lingers: "What is the nationality of an American model distilled from a Chinese model that was distilled from an American models?" At the 8k token budgets production systems actually run, our 120B scores 83.61% on FinanceReasoning. Above Kimi K3 (81.93%) and Inkling (65.13%). At 62 to 160x lower cost per query, on one H100. At unlimited budget the big models win on raw accuracy.

Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents. 2. Fine: general-purpose API settings that were not developed for ARC-AGI-3 and that are available to all API users. In the past, we've had a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction. I'm glad they're starting to figure out the answer. Of course, if each provider uses different settings when getting their model tested, it creates a potential parity issue. My take is that this is fine as long as the settings and the cost are clearly reported.
A significant portion of current AI "discourse" is less about technological capabilities and more about frontier lab employees navigating their own self-esteem and sense of identity.
In science, you have to report the experiments that didn't work, not just the ones that did. Same with AI. (I wish.)
The innateness of personality is obvious in kids because they haven't yet had the time to undergo any meaningful environment-induced (or self-induced) change. They express their raw nature. However, I believe that experience shapes personality significantly over the long term, especially between the ages of ~15 and ~25. I also believe that people are capable of deliberately building themselves, choosing to alter their own personalities through a wide range of means. The person you were born as doesn't have to be the person you die as.

We have multiple reasons for wanting to believe nurture is more important than nature and none for wanting to believe the opposite (unless we = the far right, which we ≠), so it's not surprising we overestimate its effect. But man are kids' personalities inborn.
Most people are conditioned to expect that all known problems already have canonical solutions, that these solutions are the best that can be achieved, and that attempting to reinvent them would be a pointless, quixotic effort. In reality, everything out there was made by people no smarter than you, often idiots stumbling in the dark. Not only can new solutions be found, but entirely new paradigms are absolutely possible, including ones that completely bypass the current tech tree.
Opus 5 sets a new state-of-the-art on ARC-AGI-3, at 30%. ARC-AGI-3 measures solving problems with no prior exposure -- the setting where scaling has historically bought the least. Impressive jump!

On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model.

Not sure if people in SF use the phrase "xyz-shaped" all the time because LLMs speak like that, or the other way around
Always keep in mind that the rate of change matters more than the current metric value
AI competence has always been very spiky, superhuman in some narrow domains and largely useless in others. The fundamental marketing trick of the AI industry is to make you believe the tallest spike is a floor.
