Recent posts
this is already one of the most important papers of this year. https://www.latent.space/p/ainews-how-to-steal-a-reasoning-trace the methodology doesnt seem clearly explained so here are some notes with a further distillation

We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.

gpt luna max vs claude fable ultracode sent "pls build a mostly faithful clone of grok imagine with open models via fal" i woke up to these two and assumed fable was left and luna was right i was wrong... it was the other way!! objectively, fable did the better visual clone. but luna somehow understood intent better and created the more USABLE clone given my open model bent.


muse spark open source!!! PERSONAL SUPERINTELLIGENCE FOR EVERYONE


1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon. we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
comments like this on the aie channel miss the point. - we are building a community and an industry that is bigger than any one person can hold in their head. your slop is someone's aha moment and vice versa. - speakers spend quality time coming and presenting their strongest beliefs/entire year's work in 20-180 minutes - we spend millions on union AV labor and editing to get our speakers a public record that they can then send to customers, employees, and investors - our speakers are mostly engineers, researchers, academics and founders doing the work; not polished professional talking heads doing the circuit. most talks are prepped <1 week before. most have had ~0 public speaking training. - if you want the polished ppl, many other conferences select for people whose main job it is to be great speakers who give great talks - if you only judge quality by view count, you are guaranteed to be cooked by the algorithm. you will only ever hear about things after they are popular; worse; you consider things good only because they are popular. there are entire industries dedicated to manipulating you. do better. that said: - we CAN do a better job in curation. that's on me. - we CAN do a better job in coaching. also on me. - we CAN do a better job in production. that's on our team. - we COULD publish some talks to a secondary channel... I'm just concerned for those speakers as that will start form a smaller base, advice welcome, i am constantly pressured to do this every single year and have said no so far

occasional reminder to DELETE your skills. https://t.co/eyK7DprCso when you are bombarded constantly by "this skill changed my life!! you have to try!!" on the timeline, you will pile up stuff that at best just eats context, and at worst interacts with other skills nastily in unforeseen ways if you dont stare at your traces.
i shipped first set of llm-as-judge evals for the kill my saas competition tonight. people can run this to check if their solutions at least pass the sniff test.


$10,000 kill my saas in a weekend competition is live! TECH STACK: any coding agent any model up to $500 in token spend incl subscriptions see luma for description. join waitlist if late to this - finish line extended to Wednesday. the brief is live and people are prompting their clankers already. get to it!!!

i still think @AnthropicAI ultracode is one of the most important coding mode innovations ever invented. if you havent understood the potential of dynamic workflows you should try to. just met a Kill My SaaS competitor who did a pretty good submission in 3 ultracode prompts
reading thru applications. over 600 people applied, 100 admitted last night. we are going to kill SO MUCH SAAS




$10,000 kill my saas in a weekend competition is live! TECH STACK: any coding agent any model up to $500 in token spend incl subscriptions see luma for description. join waitlist if late to this - finish line extended to Wednesday. the brief is live and people are prompting their clankers already. get to it!!!

dear openai just make a new phone everyone wants openaiphone we can read 2-4x faster than we talk and speak openai alexa reachy hybrid is fine but pls just be a stepping stone to phone we want phone signed, everybody

NEW: OpenAI’s device - a human-like smart speaker - will look like a doughnut and be the size of a hockey puck. It meant to be held. It has a camera, speakers, microphones, lights and, most notably, moving parts to show interactivity. It’ll be $300-400. https://www.bloomberg.com/news/articles/2026-08-06/what-is-openai-s-device-a-doughnut-shaped-speaker-that-costs-over-300?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc4NjA0NjY3NSwiZXhwIjoxNzg2NjUxNDc1LCJhcnRpY2xlSWQiOiJUSjlNQ01UOU5KTFUwMCIsImJjb25uZWN0SWQiOiJDNEVEQ0FFMUZBMDU0MEJFQTI0QTlGMjExQzFFOTA4MCJ9.pj0oCNz7Ez90rn67tMWib-ed2PxcUAhAG2-hlVQ_DRg
current end state of forge


the implosion of git (as a protocol) in the next 2 years will be fun to watch
if you don't have a model that escaped sandbox during cybersecurity testing are you even a frontier lab anymore
man @waterloo_intern theses keep getting challenged* > megakernels are dead cursor drops monster kitty kernel > etched asics are dead amd buys taalas, and etched is valued at $10b by SK Hynix and TSMC lol whos gonna take on the remaining two theses? *note that he was making long term calls, these are short term datapoints

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a) re megakernels are dead why are megakernels useful? you spend two months writing a kernel to save time on launch overhead and poor inter-kernel overlap. you had PDL but then people said it wasn't perfect, that you could still get some marginal gains due to straggler CTAs and therefore- wait, sorry, I forgot, Rubin fixes that (kernel two needs 10 CTAs and kernel one has seven finished and three straggling, kernel two launches seven of its CTAs). given a long enough timeline, it all evens out. no serious inference provider is using a 67k loc hand-fused forward pass kernel in production, and the teams doing that are doing so out of pure research. dead. b) re ASICs are dead i'm sorry. to be specific: data-center transformer-inference ASIC companies (not naming any) who etched the arch into silicon have bet on architectural convergence. read kimi's architecture. read deepseek. qwen. we did not converge, and probably will not. dead. c) re gpu kernel dev is dead this one kind of hurts because it is (was) my job. gpu kernel optimization is the single most RL-able task in existence correct=check_correctness(kernel, shape) for shape in shapes if all(correct): time(kernel) give an agent ncu cli and an mcp with nvidia's tribal knowledge and it's done. dead. d) re NVIDIA is scared of AMD humans hate programming AMD. i'm sorry. it's just true. fine taking a performance hit as long as i don't have to touch rocm or a programming paradigm that says a warp is 64 threads (wtf?)...but an agent does not... so assuming software no longer moat, HBM capacity and bandwidth matter, and currently on perf / price they're goated. 'bUt NvIdIa iS gOaTeD oN hArDwArE sOfTwArE cOdEsIgN' and that's the new moat. watch how much tooling they open source to get kernel devs on nvidia. apologies all.



have you noticed an interesting correspondence between the plugins spec and the @harborframework spec... you know what happens next right



Build a plugin once and use it across compatible agent clients. Introducing Agent Plugins, an open standard developed with @awsdevelopers, @cursor_ai, @github, @code, and @vercel that packages Agent Skills and supports MCP server configurations in a shared format.
a very primitive form of the near term multiagent agi future is setting up one thread to ping back once its done so you create an implicit kanban/waterfall graph of dependent threads but each preserving their own work and agents i can't wait to setup a proper ui for this pattern but you can hack it together in most koding agents of choice rn





one way i'm developing Forge (https://t.co/2x9e0oRjxQ) is to use it to host all my projects going forward. so I often find myself having to bounce back and forth between platform and product. sharing neat trick - in @openai codex you can @ a thread + queue up the @, so if your project is blocked on a platform feature, you can actually premove your project to proceed once the platform is unblocked. now of course.. i wasnt actually strictly necessary in this whole process, so an even better multiagent harness would seamlessly orchestrate work back and forth between platform and project... it's just not a very common usecase unless you're building a real platform (which most people dont - most people just build applications on top of platforms, or a small platform with only 1 proudct/application tenant)



TIL Paul Erdős prompted his fellow mathematicians with bribes like we did the early LLMs

if you have been following his excellent work, @shloked has been breaking down every frontier labs' harness engineering for the last few months. excited to publish his deepest dive into ChatGPT yet as our newest guest on @Latentspacepod!


🆕 Unpacking ChatGPT Work https://t.co/wccYg8wY9K ChatGPT Work, an extension of the @openai Codex harness to cloud and general knowledge work, launched on July 9 and crossed 10M users 3 weeks later. @shloked guests with an A+ reconstruction of how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills and Tools work in this new harness that is bringing a full agentic experience to the almost 1B weekly active users of ChatGPT!
even the best founding team with all the money in the world is rate limited by bay area real estate smh


Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at: https://t.co/Rv3LMdLluK


smol aha moment at @_chenglou’s @midjourney meetup today - one reason that ontologies and graph knowledge are finally trending now rather than 3 years ago is that “good enough” intelligence is finally too cheap to meter this is why talks like @coyle_frankp and @emileifrem’s are on fire. the hardest part of KGs is cheap now. intelligence is commoditized so complements are increasing in value




Luna is such an insane value after the 80% cost reduction. It's basically free and can do a ton of real data processing type work. I'm overhauling title generation in T3 Code to take more advantage of it. I almost want to spin it up on every prompt to generate descriptions, feedback, and statuses. Why wouldn't I? It's basically free


