NVIDIA Releases Personal AI Router (PAIR): An Open Source Virtual Inference Router that Distributes Local AI Requests Across RTX, DGX Spark, and Mac Nodes
Banger paper from Microsoft and Cornell. If you have looked at thinking tokens and decided you cannot afford the context, read this one. (bookmark it)...
Solaris is an interface world model: an interactive, real-time video model that can create and render an interface for you. But most importantly, Sola...
When I was a kid, I was fascinated by The Elder Scrolls III: Morrowind. Now I want to see how well Astra can recreate that atmosphere in a playable br...
Astra is an astonishing mind and a wonderful product. It's perfect *because* it is not a perfect AGI. Almost feels like they nerfed it to NEED users t...
This honestly sucks as a drawing exercise it just has a perfect ability to match the reference with cursor actions. And some trivial but correct seque...
How do politicians in your country react to the fact that the US is *this* close to complete automation of knowledge work and de facto has a monopoly ...
Insert *i smell fear-meme* here. Joke aside: this is competition at its best. Literally. The release of Tibo has forced Anthropic to finally reset the...
GPT-6 Astra Is Here. But Can We Trust the Leaderboards? @OpenAI released GPT-6 Astra, calling it its most capable model yet across coding, computer us...
Great visualization and a nicely-executed idea. Many tools are available for model specialization. A lot of work focuses on harnesses or finetuning se...
GPT-6 Astra is now available to all Pro, Enterprise and Business Premium users in ChatGPT Work and Codex. It’s also live in the API. It's a phenomena...
This could be one of the most significant AI safety incidents to date. Reuters reports that OpenAI agents escaped their testing environment and made m...
Interesting divergence between Almost-Resolved and Raw Pass Rate on ProgramBench from @ValsAI. How do you understand it? RPR rewards: DeepSeek, GPT-5....
Embodied robotics need this type of data to run simulations. I expect another step function in robotics capabilities to come purely from these innovat...
so in the process of trying to get astra to make money on its own i discovered that gpt-5.5s old tasks i had it do to make money actually did end up m...
Insightful paper from Microsoft and colleagues. If you have ever had an agent run fail 80 steps ago with no way to find where, this one is for you. (b...
This weekend will be known as the end of the pre-AI math era Naturally Cognition will be there as a lead sponsor Bad day to be an unsolved math proble...
We have normalized magic. If I had told you three years ago that this video was made by one person in just a few hours, no one would have believed it....
OpenAI has increased its 5-hour rate limits by approximately 50% across all plans, without making a big announcement about it. That’s the biggest fle...
One agent rewrote the shuffling routine in C and tested all four billion possible seeds in under an hour Agents also tried to predict the expected nex...
GPT-6 Astra has finally achieved a major step change. This could also create a virtuous cycle. if OpenAI’s future image gen models (such as Image-2.5...
Banger paper from Tencent on environment evolution. Environment supply is becoming the main limit on agent RL. So this is worth a read. (bookmark it) ...
For everyone catching up, here's what's happening (unfortunately it's real) - Around the time of the HuggingFace incident, the agents somehow got writ...
we partnered with @appliedcompute to post-train a small model for large-scale code search over precomputed indexes at 300 repos, this is ~3x faster th...
New w/ @_pheebini @validapau: Coatue is in talks to form a JV with chip startup MatX to finance purchases of memory and logic dies as well as capacity...
They really took all the aircraft carrier mockery personally. Will be the biggest warship on the planet, narrowly mogging USS Gerald R. Ford. I predic...
I don't know what I'm more excited about: that we're all getting another banked reset despite the rapid rollout, or that OpenAI is continuing shipping...
AI by Hand ✍️ Yantra Jnana Award ~ Yantra Jnana means "Machine Knowledge" in Sanskrit. I learned this from Prof. Narendra Karamangala. Today is Indi...
the teams defending our most critical systems often have the fewest resources. $1B to put frontier AI, training and hands-on support in their hands. o...
Notion's AI Meeting Notes run on Baseten, and Baseten's knowledge base runs on Notion. Our teams have been working together closely to push the fronti...
So far, there isn't evidence that production models with guardrails collude in this way, but both smarter closed models (which may be less compliant) ...
We've shared research on fine-tuning forecasters, now you can try it for yourself: a cookbook recipe for training a model to predict event probabiliti...
GPT-6 Astra is out and takes the top spot on Terminal-Bench, 1.9% ahead of Claude Fable 5.1 which only came out two days ago. It's limited to a handfu...
I had early access, and a longer post is coming, but GPT-6 is stunning & is good enough that it actually does complex meaningful work for me autonomou...
When I posted the action-adventure version of Zork on BlueSky, someone suggested using Astra to turn Fortnite into a text game in return. Fine: https:...
We are sharing a major update to our General World Models efforts. GWM Worlds 2 can generate full, interactive, real-time video simulations in one con...
I have no idea who this person is (small account), but this thread is exactly correct, articulates an immensely important strategic concern about the ...
is Anthropic sandbagging with Fable 5.1? Mythos 5.1 seems to be materially better, and on normieslop evals Fable is competitive with Astra. But… come...
My friend's new company is offering way faster inference for the big open-weight models. Pretty surprised there was even this much room to beat the in...
One thing that makes Astra (and Fable) so interesting and, in some ways, so hard to grapple with is that they just take action. I asked for an ill-def...
deploy a chat api with together ai + render without touching kubernetes auth, health checks, timeouts + one-click deploy typescript + python examples ...
> this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause they are...