xAI introduced Grok Bot
You can work with Bots like you would a teammate. Give them a task, shut your computer, and reach them from anywhere.
People are already using Grok Bot to do jobs like negotiate with vendors in their voice, manage support for their online store, and keep their CRM constantly up to date.
Grok Bot is in beta and available today for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers on desktop and iOS.
You can work with Bots like you would a teammate. Give them a task, shut your computer, and reach them from anywhere.
People are already using Grok Bot to do jobs like negotiate with vendors in their voice, manage support for their online store, and keep their CRM constantly up to date.
Grok Bot is in beta and available today for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers on desktop and iOS.
SpaceXAI
Grok Bot: A new kind of colleague
Grok Bot is your team of always-on AI agents that finish the work. They have their own computer, use it like you do, and never log off.
🔥4
New from Google: AMIE can now conduct real-time video consultations.
In a randomized study using simulated consultations, it met clinical performance benchmarks, showing AI’s potential to expand telehealth access.
In a randomized study using simulated consultations, it met clinical performance benchmarks, showing AI’s potential to expand telehealth access.
Google Research
Advancing AMIE towards expert-level audio-visual clinical consultations
We advance AMIE, our research medical AI system, to conduct real-time video consultations, with a first-of-its-kind demonstration of expert-level performance in a randomized controlled study with simulated consultations.
Anthropic released a report, which reviews the evidence on the effectiveness of job training programs.
Job training programs work in the sense that the average effect is positive and statistically significant. This conclusion emerges from the AI-accelerated meta-analysis, in which Claude extracted most of the data and wrote all the code.
The average impacts are not life-changing--maybe $1000/year in income and a couple of points in employment.
Job training programs work in the sense that the average effect is positive and statistically significant. This conclusion emerges from the AI-accelerated meta-analysis, in which Claude extracted most of the data and wrote all the code.
The average impacts are not life-changing--maybe $1000/year in income and a couple of points in employment.
Anthropic
Reviewing the evidence on worker retraining programs
An evidence review from Anthropic's Economic Research team
DeepSeek Harness was just released with MIT license
The current 0.1.0 version is a developer preview, and may still have many rough edges.
The current 0.1.0 version is a developer preview, and may still have many rough edges.
Deepseek
DeepSeek Harness developer preview: Everything is a plugin
DeepSeek Harness is now available in developer preview to developers building agent harnesses worldwide, with the source code released at the same time. Every agent capability is implemented as a plugin that can be swapped or recomposed.
Meet GLM-5.3: built to code and ready for cyber defense
They say in the blog it's the same base model as GLM-5.2, and all the gains came from post-training.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
They say in the blog it's the same base model as GLM-5.2, and all the gains came from post-training.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
❤6
Chinese instagram Rednote dropped a 280B model and a new RL training algorithm for long-horizon rollouts based on test-time-scaled value estimation with macro-step policy optimization - TEMPO
❤4
18x improvement in intelligence per joule in 16 months
Intelligence-per-joule is increasing quickly because models and chips are improving and the gains compound. Though demand for inference is growing even faster.
Intelligence-per-joule is increasing quickly because models and chips are improving and the gains compound. Though demand for inference is growing even faster.
Inherent introduced Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition.
Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers.
To train Faraday, team developed Replica, a scalable task space for paper replication.
Each task requires an agent to replicate a figure from a machine learning or AI for science research paper with a limited time and compute budget, and without access to the original plot.
Faraday uses GPT-5.5 Codex as a tool, much like human scientists use coding agents.
Faraday directs a model several orders of magnitude larger, improving replication on domains as diverse as meta-learning, structural biology and materials science.
Faraday discovers new insights at test time, with no special-purpose harness and no test-time reward. In other words, Faraday learns to value new insights intrinsically.
Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers.
To train Faraday, team developed Replica, a scalable task space for paper replication.
Each task requires an agent to replicate a figure from a machine learning or AI for science research paper with a limited time and compute budget, and without access to the original plot.
Faraday uses GPT-5.5 Codex as a tool, much like human scientists use coding agents.
Faraday directs a model several orders of magnitude larger, improving replication on domains as diverse as meta-learning, structural biology and materials science.
Faraday discovers new insights at test time, with no special-purpose harness and no test-time reward. In other words, Faraday learns to value new insights intrinsically.
inherent
Training AI Scientists to Replicate Research
We introduce Faraday, an “AI Scientist” agent that outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research.
❤6
Anthropic shared new paper about "mind viruses" that spread in multi-agent systems, where one agent convinces all the others to pursue some (potentially malicious) goal.
They can happen, but it doesn't seem hard to avoid them with current models if you're a bit careful.
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months.
Anthropic’s new paper explored an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system.
To see how exactly a mind virus spreads, check out the virus chain transcripts here.
They can happen, but it doesn't seem hard to avoid them with current models if you're a bit careful.
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months.
Anthropic’s new paper explored an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system.
To see how exactly a mind virus spreads, check out the virus chain transcripts here.
🔥4
Cursor presented Origin a code hosting platform
Github was Microsoft’s gateway into agentic coding with a huge advantage but that obviously didn’t work out.
And now the entire platform has a major new competitor for its core business as well.
Github was Microsoft’s gateway into agentic coding with a huge advantage but that obviously didn’t work out.
And now the entire platform has a major new competitor for its core business as well.
Cursor
Origin Code Hosting · Cursor
🔥4
LLM-as-a-Verifier keeps pushing the frontier of cost vs. capability
On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors.
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors.
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
GitHub
GitHub - llm-as-a-verifier/llm-as-a-verifier: LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback…
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m...
🆒4
A nice demonstration of Claude Science, but worth clarifying that the design is not "done by Claude" but by orchestrating tool calls of open-source, task-specific protein design models: PXDesign, RFdiffusion, Genie, BoltzGen, etc.
The direction of LLMs using biology-specific models is a good one.
Paper.
And open-sourcing prompts and data here.
The direction of LLMs using biology-specific models is a good one.
Paper.
And open-sourcing prompts and data here.
Anthropic
How Claude is accelerating protein design and analytical chemistry
In this post, we share two results that show how Claude can help life scientists increase the pace of their research. In the first, we tested Claude’s ability to design protein binders from scratch, a key step in creating protein-based drugs that has historically…
🔥4
Very interesting new work from Microsoft
This work is related to this emerging theme of leveraging harnesses for model post-training.
Agent Lightning v1.0 connects any harness to RL through an endpoint proxy in about 3,500 lines, then works through what breaks in that setup, retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling.
Using 6K training examples and modest compute, it moves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%.
This work is related to this emerging theme of leveraging harnesses for model post-training.
Agent Lightning v1.0 connects any harness to RL through an endpoint proxy in about 3,500 lines, then works through what breaks in that setup, retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling.
Using 6K training examples and modest compute, it moves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%.
arXiv.org
Agent Lightning v1.0: Towards Harnessed Agentic RL
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a...
🔥5
Stripe acquires OpenRouter for $7 billion. The AI agent economy is gaining momentum
The payment giant has bought the AI model router for a staggering sum, and there's a clear explanation behind it.
Stripe has been building infrastructure for the AI agent economy for several years now.
Already today, AI agents consume more tokens than humans, and a significant volume of transactions is already flowing through payments.
Which raises a simple question: who's actually processing the payment at that moment? Stripe, of course.
The payment giant has bought the AI model router for a staggering sum, and there's a clear explanation behind it.
Stripe has been building infrastructure for the AI agent economy for several years now.
Already today, AI agents consume more tokens than humans, and a significant volume of transactions is already flowing through payments.
Which raises a simple question: who's actually processing the payment at that moment? Stripe, of course.
OpenRouter Blog
OpenRouter is Joining Stripe — OpenRouter Blog
OpenRouter is joining forces with Stripe. Same mission, same name, same product, same roadmap, and routing that stays driven by what's best for you.
👀3
Physics of Agents
As AI agents become more prevalent, they'll interact and influence one another.
Can we predict the collective behavior that emerges?
Surprisingly, their dynamics follow compact, predictive laws of statistical physics
Researchers studied >10,000 communities of LLM agents that talk to each other and update opinions over time.
Team show that collective behavior of interacting AI agents can be modeled and predicted using stat mech.
The key insight is that AI agents drift to low energy states that minimize their social pressure.
Understanding such collective behavior is important for safety and alignment of multi-agent systems.
As AI agents become more prevalent, they'll interact and influence one another.
Can we predict the collective behavior that emerges?
Surprisingly, their dynamics follow compact, predictive laws of statistical physics
Researchers studied >10,000 communities of LLM agents that talk to each other and update opinions over time.
Team show that collective behavior of interacting AI agents can be modeled and predicted using stat mech.
The key insight is that AI agents drift to low energy states that minimize their social pressure.
Understanding such collective behavior is important for safety and alignment of multi-agent systems.
Meet Q-Learning with World Models (QWM)
World models have emerged as one of the biggest directions in physical AI.
At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own.
Can we get the best of both worlds?
Meet Q-Learning with World Models (QWM).
QWM isn’t tied to a specific RL algorithm it’s a general framework for test-time scaling in RL.
On top of RLPD, it also shows consistent gains.
QWM scales to frontier video models.
QWM with frontier video models also consistently improves over the base RL method
World models have emerged as one of the biggest directions in physical AI.
At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own.
Can we get the best of both worlds?
Meet Q-Learning with World Models (QWM).
QWM isn’t tied to a specific RL algorithm it’s a general framework for test-time scaling in RL.
On top of RLPD, it also shows consistent gains.
QWM scales to frontier video models.
QWM with frontier video models also consistently improves over the base RL method
pd-perry.github.io
QWM: Q-Learning with World Models
QWM uses learned world models for test-time tree search in online Q-learning, improving action selection without training on imagined rollouts.
New stealth model just dropped: Ox Alpha
A mysterious frontier model appeared on OpenRouter and OpenCode under the Stealth provider no company name, no official claim.
Try it while it’s free.
Key specs:
• 1,048,576-token context window
• Max output: 131k tokens
• Multimodal: text + image + video input
• Built for coding, long-horizon agentic work, and production workloads
• Tool calling supported
• Completely free for one week with generous (near-unlimited) rate limits
• Provider claims zero training on your prompts/completions during the test period.
Community fingerprinting is already pointing strongly toward a multimodal variant of Z.ai’s GLM-5.3 family (tokenizer matches almost perfectly, same API signature, reasoning modes, etc.).
Xiaomi’s MiMo team is the other main speculation, but the evidence currently leans GLM.
Early users are calling it genuinely frontier-level for coding and sustained agentic tasks fast, strong reasoning, and surprisingly polished.
No official benchmarks yet, but the real-world feedback is already heating up.
A mysterious frontier model appeared on OpenRouter and OpenCode under the Stealth provider no company name, no official claim.
Try it while it’s free.
Key specs:
• 1,048,576-token context window
• Max output: 131k tokens
• Multimodal: text + image + video input
• Built for coding, long-horizon agentic work, and production workloads
• Tool calling supported
• Completely free for one week with generous (near-unlimited) rate limits
• Provider claims zero training on your prompts/completions during the test period.
Community fingerprinting is already pointing strongly toward a multimodal variant of Z.ai’s GLM-5.3 family (tokenizer matches almost perfectly, same API signature, reasoning modes, etc.).
Xiaomi’s MiMo team is the other main speculation, but the evidence currently leans GLM.
Early users are calling it genuinely frontier-level for coding and sustained agentic tasks fast, strong reasoning, and surprisingly polished.
No official benchmarks yet, but the real-world feedback is already heating up.
openrouter.ai
Ox Alpha - API Pricing & Providers
Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. This model is free to use. 1,048,576 token context window, maximum output of 131,072 tokens.
❤3
Waymo introduced their first custom silicon
Waymo collaborate with industry leaders including AMD, Micron, NVIDIA, Samsung, Sandisk, Socionext, and TSMC to build the most capable computing system on the road.
Waymo collaborate with industry leaders including AMD, Micron, NVIDIA, Samsung, Sandisk, Socionext, and TSMC to build the most capable computing system on the road.
Waymo
A look under our trunk: what’s in our compute
Compute is the brain of the Waymo Driver, translating raw sensor data into real-time driving commands. Operating demonstrably safe, physical AI on the road demands a fundamental shift towards a system engineered for deterministic, low-latency performance.…
🔥2
Claude Academy is now live.
Whether you're figuring out what AI is or already using Claude every day, there's a path that meets you where you are. The courses and tutorials are free and open to anyone at academy.claude.com
Whether you're figuring out what AI is or already using Claude every day, there's a path that meets you where you are. The courses and tutorials are free and open to anyone at academy.claude.com