New Mistral Cookbook: a Multi-Agent Earnings Call Analysis System that turns lengthy and complex financial discussions into clear, actionable insights in minutes.
⚡6❤2🔥2👏1
Google DeepMind has started hiring for post AGI research
👍6❤2🔥2
OpenAI published 3 new guides:
AI in the Enterprise
A practical guide to building AI agents
Identifying and scaling AI use cases.
AI in the Enterprise
A practical guide to building AI agents
Identifying and scaling AI use cases.
🆒6❤3👍3🥴2
Veo 2, Google’s SOTA video model, is rolling out to Gemini Advanced + Whisk
You can create 8s, high-res videos from text prompts fluid character movement + lifelike scenes across a range of styles.
Tip: the more detailed your description, the better.
Plus, you can try Veo 2 using Whisk from Google labs.
Just input images, blend them together, and – now – “animate” to bring your creation to life. Available for all Google One AI Premium subscribers today.
You can create 8s, high-res videos from text prompts fluid character movement + lifelike scenes across a range of styles.
Tip: the more detailed your description, the better.
Plus, you can try Veo 2 using Whisk from Google labs.
Just input images, blend them together, and – now – “animate” to bring your creation to life. Available for all Google One AI Premium subscribers today.
Google
Generate videos in Gemini and Whisk with Veo 2
You can now generate videos in Gemini, powered by Veo 2.
👍4🦄3❤2👏1
Convergent Research released map of the things that need solving in science and R&D
gap-map.org is a tool to help you explore the landscape of R&D gaps holding back science - and the bridge-scale fundamental development efforts that might allow humanity to solve them, across almost two dozen fields
gap-map.org is a tool to help you explore the landscape of R&D gaps holding back science - and the bridge-scale fundamental development efforts that might allow humanity to solve them, across almost two dozen fields
❤7🔥2👏2
Goodfire released the first open-source sparse autoencoders (SAEs) trained on DeepSeek's 671B parameter reasoning model, R1—giving a new tools to understand and steer model thinking.
Reasoning models like DeepSeek R1, OpenAI’s o3, and Anthropic’s Claude 3.7 are changing how we use AI, providing more reliable and coherent responses for complex problems. But understanding their internal mechanisms remains challenging.
At Goodfire, platform for mechanistic interpretability can reverse engineer AI to understand internal representations and reasoning steps. Their interpreter models (e.g., SAEs in this instance) act as a microscope, revealing how R1 processes and responds to information.
Early insights from SAEs:
- Effective steering must wait until after phrases like “Okay, so the user has asked a question about…”—not explicit tags like "<think>"—highlighting unintuitive internal markers of reasoning
- Oversteering can paradoxically revert the model to original behaviors—hinting at deeper internal "awareness"
These insights suggest reasoning models differ fundamentally from non-reasoning language models.
GitHub
Reasoning models like DeepSeek R1, OpenAI’s o3, and Anthropic’s Claude 3.7 are changing how we use AI, providing more reliable and coherent responses for complex problems. But understanding their internal mechanisms remains challenging.
At Goodfire, platform for mechanistic interpretability can reverse engineer AI to understand internal representations and reasoning steps. Their interpreter models (e.g., SAEs in this instance) act as a microscope, revealing how R1 processes and responds to information.
Early insights from SAEs:
- Effective steering must wait until after phrases like “Okay, so the user has asked a question about…”—not explicit tags like "<think>"—highlighting unintuitive internal markers of reasoning
- Oversteering can paradoxically revert the model to original behaviors—hinting at deeper internal "awareness"
These insights suggest reasoning models differ fundamentally from non-reasoning language models.
GitHub
www.goodfire.ai
Under the Hood of a Reasoning Model
🔥7👍3
OpenAI released o3/o4-mini. The eval numbers are SOTA (2700 Elo is among the top 200 competition coders)
OpenAI’s team expect o3/o4-mini will aid scientists in their research.
And the secret trick is to talk to the models in images.
OpenAI’s team expect o3/o4-mini will aid scientists in their research.
And the secret trick is to talk to the models in images.
OpenAI
Introducing OpenAI o3 and o4-mini
Our smartest and most capable models to date with full tool access
🔥7👍2
Quantum computing firm Project Eleven has announced the “Q-Day Prize”, offering 1 BTC to the first team that can successfully break Bitcoin's ECDSA signature algorithm using a quantum computer with Shor's algorithm within one year.
The goal is to highlight the potential threat quantum technology poses to Bitcoin's cryptographic foundations.
The company estimates that around 6.2 million BTC (worth nearly $500 billion) could be vulnerable if such a breakthrough occurs.
The goal is to highlight the potential threat quantum technology poses to Bitcoin's cryptographic foundations.
The company estimates that around 6.2 million BTC (worth nearly $500 billion) could be vulnerable if such a breakthrough occurs.
The Block
Quantum computing research firm Project Eleven is offering 1 BTC to anyone who can break Bitcoin's cryptography
“The Q-Day Prize is designed to take a theoretical threat from a quantum computer, and turn that into a concrete model," CEO Pruden said.
❤4🔥3👏2🥴2💯1
Firecrawl just launched FIRE-1, a new agent-powered web-scraper.
It navigates complex websites, interacts with dynamic content, and fills forms to scrape the data you need.
It navigates complex websites, interacts with dynamic content, and fills forms to scrape the data you need.
Firecrawl
Announcing FIRE-1, Our Web Action Agent: Launch Week III - Day 2
Firecrawl's new FIRE-1 AI Agent enhances web scraping capabilities by intelligently navigating and interacting with web pages.
❤3🔥3👏2
Cohere released Embed 4, a SOTA multimodal embedding model to add frontier search and retrieval capabilities to AI apps
—128K-token context window
—Supports 100+ languages
—Optimized for data from regulated industries
—Up to 83% savings on storage costs
—128K-token context window
—Supports 100+ languages
—Optimized for data from regulated industries
—Up to 83% savings on storage costs
Cohere
Introducing Embed 4: Multimodal search for business | Cohere Blog
Embed 4 delivers state-of-the-art accuracy and efficiency, helping enterprises securely retrieve their multimodal data to build agentic AI applications.
🔥3❤2👏2
Alibaba just now introduced open-sourceWan2.1-FLF2V-14B - 14B-parameter large model for First-Last-Frame to video generation
Powered by data-driven training and DiT architecture with First-Last Frame conditional control:
‒ Perfectly replicates reference visuals
‒ Precise instruction-following
‒ Smooth transitions + real-world physics adherence
‒ Cinema-quality 720P output.
GitHub
HuggingFace
ModelScope
Powered by data-driven training and DiT architecture with First-Last Frame conditional control:
‒ Perfectly replicates reference visuals
‒ Precise instruction-following
‒ Smooth transitions + real-world physics adherence
‒ Cinema-quality 720P output.
GitHub
HuggingFace
ModelScope
wan.video
Wan AI: Leading AI Video Generation Model
Wan is an AI creative platform. It aims to lower the barrier to creative work using artificial intelligence, offering features like text-to-image, image-to-image, text-to-video, image-to-video, and image editing.
❤2🔥2👏2💯2
Yale and GoogleDeepMind introduced C2S‑Scale a family open-source LLMs trained to “read” and “write” biological data at the single-cell level.
Preprint
Preprint
❤2🔥2🥰2💯2
Economics of Minds: LLMs planning their own workload. A new paper “Self‑Resource Allocation in Multi‑Agent LLM Systems.”
A lightweight Planner beats a monolithic Orchestrator—faster, cheaper, and smarter multi‑agent coordination
A lightweight Planner beats a monolithic Orchestrator—faster, cheaper, and smarter multi‑agent coordination
arXiv.org
Self-Resource Allocation in Multi-Agent LLM Systems
With the development of LLMs as agents, there is a growing interest in connecting multiple agents into multi-agent systems to solve tasks concurrently, focusing on their role in task assignment...
🔥3❤2👏2
Meta FAIR is open sourcing Matrix (Multi-Agent daTa geneRation Infra and eXperimentation) under MIT license.
It is a versatile toolkit with a high-performance model-serving engine designed for large scale inference.
It integrates Slurm for resource management and Ray for distributed job execution. It leverages lower-level model serving engines such as vLLM, SGLang for efficient LLM inference, and support API-based services.
Code.
It is a versatile toolkit with a high-performance model-serving engine designed for large scale inference.
It integrates Slurm for resource management and Ray for distributed job execution. It leverages lower-level model serving engines such as vLLM, SGLang for efficient LLM inference, and support API-based services.
Code.
Meta
Collaborative Reasoner: Self-improving Social Agents with Synthetic Conversations | Research - AI at Meta
With increasingly more powerful large language models (LLMs) and LLM-based agents tackling an ever-growing list of tasks, we envision a future where...
👍4🔥3❤2
Google dropped Gemini 2.5 Flash, a reasoning model matching o4-mini in preview
It's 'thinking budget' (up 24k tokens), can balance between answer quality, cost, and speed
The model performs particularly well on reasoning, STEM, and visual reasoning
It's 'thinking budget' (up 24k tokens), can balance between answer quality, cost, and speed
The model performs particularly well on reasoning, STEM, and visual reasoning
Google
Google AI Studio
The fastest path from prompt to production with Gemini
🔥3❤2🆒2👏1
Meta released AI research to push the boundaries of machine intelligence (AMI)
1. Meta Perception Encoder. Powering advanced computer vision for tasks like image recognition and object detection.
2. 3D Scene Understanding. Smarter AI that can locate objects from natural language queries.
3. Collaborative Reasoner. A framework to boost the reasoning skills of large language models, paving the way for collaborative AI agents.
These open-source advancements bring us closer to machines that perceive and decide like humans.
1. Meta Perception Encoder. Powering advanced computer vision for tasks like image recognition and object detection.
2. 3D Scene Understanding. Smarter AI that can locate objects from natural language queries.
3. Collaborative Reasoner. A framework to boost the reasoning skills of large language models, paving the way for collaborative AI agents.
These open-source advancements bring us closer to machines that perceive and decide like humans.
Meta AI
Advancing AI systems through progress in perception, localization, and reasoning
Meta FAIR is releasing several new research artifacts that advance our understanding of perception and support our goal of achieving advanced machine intelligence (AMI).
🔥4❤3👏2
The_state_of_AI_1745227157.pdf
5.4 MB
McKinsey's Latest State of AI Report: Organizations Are Finally Moving From Experimentation to Value Creation
The March 2025 State of AI report is out, and it captures a pivotal moment in AI adoption. Here's what's most significant:
A few takeaways:
• 78% of organizations now use AI in at least one business function (up from 72% in early 2024)
• 71% regularly use generative AI (up from 65% six months ago)
• Only 21% have fundamentally redesigned workflows to incorporate gen AI - but these organizations are seeing the most value
The report highlights what separates organizations that are creating value from those still experimenting:
CEO-level oversight of AI governance has the strongest correlation with bottom-line impact
Workflow redesign - not just technology adoption - is critical for value creation
Tracking specific KPIs for AI solutions drives measurable results
Workforce Impact:
• AI skills shortage is easing: fewer organizations report difficulty hiring AI talent compared to previous years
• 53% of C-level executives regularly use gen AI at work (vs. 44% of middle managers)
• 38% predict gen AI will have little effect on workforce size in the next 3 years
While progress is significant, over 80% of organizations still don't see material impact on enterprise-level EBIT from gen AI. This underscores that we're still in the early stages of this transformation.
The March 2025 State of AI report is out, and it captures a pivotal moment in AI adoption. Here's what's most significant:
A few takeaways:
• 78% of organizations now use AI in at least one business function (up from 72% in early 2024)
• 71% regularly use generative AI (up from 65% six months ago)
• Only 21% have fundamentally redesigned workflows to incorporate gen AI - but these organizations are seeing the most value
The report highlights what separates organizations that are creating value from those still experimenting:
CEO-level oversight of AI governance has the strongest correlation with bottom-line impact
Workflow redesign - not just technology adoption - is critical for value creation
Tracking specific KPIs for AI solutions drives measurable results
Workforce Impact:
• AI skills shortage is easing: fewer organizations report difficulty hiring AI talent compared to previous years
• 53% of C-level executives regularly use gen AI at work (vs. 44% of middle managers)
• 38% predict gen AI will have little effect on workforce size in the next 3 years
While progress is significant, over 80% of organizations still don't see material impact on enterprise-level EBIT from gen AI. This underscores that we're still in the early stages of this transformation.
❤4🔥3👏2
Nari Labs released Dia, a SOTA open-source text-to-speech model
Built with zero funding, the 1.6B param AI supports emotional tones, multiple speakers, and nonverbal cues
Plus, it beats leaders like ElevenLabs Sesame.
GitHub
HuggingFace
Waitlist.
Built with zero funding, the 1.6B param AI supports emotional tones, multiple speakers, and nonverbal cues
Plus, it beats leaders like ElevenLabs Sesame.
GitHub
HuggingFace
Waitlist.
GitHub
GitHub - nari-labs/dia: A TTS model capable of generating ultra-realistic dialogue in one pass.
A TTS model capable of generating ultra-realistic dialogue in one pass. - nari-labs/dia
🔥3❤2🥰2👏1