Firecrawl just launched FIRE-1, a new agent-powered web-scraper.
It navigates complex websites, interacts with dynamic content, and fills forms to scrape the data you need.
It navigates complex websites, interacts with dynamic content, and fills forms to scrape the data you need.
Firecrawl
Announcing FIRE-1, Our Web Action Agent: Launch Week III - Day 2
Firecrawl's new FIRE-1 AI Agent enhances web scraping capabilities by intelligently navigating and interacting with web pages.
❤3🔥3👏2
Cohere released Embed 4, a SOTA multimodal embedding model to add frontier search and retrieval capabilities to AI apps
—128K-token context window
—Supports 100+ languages
—Optimized for data from regulated industries
—Up to 83% savings on storage costs
—128K-token context window
—Supports 100+ languages
—Optimized for data from regulated industries
—Up to 83% savings on storage costs
Cohere
Introducing Embed 4: Multimodal search for business | Cohere Blog
Embed 4 delivers state-of-the-art accuracy and efficiency, helping enterprises securely retrieve their multimodal data to build agentic AI applications.
🔥3❤2👏2
Alibaba just now introduced open-sourceWan2.1-FLF2V-14B - 14B-parameter large model for First-Last-Frame to video generation
Powered by data-driven training and DiT architecture with First-Last Frame conditional control:
‒ Perfectly replicates reference visuals
‒ Precise instruction-following
‒ Smooth transitions + real-world physics adherence
‒ Cinema-quality 720P output.
GitHub
HuggingFace
ModelScope
Powered by data-driven training and DiT architecture with First-Last Frame conditional control:
‒ Perfectly replicates reference visuals
‒ Precise instruction-following
‒ Smooth transitions + real-world physics adherence
‒ Cinema-quality 720P output.
GitHub
HuggingFace
ModelScope
wan.video
Wan AI: Leading AI Video Generation Model
Wan is an AI creative platform. It aims to lower the barrier to creative work using artificial intelligence, offering features like text-to-image, image-to-image, text-to-video, image-to-video, and image editing.
❤2🔥2👏2💯2
Yale and GoogleDeepMind introduced C2S‑Scale a family open-source LLMs trained to “read” and “write” biological data at the single-cell level.
Preprint
Preprint
❤2🔥2🥰2💯2
Economics of Minds: LLMs planning their own workload. A new paper “Self‑Resource Allocation in Multi‑Agent LLM Systems.”
A lightweight Planner beats a monolithic Orchestrator—faster, cheaper, and smarter multi‑agent coordination
A lightweight Planner beats a monolithic Orchestrator—faster, cheaper, and smarter multi‑agent coordination
arXiv.org
Self-Resource Allocation in Multi-Agent LLM Systems
With the development of LLMs as agents, there is a growing interest in connecting multiple agents into multi-agent systems to solve tasks concurrently, focusing on their role in task assignment...
🔥3❤2👏2
Meta FAIR is open sourcing Matrix (Multi-Agent daTa geneRation Infra and eXperimentation) under MIT license.
It is a versatile toolkit with a high-performance model-serving engine designed for large scale inference.
It integrates Slurm for resource management and Ray for distributed job execution. It leverages lower-level model serving engines such as vLLM, SGLang for efficient LLM inference, and support API-based services.
Code.
It is a versatile toolkit with a high-performance model-serving engine designed for large scale inference.
It integrates Slurm for resource management and Ray for distributed job execution. It leverages lower-level model serving engines such as vLLM, SGLang for efficient LLM inference, and support API-based services.
Code.
Meta
Collaborative Reasoner: Self-improving Social Agents with Synthetic Conversations | Research - AI at Meta
With increasingly more powerful large language models (LLMs) and LLM-based agents tackling an ever-growing list of tasks, we envision a future where...
👍4🔥3❤2
Google dropped Gemini 2.5 Flash, a reasoning model matching o4-mini in preview
It's 'thinking budget' (up 24k tokens), can balance between answer quality, cost, and speed
The model performs particularly well on reasoning, STEM, and visual reasoning
It's 'thinking budget' (up 24k tokens), can balance between answer quality, cost, and speed
The model performs particularly well on reasoning, STEM, and visual reasoning
Google
Google AI Studio
The fastest path from prompt to production with Gemini
🔥3❤2🆒2👏1
Meta released AI research to push the boundaries of machine intelligence (AMI)
1. Meta Perception Encoder. Powering advanced computer vision for tasks like image recognition and object detection.
2. 3D Scene Understanding. Smarter AI that can locate objects from natural language queries.
3. Collaborative Reasoner. A framework to boost the reasoning skills of large language models, paving the way for collaborative AI agents.
These open-source advancements bring us closer to machines that perceive and decide like humans.
1. Meta Perception Encoder. Powering advanced computer vision for tasks like image recognition and object detection.
2. 3D Scene Understanding. Smarter AI that can locate objects from natural language queries.
3. Collaborative Reasoner. A framework to boost the reasoning skills of large language models, paving the way for collaborative AI agents.
These open-source advancements bring us closer to machines that perceive and decide like humans.
Meta AI
Advancing AI systems through progress in perception, localization, and reasoning
Meta FAIR is releasing several new research artifacts that advance our understanding of perception and support our goal of achieving advanced machine intelligence (AMI).
🔥4❤3👏2
The_state_of_AI_1745227157.pdf
5.4 MB
McKinsey's Latest State of AI Report: Organizations Are Finally Moving From Experimentation to Value Creation
The March 2025 State of AI report is out, and it captures a pivotal moment in AI adoption. Here's what's most significant:
A few takeaways:
• 78% of organizations now use AI in at least one business function (up from 72% in early 2024)
• 71% regularly use generative AI (up from 65% six months ago)
• Only 21% have fundamentally redesigned workflows to incorporate gen AI - but these organizations are seeing the most value
The report highlights what separates organizations that are creating value from those still experimenting:
CEO-level oversight of AI governance has the strongest correlation with bottom-line impact
Workflow redesign - not just technology adoption - is critical for value creation
Tracking specific KPIs for AI solutions drives measurable results
Workforce Impact:
• AI skills shortage is easing: fewer organizations report difficulty hiring AI talent compared to previous years
• 53% of C-level executives regularly use gen AI at work (vs. 44% of middle managers)
• 38% predict gen AI will have little effect on workforce size in the next 3 years
While progress is significant, over 80% of organizations still don't see material impact on enterprise-level EBIT from gen AI. This underscores that we're still in the early stages of this transformation.
The March 2025 State of AI report is out, and it captures a pivotal moment in AI adoption. Here's what's most significant:
A few takeaways:
• 78% of organizations now use AI in at least one business function (up from 72% in early 2024)
• 71% regularly use generative AI (up from 65% six months ago)
• Only 21% have fundamentally redesigned workflows to incorporate gen AI - but these organizations are seeing the most value
The report highlights what separates organizations that are creating value from those still experimenting:
CEO-level oversight of AI governance has the strongest correlation with bottom-line impact
Workflow redesign - not just technology adoption - is critical for value creation
Tracking specific KPIs for AI solutions drives measurable results
Workforce Impact:
• AI skills shortage is easing: fewer organizations report difficulty hiring AI talent compared to previous years
• 53% of C-level executives regularly use gen AI at work (vs. 44% of middle managers)
• 38% predict gen AI will have little effect on workforce size in the next 3 years
While progress is significant, over 80% of organizations still don't see material impact on enterprise-level EBIT from gen AI. This underscores that we're still in the early stages of this transformation.
❤4🔥3👏2
Nari Labs released Dia, a SOTA open-source text-to-speech model
Built with zero funding, the 1.6B param AI supports emotional tones, multiple speakers, and nonverbal cues
Plus, it beats leaders like ElevenLabs Sesame.
GitHub
HuggingFace
Waitlist.
Built with zero funding, the 1.6B param AI supports emotional tones, multiple speakers, and nonverbal cues
Plus, it beats leaders like ElevenLabs Sesame.
GitHub
HuggingFace
Waitlist.
GitHub
GitHub - nari-labs/dia: A TTS model capable of generating ultra-realistic dialogue in one pass.
A TTS model capable of generating ultra-realistic dialogue in one pass. - nari-labs/dia
🔥3❤2🥰2👏1
Researchers introduced Robotic Tactile Simulation
Taccel Simulator is a high-performance simulation platform for vision-based tactile sensors and robots.
Boosted by Nvidia Warp, researchers optimized Taccel with highly parallelized simulations and support 900fps simulation with 4k+ parallel training envs.
Taccel is designed with user-friendly APIs and is easy to use. Open-sourced all the code and documentation.
Preprint
Code
Taccel Simulator is a high-performance simulation platform for vision-based tactile sensors and robots.
Boosted by Nvidia Warp, researchers optimized Taccel with highly parallelized simulations and support 900fps simulation with 4k+ parallel training envs.
Taccel is designed with user-friendly APIs and is easy to use. Open-sourced all the code and documentation.
Preprint
Code
arXiv.org
Taccel: Scaling Up Vision-based Tactile Robotics via...
Tactile sensing is crucial for achieving human-level robotic capabilities in manipulation tasks. As a promising solution, Vision-Based Tactile Sensors (VBTSs) offer high spatial resolution and...
🔥3🥰3👏2
This is the paper with the most Huawei Fellow authors in the company’s history
The core question it addresses is straightforward: How do you efficiently and cost-effectively build a training cluster for large-scale models that supports tens of thousands or even hundreds of thousands of AI chips?
The core question it addresses is straightforward: How do you efficiently and cost-effectively build a training cluster for large-scale models that supports tens of thousands or even hundreds of thousands of AI chips?
👍3❤2👏2
DeepMind built an AI that replicates how a fruit fly walks, flies, and behaves — using MuJoCo, its open physics simulator
The work will help scientists understand what drives specific behaviors in the fly, finding links that labs can't always measure.
DeepMind applied this approach to multiple organisms – a virtual rodent, and now a fruit fly.
So what comes next for neuroscientists? The zebrafish – a widely studied creature which shares 70% of its protein-coding genes with humans.
GitHub
The work will help scientists understand what drives specific behaviors in the fly, finding links that labs can't always measure.
DeepMind applied this approach to multiple organisms – a virtual rodent, and now a fruit fly.
So what comes next for neuroscientists? The zebrafish – a widely studied creature which shares 70% of its protein-coding genes with humans.
GitHub
Nature
Whole-body physics simulation of fruit fly locomotion
Nature - A detailed whole-body model of the fruit fly, developed using a physics-based simulation and deep reinforcement learning, accurately replicates real fly behaviour.
👏3❤2💯2🔥1
OpenAI now projects $125B in revenue in 2029, with $25B of that from new products not yet announced
The Information reports OpenAI forecasts revenue reaching $125 billion in 2029 and $174 billion in 2030, mainly from AI agents, subscriptions, monetizing free users, and potentially affiliate fees.
According to internal documents seen by The Information, OpenAI expects revenue in 2029 to include $29 billion from AI agents, $50 billion from ChatGPT subscriptions, $22 billion from API access, and $25 billion from monetizing free users and other new, unspecified products.
CEO Sam Altman mentioned recently affiliate fees or taking a percentage of sales generated through user searches as possible revenue sources, while CFO Sarah Friar told the Financial Times there are “no active plans” for selling traditional advertising.
If they hit it, the current valuation ($300B) will be a steal; Google does ~$400B in revenue and is worth $2T.
The Information reports OpenAI forecasts revenue reaching $125 billion in 2029 and $174 billion in 2030, mainly from AI agents, subscriptions, monetizing free users, and potentially affiliate fees.
According to internal documents seen by The Information, OpenAI expects revenue in 2029 to include $29 billion from AI agents, $50 billion from ChatGPT subscriptions, $22 billion from API access, and $25 billion from monetizing free users and other new, unspecified products.
CEO Sam Altman mentioned recently affiliate fees or taking a percentage of sales generated through user searches as possible revenue sources, while CFO Sarah Friar told the Financial Times there are “no active plans” for selling traditional advertising.
If they hit it, the current valuation ($300B) will be a steal; Google does ~$400B in revenue and is worth $2T.
👍3❤2👏2
Trends in AI Supercomputers
Epoch AI dropped a map of the world’s 500+ AI supercomputers.
• Performance doubling every 9 months.
• Hardware cost & power use doubling every year.
• xAI’s Colossus already gulps 300 MW — equal to 250 k homes — and that’s only 2025.
If the trendline holds, the 2030 front-runner will burn 9 GW, pack 2 M chips, and sport a $200 B price tag.
The U.S. owns 75 % of today’s compute muscle while China trails at 15 %.
Epoch AI dropped a map of the world’s 500+ AI supercomputers.
• Performance doubling every 9 months.
• Hardware cost & power use doubling every year.
• xAI’s Colossus already gulps 300 MW — equal to 250 k homes — and that’s only 2025.
If the trendline holds, the 2030 front-runner will burn 9 GW, pack 2 M chips, and sport a $200 B price tag.
The U.S. owns 75 % of today’s compute muscle while China trails at 15 %.
🔥7❤3👏2