DeepSeek just dropped another open-source AI model, Janus-Pro-7B
It's multimodal (can generate images) and beats OpenAI's DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks.
This comes on top of all the R1 hype.
It's multimodal (can generate images) and beats OpenAI's DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks.
This comes on top of all the R1 hype.
🔥15🆒3❤2
Alibaba released the SOTA open multimodal model, Qwen2.5-VL
It shows significant improvements across various aspects compared to the previous version.
Key Highlights:
1. Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all!
2. Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones.
3. Long Video Comprehension : Captures events in videos over 1 hour long
4. Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection.
5. Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more.
ModelScope.
It shows significant improvements across various aspects compared to the previous version.
Key Highlights:
1. Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all!
2. Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones.
3. Long Video Comprehension : Captures events in videos over 1 hour long
4. Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection.
5. Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more.
ModelScope.
chat.qwen.ai
Qwen Studio
Qwen Studio is an official platform from Qwen that empowers both everyday users and developers with unified access to Qwen’s series of open-source and proprietary models. It offers comprehensive functionality spanning chatbots, image and video understanding…
The key insights from Yann Lecun's speech at the #WEF2025
At a very interesting debate on technology he covered:
- The perspectives of AI in the next 3-5 years
- What intelligent AI systems need
- Why LLMs are not enough
- True diversity in AI
- Open source importance
and more
How do we make sure the AI we build is the AI we want?
Yann LeCun said that where we are now with Gen AI isn't where we want to be. Gen AI isn't very controllable, but researchers try to make it applicable to a wide range of areas.
The perspectives of AI:
In the next 3-5 years, we are going to see the emergence of a new revolutionary brand or paradigm for AI architectures.
- This new AI won't be generative as we see generative AI today.
- It will be able to plan the sequence of actions and the sequence of plans.
Then we'll come to agentic AI and then to robotics. And the coming decade may be the era of robotics.
4 things that are essential for intelligent AI systems (and are limitations for current systems):
- Understanding the physical world
- Having persistent memory
- Being capable of reasoning
- Complex planning capabilities
LLMs' place in the future of AI:
Yann Lecun considers LLMs as a part of a bigger future AI system, because they are good at manipulating language but bad at thinking.
We're never going to get to human-level AI just with text models. AI needs to know how the real world works from sensory data, as well.
Why is language simple?
Language, like DNA and proteins, is a discrete object, and it's easy to make predictions in discrete worlds.
Physical AI is much more difficult. That's why techniques used in LLMs can't be used for predicting videos.
Open source matters the most:
Yann Lecun highlighted that open source makes AI tool accessible to everyone.
It's important for cultural diversity and democracy because a wide diversity of AI assistants can only be achieved with open source.
How to achieve true diversity in AI:
True diversity is when we have models that are trained on all the languages, cultures, and values of the world.
That's where foundational models can be the base that is fine-tuned for different ideas and visions of what good value systems are, so people can choose from many options.
Federated learning:
One of the ways to achieve diversity is federated learning. It's when every region in the world has its own datasets which contribute to training a big global model.
How can AI be valued in terms of disinformation and toxic content?
With transformers and self-supervised learning appeared, the proportion of hate speech taken down by AI systems has reached 96%.
However, high false positives (good content removed) are a problem. Therefore, detection thresholds will be adjusted to enable discussions on key societal topics.
Regulations:
"Making the distribution of open source AI engine essentially illegal is way more dangerous than all the other potential dangers."
At a very interesting debate on technology he covered:
- The perspectives of AI in the next 3-5 years
- What intelligent AI systems need
- Why LLMs are not enough
- True diversity in AI
- Open source importance
and more
How do we make sure the AI we build is the AI we want?
Yann LeCun said that where we are now with Gen AI isn't where we want to be. Gen AI isn't very controllable, but researchers try to make it applicable to a wide range of areas.
The perspectives of AI:
In the next 3-5 years, we are going to see the emergence of a new revolutionary brand or paradigm for AI architectures.
- This new AI won't be generative as we see generative AI today.
- It will be able to plan the sequence of actions and the sequence of plans.
Then we'll come to agentic AI and then to robotics. And the coming decade may be the era of robotics.
4 things that are essential for intelligent AI systems (and are limitations for current systems):
- Understanding the physical world
- Having persistent memory
- Being capable of reasoning
- Complex planning capabilities
LLMs' place in the future of AI:
Yann Lecun considers LLMs as a part of a bigger future AI system, because they are good at manipulating language but bad at thinking.
We're never going to get to human-level AI just with text models. AI needs to know how the real world works from sensory data, as well.
Why is language simple?
Language, like DNA and proteins, is a discrete object, and it's easy to make predictions in discrete worlds.
Physical AI is much more difficult. That's why techniques used in LLMs can't be used for predicting videos.
Open source matters the most:
Yann Lecun highlighted that open source makes AI tool accessible to everyone.
It's important for cultural diversity and democracy because a wide diversity of AI assistants can only be achieved with open source.
How to achieve true diversity in AI:
True diversity is when we have models that are trained on all the languages, cultures, and values of the world.
That's where foundational models can be the base that is fine-tuned for different ideas and visions of what good value systems are, so people can choose from many options.
Federated learning:
One of the ways to achieve diversity is federated learning. It's when every region in the world has its own datasets which contribute to training a big global model.
How can AI be valued in terms of disinformation and toxic content?
With transformers and self-supervised learning appeared, the proportion of hate speech taken down by AI systems has reached 96%.
However, high false positives (good content removed) are a problem. Therefore, detection thresholds will be adjusted to enable discussions on key societal topics.
Regulations:
"Making the distribution of open source AI engine essentially illegal is way more dangerous than all the other potential dangers."
World Economic Forum
Debating Technology
With AI, space exploration and biotechnology advancing rapidly, some see these innovations as solutions to humanity’s greatest challenges, while others raise concerns about ethics, society and inequality.
In this town hall, leaders debate how to responsibly…
In this town hall, leaders debate how to responsibly…
❤5
OpenAI announced ChatGPT Gov, a version of ChatGPT that government agencies can deploy in their own MS Azure commercial or government cloud environment.
OpenAI
Introducing ChatGPT Gov
ChatGPT Gov is designed to streamline government agencies’ access to OpenAI’s frontier models.
💅4
Researchers has released OpenThoughts, a large-scale open-source dataset for training AI reasoning models.
Along with the dataset, they've introduced OpenThinker-7B, a new model showing promising results in mathematical and code reasoning tasks.
The project is a collaboration between researchers and engineers from Bespoke Labs, Stanford, UC Berkeley, University of Washington, Juelich Supercomputing Center, LAION, UCLA, UNC Chapel Hill, and Toyota Research Institute.
Their first release includes the OpenThoughts-114k dataset, specifically designed to improve AI reasoning capabilities.
Initial benchmarks show impressive performance, with OpenThinker-7B achieving scores of 43.3 on AIME24 and 83.0 on MATH500, approaching the performance of DeepSeek's distilled models.
The results were evaluated using their open-source tool Evalchemy.
The initiative is supported by major organizations including NSF IFML, UT Austin Machine Learning Lab, Juelich Supercomputing Center, Toyota Research Institute, and Lambda Labs.
Code
OpenThoughts-114k Dataset
OpenThinker-7B model
Along with the dataset, they've introduced OpenThinker-7B, a new model showing promising results in mathematical and code reasoning tasks.
The project is a collaboration between researchers and engineers from Bespoke Labs, Stanford, UC Berkeley, University of Washington, Juelich Supercomputing Center, LAION, UCLA, UNC Chapel Hill, and Toyota Research Institute.
Their first release includes the OpenThoughts-114k dataset, specifically designed to improve AI reasoning capabilities.
Initial benchmarks show impressive performance, with OpenThinker-7B achieving scores of 43.3 on AIME24 and 83.0 on MATH500, approaching the performance of DeepSeek's distilled models.
The results were evaluated using their open-source tool Evalchemy.
The initiative is supported by major organizations including NSF IFML, UT Austin Machine Learning Lab, Juelich Supercomputing Center, Toyota Research Institute, and Lambda Labs.
Code
OpenThoughts-114k Dataset
OpenThinker-7B model
GitHub
GitHub - open-thoughts/open-thoughts: Fully open data curation for reasoning models
Fully open data curation for reasoning models. Contribute to open-thoughts/open-thoughts development by creating an account on GitHub.
🔥6❤5🆒5
Hugging Face wants to reverse engineer DeepSeek’s R1 reasoning model
Hugging Face researchers say the Open-R1 project aims to create a fully open-source duplicate of the R1 model and make all of its components available to the AI community.
Elie Bakouch, one of the Hugging Face engineers leading the project, told TechCrunch that though DeepSeek claims R1 is open-source because it can be used without any restrictions, the truth is that it doesn’t meet the standard definition of open software. That’s because many of the components used to build it, and also the data it was trained on, have not been made publicly available.
The lack of information about what goes into DeepSeek means that it’s really just another “black box,” similar to proprietary models such as OpenAI’s GPT series, making it impossible for the AI community to build on or improve, he said.
Hugging Face says it’s attempting to replicate R1 to benefit the AI research community, and it intends to do so in just a few weeks.
To do this, it will leverage the company’s dedicated research server, the “Science Cluster,” which is powered by 768 Nvidia H100 GPUs. The plan is to try to reverse engineer the R1 model to try and understand what data was used to train it, and which components were used in its creation.
The Open-R1 project is seeking assistance from the broader AI research community to try and recreate the training datasets used by DeepSeek, and it has garnered a lot of interest so far, with its associated GitHub page getting more than 100,000 stars just three days after its launch.
Hugging Face researchers say the Open-R1 project aims to create a fully open-source duplicate of the R1 model and make all of its components available to the AI community.
Elie Bakouch, one of the Hugging Face engineers leading the project, told TechCrunch that though DeepSeek claims R1 is open-source because it can be used without any restrictions, the truth is that it doesn’t meet the standard definition of open software. That’s because many of the components used to build it, and also the data it was trained on, have not been made publicly available.
The lack of information about what goes into DeepSeek means that it’s really just another “black box,” similar to proprietary models such as OpenAI’s GPT series, making it impossible for the AI community to build on or improve, he said.
Hugging Face says it’s attempting to replicate R1 to benefit the AI research community, and it intends to do so in just a few weeks.
To do this, it will leverage the company’s dedicated research server, the “Science Cluster,” which is powered by 768 Nvidia H100 GPUs. The plan is to try to reverse engineer the R1 model to try and understand what data was used to train it, and which components were used in its creation.
The Open-R1 project is seeking assistance from the broader AI research community to try and recreate the training datasets used by DeepSeek, and it has garnered a lot of interest so far, with its associated GitHub page getting more than 100,000 stars just three days after its launch.
GitHub
GitHub - huggingface/open-r1: Fully open reproduction of DeepSeek-R1
Fully open reproduction of DeepSeek-R1. Contribute to huggingface/open-r1 development by creating an account on GitHub.
🔥5👏4
OpenAI says it has evidence DeepSeek used its model to train competitor
OpenAI suspects that DeepSeek used the "distillation" technique.
While distillation is a common practice in the industry, the issue is that DeepSeek may have been doing this to create their own competing model, which violates OpenAI's terms of service.
OpenAI suspects that DeepSeek used the "distillation" technique.
While distillation is a common practice in the industry, the issue is that DeepSeek may have been doing this to create their own competing model, which violates OpenAI's terms of service.
🐳4
Why #DeepSeek's Success Doesn't Change the AI Race: Dario Amodei's View
Anthropic's CEO explains why the apparent breakthrough fits into the expected trajectory of AI development.*
3 Laws of AI Development
1. Scaling Law:
- More resources = better results
- $1M = 20% tasks, $10M = 40%, $100M = 60%
- Progress is smooth and predictable
2. Curve Shifting:
- Innovations improve efficiency
- Typical improvements:
* Small (1.2x)
* Medium (2x)
* Large (10x)
- Overall pace: ~4x per year
3. Paradigm Shifts:
- 2020-2023: text training
- 2024: adding Reinforcement Learning
- Now: unique "crossover point"
What DeepSeek actually achieved?
- Performance similar to 7-10 month old US models
- Lower costs, but within normal trend
- Significant resources (~50,000 chips, ~$1B)
Not a revolution because:
- Cost reduction follows expected 4x/year trend
- V3 more innovative than R1
- Total company spending comparable to US labs
The Future (2026-2027)
According to Amodei, truly advanced AI will require:
- Millions of chips
- Tens of billions of dollars
- 2-3 years of work
Key Takeaway
DeepSeek demonstrates an expected point on the progress curve, not a revolutionary breakthrough. The real race for superhuman AI is just beginning, and it will require unprecedented resources.
"Making AI that is smarter than almost all humans at almost all things will require millions of chips, tens of billions of dollars (at least), and is most likely to happen in 2026-2027"- Dario Amodei
Anthropic's CEO explains why the apparent breakthrough fits into the expected trajectory of AI development.*
3 Laws of AI Development
1. Scaling Law:
- More resources = better results
- $1M = 20% tasks, $10M = 40%, $100M = 60%
- Progress is smooth and predictable
2. Curve Shifting:
- Innovations improve efficiency
- Typical improvements:
* Small (1.2x)
* Medium (2x)
* Large (10x)
- Overall pace: ~4x per year
3. Paradigm Shifts:
- 2020-2023: text training
- 2024: adding Reinforcement Learning
- Now: unique "crossover point"
What DeepSeek actually achieved?
- Performance similar to 7-10 month old US models
- Lower costs, but within normal trend
- Significant resources (~50,000 chips, ~$1B)
Not a revolution because:
- Cost reduction follows expected 4x/year trend
- V3 more innovative than R1
- Total company spending comparable to US labs
The Future (2026-2027)
According to Amodei, truly advanced AI will require:
- Millions of chips
- Tens of billions of dollars
- 2-3 years of work
Key Takeaway
DeepSeek demonstrates an expected point on the progress curve, not a revolutionary breakthrough. The real race for superhuman AI is just beginning, and it will require unprecedented resources.
"Making AI that is smarter than almost all humans at almost all things will require millions of chips, tens of billions of dollars (at least), and is most likely to happen in 2026-2027"- Dario Amodei
Darioamodei
Dario Amodei — On DeepSeek and Export Controls
❤4✍2
Voice Agents Market Map - B2B by Andreessen Horowitz - an update on AI voice agents
2024 was a huge year for voice agents, with massive model breakthroughs + hundreds of new startups.
How has the space evolved, and where do we see opportunity in 2025?
To start: voice as one of AI's biggest unlocks. Latency and reliability are now largely solved - and interruptibility and emotionality have made major strides, too. Voice AI is now nearly at human standards, allowing tech to replace labor on the phone.
As a result, there's been an explosion in startups building applications on these models. Y Combinator alone has seen 90 voice agent cos. Many are targeting specific verticals - by industry (e.g. home services, dental) or function (e.g. recruiting, customer support) - and are scaling fast!
Most companies need to tap into adjacent workflows: pushing call details to a CRM, automating follow-ups, etc.
What a16Z're looking for in voice agent startups:
1. Building in an industry where phone is the preferred or required medium, or has a much higher success rate than other modalities
2. Calls are constrained — both in length and in format/outcome
3. Voice agent delivers 50%+ cost reduction with similar success rate as a human
4. Calls are "life or death" for the customer —they will pay significant $ to get them made or answered...but not for the end consumer
5. When selling into SMB/mid-market, the agent product has easy integration. When selling into enterprise, complex integration can be a moat!
2024 was a huge year for voice agents, with massive model breakthroughs + hundreds of new startups.
How has the space evolved, and where do we see opportunity in 2025?
To start: voice as one of AI's biggest unlocks. Latency and reliability are now largely solved - and interruptibility and emotionality have made major strides, too. Voice AI is now nearly at human standards, allowing tech to replace labor on the phone.
As a result, there's been an explosion in startups building applications on these models. Y Combinator alone has seen 90 voice agent cos. Many are targeting specific verticals - by industry (e.g. home services, dental) or function (e.g. recruiting, customer support) - and are scaling fast!
Most companies need to tap into adjacent workflows: pushing call details to a CRM, automating follow-ups, etc.
What a16Z're looking for in voice agent startups:
1. Building in an industry where phone is the preferred or required medium, or has a much higher success rate than other modalities
2. Calls are constrained — both in length and in format/outcome
3. Voice agent delivers 50%+ cost reduction with similar success rate as a human
4. Calls are "life or death" for the customer —they will pay significant $ to get them made or answered...but not for the end consumer
5. When selling into SMB/mid-market, the agent product has easy integration. When selling into enterprise, complex integration can be a moat!
🔥3❤2
Trump Media and Technology Group announced the launch of the financial services and FinTech brand TruthFi.
$250 million will be invested in SMAs, ETFs, Bitcoin and similar cryptocurrencies or crypto-related securities. TrueFi will enter the field of decentralized finance.
$250 million will be invested in SMAs, ETFs, Bitcoin and similar cryptocurrencies or crypto-related securities. TrueFi will enter the field of decentralized finance.
GlobeNewswire News Room
Trump Media Announces Expansion into Financial Services
TMTG Launches Truth.Fi Brand, Plans to Build Investment Vehicles Based on America-First Principles SARASOTA, Fla., Jan. 29, 2025 (GLOBE NEWSWIRE) --...
👍3👀1
Another embodied AI evaluation suite that might be worth a look
huggingface.co
Paper page - EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
Join the discussion on this paper page
❤3
Mistral AI released a new small model
Mistral Small 3, a latency-optimized 24B-parameter model released under the Apache 2.0 license.
Mistral Small 3 is competitive with larger models such as Llama 3.3 70B or Qwen 32B, and is an excellent open replacement for opaque proprietary models like GPT4o-mini.
Mistral Small 3 is on par with Llama 3.3 70B instruct, while being more than 3x faster on the same hardware.
Mistral Small 3 is a pre-trained and instructed model catered to the ‘80%’ of generative AI tasks—those that require robust language and instruction following performance, with very low latency.
Mistral Small 3, a latency-optimized 24B-parameter model released under the Apache 2.0 license.
Mistral Small 3 is competitive with larger models such as Llama 3.3 70B or Qwen 32B, and is an excellent open replacement for opaque proprietary models like GPT4o-mini.
Mistral Small 3 is on par with Llama 3.3 70B instruct, while being more than 3x faster on the same hardware.
Mistral Small 3 is a pre-trained and instructed model catered to the ‘80%’ of generative AI tasks—those that require robust language and instruction following performance, with very low latency.
Mistral
Mistral Small 3 | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
🆒5❤2
Now available on the Allen Brain Cell Atlas - aging mouse brain data
Contains 1.2 million high-quality, single-cell transcriptomes.
Contains 1.2 million high-quality, single-cell transcriptomes.
OpenAI's chief product officer Kevin Weil: "we have a new model coming soon that is head and shoulders above anything that's out there today... the more compute you apply, the more intelligent the model is"
YouTube
OpenAI CPO Kevin Weil on DeepSeek: The 'first big salvo' in U.S.-China AI arms race
CNBC’s Kate Rooney and OpenAI chief product officer Kevin Weil join 'Squawk on the Street' to discuss the company's new initiatives and government partnerships, his thoughts on China's DeepSeek AI development, state of the AI technology race, and more. For…
🔥5🆒2
Google rolled out Gemini 2 Flash for free to its products.
— one of the most high quality non-reasoning LLMs
— super fast (150tok/s+)
— 1M tok context window.
API price isnt out but was previously $0.075/$0.30 per M input/output tokens. Big move from Google.
— one of the most high quality non-reasoning LLMs
— super fast (150tok/s+)
— 1M tok context window.
API price isnt out but was previously $0.075/$0.30 per M input/output tokens. Big move from Google.
AI Distillation Race: From $450 Berkeley Experiment to Industry Disruption
In a fascinating turn of events in AI development, UC Berkeley doctoral students demonstrated that advanced AI capabilities can be replicated for just $450 in computing costs.
This comes amid industry buzz about #DeepSeek's R1 model, which allegedly used similar distillation techniques to replicate OpenAI's reasoning capabilities.
The Berkeley breakthrough:
- Used Alibaba's Qwen model to generate 17,000 training examples
- Focused on math and coding problems with verifiable answers
- Their model outperformed OpenAI's first reasoning model on several benchmarks
The bigger picture:
- OpenAI claims Chinese quant fund DeepSeek used distillation to replicate their o1 reasoning model
- While OpenAI tries to hide their models' thought processes, DeepSeek took an open approach with R1
- This transparency means other developers could potentially replicate R1's capabilities
The Berkeley case and DeepSeek's R1 show that AI innovation can be democratized. While tech giants invest billions in AI development, smaller teams can achieve impressive results through clever engineering and efficient training methods.
In a fascinating turn of events in AI development, UC Berkeley doctoral students demonstrated that advanced AI capabilities can be replicated for just $450 in computing costs.
This comes amid industry buzz about #DeepSeek's R1 model, which allegedly used similar distillation techniques to replicate OpenAI's reasoning capabilities.
The Berkeley breakthrough:
- Used Alibaba's Qwen model to generate 17,000 training examples
- Focused on math and coding problems with verifiable answers
- Their model outperformed OpenAI's first reasoning model on several benchmarks
The bigger picture:
- OpenAI claims Chinese quant fund DeepSeek used distillation to replicate their o1 reasoning model
- While OpenAI tries to hide their models' thought processes, DeepSeek took an open approach with R1
- This transparency means other developers could potentially replicate R1's capabilities
The Berkeley case and DeepSeek's R1 show that AI innovation can be democratized. While tech giants invest billions in AI development, smaller teams can achieve impressive results through clever engineering and efficient training methods.
novasky-ai.github.io
Sky-T1: Train your own O1 preview model within $450
We introduce Sky-T1-32B-Preview, our reasoning model that performs on par with o1-preview on popular reasoning and coding benchmarks.
❤1
OpenAI just released o3-mini, available today to all users in ChatGPT (for free)!
OpenAI
OpenAI o3-mini
OpenAI announced Deep Research, a new ChatGPT model that can conduct autonomous research by searching the web and synthesizing knowledge into a research paper as output, in what is the next step towards AI discovering new knowledge for itself
OpenAI's Deep Research scores a new state-of-the-art score on Humanity's Last Exam, surpassing the accuracy of o3-mini.
Sam Altman says new Deep Research AI agent powered by their o3 model can do "a single-digit percentage of all economically valuable tasks in the world" and this step into synthesizing knowledge will soon be followed by agents that can invent new knowledge.
OpenAI's Deep Research scores a new state-of-the-art score on Humanity's Last Exam, surpassing the accuracy of o3-mini.
Sam Altman says new Deep Research AI agent powered by their o3 model can do "a single-digit percentage of all economically valuable tasks in the world" and this step into synthesizing knowledge will soon be followed by agents that can invent new knowledge.
😁4🥴2
TSMC plans to build a massive 1nm fab in southern Taiwan a type calls a Giga Fab, able to fit as many production lines as 6 normal 12-inch wafer fabs
The investment would be part of the government’s "Greater Southern New Silicon Valley Promotion Plan," aimed at focusing new semiconductor investments in southern Taiwan.
TSMC reportedly said there are many possible locations for the fab.
The investment would be part of the government’s "Greater Southern New Silicon Valley Promotion Plan," aimed at focusing new semiconductor investments in southern Taiwan.
TSMC reportedly said there are many possible locations for the fab.
經濟日報
台積電1奈米傳落腳台南沙崙 業界:根留台灣的決心 | 產業熱點 | 產業 | 經濟日報
台積電最先進的1奈米製程新廠傳將落腳台南沙崙,規劃打造可容納六座12吋廠的超大型晶圓廠(Giga-Fab),藉此放大現有...