Shanghai has established China's 1st humanoid robot training ground
This center will be able to satisfy requirement of training 1000 humanoid robot by 2027.
Goal is to achieve reconfiguration scenarios & heterogeneous robots covering > 10 scenarios in intelligent mfg, service industry & special applications.
Building a heterogenous cluster acquisition, training & promotion open src framework. Embody smart operation & task scheduling.
- 100 robots will be deployed to help w/ tech breakthrus & applications.
- Leju, Fourier, Kepler & various intelligent computing institutions are involved.
Aims to solve the problem of data collection. Challenges of low data collection efficiency, high cost & low reusability across platform & lack of data standards.
Goal of inspiring more leading enterprises to build humanoid robot innovation centers around China.
This center will be able to satisfy requirement of training 1000 humanoid robot by 2027.
Goal is to achieve reconfiguration scenarios & heterogeneous robots covering > 10 scenarios in intelligent mfg, service industry & special applications.
Building a heterogenous cluster acquisition, training & promotion open src framework. Embody smart operation & task scheduling.
- 100 robots will be deployed to help w/ tech breakthrus & applications.
- Leju, Fourier, Kepler & various intelligent computing institutions are involved.
Aims to solve the problem of data collection. Challenges of low data collection efficiency, high cost & low reusability across platform & lack of data standards.
Goal of inspiring more leading enterprises to build humanoid robot innovation centers around China.
Ithome
全国首个异构人形机器人训练场正式启用,能容纳超 100 个人形机器人同时训练 - IT之家
该训练场可容纳 100 个人形机器人同时进行智能训练的场地,到 2027 年可以满足 1000 个人形机器人同时训练。
❤5
Building Towards Computer Use with Anthropic
This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes.
Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access.
In detail, you’ll:
- Learn about Anthropic's family of models, when to use which one, and make API requests to Claude
- Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses
- Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples
- Implement prompt caching to reduce cost and latency
- Apply tool-use to build a chatbot that can call different tools to respond to queries
- See all these building blocks come together in Computer Use demo
This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes.
Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access.
In detail, you’ll:
- Learn about Anthropic's family of models, when to use which one, and make API requests to Claude
- Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses
- Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples
- Implement prompt caching to reduce cost and latency
- Apply tool-use to build a chatbot that can call different tools to respond to queries
- See all these building blocks come together in Computer Use demo
www.deeplearning.ai
Building toward Computer Use with Anthropic
Learn how an AI Assistant is built to use and accomplish tasks on computers.
ByteDance unveils Doubao-1.5-pro that seems to be world class, comparable or better to GPT-4o, latest Gemini, DS & Claude.
Its MoE architecture explores balance bw model & reasoning.
It build highly autonomous data production system & not using data from any other models.
Its MoE architecture explores balance bw model & reasoning.
It build highly autonomous data production system & not using data from any other models.
Quantum computer helps discover new drug candidates for cancer
A team of scientists from the University of Toronto, Harvard, Stanford, and other institutions developed a hybrid quantum-classical algorithm to discover inhibitors of KRAS protein - a crucial target in cancer therapy.
Using IBM's 16-qubit Guadalupe quantum computer combined with classical machine learning algorithms, the researchers were able to generate and select promising drug candidates.
Key Results:
- 15 compounds were synthesized and tested
- Two compounds showed particularly promising results:
* ISM061-018-2 demonstrated high binding affinity to KRAS-G12D (1.4 μM)
* ISM061-022 showed selective activity against specific KRAS mutations
Why This Matters:
1. First-ever instance of quantum computing leading to experimentally validated drug candidates
2. Achieved using current quantum hardware with relatively few qubits
3. Could potentially reduce preclinical drug development time from years to months
The researchers' approach combined:
- Quantum Circuit Born Machine (QCBM) on IBM's quantum hardware
- Classical LSTM neural network
- Chemistry42 platform for validation
- Experimental validation through surface plasmon resonance and cell-based assays
While the team doesn't claim "quantum advantage" (where quantum computers fundamentally outperform classical ones), their work demonstrates that even today's quantum computers can contribute meaningfully to solving complex drug discovery challenges.
A team of scientists from the University of Toronto, Harvard, Stanford, and other institutions developed a hybrid quantum-classical algorithm to discover inhibitors of KRAS protein - a crucial target in cancer therapy.
Using IBM's 16-qubit Guadalupe quantum computer combined with classical machine learning algorithms, the researchers were able to generate and select promising drug candidates.
Key Results:
- 15 compounds were synthesized and tested
- Two compounds showed particularly promising results:
* ISM061-018-2 demonstrated high binding affinity to KRAS-G12D (1.4 μM)
* ISM061-022 showed selective activity against specific KRAS mutations
Why This Matters:
1. First-ever instance of quantum computing leading to experimentally validated drug candidates
2. Achieved using current quantum hardware with relatively few qubits
3. Could potentially reduce preclinical drug development time from years to months
The researchers' approach combined:
- Quantum Circuit Born Machine (QCBM) on IBM's quantum hardware
- Classical LSTM neural network
- Chemistry42 platform for validation
- Experimental validation through surface plasmon resonance and cell-based assays
While the team doesn't claim "quantum advantage" (where quantum computers fundamentally outperform classical ones), their work demonstrates that even today's quantum computers can contribute meaningfully to solving complex drug discovery challenges.
🆒5😁2
New Google DeepMind safety paper. MONA: a method for addressing multi-step reward hacking
Key idea: Use RL training only for short horizons (myopic optimization), but have an overseer evaluate how good actions are for the long term (non-myopic approval).
The technical safety post.
Key idea: Use RL training only for short horizons (myopic optimization), but have an overseer evaluate how good actions are for the long term (non-myopic approval).
The technical safety post.
arXiv.org
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate...
Future advanced AI systems may learn sophisticated strategies through reinforcement learning (RL) that humans cannot understand well enough to safely evaluate. We propose a training method which...
🔥5
All about AI, Web 3.0, BCI
OpenAI is prepping to release "Operator," a new ChatGPT feature that will take actions on behalf of users in their browsers, this week. Interesting details: - Operator provides suggested prompts - Users can save/share tasks - Not available in API
OpenAI released a computer-using agent Operator as a research preview.
Ensuring safety for agentic models is far more complex than for chatbots.
Errors can lead to serious consequences—for instance, the agent might make costly real-world decisions, like accidentally spending money from a credit card.
Achieving full agentic safety will be as challenging as ensuring the safety of self-driving cars, but with an added layer of complexity.
Ensuring safety for agentic models is far more complex than for chatbots.
Errors can lead to serious consequences—for instance, the agent might make costly real-world decisions, like accidentally spending money from a credit card.
Achieving full agentic safety will be as challenging as ensuring the safety of self-driving cars, but with an added layer of complexity.
OpenAI
Introducing Operator
Perplexity Launches Advanced AI Assistant for Android. It's completely FREE for all Android users
Perplexity has just made a major leap forward, transforming from an answer engine into a fully integrated Android assistant.
This new release brings a powerful set of features that make daily tasks seamless and intuitive.
What's New:
• Voice-controlled actions and gesture support
• Seamless booking services (Uber, OpenTable)
• YouTube video playback
• Music streaming
• Navigation assistance
• Real-time translation capabilities
The game-changing feature? Context awareness! Start a conversation about a basketball game, and easily set up alerts for it. Research restaurants and book a table directly through the app – all in one continuous interaction.
Plus, it supports multimodal interactions, combining both camera and voice inputs for enhanced functionality.
Simply switch your default assistant from Google/Gemini to Perplexity in your settings.
While Google maintains its default search position due to Play Store requirements, Perplexity is leading the way in revolutionizing how we interact with our phones.
Note: Some features like OpenTable integration are still being refined and will see improvements in the coming months.
Perplexity has just made a major leap forward, transforming from an answer engine into a fully integrated Android assistant.
This new release brings a powerful set of features that make daily tasks seamless and intuitive.
What's New:
• Voice-controlled actions and gesture support
• Seamless booking services (Uber, OpenTable)
• YouTube video playback
• Music streaming
• Navigation assistance
• Real-time translation capabilities
The game-changing feature? Context awareness! Start a conversation about a basketball game, and easily set up alerts for it. Research restaurants and book a table directly through the app – all in one continuous interaction.
Plus, it supports multimodal interactions, combining both camera and voice inputs for enhanced functionality.
Simply switch your default assistant from Google/Gemini to Perplexity in your settings.
While Google maintains its default search position due to Play Store requirements, Perplexity is leading the way in revolutionizing how we interact with our phones.
Note: Some features like OpenTable integration are still being refined and will see improvements in the coming months.
Mistral AI is “not for sale” and instead is working toward an initial public offering, CEO Arthur Mensch told Bloomberg
“Startups are always raising money, but we have plenty,” Mensch said when asked about fundraising. Scaling to compete with larger rivals could require new funds, he added.
“Startups are always raising money, but we have plenty,” Mensch said when asked about fundraising. Scaling to compete with larger rivals could require new funds, he added.
Bloomberg.com
French AI Champion Mistral Isn’t for Sale, CEO Mensch Says
European artificial intelligence champion Mistral AI is “not for sale” and instead is working toward an initial public offering, according to Chief Executive Officer Arthur Mensch.
TypeDB isn't a relational database, or document, or graph, or vector. Built on a different mathematical model, it represents an evolution of them.
Ultimately, TypeDB is an entity-relation database, underpinned in type theory.
Ultimately, TypeDB is an entity-relation database, underpinned in type theory.
Typedb
What is an entity-relation database?
Breaking down our novel database. Learn about what it means to be an entity-relation database.
Blackrock Neurotech BCI Achieves Four-Dimensional Finger Control
Scientists have developed a high-performance BCI system that allows a paralyzed person to control individual finger movements with unprecedented precision.
Key Achievements:
1. The participant gained control of three independent finger groups:
- Thumb movement in 2 dimensions
- Index-middle finger group
- Ring-little finger group
- Total of 4 degrees of freedom
2. Outstanding Performance:
- 76 targets per minute
- Average completion time: 1.58 seconds
- Exceeds previous achievements in animal studies
Technical Details
- Two 96-channel Blackrock microelectrode arrays
- Placed in the hand 'knob' area of the left precentral gyrus
- Neural signals processed in real-time for continuous control
Practical Application
The system was successfully used to control a virtual quadcopter through obstacle courses, demonstrating real-world potential for complex control tasks.
Human Impact
The participant reported strong feelings of enablement and social connectedness, highlighting the technology's potential to meet unmet needs of people with paralysis for peer support, leisure activities, and sports.
Scientists have developed a high-performance BCI system that allows a paralyzed person to control individual finger movements with unprecedented precision.
Key Achievements:
1. The participant gained control of three independent finger groups:
- Thumb movement in 2 dimensions
- Index-middle finger group
- Ring-little finger group
- Total of 4 degrees of freedom
2. Outstanding Performance:
- 76 targets per minute
- Average completion time: 1.58 seconds
- Exceeds previous achievements in animal studies
Technical Details
- Two 96-channel Blackrock microelectrode arrays
- Placed in the hand 'knob' area of the left precentral gyrus
- Neural signals processed in real-time for continuous control
Practical Application
The system was successfully used to control a virtual quadcopter through obstacle courses, demonstrating real-world potential for complex control tasks.
Human Impact
The participant reported strong feelings of enablement and social connectedness, highlighting the technology's potential to meet unmet needs of people with paralysis for peer support, leisure activities, and sports.
❤2
DeepSeek showed us in just 4 days:
1. Open-source AI is only <6 months behind closed AI
2. China is leading the open-source AI race
3. we are entering the LLM RL golden era
4. distilled models are powerful, we'll have highly intelligent AI running locally on our phones.
Reactions:
❗️OpenAI o3-mini available for free-tier
👽hopefully we will see less AGI/ASI vagueposting.
It’s hard to predict who will ultimately win, but don’t forget the power of the last mover advantage: Google invented the Transformer, but OpenAI unlocked its true potential.
1. Open-source AI is only <6 months behind closed AI
2. China is leading the open-source AI race
3. we are entering the LLM RL golden era
4. distilled models are powerful, we'll have highly intelligent AI running locally on our phones.
Reactions:
❗️OpenAI o3-mini available for free-tier
👽hopefully we will see less AGI/ASI vagueposting.
It’s hard to predict who will ultimately win, but don’t forget the power of the last mover advantage: Google invented the Transformer, but OpenAI unlocked its true potential.
🔥4👍2
Dragon Intermittency: A New Type of Chaos in Neural Networks Discovered
Scientists have unveiled a new type of chaotic behavior in neural systems, potentially revolutionizing our understanding of brain dynamics and neural computing.
The Discovery
Imagine two neurons trying to synchronize their activity. Sometimes they work in perfect harmony, and sometimes they fall into chaos. Scientists discovered that these transitions follow a unique pattern they named "dragon intermittency" due to its distinctive probability distribution shape.
Why It Matters:
1. For Brain Science:
- New insights into neural synchronization mechanisms
- Potential explanations for certain neurodegenerative disorders
- Advanced approaches to neural network modeling
2. For Technology:
- Improvements in artificial neural network architectures
- Novel methods for neuromorphic computing security
- Energy optimization in brain-like systems
3. For Business:
- New opportunities in neuromorphic processor development
- Applications in predictive analytics
- Promising R&D directions in neurotechnology
Technical Highlights:
The phenomenon is characterized by two key scaling laws:
- Mean duration of synchronous phases follows a power law with exponent -0.5
- Probability distribution of synchronization durations follows a power law with exponent -1.7
Practical Applications:
1. Medicine:
- Enhanced brain activity monitoring methods
- More precise brain-computer interfaces
- Personalized approaches to epilepsy treatment
2. Computing:
- Optimization of neuromorphic processors
- Enhancement of machine learning algorithms
- Novel approaches to quantum computing
Investment Potential:
This discovery opens new opportunities for:
- Neurotechnology startups
- Specialized software developers
- Medical equipment manufacturers
- Companies working on neuromorphic computing
Future Implications:
The research lays groundwork for:
- More efficient neuromorphic computers
- Novel treatments for neurological disorders
- Advanced AI technologies
- Better understanding of brain dynamics
This discovery positions at the forefront of neural dynamics research, with potential applications spanning:
- Healthcare technology
- Artificial Intelligence
- Neuromorphic computing
- Brain-computer interfaces
Scientists have unveiled a new type of chaotic behavior in neural systems, potentially revolutionizing our understanding of brain dynamics and neural computing.
The Discovery
Imagine two neurons trying to synchronize their activity. Sometimes they work in perfect harmony, and sometimes they fall into chaos. Scientists discovered that these transitions follow a unique pattern they named "dragon intermittency" due to its distinctive probability distribution shape.
Why It Matters:
1. For Brain Science:
- New insights into neural synchronization mechanisms
- Potential explanations for certain neurodegenerative disorders
- Advanced approaches to neural network modeling
2. For Technology:
- Improvements in artificial neural network architectures
- Novel methods for neuromorphic computing security
- Energy optimization in brain-like systems
3. For Business:
- New opportunities in neuromorphic processor development
- Applications in predictive analytics
- Promising R&D directions in neurotechnology
Technical Highlights:
The phenomenon is characterized by two key scaling laws:
- Mean duration of synchronous phases follows a power law with exponent -0.5
- Probability distribution of synchronization durations follows a power law with exponent -1.7
Practical Applications:
1. Medicine:
- Enhanced brain activity monitoring methods
- More precise brain-computer interfaces
- Personalized approaches to epilepsy treatment
2. Computing:
- Optimization of neuromorphic processors
- Enhancement of machine learning algorithms
- Novel approaches to quantum computing
Investment Potential:
This discovery opens new opportunities for:
- Neurotechnology startups
- Specialized software developers
- Medical equipment manufacturers
- Companies working on neuromorphic computing
Future Implications:
The research lays groundwork for:
- More efficient neuromorphic computers
- Novel treatments for neurological disorders
- Advanced AI technologies
- Better understanding of brain dynamics
This discovery positions at the forefront of neural dynamics research, with potential applications spanning:
- Healthcare technology
- Artificial Intelligence
- Neuromorphic computing
- Brain-computer interfaces
❤3🔥3👎1
AI Transformation: From Hardware to Agents - Analysis of Time Horizons
DeepSeek as a Catalyst (2024)
DeepSeek has demonstrated a radically new approach to AI development:
- Reducing model training costs by 20x (from $100M to $5M)
- Decreasing required GPUs from 100,000 to 2,000
- Enabling deployment on regular gaming graphics cards
- Innovative "expert system" where only necessary parameters are activated
Transition Period: Jevons Paradox (2024-2026)
As noted by Satya Nadella, more efficient technologies will lead to explosive growth in usage:
- Democratization of AI access will increase the number of developer companies
- New use cases will emerge
- Each company will deploy multiple models
- In the short term, this will increase overall computational resource consumption
Long-term Transformation: The Era of Agents (2025-2035)
In the long term, a fundamental change in industry architecture will occur:
- From hardware and foundation models to practical applications
- From centralized systems to decentralized AI agents
- From infrastructure model to service model
Key Indicators of Change:
1. Technological:
- Emergence of efficient small models
- Development of systems with dynamic parameter activation
- Improvement in model distillation techniques
2. Economic:
- Lowering industry entry barriers
- Value shift from infrastructure to application
- Development of new AI agent-based business models
3. Structural:
- From large company monopoly to distributed ecosystem
- From vertical integration to horizontal networks
- From closed systems to open standards
We are witnessing not just technological optimization, but a comprehensive industry transformation. DeepSeek shows the technical path, Jevons paradox describes the transition period, and the AI agents concept defines the end point of this transformation.
This change is comparable to the transition from mainframes to personal computers or from local servers to cloud computing - technological breakthrough leads to temporary increase in resource consumption, but ultimately creates a fundamentally new industry architecture.
The significance of this transformation extends beyond mere efficiency gains. It represents a paradigm shift in how AI will be developed, deployed, and utilized. While the immediate effect may be increased resource consumption (as predicted by Jevons paradox), the long-term impact will be a more democratized, efficient, and accessible AI ecosystem.
Future Implications
1. Market Structure:
- Shift from hardware dominance to software and service innovation
- Emergence of specialized AI agent marketplaces
- New opportunities for smaller players and startups
2. Development Patterns:
- Focus on agent orchestration rather than model size
- Emphasis on efficiency and specialization
- Growing importance of interoperability standards
3. Resource Utilization:
- Initial surge in computational demand
- Gradual optimization through improved efficiency
- Evolution toward distributed processing models
This transformation marks a pivotal moment in AI development, where the focus shifts from raw computational power to intelligent resource utilization and practical applications. The next decade will likely be defined by how successfully we navigate this transition from centralized, hardware-dependent AI to a more distributed, agent-based ecosystem.
DeepSeek as a Catalyst (2024)
DeepSeek has demonstrated a radically new approach to AI development:
- Reducing model training costs by 20x (from $100M to $5M)
- Decreasing required GPUs from 100,000 to 2,000
- Enabling deployment on regular gaming graphics cards
- Innovative "expert system" where only necessary parameters are activated
Transition Period: Jevons Paradox (2024-2026)
As noted by Satya Nadella, more efficient technologies will lead to explosive growth in usage:
- Democratization of AI access will increase the number of developer companies
- New use cases will emerge
- Each company will deploy multiple models
- In the short term, this will increase overall computational resource consumption
Long-term Transformation: The Era of Agents (2025-2035)
In the long term, a fundamental change in industry architecture will occur:
- From hardware and foundation models to practical applications
- From centralized systems to decentralized AI agents
- From infrastructure model to service model
Key Indicators of Change:
1. Technological:
- Emergence of efficient small models
- Development of systems with dynamic parameter activation
- Improvement in model distillation techniques
2. Economic:
- Lowering industry entry barriers
- Value shift from infrastructure to application
- Development of new AI agent-based business models
3. Structural:
- From large company monopoly to distributed ecosystem
- From vertical integration to horizontal networks
- From closed systems to open standards
We are witnessing not just technological optimization, but a comprehensive industry transformation. DeepSeek shows the technical path, Jevons paradox describes the transition period, and the AI agents concept defines the end point of this transformation.
This change is comparable to the transition from mainframes to personal computers or from local servers to cloud computing - technological breakthrough leads to temporary increase in resource consumption, but ultimately creates a fundamentally new industry architecture.
The significance of this transformation extends beyond mere efficiency gains. It represents a paradigm shift in how AI will be developed, deployed, and utilized. While the immediate effect may be increased resource consumption (as predicted by Jevons paradox), the long-term impact will be a more democratized, efficient, and accessible AI ecosystem.
Future Implications
1. Market Structure:
- Shift from hardware dominance to software and service innovation
- Emergence of specialized AI agent marketplaces
- New opportunities for smaller players and startups
2. Development Patterns:
- Focus on agent orchestration rather than model size
- Emphasis on efficiency and specialization
- Growing importance of interoperability standards
3. Resource Utilization:
- Initial surge in computational demand
- Gradual optimization through improved efficiency
- Evolution toward distributed processing models
This transformation marks a pivotal moment in AI development, where the focus shifts from raw computational power to intelligent resource utilization and practical applications. The next decade will likely be defined by how successfully we navigate this transition from centralized, hardware-dependent AI to a more distributed, agent-based ecosystem.
👍5❤3👏1
DeepSeek just dropped another open-source AI model, Janus-Pro-7B
It's multimodal (can generate images) and beats OpenAI's DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks.
This comes on top of all the R1 hype.
It's multimodal (can generate images) and beats OpenAI's DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks.
This comes on top of all the R1 hype.
🔥15🆒3❤2
Alibaba released the SOTA open multimodal model, Qwen2.5-VL
It shows significant improvements across various aspects compared to the previous version.
Key Highlights:
1. Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all!
2. Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones.
3. Long Video Comprehension : Captures events in videos over 1 hour long
4. Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection.
5. Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more.
ModelScope.
It shows significant improvements across various aspects compared to the previous version.
Key Highlights:
1. Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all!
2. Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones.
3. Long Video Comprehension : Captures events in videos over 1 hour long
4. Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection.
5. Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more.
ModelScope.
chat.qwen.ai
Qwen Studio
Qwen Studio is an official platform from Qwen that empowers both everyday users and developers with unified access to Qwen’s series of open-source and proprietary models. It offers comprehensive functionality spanning chatbots, image and video understanding…
The key insights from Yann Lecun's speech at the #WEF2025
At a very interesting debate on technology he covered:
- The perspectives of AI in the next 3-5 years
- What intelligent AI systems need
- Why LLMs are not enough
- True diversity in AI
- Open source importance
and more
How do we make sure the AI we build is the AI we want?
Yann LeCun said that where we are now with Gen AI isn't where we want to be. Gen AI isn't very controllable, but researchers try to make it applicable to a wide range of areas.
The perspectives of AI:
In the next 3-5 years, we are going to see the emergence of a new revolutionary brand or paradigm for AI architectures.
- This new AI won't be generative as we see generative AI today.
- It will be able to plan the sequence of actions and the sequence of plans.
Then we'll come to agentic AI and then to robotics. And the coming decade may be the era of robotics.
4 things that are essential for intelligent AI systems (and are limitations for current systems):
- Understanding the physical world
- Having persistent memory
- Being capable of reasoning
- Complex planning capabilities
LLMs' place in the future of AI:
Yann Lecun considers LLMs as a part of a bigger future AI system, because they are good at manipulating language but bad at thinking.
We're never going to get to human-level AI just with text models. AI needs to know how the real world works from sensory data, as well.
Why is language simple?
Language, like DNA and proteins, is a discrete object, and it's easy to make predictions in discrete worlds.
Physical AI is much more difficult. That's why techniques used in LLMs can't be used for predicting videos.
Open source matters the most:
Yann Lecun highlighted that open source makes AI tool accessible to everyone.
It's important for cultural diversity and democracy because a wide diversity of AI assistants can only be achieved with open source.
How to achieve true diversity in AI:
True diversity is when we have models that are trained on all the languages, cultures, and values of the world.
That's where foundational models can be the base that is fine-tuned for different ideas and visions of what good value systems are, so people can choose from many options.
Federated learning:
One of the ways to achieve diversity is federated learning. It's when every region in the world has its own datasets which contribute to training a big global model.
How can AI be valued in terms of disinformation and toxic content?
With transformers and self-supervised learning appeared, the proportion of hate speech taken down by AI systems has reached 96%.
However, high false positives (good content removed) are a problem. Therefore, detection thresholds will be adjusted to enable discussions on key societal topics.
Regulations:
"Making the distribution of open source AI engine essentially illegal is way more dangerous than all the other potential dangers."
At a very interesting debate on technology he covered:
- The perspectives of AI in the next 3-5 years
- What intelligent AI systems need
- Why LLMs are not enough
- True diversity in AI
- Open source importance
and more
How do we make sure the AI we build is the AI we want?
Yann LeCun said that where we are now with Gen AI isn't where we want to be. Gen AI isn't very controllable, but researchers try to make it applicable to a wide range of areas.
The perspectives of AI:
In the next 3-5 years, we are going to see the emergence of a new revolutionary brand or paradigm for AI architectures.
- This new AI won't be generative as we see generative AI today.
- It will be able to plan the sequence of actions and the sequence of plans.
Then we'll come to agentic AI and then to robotics. And the coming decade may be the era of robotics.
4 things that are essential for intelligent AI systems (and are limitations for current systems):
- Understanding the physical world
- Having persistent memory
- Being capable of reasoning
- Complex planning capabilities
LLMs' place in the future of AI:
Yann Lecun considers LLMs as a part of a bigger future AI system, because they are good at manipulating language but bad at thinking.
We're never going to get to human-level AI just with text models. AI needs to know how the real world works from sensory data, as well.
Why is language simple?
Language, like DNA and proteins, is a discrete object, and it's easy to make predictions in discrete worlds.
Physical AI is much more difficult. That's why techniques used in LLMs can't be used for predicting videos.
Open source matters the most:
Yann Lecun highlighted that open source makes AI tool accessible to everyone.
It's important for cultural diversity and democracy because a wide diversity of AI assistants can only be achieved with open source.
How to achieve true diversity in AI:
True diversity is when we have models that are trained on all the languages, cultures, and values of the world.
That's where foundational models can be the base that is fine-tuned for different ideas and visions of what good value systems are, so people can choose from many options.
Federated learning:
One of the ways to achieve diversity is federated learning. It's when every region in the world has its own datasets which contribute to training a big global model.
How can AI be valued in terms of disinformation and toxic content?
With transformers and self-supervised learning appeared, the proportion of hate speech taken down by AI systems has reached 96%.
However, high false positives (good content removed) are a problem. Therefore, detection thresholds will be adjusted to enable discussions on key societal topics.
Regulations:
"Making the distribution of open source AI engine essentially illegal is way more dangerous than all the other potential dangers."
World Economic Forum
Debating Technology
With AI, space exploration and biotechnology advancing rapidly, some see these innovations as solutions to humanity’s greatest challenges, while others raise concerns about ethics, society and inequality.
In this town hall, leaders debate how to responsibly…
In this town hall, leaders debate how to responsibly…
❤5
OpenAI announced ChatGPT Gov, a version of ChatGPT that government agencies can deploy in their own MS Azure commercial or government cloud environment.
OpenAI
Introducing ChatGPT Gov
ChatGPT Gov is designed to streamline government agencies’ access to OpenAI’s frontier models.
💅4
Researchers has released OpenThoughts, a large-scale open-source dataset for training AI reasoning models.
Along with the dataset, they've introduced OpenThinker-7B, a new model showing promising results in mathematical and code reasoning tasks.
The project is a collaboration between researchers and engineers from Bespoke Labs, Stanford, UC Berkeley, University of Washington, Juelich Supercomputing Center, LAION, UCLA, UNC Chapel Hill, and Toyota Research Institute.
Their first release includes the OpenThoughts-114k dataset, specifically designed to improve AI reasoning capabilities.
Initial benchmarks show impressive performance, with OpenThinker-7B achieving scores of 43.3 on AIME24 and 83.0 on MATH500, approaching the performance of DeepSeek's distilled models.
The results were evaluated using their open-source tool Evalchemy.
The initiative is supported by major organizations including NSF IFML, UT Austin Machine Learning Lab, Juelich Supercomputing Center, Toyota Research Institute, and Lambda Labs.
Code
OpenThoughts-114k Dataset
OpenThinker-7B model
Along with the dataset, they've introduced OpenThinker-7B, a new model showing promising results in mathematical and code reasoning tasks.
The project is a collaboration between researchers and engineers from Bespoke Labs, Stanford, UC Berkeley, University of Washington, Juelich Supercomputing Center, LAION, UCLA, UNC Chapel Hill, and Toyota Research Institute.
Their first release includes the OpenThoughts-114k dataset, specifically designed to improve AI reasoning capabilities.
Initial benchmarks show impressive performance, with OpenThinker-7B achieving scores of 43.3 on AIME24 and 83.0 on MATH500, approaching the performance of DeepSeek's distilled models.
The results were evaluated using their open-source tool Evalchemy.
The initiative is supported by major organizations including NSF IFML, UT Austin Machine Learning Lab, Juelich Supercomputing Center, Toyota Research Institute, and Lambda Labs.
Code
OpenThoughts-114k Dataset
OpenThinker-7B model
GitHub
GitHub - open-thoughts/open-thoughts: Fully open data curation for reasoning models
Fully open data curation for reasoning models. Contribute to open-thoughts/open-thoughts development by creating an account on GitHub.
🔥6❤5🆒5
Hugging Face wants to reverse engineer DeepSeek’s R1 reasoning model
Hugging Face researchers say the Open-R1 project aims to create a fully open-source duplicate of the R1 model and make all of its components available to the AI community.
Elie Bakouch, one of the Hugging Face engineers leading the project, told TechCrunch that though DeepSeek claims R1 is open-source because it can be used without any restrictions, the truth is that it doesn’t meet the standard definition of open software. That’s because many of the components used to build it, and also the data it was trained on, have not been made publicly available.
The lack of information about what goes into DeepSeek means that it’s really just another “black box,” similar to proprietary models such as OpenAI’s GPT series, making it impossible for the AI community to build on or improve, he said.
Hugging Face says it’s attempting to replicate R1 to benefit the AI research community, and it intends to do so in just a few weeks.
To do this, it will leverage the company’s dedicated research server, the “Science Cluster,” which is powered by 768 Nvidia H100 GPUs. The plan is to try to reverse engineer the R1 model to try and understand what data was used to train it, and which components were used in its creation.
The Open-R1 project is seeking assistance from the broader AI research community to try and recreate the training datasets used by DeepSeek, and it has garnered a lot of interest so far, with its associated GitHub page getting more than 100,000 stars just three days after its launch.
Hugging Face researchers say the Open-R1 project aims to create a fully open-source duplicate of the R1 model and make all of its components available to the AI community.
Elie Bakouch, one of the Hugging Face engineers leading the project, told TechCrunch that though DeepSeek claims R1 is open-source because it can be used without any restrictions, the truth is that it doesn’t meet the standard definition of open software. That’s because many of the components used to build it, and also the data it was trained on, have not been made publicly available.
The lack of information about what goes into DeepSeek means that it’s really just another “black box,” similar to proprietary models such as OpenAI’s GPT series, making it impossible for the AI community to build on or improve, he said.
Hugging Face says it’s attempting to replicate R1 to benefit the AI research community, and it intends to do so in just a few weeks.
To do this, it will leverage the company’s dedicated research server, the “Science Cluster,” which is powered by 768 Nvidia H100 GPUs. The plan is to try to reverse engineer the R1 model to try and understand what data was used to train it, and which components were used in its creation.
The Open-R1 project is seeking assistance from the broader AI research community to try and recreate the training datasets used by DeepSeek, and it has garnered a lot of interest so far, with its associated GitHub page getting more than 100,000 stars just three days after its launch.
GitHub
GitHub - huggingface/open-r1: Fully open reproduction of DeepSeek-R1
Fully open reproduction of DeepSeek-R1. Contribute to huggingface/open-r1 development by creating an account on GitHub.
🔥5👏4
OpenAI says it has evidence DeepSeek used its model to train competitor
OpenAI suspects that DeepSeek used the "distillation" technique.
While distillation is a common practice in the industry, the issue is that DeepSeek may have been doing this to create their own competing model, which violates OpenAI's terms of service.
OpenAI suspects that DeepSeek used the "distillation" technique.
While distillation is a common practice in the industry, the issue is that DeepSeek may have been doing this to create their own competing model, which violates OpenAI's terms of service.
🐳4