All about AI, Web 3.0, BCI
3.88K subscribers
784 photos
29 videos
162 files
3.65K links
This channel about AI, Web 3.0 and brain computer interface(BCI)

owner @Aniaslanyan
Download Telegram
DeepSeek just dropped another open-source AI model, Janus-Pro-7B

It's multimodal (can generate images) and beats OpenAI's DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks.

This comes on top of all the R1 hype.
🔥15🆒32
Alibaba released the SOTA open multimodal model, Qwen2.5-VL

It shows significant improvements across various aspects compared to the previous version.

Key Highlights:

1. Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all!

2. Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones.

3. Long Video Comprehension : Captures events in videos over 1 hour long

4. Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection.

5. Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more.

ModelScope.
The key insights from Yann Lecun's speech at the #WEF2025

At a very interesting debate on technology he covered:

- The perspectives of AI in the next 3-5 years
- What intelligent AI systems need
- Why LLMs are not enough
- True diversity in AI
- Open source importance
and more

How do we make sure the AI we build is the AI we want?

Yann LeCun said that where we are now with Gen AI isn't where we want to be. Gen AI isn't very controllable, but researchers try to make it applicable to a wide range of areas.

The perspectives of AI:

In the next 3-5 years, we are going to see the emergence of a new revolutionary brand or paradigm for AI architectures.

- This new AI won't be generative as we see generative AI today.
- It will be able to plan the sequence of actions and the sequence of plans.

Then we'll come to agentic AI and then to robotics. And the coming decade may be the era of robotics.

4 things that are essential for intelligent AI systems (and are limitations for current systems):

- Understanding the physical world
- Having persistent memory
- Being capable of reasoning
- Complex planning capabilities

LLMs' place in the future of AI:

Yann Lecun considers LLMs as a part of a bigger future AI system, because they are good at manipulating language but bad at thinking.

We're never going to get to human-level AI just with text models. AI needs to know how the real world works from sensory data, as well.

Why is language simple?

Language, like DNA and proteins, is a discrete object, and it's easy to make predictions in discrete worlds.

Physical AI is much more difficult. That's why techniques used in LLMs can't be used for predicting videos.

Open source matters the most:

Yann Lecun highlighted that open source makes AI tool accessible to everyone.
It's important for cultural diversity and democracy because a wide diversity of AI assistants can only be achieved with open source.

How to achieve true diversity in AI:

True diversity is when we have models that are trained on all the languages, cultures, and values of the world.

That's where foundational models can be the base that is fine-tuned for different ideas and visions of what good value systems are, so people can choose from many options.

Federated learning:

One of the ways to achieve diversity is federated learning. It's when every region in the world has its own datasets which contribute to training a big global model.

How can AI be valued in terms of disinformation and toxic content?

With transformers and self-supervised learning appeared, the proportion of hate speech taken down by AI systems has reached 96%.

However, high false positives (good content removed) are a problem. Therefore, detection thresholds will be adjusted to enable discussions on key societal topics.

Regulations:

"Making the distribution of open source AI engine essentially illegal is way more dangerous than all the other potential dangers."
5
OpenAI announced ChatGPT Gov, a version of ChatGPT that government agencies can deploy in their own MS Azure commercial or government cloud environment.
💅4
Researchers has released OpenThoughts, a large-scale open-source dataset for training AI reasoning models.

Along with the dataset, they've introduced OpenThinker-7B, a new model showing promising results in mathematical and code reasoning tasks.

The project is a collaboration between researchers and engineers from Bespoke Labs, Stanford, UC Berkeley, University of Washington, Juelich Supercomputing Center, LAION, UCLA, UNC Chapel Hill, and Toyota Research Institute.

Their first release includes the OpenThoughts-114k dataset, specifically designed to improve AI reasoning capabilities.

Initial benchmarks show impressive performance, with OpenThinker-7B achieving scores of 43.3 on AIME24 and 83.0 on MATH500, approaching the performance of DeepSeek's distilled models.

The results were evaluated using their open-source tool Evalchemy.

The initiative is supported by major organizations including NSF IFML, UT Austin Machine Learning Lab, Juelich Supercomputing Center, Toyota Research Institute, and Lambda Labs.

Code
OpenThoughts-114k Dataset
OpenThinker-7B model
🔥65🆒5
Hugging Face wants to reverse engineer DeepSeek’s R1 reasoning model

Hugging Face researchers say the Open-R1 project aims to create a fully open-source duplicate of the R1 model and make all of its components available to the AI community.

Elie Bakouch, one of the Hugging Face engineers leading the project, told TechCrunch that though DeepSeek claims R1 is open-source because it can be used without any restrictions, the truth is that it doesn’t meet the standard definition of open software. That’s because many of the components used to build it, and also the data it was trained on, have not been made publicly available.

The lack of information about what goes into DeepSeek means that it’s really just another “black box,” similar to proprietary models such as OpenAI’s GPT series, making it impossible for the AI community to build on or improve, he said.

Hugging Face says it’s attempting to replicate R1 to benefit the AI research community, and it intends to do so in just a few weeks.

To do this, it will leverage the company’s dedicated research server, the “Science Cluster,” which is powered by 768 Nvidia H100 GPUs. The plan is to try to reverse engineer the R1 model to try and understand what data was used to train it, and which components were used in its creation.

The Open-R1 project is seeking assistance from the broader AI research community to try and recreate the training datasets used by DeepSeek, and it has garnered a lot of interest so far, with its associated GitHub page getting more than 100,000 stars just three days after its launch.
🔥5👏4
OpenAI says it has evidence DeepSeek used its model to train competitor

OpenAI suspects that DeepSeek used the "distillation" technique.

While distillation is a common practice in the industry, the issue is that DeepSeek may have been doing this to create their own competing model, which violates OpenAI's terms of service.
🐳4
Why #DeepSeek's Success Doesn't Change the AI Race: Dario Amodei's View

Anthropic's CEO explains why the apparent breakthrough fits into the expected trajectory of AI development.*

3 Laws of AI Development

1. Scaling Law:
- More resources = better results
- $1M = 20% tasks, $10M = 40%, $100M = 60%
- Progress is smooth and predictable

2. Curve Shifting:
- Innovations improve efficiency
- Typical improvements:
* Small (1.2x)
* Medium (2x)
* Large (10x)
- Overall pace: ~4x per year

3. Paradigm Shifts:
- 2020-2023: text training
- 2024: adding Reinforcement Learning
- Now: unique "crossover point"

What DeepSeek actually achieved?
- Performance similar to 7-10 month old US models
- Lower costs, but within normal trend
- Significant resources (~50,000 chips, ~$1B)


Not a revolution because:
- Cost reduction follows expected 4x/year trend
- V3 more innovative than R1
- Total company spending comparable to US labs

The Future (2026-2027)

According to Amodei, truly advanced AI will require:
- Millions of chips
- Tens of billions of dollars
- 2-3 years of work

Key Takeaway

DeepSeek demonstrates an expected point on the progress curve, not a revolutionary breakthrough. The real race for superhuman AI is just beginning, and it will require unprecedented resources.

"Making AI that is smarter than almost all humans at almost all things will require millions of chips, tens of billions of dollars (at least), and is most likely to happen in 2026-2027"- Dario Amodei
42
Voice Agents Market Map - B2B by Andreessen Horowitz - an update on AI voice agents

2024 was a huge year for voice agents, with massive model breakthroughs + hundreds of new startups.

How has the space evolved, and where do we see opportunity in 2025?

To start: voice as one of AI's biggest unlocks. Latency and reliability are now largely solved - and interruptibility and emotionality have made major strides, too. Voice AI is now nearly at human standards, allowing tech to replace labor on the phone.

As a result, there's been an explosion in startups building applications on these models. Y Combinator alone has seen 90 voice agent cos. Many are targeting specific verticals - by industry (e.g. home services, dental) or function (e.g. recruiting, customer support) - and are scaling fast!

Most companies need to tap into adjacent workflows: pushing call details to a CRM, automating follow-ups, etc.

What a16Z're looking for in voice agent startups:

1. Building in an industry where phone is the preferred or required medium, or has a much higher success rate than other modalities

2. Calls are constrained — both in length and in format/outcome

3. Voice agent delivers 50%+ cost reduction with similar success rate as a human

4. Calls are "life or death" for the customer —they will pay significant $ to get them made or answered...but not for the end consumer

5. When selling into SMB/mid-market, the agent product has easy integration. When selling into enterprise, complex integration can be a moat!
🔥32
Trump Media and Technology Group announced the launch of the financial services and FinTech brand TruthFi.

$250 million will be invested in SMAs, ETFs, Bitcoin and similar cryptocurrencies or crypto-related securities. TrueFi will enter the field of decentralized finance.
👍3👀1
Mistral AI released a new small model

Mistral Small 3, a latency-optimized 24B-parameter model released under the Apache 2.0 license.

Mistral Small 3 is competitive with larger models such as Llama 3.3 70B or Qwen 32B, and is an excellent open replacement for opaque proprietary models like GPT4o-mini.

Mistral Small 3 is on par with Llama 3.3 70B instruct, while being more than 3x faster on the same hardware.

Mistral Small 3 is a pre-trained and instructed model catered to the ‘80%’ of generative AI tasks—those that require robust language and instruction following performance, with very low latency.
🆒52
Now available on the Allen Brain Cell Atlas - aging mouse brain data

Contains 1.2 million high-quality, single-cell transcriptomes.
Google rolled out Gemini 2 Flash for free to its products.

— one of the most high quality non-reasoning LLMs
— super fast (150tok/s+)
— 1M tok context window.

API price isnt out but was previously $0.075/$0.30 per M input/output tokens. Big move from Google.
AI Distillation Race: From $450 Berkeley Experiment to Industry Disruption

In a fascinating turn of events in AI development, UC Berkeley doctoral students demonstrated that advanced AI capabilities can be replicated for just $450 in computing costs.

This comes amid industry buzz about #DeepSeek's R1 model, which allegedly used similar distillation techniques to replicate OpenAI's reasoning capabilities.

The Berkeley breakthrough:

- Used Alibaba's Qwen model to generate 17,000 training examples
- Focused on math and coding problems with verifiable answers
- Their model outperformed OpenAI's first reasoning model on several benchmarks

The bigger picture:
- OpenAI claims Chinese quant fund DeepSeek used distillation to replicate their o1 reasoning model
- While OpenAI tries to hide their models' thought processes, DeepSeek took an open approach with R1
- This transparency means other developers could potentially replicate R1's capabilities

The Berkeley case and DeepSeek's R1 show that AI innovation can be democratized. While tech giants invest billions in AI development, smaller teams can achieve impressive results through clever engineering and efficient training methods.
1
OpenAI just released o3-mini, available today to all users in ChatGPT (for free)!
OpenAI announced Deep Research, a new ChatGPT model that can conduct autonomous research by searching the web and synthesizing knowledge into a research paper as output, in what is the next step towards AI discovering new knowledge for itself

OpenAI's Deep Research scores a new state-of-the-art score on Humanity's Last Exam, surpassing the accuracy of o3-mini.

Sam Altman says new Deep Research AI agent powered by their o3 model can do "a single-digit percentage of all economically valuable tasks in the world" and this step into synthesizing knowledge will soon be followed by agents that can invent new knowledge.
😁4🥴2
TSMC plans to build a massive 1nm fab in southern Taiwan a type calls a Giga Fab, able to fit as many production lines as 6 normal 12-inch wafer fabs

The investment would be part of the government’s "Greater Southern New Silicon Valley Promotion Plan," aimed at focusing new semiconductor investments in southern Taiwan.

TSMC reportedly said there are many possible locations for the fab.