The next OpenAI frontier model has started training
OpenAI
OpenAI Board Forms Safety and Security Committee
RoboCasa is a large-scale simulation framework of everyday tasks
Researchers use generative AI tools to create diverse objects, scenes, and tasks. Simulation plays a pivotal role in our Data Pyramid for training generalist robots.
It's all open-source of course.
Researchers use generative AI tools to create diverse objects, scenes, and tasks. Simulation plays a pivotal role in our Data Pyramid for training generalist robots.
It's all open-source of course.
Big news โ Gemini 1.5 Flash, Pro and Advanced results are out!๐ฅ
- Gemini 1.5 Pro/Advanced at #2, closing in on GPT-4o
- Gemini 1.5 Flash at #9, outperforming Llama-3-70b and nearly reaching GPT-4-0125 (!)
Pro is significantly stronger than its April version. Flashโs cost, capabilities, and unmatched context length make it a market game-changer!
In Chinese, Gemini 1.5 Pro & Advanced are now the best #1 model in the world. Flash becomes even stronger!
Gemini family remains top in our new "Hard Prompts" category, which features more challenging, problem-solving user queries.
Learn more about Hard Prompts.
- Gemini 1.5 Pro/Advanced at #2, closing in on GPT-4o
- Gemini 1.5 Flash at #9, outperforming Llama-3-70b and nearly reaching GPT-4-0125 (!)
Pro is significantly stronger than its April version. Flashโs cost, capabilities, and unmatched context length make it a market game-changer!
In Chinese, Gemini 1.5 Pro & Advanced are now the best #1 model in the world. Flash becomes even stronger!
Gemini family remains top in our new "Hard Prompts" category, which features more challenging, problem-solving user queries.
Learn more about Hard Prompts.
lmsys.org
Introducing Hard Prompts Category in Chatbot Arena | LMSYS Org
<h3><a id="background" class="anchor" href="#background" aria-hidden="true"><svg aria-hidden="true" class="octicon octicon-link" height="16" version="1.1" vi...
Researchers introduced SignLLM.
It's the first multilingual Sign Language Production (SLP) AI model capable of generating avatar videos of sign language gestures from prompts across eight languages.
It's the first multilingual Sign Language Production (SLP) AI model capable of generating avatar videos of sign language gestures from prompts across eight languages.
What is going on? OpenAI has said its mission is not to build "superintelligence" in an apparent backtrack from previous comments by Sam Altman, as it readies its new model.
Mistral released first code model, Codestral-22B:
- Outperforms all open code models
- Fluent in 80+ languages, 32K context length
- Available on our API and for free on Le Chat
- Integrated with VS Code
Weights.
- Outperforms all open code models
- Fluent in 80+ languages, 32K context length
- Available on our API and for free on Le Chat
- Integrated with VS Code
Weights.
Mistral AI
Codestral | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
Scale launched SEAL Leaderboardsโprivate, expert evaluations of leading frontier models.
Evaluations are a critical component of the AI ecosystem.
Evals are incentives for researchers, and our evaluations set the goals for how we aim to improve our models.
Trusted 3rd party evals are a missing part of the whole ecosystem, which is why Scale built these.
They eval'd many of the leading models:
- GPT-4o
- GPT-4 Turbo
- Claude 3 Opus
- Gemini 1.5 Pro
- Gemini 1.5 Flash
- Llama3
- Mistral Large
On Coding, Math, Instruction Following, and Multilinguality (Spanish).
Evaluations are a critical component of the AI ecosystem.
Evals are incentives for researchers, and our evaluations set the goals for how we aim to improve our models.
Trusted 3rd party evals are a missing part of the whole ecosystem, which is why Scale built these.
They eval'd many of the leading models:
- GPT-4o
- GPT-4 Turbo
- Claude 3 Opus
- Gemini 1.5 Pro
- Gemini 1.5 Flash
- Llama3
- Mistral Large
On Coding, Math, Instruction Following, and Multilinguality (Spanish).
Scale Labs Leaderboards
AI Model Leaderboards & Benchmarks
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
โค4
Researchers introduced Geometry-Informed Neural Networks to train shape generative models
without any data (!!), combining learning under constraints, neural fields as a suitable representation, and generating diverse solutions to under-determined problems.
Paper.
without any data (!!), combining learning under constraints, neural fields as a suitable representation, and generating diverse solutions to under-determined problems.
Paper.
arXiv.org
Geometry-Informed Neural Networks
Geometry is a ubiquitous tool in computer graphics, design, and engineering. However, the lack of large shape datasets limits the application of state-of-the-art supervised learning methods and...
AI_AMA_1717076968.pdf
1.3 MB
๐ฅ๐ฒ๐ฝ๐ผ๐ฟ๐: ๐ง๐ต๐ฒ ๐๐๐๐๐ฟ๐ฒ ๐ผ๐ณ ๐๐ฒ๐ฎ๐น๐๐ต โ ๐ง๐ต๐ฒ ๐๐บ๐ฒ๐ฟ๐ด๐ถ๐ป๐ด ๐๐ฎ๐ป๐ฑ๐๐ฐ๐ฎ๐ฝ๐ฒ ๐ผ๐ณ ๐๐๐ด๐บ๐ฒ๐ป๐๐ฒ๐ฑ ๐๐ป๐๐ฒ๐น๐น๐ถ๐ด๐ฒ๐ป๐ฐ๐ฒ ๐ถ๐ป ๐๐ฒ๐ฎ๐น๐๐ต ๐๐ฎ๐ฟ๐ฒ
๐๐ฒ๐ ๐๐ถ๐ด๐ต๐น๐ถ๐ด๐ต๐๐:
1. ๐๐ถ๐๐๐ผ๐ฟ๐ถ๐ฐ๐ฎ๐น ๐๐ผ๐ป๐๐ฒ๐ ๐: AI has been a part of medicine since the mid-20th century, with its role significantly expanding in fields like radiology, cardiology, and neurology.
2. ๐ง๐ฒ๐ฐ๐ต๐ป๐ผ๐น๐ผ๐ด๐ถ๐ฐ๐ฎ๐น ๐๐ฑ๐๐ฎ๐ป๐ฐ๐ฒ๐บ๐ฒ๐ป๐๐: Recent innovations in deep learning and foundation models are unlocking new use cases in healthcare, from cancer prognoses to predicting adverse clinical events.
3. ๐ฃ๐ต๐๐๐ถ๐ฐ๐ถ๐ฎ๐ป ๐ฃ๐ฒ๐ฟ๐๐ฝ๐ฒ๐ฐ๐๐ถ๐๐ฒ๐: While 65% of surveyed physicians see definite or some advantage in using AI, 70% express concerns about potential biases, privacy risks, and liability issues.
4. ๐ฆ๐ฝ๐ฒ๐ฐ๐ถ๐ฎ๐น๐๐-๐ฆ๐ฝ๐ฒ๐ฐ๐ถ๐ณ๐ถ๐ฐ ๐จ๐๐ฒ ๐๐ฎ๐๐ฒ๐: Each medical specialty leverages AI differently, from emergency medicine's vitals monitoring to family medicine's focus on personalized patient education and medication adherence.
5. ๐๐ฑ๐บ๐ถ๐ป๐ถ๐๐๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐ฆ๐๐ฝ๐ฝ๐ผ๐ฟ๐: 56% of surveyed physicians identify administrative burden reduction through automation as AI's biggest opportunity.
๐๐ฒ๐ ๐๐ถ๐ด๐ต๐น๐ถ๐ด๐ต๐๐:
1. ๐๐ถ๐๐๐ผ๐ฟ๐ถ๐ฐ๐ฎ๐น ๐๐ผ๐ป๐๐ฒ๐ ๐: AI has been a part of medicine since the mid-20th century, with its role significantly expanding in fields like radiology, cardiology, and neurology.
2. ๐ง๐ฒ๐ฐ๐ต๐ป๐ผ๐น๐ผ๐ด๐ถ๐ฐ๐ฎ๐น ๐๐ฑ๐๐ฎ๐ป๐ฐ๐ฒ๐บ๐ฒ๐ป๐๐: Recent innovations in deep learning and foundation models are unlocking new use cases in healthcare, from cancer prognoses to predicting adverse clinical events.
3. ๐ฃ๐ต๐๐๐ถ๐ฐ๐ถ๐ฎ๐ป ๐ฃ๐ฒ๐ฟ๐๐ฝ๐ฒ๐ฐ๐๐ถ๐๐ฒ๐: While 65% of surveyed physicians see definite or some advantage in using AI, 70% express concerns about potential biases, privacy risks, and liability issues.
4. ๐ฆ๐ฝ๐ฒ๐ฐ๐ถ๐ฎ๐น๐๐-๐ฆ๐ฝ๐ฒ๐ฐ๐ถ๐ณ๐ถ๐ฐ ๐จ๐๐ฒ ๐๐ฎ๐๐ฒ๐: Each medical specialty leverages AI differently, from emergency medicine's vitals monitoring to family medicine's focus on personalized patient education and medication adherence.
5. ๐๐ฑ๐บ๐ถ๐ป๐ถ๐๐๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐ฆ๐๐ฝ๐ฝ๐ผ๐ฟ๐: 56% of surveyed physicians identify administrative burden reduction through automation as AI's biggest opportunity.
Astronomers are preparing to use AI to tackle 300 petabytes of data annually.
Cecilia Garraffo's AstroAI initiative is pioneering the fusion of AI and astronomy to explore deep cosmic questions, already planning dozens of projects with a 50-member interdisciplinary team.
Cecilia Garraffo's AstroAI initiative is pioneering the fusion of AI and astronomy to explore deep cosmic questions, already planning dozens of projects with a 50-member interdisciplinary team.
MIT Technology Review
Astronomers are enlisting AI to prepare for a data downpour
Tailored algorithms will help filter a coming flood of astronomical observations, helping scientists make new discoveries about the universe.
The Simulationโs new Showrunner platform lets you direct, star in, and even get paid for your own AI-generated TV shows
The lines are blurring fast between creators and audiences โ and the traditional Hollywood media model is changing in front of our eyes.
The lines are blurring fast between creators and audiences โ and the traditional Hollywood media model is changing in front of our eyes.
Showrunner
Gossip Goblin Competition
One Client. One Doc. One Hidden Desire. Be a part of AI filmmaking history.
Similarity is Not All You Need: Endowing Retrieval-Augmented Generation with Multiโlayered Thoughts.
Microsoft presents Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
SELM significantly boosts the performance on instructionfollowing benchmarks such as MT-Bench and AlpacaEval 2.0
Repo.
SELM significantly boosts the performance on instructionfollowing benchmarks such as MT-Bench and AlpacaEval 2.0
Repo.
arXiv.org
Self-Exploring Language Models: Active Preference Elicitation for...
Preference optimization, particularly through Reinforcement Learning from Human Feedback (RLHF), has achieved significant success in aligning Large Language Models (LLMs) to adhere to human...
Google DeepMind released new demos of its Veo AI video generation model.
The demo showcases the ability to turn single reference images into new videos with simple text instructions.
The demo showcases the ability to turn single reference images into new videos with simple text instructions.
Google
Introducing VideoFX, plus new features for ImageFX and MusicFX
Today weโre introducing VideoFX, plus new features for ImageFX and MusicFX that are now available in 110 countries.
Revolutionary work from Huggingface: FineWeb
FineWeb:
0. Pretraining is far less intuitive than instruct finetune
1. Unclear what data to include to boost performance
2. HF's team simply tried different filters [3] -> trained small models on the filtered dataset -> measured eval scores.
Result:
The first version of FineWeb. A pre-training dataset that is was methodically selected for the exact purpose of making the models you train on it get the best eval scores as possible.
FineWeb-Edu:
0. Filtered the dataset even more to include only "high educational value" texts. [1]
2. Used Llama-3-70B-Instruct [2] to extract "educational score" for 500K texts.
3. Trained a classifier to classify the rest of the 15T tokens.
Result:
FineWeb-Edu outperform every other open pre-training dataset by a huge margin.
Other treasures in the blog.
In the blog there is a ton of useful material: how to filter at scale? how to run LLMs at scale? breakdown of every single filter contribution to the score..
To the best this is the first large scale empirical study of pre-training data trying to improve an underlying model's performance.
[1] As proposed on "Textsbooks are all you need" (phi-1): arxiv.org/abs/2306.11644
[2] Language filter -> Gopher filter (from it's paper) -> MinHash -> C4 filter (simple rules in the paper) -> PII filter.
[3] Prompt from arxiv.org/abs/2401.10020
FineWeb:
0. Pretraining is far less intuitive than instruct finetune
1. Unclear what data to include to boost performance
2. HF's team simply tried different filters [3] -> trained small models on the filtered dataset -> measured eval scores.
Result:
The first version of FineWeb. A pre-training dataset that is was methodically selected for the exact purpose of making the models you train on it get the best eval scores as possible.
FineWeb-Edu:
0. Filtered the dataset even more to include only "high educational value" texts. [1]
2. Used Llama-3-70B-Instruct [2] to extract "educational score" for 500K texts.
3. Trained a classifier to classify the rest of the 15T tokens.
Result:
FineWeb-Edu outperform every other open pre-training dataset by a huge margin.
Other treasures in the blog.
In the blog there is a ton of useful material: how to filter at scale? how to run LLMs at scale? breakdown of every single filter contribution to the score..
To the best this is the first large scale empirical study of pre-training data trying to improve an underlying model's performance.
[1] As proposed on "Textsbooks are all you need" (phi-1): arxiv.org/abs/2306.11644
[2] Language filter -> Gopher filter (from it's paper) -> MinHash -> C4 filter (simple rules in the paper) -> PII filter.
[3] Prompt from arxiv.org/abs/2401.10020
arXiv.org
Textbooks Are All You Need
We introduce phi-1, a new large language model for code, with significantly smaller size than competing models: phi-1 is a Transformer-based model with 1.3B parameters, trained for 4 days on 8...
ElevenLabs launched Text to Sound Effects.
It's a new AI model that generates sound effects, short instrumental tracks, and soundscapes from text prompts.
The lines have officially fully blurred.
API access is coming soon.
It's a new AI model that generates sound effects, short instrumental tracks, and soundscapes from text prompts.
The lines have officially fully blurred.
API access is coming soon.
ElevenLabs
AI Voice Generator & Text to Speech
Rated the best text to speech (TTS) software online. Create premium AI voices for free and generate text to speech voiceovers in minutes with our character AI voice generator. Use free text to speech AI to convert text to mp3 in 29 languages with 100+ voices.
Real_Estate_Toksenization_1717419791.pdf
3 MB
Real estate tokenisation will unlock $3.5 trillion in global liquidity by 2027. This could solve one of the biggest issues in real estate โ
Illiquidity.
This report explores the detailed process of tokenising real estate, the legal and regulatory challenges, and what this could mean for the industry.
Here are key takeaways:
1. Tokenisation offers many benefits, such as increased liquidity, fractional ownership, and streamlined transactions.
2. The tokenisation process involves multiple phases, including deal structuring, digitisation, primary distribution, post-tokenisation management, and secondary trading.
3. Regulations for tokenised securities differ by country, with some countries being more open to these innovations than others.
4. Determining the value and tax obligations of tokenised assets is complex. We need to tackle these issues very carefully.
5. Challenges also include confidentiality and legal uncertainties as the legal frameworks are still evolving.
6. Opportunities extend beyond traditional investments to things like employee incentives, lease-to-own models, and co-working spaces.
Illiquidity.
This report explores the detailed process of tokenising real estate, the legal and regulatory challenges, and what this could mean for the industry.
Here are key takeaways:
1. Tokenisation offers many benefits, such as increased liquidity, fractional ownership, and streamlined transactions.
2. The tokenisation process involves multiple phases, including deal structuring, digitisation, primary distribution, post-tokenisation management, and secondary trading.
3. Regulations for tokenised securities differ by country, with some countries being more open to these innovations than others.
4. Determining the value and tax obligations of tokenised assets is complex. We need to tackle these issues very carefully.
5. Challenges also include confidentiality and legal uncertainties as the legal frameworks are still evolving.
6. Opportunities extend beyond traditional investments to things like employee incentives, lease-to-own models, and co-working spaces.
Fascinating finding! Research suggests that the #brain language regions are hardly activated when "reading" python and other programming languages
Instead, reading computer code appears to engage both left and right sides of the Multiple Demand Network (MDN).
The MDN, whose activity is spread throughout the frontal and parietal lobes of the brain, is typically recruited for tasks that require holding many pieces of information in mind at once, and is responsible for our ability to perform a wide variety of mental tasks.
Cognitive skills, such as reading, writing, map-based navigation, mathematical reasoning, and scientific logic, draw heavily on the MDN.
โThe left hemisphere is activated more when solving math and logic problems .
โ The right hemisphere activates more when doing tasks involving spatial navigation.
In a companion paper appearing in the same issue of eLife, a team of researchers from Johns Hopkins University also reported that solving code problems activates the multiple demand network rather than the language regions.
โThese findings suggest there isnโt a definitive answer to whether coding should be taught as a math-based skill or a language-based skill. In part, thatโs because learning to program may draw on both language and multiple demand systems, even if โ once learned โ programming doesnโt rely on the language regions,โ the researchers say.
Instead, reading computer code appears to engage both left and right sides of the Multiple Demand Network (MDN).
The MDN, whose activity is spread throughout the frontal and parietal lobes of the brain, is typically recruited for tasks that require holding many pieces of information in mind at once, and is responsible for our ability to perform a wide variety of mental tasks.
Cognitive skills, such as reading, writing, map-based navigation, mathematical reasoning, and scientific logic, draw heavily on the MDN.
โThe left hemisphere is activated more when solving math and logic problems .
โ The right hemisphere activates more when doing tasks involving spatial navigation.
In a companion paper appearing in the same issue of eLife, a team of researchers from Johns Hopkins University also reported that solving code problems activates the multiple demand network rather than the language regions.
โThese findings suggest there isnโt a definitive answer to whether coding should be taught as a math-based skill or a language-based skill. In part, thatโs because learning to program may draw on both language and multiple demand systems, even if โ once learned โ programming doesnโt rely on the language regions,โ the researchers say.
eLife
Comprehension of computer code relies primarily on domain-general executive brain regions
The domain-general executive brain regions support the use of a novel cognitive tool even when it is structurally similar to natural language.
๐ฆ3
So it begins, Hollywood studios being open about their leaning into AI strategies.
A cost reducing measure, going to be interesting to see the impacts of this.
A cost reducing measure, going to be interesting to see the impacts of this.
Perplexity AI
Sony Pictures Uses AI to Cut Film Costs
Sony Pictures Entertainment is embracing artificial intelligence (AI) to streamline film and television production and reduce costs. CEO Tony Vinciquerra...
๐
3
How the brain compose minimal phrases?
In this new study on how negation transforms the neural representation of adjectives, researchers combine time-resolved behavioral and neuroimaging methods to track and decode changes in meaning representation.
In this new study on how negation transforms the neural representation of adjectives, researchers combine time-resolved behavioral and neuroimaging methods to track and decode changes in meaning representation.
journals.plos.org
Negation mitigates rather than inverts the neural representations of adjectives
Negation is an important concept in natural language but we still know relatively little about the cognitive and neural mechanisms underpinning negation. Combining behavioral and neurophysiological data, this study shows that humans first interpret negatedโฆ
The Bank for International Settlements (BIS) is launching Project Rialto to explore how instant cross-border payments could be improved using a modular foreign exchange component combined with settlement in wholesale central bank digital currencies ( #wCBDC ).
www.bis.org
Project Rialto: improving instant cross-border payments using central bank money settlement
Project Rialto explored how instant cross-border payments could be improved using a modular foreign exchange (FX) component combined with settlement in tokenised wholesale central bank money.