AI & ML Papers
33.2K subscribers
7.16K photos
550 videos
24 files
7.85K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages

📝 Summary:
IndicParam is a new benchmark with over 13000 multiple-choice questions for 11 low-resource Indic languages. It reveals that even top LLMs achieve only ~45% accuracy, showing limitations in cross-lingual transfer and grammatical proficiency. The benchmark also assesses diverse question formats.

🔹 Publication Date: Published on Nov 29

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.00333
• PDF: https://arxiv.org/pdf/2512.00333
• Project Page: https://huggingface.co/datasets/bharatgenai/IndicParam
• Github: https://github.com/ayushbits/IndicParam

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#LLM #NLP #LowResourceLanguages #IndicLanguages #AIResearch
1
🤖🧠 Omnilingual ASR: Meta’s Breakthrough in Multilingual Speech Recognition for 1600+ Languages

🗓️ 24 Nov 2025
📚 AI News & Trends

In an increasingly connected world, speech technology plays a vital role in bridging communication gaps across languages and cultures. Yet, despite rapid progress in Automatic Speech Recognition (ASR), most commercial systems still cater to only a few dozen major languages. Billions of people who speak lesser-known or low-resource languages remain excluded from the benefits of ...

#OmnilingualASR #MultilingualSpeechRecognition #MetaAI #LowResourceLanguages #SpeechTechnology #GlobalCommunication
1
Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices

📝 Summary:
Flavors of Moonshine are tiny monolingual ASR models for underrepresented languages. They outperform larger multilingual models by using balanced data, achieving 48% lower error rates. This enables accurate on-device speech recognition.

🔹 Publication Date: Published on Sep 2, 2025

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2509.02523
• PDF: https://arxiv.org/pdf/2509.02523
• Github: https://github.com/moonshine-ai/moonshine

🔹 Models citing this paper:
https://huggingface.co/UsefulSensors/moonshine-tiny-ja
https://huggingface.co/UsefulSensors/moonshine-tiny-ar
https://huggingface.co/UsefulSensors/moonshine-tiny-zh

Spaces citing this paper:
https://huggingface.co/spaces/wmoto-ai/moonshine-tiny-ja-demo

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#ASR #EdgeAI #LowResourceLanguages #MachineLearning #TinyML
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report

📝 Summary:
OpenLIDv3 improves language identification for closely related and low resource languages. It uses enhanced training data, cluster merging, and noise detection. This significantly boosts precision over prior tools.

🔹 Publication Date: Published on Feb 13

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.13139
• PDF: https://arxiv.org/pdf/2602.13139
• Project Page: https://huggingface.co/HPLT/OpenLID-v3
• Github: https://github.com/hplt-project/openlid

🔹 Models citing this paper:
https://huggingface.co/HPLT/OpenLID-v3

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#LanguageIdentification #NLP #LowResourceLanguages #MachineLearning #AIResearch
👍1
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language

📝 Summary:
Yor-Sarc introduces the first gold-standard dataset for sarcasm detection in Yorùbá, a low-resource African language. It offers 436 expertly annotated instances with high inter-annotator agreement and soft labels, designed to advance NLP for African languages.

🔹 Publication Date: Published on Feb 21

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.18964
• PDF: https://arxiv.org/pdf/2602.18964
• Project Page: https://arxiv.org/abs/2602.18964
• Github: https://github.com/toheebadura/yor-sarc

Datasets citing this paper:
https://huggingface.co/datasets/toheebadura/yor-sarc

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#NLP #SarcasmDetection #Yoruba #LowResourceLanguages #AfricanLanguages
1
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation

📝 Summary:
LLMs struggle with low-resource language translation due to data scarcity. WALAR, a novel RL method, uses only monolingual text to improve LLM translation by mitigating reward hacking in quality estimation models. This significantly outperforms existing multilingual LLMs.

🔹 Publication Date: Published on Mar 13

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.13045
• PDF: https://arxiv.org/pdf/2603.13045
• Github: https://github.com/LeiLiLab/WALAR

🔹 Models citing this paper:
https://huggingface.co/lyf07/LLaMAX3-8B-Alpaca-WALAR
https://huggingface.co/lyf07/Translategemma-4B-it-WALAR
https://huggingface.co/lyf07/Qwen3-8B-WALAR

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#ReinforcementLearning #LLM #MultilingualTranslation #NLP #LowResourceLanguages
Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality

📝 Summary:
XBridge combines LLMs with translation models to boost multilingual performance, especially for low-resource languages. It keeps the LLM as an English knowledge core, bridging model misalignment with lightweight mapping layers for semantic consistency without retraining the LLM.

🔹 Publication Date: Published on Mar 18

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.17512
• PDF: https://arxiv.org/pdf/2603.17512
• Github: https://github.com/ictnlp/XBridge

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#LLM #MultilingualAI #NLP #LowResourceLanguages #AIResearch