✨IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
📝 Summary:
IndicParam is a new benchmark with over 13000 multiple-choice questions for 11 low-resource Indic languages. It reveals that even top LLMs achieve only ~45% accuracy, showing limitations in cross-lingual transfer and grammatical proficiency. The benchmark also assesses diverse question formats.
🔹 Publication Date: Published on Nov 29
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.00333
• PDF: https://arxiv.org/pdf/2512.00333
• Project Page: https://huggingface.co/datasets/bharatgenai/IndicParam
• Github: https://github.com/ayushbits/IndicParam
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#LLM #NLP #LowResourceLanguages #IndicLanguages #AIResearch
📝 Summary:
IndicParam is a new benchmark with over 13000 multiple-choice questions for 11 low-resource Indic languages. It reveals that even top LLMs achieve only ~45% accuracy, showing limitations in cross-lingual transfer and grammatical proficiency. The benchmark also assesses diverse question formats.
🔹 Publication Date: Published on Nov 29
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.00333
• PDF: https://arxiv.org/pdf/2512.00333
• Project Page: https://huggingface.co/datasets/bharatgenai/IndicParam
• Github: https://github.com/ayushbits/IndicParam
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#LLM #NLP #LowResourceLanguages #IndicLanguages #AIResearch
❤1
🤖🧠 Omnilingual ASR: Meta’s Breakthrough in Multilingual Speech Recognition for 1600+ Languages
🗓️ 24 Nov 2025
📚 AI News & Trends
In an increasingly connected world, speech technology plays a vital role in bridging communication gaps across languages and cultures. Yet, despite rapid progress in Automatic Speech Recognition (ASR), most commercial systems still cater to only a few dozen major languages. Billions of people who speak lesser-known or low-resource languages remain excluded from the benefits of ...
#OmnilingualASR #MultilingualSpeechRecognition #MetaAI #LowResourceLanguages #SpeechTechnology #GlobalCommunication
🗓️ 24 Nov 2025
📚 AI News & Trends
In an increasingly connected world, speech technology plays a vital role in bridging communication gaps across languages and cultures. Yet, despite rapid progress in Automatic Speech Recognition (ASR), most commercial systems still cater to only a few dozen major languages. Billions of people who speak lesser-known or low-resource languages remain excluded from the benefits of ...
#OmnilingualASR #MultilingualSpeechRecognition #MetaAI #LowResourceLanguages #SpeechTechnology #GlobalCommunication
❤1
✨Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
📝 Summary:
Flavors of Moonshine are tiny monolingual ASR models for underrepresented languages. They outperform larger multilingual models by using balanced data, achieving 48% lower error rates. This enables accurate on-device speech recognition.
🔹 Publication Date: Published on Sep 2, 2025
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2509.02523
• PDF: https://arxiv.org/pdf/2509.02523
• Github: https://github.com/moonshine-ai/moonshine
🔹 Models citing this paper:
• https://huggingface.co/UsefulSensors/moonshine-tiny-ja
• https://huggingface.co/UsefulSensors/moonshine-tiny-ar
• https://huggingface.co/UsefulSensors/moonshine-tiny-zh
✨ Spaces citing this paper:
• https://huggingface.co/spaces/wmoto-ai/moonshine-tiny-ja-demo
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ASR #EdgeAI #LowResourceLanguages #MachineLearning #TinyML
📝 Summary:
Flavors of Moonshine are tiny monolingual ASR models for underrepresented languages. They outperform larger multilingual models by using balanced data, achieving 48% lower error rates. This enables accurate on-device speech recognition.
🔹 Publication Date: Published on Sep 2, 2025
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2509.02523
• PDF: https://arxiv.org/pdf/2509.02523
• Github: https://github.com/moonshine-ai/moonshine
🔹 Models citing this paper:
• https://huggingface.co/UsefulSensors/moonshine-tiny-ja
• https://huggingface.co/UsefulSensors/moonshine-tiny-ar
• https://huggingface.co/UsefulSensors/moonshine-tiny-zh
✨ Spaces citing this paper:
• https://huggingface.co/spaces/wmoto-ai/moonshine-tiny-ja-demo
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ASR #EdgeAI #LowResourceLanguages #MachineLearning #TinyML
✨OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
📝 Summary:
OpenLIDv3 improves language identification for closely related and low resource languages. It uses enhanced training data, cluster merging, and noise detection. This significantly boosts precision over prior tools.
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.13139
• PDF: https://arxiv.org/pdf/2602.13139
• Project Page: https://huggingface.co/HPLT/OpenLID-v3
• Github: https://github.com/hplt-project/openlid
🔹 Models citing this paper:
• https://huggingface.co/HPLT/OpenLID-v3
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#LanguageIdentification #NLP #LowResourceLanguages #MachineLearning #AIResearch
📝 Summary:
OpenLIDv3 improves language identification for closely related and low resource languages. It uses enhanced training data, cluster merging, and noise detection. This significantly boosts precision over prior tools.
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.13139
• PDF: https://arxiv.org/pdf/2602.13139
• Project Page: https://huggingface.co/HPLT/OpenLID-v3
• Github: https://github.com/hplt-project/openlid
🔹 Models citing this paper:
• https://huggingface.co/HPLT/OpenLID-v3
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#LanguageIdentification #NLP #LowResourceLanguages #MachineLearning #AIResearch
👍1
✨Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
📝 Summary:
Yor-Sarc introduces the first gold-standard dataset for sarcasm detection in Yorùbá, a low-resource African language. It offers 436 expertly annotated instances with high inter-annotator agreement and soft labels, designed to advance NLP for African languages.
🔹 Publication Date: Published on Feb 21
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.18964
• PDF: https://arxiv.org/pdf/2602.18964
• Project Page: https://arxiv.org/abs/2602.18964
• Github: https://github.com/toheebadura/yor-sarc
✨ Datasets citing this paper:
• https://huggingface.co/datasets/toheebadura/yor-sarc
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#NLP #SarcasmDetection #Yoruba #LowResourceLanguages #AfricanLanguages
📝 Summary:
Yor-Sarc introduces the first gold-standard dataset for sarcasm detection in Yorùbá, a low-resource African language. It offers 436 expertly annotated instances with high inter-annotator agreement and soft labels, designed to advance NLP for African languages.
🔹 Publication Date: Published on Feb 21
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.18964
• PDF: https://arxiv.org/pdf/2602.18964
• Project Page: https://arxiv.org/abs/2602.18964
• Github: https://github.com/toheebadura/yor-sarc
✨ Datasets citing this paper:
• https://huggingface.co/datasets/toheebadura/yor-sarc
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#NLP #SarcasmDetection #Yoruba #LowResourceLanguages #AfricanLanguages
❤1
✨Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
📝 Summary:
LLMs struggle with low-resource language translation due to data scarcity. WALAR, a novel RL method, uses only monolingual text to improve LLM translation by mitigating reward hacking in quality estimation models. This significantly outperforms existing multilingual LLMs.
🔹 Publication Date: Published on Mar 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.13045
• PDF: https://arxiv.org/pdf/2603.13045
• Github: https://github.com/LeiLiLab/WALAR
🔹 Models citing this paper:
• https://huggingface.co/lyf07/LLaMAX3-8B-Alpaca-WALAR
• https://huggingface.co/lyf07/Translategemma-4B-it-WALAR
• https://huggingface.co/lyf07/Qwen3-8B-WALAR
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ReinforcementLearning #LLM #MultilingualTranslation #NLP #LowResourceLanguages
📝 Summary:
LLMs struggle with low-resource language translation due to data scarcity. WALAR, a novel RL method, uses only monolingual text to improve LLM translation by mitigating reward hacking in quality estimation models. This significantly outperforms existing multilingual LLMs.
🔹 Publication Date: Published on Mar 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.13045
• PDF: https://arxiv.org/pdf/2603.13045
• Github: https://github.com/LeiLiLab/WALAR
🔹 Models citing this paper:
• https://huggingface.co/lyf07/LLaMAX3-8B-Alpaca-WALAR
• https://huggingface.co/lyf07/Translategemma-4B-it-WALAR
• https://huggingface.co/lyf07/Qwen3-8B-WALAR
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ReinforcementLearning #LLM #MultilingualTranslation #NLP #LowResourceLanguages
✨Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality
📝 Summary:
XBridge combines LLMs with translation models to boost multilingual performance, especially for low-resource languages. It keeps the LLM as an English knowledge core, bridging model misalignment with lightweight mapping layers for semantic consistency without retraining the LLM.
🔹 Publication Date: Published on Mar 18
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.17512
• PDF: https://arxiv.org/pdf/2603.17512
• Github: https://github.com/ictnlp/XBridge
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#LLM #MultilingualAI #NLP #LowResourceLanguages #AIResearch
📝 Summary:
XBridge combines LLMs with translation models to boost multilingual performance, especially for low-resource languages. It keeps the LLM as an English knowledge core, bridging model misalignment with lightweight mapping layers for semantic consistency without retraining the LLM.
🔹 Publication Date: Published on Mar 18
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.17512
• PDF: https://arxiv.org/pdf/2603.17512
• Github: https://github.com/ictnlp/XBridge
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#LLM #MultilingualAI #NLP #LowResourceLanguages #AIResearch