AI & ML Papers
33.2K subscribers
7.14K photos
547 videos
24 files
7.83K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation

📝 Summary:
MultiMed-ST, a large-scale multilingual medical speech translation dataset, is introduced. With 290,000 samples in five languages, it is the largest medical MT and multilingual ST dataset. This work also provides an extensive comparative analysis.

🔹 Publication Date: Published on Apr 4

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2504.03546
• PDF: https://arxiv.org/pdf/2504.03546
• Project Page: https://github.com/leduckhai/MultiMed-ST
• Github: https://github.com/leduckhai/MultiMed-ST

🔹 Models citing this paper:
https://huggingface.co/leduckhai/MultiMed-ST

Datasets citing this paper:
https://huggingface.co/datasets/leduckhai/MultiMed-ST

Spaces citing this paper:
https://huggingface.co/spaces/HaoVuong/MedicalASR

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#SpeechTranslation #MedicalAI #MultilingualNLP #MachineTranslation #Dataset
MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

📝 Summary:
MeViS is a multi-modal dataset for referring motion expression video segmentation, addressing the need to segment and track objects based on their motion descriptions. It provides text and audio annotations for complex videos, enabling research into motion-guided video understanding.

🔹 Publication Date: Published on Dec 11

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.10945
• PDF: https://arxiv.org/pdf/2512.10945
• Project Page: https://henghuiding.com/MeViS/

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoSegmentation #MultiModalAI #ComputerVision #Dataset #MotionUnderstanding
2
SecureCode v2.0: A Production-Grade Dataset for Training Security-Aware Code Generation Models

📝 Summary:
SecureCode v2.0 is a production-grade dataset of 1215 security-focused coding examples. It trains AI models to generate secure code by providing real-incident examples with vulnerable and secure implementations, attacks, defense, and operational security context across 11 languages, using a conve...

🔹 Publication Date: Published on Dec 20

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.18542
• PDF: https://arxiv.org/pdf/2512.18542
• Project Page: https://perfecxion.ai/
• Github: https://github.com/scthornton/securecode-v2

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#Cybersecurity #CodeSecurity #AI #CodeGeneration #Dataset
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation

📝 Summary:
This work presents an automated rubric generation framework and RubricHub dataset for open-ended AI generation. RubricHub enables significant performance gains, achieving state-of-the-art results on HealthBench and surpassing GPT-5.

🔹 Publication Date: Published on Jan 13

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.08430
• PDF: https://arxiv.org/pdf/2601.08430
• Project Page: https://huggingface.co/datasets/sojuL/RubricHub_v1
• Github: https://github.com/teqkilla/RubricHub

Datasets citing this paper:
https://huggingface.co/datasets/sojuL/RubricHub_v1

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AI #GenerativeAI #MachineLearning #NLP #Dataset
PubMed-OCR: PMC Open Access OCR Annotations

📝 Summary:
PubMed-OCR is a corpus of 209.5K scientific articles from PubMed Central with Google Cloud Vision OCR annotations. It provides word, line, and paragraph bounding boxes to support layout-aware modeling and OCR evaluation. This data is publicly released.

🔹 Publication Date: Published on Jan 16

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.11425
• PDF: https://arxiv.org/pdf/2601.11425

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#OCR #Dataset #ComputerVision #MachineLearning #DataScience
3
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images

📝 Summary:
ExStrucTiny is a new benchmark dataset for structured information extraction from document images. It addresses limitations of existing datasets by covering diverse document types and flexible schemas. This aims to improve generalist models for structured information extraction.

🔹 Publication Date: Published on Feb 12

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12203
• PDF: https://arxiv.org/pdf/2602.12203

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#InformationExtraction #DocumentAI #MachineLearning #Dataset #ComputerVision
StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation

📝 Summary:
StereoAdapter-2 improves underwater stereo depth estimation by replacing ConvGRU with a ConvSS2D operator for efficient, long-range disparity propagation. It also introduces UW-StereoDepth-80K, a new large-scale synthetic dataset. This approach achieves state-of-the-art zero-shot performance on u...

🔹 Publication Date: Published on Feb 18

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.16915
• PDF: https://arxiv.org/pdf/2602.16915
• Project Page: https://aigeeksgroup.github.io/StereoAdapter-2
• Github: https://aigeeksgroup.github.io/StereoAdapter-2

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#UnderwaterAI #ComputerVision #DeepLearning #StereoVision #Dataset
1
DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

📝 Summary:
DLT-Corpus is a large new dataset for Distributed Ledger Technology research from scientific literature, patents, and social media. It reveals technologies originate in science before reaching patents and social media. Scientific and patent activity independently drive economic growth, unaffected...

🔹 Publication Date: Published on Feb 25

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.22045
• PDF: https://arxiv.org/pdf/2602.22045
• Github: https://github.com/dlt-science/DLT-Corpus

🔹 Models citing this paper:
https://huggingface.co/ExponentialScience/LedgerBERT
https://huggingface.co/ExponentialScience/LedgerBERT-Market-Sentiment

Datasets citing this paper:
https://huggingface.co/datasets/ExponentialScience/DLT-Patents
https://huggingface.co/datasets/ExponentialScience/DLT-Tweets
https://huggingface.co/datasets/ExponentialScience/DLT-Sentiment-News

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#DLT #Dataset #DataScience #Research #Blockchain
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

📝 Summary:
SWE-rebench V2 presents a new language-agnostic automated pipeline to create a large-scale dataset of over 32,000 software engineering tasks across 20 languages and 3,600 repositories. It provides reproducible environments and reliable tests, validated by LLMs, to advance training for SWE agents.

🔹 Publication Date: Published on Feb 27

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.23866
• PDF: https://arxiv.org/pdf/2602.23866
• Github: https://huggingface.co/collections/nebius/swe-rebench-v2

Datasets citing this paper:
https://huggingface.co/datasets/nebius/SWE-rebench-V2
https://huggingface.co/datasets/nebius/SWE-rebench-V2-PRs

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#SoftwareEngineering #LLMs #AI #Dataset #SWEAgents
Do What I Say: A Spoken Prompt Dataset for Instruction-Following

📝 Summary:
DoWhatISay is a new multilingual dataset of human-recorded spoken and written prompts for evaluating Speech Large Language Models. It reveals text prompts consistently outperform spoken prompts, except in speech-output tasks. This highlights the need for speech-based SLLM evaluation.

🔹 Publication Date: Published on Mar 10

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.09881
• PDF: https://arxiv.org/pdf/2603.09881
• Project Page: https://huggingface.co/collections/meetween/meetweens-research-papers
• Github: https://github.com/MaikeZuefle/DOWIS

Datasets citing this paper:
https://huggingface.co/datasets/maikezu/dowis

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#SLLM #SpeechAI #LLM #PromptEngineering #Dataset