✨MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
📝 Summary:
MultiMed-ST, a large-scale multilingual medical speech translation dataset, is introduced. With 290,000 samples in five languages, it is the largest medical MT and multilingual ST dataset. This work also provides an extensive comparative analysis.
🔹 Publication Date: Published on Apr 4
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2504.03546
• PDF: https://arxiv.org/pdf/2504.03546
• Project Page: https://github.com/leduckhai/MultiMed-ST
• Github: https://github.com/leduckhai/MultiMed-ST
🔹 Models citing this paper:
• https://huggingface.co/leduckhai/MultiMed-ST
✨ Datasets citing this paper:
• https://huggingface.co/datasets/leduckhai/MultiMed-ST
✨ Spaces citing this paper:
• https://huggingface.co/spaces/HaoVuong/MedicalASR
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SpeechTranslation #MedicalAI #MultilingualNLP #MachineTranslation #Dataset
📝 Summary:
MultiMed-ST, a large-scale multilingual medical speech translation dataset, is introduced. With 290,000 samples in five languages, it is the largest medical MT and multilingual ST dataset. This work also provides an extensive comparative analysis.
🔹 Publication Date: Published on Apr 4
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2504.03546
• PDF: https://arxiv.org/pdf/2504.03546
• Project Page: https://github.com/leduckhai/MultiMed-ST
• Github: https://github.com/leduckhai/MultiMed-ST
🔹 Models citing this paper:
• https://huggingface.co/leduckhai/MultiMed-ST
✨ Datasets citing this paper:
• https://huggingface.co/datasets/leduckhai/MultiMed-ST
✨ Spaces citing this paper:
• https://huggingface.co/spaces/HaoVuong/MedicalASR
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SpeechTranslation #MedicalAI #MultilingualNLP #MachineTranslation #Dataset
✨MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
📝 Summary:
MeViS is a multi-modal dataset for referring motion expression video segmentation, addressing the need to segment and track objects based on their motion descriptions. It provides text and audio annotations for complex videos, enabling research into motion-guided video understanding.
🔹 Publication Date: Published on Dec 11
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.10945
• PDF: https://arxiv.org/pdf/2512.10945
• Project Page: https://henghuiding.com/MeViS/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoSegmentation #MultiModalAI #ComputerVision #Dataset #MotionUnderstanding
📝 Summary:
MeViS is a multi-modal dataset for referring motion expression video segmentation, addressing the need to segment and track objects based on their motion descriptions. It provides text and audio annotations for complex videos, enabling research into motion-guided video understanding.
🔹 Publication Date: Published on Dec 11
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.10945
• PDF: https://arxiv.org/pdf/2512.10945
• Project Page: https://henghuiding.com/MeViS/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoSegmentation #MultiModalAI #ComputerVision #Dataset #MotionUnderstanding
❤2
✨SecureCode v2.0: A Production-Grade Dataset for Training Security-Aware Code Generation Models
📝 Summary:
SecureCode v2.0 is a production-grade dataset of 1215 security-focused coding examples. It trains AI models to generate secure code by providing real-incident examples with vulnerable and secure implementations, attacks, defense, and operational security context across 11 languages, using a conve...
🔹 Publication Date: Published on Dec 20
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.18542
• PDF: https://arxiv.org/pdf/2512.18542
• Project Page: https://perfecxion.ai/
• Github: https://github.com/scthornton/securecode-v2
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Cybersecurity #CodeSecurity #AI #CodeGeneration #Dataset
📝 Summary:
SecureCode v2.0 is a production-grade dataset of 1215 security-focused coding examples. It trains AI models to generate secure code by providing real-incident examples with vulnerable and secure implementations, attacks, defense, and operational security context across 11 languages, using a conve...
🔹 Publication Date: Published on Dec 20
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.18542
• PDF: https://arxiv.org/pdf/2512.18542
• Project Page: https://perfecxion.ai/
• Github: https://github.com/scthornton/securecode-v2
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Cybersecurity #CodeSecurity #AI #CodeGeneration #Dataset
✨RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation
📝 Summary:
This work presents an automated rubric generation framework and RubricHub dataset for open-ended AI generation. RubricHub enables significant performance gains, achieving state-of-the-art results on HealthBench and surpassing GPT-5.
🔹 Publication Date: Published on Jan 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.08430
• PDF: https://arxiv.org/pdf/2601.08430
• Project Page: https://huggingface.co/datasets/sojuL/RubricHub_v1
• Github: https://github.com/teqkilla/RubricHub
✨ Datasets citing this paper:
• https://huggingface.co/datasets/sojuL/RubricHub_v1
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#AI #GenerativeAI #MachineLearning #NLP #Dataset
📝 Summary:
This work presents an automated rubric generation framework and RubricHub dataset for open-ended AI generation. RubricHub enables significant performance gains, achieving state-of-the-art results on HealthBench and surpassing GPT-5.
🔹 Publication Date: Published on Jan 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.08430
• PDF: https://arxiv.org/pdf/2601.08430
• Project Page: https://huggingface.co/datasets/sojuL/RubricHub_v1
• Github: https://github.com/teqkilla/RubricHub
✨ Datasets citing this paper:
• https://huggingface.co/datasets/sojuL/RubricHub_v1
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#AI #GenerativeAI #MachineLearning #NLP #Dataset
✨PubMed-OCR: PMC Open Access OCR Annotations
📝 Summary:
PubMed-OCR is a corpus of 209.5K scientific articles from PubMed Central with Google Cloud Vision OCR annotations. It provides word, line, and paragraph bounding boxes to support layout-aware modeling and OCR evaluation. This data is publicly released.
🔹 Publication Date: Published on Jan 16
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.11425
• PDF: https://arxiv.org/pdf/2601.11425
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#OCR #Dataset #ComputerVision #MachineLearning #DataScience
📝 Summary:
PubMed-OCR is a corpus of 209.5K scientific articles from PubMed Central with Google Cloud Vision OCR annotations. It provides word, line, and paragraph bounding boxes to support layout-aware modeling and OCR evaluation. This data is publicly released.
🔹 Publication Date: Published on Jan 16
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.11425
• PDF: https://arxiv.org/pdf/2601.11425
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#OCR #Dataset #ComputerVision #MachineLearning #DataScience
❤3
✨ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
📝 Summary:
ExStrucTiny is a new benchmark dataset for structured information extraction from document images. It addresses limitations of existing datasets by covering diverse document types and flexible schemas. This aims to improve generalist models for structured information extraction.
🔹 Publication Date: Published on Feb 12
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12203
• PDF: https://arxiv.org/pdf/2602.12203
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#InformationExtraction #DocumentAI #MachineLearning #Dataset #ComputerVision
📝 Summary:
ExStrucTiny is a new benchmark dataset for structured information extraction from document images. It addresses limitations of existing datasets by covering diverse document types and flexible schemas. This aims to improve generalist models for structured information extraction.
🔹 Publication Date: Published on Feb 12
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12203
• PDF: https://arxiv.org/pdf/2602.12203
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#InformationExtraction #DocumentAI #MachineLearning #Dataset #ComputerVision
✨StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation
📝 Summary:
StereoAdapter-2 improves underwater stereo depth estimation by replacing ConvGRU with a ConvSS2D operator for efficient, long-range disparity propagation. It also introduces UW-StereoDepth-80K, a new large-scale synthetic dataset. This approach achieves state-of-the-art zero-shot performance on u...
🔹 Publication Date: Published on Feb 18
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.16915
• PDF: https://arxiv.org/pdf/2602.16915
• Project Page: https://aigeeksgroup.github.io/StereoAdapter-2
• Github: https://aigeeksgroup.github.io/StereoAdapter-2
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#UnderwaterAI #ComputerVision #DeepLearning #StereoVision #Dataset
📝 Summary:
StereoAdapter-2 improves underwater stereo depth estimation by replacing ConvGRU with a ConvSS2D operator for efficient, long-range disparity propagation. It also introduces UW-StereoDepth-80K, a new large-scale synthetic dataset. This approach achieves state-of-the-art zero-shot performance on u...
🔹 Publication Date: Published on Feb 18
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.16915
• PDF: https://arxiv.org/pdf/2602.16915
• Project Page: https://aigeeksgroup.github.io/StereoAdapter-2
• Github: https://aigeeksgroup.github.io/StereoAdapter-2
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#UnderwaterAI #ComputerVision #DeepLearning #StereoVision #Dataset
❤1
✨DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain
📝 Summary:
DLT-Corpus is a large new dataset for Distributed Ledger Technology research from scientific literature, patents, and social media. It reveals technologies originate in science before reaching patents and social media. Scientific and patent activity independently drive economic growth, unaffected...
🔹 Publication Date: Published on Feb 25
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.22045
• PDF: https://arxiv.org/pdf/2602.22045
• Github: https://github.com/dlt-science/DLT-Corpus
🔹 Models citing this paper:
• https://huggingface.co/ExponentialScience/LedgerBERT
• https://huggingface.co/ExponentialScience/LedgerBERT-Market-Sentiment
✨ Datasets citing this paper:
• https://huggingface.co/datasets/ExponentialScience/DLT-Patents
• https://huggingface.co/datasets/ExponentialScience/DLT-Tweets
• https://huggingface.co/datasets/ExponentialScience/DLT-Sentiment-News
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#DLT #Dataset #DataScience #Research #Blockchain
📝 Summary:
DLT-Corpus is a large new dataset for Distributed Ledger Technology research from scientific literature, patents, and social media. It reveals technologies originate in science before reaching patents and social media. Scientific and patent activity independently drive economic growth, unaffected...
🔹 Publication Date: Published on Feb 25
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.22045
• PDF: https://arxiv.org/pdf/2602.22045
• Github: https://github.com/dlt-science/DLT-Corpus
🔹 Models citing this paper:
• https://huggingface.co/ExponentialScience/LedgerBERT
• https://huggingface.co/ExponentialScience/LedgerBERT-Market-Sentiment
✨ Datasets citing this paper:
• https://huggingface.co/datasets/ExponentialScience/DLT-Patents
• https://huggingface.co/datasets/ExponentialScience/DLT-Tweets
• https://huggingface.co/datasets/ExponentialScience/DLT-Sentiment-News
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#DLT #Dataset #DataScience #Research #Blockchain
arXiv.org
DLT-Corpus: A Large-Scale Text Collection for the Distributed...
We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents spanning scientific...
✨SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
📝 Summary:
SWE-rebench V2 presents a new language-agnostic automated pipeline to create a large-scale dataset of over 32,000 software engineering tasks across 20 languages and 3,600 repositories. It provides reproducible environments and reliable tests, validated by LLMs, to advance training for SWE agents.
🔹 Publication Date: Published on Feb 27
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.23866
• PDF: https://arxiv.org/pdf/2602.23866
• Github: https://huggingface.co/collections/nebius/swe-rebench-v2
✨ Datasets citing this paper:
• https://huggingface.co/datasets/nebius/SWE-rebench-V2
• https://huggingface.co/datasets/nebius/SWE-rebench-V2-PRs
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SoftwareEngineering #LLMs #AI #Dataset #SWEAgents
📝 Summary:
SWE-rebench V2 presents a new language-agnostic automated pipeline to create a large-scale dataset of over 32,000 software engineering tasks across 20 languages and 3,600 repositories. It provides reproducible environments and reliable tests, validated by LLMs, to advance training for SWE agents.
🔹 Publication Date: Published on Feb 27
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.23866
• PDF: https://arxiv.org/pdf/2602.23866
• Github: https://huggingface.co/collections/nebius/swe-rebench-v2
✨ Datasets citing this paper:
• https://huggingface.co/datasets/nebius/SWE-rebench-V2
• https://huggingface.co/datasets/nebius/SWE-rebench-V2-PRs
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SoftwareEngineering #LLMs #AI #Dataset #SWEAgents
✨Do What I Say: A Spoken Prompt Dataset for Instruction-Following
📝 Summary:
DoWhatISay is a new multilingual dataset of human-recorded spoken and written prompts for evaluating Speech Large Language Models. It reveals text prompts consistently outperform spoken prompts, except in speech-output tasks. This highlights the need for speech-based SLLM evaluation.
🔹 Publication Date: Published on Mar 10
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.09881
• PDF: https://arxiv.org/pdf/2603.09881
• Project Page: https://huggingface.co/collections/meetween/meetweens-research-papers
• Github: https://github.com/MaikeZuefle/DOWIS
✨ Datasets citing this paper:
• https://huggingface.co/datasets/maikezu/dowis
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SLLM #SpeechAI #LLM #PromptEngineering #Dataset
📝 Summary:
DoWhatISay is a new multilingual dataset of human-recorded spoken and written prompts for evaluating Speech Large Language Models. It reveals text prompts consistently outperform spoken prompts, except in speech-output tasks. This highlights the need for speech-based SLLM evaluation.
🔹 Publication Date: Published on Mar 10
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.09881
• PDF: https://arxiv.org/pdf/2603.09881
• Project Page: https://huggingface.co/collections/meetween/meetweens-research-papers
• Github: https://github.com/MaikeZuefle/DOWIS
✨ Datasets citing this paper:
• https://huggingface.co/datasets/maikezu/dowis
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SLLM #SpeechAI #LLM #PromptEngineering #Dataset