✨On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
📝 Summary:
RL-finetuned VLMs are highly vulnerable to misleading text, severely impacting robustness and confidence. RL fine-tuning presents an accuracy-faithfulness trade-off, eroding reasoning reliability despite accuracy gains. This necessitates joint evaluation of correctness, robustness, and reasoning ...
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12506
• PDF: https://arxiv.org/pdf/2602.12506
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VLM #Robustness #ReinforcementLearning #ChainOfThought #AI
📝 Summary:
RL-finetuned VLMs are highly vulnerable to misleading text, severely impacting robustness and confidence. RL fine-tuning presents an accuracy-faithfulness trade-off, eroding reasoning reliability despite accuracy gains. This necessitates joint evaluation of correctness, robustness, and reasoning ...
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12506
• PDF: https://arxiv.org/pdf/2602.12506
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VLM #Robustness #ReinforcementLearning #ChainOfThought #AI
✨SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
📝 Summary:
SQuTR is a new robustness benchmark for spoken query to text retrieval. It uses 37k diverse queries, real speaker profiles, and 17 noise categories to test systems. Experiments show all systems struggle under extreme noise, making robustness a key bottleneck.
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12783
• PDF: https://arxiv.org/pdf/2602.12783
• Github: https://github.com/ttoyekk1a/SQuTR-Spoken-Query-to-Text-Retrieval
✨ Datasets citing this paper:
• https://huggingface.co/datasets/SLLMCommunity/SQuTR
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SQTR #Robustness #NLP #SpeechRecognition #Benchmarking
📝 Summary:
SQuTR is a new robustness benchmark for spoken query to text retrieval. It uses 37k diverse queries, real speaker profiles, and 17 noise categories to test systems. Experiments show all systems struggle under extreme noise, making robustness a key bottleneck.
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12783
• PDF: https://arxiv.org/pdf/2602.12783
• Github: https://github.com/ttoyekk1a/SQuTR-Spoken-Query-to-Text-Retrieval
✨ Datasets citing this paper:
• https://huggingface.co/datasets/SLLMCommunity/SQuTR
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SQTR #Robustness #NLP #SpeechRecognition #Benchmarking
👍1
✨Do World Action Models Generalize Better than VLAs? A Robustness Study
📝 Summary:
World Action Models WAMs show superior robustness in robot action planning compared to Vision-Language-Action VLAs. WAMs achieve higher success rates on benchmarks under various perturbations, benefiting from video-based dynamic prediction.
🔹 Publication Date: Published on Apr 1
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.22078
• PDF: https://arxiv.org/pdf/2603.22078
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #MachineLearning #Robustness #ComputerVision
📝 Summary:
World Action Models WAMs show superior robustness in robot action planning compared to Vision-Language-Action VLAs. WAMs achieve higher success rates on benchmarks under various perturbations, benefiting from video-based dynamic prediction.
🔹 Publication Date: Published on Apr 1
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.22078
• PDF: https://arxiv.org/pdf/2603.22078
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #MachineLearning #Robustness #ComputerVision