AI & ML Papers
34K subscribers
7.37K photos
593 videos
24 files
8.11K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
AI & ML Papers
Photo
🔥 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

💡 The paper introduces OSReward, a benchmark for evaluating the reliability of vision-language models as judges of computer-using agent trajectories. The trajectories are generated by diverse agents executing human-verified instructions across platforms and are rigorously labeled with ground-truth verdicts through multi-stage human annotation. The benchmark is used to evaluate the performance of state-of-the-art vision-language models, which are found to fall short of ideal judges, sharing a systematic leniency bias that mislabels failed runs as successes.

To address this issue, the authors derive two challenge sets from the OSReward benchmark: OSReward-Hard, which focuses on genuinely hard cases, and OSReward-Multi, which is designed for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of vision-language models to date finds that even state-of-the-art models are not reliable enough to trust, with the few reliable models being too expensive to run at scale, while affordable open models trail far behind.

To close this gap, the authors construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the computer-using agent community. They train OS-Shepherd, an open reward model that supplies low-cost, stable, and reliable reward signals, matching commercial judges at 30-60 percent lower cost than the frontier. Extensive analyses further inform the design of reliable computer-using agent rewards at scale. The code, benchmark, dataset, and model checkpoints are made available to facilitate future research.


📅 Published on Jul 30

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.28609
• PDF: https://arxiv.org/pdf/2607.28609
• Project Page: https://os-copilot.github.io/OSReward-Home/

🤖 Models citing this paper:
• https://huggingface.co/OS-Copilot/OS-Shepherd-9B
• https://huggingface.co/OS-Copilot/OS-Shepherd-35B-A3B

📊 Datasets citing this paper:
• https://huggingface.co/datasets/OS-Copilot/OSReward

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#ComputerUseRewardModels #VisionLanguageModels #CrossPlatformEvaluation #AgentTrajectoryAnalysis #RewardModelBenchmarking
❤1