Forwarded from AI Scope
With so many LLM papers being published, it's hard to keep up and compare results. This study introduces a semi-automated method that uses LLMs to extract and organize experimental results from arXiv papers into a structured dataset called LLMEvalDB. This process cuts manual effort by over 93%. It reproduces key findings from earlier studies and even uncovers new insights—like how in-context examples help with coding and multimodal tasks, but not so much with math reasoning. The dataset updates automatically, making it easier to track LLM performance over time and analyze trends.
📂 Paper: https://arxiv.org/pdf/2502.18791
▫️@scopeofai
▫️@LLM_learning
📂 Paper: https://arxiv.org/pdf/2502.18791
▫️@scopeofai
▫️@LLM_learning
❤3
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
https://arxiv.org/abs/2602.24286
@LLM_learning
https://arxiv.org/abs/2602.24286
@LLM_learning
arXiv.org
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA...
GPU kernel optimization is fundamental to modern deep learning but remains a highly specialized task requiring deep hardware expertise. Despite strong performance in general programming, large...
❤3