AI & ML Papers
Photo
🔥 EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
📅 Published on May 18
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.18703
• PDF: https://arxiv.org/pdf/2605.18703
🤖 Models citing this paper:
• https://huggingface.co/LARK-Lab/EnvFactory-1.7B
• https://huggingface.co/LARK-Lab/EnvFactory-4B
• https://huggingface.co/LARK-Lab/EnvFactory-8B
📊 Datasets citing this paper:
• https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL
• https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED
• https://huggingface.co/datasets/LARK-Lab/EnvFactory-RL
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#ExecutableEnvironments #ToolUseAgents #AgenticReinforcementLearning #RobustRL #LanguageModelTraining
💡 The paper introduces EnvFactory, a framework that automates the creation of executable tool environments and natural multi-turn trajectories for training large language models with agentic reinforcement learning. The problem addressed is that current approaches to equip large language models with tool-use capabilities are limited by the lack of scalable and robust execution environments and the scarcity of realistic training data. Existing methods rely on costly real-world APIs, simulators that are prone to hallucination, or synthetic environments that are often single-turn or based on pre-collected documents.
EnvFactory addresses these challenges by autonomously exploring and verifying stateful, executable tool environments from authentic resources, and synthesizing natural multi-turn trajectories through topology-aware sampling and calibrated refinement. This approach produces grounded queries with implicit intents, which are more effective for reinforcement learning training.
The method involves using a fully automated framework to generate environments and trajectories. The results show that using only 85 verified environments across 7 domains, EnvFactory generates a large number of trajectories, achieving superior training efficiency and downstream performance. The framework improves the performance of Qwen3-series models by up to 15 percent on certain benchmarks, and by up to 8.6 percent and 6 percent on other conversational benchmarks.
The contributions of the paper are that EnvFactory provides a scalable, extensible, and robust foundation for agentic reinforcement learning, and that it achieves superior performance with fewer resources compared to prior work. The framework has the potential to advance the field of large language models and their application to real-world problems. Overall, the paper presents a significant contribution to the field of artificial intelligence and natural language processing.
📅 Published on May 18
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.18703
• PDF: https://arxiv.org/pdf/2605.18703
🤖 Models citing this paper:
• https://huggingface.co/LARK-Lab/EnvFactory-1.7B
• https://huggingface.co/LARK-Lab/EnvFactory-4B
• https://huggingface.co/LARK-Lab/EnvFactory-8B
📊 Datasets citing this paper:
• https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL
• https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED
• https://huggingface.co/datasets/LARK-Lab/EnvFactory-RL
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#ExecutableEnvironments #ToolUseAgents #AgenticReinforcementLearning #RobustRL #LanguageModelTraining
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
📅 Published on Jul 24
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.22529
• PDF: https://arxiv.org/pdf/2607.22529
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#LargeLanguageModels #CoEvolvingSkills #SkillSelfPlay #LanguageModelTraining #ArtificialIntelligenceAdvances
💡 The paper introduces a new framework called Skill Self-Play that aims to improve the capability of large language models through co-evolving skills. The existing self-evolutionary methods face a dilemma between task diversity and verification reliability, where environment-bound methods provide precise feedback but confine learning to narrow domains, while open-ended self-generation broadens the task space but lacks reliable verification. The proposed framework identifies agent skills as a middle ground to reconcile this tension, ensuring deep and verifiable execution in specific scenarios while maintaining open-ended task variety through dynamic routing across skills.
The Skill Self-Play framework consists of a proposer, a solver, and a dynamic skill controller, which co-evolve in a continuous self-play loop orchestrated by a reinforcement learning loop. The proposer generates challenging tasks conditioned on dynamically sampled skills, the solver explores candidate solutions to push its capability boundaries, and the skill controller collects execution feedback to update and expand the skill library.
The empirical evaluations on tool-use and reasoning benchmarks demonstrate that Skill Self-Play effectively bridges the gap between structured verification and open-ended exploration, consistently pushing the performance ceiling of competent backbones while catalyzing striking turnarounds for initially misaligned models. The framework serves as a robust evolution engine, and the code is available for further research and development. Overall, the paper contributes a novel approach to improve the capability of large language models through co-evolving skills, addressing the long-standing dilemma between task diversity and verification reliability in self-evolutionary methods.
📅 Published on Jul 24
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.22529
• PDF: https://arxiv.org/pdf/2607.22529
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#LargeLanguageModels #CoEvolvingSkills #SkillSelfPlay #LanguageModelTraining #ArtificialIntelligenceAdvances
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 SPADE: Self-Play in Adaptive Synthetic Executable Environments
📅 Published on Aug 19
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2608.19197
• PDF: https://arxiv.org/pdf/2608.19197
• Project Page: https://spade-rl.github.io/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#SelfPlayRL #SyntheticEnvironments #LanguageModelTraining #AdaptiveTaskGeneration #ExecutableAI
💡 The paper addresses the need for ever‑expanding, diverse training goals that can keep pace with a language model’s growing abilities. Existing collections of training environments are either hand‑crafted, generated once and frozen, or verified by a static system, so the distribution of tasks does not change as the learner improves. This limits the potential for open‑ended self‑improvement.
SPADE Self‑Play in Adaptive Synthetic Executable Environments proposes a self‑play reinforcement‑learning framework in which a single large language model assumes two complementary roles. The first role, the Environment Designer, writes complete, long‑horizon training environments as executable code that follows an OpenAI‑Gym‑style reset and step interface. The second role, the Reasoning Agent, interacts with those environments, learning to act, reason, and use tools over multiple steps. Both roles are stateful and involve multi‑turn interactions, allowing the same interface to cover pure reasoning problems as well as tool‑use scenarios.
A key innovation is the use of a regret‑based signal to guide environment creation. The Reasoning Agent’s regret is estimated as the difference between the reward it obtains when it receives privileged hints and the reward it obtains without those hints. The Environment Designer is trained to maximize this regret, thereby generating environments that sit at the edge of the agent’s current capabilities while remaining solvable. The authors find that two components are critical for success: grounding the Designer on documents sampled from a large pre‑training corpus, and providing the Designer with an accumulated memory of previously created environments so it can build on past experience.
Experiments scale the framework up to 30‑billion‑parameter models and compare against the strongest fixed‑environment baselines across eight held‑out benchmarks covering mathematics, science, code, and general reasoning. SPADE yields an average improvement of 5.3 points. In tool‑use settings it raises performance by 5.7 points on the BFCL‑v4 multi‑turn benchmark and by 13.9 points on ACEBench‑Agent. In game‑like environments the performance gap over baselines grows with model size. The results demonstrate that making environment design a learnable component enables a concrete step toward open‑ended self‑improvement for language models.
📅 Published on Aug 19
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2608.19197
• PDF: https://arxiv.org/pdf/2608.19197
• Project Page: https://spade-rl.github.io/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#SelfPlayRL #SyntheticEnvironments #LanguageModelTraining #AdaptiveTaskGeneration #ExecutableAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
👍1