AI & ML Papers
Photo
π₯ EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
π Published on Mar 9
π Links:
β’ arXiv: https://arxiv.org/abs/2603.08127
β’ PDF: https://arxiv.org/pdf/2603.08127
β’ GitHub: https://github.com/EvoScientist/EvoScientist β 2.6k
ββββββββββββββββββββββββ
π’ By: https://xn--r1a.website/PaperNexus
#MultiAgentSystems #EvolvingAI #ScientificDiscovery #ArtificialIntelligenceResearch #AutonomousScience
π‘ The paper introduces EvoScientist, a multi-agent framework designed to enhance scientific discovery by learning from past interactions. The problem with current AI scientist systems is that they rely on static pipelines and fail to adapt based on accumulated interaction histories, leading to overlooked research directions, repeated failed experiments, and pursuit of infeasible ideas. To address this, EvoScientist uses three specialized agents: a Researcher Agent for idea generation, an Engineer Agent for experiment implementation, and an Evolution Manager Agent that distills insights from prior interactions into reusable knowledge. The framework also includes two persistent memory modules: an ideation memory that summarizes feasible research directions and records unsuccessful ones, and an experimentation memory that captures effective data processing and model training strategies. These modules enable the agents to retrieve relevant prior strategies, improving idea quality and code execution success rates over time. The results show that EvoScientist outperforms seven state-of-the-art systems in scientific idea generation, achieving higher novelty, feasibility, relevance, and clarity, and also improves code execution success rates through multi-agent evolution, demonstrating the effectiveness of persistent memory for end-to-end scientific discovery. Overall, the paper contributes a novel framework that enables AI scientists to learn from their past interactions and adapt their research strategies, leading to more effective and efficient scientific discovery.
π Published on Mar 9
π Links:
β’ arXiv: https://arxiv.org/abs/2603.08127
β’ PDF: https://arxiv.org/pdf/2603.08127
β’ GitHub: https://github.com/EvoScientist/EvoScientist β 2.6k
ββββββββββββββββββββββββ
π’ By: https://xn--r1a.website/PaperNexus
#MultiAgentSystems #EvolvingAI #ScientificDiscovery #ArtificialIntelligenceResearch #AutonomousScience
arXiv.org
EvoScientist: Towards Multi-Agent Evolving AI Scientists for...
The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including...
AI & ML Papers
Photo
π₯ ASI-Bench: At the Dawn of Artificial Superintelligence
π Published on Aug 18
π Links:
β’ GitHub: https://github.com/huggingface
β’ arXiv: https://arxiv.org/abs/2608.17271
β’ PDF: https://arxiv.org/pdf/2608.17271
β’ Project Page: https://asibench.apexin.ai/
π Datasets citing this paper:
β’ https://huggingface.co/datasets/Apexintelligence-AI/ASI-Bench-seed31415
β’ https://huggingface.co/datasets/Apexintelligence-AI/ASI-Bench-seed42
ββββββββββββββββββββββββ
π’ By: https://xn--r1a.website/PaperNexus
#ArtificialSuperintelligence #ASIBench #AIResearchBenchmark #AutonomousScience #ExploratoryAI
π‘ The paper addresses the gap between current artificial intelligence systems, which are largely built to learn, compress and apply existing human knowledge, and the capabilities required for artificial superintelligence, which must explore unknown domains, generate new knowledge and produce verifiable results with little or no human direction. Existing benchmarks only test whether AI can answer questions based on learned knowledge or follow extensive human instructions, leaving the ability to conduct independent scientific research largely unmeasured.
To fill this gap the authors introduce ASIβBench, the first benchmark designed to evaluate AI systems on innovative exploration and autonomous scientific execution across a wide range of research areas. ASIβBench comprises sixty projectβlevel tasks drawn from eleven scientific domains. Each task is built to require the AI to select appropriate methods, conduct the research, and generate results that can be independently verified. The benchmark reduces the amount of methodological guidance provided to the AI in three stages: full logical guidance, only the method name specified, and no guidance at all, forcing the system to determine the method itself. All tasks undergo expert review, AIβassisted auditing, sandboxed execution, and scorer validation to ensure fairness and reproducibility.
The results show a clear decline in performance as guidance is removed. With full logical guidance the average score across 18 stateβofβtheβart model configurations is 50.91. When only the method is specified the average drops to 29.10, and when agents must choose the method themselves the average falls further to 26.62. This sharp decline demonstrates that current AI systems remain heavily dependent on human direction and are far from being able to conduct endβtoβend scientific projects autonomously. The authors make ASIβBench publicly available and invite researchers and developers to contribute new tasks, push the limits of todayβs AI, and accelerate the collective progress toward artificial superintelligence.
π Published on Aug 18
π Links:
β’ GitHub: https://github.com/huggingface
β’ arXiv: https://arxiv.org/abs/2608.17271
β’ PDF: https://arxiv.org/pdf/2608.17271
β’ Project Page: https://asibench.apexin.ai/
π Datasets citing this paper:
β’ https://huggingface.co/datasets/Apexintelligence-AI/ASI-Bench-seed31415
β’ https://huggingface.co/datasets/Apexintelligence-AI/ASI-Bench-seed42
ββββββββββββββββββββββββ
π’ By: https://xn--r1a.website/PaperNexus
#ArtificialSuperintelligence #ASIBench #AIResearchBenchmark #AutonomousScience #ExploratoryAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
π2β€1