AI & ML Papers
34K subscribers
7.37K photos
593 videos
24 files
8.11K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
AI & ML Papers
Photo
🔥 FastContext: Training Efficient Repository Explorer for Coding Agents

💡 The paper introduces FastContext, a dedicated exploration subagent designed to improve the efficiency of repository exploration in large language model coding agents. The problem addressed is that repository exploration is a major bottleneck in coding agents, consuming a substantial token budget and polluting the agent's context with irrelevant code snippets.

The method involves separating repository exploration from code solving using specialized exploration models. FastContext is invoked on demand and issues parallel tool calls to return concise file paths and line ranges as focused context. The exploration models used in FastContext are powered by 4B-30B parameters and are bootstrapped from strong reference-model trajectories. They are then refined with task-grounded rewards for broad first-turn search, multi-turn evidence gathering, and precise citation generation.

The results show that integrating FastContext into a coding agent improves end-to-end resolution rates by up to 5.5 percent while reducing coding-agent token consumption by up to 60 percent, with minimal overhead. The paper demonstrates that repository exploration can be effectively handled by specialized models, separate from the code solving process. The code and data for FastContext are made available, allowing for further research and development in this area. Overall, the paper presents a significant contribution to the field of coding agents and software engineering, providing a more efficient and effective approach to repository exploration.


📅 Published on Jun 12

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.14066
• PDF: https://arxiv.org/pdf/2606.14066
• Project Page: https://huggingface.co/microsoft/FastContext-1.0-4B-SFT

🤖 Models citing this paper:
• https://huggingface.co/microsoft/FastContext-1.0-4B-SFT
• https://huggingface.co/microsoft/FastContext-1.0-4B-RL

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#EfficientRepositoryExploration #CodingAgents #LargeLanguageModels #RepositoryExplorationSubagents #SpecializedExplorationModels
❤1
AI & ML Papers
Photo
🔥 From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

💡 The paper discusses the challenges of continual learning in large language models and how current methods such as retrieval-augmented generation have limitations in mimicking human long-term memory. The authors propose a new framework called HippoRAG 2 which builds upon previous work and enhances it with deeper passage integration and more effective online use of a large language model. This approach improves performance across factual, sense-making, and associative memory tasks, addressing the deterioration in performance seen in previous methods that tried to augment vector embeddings with structures like knowledge graphs. The results show that HippoRAG 2 outperforms standard retrieval-augmented generation comprehensively, achieving a 7 percent improvement in associative memory tasks over the state-of-the-art embedding model, while also exhibiting superior factual knowledge and sense-making memory capabilities. The work contributes to non-parametric continual learning for large language models, paving the way for more effective and human-like memory capabilities in artificial intelligence systems.


📅 Published on Feb 20, 2025

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2502.14802
• PDF: https://arxiv.org/pdf/2502.14802

🤖 Models citing this paper:
• https://huggingface.co/muthuk1/graphrag-inference-hackathon

📊 Datasets citing this paper:
• https://huggingface.co/datasets/osunlp/HippoRAG_2
• https://huggingface.co/datasets/g7haha/HippoRAG_2

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#ContinualLearning #LargeLanguageModels #NonParametricLearning #RetrievalAugmentedGeneration #LongTermMemory
AI & ML Papers
Photo
🔥 Enhancing Financial Sentiment Analysis via Retrieval Augmented Large Language Models

💡 The paper addresses the challenge of financial sentiment analysis, which is crucial for investment decision-making. Traditional natural language processing models are limited by their size and training data, resulting in poor generalization and effectiveness. Large Language Models, despite their superior performance in various NLP tasks, also face challenges in financial sentiment analysis due to the discrepancy between their pre-training objective and the task of predicting sentiment labels. Additionally, the concise nature of financial news often lacks sufficient context, which can compromise the reliability of Large Language Models' sentiment analysis.

To overcome these challenges, the authors propose a retrieval-augmented Large Language Model framework. This framework consists of two modules: an instruction-tuned Large Language Model module that ensures the model behaves as a predictor of sentiment labels, and a retrieval-augmentation module that retrieves additional context from reliable external sources. This approach enables the model to leverage external context to improve its sentiment analysis capabilities.

The authors evaluate their framework against traditional models and other Large Language Models, such as ChatGPT and LLaMA. The results show that their approach achieves a significant performance gain, with improvements in accuracy and F1 score ranging from 15% to 48%. This demonstrates the effectiveness of the proposed retrieval-augmented Large Language Model framework in enhancing financial sentiment analysis. Overall, the paper contributes to the development of more accurate and reliable financial sentiment analysis models, which can inform better investment decisions.


📅 Published on Oct 6, 2023

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2310.04027
• PDF: https://arxiv.org/pdf/2310.04027

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#FinancialSentimentAnalysis #RetrievalAugmentedModels #LargeLanguageModels #NaturalLanguageProcessing #FinancialTextAnalysis
AI & ML Papers
Photo
🔥 Efficient Guided Generation for Large Language Models

💡 The paper presents an efficient method for guiding large language model text generation using regular expressions and context-free grammars. The problem addressed is that guided generation can be impractical due to significant overhead. The authors propose an approach that adds minimal overhead to the token sequence generation process. This method makes guided generation feasible in practice. The approach is implemented in the open source Python library Outlines, providing a practical solution for efficient guided generation. The results indicate that the method is effective, allowing for guided generation with little to no overhead, which is a significant contribution to the field of natural language processing.


📅 Published on Jul 19, 2023

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2307.09702
• PDF: https://arxiv.org/pdf/2307.09702

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#LargeLanguageModels #GuidedTextGeneration #RegularExpressions #ContextFreeGrammars #EfficientGenerationMethods
❤2
AI & ML Papers
Photo
🔥 JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

💡 The paper introduces JetSpec, a speculative decoding framework designed to improve the inference speed and acceptance rates of large language models. The problem addressed is the scaling limitation of speculative decoding, which accelerates autoregressive large language models by drafting multiple tokens and verifying them in parallel. However, increasing the draft budget only improves speed when acceptance remains high and drafting overhead stays low, creating a scaling ceiling.

The proposed JetSpec framework combines efficient forward drafting with causal conditioning to break this ceiling. It trains a causal parallel draft head over fused hidden states from the frozen target model, producing candidate trees whose scores align with the target model's autoregressive factorization. This approach enables JetSpec to convert larger draft budgets into longer accepted prefixes and higher end-to-end speedup.

The method is compared to bidirectional-head and tree-based speculative decoding baselines across various benchmarks, including math, coding, and chat tasks on dense and MoE models. The results show that JetSpec consistently outperforms these baselines, achieving significant speedup on different workloads. Specifically, JetSpec achieves up to 9.64x speedup on math tasks and 4.58x on open-ended conversational workloads, with further latency gains demonstrated through integration with virtual large language models under realistic serving loads.

Overall, the paper contributes a novel speculative decoding framework that breaks the scaling ceiling of prior methods, enabling faster and more efficient large language model inference. The code and models are made available for further research and development.


📅 Published on Jun 25

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.18394
• PDF: https://arxiv.org/pdf/2606.18394
• Project Page: https://jetspec-project.github.io/jetspec-web/

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#SpeculativeDecoding #LargeLanguageModels #AutoregressiveModeling #ParallelTreeDrafting #CausalConditioning
AI & ML Papers
Photo
🔥 Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

💡 The paper introduces Agents-A1, a 35 billion parameter Mixture-of-Experts Agentic Model that achieves performance comparable to trillion-parameter models by scaling the agent horizon instead of the parameters. The problem addressed is how to improve the performance of large language models on long-horizon tasks without increasing the number of parameters. The method used is a three-stage training approach, which includes supervised fine-tuning, domain-level teacher models, and multi-teacher distillation. The model is trained on a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45,000 tokens. The results show that Agents-A1 achieves strong and broad performance on long-horizon agent benchmarks, outperforming or matching the results of 1 trillion parameter models on several tasks, including SEAL-0, IFBench, HiPhO, FrontierScience-Olympiad, and MolBench-Bind. The paper provides a practical path for scaling the horizon using a smaller model that can reach or match the performance of larger models on long-horizon tasks.


📅 Published on Jun 29

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.30616
• PDF: https://arxiv.org/pdf/2606.30616
• Project Page: https://internscience.github.io/Agents-A1/

🤖 Models citing this paper:
• https://huggingface.co/InternScience/Agents-A1
• https://huggingface.co/InternScience/Agents-A1-FP8-dynamic
• https://huggingface.co/Abiray/Agents-A1-Q4_K_M-GGUF

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#MixtureOfExperts #AgenticModels #LongHorizonTasks #LargeLanguageModels #ParameterEfficientTraining
AI & ML Papers
Photo
🔥 VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

💡 The paper introduces VideoChat3, a fully open, efficient, and generalist video-centric multimodal large language model for video understanding. Current open-source models are limited in several ways, struggling to generalize across diverse video types and being computationally demanding, which restricts their efficiency and scalability. Most models are also only partially open, with key components such as training code, strategy, or datasets unavailable, hindering reproducibility and slowing community-driven development.

To address these issues, VideoChat3 advances video understanding through two complementary designs. For efficiency, it introduces the Inflated 3D Vision Transformer and Adaptive Frame Resolution for Streaming Video Perception, enabling efficient spatiotemporal representation and reducing the cost of processing video inputs during training and inference. For effectiveness, it develops a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets, covering general, long-form, and streaming video scenarios, which improves the model's generalization across domains.

By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts, using only 4B parameters and achieving higher efficiency. The paper's contributions include a fully open and efficient video understanding model, a scalable video data synthesis pipeline, and state-of-the-art results on various video understanding benchmarks.


📅 Published on Jul 16

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.14935
• PDF: https://arxiv.org/pdf/2607.14935
• Project Page: https://mcg-nju.github.io/VideoChat3

🤖 Models citing this paper:
• https://huggingface.co/MCG-NJU/VideoChat3-4B
• https://huggingface.co/MCG-NJU/I3D-ViT

📊 Datasets citing this paper:
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-Academic2M
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-OL617k

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#VideoUnderstanding #MultimodalLearning #LargeLanguageModels #VideoCentricAI #EfficientComputerVision
AI & ML Papers
Photo
🔥 Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

💡 The paper introduces a new framework called Skill Self-Play that aims to improve the capability of large language models through co-evolving skills. The existing self-evolutionary methods face a dilemma between task diversity and verification reliability, where environment-bound methods provide precise feedback but confine learning to narrow domains, while open-ended self-generation broadens the task space but lacks reliable verification. The proposed framework identifies agent skills as a middle ground to reconcile this tension, ensuring deep and verifiable execution in specific scenarios while maintaining open-ended task variety through dynamic routing across skills.

The Skill Self-Play framework consists of a proposer, a solver, and a dynamic skill controller, which co-evolve in a continuous self-play loop orchestrated by a reinforcement learning loop. The proposer generates challenging tasks conditioned on dynamically sampled skills, the solver explores candidate solutions to push its capability boundaries, and the skill controller collects execution feedback to update and expand the skill library.

The empirical evaluations on tool-use and reasoning benchmarks demonstrate that Skill Self-Play effectively bridges the gap between structured verification and open-ended exploration, consistently pushing the performance ceiling of competent backbones while catalyzing striking turnarounds for initially misaligned models. The framework serves as a robust evolution engine, and the code is available for further research and development. Overall, the paper contributes a novel approach to improve the capability of large language models through co-evolving skills, addressing the long-standing dilemma between task diversity and verification reliability in self-evolutionary methods.


📅 Published on Jul 24

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.22529
• PDF: https://arxiv.org/pdf/2607.22529

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#LargeLanguageModels #CoEvolvingSkills #SkillSelfPlay #LanguageModelTraining #ArtificialIntelligenceAdvances
AI & ML Papers
Photo
🔥 DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

💡 The paper introduces DataFlex, a unified framework for dynamic data-centric training of large language models. The problem addressed is that existing approaches to data selection, data mixture optimization, and data reweighting are often developed in isolated codebases, making it difficult to reproduce, compare, and integrate them. DataFlex solves this problem by providing a unified framework that supports three major paradigms of dynamic data optimization: sample selection, domain mixture adjustment, and sample reweighting.

The method involves building DataFlex upon the LLaMA-Factory framework, which allows for extensible trainer abstractions and modular components. This enables a drop-in replacement for standard large language model training and unifies key model-dependent operations such as embedding extraction, inference, and gradient computation. DataFlex is also compatible with large-scale settings, including DeepSpeed ZeRO-3.

The results show that DataFlex provides an effective, efficient, and reproducible infrastructure for data-centric dynamic training of large language models. Comprehensive experiments demonstrate that dynamic data selection consistently outperforms static full-data training, and data mixture methods improve both accuracy and perplexity over default proportions. Additionally, DataFlex achieves consistent runtime improvements over original implementations. Overall, the paper contributes a unified framework that enables efficient large-scale deployment of data-centric dynamic training methods for large language models.


📅 Published on Mar 27

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2603.26164
• PDF: https://arxiv.org/pdf/2603.26164
• Project Page: https://opendcai.github.io/DataFlex-Doc/en/

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#LargeLanguageModels #DataCentricTraining #DynamicTrainingMethods #LanguageModelOptimization #DataDrivenAI
AI & ML Papers
Photo
🔥 K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

💡 The paper introduces K12-KGraph, a curriculum-aligned knowledge graph for benchmarking and training educational large language models. The existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented, which is referred to as curriculum cognition. K12-KGraph is extracted from official textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school, containing nine node types and fourteen relation types covering curriculum structure and visual grounding.

From this graph, the authors derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. They also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs.

The results show that existing large language models, such as Gemini-3-Flash and Gemma-4-31B-IT, achieve only 57 percent and 46 percent exact match on K12-Bench, with Prereq and Neighbor being the hardest tasks. The training experiments demonstrate that domain-specific supervision can reduce this gap. Under an unmatched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaoKao Bench and EdEval. For vision-language models, K12-Train-Full achieves the best overall results on GaoKao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full Data Flow and Wizard LM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary.

The authors release the graph, benchmark, training data, and complete construction pipeline, providing a valuable resource for developing and evaluating educational large language models. The paper's contributions include the creation of a curriculum-aligned knowledge graph, a comprehensive benchmark, and a graph-guided supervised fine-tuning corpus, which can help improve the performance of large language models in educational settings.


📅 Published on Jul 23

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.09635
• PDF: https://arxiv.org/pdf/2605.09635
• Project Page: https://haolpku.github.io/K12-KGraph-page/

📊 Datasets citing this paper:
• https://huggingface.co/datasets/lhpku20010120/K12-KGraph

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#EducationalLanguageModels #CurriculumAlignedKnowledgeGraph #LargeLanguageModels #KnowledgeGraphEmbeddings #BenchmarkingEducationalAI
❤1
AI & ML Papers
Photo
🔥 Kimi K3: Open Frontier Intelligence

💡 The paper introduces Kimi K3, a 2.8 trillion parameter mixture of experts model with native vision capabilities and a 1 million token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. The model also incorporates Stable Latent MoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes. These advances yield an approximately 2.5 times improvement in overall scaling efficiency over Kimi K2.

The model achieves frontier level performance across long horizon coding, genetic, knowledge, reasoning, and vision tasks. Although its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT 5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in the study.

The key contributions of the paper include the introduction of Kimi K3, which is supported by infrastructure advances in multiple areas, such as algorithm system co design for KD, perfectly balanced expert parallel training with efficient memory management, million token genetic RL with persistent rollout and sandbox states, and deployment innovations. The full Kimi K3 model weights are released to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

The problem addressed in the paper is the development of a highly efficient and scalable model that can achieve state of the art performance across a wide range of tasks. The method used to address this problem is the introduction of Kimi K3, which incorporates several key innovations, including Kimi Delta Attention, Attention Residuals, and Stable Latent MoE. The results of the study demonstrate the effectiveness of Kimi K3, which achieves frontier level performance across a range of tasks and outperforms other open and proprietary models. Overall, the paper contributes to the development of highly efficient and scalable models that can achieve state of the art performance across a wide range of tasks.


📅 Published on Jul 27

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24653
• PDF: https://arxiv.org/pdf/2607.24653
• Project Page: https://www.kimi.com/blog/kimi-k3

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#MixtureOfExpertsModel #VisionCapabilitiesInAI #LargeLanguageModels #AttentionMechanismsInDeepLearning #ScalingEfficiencyInAI
AI & ML Papers
Photo
🔥 GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation

💡 The paper introduces GraphGen, a framework designed to improve supervised fine-tuning for large language models by generating high-quality synthetic data. The problem addressed is that fine-tuning large language models requires substantial amounts of high-quality supervised data, which is costly and labor-intensive to acquire. Existing synthetic data generation approaches often suffer from factual inaccuracies, insufficient long-tail coverage, simplistic knowledge structures, and homogenized outputs.

To address these challenges, GraphGen constructs a fine-grained knowledge graph from the source text and identifies knowledge gaps in large language models using the expected calibration error metric. It prioritizes the generation of question-answering pairs that target high-value, long-tail knowledge. GraphGen also incorporates multi-hop neighborhood sampling to capture complex relational information and employs style-controlled generation to diversify the resulting question-answering data.

The framework is designed for three key question-answering scenarios: atomic question-answering, aggregated question-answering, and multi-hop question-answering. Experimental results on knowledge-intensive tasks under closed-book settings demonstrate that GraphGen outperforms conventional synthetic data methods, offering a more reliable and comprehensive solution to the data scarcity challenge in supervised fine-tuning. The code and data are publicly available, making it a valuable resource for the research community. Overall, GraphGen provides a novel approach to synthetic data generation, addressing the limitations of existing methods and improving the performance of large language models.


📅 Published on May 26, 2025

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2505.20416
• PDF: https://arxiv.org/pdf/2505.20416
• Project Page: https://huggingface.co/spaces/chenzihong/GraphGen

📊 Datasets citing this paper:
• https://huggingface.co/datasets/chenzihong/GraphGen-Data

🚀 Spaces citing this paper:
• https://huggingface.co/spaces/chenzihong/GraphGen

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#KnowledgeGraphConstruction #SyntheticDataGeneration #LargeLanguageModels #SupervisedFineTuning #LLMSupervision
AI & ML Papers
Photo
🔥 Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation

💡 The paper introduces Decompile-Bench, a large-scale open-source dataset designed to improve the accuracy of large language model-based decompilers. The problem addressed is the lack of a comprehensive benchmark for evaluating decompilation technology, which is necessary for converting low-level binaries into human-readable source code. Previous efforts have relied on limited or synthetic benchmarks that do not accurately represent real-world binary-source mappings.

To address this issue, the authors created Decompile-Bench, which consists of two million binary-source function pairs generated from 100 million collected function pairs, totaling 450GB of binaries compiled from permissively licensed GitHub projects. The dataset is accompanied by a benchmark, Decompile-Bench-Eval, which includes manually crafted binaries from established datasets and compiled GitHub repositories released after 2025 to mitigate data leakage issues.

The results show that fine-tuning large language model-based decompilers with Decompile-Bench leads to a 20% improvement in re-executability rate compared to previous benchmarks. The authors also explored commonly used evaluation metrics to provide a thorough assessment of the studied decompilers. The code and data are publicly available on HuggingFace and GitHub, making it a valuable resource for advancing decompilation technology. Overall, Decompile-Bench provides a significant contribution to the field by offering a large-scale, real-world dataset for evaluating and improving decompilation accuracy.


📅 Published on May 19, 2025

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2505.12668
• PDF: https://arxiv.org/pdf/2505.12668

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#BinaryDecompilation #DecompilationBenchmarks #LargeLanguageModels #BinarySourceMapping #DecompilerEvaluation
❤1
AI & ML Papers
Photo
🔥 Orchard: An Open-Source Agentic Modeling Framework

💡 The paper introduces Orchard, an open source framework for scalable agentic modeling, which aims to transform large language models into autonomous agents capable of solving complex tasks. The problem addressed is that current research in agentic modeling is constrained by infrastructure and training gaps, with many high performing systems relying on proprietary codebases, models, or services. Most open source frameworks focus on orchestration and evaluation rather than scalable agent training.

The method presented is the Orchard framework, which consists of a lightweight environment service called Orchard Env, providing reusable primitives for sandbox lifecycle management across task domains, agent harnesses, and pipeline stages. On top of Orchard Env, three agentic modeling recipes are built: Orchard-SWE for coding agents, Orchard-GUI for vision language computer use agents, and Orchard-Claw for personal assistant agents.

The results show that Orchard achieves state of the art performance among open source models of comparable size. Specifically, Orchard-SWE achieves 64.3% and 67.5% on SWE-bench Verified after applying credit assignment and reinforcement learning. Orchard-GUI achieves 74.1%, 67.0%, and 64.0% success rates on WebVoyager, Online-Mind2Web, and DeepShop, respectively. Orchard-Claw achieves 59.6% pass@3 on Claw-Eval and 73.9% when paired with a stronger ZeroClaw harness.

The contributions of the paper are the introduction of the Orchard framework, which enables reusable agentic data, training recipes, and evaluations across domains, and the demonstration of its effectiveness in achieving state of the art performance in various tasks. The paper shows that a lightweight, open, harness agnostic environment layer can enable scalable agentic modeling, making it a significant contribution to the field of artificial intelligence.


📅 Published on May 14

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.15040
• PDF: https://arxiv.org/pdf/2605.15040

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#AgenticModeling #AutonomousAgents #ScalableAgentTraining #OpenSourceFrameworks #LargeLanguageModels
AI & ML Papers
Photo
🔥 MetaChain: A Fully-Automated and Zero-Code Framework for LLM Agents

💡 The paper introduces MetaChain, a fully automated framework that enables non technical users to create and deploy large language model agents using natural language alone. The current state of agent development frameworks such as LangChain and AutoGen requires extensive technical expertise, limiting their accessibility to only a small percentage of the global population. To address this challenge, MetaChain is designed as an autonomous agent operating system that comprises four key components: agentic system utilities, LLM powered actionable engine, self managing file system, and self play agent customization module. This system allows for efficient and dynamic creation and modification of tools, agents, and workflows without requiring coding skills or manual intervention. The framework is evaluated on the GAIA benchmark and demonstrates superior performance in generalist multi agent tasks, surpassing existing state of the art methods. Additionally, MetaChain shows consistently superior performance in retrieval augmented generation tasks compared to other large language model based solutions. The contributions of MetaChain are significant as it enables non technical users to build and deploy large language model agents, bridging the accessibility gap and opening up new possibilities for the widespread adoption of agent development frameworks.


📅 Published on Feb 9, 2025

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2502.05957
• PDF: https://arxiv.org/pdf/2502.05957

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#LLMAgents #ZeroCodeDevelopment #NaturalLanguageProcessing #AutonomousAgentSystems #LargeLanguageModels
AI & ML Papers
Photo
🔥 MemSFT: Mitigating Alignment Tax with an External Parametric Memory

💡 The paper addresses the problem of adapting large language models to specialized domains, which often results in a significant decrease in performance on general tasks due to catastrophic forgetting. This issue is known as the alignment tax. To mitigate this problem, the authors propose MemSFT, a method that decouples domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval.

Once trained on a specific domain, the memory can be reused across large language models of different sizes. During generation, a learned router dynamically fuses the output distributions of the memory and backbone at each decoding step, allowing domain expertise to be invoked selectively. The authors evaluate MemSFT across biology, geoscience, and law domains using models ranging from 8 billion to 235 billion parameters. The results show that MemSFT consistently improves domain performance with negligible degradation in general performance, whereas full fine-tuning suffers from severe forgetting on general tasks.

Overall, the paper demonstrates a practical approach to decoupling general model capabilities from domain-specific knowledge at the parameter level, thereby equipping large language models with new specialized capabilities without compromising their general capabilities. The proposed method provides a solution to the alignment tax problem, enabling large language models to adapt to specialized domains without sacrificing their performance on general tasks.


📅 Published on Jul 28

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.25614
• PDF: https://arxiv.org/pdf/2607.25614

🤖 Models citing this paper:
• https://huggingface.co/Jiarui-Wang/MemSFT-Qwen3-Bio-Memory-1.7B
• https://huggingface.co/Jiarui-Wang/MemSFT-Qwen3-Bio-Memory-4B
• https://huggingface.co/Jiarui-Wang/MemSFT-Qwen3-Bio-Memory-8B

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#DomainAdaptation #CatastrophicForgetting #ParametricMemory #LargeLanguageModels #AlignmentTax