AI & ML Papers
Photo
🔥 RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
📅 Published on Jul 13
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.11683
• PDF: https://arxiv.org/pdf/2607.11683
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#GraphRetrievalAugmentedGeneration #DomainAdaptedLLM #MultiStepGraphEngine #KnowledgeGraphConstruction #RetrievalAugmentedGenerationModels
💡 The paper introduces RAGU, a multi-step graph retrieval-augmented generation engine that enhances large language models with structured knowledge. Existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU addresses this by separating extraction from consolidation, using a two-stage typed extraction process, DBSCAN-backed deduplication, LLM summarization, and Leiden community detection.
A key insight is that the skills an in-pipeline LLM needs, such as comprehension, extraction, and reasoning over context, are language skills that grow only weakly with model size, unlike factual world knowledge. Therefore, the authors train MENO-LITE-0.1, a 7B model optimized for language skills, which outperforms Qwen 2.5-32B on knowledge-graph construction and matches it on English GraphRAG tasks.
The results show that RAGU retrieves the most complete context at every factoid level, with evidence recall up to 0.84, and overtakes HippoRAG 2 on synthesis tasks. On multi-hop factoid QA, the apparent HippoRAG 2 advantage is shown to be largely an answer-format artifact. RAGU is installable via pip install graphragu, runs on a single GPU, and is released under MIT. The source code and MENO-LITE-0.1 model are publicly available. Overall, RAGU provides a more effective and efficient approach to graph retrieval-augmented generation, with significant improvements in knowledge-graph construction and question answering tasks.
📅 Published on Jul 13
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.11683
• PDF: https://arxiv.org/pdf/2607.11683
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#GraphRetrievalAugmentedGeneration #DomainAdaptedLLM #MultiStepGraphEngine #KnowledgeGraphConstruction #RetrievalAugmentedGenerationModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
📅 Published on May 26, 2025
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2505.20416
• PDF: https://arxiv.org/pdf/2505.20416
• Project Page: https://huggingface.co/spaces/chenzihong/GraphGen
📊 Datasets citing this paper:
• https://huggingface.co/datasets/chenzihong/GraphGen-Data
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/chenzihong/GraphGen
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#KnowledgeGraphConstruction #SyntheticDataGeneration #LargeLanguageModels #SupervisedFineTuning #LLMSupervision
💡 The paper introduces GraphGen, a framework designed to improve supervised fine-tuning for large language models by generating high-quality synthetic data. The problem addressed is that fine-tuning large language models requires substantial amounts of high-quality supervised data, which is costly and labor-intensive to acquire. Existing synthetic data generation approaches often suffer from factual inaccuracies, insufficient long-tail coverage, simplistic knowledge structures, and homogenized outputs.
To address these challenges, GraphGen constructs a fine-grained knowledge graph from the source text and identifies knowledge gaps in large language models using the expected calibration error metric. It prioritizes the generation of question-answering pairs that target high-value, long-tail knowledge. GraphGen also incorporates multi-hop neighborhood sampling to capture complex relational information and employs style-controlled generation to diversify the resulting question-answering data.
The framework is designed for three key question-answering scenarios: atomic question-answering, aggregated question-answering, and multi-hop question-answering. Experimental results on knowledge-intensive tasks under closed-book settings demonstrate that GraphGen outperforms conventional synthetic data methods, offering a more reliable and comprehensive solution to the data scarcity challenge in supervised fine-tuning. The code and data are publicly available, making it a valuable resource for the research community. Overall, GraphGen provides a novel approach to synthetic data generation, addressing the limitations of existing methods and improving the performance of large language models.
📅 Published on May 26, 2025
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2505.20416
• PDF: https://arxiv.org/pdf/2505.20416
• Project Page: https://huggingface.co/spaces/chenzihong/GraphGen
📊 Datasets citing this paper:
• https://huggingface.co/datasets/chenzihong/GraphGen-Data
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/chenzihong/GraphGen
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#KnowledgeGraphConstruction #SyntheticDataGeneration #LargeLanguageModels #SupervisedFineTuning #LLMSupervision
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.