AI & ML Papers
Photo
🔥 K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
📅 Published on Jul 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.09635
• PDF: https://arxiv.org/pdf/2605.09635
• Project Page: https://haolpku.github.io/K12-KGraph-page/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/lhpku20010120/K12-KGraph
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#EducationalLanguageModels #CurriculumAlignedKnowledgeGraph #LargeLanguageModels #KnowledgeGraphEmbeddings #BenchmarkingEducationalAI
💡 The paper introduces K12-KGraph, a curriculum-aligned knowledge graph for benchmarking and training educational large language models. The existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented, which is referred to as curriculum cognition. K12-KGraph is extracted from official textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school, containing nine node types and fourteen relation types covering curriculum structure and visual grounding.
From this graph, the authors derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. They also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs.
The results show that existing large language models, such as Gemini-3-Flash and Gemma-4-31B-IT, achieve only 57 percent and 46 percent exact match on K12-Bench, with Prereq and Neighbor being the hardest tasks. The training experiments demonstrate that domain-specific supervision can reduce this gap. Under an unmatched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaoKao Bench and EdEval. For vision-language models, K12-Train-Full achieves the best overall results on GaoKao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full Data Flow and Wizard LM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary.
The authors release the graph, benchmark, training data, and complete construction pipeline, providing a valuable resource for developing and evaluating educational large language models. The paper's contributions include the creation of a curriculum-aligned knowledge graph, a comprehensive benchmark, and a graph-guided supervised fine-tuning corpus, which can help improve the performance of large language models in educational settings.
📅 Published on Jul 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.09635
• PDF: https://arxiv.org/pdf/2605.09635
• Project Page: https://haolpku.github.io/K12-KGraph-page/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/lhpku20010120/K12-KGraph
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#EducationalLanguageModels #CurriculumAlignedKnowledgeGraph #LargeLanguageModels #KnowledgeGraphEmbeddings #BenchmarkingEducationalAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
❤1