SciMKG: A Multimodal Knowledge Graph for Science Education with Text, Image, Video and Audio

Jan 1, 2026·

Tong Lu

Zhichun Wang*

Yaoyu Zhou

Yiming Guan

Zhiyong Bai

Junsheng Du

· 0 min read

PDF Post Poster Website Code Dataset Demo

Abstract

Knowledge graphs (KGs) play a vital role in intelligent education by offering structured representations of educational content. However, constructing multimodal educational knowledge graphs (EKGs) from diverse open educational resources remains a challenge due to the reliance on costly manual annotations and the lack of multimodal integration. In this work, we propose an automated framework that harnesses the reasoning capabilities of large language models (LLMs) to construct multimodal EKGs from open courses efficiently. In our framework, an Extraction-Verification-Integration-Augmentation pipeline is designed to incrementally extract and refine disciplinary concepts from learning resources. Texts, images, videos and audios are aligned with their corresponding concepts. To ensure semantic consistency across modalities, we propose a cross-modal alignment method based on shared structural and semantic features. Using our framework, we build SciMKG, a large-scale multimodal EKG for Chinese K12 education in sciences (biology, physics, and chemistry), encompassing 1,356 knowledge points, 34,630 multimodal concepts, and 403,400 relational triples. Experimental results show that our method improves concept extraction F1 score by 9% over state-of-the-art baselines; both automatic and human evaluations confirm the robustness of our multimodal alignment method. SciMKG and our construction toolkit will be publicly released to support further research and applications in AI-driven education.

Type

Conference paper

Publication

In the Fortieth AAAI Conference on Artificial Intelligence

Last updated on Jan 1, 2026

Multimodal Educational Knowledge Graph Concept Extraction Multimodal Alignment Large Language Models

Authors

Tong Lu

PhD candidate

I am a second year Ph.D. candidate in the School of Artificial Intelligence at Beijing Normal University, Beijing, China. I obtained a Bachelor of Science degree from Hebei GEO University and a Master of Engineering degree from Yunnan University. Now, I engage in research related to the field of natural language processing.

Authors

Zhichun Wang*

Authors

Yaoyu Zhou

Authors

Yiming Guan

Authors

Zhiyong Bai

Authors

Junsheng Du

No results found

SciMKG: A Multimodal Knowledge Graph for Science Education with Text, Image, Video and Audio