UIST 2026 (Accepted)
MAIC-UI: Making Interactive Courseware with Generative UI
About this paper
A system for generating and incrementally editing interactive courseware from textbooks, slides, and PDFs.
THU-MAIC · AI for education
Research on educational agents, language models, and learning environments.
Browse selected work below, or visit Google Scholar ↗ for the full list. Preprints are labelled explicitly.
UIST 2026 (Accepted)
A system for generating and incrementally editing interactive courseware from textbooks, slides, and PDFs.
Journal of Computer Science and Technology (JCST), 41(1): 394–414
A multi-agent approach to interactive online teaching and learning, extending the MOOC model with LLM-driven educational agents.
ACL 2026
ACL 2026
ACL 2026
The Web Conference (WWW) 2026
AIED 2026
AI Open, 7: 142–151
Computers & Education, 239: 105472
An analysis of student–AI interaction patterns in a multi-agent course and their relationships with learning gains and prior knowledge.
Journal of Research on Technology in Education
A study of human–AI collaborative problem solving through the lenses of distributed cognition and regulation of learning.
arXiv preprint, 2026
Models for generating complete slides and interactive learning pages in a single pass.
NAACL 2025
Corresponding author
KDD 2025
Corresponding author
Slide2Lecture turns existing lecture slides into structured teaching actions and interactive tutoring sessions.
CIKM 2025 Demo
Co-first author
A system that converts long, multimodal presentations into pedagogically structured lecture scripts.
CIKM 2025 Demo
Corresponding author
Frontiers of Digital Education, 2(4): 34
Co-first author
An LLM-based knowledge-tracing framework for interpreting student mastery from a small number of learning records.
ACM Multimedia 2025
Corresponding author
Training data and iterative preference optimization for long, faithful outputs from vision-language models.
KDD 2025
Corresponding author
KDD 2025
ACL 2025
ACL 2025
Findings of ACL 2025
学位与研究生教育, 2025年第9期
arXiv preprint, 2025
arXiv preprint, 2025
arXiv preprint, 2025
ICLS 2025
This study explores how students interact with and perceive multiple AI agents in a simulated classroom environment. Findings reveal that students actively engage in both cognitive and regulatory activities, holding generally positive perceptions of AI agents' cognitive support and personalization capabilities. However, limitations in emotional resonance were identified, highlighting the potential of multi-agent systems to enrich educational experiences while underscoring the need for improved emotional engagement design.
EMNLP 2024 Industry
LREC-COLING 2024
Expert Systems with Applications, 239: 122321
KDD 2024
NeurIPS 2024 Datasets and Benchmarks
LREC-COLING 2024
ACL 2024
现代教育技术, 2024年第12期
EMNLP 2024 Demo
We present LM-Interview, a knowledge-guided language model system that automates the full pipeline of semi-structured interviews, including interview guide construction, intelligent dialogue execution, and multimodal data analysis. By adopting a state-action-reward paradigm for fine-grained control of interview flow, the system supports both qualitative and quantitative analysis dimensions. Experiments in real-world scenarios demonstrate that LM-Interview achieves performance comparable to experienced human interviewers, offering an efficient and scalable solution for qualitative research data collection.
ICLR 2024
We construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) Ability Modeling (2) Evolving Data, (3) Standardized Evaluation. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings.
COLING 2024
In this paper, we analyze the causal story behind this problem with counterfactual reasoning methods. Based on the causal effect analysis, we propose a possible solution for alleviating the hallucination in KGD by exploiting the dialogue-knowledge interaction.
ACL 2024
This paper proposes PINOSE, which trains a probing model on offline self-consistency checking results, thereby circumventing the need for human-annotated data and achieving transferability across diverse data distributions.
SIGIR 2023
In this paper, we present MoocRadar, a fine-grained, multi-aspect knowledge repository consisting of 2,513 exercise questions, 5,600 knowledge concepts, and over 12 million behavioral records. Specifically, we propose a framework to guarantee a high-quality and comprehensive annotation of fine-grained concepts and cognitive labels.
ACL 2023
We present a novel three-stage framework DS-MOCE, which leverages the power of pre-trained language models explicitly and implicitly and employs discipline-embedding models with a self-train strategy based on label generation refinement across different domains.
ACL 2023 Demo
We present Visual Knowledge oriented Programming platform (VisKoP), a knowledge base question answering (KBQA) system that integrates human into the loop to edit and debug the knowledge base (KB) queries.
EMNLP 2023
In this paper, based on educational diagnostic assessment method, we conduct an evaluation using MoocRadar, a meticulously annotated human test dataset based on Bloom Taxonomy.
EMNLP 2023
As the large language models (LLMs) have exhibited remarkable in-context learning capabilities, a question arises as to whether the task of OIE can be effectively tackled with this paradigm? In this paper, we explore solving the OIE problem by constructing an appropriate reasoning environment for LLMs.
KDD 2023
We present GLM-Dialog, a large-scale language model (LLM) with 10B parameters capable of knowledge-grounded conversation in Chinese using a search engine to access the Internet knowledge. GLM-Dialog offers a series of applicable techniques for exploiting various external knowledge including both helpful and noisy knowledge, enabling the creation of robust knowledge-grounded dialogue LLMs with limited proper datasets.
CIKM 2023
In this paper, we present a virtual MOOC teaching assistant, LittleMu with minimum labeled training data, to provide question answering and chit-chat services.
CIKM 2023
In this paper, we present GOAL, a benchmark of over 8.9k soccer video clips, 22k sentences, and 42k knowledge triples for proposing a challenging new task setting as Knowledge-grounded Video Captioning (KGVC).
NeuraIPS 2023 Benchmarking Track
(1) We instruct the LM examiner to generate questions across a multitude of domains to probe for a broad acquisition, and raise follow-up questions to engage in a more in-depth assessment. (2) Upon evaluation, the examiner combines both scoring and ranking measurements, providing a reliable result as it aligns closely with human annotations. (3) We additionally propose a decentralized Peer-examination method to address the biases in a single examiner.
Findings of ACL 2023
We propose an entity linking dataset NEL focuses on the NIL prediction problem. NEL takes entities that share an alias with other entities as seeds, collects relevant mention context in the Wikipedia corpus, and ensures the presence of mentions linking to NIL by human annotation and entity masking.
KDD 2022
We propose XDAI, a knowledge-grounded dialogue system that is equipped with the prompt-aware tuning-free PLM exploitation and supported by the ready-to-use open-domain external knowledge resources plus the easy-to-change domain-specific mechanism.
ACL 2022
In this paper, we propose the approach of program transfer, which aims to leverage the valuable program annotations on the rich-resourced KBs as external supervision signals to aid program induction for the low-resourced KBs that lack program annotations.
ACL 2022
This paper proposes a trainable subgraph retriever (SR) decoupled from the subsequent reasoning process, which enables a plug-and-play framework to enhance any subgraph-oriented KBQA model.
ACL 2022
We investigate the usage of entity linking (EL) in downstream tasks and present the first modularized EL toolkit for easy task adaptation. Different from the existing EL methods that deal with all the features simultaneously, we modularize the whole model into separate parts with each feature.
COLING 2022
To extract documents effectively, we construct prompting templates that invoke the underlying knowledge in Pre-trained Language Model (PLM) to calculate the document and keyword’s perplexity, which can assess the document’s semantic salience. Our unsupervised approach can be applied as a plug-in to boost other metrics for evaluating a document’s salience, thus improving the subsequent abstract generation.
CIKM 2022
In this paper, we construct CStory, a large-scale Chinese news storyline dataset, which contains 11, 978 news articles, 112, 549 manually labeled storyline relation pairs, and 49, 832 evidence sentences for annotation judgment. We conduct extensive experiments on CStory using various algorithms and find that constructing news storylines is challenging even for pre-trained language models.
CIKM 2022
In this paper, we propose an interactive contrastive learning model for self-supervised EA. The model encodes not only structures and semantics of entities (including entity name, entity description, and entity neighborhood), but also conducts cross-KG contrastive learning by building pseudo-aligned entity pairs.
Arxiv
To realize the idea of general adaptive systems proposed in pedagogical theory, with the emerging pre-training techniques in NLP, we try to conduct a practical exploration on applying pre-training to adaptive learning, to propose a unified framework based on data observation and learning style analysis, properly leveraging heterogeneous learning elements.
CIKM 2021
We present MOOCCubeX, a large, knowledge-centered repository consisting of 4,216 courses, 230,263 videos, 358,265 exercises, 637,572 fine-grained concepts and over 296 million behavioral data of 3,330,294 students, for supporting the research topics on adaptive learning in MOOCs.
ACL 2021
We propose to decouple the representation learning stage and the decision making stage to fully utilize unlabeled data for entity matching task.
WISE 2021
In this work, we propose a unified crowdsourcing framework to mitigate both challenges. It leverages the skill locality of workers with a Graph Gaussian Process model.
ICPCSEE 2021
In this paper, a learning behavior-aware cognitive diagnosis (LCD) framework is proposed for students’ cognitive modeling with both learning behavior records and exercising records.
ACL 2020
We present MOOCCube, a large-scale data repository of over 700 MOOC courses, 100k concepts, 8 million student behaviors with an external resource. Moreover, we conduct a prerequisite discovery task as an example application to show the potential of MOOCCube in facilitating relevant research.
AACL 2020
We present ExpanRL, an end-to-end hierarchical reinforcement learning (HRL) model for concept expansion in MOOCs. Employing a two-level HRL mechanism of seed selection and concept expansion, ExpanRL is more feasible to adjust the expansion strategy to find new concepts based on the students’ feedback on expansion results.
AP-Web 2020
We propose an Ensemble POI Hierarchical Classification framework (EHC) consisting of three components: Textual and Geographic Feature Extraction, Hierarchical Classifier, and Soft Voting Ensemble Model.
ACL 2019
In this paper, we first build a novel boundary during searching for new concepts via external knowledge base and then utilize heterogeneous features to verify the high-quality results. In addition, to involve human efforts in our model, we design an interactive optimization mechanism based on a game.
CCKS 2018
The existing researches mainly use topics extracted from literatures as objects to build predicting model. To get more accurate results, we use concepts instead of topics constructing a model to predict their rise and fall trends, considering the rhetorical characteristics of them.