THU-MAIC · AI for education

Publications

Research on educational agents, language models, and learning environments.

Browse selected work below, or visit Google Scholar ↗ for the full list. Preprints are labelled explicitly.

2026

UIST 2026 (Accepted)

MAIC-UI: Making Interactive Courseware with Generative UI

Shangqing Tu, Yanjia Li, Keyu Chen, Sichen Zhang, Jifan Yu, Daniel Zhang-Li, Lei Hou, Juanzi Li, Yu Zhang, Huiqin Liu

About this paper

A system for generating and incrementally editing interactive courseware from textbooks, slides, and PDFs.

Journal of Computer Science and Technology (JCST), 41(1): 394–414

From MOOC to MAIC: Reimagine Online Teaching and Learning Through LLM-Driven Agents

Jifan Yu, Daniel Zhang-Li, Zheyuan Zhang, Yucheng Wang, Haoxuan Li, Joy Jia Yin Lim, Zhanxin Hao, Shangqing Tu, Lu Zhang, Xusheng Dai, Jianxiao Jiang, Shen Yang, Fei Qin, Zekun Li, Binglin Liu, Xin Cong, Bin Xu, Lei Hou, Manli Li, Juanzi Li, Huiqin Liu, Yu Zhang, Zhiyuan Liu, Maosong Sun

About this paper

A multi-agent approach to interactive online teaching and learning, extending the MOOC model with LLM-driven educational agents.

ACL 2026

SimPBL: A Multi-Agent Framework for Project-Based Learning

Daniel Zhang-Li, Joy Jia Yin Lim, Binglin Liu, Shangqing Tu, Zijun Yao, Hao Peng, Jifan Yu, Haoxuan Li, Zhanxin Hao, Ye He, Zekun Li, Jiangyi Wang, Lei Hou, Bin Xu, Xin Cong, Zhiyuan Liu, Huiqin Liu, Yu Zhang, Juanzi Li

ACL 2026

From Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM Agent

Yucheng Wang, Shen Yang, Jifan Yu, Haoxuan Li, Joy Jia Yin Lim, Daniel Zhang-Li, Huiqin Liu, Lei Hou, Juanzi Li, Bin Xu

ACL 2026

Beyond Self-Report: Bridging the Intention-Behavior Gap in Critical Thinking Assessment via Interpretable Multi-Agent System

Zekun Li, Jifan Yu, Haoxuan Li, Ye He, Daniel Zhang-Li, Shangqing Tu, Joy Jia Yin Lim, Yikun Jiang, Jiaxin Yuan, Yu Zhang

The Web Conference (WWW) 2026

Personalized Learning Path Planning through Goal-Driven Learner State Modeling

Joy Jia Yin Lim, Ye He, Jifan Yu, Xin Cong, Daniel Zhang-Li, Zhiyuan Liu, Huiqin Liu, Lei Hou, Juanzi Li, Bin Xu

AIED 2026

Decoding Student Dialogue: A Multi-Dimensional Comparison and Bias Analysis of Large Language Models as Annotation Tools

Jie Cao, Zhanxin Hao, Jifan Yu

AI Open, 7: 142–151

Which type of students can LLMs act? Investigating authentic simulation with graph-based Human–AI collaborative system

Haoxuan Li, Jifan Yu, Xin Cong, Yang Dang, Daniel Zhang-Li, Lu Mi, Yisi Zhan, Huiqin Liu, Zhiyuan Liu

Computers & Education, 239: 105472

Mapping student-AI interaction dynamics in multi-agent learning environments: Supporting personalized learning and reducing performance gaps

Zhanxin Hao, Jie Cao, Ruimiao Li, Jifan Yu, Zhiyuan Liu, Yu Zhang

About this paper

An analysis of student–AI interaction patterns in a multi-agent course and their relationships with learning gains and prior knowledge.

Journal of Research on Technology in Education

Unpacking interaction profiles and strategies in human-AI collaborative problem solving: a cognitive distribution and regulation perspective

Zhanxin Hao, Xiaobo Liu, Jiaxin Fan, Yun Long, Jifan Yu, Wenli Chen, Yu Zhang

About this paper

A study of human–AI collaborative problem solving through the lenses of distributed cognition and regulation of learning.

arXiv preprint, 2026

CogEvol: Towards Efficient and Reliable Learning Environment Generation

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang

About this paper

Models for generating complete slides and interactive learning pages in a single pass.

2025

NAACL 2025

Simulating Classroom Education with LLM-Empowered Agents

Zheyuan Zhang, Daniel Zhang-Li, Jifan Yu, Linlu Gong, Jinchang Zhou, Zhanxin Hao, Jianxiao Jiang, Jie Cao, Huiqin Liu, Zhiyuan Liu, Lei Hou, Juanzi Li

Corresponding author

KDD 2025

Awaking the Slides: A Tuning-free and Knowledge-regulated AI Tutoring System via Language Model Coordination

Daniel Zhang-Li, Zheyuan Zhang, Jifan Yu, Joy Jia Yin Lim, Shangqing Tu, Linlu Gong, Haohua Wang, Zhiyuan Liu, Huiqin Liu, Lei Hou, Juanzi Li

Corresponding author

About this paper

Slide2Lecture turns existing lecture slides into structured teaching actions and interactive tutoring sessions.

CIKM 2025 Demo

EduCraft: A System for Generating Pedagogical Lecture Scripts from Long-Context Multimodal Presentations

Yucheng Wang, Jifan Yu, Daniel Zhang-Li, Joy Jia Yin Lim, Shangqing Tu, Haoxuan Li, Zhiyuan Liu, Huiqin Liu, Lei Hou, Juanzi Li, Bin Xu

Co-first author

About this paper

A system that converts long, multimodal presentations into pedagogically structured lecture scripts.

CIKM 2025 Demo

VocQuiz: Vocabulary Question Generation for English Language Education

Yongqi Li, Jiajun Wu, Shangqing Tu, Jifan Yu, Huiqin Liu, Lei Hou, Juanzi Li

Corresponding author

Frontiers of Digital Education, 2(4): 34

Explainable Few-Shot Knowledge Tracing

Haoxuan Li, Jifan Yu, Yuanxin Ouyang, Zhuang Liu, Wenge Rong, Huiqin Liu, Juanzi Li, Zhang Xiong

Co-first author

About this paper

An LLM-based knowledge-tracing framework for interpreting student mastery from a small number of learning records.

ACM Multimedia 2025

LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models

Shangqing Tu, Yucheng Wang, Daniel Zhang-Li, Yushi Bai, Jifan Yu, Yuhao Wu, Lei Hou, Huiqin Liu, Zhiyuan Liu, Bin Xu, Juanzi Li

Corresponding author

About this paper

Training data and iterative preference optimization for long, faithful outputs from vision-language models.

KDD 2025

Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models

Shangqing Tu, Zhuoran Pan, Wenxuan Wang, Zhexin Zhang, Yuliang Sun, Jifan Yu, Hongning Wang, Lei Hou, Juanzi Li

Corresponding author

KDD 2025

SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking

Yuanchun Wang, Jifan Yu, Zijun Yao, Jing Zhang, Yuyang Xie, Shangqing Tu, Yiyang Fu, Youhe Feng, Jinkai Zhang, Jingyao Zhang, Bowen Huang, Yuanyao Li, Huihui Yuan, Lei Hou, Juanzi Li, Jie Tang

ACL 2025

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis

Bohan Zhang, Xiaokang Zhang, Jing Zhang, Jifan Yu, Sijia Luo, Jie Tang

ACL 2025

Dynamic Scaling of Unit Tests for Code Reward Modeling

Zeyao Ma, Xiaokang Zhang, Jing Zhang, Jifan Yu, Sijia Luo, Jie Tang

Findings of ACL 2025

TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios

Xiaokang Zhang, Sijia Luo, Bohan Zhang, Zeyao Ma, Jing Zhang, Yang Li, Guanlin Li, Zijun Yao, Kangli Xu, Jinchang Zhou, Daniel Zhang-Li, Jifan Yu, Shu Zhao, Juanzi Li, Jie Tang

学位与研究生教育, 2025年第9期

基于生成式人工智能新特征的研究生教育融合创新探析

于济凡, 刘惠琴

arXiv preprint, 2025

Addressing Situated Teaching Needs: A Multi-Agent Framework for Automated Slide Adaptation

Binglin Liu, Yucheng Wang, Zheyuan Zhang, Jiyuan Lu, Shen Yang, Daniel Zhang-Li, Huiqin Liu, Jifan Yu

arXiv preprint, 2025

Handling Students Dropouts in an LLM-driven Interactive Online Course Using Language Models

Yuanchun Wang, Yiyang Fu, Jifan Yu, Daniel Zhang-Li, Zheyuan Zhang, Joy Lim Jia Yin, Yucheng Wang, Peng Zhou, Jing Zhang, Huiqin Liu

arXiv preprint, 2025

Learning in Context: Personalizing Educational Content with Large Language Models to Enhance Student Learning

Joy Jia Yin Lim, Daniel Zhang-Li, Jifan Yu, Xin Cong, Ye He, Zhiyuan Liu, Huiqin Liu, Lei Hou, Juanzi Li, Bin Xu

ICLS 2025

AI as Learning Partners: Students' Interactions and Perceptions in a Simulated Classroom with Multiple LLM-Powered Agents

Zhanxin Hao, Fei Qin, Jianxiao Jiang, Jie Cao, Jifan Yu, Zhiyuan Liu, Yu Zhang

About this paper

This study explores how students interact with and perceive multiple AI agents in a simulated classroom environment. Findings reveal that students actively engage in both cognitive and regulatory activities, holding generally positive perceptions of AI agents' cognitive support and personalization capabilities. However, limitations in emotional resonance were identified, highlighting the potential of multi-agent systems to enrich educational experiences while underscoring the need for improved emotional engagement design.

2024

EMNLP 2024 Industry

CharacterGLM: Customizing Social Characters with Large Language Models

Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Pei Ke, Guanqun Bi, Libiao Peng, JiaMing Yang, Xiyao Xiao, Sahand Sabour, Xiaohan Zhang, Wenjing Hou, Yijia Zhang, Yuxiao Dong, Hongning Wang, Jie Tang, Minlie Huang

LREC-COLING 2024

Evaluating Generative Language Models in Information Extraction as Subjective Question Correction

Yuchen Fan, Yantao Liu, Zijun Yao, Jifan Yu, Lei Hou, Juanzi Li

Expert Systems with Applications, 239: 122321

Exploring sequence-to-sequence taxonomy expansion via language model probing

Kai Sun, Jifan Yu, Juanzi Li, Lei Hou

KDD 2024

OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining

Fanjin Zhang, Shijie Shi, Yifan Zhu, Bo Chen, Yukuo Cen, Jifan Yu, Yelin Chen, Lulu Wang, Qingfei Zhao, Yuqing Cheng, Tianyi Han, Yuwei An, Dan Zhang, Weng Lam Tam, Kun Cao, Yunhe Pang, Xinyu Guan, Huihui Yuan, Jian Song, Xiaoyan Li, Yuxiao Dong, Jie Tang

NeurIPS 2024 Datasets and Benchmarks

SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, Xiaokang Zhang, Xiaohan Zhang, Sijia Luo, Xi Wang, Jie Tang

LREC-COLING 2024

Untangle the KNOT: Interweaving Conflicting Knowledge and Reasoning Skills in Large Language Models

Yantao Liu, Zijun Yao, Xin Lv, Yuchen Fan, Shulin Cao, Jifan Yu, Lei Hou, Juanzi Li

ACL 2024

WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models

Shangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu, Lei Hou, Juanzi Li

现代教育技术, 2024年第12期

多智能体协同交互的高临场感在线学习环境构建

于济凡, 李睿淼, 李曼丽, 刘惠琴

EMNLP 2024 Demo

LM-Interview: An Easy-to-use Smart Interviewer System via Knowledge-guided Language Model Exploitation

Hanming Li, Jifan Yu, Ruimiao Li, Zhanxin Hao, Xuan Yan, Jiaxin Yuan, Bin Xu, Juanzi Li, Zhiyuan Liu

About this paper

We present LM-Interview, a knowledge-guided language model system that automates the full pipeline of semi-structured interviews, including interview guide construction, intelligent dialogue execution, and multimodal data analysis. By adopting a state-action-reward paradigm for fine-grained control of interview flow, the system supports both qualitative and quantitative analysis dimensions. Experiments in real-world scenarios demonstrate that LM-Interview achieves performance comparable to experienced human interviewers, offering an efficient and scalable solution for qualitative research data collection.

ICLR 2024

KoLA: Carefully Benchmarking World Knowledge of Large Language Models

Jifan Yu*, Xiaozhi Wang*, Shangqing Tu*, Shulin Cao, Daniel Zhang-Li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, Chunyang Li, Zheyuan Zhang, Yushi Bai, Yantao Liu, Amy Xin, Nianyi Lin, Kaifeng Yun, Linlu Gong, Jianhui Chen, Zhili Wu, Yunjia Qi, Weikai Li, Yong Guan, Kaisheng Zeng, Ji Qi, Hailong Jin, Jinxi Liu, Yu Gu, Yuan Yao, Ning Ding, Lei Hou, Zhiyuan Liu, Bin Xu, Jie Tang, Juanzi Li

About this paper

We construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) Ability Modeling (2) Evolving Data, (3) Standardized Evaluation. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings.

COLING 2024

A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation

Jifan Yu, Xiaohan Zhang, Yifan Xu, Xuanyu Lei, Zijun Yao, Jing Zhang, Lei Hou, Juanzi Li

About this paper

In this paper, we analyze the causal story behind this problem with counterfactual reasoning methods. Based on the causal effect analysis, we propose a possible solution for alleviating the hallucination in KGD by exploiting the dialogue-knowledge interaction.

ACL 2024

Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking

Xiaokang Zhang, Zijun Yao, Jing Zhang, Kaifeng Yun, Jifan Yu, Juanzi Li, Jie Tang

About this paper

This paper proposes PINOSE, which trains a probing model on offline self-consistency checking results, thereby circumventing the need for human-annotated data and achieving transferability across diverse data distributions.

2023

SIGIR 2023

MoocRadar: A Fine-grained and Multi-aspect Knowledge Repository for Improving Cognitive Student Modeling in MOOCs

Jifan Yu, Mengying Lu, Qingyang Zhong, Zijun Yao, Shangqing Tu, Zhengshan Liao, Xiaoya Li, Manli Li, Lei Hou, Hai-Tao Zheng, Juanzi Li, Jie Tang

About this paper

In this paper, we present MoocRadar, a fine-grained, multi-aspect knowledge repository consisting of 2,513 exercise questions, 5,600 knowledge concepts, and over 12 million behavioral records. Specifically, we propose a framework to guarantee a high-quality and comprehensive annotation of fine-grained concepts and cognitive labels.

ACL 2023

Distantly Supervised Course Concept Extraction in MOOCs with Academic Discipline

Mengying Lu, Yuquan Wang, Jifan Yu (Corresponding Author), Yexing Du, Lei Hou, Juanzi Li

About this paper

We present a novel three-stage framework DS-MOCE, which leverages the power of pre-trained language models explicitly and implicitly and employs discipline-embedding models with a self-train strategy based on label generation refinement across different domains.

ACL 2023 Demo

VisKoP: Visual Knowledge oriented Programming for Interactive Knowledge Base Question Answering (Best Demo Paper Award)

Zijun Yao, Yuanyong Chen, Xin Lv, Shulin Cao, Amy Xin, Jifan Yu, Hailong Jin, Jianjun Xu, Peng Zhang, Lei Hou, Juanzi Li

About this paper

We present Visual Knowledge oriented Programming platform (VisKoP), a knowledge base question answering (KBQA) system that integrates human into the loop to edit and debug the knowledge base (KB) queries.

EMNLP 2023

Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment Approach

Zheyuan Zhang*, Jifan Yu*, Juanzi Li, Lei Hou

About this paper

In this paper, based on educational diagnostic assessment method, we conduct an evaluation using MoocRadar, a meticulously annotated human test dataset based on Bloom Taxonomy.

EMNLP 2023

Mastering the Task of Open Information Extraction with Large Language Models and Consistent Reasoning Environment (Outstanding Paper)

Ji Qi, Kaixuan Ji, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Lei Hou, Juanzi Li, Bin Xu

About this paper

As the large language models (LLMs) have exhibited remarkable in-context learning capabilities, a question arises as to whether the task of OIE can be effectively tackled with this paradigm? In this paper, we explore solving the OIE problem by constructing an appropriate reasoning environment for LLMs.

KDD 2023

GLM-Dialog: Noise-tolerant Pre-training for Knowledge-grounded Dialogue Generation

Jing Zhang*, Xiaokang Zhang*, Daniel Zhang-Li*, Jifan Yu*, Zijun Yao, Zeyao Ma, Yiqi Xu, Haohua Wang, Xiaohan Zhang, Nianyi Lin, Sunrui Lu, Juanzi Li, Jie Tang

About this paper

We present GLM-Dialog, a large-scale language model (LLM) with 10B parameters capable of knowledge-grounded conversation in Chinese using a search engine to access the Internet knowledge. GLM-Dialog offers a series of applicable techniques for exploiting various external knowledge including both helpful and noisy knowledge, enabling the creation of robust knowledge-grounded dialogue LLMs with limited proper datasets.

CIKM 2023

LittleMu: Deploying an Online Virtual Teaching Assistant via Heterogeneous Sources Integration and Chain of Teach Prompts

Shangqing Tu, Zheyuan Zhang, Jifan Yu, Chunyang Li, Siyu Zhang, Zijun Yao, Lei Hou, Juanzi Li

About this paper

In this paper, we present a virtual MOOC teaching assistant, LittleMu with minimum labeled training data, to provide question answering and chit-chat services.

CIKM 2023

GOAL: A Challenging Knowledge-grounded Video Captioning Benchmark for Real-time Soccer Commentary Generation

Ji Qi*, Jifan Yu*, Teng Tu, Kunyu Gao, Yifan Xu, Xinyu Guan, Xiaozhi Wang, Bin Xu, Lei Hou, Juanzi Li, Jie Tang

About this paper

In this paper, we present GOAL, a benchmark of over 8.9k soccer video clips, 22k sentences, and 42k knowledge triples for proposing a challenging new task setting as Knowledge-grounded Video Captioning (KGVC).

NeuraIPS 2023 Benchmarking Track

Benchmarking Foundation Models with Language-Model-as-an-Examiner

Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, Lei Hou

About this paper

(1) We instruct the LM examiner to generate questions across a multitude of domains to probe for a broad acquisition, and raise follow-up questions to engage in a more in-depth assessment. (2) Upon evaluation, the examiner combines both scoring and ranking measurements, providing a reliable result as it aligns closely with human annotations. (3) We additionally propose a decentralized Peer-examination method to address the biases in a single examiner.

Findings of ACL 2023

Learn to Not Link: Exploring NIL Prediction in Entity Linking

Fangwei Zhu, Jifan Yu*, Hailong Jin, Juanzi Li, Lei Hou, Zhifang Sui

About this paper

We propose an entity linking dataset NEL focuses on the NIL prediction problem. NEL takes entities that share an alias with other entities as seeds, collects relevant mention context in the Wikipedia corpus, and ensures the presence of mentions linking to NIL by human annotation and entity masking.

2022

KDD 2022

XDAI: A Tuning‑free Framework for Exploiting the Pre‑trained Language Models in Knowledge Grounded Dialogue Generation

Jifan Yu, Xiaohan Zhang, Yifan Xu, Xuanyu Lei, Xinyu Guan, Jing Zhang, Lei Hou, Juanzi Li, Jie Tang

About this paper

We propose XDAI, a knowledge-grounded dialogue system that is equipped with the prompt-aware tuning-free PLM exploitation and supported by the ready-to-use open-domain external knowledge resources plus the easy-to-change domain-specific mechanism.

ACL 2022

Program Transfer for Answering Complex Questions over Knowledge Bases

Shulin Cao, Jiaxin Shi, Zijun Yao, Xin Lv, Jifan Yu, Lei Hou, Juanzi Li, Zhiyuan Liu, Jinghui Xiao

About this paper

In this paper, we propose the approach of program transfer, which aims to leverage the valuable program annotations on the rich-resourced KBs as external supervision signals to aid program induction for the low-resourced KBs that lack program annotations.

ACL 2022

Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question Answering

Jing Zhang, Xiaokang Zhang, Jifan Yu, Jian Tang, Jie Tang, Cuiping Li, Hong Chen

About this paper

This paper proposes a trainable subgraph retriever (SR) decoupled from the subsequent reasoning process, which enables a plug-and-play framework to enhance any subgraph-oriented KBQA model.

ACL 2022

HOSMEL: A Hot-Swappable Modularized Entity Linking Toolkit for Chinese

Daniel Zhang-Li, Jing Zhang, Jifan Yu, Xiaokang Zhang, Peng Zhang, Jie Tang, Juanzi Li

About this paper

We investigate the usage of entity linking (EL) in downstream tasks and present the first modularized EL toolkit for easy task adaptation. Different from the existing EL methods that deal with all the features simultaneously, we modularize the whole model into separate parts with each feature.

COLING 2022

UPER: Boosting Multi-Document Summarization with an Unsupervised Prompt-based Extractor

Shangqing Tu, Jifan Yu, Fangwei Zhu, Juanzi Li, Lei Hou and Jian-Yun Nie

About this paper

To extract documents effectively, we construct prompting templates that invoke the underlying knowledge in Pre-trained Language Model (PLM) to calculate the document and keyword’s perplexity, which can assess the document’s semantic salience. Our unsupervised approach can be applied as a plug-in to boost other metrics for evaluating a document’s salience, thus improving the subsequent abstract generation.

CIKM 2022

CStory: A Chinese Large-scale News Storyline Dataset

Kaijie Shi, Xiaozhi Wang, Jifan Yu, Lei Hou, Juanzi Li, Jingtong Wu, Dingyu Yong, Jinghui Xiao, Qun Liu

About this paper

In this paper, we construct CStory, a large-scale Chinese news storyline dataset, which contains 11, 978 news articles, 112, 549 manually labeled storyline relation pairs, and 49, 832 evidence sentences for annotation judgment. We conduct extensive experiments on CStory using various algorithms and find that constructing news storylines is challenging even for pre-trained language models.

CIKM 2022

ICLEA: Interactive Contrastive Learning for Self-supervised Entity Alignment

Kaisheng Zeng, Zhenhao Dong, Lei Hou, Yixin Cao, Minghao Hu, Jifan Yu, Xin Lv, Juanzi Li, Ling Feng

About this paper

In this paper, we propose an interactive contrastive learning model for self-supervised EA. The model encodes not only structures and semantics of entities (including entity name, entity description, and entity neighborhood), but also conducts cross-KG contrastive learning by building pseudo-aligned entity pairs.

Arxiv

Towards a General Pre-training Framework for Adaptive Learning in MOOCs

Qingyang Zhong*, Jifan Yu*, Zheyuan Zhang, Yiming Mao, Yuquan Wang, Yankai Lin, Lei Hou, Juanzi Li, Jie Tang

About this paper

To realize the idea of general adaptive systems proposed in pedagogical theory, with the emerging pre-training techniques in NLP, we try to conduct a practical exploration on applying pre-training to adaptive learning, to propose a unified framework based on data observation and learning style analysis, properly leveraging heterogeneous learning elements.

2021

CIKM 2021

MOOCCubeX: A Large Knowledge-centered Repository for Adaptive Learning in MOOCs (Best Resource Paper Nomination)

Jifan Yu, Yuquan Wang, Qingyang Zhong, Gan Luo, Yiming Mao, Kai Sun, Wenzheng Feng, Wei Xu, Shulin Cao, Kaisheng Zeng, Zijun Yao, Lei Hou, Yankai Lin, Peng Li, Jie Zhou, Bin Xu, Juanzi Li, Jie Tang, Maosong Sun

About this paper

We present MOOCCubeX, a large, knowledge-centered repository consisting of 4,216 courses, 230,263 videos, 358,265 exercises, 637,572 fine-grained concepts and over 296 million behavioral data of 3,330,294 students, for supporting the research topics on adaptive learning in MOOCs.

ACL 2021

Interpretable and Low-Resource Entity Matching via Decoupling Feature Learning from Decision Making

Zijun Yao, Chengjiang Li, Tiansi Dong, Xin Lv, Jifan Yu, Lei Hou, Juanzi Li, Yichi Zhang and Zelin Dai

About this paper

We propose to decouple the representation learning stage and the decision making stage to fully utilize unlabeled data for entity matching task.

WISE 2021

Expertise-Aware Crowdsourcing Taxonomy Enrichment

Yuquan Wang, Yanpeng Wang, Yiming Mao, Jifan Yu, Kaisheng Zeng, Lei Hou, Juanzi Li, Jie Tang

About this paper

In this work, we propose a unified crowdsourcing framework to mitigate both challenges. It leverages the skill locality of workers with a Graph Gaussian Process model.

ICPCSEE 2021

Learning Behavior-Aware Cognitive Diagnosis for Online Education Systems

Yiming Mao, Bin Xu, Jifan Yu, Yifan Fang, Jie Yuan, Juanzi Li, Lei Hou

About this paper

In this paper, a learning behavior-aware cognitive diagnosis (LCD) framework is proposed for students’ cognitive modeling with both learning behavior records and exercising records.

2020

ACL 2020

MOOCCube: A Large-scale Data Repository for NLP Applications in MOOCs

Jifan Yu, Gan Luo, Tong Xiao, Qingyang Zhong, Yuquan Wang, Wenzheng Feng, Junyi Luo, Chenyu Wang, Lei Hou, Juanzi Li, Zhiyuan Liu, Jie Tang

About this paper

We present MOOCCube, a large-scale data repository of over 700 MOOC courses, 100k concepts, 8 million student behaviors with an external resource. Moreover, we conduct a prerequisite discovery task as an example application to show the potential of MOOCCube in facilitating relevant research.

AACL 2020

Expanrl: Hierarchical Reinforcement Learning for Course Concept Expansion in MOOCs

Jifan Yu, Chenyu Wang, Gan Luo, Lei Hou, Juanzi Li, Jie Tang, Minlie Huang, Zhiyuan Liu

About this paper

We present ExpanRL, an end-to-end hierarchical reinforcement learning (HRL) model for concept expansion in MOOCs. Employing a two-level HRL mechanism of seed selection and concept expansion, ExpanRL is more feasible to adjust the expansion strategy to find new concepts based on the students’ feedback on expansion results.

AP-Web 2020

Geographical Information Enhanced POI Hierarchical Classification

Shaopeng Liu, Jifan Yu, Juanzi Li, Lei Hou

About this paper

We propose an Ensemble POI Hierarchical Classification framework (EHC) consisting of three components: Textual and Geographic Feature Extraction, Hierarchical Classifier, and Soft Voting Ensemble Model.

2019

ACL 2019

Course Concept Expansion in MOOCs with External Knowledge and Interactive Game

Jifan Yu, Chenyu Wang, Gan Luo, Lei Hou, Juanzi Li, Jie Tang, Zhiyuan Liu

About this paper

In this paper, we first build a novel boundary during searching for new concepts via external knowledge base and then utilize heterogeneous features to verify the high-quality results. In addition, to involve human efforts in our model, we design an interactive optimization mechanism based on a game.

2018

CCKS 2018

Predicting Concept-based Research Trends with Rhetorical Framing

Jifan Yu, Liangming Pan, Juanzi Li, Xiaoping Du

About this paper

The existing researches mainly use topics extracted from literatures as objects to build predicting model. To get more accurate results, we use concepts instead of topics constructing a model to predict their rise and fall trends, considering the rhetorical characteristics of them.