Skip to content

AI-Driven System for Automated Generation and Evaluation of Multiple-Choice Questions

Sep 2026 · Communications of International Proceedings · 0 citations

TL;DR

The results showed that the BankGen-generated MCQs are curriculum-related, aligned with learning outcomes, and free from spelling and grammatical errors, indicating that BankGen can reduce the effort required to evaluate and improve MCQs.

Abstract

Student assessment is a fundamental aspect of the learning process in higher education, enabling faculty members to measure learning outcomes and obtain feedback on their teaching. Multiple-choice questions MCQs are widely used in higher education assessments due to their low cost, ease of preparation, accuracy, and ability to measure various learning outcomes. However, the effectiveness of this tool in measuring learning outcomes is directly related to the quality of the MCQs. Student responses are used to calculate both the difficulty index (P) and the discrimination index (D), which helps in evaluating and continuously improving the quality of MCQs. However, the process of creating and writing high-quality MCQs and collecting and analyzing student results is time-consuming and labor-intensive, often resulting in low-quality questions or overlooking feedback calculations. This paper presents BankGen, an AI-powered framework that automates the creation, management, and evaluation of MCQs using generative AI. BankGen is developed using the Replit platform, it integrates Google’s Gemini engine with a user-friendly interface to generate curriculum-aligned and learning-outcome-related MCQs. BankGen supports the creation of question banks, creation of test, and analysis using difficulty and discrimination indicators. It also allows users to delete or modify low-quality MCQs. The results showed that the BankGen-generated MCQs are curriculum-related, aligned with learning outcomes, and free from spelling and grammatical errors. These results indicate that BankGen can reduce the effort required to evaluate and improve MCQs.

View source

Similar papers

Open access Sep 2026

Design and Development of an AI-Assisted Web Learning Platform for Exam Planning, Self-Assessment and Personalized Gap Detection

The experimental evaluation shows an improvement in student performance between the first and last attempt, high structural validity of the generated questions, concise summaries and low response times for data access operations, which support the usefulness of the platform as a complementary tool for active learning,...

Aris-Georgian Ilie, Mihalache Papuc, B. Dobre et al. · 0 citations
Review Open access Sep 2026

Artificial Intelligence in the Classroom: Tools for Personalizing the Educational Experience in Higher Education

Artificial intelligence (AI) is transforming higher education, establishing itself as a catalyst for change that enables the personalization of the educational experience. In contrast to traditional, one-size-fits-all models, AI tailors learning to the individual needs, paces, and styles of students through tools such...

Sandra Rebeca Bajaña Cadena, Juri Evelyn Núñez Portilla, Juliana Karina Zapa Cedeño et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Simple Input-Output Assessment Tasks: Leveraging Automated Programming Assessment for Non-Trivial Courses

The public visibility of Artificial Intelligence (AI) is growing rapidly, driven by the positive impact of its applications across diverse fields of knowledge. In this new chapter, courses that cover the foundations of AI and machine learning become essential for understanding their role and potential in contemporary s...

Artur Jordão · 0 citations
Conference Aug 2026

Dual-Signal Explainability for LLM-Based Automated Grading and Feedback: A Pilot Study

The provision of prompt and customized feedback on student responses in descriptive questions continues to be a difficult problem for educators in virtual and hybrid education systems. The paper introduces the Explainable AI Grading Framework (XAIGF), which is a fusion of two paradigms - semantic evaluation through Lar...

Pratap P. Nayadkar, Shilpa Gite, Swati Sharma et al. · 0 citations
Review Open access Aug 2026

Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation

The system attained a System Usability Scale (SUS) score of 88.5 and cut grading time by 87.5 percent, facilitating focused instructor review in LLM-supported programming assessments, facilitating focused instructor review in LLM-supported programming assessments.

A. Ibrahim, Runal Rezkiawan · 0 citations
Sep 2026

QUESTIFY: An Intelligent Application for Creating, Submitting and Assessing Open-Ended Questions

The assessment of open-ended questions remains a challenging and time-consuming task in higher education, requiring significant teacher effort and often leading to inconsistencies in feedback and evaluation. This paper presents Questify, an intelligent educational platform designed to support the complete lifecycle of...

Ana-Ioana Popescu, Livia Olaru, B. Dobre et al. · 0 citations

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.