Add BrightSurf on Google Email

Can an AI knowledge base teach medicine? A VTE Education study shows strong performance in standardized tasks, but creative test writing still needs teacher oversight

07.25.26 | Society of China University Journals
GoPro HERO13 Black

GoPro HERO13 Black records stabilized 5.3K video for instrument deployments, field notes, and outreach, even in harsh weather and underwater conditions.

Can an AI knowledge base teach medicine?

The answer depends on what we mean by “teach.” If the question is whether a large language model can produce fluent medical explanations, the answer may seem straightforward. But medical education is not only about generating answers. It also requires reliable evidence, clinical context, judgment, supervision, and the ability to design learning tasks that help students think like clinicians.

A recent study published in the Chinese Journal of Medical Education Research offers a more measured answer. The research team constructed an intelligent knowledge base for venous thromboembolism, or VTE, using retrieval-augmented generation, and evaluated its application in medical education. The findings suggest that a specialty AI knowledge base can be useful for standardized learning tasks, but it also has clear boundaries—especially when the task requires creativity, educational measurement, and teacher judgment.

VTE is a meaningful test case for this kind of tool. It is a major preventable cause of in-hospital death and appears across many clinical settings, including surgery, oncology, obstetrics, intensive care, and perioperative care. Its diagnosis and treatment are not governed by a single simple rule. Clinicians often need to balance thrombosis risk against bleeding risk, interpret updated guidelines, and make decisions in highly specific patient contexts.

That makes VTE difficult to teach. The knowledge base is broad, evidence changes quickly, and learners may struggle to integrate guidelines, expert consensus, clinical studies, and real-world cases into practical reasoning. Traditional classroom teaching may not be able to cover this complexity within limited time. General-purpose large language models can generate quick responses, but in serious medical education settings, hallucinations, unclear citations, and weak traceability remain major concerns.

To address this problem, the research team used retrieval-augmented generation, or RAG. In simple terms, RAG does not ask the model to answer only from its internal training. Instead, the system first retrieves relevant information from a curated local knowledge base, then generates an answer based on those retrieved materials. The goal is to anchor AI output in selected medical evidence rather than allowing the model to respond freely.

The VTE knowledge base was built with a three-layer corpus. The first layer consisted of evidence-based materials, including clinical guidelines, expert consensus documents, high-quality meta-analyses, randomized controlled trials, and core journal articles. The second layer added expert experience, such as teaching slides, typical consultation scenarios, summaries of common clinical pitfalls, and educational materials accumulated by vascular surgery specialists. The third layer included de-identified real-world cases, organized into structured files for case-based teaching and contextual learning.

The team also added evidence grading, version control, manual tagging, citation traceability, and human–AI collaborative review. These steps are particularly important in medical education, where a convincing but unsupported answer can be more dangerous than no answer at all.

By January 18, 2026, the knowledge base had included 688 core documents, including 151 structured de-identified real-world cases. Platform-side usage data showed 5,424 visits and 915 deep question-answer interactions. These figures suggest that the system was not only a conceptual prototype, but had already attracted real user activity.

The study evaluated the tool through two pathways: one for learners and one for teachers.

The learner pathway included 131 valid questionnaires from medical students and non-specialist healthcare workers. Participants evaluated the knowledge base in three standardized scenarios: clinical decision support, knowledge question answering, and learning resource recommendation. The results were generally positive. Median satisfaction scores across evaluation dimensions were all 4.00 on a five-point scale.

Two areas were especially well received. The Top-box proportion, meaning the percentage of ratings at 4 or 5, was 83.97% for continued-use value and 82.44% for understanding support. In other words, learners felt that the knowledge base helped them approach a complex topic more efficiently and supported their understanding of VTE-related clinical knowledge.

This is one of the main educational values of a specialty AI knowledge base. It does not simply provide an answer; it can reduce the burden of searching, sorting, and integrating scattered information. For a topic like VTE, that may allow learners to spend more cognitive effort on clinical reasoning rather than on locating materials.

At the same time, the results also showed caution. The Top-box proportion for evidence-based credibility was 74.05%, lower than several other dimensions. This does not necessarily mean the knowledge base was unreliable. Rather, it may reflect a healthy skepticism among medical learners. In clinical topics such as anticoagulation, thrombolysis, perioperative management, and care for special populations, learners want to know not only what the answer is, but where it comes from, which guideline or study supports it, and under what conditions it applies.

This is what separates a medical education AI tool from a general chatbot. In medicine, an answer cannot merely sound right. It must be traceable, verifiable, and open to further professional judgment.

The teacher pathway revealed an even more important point: AI knowledge bases have task boundaries.

The study included 19 vascular surgery teachers and used a single-blind randomized controlled design. Teachers evaluated AI-generated outputs for three types of teaching preparation tasks: lesson planning, difficult case discussion, and assisted exam item generation. The system was tested under different levels of use depth, from basic use to more intensive knowledge base use.

For relatively convergent tasks such as lesson planning and difficult case discussion, deeper use of the knowledge base produced slightly higher scores than basic use, although the differences were not statistically significant. This suggests that when the task is clearly defined and mainly requires organizing professional knowledge, a specialty knowledge base can provide useful support.

The result was different for assisted exam item generation. In this more creative task, deeper use of the single VTE knowledge base actually performed worse than basic use, and the difference reached statistical significance.

This finding is important because it challenges a simple assumption: more specialty knowledge does not automatically make AI better at every educational task.

Writing good medical exam questions is not just a matter of turning expert knowledge into questions. It also requires educational measurement, difficulty control, distractor design, alignment with learning objectives, clarity of wording, and the ability to assess whether a question can distinguish between levels of learner understanding. A single VTE clinical knowledge base may contain rich professional content, but it may lack the pedagogical and assessment-related materials needed for high-quality item writing.

In some cases, a strongly constrained RAG system may even narrow the model’s response too much. By forcing the model to rely mainly on clinical documents, the system may reduce the broader reasoning, transfer, and creative flexibility that exam writing requires.

This is perhaps the most useful message of the study: a specialty AI knowledge base is not “more professional, therefore more universal.” It may help learners understand complex medical content and help teachers prepare certain types of teaching materials. But for tasks that require creativity, educational design, value judgment, or assessment expertise, teachers still need to remain in control.

The research team therefore proposes a knowledge base matrix as a future direction. Instead of relying on a single disease-specific knowledge base, medical education AI systems may need multiple coordinated knowledge bases: specialty medical knowledge, teaching resources, research materials, tool and prompt libraries, and updated professional information. For example, when generating VTE exam questions, a system would need not only VTE guidelines and cases, but also curriculum objectives, classic question formats, educational measurement principles, and item-writing standards.

This shifts the discussion from a single AI tool to a structured knowledge system. The model matters, but the knowledge foundation matters just as much. The amount of material matters, but so do its organization, governance, traceability, and suitability for the task.

The study also reminds us that evaluating AI in medical education requires more than user satisfaction. Important questions remain: Does the tool improve learning outcomes? Does it reduce errors? Does it support clinical reasoning? Can it be used in real teaching settings over time? Can its performance be verified across multiple centers, larger samples, and objective educational outcomes?

The authors acknowledge several limitations. This was a single-center preliminary application study with a relatively small sample, especially in the teacher pathway. The evaluation relied mainly on subjective satisfaction and perceived benefit, while objective learning outcomes were limited. The study also did not include long-term follow-up, so the sustained educational impact of the knowledge base remains to be tested.

Even with these limitations, the study offers a practical direction. Medical education does not have to choose between using a general-purpose large language model without constraints and rejecting AI altogether. A more realistic path may be to build curated, governed specialty knowledge bases that allow large language models to work within evidence-based, version-controlled, and human-reviewed environments.

So, can an AI knowledge base teach medicine?

It can help learners enter complex knowledge systems more efficiently. It can support teachers in preparing some structured teaching tasks. It can connect guidelines, expert experience, and real cases in ways that may reduce search and integration burden.

But it cannot replace the teacher’s role in judgment, design, and oversight. In medical education, the value of AI may lie less in automatic teaching and more in organizing knowledge, supporting reasoning, and making evidence easier to access.

The next question is not whether AI knowledge bases should be used in medical education. It is where they work well, where they should be limited, and where human teachers must remain responsible for the final educational decision.

Article information

Title: Construction of an Intelligent Knowledge Base for Venous Thromboembolism and Evaluation of Its Application in Medical Education

Journal: Chinese Journal of Medical Education Research

DOI: 10.3760/cma.j.cn116021-20260128-02296

Chinese Journal of Medical Education Research is a monthly peer-reviewed journal sponsored by the Chinese Medical Association and hosted by Chongqing Medical University, under the supervision of the China Association for Science and Technology.
Launched in 2002, the journal publishes research and practice-oriented studies on medical education, with a focus on teaching reform, clinical teaching, curriculum development, residency training, graduate education, nursing education, educational technology, and international medical education. It also features themed columns and special issues on emerging topics and institutional innovations in medical education.
The journal is recognized as a Source Journal for Chinese Scientific and Technical Papers and Citations and is included in the Chinese Science and Technology Core Journals. It is indexed in Wanfang Data, Index Copernicus, the WHO Western Pacific Region Index Medicus, and Ulrich’s Periodicals Directory.
Journal page: http://yxjyts.alljournals.ac.cn/homeNav?lang=zh

10.3760/cma.j.cn116021-20260128-02296

Survey

People

Construction of an intelligent knowledge base for venous thromboembolism and evaluation of its application in medical education

20-Jun-2026

Keywords

Article Information

Contact Information

Linda Liu
iesResearch
linda@igroup.com.cn
Cong Ma
Society of China University Journals
cujs-office@cujs.org.cn

Source

This article is based on a news release from Society of China University Journals. BrightSurf curates and republishes science news from research institutions worldwide; the original release is linked below.

How to Cite This Article

APA:
Society of China University Journals. (2026, July 25). Can an AI knowledge base teach medicine? A VTE Education study shows strong performance in standardized tasks, but creative test writing still needs teacher oversight. Brightsurf News. https://www.brightsurf.com/news/1WR4VN9L/can-an-ai-knowledge-base-teach-medicine-a-vte-education-study-shows-strong-performance-in-standardized-tasks-but-creative-test-writing-still-needs-teacher-oversight.html
MLA:
"Can an AI knowledge base teach medicine? A VTE Education study shows strong performance in standardized tasks, but creative test writing still needs teacher oversight." Brightsurf News, Jul. 25 2026, https://www.brightsurf.com/news/1WR4VN9L/can-an-ai-knowledge-base-teach-medicine-a-vte-education-study-shows-strong-performance-in-standardized-tasks-but-creative-test-writing-still-needs-teacher-oversight.html.