Computational pathology has advanced from narrow machine-learning tools to foundation models trained across organs, diseases and tasks. Yet most systems still process a slide or image in a single pass and return a classification, score or short description. That design overlooks how pathologists actually work: they search across large tissue areas, change magnification, compare multiple slides, combine morphology with immunohistochemistry, molecular tests and clinical information, consider competing diagnoses, and express uncertainty in a carefully structured report. Even multimodal models commonly rely on fixed input combinations and static prediction pipelines. Against these challenges, deeper investigation is needed into agent-based systems that can model pathology as a sequential, evidence-driven and revisable diagnostic process.
Researchers from the Department of Pathology at Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, published (DOI: 10.12290/xhyxzz.2026-0402) the review online on July 7, 2026, in the Medical Journal of Peking Union Medical College Hospital. The article examines how intelligent agents, powered by large language models (LLMs), vision-language models (VLMs) and specialized computational tools, could be organized around three connected stages of clinical pathology: low-magnification slide overview, diagnostic reasoning and report generation, with memory and dynamic revision supporting the workflow across time.
At the overview stage, digital specimens are commonly stored as whole-slide images (WSIs), whose enormous size makes direct analysis difficult. Conventional systems divide WSIs into patches and use methods such as multiple instance learning (MIL) to highlight suspicious regions, but they rarely model the search strategy itself. Navigation agents instead treat slide review as a sequence of decisions, choosing where to move, when to zoom and whether enough information has been collected.
At the diagnostic stage, visual models can extract morphological features while LLMs coordinate higher-level reasoning. Chain-of-thought (CoT) strategies can organize observed features, supporting evidence and exclusion of alternatives, while tool use allows the agent to call dedicated image-analysis, statistical or multimodal modules when needed. Retrieval-augmented generation (RAG) can bring guidelines, textbooks or annotated cases into the reasoning process, particularly for rare or atypical presentations.
For reporting, emerging systems attempt to combine information from several slides rather than generating text from one image. Memory mechanisms may preserve previously observed features or retrieve similar historical cases, while dynamic adaptation can incorporate pathologist feedback. However, incorrect memories, model drift, uncertain evidence tracing, privacy risks and inconsistent outputs remain major barriers.
The review highlights that the potential value of intelligent agents in pathology lies in their ability to integrate information, coordinate analytical tools and support complex diagnostic workflows. Effective systems will need to identify relevant evidence, recognize missing information, select appropriate tools and appropriately represent uncertainty rather than generate unsupported conclusions. However, the authors also emphasized that the field has not yet demonstrated complete, dependable performance across real diagnostic workflows. Clinical usefulness will depend on transparent reasoning, traceable evidence, controlled updating and rigorous evaluation alongside practicing pathologists.
Agent-based pathology could eventually reduce time spent scanning large slides, help organize evidence from multiple specimens and tests, support differential diagnosis, and produce more consistent reports that clearly communicate uncertainty. Such systems may be especially valuable in complex oncology cases, rare diseases and settings where specialist expertise is limited. The review nevertheless frames this as a roadmap, not a clinically validated product: most studies have focused on visual question answering (VQA) or restricted diagnostic tasks rather than end-to-end care. Future work must integrate modules across stages and test stability, reproducibility, data security, auditability and clinical benefit in systematic trials before pathology agents can become trusted components of routine diagnosis.
###
References
DOI
10.12290/xhyxzz.2026-0402
Original Source URL
https://xhyxzz.pumch.cn/article/doi/10.12290/xhyxzz.2026-0402
Funding Information
Capital’s Funds for Health Improvement and Research (2024-2-4012); National High Level Hospital Clinical Research Funding (2025-PUMCH-D-002); Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences (2024-I2M-C&T-C-001); National Key Clinical Specialty Construction Project (U114000); and Beijing Municipal Natural Science Foundation (L252174).
About Medical Journal of Peking Union Medical College Hospital
Medical Journal of Peking Union Medical College Hospital is a leading clinical medicine publication, supported by the multidisciplinary expertise of Peking Union Medical College Hospital. It features the latest research, advancements, and academic trends in clinical and translational medicine, pharmacy, and related interdisciplinary fields, catering to clinicians and medical students across China. The journal aims to promote the exchange of medical knowledge and serve as a high-quality platform for leading academic discussions and fostering scholarly debate in clinical medicine. The journal is listed in China's Core Journals of Science and Technology (CSTPCD), Chinese Science Citation Database (CSCD), A Guide to the Core Journals of China, and the Chinese Biomedical Literature Database (CMCC). Full-text content is accessible on platforms such as Wanfang Data, CNKI, and Chongqing VIP Database. It is indexed in Scopus (Netherlands), the Directory of Open Access Journals (DOAJ) in Sweden, and the Japan Science and Technology Agency Database (JST).