TY - JOUR
T1 - Agentic AI in Healthcare and Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-Based Agents
AU - Vatsal, Shubham
AU - Dubey, Harsh
AU - Singh, Aditi
PY - 2026/1/1
Y1 - 2026/1/1
N2 - Large Language Model (LLM)-based agents that plan, use tools and act has begun to shape healthcare and medicine. Reported studies demonstrate competence on various tasks ranging from EHR analysis and differential diagnosis to treatment planning and research workflows. Yet the literature largely consists of overviews which are either broad surveys or narrow dives into a single capability (e.g., memory, planning, reasoning), leaving healthcare work without a common frame. We address this by reviewing 49 studies using a seven-dimensional taxonomy: Cognitive Capabilities, Knowledge Management, Interaction Patterns, Adaptation & Learning, Safety & Ethics, Framework Typology and Core Tasks & Subtasks with 29 operational sub-dimensions. Using explicit inclusion and exclusion criteria and a labeling rubric (Fully Implemented ✓, Partially Implemented Δ, Not Implemented ✗), we map each study to the taxonomy and report quantitative summaries of capability prevalence and co-occurrence patterns. Our empirical analysis surfaces clear asymmetries. For instance, the External Knowledge Integration sub-dimension under Knowledge Management is commonly realized (~76% ✓) whereas Event-Triggered Activation sub-dimenison under Interaction Patterns is largely absent (~92% ✗) and Drift Detection & Mitigation sub-dimension under Adaptation & Learning is rare (~98% ✗). Architecturally, Multi-Agent Design sub-dimension under Framework Typology is the dominant pattern (~82% ✓) while orchestration layers remain mostly partial. Across Core Tasks & Subtasks, information centric capabilities lead e.g., Medical Question Answering & Decision Support and Benchmarking & Simulation, while action and discovery oriented areas such as Treatment Planning & Prescription still show substantial gaps (~59% ✗). Together, these findings provide an empirical baseline indicating that current agents excel at retrieval-grounded advising but require stronger adaptation and compliance platforms to move from early-stage systems to dependable systems.
AB - Large Language Model (LLM)-based agents that plan, use tools and act has begun to shape healthcare and medicine. Reported studies demonstrate competence on various tasks ranging from EHR analysis and differential diagnosis to treatment planning and research workflows. Yet the literature largely consists of overviews which are either broad surveys or narrow dives into a single capability (e.g., memory, planning, reasoning), leaving healthcare work without a common frame. We address this by reviewing 49 studies using a seven-dimensional taxonomy: Cognitive Capabilities, Knowledge Management, Interaction Patterns, Adaptation & Learning, Safety & Ethics, Framework Typology and Core Tasks & Subtasks with 29 operational sub-dimensions. Using explicit inclusion and exclusion criteria and a labeling rubric (Fully Implemented ✓, Partially Implemented Δ, Not Implemented ✗), we map each study to the taxonomy and report quantitative summaries of capability prevalence and co-occurrence patterns. Our empirical analysis surfaces clear asymmetries. For instance, the External Knowledge Integration sub-dimension under Knowledge Management is commonly realized (~76% ✓) whereas Event-Triggered Activation sub-dimenison under Interaction Patterns is largely absent (~92% ✗) and Drift Detection & Mitigation sub-dimension under Adaptation & Learning is rare (~98% ✗). Architecturally, Multi-Agent Design sub-dimension under Framework Typology is the dominant pattern (~82% ✓) while orchestration layers remain mostly partial. Across Core Tasks & Subtasks, information centric capabilities lead e.g., Medical Question Answering & Decision Support and Benchmarking & Simulation, while action and discovery oriented areas such as Treatment Planning & Prescription still show substantial gaps (~59% ✗). Together, these findings provide an empirical baseline indicating that current agents excel at retrieval-grounded advising but require stronger adaptation and compliance platforms to move from early-stage systems to dependable systems.
KW - Agentic AI
KW - clinical trial
KW - diagnostic reasoning
KW - empirical analysis
KW - healthcare
KW - large language model
KW - medicine
KW - multi-agent
KW - patient interaction
KW - prompt engineering
KW - survey
KW - taxonomy
UR - https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=105026980532&origin=inward
UR - https://www.scopus.com/inward/citedby.uri?partnerID=HzOxMe3b&scp=105026980532&origin=inward
U2 - 10.1109/ACCESS.2026.3651218
DO - 10.1109/ACCESS.2026.3651218
M3 - Review article
SN - 2169-3536
VL - 14
SP - 4840
EP - 4863
JO - IEEE Access
JF - IEEE Access
ER -