Generative AI answers questions. Agentic AI executes tasks. That is the shift and it deserves a precise name. This is not a faster chatbot. This is a different category of system.
An agent perceives, reasons and acts in a continuous loop. It plans a sequence of steps. It uses external tools. It corrects its own errors. It coordinates with other agents when the task demands it (5,6). The assistant waits for your question. The agent pursues a goal.
The leap that matters
Oncology gives us the first serious validation. Ferber and colleagues built an agent on GPT-4 with precision tools. Vision transformers for mutations in histology. Radiological image segmentation. Search across OncoKB and PubMed. Across twenty multimodal cases, the agent selected the correct tool 87.5 percent of the time and reached the correct clinical conclusion in 91 percent (1).
The number that changes the conversation is another one. GPT-4 alone scored 30.3 percent. The same model, inside an agentic architecture, rose to 87.2 percent (1). The model did not get smarter. The architecture did the work. That is the lesson.
Where it is landing
The evidence concentrates in 2025 and in a few domains. These systems share three traits: planning, autonomous tool use and self-correction (3). They appear in oncology, radiology, neurophysiology and hospital operations. A multi-agent system optimized clinical order sets in a real setting (4). Agents outperform the single-pass model when the task has several steps (7).
The uncomfortable part
The evidence is thin. One scoping review identified seven eligible studies. Only one involved real patients (2). The rest were exploratory exercises. Autonomy without validation is not progress. It is risk transferred to the patient.
And autonomy multiplies error. A model that answers makes a single mistake. An agent that chains ten steps propagates that mistake through the entire chain. Safety, transparency, bias, traceability and interoperability remain unresolved, with a risk of over-reliance and error propagation (6).
The question leaders must ask
It is not whether the agent works in a demo. It works. The question is who answers when an autonomous chain fails. Who validates. Who audits. Who decides.
The operational answer is easy to state and hard to sustain. The agent assists. The human decides. Truhn and colleagues frame it this way: autonomous systems that execute complex tasks with minimal oversight, and that require solving safety challenges before efficient clinical use (5).
What to do now
Build the governance before the agent. Define the accountability circuit. Demand clinical validation with patients, not benchmarks. Measure value, not novelty. The agent that acts needs an institution that governs. Without that you do not deploy intelligence. You deploy automated risk.
References
- Ferber D, El Nahhas OSM, Wölflein G, Wiest I, Clusmann J, Leßmann ME, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat Cancer. 2025;6(8):1337-1349.
- The role of agentic artificial intelligence in healthcare: a scoping review. npj Digit Med. 2026. doi:10.1038/s41746-026-02517-5.
- Artificial intelligence agents in healthcare research: a scoping review. PLoS One. 2026. doi:10.1371/journal.pone.0342182.
- Liu S, Huang SS, McCoy AB, Wright AP, Horst S, Wright A. Optimizing order sets with a large language model-powered multiagent system. JAMA Netw Open. 2025;8(9):e2533277.
- Truhn D, et al. Artificial intelligence agents in cancer research and oncology. Nat Rev Cancer. 2026. doi:10.1038/s41568-025-00900-0.
- Banerjie S, Zhu Y, Freeman I, Villa Machado J, Ahmed A, Sarker A, et al. Agentic AI in healthcare: a comprehensive survey. TechRxiv. 2025. Preprint. doi:10.36227/techrxiv.176238073.31262603/v1.
- AI agents in clinical medicine: a systematic review. medRxiv. 2025. Preprint. doi:10.1101/2025.08.22.25334232.

Leave a Reply