ABSTRACT
Artificial intelligence (AI) in radiology is often described as a sequence of architectures. This conceptual narrative review instead organizes its evolution around two questions: Which computational constraint was relaxed, and where did the resulting capability enter the radiologic chain from signal formation to recommendation? We identify five analytical epochs: handcrafted computer-aided detection, deep and volumetric learning, a branching stage of global-context modeling and learned reconstruction, multimodal foundation models, and an emerging stage of inference-stage reasoning. Epochs I–IV are characterized primarily by changes in representation or image formation; Epoch V shifts the frontier toward adaptive case-specific inference, for which radiology-specific evidence remains largely preclinical. The framework also generates a clinical prediction: Verification should be matched to the level of participation and its characteristic failure mode. Existing studies provide partial empirical support, but later-stage verification requirements remain framework-derived proposals rather than established standards.
Main points
• Radiologic artificial intelligence (AI) is better understood by the computational bottleneck each generation relaxed than by a simple chronology of model names.
• Radiology is especially informative because clinical work can be analyzed from signal to image, finding, diagnostic concept, and recommendation; different AI systems enter this chain at different levels and times.
• Epoch V changes the inference regime more than image representation, and current radiology-specific evidence for this transition remains early and predominantly preclinical.
• As machine participation becomes less directly inspectable, oversight should shift from a final-glance model toward verification matched to chain position and failure mode.
1. Introduction
Radiology has repeatedly been a prominent clinical setting for artificial intelligence (AI) translation, from computer-aided detection (CAD) through deep learning, learned reconstruction, transformers, and multimodal foundation models. Reasoning-oriented systems now raise a more difficult question: Can AI participate not only in recognition but also in constructing, comparing, and verifying diagnostic hypotheses?
Previous histories have organized medical imaging AI around chronological milestones, technical families, and clinical applications.1, 2 The present framework instead asks which constraint warrants separate historical status and what verification problem follows from the resulting chain position.
The central thesis is that radiologic AI advances when relaxing a computational constraint makes a new class of machine participation tractable. Here, “bottleneck” denotes that constraint; “analytical epoch” denotes a broad historical-computational period; “frontier” denotes the leading boundary of capabilities made tractable during that period; and “branch” denotes a distinct bottleneck-defined path within an epoch. Paradigm is descriptive only, not a formal taxonomic level. Bottleneck identity classifies; chain position predicts the dominant verification problem. Independent bottlenecks may be grouped as branches within one epoch only when neither is a necessary predecessor of the other and forcing a sequence adds no explanatory value; temporal overlap alone is insufficient. Research frontier status does not imply clinical maturity.
The framework has explicit limits. A bottleneck is weakened if its relaxation does not open a materially new capability, if that capability was already routine, or if gains arise mainly from scale, data, or engineering. Older architectures can persist within the framework when new strategies complement or hybridize with them.
Sutton’s “Bitter Lesson” and scaling-law work emphasize general methods, data, and computation; here, the unit of analysis is the constraint whose relaxation changes what a machine can contribute clinically.3, 4
The criterion also has retrospective discriminatory value. Classical handcrafted radiomics retains predefined quantitative descriptors and therefore remains within Epoch I rather than earning separate epoch status.5 The fuller classification is developed in Section 4.
Recent radiology literature has synthesized foundation models, agentic systems, and human–AI collaboration.6-8 Those reviews characterize technological families and implementation relationships; the present review asks why the frontier of machine participation moves and how that changes verification.
Rahmim al.’s9 2026 perspective uses bottleneck language for deployment and clinical-alignment constraints, including multimodal integration, data sharing, trust, deployability, and actionable guidance. Here, computational bottleneck instead denotes what representation or inference becomes technically tractable. Clinical alignment asks whether that capability becomes useful, trusted, deployable, and actionable. Either can remain binding after the other is relaxed; the accounts are complementary, not interchangeable.
Epochs I–IV are characterized primarily by changes in what is represented or how the image itself is formed: Engineered features give way to learned visual patterns, global context, learned reconstruction priors, and multimodal concepts. Epoch V is classified separately because the limiting problem shifts toward how available representations are interrogated during a case; Section 8 specifies the admission threshold.
Inference is not representation free: Case-specific hidden states, retrieved context, scratchpads, and tool outputs may still be generated. The distinction concerns which constraint makes a new class of participation tractable, not whether internal representations ever change (Figure 1).
2. Review approach
The literature search was last updated on September 4, 2026. Searches were iterative and question driven rather than systematic. PubMed, IEEE Xplore, arXiv, primary regulatory sources, and backward citation chains were used for landmark technical papers, representative radiologic translations, governance documents, and recent reasoning/agentic work. Primary technical papers, original radiology studies, and official documents were prioritized; peer-reviewed sources were preferred, with preprints retained where mature evidence was limited. No formal screening counts, meta-analysis, or risk-of-bias assessment were undertaken. Epoch boundaries were adjudicated using the classification rule outlined above; overlap, coexistence, and hybridization were permitted.
3. Why radiology is an informative clinical laboratory
Radiology provided unusually favorable infrastructure for AI development. The transition from film to picture archiving and communication systems converted examinations into durable, transferable, machine-accessible objects, whereas the Digital Imaging and Communications in Medicine standard imposed common structures for images, metadata, acquisition parameters, and study organization.10, 11 Each examination also routinely generated an expert report, creating large image–text corpora later used for weak supervision, alignment, report generation, and foundation-model training.12 Lesions can be detected, segmented, measured, and followed, and diagnostic performance can be tested in reader studies. These properties make many radiologic tasks unusually operationalizable. In one systematic review of 950 Food and Drug Administration (FDA)-authorized AI and machine-learning devices through June 2024, 723 (76%) were radiology devices.13
Infrastructure, however, explains only why translation was comparatively easy. A deeper reason radiology is analytically informative is that its work can be decomposed into transformations from physical measurement to reconstructed image, finding, diagnostic concept, and recommendation. Major modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) also require computational reconstruction from acquired measurements, creating a pre-interpretive intervention point at which AI can modify the evidentiary object before diagnostic interpretation begins. The chain is an analytical scaffold rather than a literal one-way workflow: Clinical context, prior imaging, and provisional interpretation can feed backward across protocol selection, reconstruction, image search, and interpretation. AI systems have entered this scaffold at different levels, from marks on finished images through learned patterns, pre-finalization reconstruction, multimodal reporting, and reasoning or coordinated action.
This correspondence makes radiology more than an early adopter. From the framework we derive a prospective, normative prediction: The dominant verification problem should track chain position more closely than architecture. Systems that alter image formation should be judged primarily by evidence preservation and lesion fidelity; systems that produce semantic assertions by factuality, omission, and contextual consistency; and systems that recommend actions by calibration, provenance, and downstream clinical effects. These are proposed verification priorities, not established universal standards.
The literature provides partial convergent support. At the finding level, handcrafted CAD and deep-learning detection are architecturally distinct yet judged through reader performance and workflow effects.14-18 In the retrospective paired-reader PI-CAI study, an AI system detected clinically significant prostate cancer more accurately than 62 radiologists reading with PI-RADS 2.1, although non-inferiority to standard-of-care multidisciplinary routine practice was not confirmed.19 At image formation, adversarial CT reconstruction and diffusion MRI converge on fidelity, lesion preservation, and perturbation stability.20-22 At the semantic level, generalist multimodal and radiology-specific report generators converge on factuality and omission errors.23-26 These studies compare different architectures at one chain level. The reciprocal prediction—that one architectural lineage should require different verification across levels—remains prospective because no common design has yet tested it across the chain.
The fit is not deterministic: Reports are imperfect labels, archives reflect who was imaged, and measurable tasks can dominate research without being the most clinically important. The chain should therefore be treated as an explanatory scaffold whose predictions require empirical testing, not as a universal workflow law (Table 1).
4. Epoch I: handcrafted features and classical computer-aided detection
The first bottleneck was representational rigidity. Early systems classified only what engineers had formalized in advance: edges, textures, shapes, intensity distributions, or geometric descriptors passed to shallow classifiers. In radiology, this produced CAD for mammography and pulmonary nodules.27 The representational unit was the predefined feature, and the role was narrow: Mark suspicious locations for radiologist review.
CAD nevertheless established an enduring human–machine arrangement: The algorithm functioned as a second set of eyes rather than an interpreter of the examination as a whole. Its limitations exposed the cost of explicit feature engineering. In screening mammography, large real-world studies showed inconsistent or absent improvements and the burden of false-positive marks.14-16 Radiomics illustrates the same boundary: Classical pipelines expanded the inventory of predefined descriptors without changing the representational class or machine role. Medical images contained more variation than engineers could anticipate, making learned representation, rather than more elaborate feature engineering, the next necessary step.
5. Epoch II: deep learning and volumetric representation
Deep convolutional neural networks (CNNs) addressed that limitation through representation learning. AlexNet showed that layered networks could learn increasingly abstract visual features from raw pixels;28 residual connections enabled deeper networks;29 and U-Net and its three-dimensional (3D) variants provided encoder–decoder designs for biomedical segmentation.30, 31 Representation expanded from engineered features to learned patterns and anatomically meaningful objects.
Radiologic translation was rapid; CNNs classified examinations, detected urgent findings, segmented organs and tumors, quantified disease burden, and prioritized worklists. U-Net variants made voxel-level delineation central to imaging AI, whereas nnU-Net demonstrated how a common framework could self-configure across heterogeneous segmentation tasks.32 Systems for intracranial hemorrhage and other critical findings showed how image analysis could alter workflow before the examination was opened.33, 34
Generalization remained unresolved. Supervised models were label dependent and could degrade across scanners, protocols, institutions, prevalence conditions, and populations.35 Self-supervised and label-efficient pretraining partly relaxed annotation dependence by learning reusable representations from unlabeled or weakly labeled data.36 Self-supervised learning nevertheless fails the primary epoch criterion: Annotation dependence constrains the efficiency and data requirements for acquiring an existing visual representational class; it does not make a new representational relation or inferential organization tractable. Unchanged chain position corroborates rather than determines the classification. Self-supervised learning therefore remains a bridge within Epoch II and a scaling condition for later foundation models. Convolutional architectures can still make long-range relationships cumbersome to model; broader context and image formation became the next distinct constraints.
6. Epoch III: Global context and learned reconstruction—parallel branches
The third period is best understood as a branching transition. One branch addressed contextual integration. Transformers used self-attention to model relationships dynamically across a sequence,37 and vision transformers allowed image-region tokens to interact beyond fixed local receptive fields.38 Unlike self-supervised learning, attention changes the representational relation itself: Distant elements can interact directly rather than only through repeated local operations. The UNETR model provides a radiology-specific instantiation: It reformulated 3D medical-image segmentation as sequence-to-sequence prediction because localized convolutional receptive fields limited long-range spatial dependency modeling, using a transformer encoder to capture global multiscale information.39 Medical imaging remains heavily hybridized with convolution because global modeling and fine localization impose different demands; III-A is defined by tractable global context, not CNN displacement.
A parallel branch addressed inverse imaging. Variational networks, adversarial methods, and later score-based diffusion models learned priors for producing clinically usable images from incomplete or degraded measurements.20, 21, 40, 41 In accelerated MRI and low-dose CT, AI entered the chain before interpretation by participating in image formation. Failure could therefore become embedded in the evidence rather than remain a visible mark a reader could dismiss. Learned priors can suppress noise but may remove subtle pathology or introduce unsupported plausible structure, and perturbation instability remains an important concern.22
Epoch III is an umbrella analytical period, not a claim of shared mechanism or simultaneous invention. Both branches occupy a post–Epoch-II research frontier in which learned models extend beyond discriminative analysis of finalized images; III-A makes direct non-local context tractable within interpretation, whereas III-B moves learned modeling upstream into the inverse problem that constructs the evidentiary image, replacing or augmenting fixed reconstruction priors with learned priors. Therefore, III-B is not another Epoch II application merely because it uses deep learning: Its bottleneck is image formation, not representation of an already reconstructed image. The branches remain distinct; their common epoch label is an analytical grouping, not a chronology or causal sequence.
7. Epoch IV: Multimodal foundation models
Foundation models addressed task and modality isolation through large-scale pretraining and shared representation spaces. Visual encoders and language models could align images, text, and other clinical information within a common semantic framework.23 Recent radiology reviews describe the resulting adaptability, multimodal integration, opportunities, and risks in detail.6 Within the present framework, their historical significance is narrower and more specific: They relax semantic and task isolation sufficiently for a model to operate across multiple levels of the radiologic chain.
Multimodal systems now support image interpretation, retrieval, visual question answering, and report generation. Generalist biomedical models demonstrated cross-modality transfer,23 while radiology-specific work advanced clinician–AI report composition, large-scale two-dimensional/3D representation learning, and 3D CT report generation.24-26 Their most plausible current role remains collaborative report and context assistance rather than autonomous interpretation.
This transition moves AI into radiology’s principal professional assertion. A CAD mark is an intermediate prompt; a report can redirect care. Standard next-token training is not directly optimized to minimize clinically consequential omissions, unsupported findings, incorrect negation, or misstated interval change. A report can therefore be fluent while clinically wrong, and multiple phrasings may be equivalent, but one omitted finding can matter greatly. Foundation models qualify as a distinct epoch not because they are larger, but because relaxing semantic and task isolation makes cross-modal interpretation and report-level participation tractable. Their success also exposes the next bottleneck: Broad representation does not guarantee controlled case-specific inference.
8. Epoch V: Inference-stage reasoning and test-time scaling
Epoch V contains mechanisms that operate at different stages. Reasoning-oriented post-training can shape a policy for generating, selecting, or revising responses; reinforcement learning is one route to encouraging reflection and verification behavior.42 Radiology-specific work now provides a closer analog. Fan et al. trained ChestX-Reasoner with supervised fine-tuning and reinforcement learning using process supervision and introduced RadRBench-CXR, comprising 59,000 visual-question-answering samples and 301,000 clinically validated reasoning steps, together with RadRScore, which evaluates reasoning factuality, completeness, and effectiveness. The model improved benchmark reasoning and diagnostic performance, but the evidence remains benchmark-based rather than prospective clinical reasoning.43
Test-time scaling adds per-case computation through longer budgets, multiple candidates, self-consistency, revision, retrieval, or tools. In medical question answering, m1, a text-only test-time scaling framework for large language models (LLMs), found gains only up to an optimal budget; beyond it, overthinking could degrade performance, while insufficient knowledge remained a separate bottleneck.44 Byun et al.45 found gains using multiple candidates and self-consistency across text and medical imaging tasks. Oh et al. found model- and task-dependent benefits, limited gains from simple token expansion in current vision-language models, and vulnerability to misleading expert-physician cues.46 These studies establish case-specific computation as a conditional lever but do not all clear our threshold: m1, Byun,45 and much of Oh mainly test budget allocation, sampling, or revision and are enabling evidence rather than standalone demonstrations.
Epoch V is therefore different from its predecessors in what primarily limits progress. The stronger claim rests on reorganized case-level computation—iterative hypothesis generation, evidence acquisition, verification, or tool-mediated action—not simply more tokens, samples, thresholds, or prompts. ChestX-Reasoner moves closer to this threshold because stepwise verification is part of the learned reasoning policy. Routine test-time augmentation, ensembling, threshold adjustment, or prompt variation may improve outputs but do not qualify on their own. Reasoning-oriented post-training lets policies exploit this organization. This is a research frontier claim, not a claim of equivalent clinical maturity.
Radiology-specific studies occupy different positions relative to this threshold. ChestX-Reasoner is principally training-side evidence: Supervised fine-tuning and reinforcement learning with process supervision train a policy to produce explicit intermediate reasoning, whereas RadRBench-CXR and RadRScore provide reasoning-focused evaluation.43 Likewise, CMR-R is training-side. A multistage framework used quantitative cardiac MR parameters and semantic report descriptions from 16,104 multicenter cases to generate diagnostic chains across eight cardiac disease categories; because its inputs are structured and semantic rather than raw images, it tests semantic diagnostic reasoning rather than end-to-end image reasoning.47 By contrast, MedRAX is a cleaner inference-side example: Without additional training, it dynamically selects specialized chest X-ray tools and a multimodal language model to answer case-specific queries.48 Thus, training-side studies show that structured reasoning policies can be learned, whereas MedRAX illustrates case-specific orchestration. All remain benchmark based or preclinical; none establishes prospective autonomous radiologic reasoning in clinical workflow.
This remains an emerging research direction rather than established autonomous clinical practice. General LLM evidence shows that chain-of-thought explanations can be plausible yet unfaithful, irrelevant retrieved passages can distract generation, and intrinsic self-correction can fail or even degrade reasoning.49-51 These are not radiology-specific validation studies; reasoning over an incorrect perceptual premise is therefore best treated as an anticipated compound failure mode. Latency, cost, provenance, calibration, and auditability remain part of the clinical problem.
9. Cross-epoch synthesis: From representation to clinical participation
Three linked but non-identical trajectories emerge. Representation changes from predefined features to learned patterns and objects, broader context, reconstruction priors, and multimodal concepts. Inference becomes more flexible through hypothesis comparison, retrieval, critique, and verification. The scope of clinical participation extends from detection and measurement through reconstruction and reporting toward reasoning and controlled coordination. These trajectories are related but not synchronous: Representational capability can advance without equivalent inferential reliability or clinical readiness.
The important change is not simply increasing accuracy but changing where machine contribution sits in the radiologic representational chain. Characteristic errors therefore change in both consequence and inspectability. A false-positive CAD mark is usually visible and can often be readily dismissed by the reader. A segmentation error may distort a measurement but remain geometrically reviewable. A reconstruction error can be embedded in the evidence itself. A generated report may introduce semantic errors that require comparison with the images and record. A reasoning system may organize several correct observations into a persuasive but incorrect conclusion. Increasing abstraction does not make every error less visible, but it weakens the assumption that a final human glance will reliably catch failure.
Post-hoc attribution is one cross-cutting response to the inspectability problem, not a separate analytical epoch. Saliency maps and gradient-weighted class activation mapping can expose image regions associated with an output, but they do not independently establish that the attribution faithfully tracks the computation or that the conclusion is correct.52 Radiology studies have shown limited sensitivity and robustness of popular saliency methods, whereas experimental mammography work found that explanation inputs did not reliably overcome radiologists’ tendency to follow incorrect AI suggestions.53, 54 These post-hoc attribution tools can support interrogation, localization, and error discovery, but they cannot substitute for outcome-appropriate testing or independent verification.
The verification requirements that follow are framework-derived proposals, not established universal standards. When direct inspection becomes less reliable, systematic verification should increasingly replace reliance on superficial oversight. Detection tools warrant reader and workflow studies; reconstruction systems warrant signal-fidelity and lesion-preservation testing; report generators warrant factuality and omission analysis; and reasoning or agentic systems warrant calibration, robustness, provenance, tool-use auditing, and assessment of effects on human judgment. Of 723 FDA-authorized radiology devices in the cited review, submission documentation was available for 717; 208 (29%) incorporated clinical testing, 33 (5%) reported prospective testing, and 56 (8%) included human-in-the-loop testing.13 Given the authorization era and task mix, these figures mainly characterize earlier radiology-AI functions and should not be read as prevalence estimates for foundation, reasoning, or agentic systems.
Clinical and implementation evidence illustrates why the tested endpoint matters. ScreenTrustCAD showed non-inferiority of one radiologist plus AI compared with standard double reading for screening mammography.17 The randomized MASAI trial subsequently found a non-inferior interval-cancer rate, higher sensitivity, and unchanged specificity with AI-supported screening, and PRAIM provided large-scale real-world implementation evidence across 12 German sites; because radiologists voluntarily chose AI use in PRAIM, residual selection effects remain possible.55, 56 LungIMPACT provides the complementary negative result: Across 93,326 chest radiographs, AI prioritization accelerated reporting but did not significantly shorten time to CT, diagnosis, referral, treatment, or stage.18 Together, these studies show that evidentiary maturity depends on testing the workflow or patient-level endpoint actually claimed, not on whether the result is positive.
Architecture determines computational possibility, not clinical value. A technical advance may relax a computational bottleneck yet fail to solve the clinically binding one if it cannot survive external testing, workflow friction, procurement, reimbursement, regulation, liability, or professional trust. Research prominence is not evidence of clinical progress. Predetermined change control plans, the European Union AI Act, general health–AI guidance, and World Health Organization guidance specific to large multimodal models increasingly address lifecycle governance.57-60
Human–AI collaboration work emphasizes complementarity, automation bias, oversight, and workflow effects, supporting the view that AI should currently function primarily as a clinical teammate rather than an autonomous replacement for physicians.8, 61 Our framework adds a location-specific question: At which chain level does the machine contribute, and what evidence verifies it? Earlier systems mainly marked, measured, or prioritized; foundation and reasoning systems can help formulate meaning. Collaboration therefore requires provenance, recognition of mismatch between formal patterns and patient context, communication of uncertainty, and explicit responsibility for consequential judgment. The issue is not benchmark performance alone, but whether the sociotechnical system performs reliably across institutions while preserving evidence and enabling meaningful verification (Figure 2, a conceptual and qualitative map).
10. Beyond reasoning: The emerging frontier
The framework is also useful prospectively: What unresolved constraint does an emerging system actually relax? Agentic radiology already combines language models with tools, databases, planning, memory, and multistep orchestration.7, 62-64 Tool invocation alone, however, does not earn a new epoch. Agentic radiology would warrant promotion only if controlled delegation makes a genuinely new class of consequential coordination tractable and exposes a distinct failure regime involving permissions, provenance, recovery, tool boundaries, and authorization. Otherwise, it remains an extension of Epoch V plus workflow engineering.
Other emerging directions are useful mainly as stress tests of the criterion. Predictive world models target forward prediction and planning; radiologic forecasting of tumor evolution or treatment response remains extrapolative.65, 66 Neuro-symbolic systems target verification through rules or knowledge structures;67 continual learning targets model stasis but introduces drift, forgetting, versioning, and governance problems:68 and interventional AI pushes participation from interpretation toward action, where real-time imaging, planning, robotics, and procedural safety become inseparable.69 None should be granted epoch status by name alone.
These trajectories share one test: Which bottleneck do they remove? A larger model is not a revolution if it reproduces the same capability at greater cost. A meaningful advance should expand representational access, change the organization or controllability of inference, remove a genuine workflow constraint, or make a previously inaccessible task safe and tractable. A computational constraint may be relaxed while data access, trust, interoperability, regulation, or workflow integration remains clinically binding.9 The framework would be weakened if future systems repeatedly acquired genuinely new clinical roles without relaxing an identifiable representational, inferential, or coordination constraint. Promotion therefore depends on demonstrated constraint relaxation, not marketing vocabulary.
11. Limitations of the framework
This framework is explanatory rather than exhaustive. The review is selective and non-systematic, and its epochs are retrospective constructs rather than discrete events. Bottleneck identity is the primary taxonomic rule; chain position is a separate predictive dimension. Epoch III deliberately groups two distinct branches within one analytical period; treating them as occupants of a shared post–Epoch-II research frontier is an explanatory choice, not a claim of shared mechanism, synchronous invention, or causal sequence. The exact number of analytical epochs is therefore not a strong theoretical claim. Data, hardware, workflow, regulation, reimbursement, and professional practice may remain binding. Epoch V is especially provisional because its evidence mainly concerns the research frontier. Later-stage evidentiary demands are normative inferences from failure modes and clinical roles, and Figure 2 presents a conceptual qualitative hypothesis rather than a quantitative risk model.
12. Conclusion
The history of AI in radiology is therefore not merely a sequence of architectures. Bottleneck identity explains why certain technical changes deserve separate historical status; chain position generates a distinct clinical prediction about how to verify those changes.
Radiology makes this relationship unusually visible because its practice can be analyzed from signal and image formation through finding, concept, and recommendation while retaining recursive clinical feedback. Successive AI paradigms have entered that scaffold at different levels and times, and older architectures continue to coexist. Epoch V remains the least mature empirically: Radiology-specific reasoning evidence is still predominantly preclinical. The broader implication is methodological. As machine participation becomes more embedded, semantically richer, and less directly inspectable, safety cannot rest on the assumption that a supervising radiologist will simply see the error. The governing question should be whether evidence, provenance, and verification are appropriate to the level at which the system participates in care.


