Abstract
Invasive physiologic assessment is often needed for moderately stenotic coronary lesions, although it remains underused in routine practice. It remains unknown whether GPT-based large language models (LLMs) can estimate coronary physiology from coronary angiographic images, and whether retrieval augmentation can improve this task.
We performed a retrospective pilot study of consecutive cases undergoing coronary angiography with invasive instantaneous wave-free ratio (iFR) assessment between 2023 and 2025. Eligible cases required invasive iFR and two orthogonal end-diastolic still frames of the target vessel at maximal opacification. We compared a baseline GPT-5.2 model without retrieval-augmented generation (RAG), termed No-RAG, with the same GPT-5.2 model using RAG, termed RAG. Both conditions received identical angiographic frames and structured clinical text. The RAG modification added the top five case-specific text chunks retrieved from four coronary physiology/revascularization documents to provide physiologic thresholds, guideline context, and uncertainty framing; no additional angiographic images or lesion-specific iFR information were provided. Frame-level predictions were averaged to derive a case-level predicted iFR. The primary endpoint was agreement between predicted and measured iFR.
Of 34 eligible cases screened, 32 vessels were included. The cohort comprised 25/32 cases (78.1%) with significant disease (iFR ≤ 0.89) and 7 cases classified as non-ischemic. Without RAG, mean absolute error (MAE) was 0.064 and the root mean square error (RMSE) was 0.083, with weak correlation with invasive iFR (r = 0.205,
= 0.259). With RAG, point estimates favored improved continuous agreement, with MAE decreasing to 0.029, RMSE to 0.038, and correlation increasing to r = 0.830 (
< 0.001). Threshold-based classification also yielded higher point estimates for accuracy, increasing from 0.750 to 0.906.
In this small pilot study, improved point estimates for agreement between LLM-predicted and invasively measured iFR were seen after adding RAG to a GPT-based model for estimating iFR from angiographic imaging. These findings suggest that functionally classifying coronary stenoses is limited by overestimating severity in less severe stenoses, and that a scaling correction is needed. The results, however, require validation in larger, more balanced cohorts.