Ophthalmology

Ophthalmology contains one of the clearest regulated uses of autonomous clinical AI: diabetic retinopathy screening in a defined primary-care workflow. FDA granted De Novo authorization to IDx-DR, now LumineticsCore, in 2018 as a Class II retinal diagnostic software device (FDA DEN180001). That authorization applies to the labeled screening task, population, camera, operator, and workflow, not to comprehensive eye care.

Learning Objectives

After reading this chapter, you will be able to:

  • Evaluate FDA-authorized autonomous diabetic retinopathy screening systems and understand why the bounded workflow translated
  • Understand AI applications for age-related macular degeneration monitoring and glaucoma screening
  • Assess retinopathy of prematurity AI systems and their deployment challenges
  • Navigate integration of screening AI with specialty referral pathways
  • Recognize that a labeled screening output is not a comprehensive eye examination
  • Apply evidence-based frameworks for ophthalmology AI implementation
  • Understand reimbursement pathways for autonomous AI diagnostics

The Clinical Context:

Ophthalmology relies on fundus photography, optical coherence tomography (OCT), visual field testing, and slit-lamp photography. Under defined acquisition and quality-control protocols, these modalities can produce structured images and measurements suitable for bounded AI tasks. Reproducibility, image quality, reference standards, and outcome relevance still vary by device, disease, population, and workflow.

What Actually Works:

Application System Performance Status Deployment
Diabetic retinopathy screening IDx-DR / LumineticsCore 87.2% sensitivity, 90.7% specificity in the pivotal study FDA De Novo authorization, DEN180001 (2018) Defined autonomous screening workflow
Diabetic retinopathy screening EyeArt Product- and cohort-specific performance in FDA record FDA 510(k), K200667 (2020) Defined cameras and intended population
AMD home monitoring ForeseeHome Median 5 fewer letters lost at CNV detection in HOME FDA 510(k), K091579 Preferential hyperacuity monitoring, not evidence for image-based AI
ROP image classification i-ROP research system 93% sensitivity and 94% specificity for plus disease on a 100-image expert comparison Research evidence Requires complete neonatal screening pathway
Glaucoma image analysis Research systems Performance depends on target and test set Task-specific research Fundus morphology alone is not comprehensive diagnosis

IDx-DR: An Autonomous AI Deployment Case

Why diabetic retinopathy screening succeeded:

  1. Clear clinical need: Diabetic eye-examination completion remains incomplete and unequal across settings and populations
  2. Access barrier: Primary care lacks retinal specialists, particularly in rural areas
  3. Standardized imaging: Topcon fundus camera with defined protocol
  4. Binary decision: Refer vs. rescreen (not nuanced diagnosis)
  5. Rigorous validation: Prospective multicenter study, 900 enrolled participants, and 819 with both reference and device results at 10 primary-care sites
  6. Safety mechanism: All positive screens go to ophthalmologists
  7. Reimbursement pathway: CPT code 92229 for autonomous AI examination

Performance (Abramoff et al., NPJ Digital Medicine, 2018): - Sensitivity: 87.2% for more-than-mild diabetic retinopathy - Specificity: 90.7% - Imageability: 96.1% (images adequate for analysis) - Prospective validation with primary care operators (not ophthalmology experts)

What Does Not Work:

  • Autonomous glaucoma diagnosis: Glaucoma requires visual fields, intraocular pressure, gonioscopy. Fundus imaging AI cannot replace comprehensive exam.
  • One-size-fits-all algorithms: Performance varies by population, imaging equipment, and clinical setting.
  • End-to-end 3D OCT automation: FOCUS reported high multicenter F1 on retrospective imaging; that is not a vision-outcome or unmanned-clinic claim (Zhang et al., 2026).

AI-agent clinics vs autonomous screening:

Yan et al. (2026) Comment in Nature Medicine describes initial lessons from a real-world AI-agent eye clinic (AI-TEC) at Beijing Tsinghua Changgung Hospital Eye Center / BERI (Yan et al., 2026). The standfirst argues that moving from AI-assisted tools to more AI-native ophthalmic care depends on workflow integration, clinician engagement, and measurable clinical value, not model performance alone. Published figures frame a multi-agent pathway (pre-consultation, triage, diagnosis, ophthalmologist copilot, patient support) with the ophthalmologist retaining the final decision, and show that expert-reviewed training data improved discrimination versus a larger raw set on a prospective panel (n = 472). The official supplement describes a pilot where clinicians diagnose first without AI, then review AI predictions and indicate whether AI would influence their decisions before the final diagnosis, so the evaluation targets decision influence rather than algorithm scores alone. Treat this as early tertiary-hospital implementation experience: it is not FDA-style autonomous diabetic retinopathy screening evidence, figure AUROCs are not vision-outcome proof, and the Comment does not show that multi-agent clinics replace ophthalmologists or improve patient outcomes versus usual care.

Critical Insights:

  • AAO Diabetic Retinopathy PPP: The formal condition-specific guideline discusses validated digital imaging within diabetic-retinopathy care. Its recommendations should be read directly rather than inferred from product marketing or educational summaries (AAO Diabetic Retinopathy PPP).
  • 2026 diagnostic-accuracy meta-analysis: Across 28 studies, autonomous AI showed higher pooled sensitivity than store-and-forward pathways for any DR, referable DR, vision-threatening DR, and diabetic macular edema. The authors emphasized decision consequences rather than simple superiority: pathway choice depends on disease prevalence, decision thresholds, referral capacity, and the relative harm of missed disease versus excess referrals (Chen et al., 2026).
  • 2026 health-system CEA: Preferred diabetic-eye strategy flips with scale and willingness-to-pay; ECP can dominate in integrated systems at $800 per patient over 5 years (Ahmed et al., 2026).
  • Screening vs. diagnosis: DR screening AI identifies who needs referral, not what treatment they need.
  • Referral completion: DR screening value depends on completed eye-care follow-up, not a screening score alone.
  • Glaucoma is different: Cannot be diagnosed from fundus photos alone, requires multimodal assessment.
  • Oculomics / AF: Routine macular OCT can carry AF-risk signal. See Cardiology for the two-cohort association and its limits.

The Bottom Line:

Autonomous diabetic retinopathy screening has prospective, randomized workflow, regulatory, and observational implementation evidence, but each claim must remain tied to its population and endpoint. AMD monitoring can improve detection of conversion to neovascular AMD, but ForeseeHome is preferential hyperacuity perimetry and should not automatically be categorized as machine-learning evidence. Glaucoma and ROP image analysis require task-specific confirmation and complete care pathways. No screening output replaces a clinically indicated comprehensive eye examination.


Introduction: Why Diabetic Retinopathy Screening Translated

Ophthalmology occupies a distinctive position in the history of medical AI. In April 2018, FDA granted De Novo authorization to IDx-DR for autonomous diabetic retinopathy screening (FDA DEN180001). The system could return its labeled screening result without specialist interpretation, but the authorization did not cover unrelated eye disease, treatment selection, or every patient with diabetes.

Why did ophthalmology succeed where radiology, pathology, and other specialties have struggled to achieve truly autonomous AI?

Five factors created the perfect conditions:

  1. Standardized imaging: Fundus photography is highly reproducible. The Topcon fundus camera used by IDx-DR captures the same retinal fields at the same resolution every time.

  2. Well-defined disease criteria: Diabetic retinopathy staging (none, mild, moderate, severe, proliferative) is standardized and validated against clinical outcomes.

  3. Clear clinical need: Diabetic eye-examination completion remains incomplete and varies substantially by population and setting. Point-of-care testing addresses a measurable access and follow-up problem.

  4. Binary output: Screening AI does not need to stage retinopathy severity or recommend treatment. It answers one question: “Does this patient need to see an ophthalmologist?”

  5. Defined failure pathways: The label specifies what to do after a positive, negative, or insufficient-quality output. A false negative is not harmless, and new symptoms or another clinical indication override a prior screening result.

These conditions describe one bounded ophthalmic task. They do not prove that ophthalmology as a whole is uniquely suited to autonomy, and they should not be used to dismiss regulated or assistive AI in radiology, pathology, or other specialties.

But success came with caveats. Autonomous DR screening works because it operates in a narrow, well-defined lane: detecting referable retinopathy in primary care diabetic patients. Expand beyond that lane (glaucoma, cataracts, retinal detachment), and AI quickly hits limits.


Part 1: Autonomous Diabetic Retinopathy Screening

The Clinical Problem

Diabetic retinopathy (DR) is the leading cause of blindness in working-age adults in developed countries. Early detection and treatment (laser photocoagulation, anti-VEGF injections) prevent vision loss. But detection requires screening.

Current screening rates are inadequate: - American Diabetes Association recommends annual dilated eye exams for all diabetics - Published U.S. studies have reported substantial care gaps, with adherence in some settings as low as 20% (Leong et al., 2026) - Rural areas have severe ophthalmologist shortages - Primary care clinics lack access to retinal specialists

Why diabetics do not get screened: - Geographic barriers (no nearby ophthalmologist) - Long wait times for appointments (6-12 months in some areas) - Cost and insurance barriers - Lack of symptoms in early DR (patients do not perceive urgency)

The result: Patients present with advanced retinopathy when vision loss is irreversible.

The Solution: Autonomous Screening in Primary Care

If ophthalmologists cannot reach diabetic patients, bring DR screening to where diabetics already receive care: primary care clinics.

The workflow: 1. Medical assistant takes retinal photos during routine diabetes visit 2. AI analyzes images immediately 3. Result delivered in minutes: “Referable DR detected, refer to ophthalmologist” or “Negative for referable DR, rescreen in 12 months” 4. No ophthalmologist in the loop for negative results (autonomous operation)

This model addresses the access barrier by embedding screening in existing workflows, using non-specialist staff, and eliminating wait times.

IDx-DR: The First FDA-Authorized Autonomous AI Diagnostic

FDA status: De Novo authorization, DEN180001, April 11, 2018 (FDA De Novo Summary).

How IDx-DR works:

  • Requires Topcon NW400 fundus camera (standardized hardware)
  • Medical assistant captures two 45-degree images per eye, one macula-centered and one optic-disc-centered
  • Images uploaded to IDx-DR cloud platform
  • The authorized software analyzes image quality and the presence or absence of more-than-mild diabetic retinopathy within the labeled workflow
  • Binary output within 60 seconds:
    • “More than mild diabetic retinopathy detected. Refer to an eye care professional.”
    • “Negative for more than mild diabetic retinopathy. Rescreen in 12 months.”

Validation study (Abramoff et al., NPJ Digital Medicine, 2018):

Prospective, multicenter study at 10 primary-care sites: - Population: 900 enrolled adults with diabetes; 819 had both an interpretable reference-standard result and a diagnostic IDx-DR result - Operators: Primary care medical assistants (not ophthalmic photographers) - Reference standard: Wisconsin Fundus Photograph Reading Center grading by certified graders

Performance:

Metric IDx-DR Result Prespecified target Clinical Significance
Sensitivity 87.2% (95% CI: 81.8-91.2%) >85% Detects 87 of 100 referable DR cases
Specificity 90.7% (95% CI: 88.3-92.7%) >82.5% Correctly identifies 91 of 100 negative cases
Imageability 96.1% >85% 96% of images adequate for analysis

Key validation features:

  • Real-world setting: Primary care clinics, not academic ophthalmology centers
  • Non-specialist operators: Medical assistants, not ophthalmic photographers
  • Prospective design: Consecutive eligible patients, not retrospective image databases
  • Diverse population: Multiple sites, varied demographics, real-world diabetic patients

Why this validation matters:

Many AI systems show excellent performance on curated research datasets but fail in clinical practice. IDx-DR was validated in the exact deployment environment: primary care, non-specialist operators, unselected diabetic patients. This rigorous approach enabled FDA autonomous clearance.

EyeArt: The Second FDA-Cleared System

FDA status: 510(k) clearance, K200667, August 3, 2020 (FDA K200667).

How EyeArt differs from IDx-DR:

  • The FDA summary specifies Canon CR-2 AF and Canon CR-2 Plus AF cameras
  • Produces eye-level results for more-than-mild and vision-threatening diabetic retinopathy, including an ungradable category
  • Device comparison must use the intended population, camera, output definition, and handling of ungradable images rather than transferring a single headline metric

Validation study:

The FDA record describes sequential and enrichment-permitted cohorts at U.S. primary-care and ophthalmology sites, a reading-center reference standard, human-factors testing, and results stratified by output and setting (FDA K200667). It does not support the restored claim that one retrospective 24,803-encounter analysis established noninferiority.

Deployment:

EyeArt has a commercial deployment pathway, but current local availability, compatible hardware, contracting, and workflow should be confirmed directly rather than inferred from the authorization record.

LumineticsCore and AEYE-DS: Additional FDA-Cleared Systems

The U.S. regulatory record includes several authorized autonomous diabetic-retinopathy screening systems with distinct labeling:

  • LumineticsCore: the current trade name of the system originally authorized as IDx-DR under DEN180001. Real-world performance varies by site, operator, camera, and whether ungradable examinations are included (Teng et al., 2025).
  • AEYE-DS: 510(k) K240058, cleared April 23, 2024 as a diabetic retinopathy detection device (FDA K240058).
  • EyeArt: 510(k) K200667, with the intended population, outputs, and cameras described above.

Authorization is product-specific. The IDx-DR pivotal study’s prespecified performance targets should not be converted into a universal FDA threshold for every autonomous system.

Real-World Deployment and Outcomes

Scale and evidence: Commercial adoption is expanding, but a deployment count is not a clinical endpoint and changes over time.

Implementation settings: - Federally qualified health centers - Indian Health Service clinics - Veterans Affairs primary care clinics - Rural health networks - Endocrinology and diabetes specialty clinics

Randomized workflow evidence:

The ACCESS trial randomized 164 youth with diabetes to an immediate point-of-care autonomous examination or an enhanced referral-and-education control. Six-month diabetic eye-examination completion was 100% (81/81) in the intervention group and 22% (18/82) in the control group. Among 25 participants with an abnormal autonomous result, 64% completed indicated eye-care follow-through. The system was used outside its adult FDA label, and all images were overread by a retina specialist, so the trial supports the tested pediatric workflow rather than label expansion (Wolf et al., 2024).

Observational implementation evidence:

A 2026 retrospective, single-health-system study compared 3,745 patients referred through autonomous AI and clinician-referral pathways. In an exploratory propensity-weighted analysis, the AI pathway was associated with a higher probability that referred patients were African American (OR 1.15, 95% CI 1.02–1.29), but not with Medicaid coverage. The authors explicitly cautioned that the analysis was noncausal and could partly reflect office-based screening rather than autonomous AI itself (Leong et al., 2026).

Diagnostic synthesis:

A 2026 PRISMA-DTA systematic review and meta-analysis included 28 diagnostic-accuracy studies comparing autonomous AI and store-and-forward teleophthalmology. Pooled sensitivity favored AI across any, referable, and vision-threatening diabetic retinopathy and diabetic macular edema, but the authors emphasized that pathway preference depends on prevalence, referral capacity, thresholds, and the harms assigned to missed disease and excess referral (Chen et al., 2026). This supports pathway-specific evaluation, not a universal superiority claim.

A 5-year TreeAge Markov microsimulation compared autonomous AI, teleretinal imaging, and eye-care-professional (ECP) diabetic eye exams from a health-system perspective (Ahmed et al., 2026). Modeled AI strategies completed 3 times as many screenings, detected 3.6–3.8 times as many true positives, and started treatment in 7.5–8.0 times as many patients as ECP; at $800 willingness-to-pay per patient over 5 years, ECP dominated across 250–14,000 patients in the integrated-system scenario, while handheld AI won at higher primary-care volume or higher willingness-to-pay (M.D. Abramoff is an investor, director, and consultant at Digital Diagnostics). This is a cost-effectiveness model, not measured ROI and not a screening RCT.

Barriers to completion of care: - Positive and ungradable results can still fail to reach specialist evaluation; the completion rate must be measured locally rather than assumed from a universal range - Transportation barriers, appointment availability, and insurance issues - Tracking and care coordination systems essential to success

Cross-country scale (Comment, not a trial): Tiwari, Widner, Ruamviboonsuk, Turner, and colleagues describe practical lessons from scaling one diabetic-retinopathy deep-learning tool from a single hospital into three highly distinct settings - India, Thailand, and Australia - to more than a million patients screened (Tiwari et al., Nat Med, 2026). Treat this as an implementation narrative from the Google/partner program family (see also Widner et al., 2023; Ruamviboonsuk et al., 2022), not as new randomized evidence of vision outcomes and not as a substitute for the FDA-authorized IDx-DR/EyeArt pivotal pathways above.

Reimbursement: CPT Code 92229

CPT 92229: “Imaging of retina for detection or monitoring of disease; remote, autonomous analysis with report”

  • Approved by CMS in 2021
  • Covers autonomous AI DR screening
  • Coverage, payment, eligible devices, and billing requirements vary by payer and jurisdiction
  • The code describes a service; it does not itself prove coverage, profitability, or clinical value

This reimbursement pathway can support DR screening AI in some primary-care contracts, but a billing code does not settle cost-effectiveness; preferred strategy still depends on health-system scale and willingness-to-pay (Ahmed et al., 2026).

Why Autonomous DR Screening Works

The success of IDx-DR, EyeArt, and similar systems rests on several factors:

  1. Addresses real access barrier: Not a workflow optimization, but genuine expansion of care to underserved populations
  2. Narrow scope: Binary referable/non-referable decision, not nuanced staging or treatment planning
  3. Standardized input: High-quality fundus photography with defined protocols
  4. Acceptable performance: 87-91% sensitivity meets clinical needs (no screening test is perfect)
  5. Safety mechanism: Positive screens go to ophthalmologists for confirmation and treatment
  6. Defined negative pathway: The labeled output may recommend retesting in 12 months, but new symptoms, a different disease concern, or another clinical indication requires examination regardless of the screening result
  7. Prospective validation: Real-world performance data in deployment settings

What autonomous DR screening does NOT do:

  • Stage severity of DR (that’s ophthalmologist’s job)
  • Recommend treatment (laser, anti-VEGF, surgery)
  • Detect other eye conditions (glaucoma, cataracts, retinal detachment)
  • Replace comprehensive eye examination

Autonomous DR screening is successful precisely because it operates within carefully defined boundaries.

Implementation Considerations for Primary Care

Before deploying DR screening AI:

  1. Establish referral pathway with ophthalmology or optometry:
    • Identified ophthalmologist/optometrist to receive referrals
    • Process for urgent referrals (severe DR requiring immediate treatment)
    • Tracking system for referral completion
  2. Train medical assistants on imaging protocol:
    • Proper patient positioning
    • Camera focus and alignment
    • Recognizing inadequate images (motion artifact, small pupil, media opacity)
    • When to reattempt imaging vs. refer for dilated exam
  3. Integrate with EHR:
    • Results flow directly to patient chart
    • Automated referral order generation for positive screens
    • Recall system for 12-month rescreening
  4. Track quality metrics:
    • Imageability and ungradable-examination rates by operator, camera, site, and subgroup
    • Percentage of positive and ungradable screens completing the indicated eye-care pathway
    • Time from screen to evaluation, stratified by output urgency and local capacity
  5. Educate patients:
    • DR screening is not comprehensive eye exam
    • Negative screen does not mean eyes are healthy overall
    • Annual comprehensive eye exams still recommended for diabetics

Limitations and Failure Modes

When DR screening AI fails:

1. Image quality issues: - Small pupils (diabetics often have autonomic neuropathy causing poor dilation) - Media opacities (cataracts, vitreous hemorrhage) - Patient cooperation (dementia, visual impairment preventing fixation) - Operator inexperience

Mitigation: Pharmacologic dilation (tropicamide 1%) improves imageability but adds time and requires provider order.

2. False negatives: - Subtle DR changes near threshold - Peripheral retinopathy beyond imaging field - Early proliferative changes

Mitigation: Follow the labeled retesting pathway for eligible asymptomatic patients and make clear that a negative result does not rule out other ocular disease. New symptoms or a separate clinical concern require timely examination. The pivotal study did not establish that every missed case would remain safe until the next annual screen.

3. False positives (91% specificity means 9% false alarms): - Other retinal conditions mimicking DR (hypertensive retinopathy, vein occlusions) - Image artifacts

Mitigation: Ophthalmologist evaluation of all positive screens confirms or refutes AI finding. Some false positives beneficial (detect other pathology requiring treatment).

4. Conditions outside scope of algorithm: - Glaucoma - Age-related macular degeneration - Cataracts - Retinal detachment - Optic nerve disorders

Mitigation: Educate patients and providers that DR screening does not replace comprehensive eye examination for other conditions.


Part 2: Age-Related Macular Degeneration (AMD) Monitoring

ForeseeHome AMD Monitoring Program

Age-related macular degeneration is the leading cause of blindness in adults over 65. The disease progresses from dry AMD (drusen, retinal pigment epithelium changes) to wet AMD (choroidal neovascularization) in 10-15% of patients. Once wet AMD develops, rapid vision loss occurs unless treated promptly with anti-VEGF injections.

Clinical challenge: Detecting conversion from dry to wet AMD early enough to preserve vision.

Standard monitoring: Patients perform daily Amsler grid testing at home, looking for distortions indicating new neovascularization. Compliance is poor, and sensitivity is limited.

How ForeseeHome Works

Device: FDA 510(k)-cleared preferential hyperacuity perimeter, K091579 (FDA K091579).

Protocol: - Patient performs 3-minute visual test daily - Device presents hyperacuity grid pattern - Patient marks perceived distortions using touchscreen - The device and monitoring center analyze serial hyperacuity-test results for changes suggesting conversion to neovascular AMD

Category boundary:

ForeseeHome is a remote visual-function monitoring system. Its randomized evidence should not be presented as proof for image-based machine learning, autonomous retinal diagnosis, or treatment selection unless the evaluated computational component is explicitly identified.

Evidence: HOME Study

HOME (Home Monitoring of the Eye) Study (AREDS2-HOME Study Research Group, 2014):

Unmasked randomized trial of 1,520 participants at high risk of choroidal neovascularization:

Design: - Intervention: ForeseeHome daily testing + standard care - Control: Standard care alone (Amsler grid) - Primary endpoint: Visual acuity at wet AMD detection

Results:

Outcome ForeseeHome Standard Care Difference
Visual acuity decline at CNV detection Median -4 letters Median -9 letters P=0.021
Visual acuity at detection 20/40 or better in 94% (frequent users) 20/40 or better in 62% 32% absolute difference

Interpretation: The device-plus-standard-care strategy detected conversion with less visual-acuity loss at the time of detection. The trial did not test an image classifier or autonomous treatment decision, and it was stopped early for efficacy after 82 conversion events.

Medicare Coverage and Implementation

Medicare coverage: Approved for high-risk patients: - Bilateral intermediate AMD (large drusen in both eyes), OR - Wet AMD in one eye with intermediate AMD in fellow eye

Reimbursement: Coverage and payment depend on current eligibility rules, payer policy, and supplier arrangements. They should be verified rather than represented by a fixed monthly amount.

Barriers: - Requires patient compliance with daily testing - Device setup and training needed - Not all patients technologically comfortable - Requires ophthalmology follow-up infrastructure for alerts

Real-world effectiveness:

Real-world adherence and alert yield vary by patient selection and implementation. The HOME trial itself screened 1,970 people and enrolled 1,520; inability to pass device qualification was the most common reason for screening failure. Effectiveness therefore depends on qualification, sustained testing, alert handling, and access to prompt examination.

Implementation Considerations

Patient selection: - High-risk AMD (bilateral intermediate or unilateral wet) - Cognitively intact and able to perform daily testing - Access to ophthalmology for urgent evaluation if alert triggered

Infrastructure: - Process for urgent ophthalmology appointments (within 1-3 days of alert) - Staff to handle device setup and troubleshooting - Patient education on importance of daily testing

Monitoring: - Track compliance (percentage of days tested) - Response time from alert to ophthalmology evaluation - Conversion rate and visual acuity outcomes

Color-Fundus Conversational AMD Models

An August 2026 Ophthalmology Retina intramural NIH study (National Library of Medicine and National Eye Institute) evaluated OcularChat, a multimodal large language model fine-tuned from Qwen2.5-VL on AREDS color fundus photographs paired with 705,850 GPT-4V-simulated patient-physician dialogues, using 46,167 training CFPs from 3,192 participants and 13,166 test CFPs from 915 participants under a participant-level split (Gu et al., 2026). On the AREDS test set, accuracy was 0.954 for advanced AMD, 0.849 for pigmentary abnormalities, and 0.678 for drusen size, higher than the general vision-language models in the comparison; on external AREDS2 images, advanced-AMD accuracy was 0.892 while F1 fell to 0.494 (Gu et al., 2026). Three ophthalmologists graded 120 dialogues and assigned OcularChat higher mean rubric scores than the Qwen2.5-VL-32B baseline (Gu et al., 2026). The training dialogues were simulated and are not a substitute for real patient counseling. The model is CFP-only: it does not use OCT and is not a full multimodal ophthalmic workup. The authors noted that correct advanced-AMD labels can still emphasize earlier features (drusen or pigment) rather than geographic atrophy or neovascular criteria. A correct classification label is not clinically valid reasoning, and the study does not demonstrate workflow benefit or autonomous screening.


Part 3: Glaucoma Detection AI

Why Glaucoma AI Is Fundamentally Different

Glaucoma is progressive optic neuropathy causing irreversible vision loss. Diagnosis requires:

  1. Structural changes: Optic disc cupping, nerve fiber layer thinning
  2. Functional changes: Visual field defects
  3. Risk factors: Elevated intraocular pressure, family history, age
  4. Exclusion of other causes: Neurologic disease, vascular occlusion

Fundus-photography AI can detect optic-disc structural changes such as cup-to-disc ratio and rim thinning. Performance varies with the target definition, image quality, population, reference standard, threshold, and whether unexpected images are included.

But structural changes alone cannot diagnose glaucoma.

A patient with suspicious optic disc cupping requires: - Intraocular pressure measurement (tonometry) - Visual field testing (perimetry) - Gonioscopy (anterior chamber angle assessment) - OCT of optic nerve and retinal nerve fiber layer - Clinical assessment of progression over time

Fundus photo AI cannot replace this multimodal assessment.

Current Role: Triage and Screening, Not Diagnosis

The research systems discussed here are best interpreted as triage tools unless an exact authorization record establishes a broader intended use:

Screening programs: - Community health fairs - Diabetic eye screening (opportunistic glaucoma screening) - Primary care clinics

Output: “Suspicious optic disc changes, recommend comprehensive eye examination”

Not autonomous diagnosis: All flagged patients require full ophthalmology evaluation.

Research Systems

AIROGS (AI for Robust Glaucoma Screening): - The challenge dataset included approximately 113,000 images from about 60,000 patients and 500 screening centers - It explicitly tested robustness to ungradable and unexpected images; the reported 0.99 AUC for the highest-scoring team concerned on-the-fly detection of ungradable images, not a general glaucoma-diagnosis endpoint - The challenge illustrates that out-of-distribution and low-quality images are part of the clinical task, not preprocessing details (de Vente et al., 2024)

Research fundus-photo classifiers can reach high discrimination on selected test sets. This does not establish comprehensive diagnosis, safe autonomous use, or improved visual outcomes.

Unlike DR screening (binary refer/rescreen decision), glaucoma screening requires comprehensive follow-up. A positive glaucoma AI screen cannot be acted upon without full workup. This limits autonomous deployment potential.

Evidence Required for Autonomous Glaucoma Diagnosis

Comparison to DR screening:

Factor DR Screening (Works) Glaucoma Screening (Limited)
Imaging input Fundus photos alone Fundus photos + OCT + visual fields + IOP
Decision Binary (refer/rescreen) Nuanced (suspect vs. definite, progression assessment)
Intervention Refer to ophthalmology Requires IOP lowering (drops, laser, surgery)
Monitoring Annual rescreening Frequent monitoring (3-6 months) for progression
Risk of false negative Missed referable DR may progress before retesting Missed glaucoma may progress silently to irreversible vision loss

Autonomous glaucoma diagnosis would require evidence for the complete input set, intended population, reference standard, ungradable pathway, longitudinal progression, and downstream treatment decision. A fundus-photo classifier alone remains a structural triage signal.


Part 4: Retinopathy of Prematurity (ROP)

The Clinical Problem

Retinopathy of prematurity (ROP) affects premature infants, particularly those born <32 weeks gestation and <1,500 grams. Abnormal retinal vascularization can progress to retinal detachment and blindness.

Screening requirements: - Serial dilated fundus exams every 1-2 weeks until retinal vascularization complete - Exam requires neonatal ophthalmologist expertise - Treatment window is narrow (hours to days for severe ROP)

Challenges: - Neonatal ophthalmologist shortage - Exams stressful for fragile infants (bradycardia, desaturation during exam) - Variability in ROP grading between examiners

i-ROP Deep Learning System

i-ROP (Imaging and Informatics in ROP):

Developed at Oregon Health & Science University and validated in multicenter studies:

Input: Wide-angle retinal images (RetCam imaging system)

Output: - Plus disease detection (abnormal vascular tortuosity/dilation) - Pre-plus vs. plus vs. normal classification - Treatment-requiring ROP (Type 1 ROP) detection

Performance (Brown et al., JAMA Ophthalmology, 2018):

The development analysis used five-fold cross-validation on 5,511 retinal images. Mean AUC was 0.94 for normal versus pre-plus/plus and 0.98 for plus versus normal/pre-plus. A separate 100-image comparison against the reference-standard diagnosis produced 93% sensitivity and 94% specificity for plus disease, with quadratic-weighted kappa 0.92 (Brown et al., 2018). The 5,511-image cross-validation and 100-image expert comparison are different analyses and should not be combined into one external-validation claim.

Barriers to Deployment

Why i-ROP is not yet widely deployed despite strong performance:

  1. Infrastructure: Requires RetCam imaging system (expensive, not available at all NICUs)
  2. Expertise: Still requires ophthalmologist confirmation for treatment decisions
  3. Medicolegal risk: Missing treatment-requiring ROP has severe consequences (blindness)
  4. Low-volume application: Only affects very premature infants (smaller market than DR)

Current status: The cited publication supports research classification performance, not autonomous screening authorization or treatment selection. Current regulatory status must be verified against an exact public device record before clinical deployment.

Potential Future Role

Telemedicine screening model: - NICU staff perform RetCam imaging - Images analyzed by i-ROP AI - Flagged cases reviewed remotely by neonatal ophthalmologist via telemedicine - In-person exam for treatment planning if ROP detected

This model addresses ophthalmologist shortage while maintaining safety oversight.


Part 5: OCT Analysis and Imaging AI

Optical Coherence Tomography (OCT) AI

OCT provides cross-sectional imaging of the retina with micrometer resolution. AI applications in OCT analysis include:

1. Automated retinal layer segmentation: - Measure retinal thickness in diabetic macular edema - Track progression in glaucoma (ganglion cell layer thinning) - Monitor treatment response to anti-VEGF injections

The same macular ganglion-cell measurements can also carry extracardiac AF-risk signal across AlzEye and UK Biobank. That opportunistic oculomics association is covered in the cardiology wearable AF section (Huemer et al., 2026). Two-cohort association, not an AF diagnostic pivotal; hospital-coded AF undercounts silent AF.

Performance and regulation: Segmentation performance and authorization are product-, software-version-, scanner-, and task-specific. Automated measurements require review of segmentation quality before they inform a clinical decision. A manufacturer name or integration into an OCT platform does not establish that every algorithmic function is FDA-authorized.

2. Disease classification from OCT: - Differentiate wet vs. dry AMD - Detect macular holes, epiretinal membranes - Identify fluid in central serous chorioretinopathy

Evidence: Research studies report task-specific discrimination and segmentation results, but accuracy cannot be transferred across disease targets, scanners, populations, or reference standards. Device integration is not itself evidence of improved patient outcomes.

Generative retinal angiography reporting: InterpreFFA, a diagnosis-supervised contrastive learning framework for fundus fluorescein angiography interpretation, was validated on multicenter datasets and tested in a simulated clinical setting. With the tool, two residents improved diagnostic accuracy from 85.55% to 90.34% and shortened report time from 153.93 to 108.08 seconds, while AI-generated reports scored slightly lower than manual reports from ophthalmologists (Shao et al., 2025). This is supervised report-assistance evidence, not autonomous angiography interpretation.

A 19 August 2026 npj Digital Medicine multicenter study evaluated FOCUS, a foundation-model 3D OCT pipeline that chains quality assessment, abnormality detection, multi-disease classification, and 2D-to-3D aggregation (Zhang et al., 2026). Trained and tested on 3,300 patients (40,672 slices) and externally validated on 1,345 patients (18,498 slices) across four centers and devices, reported F1 was 99.01% for quality assessment, 97.46% for abnormality detection, and 94.39% for patient-level diagnosis, with multi-center F1 90.22–95.24%. In the authors’ human-machine comparison, FOCUS F1 was 95.47% versus 90.91% for experts on abnormality detection and 93.49% versus 91.35% on multi-disease diagnosis. This is retrospective diagnostic-workflow discrimination, not a vision-outcome or clearance claim for unmanned eye care.

3. Treatment decision support: - Predict response to anti-VEGF therapy based on OCT features - Optimize injection intervals for wet AMD

Status: Research stage. Prospective validation needed before clinical deployment.

Inherited Retinal Disease Diagnostic Support

Retina4IRD combines color fundus photographs and OCT to prioritize 17 genotype categories before confirmatory genetic testing. In a 2026 multicenter randomized trial, 300 participants with suspected inherited retinal disease were assigned to AI-assisted or specialist-only assessment, with 295 included in the final analysis. The primary top-five genetic diagnostic accuracy endpoint was 88.5% with assistance and 67.3% without assistance. Top-one through top-four accuracy also favored assistance. A downstream management composite was post hoc, and next-generation sequencing remained the reference standard (Jia et al., 2026). This is randomized evidence for diagnostic prioritization before genetic testing, not replacement of molecular confirmation or genetic counseling.

Integration with Clinical Workflow

OCT AI can reduce workflow friction when integrated directly into imaging devices: - Automated segmentation runs during image acquisition - Results displayed immediately for ophthalmologist review - No separate system or workflow disruption

Unlike standalone DR screening AI, OCT AI augments ophthalmologist interpretation rather than operating autonomously.


Part 6: Professional Society Guidelines and Position Statements

American Academy of Ophthalmology (AAO)

AAO Diabetic Retinopathy Preferred Practice Pattern:

Key recommendations: - Dilated fundus examination remains gold standard for DR screening - Validated digital imaging (including AI) may be an effective detection method - Authorized autonomous screening systems operate within condition- and product-specific intended uses - AI screening appropriate for primary care settings where dilated exams unavailable - Positive AI screens require ophthalmology referral - AI screening does not replace comprehensive eye examination for other conditions

The complete AAO Diabetic Retinopathy Preferred Practice Pattern should be consulted for the condition-specific recommendations. Educational pages and product summaries do not carry the same authority as a formal Preferred Practice Pattern.

The following chapter principles preserve the practical value of the restored checklist without attributing each statement to a formal AAO AI position:

Chapter Principles for Ophthalmology AI

Validation and Transparency: - Demand high-quality evidence for AI systems before clinical adoption - Prospective validation in real-world settings required - Training data demographics and performance across subgroups should be reported

Clinical Integration: - AI should augment, not replace, ophthalmologist clinical judgment - Maintain physician accountability for AI-assisted decisions - Clear communication to patients about AI use

Equity and Access: - AI has potential to expand access to eye care in underserved areas - Address algorithmic bias and ensure equitable performance across populations - Reimbursement models should support deployment in safety-net settings

Patient-Centered Care: - AI should enhance patient-clinician relationship, not replace it - Informed consent for AI-assisted diagnosis - Patient preferences respected

American Society of Retina Specialists (ASRS)

ASRS Artificial Intelligence Task Force review:

The ASRS AI Committee published a task-force review summarizing AI developments for retinal conditions. The 2024 update is an expert review, not a clinical practice guideline. It addresses: - The regulatory and implementation landscape for autonomous DR screening systems - AI applications across diabetic retinopathy, AMD, and other retinal conditions - Implementation considerations for retina practices - Emerging research directions in retinal AI

Reporting Expectations for Ophthalmic AI Research

The restored ARVO-attributed list is best treated as a general appraisal checklist unless a specific society policy is linked: - Report training and test data sources, eligibility, demographics, and image acquisition - Use external validation when the claim concerns transportability - Report ungradable and unexpected inputs, failure modes, and uncertainty - Match the study design and endpoint to the clinical claim rather than treating technical discrimination as outcome benefit


Part 7: Implementation Framework for Ophthalmology AI

Before Adopting Ophthalmology AI

1. Define the clinical need:

Ask: What problem are we solving? - Access barrier: Diabetics not receiving annual screening → DR screening AI appropriate - Workflow efficiency: Reducing ophthalmologist time on routine screening → Consider AI - Diagnostic accuracy: Improving detection of subtle disease → Evaluate evidence carefully

Do NOT adopt AI for the sake of adopting AI. Technology must address genuine clinical need.

2. Evaluate the evidence:

Questions to ask vendors: - What is sensitivity and specificity in prospective validation studies? - Performance in diverse populations (race, ethnicity, age, disease severity)? - External validation (not just development site performance)? - Real-world deployment data (not just research datasets)?

Red flags: - Only retrospective performance data - No demographic subgroup analysis - Refusal to share validation study details - Accuracy claims presented without the intended population, reference standard, uncertainty, and ungradable-image handling

3. Validate on your population:

Pilot testing: - Prespecify a sample and event count appropriate to the intended use, local prevalence, and uncertainty required for a decision - Compare the complete local workflow against an appropriate reference standard - Calculate site-specific performance, imageability, ungradable rate, subgroup uncertainty, and referral completion

If the intended workflow fails prespecified safety or performance criteria: Do not deploy. Investigate camera compatibility, operator training, population differences, image quality, referral capacity, and software version.

4. Establish integration with care pathways:

For DR screening: - Identified ophthalmologist/optometrist for referrals - Appointment availability matched to the device output, clinical findings, and local urgency policy - Tracking system for referral completion - Process for urgent referrals (proliferative DR, macular edema)

For AMD monitoring: - Process for urgent ophthalmology appointments (1-3 days) when alert triggered - Staff to troubleshoot device issues - Patient education and compliance monitoring

5. Train staff:

Medical assistants operating imaging equipment: - Proper patient positioning and camera alignment - Recognizing inadequate images - When to attempt pharmacologic dilation - Explaining results to patients

Ophthalmologists/optometrists receiving referrals: - Understand AI system limitations - Aware of false positive and false negative rates - Prepared to confirm or refute AI findings

6. Monitor quality metrics:

Track: - Imageability rate (percentage of exams producing gradable images) - Positive screen rate (percentage flagged for referral) - Referral completion rate (percentage of positive screens completing ophthalmology visit) - Time from positive screen to ophthalmology appointment - Confirmation rate (percentage of AI-positive screens confirmed by ophthalmologist)

Action thresholds: - Set local thresholds before deployment from device labeling, baseline variation, clinical consequences, referral capacity, and statistical uncertainty - Investigate imaging technique, equipment, population shift, transportation barriers, and appointment availability when monitored results cross those thresholds - Do not substitute universal percentages for a locally governed safety and effectiveness decision


Check Your Understanding

The four scenarios below are fictional teaching exercises. Patient histories, institutional metrics, vendor claims, product use, outcomes, prices, reimbursement rates, legal conclusions, and projected savings are illustrative unless an external source is cited in the same sentence. The exercises do not report actual events, product performance, payer contracts, or financial results.

Scenario 1: Primary Care DR Screening Implementation and Referral Pathway Failure

Clinical situation: You are a family medicine physician at a federally qualified health center (FQHC) serving a predominantly Hispanic and African American community. Your clinic implemented IDx-DR six months ago to improve diabetic retinopathy screening rates.

Baseline context: - 2,500 diabetic patients in your panel - Before IDx-DR: 35% received annual eye screening - After IDx-DR: 68% screened (major improvement!)

Initial success metrics: - Imageability rate: 94% (excellent) - Positive screen rate: 18% (consistent with literature) - 450 patients screened positive for referable DR in 6 months

Problem emerges:

You conduct a 6-month audit and discover: - Of 450 patients with positive IDx-DR screens, only 120 (27%) completed ophthalmology appointments - 330 patients (73%) did not follow up despite referrals

Reasons for missed appointments (chart review): - 180 patients: “Couldn’t get appointment within reasonable timeframe” (6-9 month wait) - 85 patients: Transportation barriers - 45 patients: Insurance issues (referral denied or high copay) - 20 patients: Language barriers (Spanish-speaking, ophthalmology office English-only)

6 months later:

Three patients from your clinic present to emergency department with vision loss:

Patient A: - 52-year-old woman with Type 2 diabetes - IDx-DR positive screen 8 months ago - Did not complete ophthalmology referral (6-month wait for appointment, could not take time off work) - Presents with sudden vision loss: proliferative diabetic retinopathy with vitreous hemorrhage - Requires urgent vitrectomy surgery - Visual prognosis poor (likely permanent vision impairment)

Patient B: - 67-year-old man with Type 2 diabetes - IDx-DR positive screen 10 months ago - Scheduled ophthalmology appointment but insurance denied referral (prior authorization issue) - Presents with bilateral macular edema - Requires anti-VEGF injections (months of treatment) - Vision recovery uncertain

Question 1: Did IDx-DR implementation succeed or fail at your clinic?

Answer: Partial success, systemic failure.

Success components: - Screening rate improved from 35% to 68% (33% absolute increase) - 450 patients identified with referable DR who might otherwise have been missed

Failure components: - 73% of positive screens did not complete ophthalmology follow-up - Patients experienced preventable vision loss despite being identified by AI - Screening without treatment pathway is detection without intervention

Root cause: Implementation focused on technology, ignored systemic barriers

Critical insight: AI does not solve access problems if downstream resources such as ophthalmology appointments, insurance navigation, transportation, and language access are inadequate. Detection alone does not establish patient benefit.

Question 2: What went wrong, and who is responsible?

Systems failures:

1. Referral pathway capacity not assessed before implementation: - Clinic deployed IDx-DR without ensuring ophthalmology availability - Local ophthalmology practices already had 6-9 month wait times - Adding 450 urgent referrals overwhelmed system

2. Transportation barriers not addressed: - FQHC serves low-income patients, many without cars - Ophthalmology office 20 miles away (no public transit access) - No transportation assistance program

3. Insurance navigation support absent: - Many patients on Medicaid with prior authorization requirements - Referral denials common (administrative barriers) - No dedicated staff to resolve insurance issues

4. Language barriers: - Ophthalmology office predominantly English-speaking staff - Spanish-speaking patients had difficulty scheduling, navigating appointments

Responsibility:

Clinic leadership: - Implemented technology without ensuring care pathway functionality - Did not conduct pre-implementation capacity assessment - Failed to track referral completion rates until audit

Health system: - Inadequate ophthalmology capacity for safety-net population - Insurance barriers unaddressed

Payers: - Prior authorization delays for urgent referrals - Denials for medically necessary ophthalmology visits

Question 3: How should IDx-DR have been implemented to avoid these failures?

Pre-implementation requirements:

1. Capacity assessment:

Before deploying IDx-DR, assess: - Current ophthalmology referral volume and wait times - Expected positive screen rate (15-20% of diabetics) - Can local ophthalmology handle increased volume?

If capacity inadequate: - Contract with ophthalmology to guarantee appointment slots for positive screens - Establish telemedicine ophthalmology for initial evaluation - Recruit additional ophthalmology providers - Do not deploy screening AI if treatment pathway broken

2. Care coordination infrastructure:

Hire care coordinator (navigator): - Tracks all positive IDx-DR screens - Schedules ophthalmology appointments - Arranges transportation - Resolves insurance issues - Follows up on missed appointments

Workflow:

IDx-DR positive screen
  ↓
Care coordinator contacted same day
  ↓
Care coordinator:
  - Calls patient within 24 hours
  - Explains need for urgent ophthalmology appointment
  - Schedules appointment (target <4 weeks)
  - Arranges transportation if needed
  - Submits prior authorization if required
  - Confirms appointment 48 hours before
  ↓
Patient attends ophthalmology appointment
  ↓
Care coordinator follows up post-visit (treatment plan, compliance)

3. Address transportation barriers:

Options: - Rideshare vouchers (Lyft, Uber) - Partnership with community transportation services - Mobile ophthalmology van (brings specialist to clinic) - Telemedicine initial consultation (reserve in-person for treatment)

4. Resolve insurance barriers:

Dedicated staff for: - Prior authorization submission (same day as positive screen) - Appeal denials - Financial assistance applications - Medicaid enrollment support

5. Language-concordant care:

  • Ensure ophthalmology office has Spanish-speaking staff
  • Provide interpreter services
  • Translated patient education materials

6. Track and respond to metrics:

Monitor monthly using locally prespecified targets: - Percentage of positive and ungradable screens completing ophthalmology evaluation within the urgency window - Percentage completing the full indicated pathway within the follow-up horizon - Reasons for missed appointments (identify systemic barriers)

If referral completion <80%: Investigate and address barriers immediately. Do NOT continue screening if pathway broken.

Question 4: Is screening without accessible treatment ethical?

Ethical framework:

Beneficence (do good): - Screening identifies disease early → enables treatment → prevents blindness - BUT only if treatment accessible

Non-maleficence (do no harm): - Screening without accessible treatment causes harm: - Anxiety from knowing disease exists but cannot access care - False reassurance (patients believe they’ve addressed problem by screening) - Opportunity cost (resources spent on screening could fund treatment access)

Justice (fair distribution): - Deploying AI screening in underserved communities without ensuring treatment access exacerbates disparities - May impose screening burden without delivering the expected downstream benefit

Autonomy: - Patients consent to screening expecting treatment will be available if needed - Screening without treatment pathway violates reasonable expectations

Conclusion: Screening without a credible pathway to timely assessment and treatment is ethically problematic.

Before deploying DR screening AI, ensure: 1. Ophthalmology capacity adequate 2. Care coordination support in place 3. Transportation barriers addressed 4. Insurance barriers resolved 5. Referral completion tracked against a prespecified, locally justified threshold

If these conditions are not met: do not expand the program until the care pathway is repaired, and direct implementation resources toward referral access and treatment capacity.

Lesson: Technology alone does not improve health outcomes. IDx-DR and similar AI systems are tools that must be integrated into functional care pathways. Screening identifies disease, but care coordination, transportation, insurance navigation, and specialist availability determine whether patients benefit. Implementation requires systemic assessment and investment beyond the AI system itself.

Scenario 2: Glaucoma AI in Primary Care and the Limits of Fundus Photo Screening

Clinical situation: You are an ophthalmologist at an academic medical center. Your hospital’s primary care network is considering deploying a glaucoma detection AI system (similar to AIROGS) to opportunistically screen diabetic patients during IDx-DR imaging.

Proposed workflow: 1. Medical assistant captures fundus photos for DR screening (IDx-DR) 2. Same images analyzed by glaucoma AI 3. If glaucoma AI flags “suspicious optic disc,” patient referred to ophthalmology

System performance (from vendor): - Sensitivity: 92% for glaucoma detection - Specificity: 88% - AUC: 0.96 - Published in peer-reviewed journal

Primary care leadership pitch: - “We’re already taking fundus photos for DR screening. Why not screen for glaucoma at the same time? It’s opportunistic screening at minimal added cost.”

Your concern: Glaucoma cannot be diagnosed from fundus photos alone.

You request pilot data: Primary care runs glaucoma AI on 1,000 patients already screened with IDx-DR.

Pilot results:

Outcome Number Percentage
Glaucoma AI positive 180 18%
Referred to ophthalmology 180 18%
Completed ophthalmology appointment 135 75% of positive
Glaucoma diagnosed 22 16% of those seen
Glaucoma suspect (monitoring needed) 48 36% of those seen
No glaucoma (false positive) 65 48% of those seen

Breakdown: - 180 patients flagged by AI - 135 completed appointments - 22 confirmed glaucoma (true positives) - 48 glaucoma suspects (borderline findings requiring monitoring) - 65 false positives (no glaucoma)

Positive predictive value: 22/135 = 16% (of those who completed appointments, only 16% had glaucoma)

Additional findings:

Of the 22 confirmed glaucoma cases: - 8 already knew they had glaucoma (on treatment, regular ophthalmology follow-up) - 14 new diagnoses

Of the 48 glaucoma suspects: - 30 had normal IOP, normal visual fields → likely false positives, but need monitoring to confirm - 18 had borderline findings (slightly elevated IOP or early field defects)

Primary care reaction: - “We found 14 new glaucoma cases! This is success!”

Your reaction: - “We generated 180 referrals, 135 appointments, to find 14 new glaucoma cases. That’s 9.6 appointments per true new diagnosis. We also flagged 8 patients already in our care, and created 65 false positives consuming appointment slots.”

Question 1: Should glaucoma AI be deployed in this primary care setting?

Answer: The hypothetical pilot does not justify unrestricted deployment. A limited redesign or discontinuation should be considered against prespecified local benefit-harm and capacity criteria.

Arguments against deployment:

1. Low positive predictive value (16%): - 84% of patients referred did not have glaucoma - High false positive rate consumes ophthalmology appointment capacity - Patients experience anxiety, cost, and time burden for false alarms

2. Ophthalmology capacity constraints: - Adding 180 glaucoma referrals per 1,000 screened overwhelms system - Delays appointments for patients with known urgent conditions - Opportunity cost: Could those appointment slots serve patients with symptomatic disease?

3. Incremental yield low: - 14 new glaucoma diagnoses per 1,000 screened (1.4% yield) - Many of those 14 may have been detected through routine care eventually - No evidence that AI screening improves long-term outcomes (no RCT data)

4. Time horizon and outcome benefit remain uncertain: - Glaucoma progression varies across patients, and irreversible damage can occur without symptoms - The hypothetical pilot does not measure time to confirmed diagnosis, treatment, progression, or preserved vision - Technical detection yield alone cannot establish that opportunistic screening improves prognosis

Arguments for deployment (with modifications):

1. Opportunistic detection: - Images already being captured for DR screening - Marginal cost low (AI analysis cheap) - Identified 14 patients who might not otherwise have been diagnosed

2. High-risk population: - Diabetics have higher glaucoma risk - African American patients (if significant proportion of screened population) have higher glaucoma risk - Screening targeted to high-risk may have better yield

3. Refining thresholds: - Vendor’s 92% sensitivity / 88% specificity may not be optimal for this setting - Could adjust threshold to increase specificity (reduce false positives) - Example: At 85% sensitivity / 95% specificity, might reduce referrals by 30-40%

Question 2: What would make glaucoma AI screening acceptable?

Requirements for responsible deployment:

1. Adjust AI threshold to increase specificity:

Work with the vendor and clinical governance team to prespecify a threshold that reflects local prevalence, the harms of missed disease and excess referral, uncertainty, and ophthalmology capacity. Any threshold change requires validation in the intended population; a higher positive predictive value cannot be assumed to preserve an acceptable sensitivity.

2. Two-stage screening:

Stage 1: Glaucoma AI flags suspicious cases Stage 2: On-site screening by trained technician: - IOP measurement (tonometry) - Apply a validated, guideline-consistent combination of tonometry, image review, history, and escalation criteria - If IOP normal AND optic disc unremarkable on technician review → routine follow-up

Potential benefit: Secondary triage may reduce unnecessary referrals, but the complete pathway must be evaluated prospectively for missed disease, delay, workload, and patient outcomes.

3. Risk stratification:

Only screen high-risk patients: - Age >60 - African American or Hispanic ethnicity - Family history of glaucoma - High myopia

Potential benefit: Clinically justified risk-based eligibility can increase pretest probability, but eligibility criteria must be validated and monitored for inequitable exclusion.

4. Capacity expansion:

Before deploying: - Ensure ophthalmology can handle increased volume - Create “glaucoma suspect clinic” with less urgent appointment tier - Use optometrists for initial suspect evaluation (reserve ophthalmologist time for confirmed cases)

5. Outcome tracking:

Track: - New glaucoma diagnoses per 1,000 screened - Progression to vision loss prevented (requires long-term follow-up) - Appointment burden on ophthalmology - Patient costs and anxiety from false positives

If harms outweigh benefits: Discontinue program

Question 3: Why cannot glaucoma AI be autonomous like DR screening AI?

Fundamental differences:

Diabetic retinopathy screening (autonomous): - Binary decision: Referable vs. not referable - Single modality: Fundus photos sufficient for screening decision - Low-risk false negatives: Rescreen in 12 months - Clear action: Refer to ophthalmology for all positives

Glaucoma screening (cannot be autonomous): - Multimodal required: Fundus photos + IOP + visual fields + OCT + gonioscopy - Nuanced diagnosis: Glaucoma suspect vs. early glaucoma vs. moderate/severe, different subtypes (open-angle, angle-closure, normal-tension) - High-risk false negatives: Missed glaucoma progresses silently to irreversible vision loss - Complex action: IOP lowering requires medication choice, monitoring for adherence and side effects, surgical decisions

Cannot diagnose glaucoma from fundus photo alone.

A patient with suspicious optic disc requires full workup. Glaucoma AI serves as triage tool, not diagnostic tool.

Lesson: Not all screening is beneficial. Glaucoma AI has high technical performance (92% sensitivity, 88% specificity) but low positive predictive value (16%) in unselected primary care populations, generating high false positive rates and consuming limited ophthalmology resources. Deployment should be limited to high-risk populations, with threshold adjustments to increase specificity, two-stage triage to reduce false positives, and capacity planning to ensure ophthalmology can handle referrals. Unlike DR screening, glaucoma AI cannot operate autonomously because glaucoma diagnosis requires multimodal assessment beyond fundus photography.

Scenario 3: AMD Monitoring Compliance and the Challenge of Daily Testing

Clinical situation: You are a retina specialist. You enroll a 72-year-old woman with bilateral intermediate age-related macular degeneration in the ForeseeHome monitoring program.

Patient background (invented for the exercise): - AMD stage: Large drusen in both eyes (intermediate dry AMD) - Risk: Considered high risk by the treating specialist; an individual annual probability is not assigned in this fictional case - Vision: 20/25 both eyes (excellent currently) - Medicare coverage: Approved (meets criteria)

ForeseeHome setup: - Device delivered to patient’s home - Nurse conducts in-home training (1 hour) - Patient demonstrates competence performing test - Instructed to test daily (3 minutes per day)

First 3 months: - Compliance: 85% (tested 26 days/month average) - No alerts triggered - Patient reports test “easy to use”

Months 4-6: - Compliance drops to 60% (tested 18 days/month) - Inconsistent testing (some weeks daily, some weeks skipped entirely)

Month 7: - Compliance: 40% (tested 12 days/month) - You call patient to discuss compliance

Patient: “I’m sorry, doctor. I start out doing it every day, but then I forget. It’s hard to remember to do it at the same time every day. And honestly, when nothing happens for months, it feels like it’s not necessary.”

You explain: “The test is most valuable when done consistently. If we miss conversion to wet AMD, you could lose vision rapidly. The goal is early detection.”

Patient: “I understand. I’ll try to do better.”

Month 8: - Compliance: 35% - Patient tests inconsistently

Month 10: - Patient presents to your clinic with sudden vision distortion in right eye - Onset 5 days ago, progressively worsening - She did NOT test with ForeseeHome during this time (had not tested for 2 weeks prior to symptom onset)

Exam findings: - Right eye: New subfoveal choroidal neovascularization (wet AMD conversion) - Left eye: Stable intermediate dry AMD - Visual acuity: Right eye 20/80 (down from 20/25)

OCT: - Subretinal fluid, intraretinal fluid - Central foveal involvement

Treatment: - Immediate anti-VEGF injection (ranibizumab) - Plan for monthly injections × 3, then assess

Visual outcome: - After 3 months of treatment: Right eye vision improves to 20/40 (better than presentation, but worse than baseline 20/25) - Permanent mild central vision loss

Question 1: Did ForeseeHome fail, or did the patient fail?

Answer: The device-dependent care pathway failed to maintain engagement and detect conversion before symptomatic presentation. The analysis should avoid blaming the patient and examine both usability and clinical support.

Patient factors: - Testing fell well below the prescribed daily schedule - Did not use device when symptoms started (could have triggered alert)

System factors: - Compliance challenge is inherent to home monitoring devices - Daily testing for months without events creates “alarm fatigue” in reverse (nothing happens, so seems unnecessary) - Device provides no feedback when no alert (feels like “wasted” time)

Comparison: - Taking daily medication has clear rationale (drug effect) - Daily ForeseeHome testing has unclear immediate benefit (only valuable if conversion occurs) - Humans are poor at sustained vigilance for rare events

Key insight: Technology effectiveness depends on human behavior. A device that requires daily compliance for months to years will have compliance issues regardless of how well the technology works.

Question 2: Could this outcome have been prevented?

Possible interventions:

1. Compliance monitoring and outreach:

Current approach: Passive (you noticed compliance drop but intervention was minimal)

Better approach: - Prespecified alerts for a clinically meaningful decline in testing - Timely outreach when the decline crosses that threshold - Identify barriers (forgetting, schedule changes, motivation) - Problem-solve with patient

Example: - Patient forgets to test → Set daily phone alarm reminder - Patient travels frequently → Discuss testing during travel - Patient questions value → Re-education on rapid progression risk

2. Gamification and engagement:

Make testing more engaging: - App shows “streak” of consecutive days tested - Positive feedback for sustained compliance (“You’ve tested 90% of days this month. Great job!”) - Educational content (“Did you know testing takes 3 minutes but could save your sight?”)

3. Simplify testing protocol:

Question: Does daily testing outperform every-other-day or 3x/week?

Current protocol: Daily testing Evidence: HOME study used daily testing, but unclear if less frequent testing would be non-inferior

If 3x/week testing is non-inferior: - Easier compliance (reduce burden from 365 tests/year to 156 tests/year) - Might improve adherence

Requires research to validate

4. Hybrid approach (testing + symptoms):

Educate patient: - “ForeseeHome detects early changes. But if you notice ANY vision changes (distortion, blur, dark spots), call immediately and test with ForeseeHome.” - Symptom recognition as backup for testing non-compliance

In this case: Patient noticed symptoms 5 days before presentation but did not test or call. Symptom education could have prompted earlier presentation.

Question 3: Should ForeseeHome be recommended given compliance challenges?

Answer: Yes, for selected patients, with realistic expectations.

Ideal candidates: - High-risk AMD (bilateral intermediate or unilateral wet) - Cognitively intact - Motivated and technologically comfortable - Able to commit to daily testing - Strong health literacy (understands rationale)

Less ideal candidates: - Cognitively impaired (will forget testing) - Low health literacy (does not understand why testing matters) - Chaotic lifestyle (inconsistent daily routine) - Lacks intrinsic motivation

Pre-enrollment counseling:

“ForeseeHome can detect wet AMD conversion earlier than you would notice symptoms, giving us the best chance to preserve your vision. But it only works if you test every day. That’s 3 minutes every single day, for months or years.

Can you commit to that? If daily testing feels like too much, that’s okay. We have other monitoring options. But if you enroll, consistency is critical.”

Some patients will self-select out, and that’s appropriate.

Question 4: What are alternatives to ForeseeHome for AMD monitoring?

Options:

1. Amsler grid (low cost, but limited and variable sensitivity): - Daily self-testing at home - Patient looks at grid, identifies distortions - Diagnostic performance depends on patient technique, disease stage, and the comparison standard - Compliance likely similar or worse (same daily burden, less engaging)

2. More frequent clinic-based OCT: - Every 3 months instead of every 6-12 months - Higher sensitivity than Amsler grid - No daily patient burden - More expensive (clinic visit + OCT) - Detects conversion later than daily home monitoring

3. Symptom-based monitoring: - Educate patient on wet AMD symptoms (distortion, blur, central scotoma) - Instruct to call immediately if symptoms occur - Rely on patient vigilance - Risk: Symptoms may not appear until advanced conversion

4. Home OCT: - Portable OCT devices for home use - Patient performs OCT weekly or monthly (less frequent than ForeseeHome) - Images transmitted to ophthalmologist - Regulatory status, availability, monitoring workflow, and evidence are product-specific and should be checked before use

No perfect solution. Each has tradeoffs between sensitivity, burden, cost, and compliance.

Lesson: ForeseeHome is effective technology (HOME study demonstrated earlier detection and better visual outcomes), but effectiveness depends on patient compliance. Daily testing for months to years is challenging for many patients. Pre-enrollment counseling should set realistic expectations and assess patient commitment. Compliance monitoring and support (automated alerts, nurse outreach) can improve adherence. Alternative monitoring strategies (more frequent OCT, symptom education) may be appropriate for patients unable to commit to daily testing. Technology alone does not guarantee better outcomes; human behavior is integral to success.

Scenario 4: Reimbursement and Business Model Challenges for Autonomous AI

All financial figures below are invented planning assumptions for the exercise. They are not vendor quotes, payer rates, measured savings, or a current business case. A real analysis requires current contracts, compatible-device costs, staffing, denial rates, referral costs, and the health system’s payment model.

Clinical situation: You are the chief medical officer of a regional health system serving rural and underserved communities. Your primary care network wants to implement IDx-DR diabetic retinopathy screening.

Context: - 15 primary care clinics - 8,000 diabetic patients - Current DR screening rate: 40% - Goal: Increase screening rate to 80%

Implementation costs:

Upfront: - Topcon NW400 fundus camera: $20,000 per clinic × 15 clinics = $300,000 - IDx-DR software licensing: $1,000/month per clinic × 15 clinics = $15,000/month ($180,000/year) - Staff training: $50,000 (one-time)

Total first-year cost: $530,000

Ongoing annual cost: $180,000 (software licensing)

Projected screening volume: - Target: 80% of 8,000 diabetics = 6,400 patients screened per year - Cost per screen (excluding camera amortization): $180,000 / 6,400 = $28 per patient

Reimbursement:

Hypothetical CPT 92229 payment assumptions for the worksheet: - Medicare: $60 per screen - Medicaid: $45 per screen - Commercial insurance: $80-120 per screen (varies by payer)

Your patient mix: - 40% Medicare - 30% Medicaid - 20% Commercial insurance - 10% Uninsured

Weighted average reimbursement: - 0.4 × $60 = $24 - 0.3 × $45 = $13.50 - 0.2 × $100 = $20 - 0.1 × $0 = $0 - Total: $57.50 per screen average

Revenue projection: - 6,400 screens × $57.50 = $368,000/year

Cost: - Software: $180,000/year - Camera amortization (5-year): $300,000 / 5 = $60,000/year - Staff time (0.15 FTE per clinic for imaging and coordination): 15 × 0.15 × $50,000 = $112,500/year - Total cost: $352,500/year

Net margin: $368,000 - $352,500 = $15,500/year profit (4% margin)

Your reaction: “We’re spending $530,000 upfront and breaking even annually to improve diabetic eye screening. Is this worth it?”

Question 1: Is this a good investment?

Answer: Depends on how you define “good.”

Financial result under the invented assumptions: Barely positive (4% modeled margin after break-even in year 5)

Not a lucrative business case. Health system investing $530,000 for $15,500 annual profit is weak financial justification.

Clinical and population health ROI: Potentially excellent

Value-based care perspective: - Earlier detection and treatment may avoid visual impairment, disability, and downstream care, but the chapter’s fictional worksheet does not estimate those outcomes - A defensible model would use locally measured screening uptake, disease prevalence, referral completion, treatment effectiveness, quality-adjusted outcomes, and uncertainty - Do not assume that a screening program prevents vision loss in a fixed percentage of screened patients

HEDIS quality metrics: - Diabetic retinopathy screening is HEDIS measure (Comprehensive Diabetes Care) - Health plans penalize/reward based on screening rates - Improving from 40% to 80% screening substantially improves HEDIS scores - Quality-measure performance may affect contract payments, but the terms and amounts are contract-specific

ACO shared savings: - If health system is in ACO (Accountable Care Organization), preventing diabetic complications generates shared savings - Any shared-savings estimate should be modeled from the actual attributed population, contract, downstream utilization, and counterfactual care pathway

Community benefit and mission: - FQHC mission is to serve underserved populations - Reducing preventable blindness aligns with mission - Intangible value (reputation, community trust)

Conclusion: Weak financial ROI, strong clinical and value-based care ROI

If your health system is: - Fee-for-service only, focused on short-term revenue → Marginal investment - Value-based care, at-risk contracts, HEDIS-focused → Strong investment - Mission-driven (FQHC, safety-net) → Aligned with mission

Question 2: How could the business model be improved?

Strategies to improve financial sustainability:

1. Negotiate better reimbursement:

  • Commercial payers: Verify coverage and negotiate rates using current payer contracts
  • Medicaid: Advocate for rate increase (Medicaid often undervalues preventive services)
  • Bundled payment: Negotiate diabetic care bundles including DR screening

2. Reduce costs:

Software: - Negotiate multi-year contract with IDx-DR (discount for volume and commitment) - Explore EyeArt (may have lower licensing fees, works with multiple cameras)

Hardware: - Portable fundus cameras (lower cost than Topcon NW400) - Shared equipment model (1 camera per 2-3 clinics, rotate)

Staffing: - Cross-train existing MAs (no dedicated imaging staff)

Cost-reduction target: Set a target only after obtaining current quotes and determining which costs are fixed, variable, and clinically necessary

3. Increase revenue:

Higher screening volume: - Screen non-diabetic patients opportunistically (hypertensive retinopathy, other conditions) - Bill additional CPT codes (fundus photography 92250 for other indications)

Ancillary revenue: - Contract with external primary care practices (provide DR screening service for fee) - Become regional DR screening hub

4. Alternative payment models:

Grant funding: - HRSA grants for FQHC diabetes programs - State public health department funding for preventive services - Philanthropic support

Population health payments: - Health plan contracts: Fixed per-member-per-month for diabetic care management (includes DR screening)

Question 3: What if reimbursement does not cover costs?

Scenario: Medicaid-dominant patient population (60% Medicaid, 30% uninsured, 10% Medicare)

Revised reimbursement: - 0.6 × $45 = $27 - 0.3 × $0 = $0 - 0.1 × $60 = $6 - Total: $33 per screen average

Revenue: 6,400 × $33 = $211,200/year Cost: $352,500/year Loss: -$141,300/year

This model is not financially sustainable without subsidy.

Options:

1. Seek grant funding to cover gap: - HRSA, state public health, philanthropy - $140,000/year subsidy needed

2. Reduce deployment scope: - Deploy at 5 highest-volume clinics instead of 15 - Reduce fixed costs proportionally

3. Advocate for policy change: - Medicaid reimbursement increase for DR screening - State-level preventive care funding

4. Accept loss as mission-aligned: - If preventing blindness is core mission, subsidize program from other revenue

Question 4: Why is autonomous AI reimbursement challenging?

Structural barriers:

1. The service differs from traditional interpretation: - CPT 92229 describes retinal imaging with remote autonomous analysis and report - Coverage and payment policy determine whether and how the service is reimbursed - A billing code does not guarantee payment or define the clinical value of the complete pathway

2. Preventive services undervalued: - Fee-for-service rewards treatment, not prevention - Preventing vision loss 10 years in future does not generate immediate revenue - Misaligned incentives

3. Payer fragmentation: - Medicare, Medicaid, commercial payers have different reimbursement rates - High administrative burden to negotiate with each payer

4. Cost-effectiveness uncertainty: - Long-term outcomes (vision preservation) not yet proven in real-world deployment at scale - A 2026 Markov model finds that ECP can dominate in integrated systems at modest willingness-to-pay; this does not close the vision-outcome gap (Ahmed et al., 2026) - Payers hesitant to pay for unproven ROI

Policy solutions:

1. Clear, current payment policy: - Payers can publish transparent coverage, documentation, and payment requirements for 92229 - Medicaid and commercial policy remains jurisdiction- and plan-specific; the chapter does not propose a universal legal rate mandate

2. Value-based payment: - Pay for outcomes (diabetics screened, vision loss prevented) not just services delivered - Incentivize prevention

3. Bundled payments: - Diabetic care bundles include DR screening - Simplifies payment model

Lesson: Financial sustainability of autonomous AI depends on reimbursement rates, patient payer mix, deployment costs, and health system incentive structure. In fee-for-service environments with low Medicaid/uninsured populations, margins are thin or negative. In value-based care, ACO, or HEDIS-driven systems, clinical and quality ROI justifies investment even if direct revenue is modest. Health systems must evaluate not just technology performance but business model viability before deployment. Policy advocacy for adequate reimbursement and value-based payment models is essential for widespread adoption.


Key Takeaways

Clinical Bottom Line for Ophthalmology AI

1. Diabetic retinopathy screening is a leading autonomous AI use case. IDx-DR/LumineticsCore and EyeArt have product-specific prospective and regulatory evidence. Randomized workflow evidence shows that immediate point-of-care testing can close a screening gap in a defined pediatric trial, while observational evidence remains vulnerable to confounding. The translated pattern is narrow scope, standardized input, a clear need, an actionable output, and a complete referral pathway.

2. Success requires functional care pathways, not just technology. Screening without accessible ophthalmology follow-up causes harm. Before deploying DR screening, ensure referral capacity, care coordination, transportation support, and insurance navigation are in place.

3. Glaucoma image analysis is a triage signal, not a comprehensive diagnosis. Fundus-photo analysis cannot replace assessment of intraocular pressure, visual fields, angle anatomy, OCT, and progression. Positive predictive value depends on population, threshold, and reference standard.

4. AMD home monitoring works but requires patient participation. In the HOME trial, ForeseeHome plus standard care was associated with less visual-acuity loss at conversion detection (median -4 versus -9 letters), but qualification, daily testing, alert handling, and prompt examination are part of the intervention. The device is remote hyperacuity monitoring and should not automatically be categorized as machine-learning evidence.

5. ROP screening AI shows promise but needs infrastructure. i-ROP achieves 93% sensitivity for plus disease detection, but deployment limited by RetCam availability, ophthalmologist confirmation requirements, and medicolegal risk.

6. Screening is not diagnosis. AI systems identify who needs further evaluation, not what treatment they need.

7. OCT analysis AI is most successful when integrated into imaging devices. Automated retinal layer segmentation, diabetic macular edema quantification, and treatment response monitoring work well when embedded in clinical workflow.

8. Reimbursement challenges affect deployment. CPT 92229 enables billing for autonomous AI DR screening, but rates vary by payer. Thin margins in Medicaid/uninsured populations require value-based care models, grants, or mission-aligned subsidies.

9. Use the formal AAO condition-specific guideline. The Diabetic Retinopathy Preferred Practice Pattern addresses validated digital imaging within the broader care pathway. It should be read directly rather than reduced to a product list.

10. Evidence does not transfer across ophthalmic tasks. Standardized imaging, defined criteria, access barriers, an actionable output, and governed failure pathways helped diabetic-retinopathy screening translate. They do not establish autonomy for glaucoma, ROP, OCT treatment selection, or comprehensive eye care.

Questions About AI in Ophthalmology

IDx-DR, now LumineticsCore, received FDA De Novo authorization in 2018 for autonomous diabetic retinopathy screening within a defined adult population, camera, operator, and workflow. In the pivotal prospective study, sensitivity was 87.2%, specificity was 90.7%, and the system produced a diagnostic output for 96.1% of participants.

A fundus-photograph classifier can identify suspicious structural features, but it does not by itself establish glaucoma diagnosis, progression, or treatment need. Clinical assessment commonly integrates intraocular pressure, visual fields, angle anatomy, OCT, and change over time.

CPT code 92229 describes retinal imaging with remote autonomous analysis and report. Coverage, payment, eligible devices, and billing requirements vary by payer and jurisdiction and should be verified before deployment.

The translated use case is narrow: standardized fundus imaging, a defined target condition, an immediate refer-or-rescreen output, prospective evaluation in the intended setting, and a specified pathway for positive or ungradable results. These features do not make failures harmless or support autonomy for other eye diseases.