Learning objectives#
- Separate screening eligibility, adequate image acquisition, an ungradable result, a negative result, a positive result, and a clinical diagnosis.
- Interpret sensitivity, specificity, predictive values, prevalence, reference standards, and ungradable-image handling within the actual care pathway.
- Verify intended use, labeling, software version, operator workflow, human responsibilities, result communication, and referral closure before relying on an AI-enabled screen.
- Respond to eye symptoms, image-quality failure, out-of-scope patients, or discordant findings through clinical assessment rather than automated reassurance.
- Audit access, subgroup performance, privacy, overrides, and missed follow-up as patient-safety outcomes rather than treating deployment as a one-time technical event.
Initial presentation#
A 61-year-old adult with type 2 diabetes attends a primary-care visit for medication reconciliation and preventive care. The electronic record shows no completed retinal examination in the past twenty months. A prior referral was closed after the patient missed an appointment at an eye clinic seventy miles away. The patient works seasonal jobs, no longer drives at night, and cares for a spouse with limited mobility. The new primary-care site offers retinal photography with an AI-enabled screening system during the same visit.
The medical assistant describes the service as a picture that can check for diabetic eye disease. The patient asks whether it replaces an eye doctor and whether the images will be used to train software. These questions expose two risks before a photograph is taken. Screening may be mistaken for comprehensive eye care, and privacy or secondary-use expectations may be unclear. The clinician explains that this is a bounded screen for a defined finding in eligible people, not a complete eye examination, and reviews how images and results are stored and used under the clinic's current policy.
Eligibility is checked against the current device labeling and local protocol rather than inferred from the diagnosis code alone. The patient is within the intended adult population, has a documented diabetes diagnosis, has not previously been diagnosed with diabetic retinopathy, and is not under active ophthalmology management for it. There has been no prior retinal laser treatment or eye injection. The record does list cataracts noted several years earlier, which may affect image quality but does not automatically settle eligibility. The software version and camera configuration match the authorized local setup.
The symptom screen changes the conversation. The patient denies new loss of vision, a curtain or shadow, flashes, a sudden burst of floaters, severe pain, marked redness, trauma, double vision, or a severe headache with visual change. Reading has become gradually harder over two years, and glare at night has increased. Those chronic concerns still warrant eye-care assessment, but there is no current acute symptom requiring an emergency pathway. The clinician documents that screening does not evaluate every possible cause of those symptoms.
The patient uses a wheelchair for longer distances and transfers to a stable chair with armrests. The camera station has a fixed-height table, so staff reposition the camera on an adjustable platform rather than asking the patient to stand. Instructions are given in the patient's preferred language with a qualified interpreter. The operator confirms that the patient can rest the chin and forehead without neck pain. The room can be darkened, and the operator follows the acquisition sequence in the validated protocol.
The right-eye images pass the camera's quality check. The left-eye images show glare and poor visualization through a small pupil and probable lens opacity. The operator repeats the images according to protocol after allowing time for dark adaptation and correcting alignment. The automated system then returns a result that the right eye does not meet its referral threshold and the left eye is insufficient for analysis. A banner in the record displays the overall encounter as incomplete.
During a busy afternoon, a staff member suggests documenting the screen as negative because one eye passed and no abnormality was flagged. Another proposes taking repeated images until the system gives a result. The clinician stops both shortcuts. An ungradable image is not negative. Repeated acquisition beyond the validated and tolerated workflow can waste time, cause discomfort, and create selective acceptance of the first output that looks convenient. The result must be interpreted as an incomplete bilateral screening episode, with a human-owned next step.
The local protocol says an insufficient-quality result should lead to a comprehensive dilated eye examination. The nearest appointment in the usual network is nine weeks away. A mobile eye clinic visits the county twice monthly, but its referral pathway is not connected to the electronic order. The patient can attend only if transportation and caregiver coverage are coordinated. The technical output therefore becomes a systems case: can the clinic convert image-quality failure into timely, accessible care and verify that the loop closes?
Problem representation#
This is an adult with diabetes and overdue retinal evaluation who is eligible for a locally authorized AI-supported screening workflow and has no acute eye warning symptom. Adequate right-eye images generated a below-threshold screening output, while repeated left-eye acquisition remained ungradable because of glare, a small pupil, and possible media opacity. The bilateral screen is incomplete, not negative, and it does not diagnose or exclude diabetic retinopathy or other eye disease.
The risk is not merely algorithm error. Failure can occur at eligibility review, consent and explanation, positioning, or image acquisition. It can occur at quality assessment, result display, or clinician interpretation. It can also occur at referral creation, appointment access, result return, or longitudinal audit. The patient's gradual glare and reading difficulty may reflect cataract, refractive change, retinal disease, or another process outside the automated target. Distance, transport, and caregiver duties all affect whether the theoretical benefit becomes care. So do language, physical accessibility, and data-use concerns.
The next decision is a human clinical decision: arrange an eye examination through a feasible closed loop, maintain symptom-based safety instructions, and document the incomplete screening state. At the program level, the clinic must know how often images are ungradable, who is affected, whether repeats and referrals are completed, and whether performance remains acceptable after changes in software, camera, operators, or patient mix.
Prioritized differential#
1. Ungradable image caused by media opacity, small pupil, alignment, or acquisition conditions#
The most direct explanation is inadequate retinal visualization. Cataract can scatter light and obscure detail. A small pupil, corneal opacity, dry ocular surface, and eyelid position can also contribute. So can fixation difficulty, tremor, limited neck movement, and camera alignment. Lighting, focus, operator technique, and device cleanliness can contribute too. The system's quality-control outcome is useful only if it leads to an action rather than being hidden.
An image-quality failure is not itself a diagnosis of cataract. The patient needs an eye examination that can assess the lens, retina, and optic nerve. That examination can also assess pressure, visual acuity, and other relevant structures. Nor should staff assume that poor quality is a patient failure. Station design, operator training, time pressure, and accessibility can shape gradability.
2. Diabetic retinopathy not assessable in the left eye#
Retinopathy remains possible because the retina was not adequately visualized. Duration of diabetes, glycemic history, and blood pressure affect risk. So do kidney disease, pregnancy status where relevant, and prior eye findings. But those factors cannot convert an ungradable study into a result. Vision can remain good despite clinically important retinopathy, so lack of symptoms is not sufficient reassurance.
The right-eye below-threshold output lowers but does not eliminate the chance that the target disease is present in that eye. No screening test has perfect sensitivity. The left eye has no interpretable automated result at all. A bilateral disease process cannot be inferred from one eye's output.
3. Cataract or refractive change causing gradual glare and reading difficulty#
The chronic, slowly progressive glare and near-vision difficulty may be explained by lens opacity or refractive change. Diabetes can be associated with lens and refractive changes, but an eye examination is needed. These concerns lie outside a narrow retinopathy screen even if image quality were adequate. Improving one target-disease workflow should not narrow clinical attention to that target alone.
4. Other retinal, optic-nerve, or ocular disease outside the automated target#
Age-related macular disease, retinal vascular occlusion, glaucoma, and optic neuropathy may be present with or without diabetic retinopathy. So may epiretinal membrane, corneal disease, and other conditions. Whether a system can incidentally flag some findings does not expand its authorized intended use. An unvalidated secondary observation should not be promoted as a diagnosis or used to defer care.
5. Acute eye or neurologic disease if symptoms emerge#
Sudden vision loss, a new field defect or curtain, flashes with many new floaters, severe pain, marked redness, trauma, diplopia, acute neurologic symptoms, or severe headache changes the pathway. Retinal detachment, vascular occlusion, acute glaucoma, infection, inflammation, stroke, and other urgent conditions require diagnostic triage. A recent negative screen would not rule them out because screening is neither designed nor timed for those presentations.
6. Workflow or data-integrity failure#
The source image might be assigned to the wrong eye or person, corrupted in transfer, processed by an unapproved software version, or separated from the clinical context. A display may show only the analyzable eye, while a summary field incorrectly implies overall completion. An interface can drop the referral recommendation. Human review of identifiers, laterality, image count, and quality status is therefore part of safe use. That review also covers software version, output, and downstream order.
7. Performance shift in the local population or setting#
Published validation performance may not reproduce in a clinic with different disease prevalence, camera conditions, operators, or age distribution. It may not reproduce where cataract burden, pupil characteristics, comorbidity, or follow-up pathways differ. Real-world head-to-head studies have shown substantial performance variation among systems and sites. A product can satisfy an authorization standard and still require local monitoring within its intended use.
Focused history and examination#
The clinician first determines whether the encounter is screening or diagnosis. Ask about sudden or progressive visual change, field loss, flashes, and floaters. Ask about pain, redness, discharge, and photophobia. Then ask about trauma, diplopia, and headache. Ask about neurologic symptoms and prior eye diagnoses or procedures. New acute symptoms bypass the routine automated screen. Chronic symptoms may allow a nonemergency pathway but still require a diagnostic eye examination.
Diabetes history includes type, duration, glycemic trajectory, and blood pressure. It includes kidney disease, pregnancy or plans where relevant, and tobacco use. It also includes prior retinal screening and any earlier retinopathy. The medicine list is reviewed for diabetes treatment and agents that may affect ocular care or procedural planning, but the screen is not used to adjust treatment by itself. Secondary prevention is coordinated through the patient's clinicians without supplying a generic dosing plan.
Eye history covers glasses or contact lenses, cataract, glaucoma, and macular disease. It covers retinal vascular disease, surgery, injections, and laser treatment. It also covers poor dilation, light sensitivity, and family history. Prior records are retrieved because known disease may place the patient outside a screening-only pathway. The team also checks whether an ophthalmology appointment is already pending so that automated screening does not fragment existing care.
Access history is operational. Can the patient approach and align with the camera? Is seated imaging possible? Are rests needed? Can instructions be heard, seen, and understood? Does an interpreter or communication aid need to remain present? Can the person tolerate a dark room? Are there cultural or privacy concerns about photographs? Does the referral destination accommodate mobility, sensory, cognitive, and language needs?
The primary-care examination is proportionate. Visual acuity is measured with correction when feasible. Gross pupils, fields, and eye movements are assessed according to symptoms. So are external structures and neurologic status. A normal gross examination does not validate an ungradable retinal image. Conversely, an acute abnormality prompts diagnostic escalation without waiting for a software output.
Before acquisition, the operator checks device cleanliness, calibration status where applicable, and network connection. The operator checks patient identity, eye laterality, and camera settings. The operator also checks software version and the specific acquisition protocol. The protocol determines number and field of images, whether dilation is permitted, how many repeat attempts are reasonable, and what constitutes an analyzable encounter. Local improvisation can move the workflow outside the evidence base.
After acquisition, a human checks that both eyes are represented and that images match the intended fields. The human also checks that quality messages are visible and that the output transfers correctly to the record. The operator does not interpret retinal pathology beyond training and role. The responsible clinician reads the final output in context, confirms the next action, communicates it, and ensures the order exists.
Diagnostic strategy#
Establish the intended-use boundary#
The clinic maintains the current labeling, authorization information, user instructions, validated camera and software configuration, eligible population, contraindications or exclusions, target condition, output categories, quality-failure pathway, and required operator training. Inclusion on an FDA list is not an endorsement and does not mean the tool can be used for every retinal question. The exact version matters because performance evidence and authorized claims attach to a specific system state.
This case meets the local eligibility criteria, but the left image does not meet quality requirements. The correct coded disposition is incomplete or ungradable according to the system's language. A normal or below-threshold right-eye output remains documented, yet the overall screening need is not satisfied.
Interpret performance as a pathway, not one headline number#
Sensitivity is the proportion of people with the target condition, as defined by the study reference standard, who receive a positive screening result. Specificity is the proportion without that condition who receive a negative result. Neither number tells the patient the chance that this particular result is correct without considering disease prevalence, population, thresholds, image quality, and how ungradable studies were handled.
Predictive values change with prevalence. In a low-prevalence setting, even a reasonably specific test can produce many false positives relative to true positives. In a higher-risk population, the negative predictive value may be lower even if sensitivity and specificity are unchanged. The referral threshold also determines which severity of retinopathy counts as target disease. A study detecting any retinopathy is not directly interchangeable with one detecting more-than-mild or vision-threatening disease.
Ungradable images create a denominator problem. If a performance report excludes poor-quality images, headline sensitivity and specificity apply only after successful acquisition. The program still owes care to everyone whose study failed. Counting ungradable studies as negative is unsafe. Treating all as positive may protect sensitivity but increases referrals and burden. The labeled and clinically appropriate pathway should be explicit, and the ungradable rate should be reported separately.
Reference standards matter. Human graders can disagree, and grading a limited photograph is not identical to a comprehensive eye examination. Prospective and real-world studies vary in camera, dilation, and image fields. They vary in disease definitions, populations, and whether the same images were used for both system and reference grading. The clinician therefore avoids transferring one study's number into a promise to this patient.
Choose the next diagnostic action#
Because the left eye remains ungradable after protocol-compliant attempts and the patient has gradual visual concerns, the clinic arranges a comprehensive eye examination rather than scheduling endless repeats. If the labeling allowed a repeat with a trained operator or permitted dilation and a timely repeat were feasible, that could be a branch for an asymptomatic patient. Dilation has contraindications, risks, workflow requirements, and scope-of-practice considerations, so it is not improvised.
The referral communicates diabetes history, the ungradable left image, and the right-eye output. It communicates chronic glare and reading difficulty, accessibility needs, and language. It also communicates transport constraints and how results should return. The eye clinician determines diagnostic examination and treatment. The primary-care team does not claim cataract or retinopathy from image quality alone.
Close the loop#
An order is not completion. The system records referral date, urgency, scheduled appointment, and barriers. It records reminder route, attendance, and returned report. It also records diagnosis and next interval. One person or team owns the queue. If the patient cannot attend the distant clinic, staff connect the mobile clinic through a manual but tracked pathway and arrange accessible transport and caregiver respite support.
Safety instructions remain active while the appointment is pending. Any new acute visual or neurologic symptom triggers same-day or emergency assessment regardless of the queue status.
Progressive results and interpretation#
The clinic's referral coordinator reaches the mobile eye service and secures an appointment in twelve days. The patient chooses that option because it reduces travel and offers an accessible examination chair. Transportation is arranged, and the primary-care clinic sends the screening report and relevant history through an approved channel. The appointment is visible on the clinic's unresolved-referral dashboard until a result returns.
The comprehensive examination finds no referable diabetic retinopathy in the right eye, mild nonproliferative changes in the left eye below the treatment threshold, and a visually important left cataract that likely contributed to the ungradable photograph. The eye clinician also updates refraction and discusses cataract options. The report includes a follow-up interval based on the full examination and diabetes context.
This result illustrates two separate screening limitations. The below-threshold right output was concordant with the eye examination, but that does not make every future negative certain. The ungradable left output concealed both retinopathy and cataract. The automated system did not miss a positive analyzable image; it correctly declined to interpret poor input. The unsafe event would have been a human decision to relabel that decline as negative.
The clinic reviews the previous quarter after this case. Eleven percent of attempted encounters had at least one ungradable eye. Completion after an ungradable output was only fifty-eight percent within sixty days. Rates were higher among older adults, people with documented cataract, and people needing positioning accommodations. The audit cannot determine from these observations whether the algorithm itself performs differently by demographic group because acquisition, referral access, and small subgroup sizes are intertwined. It can establish that the pathway produces unequal completion and needs correction.
The record also shows that two clinicians overrode incomplete status to negative in separate encounters. Both cases are reviewed promptly, affected patients are contacted, and the display logic is changed so that one insufficient eye cannot populate an overall negative field. Operator retraining alone would be inadequate because the interface invited the error.
After workflow changes, the clinic tracks gradability, referral creation, and completion. It breaks those down by age, language, disability accommodation, and site. It also breaks them down by operator, camera, software version, and other ethically justified categories with adequate governance. Small numbers are suppressed in broad reports to reduce reidentification risk. The goal is to locate failure points and improve care, not to rank patients or blame staff.
Management plan#
Communicate the result precisely#
The patient hears three distinct statements. The right eye produced a screening result below the system's referral threshold. The left eye could not be evaluated by the system. The overall screen is incomplete, and an eye examination is needed. The clinician adds that the screen does not assess every eye condition and that chronic glare deserves evaluation even apart from diabetes.
The after-visit summary avoids the words normal, clear, passed, or no eye disease. It names the eye, the quality problem, the next appointment, and urgent symptoms. The interpreter remains for teach-back. The patient explains in return that one image was not readable, an eye appointment is still needed, and sudden vision change should not wait.
Coordinate diabetes and vascular risk care without overclaiming the screen#
Primary care reviews glycemic management, blood pressure, lipids, and smoking status as part of comprehensive diabetes care. It also reviews kidney health, nutrition, activity, and medicine adherence. These decisions use the whole clinical picture and current guidelines, not the automated output alone. The case does not provide a personal prescription or exact target.
The later eye report is reconciled into the diabetes plan. A named clinician documents the next recommended retinal assessment date and whether follow-up belongs with eye care or screening. Cataract evaluation and retinopathy surveillance are kept distinct so that one completed task does not close the other.
Maintain human oversight at every handoff#
The operator owns correct identity, laterality, acquisition, and quality workflow within training. The ordering clinician owns eligibility, symptom triage, interpretation in context, and the clinical action. The referral coordinator owns scheduling and barrier resolution. The eye clinician owns diagnostic examination and eye treatment. The digital-safety team owns version control, interface validation, and performance monitoring. It also owns incident response and change management. Shared work still requires one named owner at each step.
Overrides are allowed only through a documented pathway. A clinician may escalate despite a below-threshold output because of symptoms or another concern. A clinician should not downgrade an ungradable or positive result merely to reduce referrals. Every override records who, why, what action followed, and whether later findings reveal a safety signal.
Design for quality rather than repeated convenience sampling#
Operators receive competency-based training in positioning, focus, and lighting. The training also covers image fields, accessibility, and quality messages. The clinic monitors retraining needs by pattern. It also examines equipment height, room darkness, and visit time. It examines camera maintenance and staffing. If one site has high failure across operators, a system fix is more likely than individual blame.
Repeat attempts follow the validated protocol and stop when further acquisition is unlikely to help or burdens the patient. The system must preserve every required quality state. Staff cannot select only the image or output that gives the desired answer. Image deletion, replacement, and laterality correction require an audit trail.
Protect privacy and data integrity#
Retinal images are health data linked to identity and clinical context. The clinic explains collection, storage, and access. It also explains retention, result transfer, and any secondary use under applicable policy. Care-related authorization is not silently converted into permission for unrelated training, marketing, or research. De-identification reduces but may not eliminate risk, especially when images or linked data are rich.
Access follows role and need. Data are encrypted in transit and storage according to the organization's controls. Vendor or service-provider agreements define responsibilities, incident reporting, and subcontractors. Those agreements also define deletion and return. The team avoids promising that de-identified means impossible to reconnect. Privacy, security, and legal personnel review the actual workflow; this educational case does not offer legal advice.
Identity and laterality checks occur before acquisition and after interface transfer. Downtime procedures define whether imaging pauses, stores locally, or shifts to referral. No one copies images to a personal device to work around a network problem. A cybersecurity or data-mismatch event triggers the organization's incident process and clinical review of affected results.
Make referral access part of test performance#
The clinic reserves pathways for urgent symptoms, positive screens, and ungradable studies. It maps capacity before scaling acquisition. Transportation, physical access, interpreter availability, and cost are assessed at the point of referral. So are caregiver duties, work schedules, and digital access. A mobile service is useful only if orders, reports, and urgent escalation connect reliably.
When capacity is constrained, the team does not solve the problem by redefining incomplete as negative. It triages according to clinical urgency, communicates delays, offers safe alternatives, and records residual risk. Screening volume without follow-up capacity can increase documented activity while leaving disease untreated.
Govern updates and model drift#
The software version, camera, protocol, and key configuration are recorded for each result. Before an update, the clinic reviews what changed and whether authorization and labeling still apply. It reviews what validation supports the change, how interface fields behave, and whether local verification is required. A silent update that alters performance or output categories is a clinical change.
Ongoing monitoring includes data completeness, gradability, and positive and negative distributions. It includes overrides, referral completion, and later diagnostic concordance when available. It also includes incidents, complaints, and subgroup outcomes. Control limits and response owners are defined in advance. A dashboard without an action threshold is observation, not governance.
Escalation, referral, and safety net#
The patient is told to seek emergency assessment for sudden vision loss, a new curtain or major field defect, severe eye pain with visual change, significant trauma, acute neurologic symptoms, or other rapidly evolving concern. New flashes or many floaters, marked redness, or photophobia requires prompt same-day eye triage according to local pathways. So does painful vision change or a sudden reduction in vision. Screening appointments and portal messages are not substitutes for emergency evaluation.
An ungradable image requires a time-bound repeat or diagnostic eye examination according to intended use and context. A positive screen requires a closed referral at the recommended urgency. A below-threshold result still requires the guideline-consistent future screening interval and symptom safety net. Known retinopathy, prior treatment, pregnancy-related risk, or another out-of-scope condition may require direct specialty care rather than automated screening.
At the program level, escalation occurs when ungradable rates rise, output distributions shift unexpectedly, one group has worse completion, interface errors alter status, data are assigned to the wrong person or eye, referrals fail, or later diagnoses suggest missed disease. The digital-safety lead can pause use while clinical owners contact affected patients and create an alternative pathway. Continuing because the tool is convenient is not a safety response.
The clinic also has an emergency plan for service loss. If the camera or software is unavailable, overdue patients are not left in an invisible queue. They receive standard eye referrals or a documented rescheduling pathway, with urgency based on symptoms and history.
Communication, shared decisions, and equity#
The service is introduced without hype. The clinician says that software analyzes specific retinal photographs for a defined screening purpose and that a clinician remains responsible for deciding what the output means and what happens next. The patient may ask questions or choose a standard eye referral. Refusal of automated screening does not reduce access to usual care.
The patient receives understandable information about possible results before imaging: below threshold, referral indicated, or insufficient image quality. This prevents the ungradable state from feeling like an unexpected personal failure. Staff explain that small pupils, cataract, positioning, and equipment factors can affect quality and that the next step is still useful care.
Equity review separates several layers. Does the camera fit different bodies and mobility needs? Are instructions accessible across language, vision, hearing, cognition, and literacy? Are quality failures more common in groups with cataract or small pupils? Does the algorithm perform comparably on analyzable images across clinically relevant subgroups? Can referred patients reach eye care? Does insurance or cost alter completion? Combining all of these into one fairness score would hide actionable causes.
Subgroup analysis is preplanned, privacy-protective, and interpreted with sample size and uncertainty. A small observed difference does not prove bias, while absence of a statistically clear difference does not prove equity. The team examines confidence intervals, missingness, quality failure, outcome definition, and downstream completion. Community and patient representatives help interpret burdens and acceptable alternatives.
Teach-back covers both clinical and data understanding. The patient states what the system did, what it did not establish, which eye was ungradable, why the eye visit remains necessary, what urgent symptoms mean, and how images are handled under the clinic's policy. The clinic provides a phone route as well as a portal because digital-only follow-up would reproduce the access problem.
Follow-up and contingencies#
The referral queue is reviewed weekly. An unreturned result remains open until the eye service report is received and reconciled. If the patient misses the appointment, the coordinator calls using the preferred method and asks about new symptoms. The coordinator identifies the barrier and reschedules or offers another pathway. A missed visit is not silently recorded as patient refusal.
After the diagnostic examination, the primary-care clinician communicates the findings in plain language, updates the disease list only from confirmed diagnoses, and records ownership of follow-up. If cataract treatment proceeds, retinal surveillance does not disappear. If retinopathy is found, the eye clinician's interval and treatment plan are coordinated with diabetes care.
If repeat imaging later becomes an option, the clinic verifies that the patient is still eligible and asymptomatic and that the system version and protocol remain valid. A prior ungradable image is a reason to consider whether another acquisition attempt is useful, not an obligation to repeat the same failed process.
The quality team reviews monthly rates and quarterly trends. Any version change creates a new monitoring annotation so pre-change and post-change results are not blended. The denominator includes every attempted eligible screen, not only images the system analyzed. Reports show incomplete, positive, and below-threshold as distinct states. Referral ordered, appointment completed, and diagnosis returned are distinct states too.
If local sensitivity cannot be calculated because most negative screens do not receive a reference examination, the limitation is stated. The clinic can monitor later diagnoses, incident reports, sampled concordance where ethically and operationally appropriate, and pathway outcomes, but it should not invent a local accuracy estimate. If a safety signal emerges, prospective evaluation and external expertise may be needed.
Reasoning traps and alternative pathways#
Trap: treating ungradable as negative. Quality failure means the target determination was not made. Preserve the incomplete status and initiate the labeled clinical path.
Trap: calling screening a diagnosis. A positive output identifies a referral need; a below-threshold output does not certify a healthy eye. Both remain within a bounded intended use.
Trap: quoting sensitivity without the denominator. Disease definition, reference standard, and prevalence determine what the number means. So do image-quality exclusions, population, camera, and workflow.
Trap: optimizing accuracy while ignoring referral closure. A technically strong system cannot prevent vision loss if positive and incomplete results never reach eye care.
Trap: assuming authorization equals endorsement or universal fit. Authorization is tied to evidence, labeling, configuration, and intended use. Local implementation still requires governance.
Trap: endless repeat imaging. Repeating until a convenient result appears can introduce selection and delay. Follow the protocol and move to diagnostic care when quality remains inadequate.
Trap: blaming the patient for image quality. Fixed camera height, time pressure, and poor positioning support all contribute. So do operator training, pupil size, and media opacity. Improve the system.
Trap: using one overall fairness metric. Acquisition access, gradability, and algorithm performance are separate questions with different remedies. So are referral completion and treatment access.
Alternative pathway: acute symptoms. Stop screening and use urgent clinical triage or emergency eye care. A software output cannot overrule the presentation.
Alternative pathway: known retinopathy or prior treatment. Route to ongoing diagnostic eye care according to the patient's established plan rather than using a screening-only system outside scope.
Alternative pathway: interface or identity failure. Quarantine the result, contact affected patients when indicated, and use a safe alternate pathway. Investigate the incident and verify correction before resuming routine use.
Evidence limits and what could change#
Evidence for AI-supported diabetic-retinopathy screening includes prospective studies, retrospective image sets, and real-world comparisons. It also includes meta-analyses, regulatory review, and implementation reports. Performance varies with the target definition, reference grading, and population. It varies with camera, dilation, operator, and handling of ungradable images. Some studies use enriched data sets or exclude poor images, while real clinics must manage every attempted encounter.
Sensitivity and specificity are not permanent properties of a broad concept called AI. They attach to a particular system, threshold, version, inputs, and use conditions. Predictive values change with prevalence. Human reference grading can disagree, and photographs do not reproduce every element of a comprehensive examination. Later diagnostic concordance is useful but may be affected by who receives follow-up.
Real-world literature reports that image-quality failure is material and that system performance can vary across algorithms and settings. Newer studies may improve estimates, but a high aggregate number does not answer whether the current clinic can acquire analyzable images equitably or close referrals. Local audit is therefore complementary to external validation, not a replacement for it.
Privacy and fairness evidence also has limits. De-identification methods reduce risk but do not justify casual reuse. Subgroup estimates may be unstable, and categories may be socially or clinically incomplete. Differences can arise before, within, or after the algorithm. Transparent uncertainty and patient involvement are necessary.
This case would change with any acute symptom, a new prior eye diagnosis, pregnancy-related risk, a different authorized use, a software or camera update, an interface incident, a higher local failure rate, or loss of eye-care capacity. The safe response is to revalidate the pathway, not to preserve the technology at all costs.
Key points#
- An ungradable retinal image is an incomplete screen, never a negative result.
- Screening output is not a diagnosis and does not replace symptom triage or comprehensive eye care.
- Sensitivity, specificity, predictive values, reference standards, prevalence, and image-quality handling must be interpreted together.
- Human accountability spans eligibility, acquisition, interpretation, referral, result return, privacy, version control, and incident response.
- Accessibility, gradability, subgroup performance, and referral completion are separate equity outcomes that all require audit.
- A screening program succeeds only when every positive or incomplete result has a feasible, verified path to care.
For your own health, talk with your clinician.*
Sources and further reading
- ADA Standards of Care in Diabetes 2026, Retinopathy
- Artificial Intelligence and Diabetic Retinopathy, Prospective and Real-World Evidence
- Systematic Review of Artificial Intelligence for Diabetic Retinopathy Screening
- Multicenter Real-World Validation of Automated Retinopathy Screening Systems
- Primary-Care Validation Including Poor-Quality Retinal Images
- Meta-Analysis of Artificial Intelligence for Diabetic Retinopathy Screening
- FDA Artificial Intelligence-Enabled Medical Devices
- FDA Transparency for Machine Learning-Enabled Medical Devices
- FDA Artificial Intelligence in Software as a Medical Device
- CDC Diabetes and Vision Loss
- HHS Guidance on De-Identification of Protected Health Information
Questions and answers
Is an AI retinal screening result a diagnosis?
No. It is a screening output produced for a defined intended use. Symptoms, an out-of-scope patient, an incomplete study, or an ungradable image requires clinical assessment and often a diagnostic eye examination.
Does an ungradable image mean no diabetic retinopathy was found?
No. It means the system could not make the intended determination from the available input. The pathway must record the screen as incomplete and arrange a repeat or eye-care examination according to labeling and clinical context.
Can high sensitivity eliminate false-negative results?
No. Sensitivity is measured against a reference standard in a particular population and workflow. Even strong performance leaves false negatives, and real-world case mix, image quality, disease definition, and implementation can change results.
Why does specificity matter in a screening program?
Lower specificity can send more people without the target disease for follow-up, adding cost, travel, anxiety, and specialist workload. Sensitivity, specificity, ungradable rate, prevalence, and referral capacity must be interpreted together.
Should visual symptoms wait for an automated screening appointment?
No. New vision loss, flashes, a curtain or shadow, severe eye pain, marked redness, trauma, or other acute symptoms require timely diagnostic assessment and may need emergency eye care regardless of any screening result.
What should a clinic audit after introducing AI-supported retinal screening?
Audit eligibility, image-quality failure, repeat attempts, output distribution, overrides, referral orders, referral completion, time to care, missed symptoms, subgroup performance, accessibility, privacy events, software versions, and patient understanding.