A smartphone can photograph a skin lesion, compare images over time, route a picture to a clinician, or run software that assigns a risk category, and those are different functions with different evidence needs. A calendar reminder does not make a diagnosis. A teledermatology service includes a human pathway. An automated output such as “low risk” or “high risk” behaves like a medical test and can change whether you seek care.
Earlier research found wide and sometimes unsafe variation in app sensitivity, but it tested products and versions that may no longer be available. The valid lesson is not a permanent score for every smartphone tool. It is that each current version needs prospective external validation in people and lesions like the intended users, and even then its labeled role matters.
First define what the application claims to do#
An image diary helps a user compare a lesion over time. A teledermatology platform transmits information to a clinician. An automated classifier converts pixels and user-entered data into a risk category. An educational app may list warning signs without analyzing a photograph. Each has a distinct intended use, evidence need, and regulatory position.
“Skin app” is therefore too broad for an accuracy claim. Ask for the product name, software version, country, supported phone and camera, target age, target lesion types, body sites, user instructions, output, decision threshold, and recommended next action. An app that says “seek review” is not necessarily claiming to diagnose melanoma. An app that labels a lesion low risk may influence delay even if a disclaimer calls the output educational.
Regulation follows intended use and risk, not the presence of artificial intelligence as a label. In the United States, some software functions meet the definition of a medical device, some fall within areas where FDA intends to exercise enforcement discretion, and others are not device functions. An app-store listing is not evidence of FDA clearance or approval. A user can search FDA databases and inspect the exact manufacturer and model, but absence or presence in a database still does not replace appraisal of the performance study.
What the early studies found#
A 2013 JAMA Dermatology study tested four applications using images of 188 lesions, including 60 melanomas, with pathology as the reference standard; sensitivity across apps ranged from 6.8% to 98.1%, and specificity ranged from 30.4% to 93.7%. Three of the four automated approaches classified at least 30% of melanomas as unconcerning. The most sensitive service used image review by a board-certified dermatologist rather than a fully automated algorithm.
Those numbers revealed risk, but they were not population screening estimates. The images represented known, selected lesions and did not reproduce the full home-use pathway. Products also change, so the performance belongs to the tested versions and dataset.
The 2018 Cochrane review found only two eligible studies covering five applications and 332 suspicious lesions, of which 86 were melanomas. Study quality was poor and applicability to routine users was limited. Across the evaluated automated apps, false reassurance affected between 7 and 55 of every 100 melanomas in the study sets, depending on the app and threshold, and the review concluded that available evidence was insufficient to support app recommendations and highlighted the danger of missed melanoma.
A 2020 BMJ systematic review found nine relevant studies. Six studies with a pathology or follow-up reference standard contributed 725 lesions. In one evaluated subset, SkinVision sensitivity was estimated at 80% with a 95% confidence interval from 63% to 92%, and specificity at 78% with a confidence interval from 67% to 87%. The authors emphasized small samples, selected images, incomplete or inappropriate reference standards, and poor reporting; a wide confidence interval means the true miss rate compatible with the data could be materially worse than the point estimate suggests.
These reviews remain historically important but cannot certify a current version. A machine-learning model may be retrained, a threshold altered, an image-quality filter changed, or a phone operating system updated, and even without an algorithm change, the users and lesions encountered after launch may differ from the validation set.
Sensitivity and specificity do not tell the whole story#
Sensitivity is the proportion of cancers correctly flagged among people or lesions with cancer. Specificity is the proportion correctly reassured among those without cancer. A test can raise sensitivity by referring more lesions, but that usually lowers specificity and increases false alarms.
Positive and negative predictive values also depend on prevalence. A model tested in a specialist clinic containing many cancers may have a different reassuring value in a consumer population where cancer is uncommon. The relevant unit may be a lesion, person, image, or episode of care. Treating multiple lesions from one person as independent can make confidence intervals too narrow.
Unreadable images require explicit accounting. If the app rejects blurry, dark, hairy, curved, nail, or otherwise difficult images, analysts should not simply remove them and report accuracy only among accepted photographs. Image rejection affects usability and could be related to cancer status, skin tone, body site, device, or user skill. An indeterminate result should have a safe action pathway.
The threshold also affects downstream care. Useful evaluation measures time to clinical review, biopsy patterns, stage at diagnosis, unnecessary procedures, anxiety, cost, and whether reassurance caused delay. A classifier can perform well on a static dataset yet fail to improve outcomes if people photograph the wrong lesion or do not follow the advice.
The spectrum and image-acquisition problems#
Specialist datasets often contain lesions selected for biopsy or teaching. Cancers may be more advanced or visually distinctive than the subtle lesions encountered in primary care or at home. Benign controls may be unusually clear. This case-mix difference can make sensitivity and specificity shift even when code is unchanged. The spectrum effect in diagnostic accuracy explains why performance is a property of a test in a population and pathway.
Home acquisition adds another layer. Lighting, focus, compression, scale, skin preparation, camera optics, and body-site access vary. Users may select a familiar mole while missing a new lesion on the scalp, back, sole, or nail. Many research images are dermoscopic, captured with magnification and controlled illumination, while a consumer submits an ordinary phone photograph. Validation using one image type cannot establish accuracy for another.
Skin-tone representation must be reported. Aggregate performance can hide weak sensitivity or high rejection rates in a smaller subgroup. Race and ethnicity do not directly measure image appearance, and a tone category does not capture lesion site, device quality, access, or tumor subtype. Developers should report sample composition, subgroup uncertainty, calibration, and failure modes without treating demographic labels as biological shortcuts. Skin cancer in skin of color covers the clinical pattern.
A current evidence checklist#
Before relying on a study, check whether it evaluated the exact available version and intended use. A credible assessment should include:
- Prospective recruitment from the target pathway rather than only a convenient image archive.
- External testing at sites that did not develop the model.
- An appropriate reference standard, usually pathology for biopsied lesions plus adequate clinical follow-up for relevant non-biopsied lesions.
- Consecutive or clearly sampled participants, with exclusions and indeterminate outputs reported.
- Prespecified thresholds, locked before final testing.
- Sensitivity, specificity, predictive values, calibration, confidence intervals, and clinically relevant subgroup results.
- Documentation of who chose the lesion and captured the image.
- A comparison that reflects actual care, such as usual assessment or a defined referral pathway.
- Version traceability and a plan to detect post-launch performance change.
STARD reporting helps readers see whether these elements are present, but complete reporting does not remove bias. An independent study can still be unrepresentative, and a manufacturer study can still be informative if methods and data are rigorous. Funding, conflicts, protocol, and analytic access should be transparent.
Regulation and advertising are separate safeguards#
In 2015, the US Federal Trade Commission challenged marketers of melanoma-detection apps over claims that were not supported by adequate evidence. Advertising enforcement asks whether promotional claims are substantiated. FDA device oversight asks about intended use, risk, premarket pathway where applicable, quality systems, reporting, and other device requirements. Privacy and consumer-protection laws address additional concerns. One safeguard does not substitute for the others.
A regulatory authorization is specific to indications, inputs, users, and labeling. It does not mean a product finds every melanoma or works outside its supported phones, ages, lesion types, and settings. Conversely, a general image-storage or communication function may not require the same pathway because it is not making an automated diagnostic claim.
After launch, developers should track complaints, false reassurance, false alarms, image failures, subgroup performance, software changes, cybersecurity, and clinical outcomes where feasible. A high-performing model can become less reliable if camera processing changes or use expands into populations absent from development.
Safer use in a real decision#
Do not use an app result to cancel an appointment for a lesion that is new, changing, bleeding, painful, unlike other lesions, beneath a nail, on a palm or sole, or not healing. A clinician may use your history, the full skin context, dermoscopy, and biopsy. An application sees only the submitted information.
Date-stamped photographs can help document evolution while care is arranged. Use consistent lighting and include a size reference if instructed, but do not repeatedly monitor a concerning lesion instead of seeking review. If access is difficult for you, primary care or a legitimate teledermatology service may help route assessment. Rapid progression, substantial bleeding, or systemic illness can require faster care.
The safest interpretation of current evidence is not that all apps are useless. It is that utility is function-specific and version-specific. Tracking, reminders, education, clinician communication, and automated triage each require evidence matched to the actual role and consequences.
References#
- Cochrane review of smartphone apps for suspicious skin lesions
- BMJ 2020 systematic review of algorithm-based melanoma apps
- JAMA Dermatology 2013 diagnostic-accuracy study
- FTC actions involving melanoma-detection apps
- FDA guidance on device software functions and mobile medical applications
- STARD 2015 reporting guideline
For your own health, talk with your clinician.*
Questions and answers
Can a melanoma app rule out cancer if it says low risk?
No consumer result should be treated as a definitive rule-out. Earlier automated apps missed melanomas, and current accuracy must be shown for the exact version and user pathway. A concerning lesion needs clinical assessment regardless of the label.
Does FDA clearance mean an app is always accurate?
No. Authorization applies to a defined intended use and evidence package, with known limitations. Performance is not perfect and may not transfer to unsupported devices, lesion types, ages, body sites, or users.
Why are pathology-confirmed image sets not enough?
They establish strong labels for sampled lesions, but often contain selected specialist cases. They may omit the steps where consumers choose a lesion, take a usable image, interpret output, and act. Those steps can change safety and effectiveness.
Is dermatologist teleconsultation the same as automated analysis?
No. Teledermatology transmits information for human assessment, while automated software generates an output from programmed or learned rules. Both need evidence, but their failure modes, workflows, and regulatory claims differ.
What should I do with a changing mole while waiting for care?
Arrange clinical review, record a dated photograph if useful, and note the change and symptoms. Do not cut, treat, or repeatedly test it with apps. Seek faster help for rapid growth, significant bleeding, infection signs, or other urgent symptoms.