AI in Dermatology: Evidence, Devices, and Equity

Dermatology is an important test case for clinical AI because the same image can support a narrow referral decision, a broad differential diagnosis, patient education, or longitudinal monitoring. Those are different tasks with different reference standards and harms. Evidence from curated image studies should therefore not be presented as evidence that an algorithm improves patient outcomes or can replace clinical assessment.

Learning Objectives

After reading this chapter, readers will be able to:

  • Distinguish lesion-triage evidence from evidence of diagnostic or patient benefit
  • Interpret regulatory status and intended use for dermatology devices
  • Evaluate image-based AI across skin tone, image capture, disease subtype, and care setting
  • Counsel patients about consumer dermatology tools without equating information support with safe self-triage
  • Assess emerging AI approaches for atopic dermatitis, psoriasis, and acne severity monitoring
  • Suspicious lesions: Dermatology AI can support defined referral decisions, but image-reader performance is not clinical-outcome evidence. Current FDA-authorized lesion devices have narrow intended uses and require clinician oversight.
  • Equity: Fitzpatrick skin type is not a skin-tone proxy. Evaluation should report a justified measurement approach, image capture conditions, subgroup representation, calibration, and external validation.
  • Patient-facing tools: Informational AI may help users name possible conditions, but that does not establish safe self-triage or treatment selection. Unvalidated app outputs must not delay evaluation of concerning lesions.
  • Inflammatory disease: Image-based severity assessment for atopic dermatitis, psoriasis, and acne is moving toward longitudinal measurement, but prospective evidence of improved management or outcomes remains limited.

Introduction

In routine care, a lesion classifier, a clinician-facing image system, an asynchronous teledermatology service, and a patient-facing informational tool have different intended users and failure modes. Evaluation should match the evidence to the decision being supported, then require prospective evidence before claiming that a system improves care.

Suspicious Lesion Assessment

Skin cancer AI should be assessed as a clinical decision-support problem, not as a contest against clinicians on a curated image set. The relevant questions are whether a tool improves referral or biopsy decisions in its intended population, what additional false-positive burden it creates, and whether performance generalizes across image capture, anatomic sites, disease subtypes, and patient groups.

Evidence Beyond Benchmark Claims

The early literature established that image classifiers can approach clinician performance on constrained lesion-classification tasks. More recent realistic reader studies are more informative for implementation. In a multi-institutional study of 1,117 clinical and dermoscopic cases, a foundation model exceeded readers with less than three years of dermatology experience, while dermatologists with more than 10 years of experience achieved the highest multiclass accuracy (Anriot et al., 2026). Image-reader performance is not prospective deployment or evidence of better patient outcomes.

A 2025 systematic review and meta-analysis of studies comparing AI with family physicians and dermatologists reported pooled melanoma sensitivity of 0.86 and specificity of 0.94. However, 25 of 38 included studies had high risk of bias, most often because the selected cases did not represent outpatient populations (Nadour et al., 2025). Spectrum bias matters: a system evaluated on selected lesions cannot be assumed to perform similarly in routine primary care, teledermatology, or patient-submitted photographs.

FDA-Authorized Lesion Devices

DermaSensor is a Class II De Novo device, authorized January 12, 2024 (DEN230008), for physicians who are not dermatologists to assist referral decisions for lesions already considered suspicious for melanoma, basal cell carcinoma, or squamous cell carcinoma in patients aged 40 years or older. It is not a screening tool or a standalone diagnostic (FDA, DEN230008). In the FDA-reviewed pivotal study, sensitivity was 96.3% and specificity was 20.3% in the indicated population. The tradeoff is intrinsic to a high-sensitivity referral aid and should be made explicit in local capacity planning.

The 2025 multi-reader companion study, sponsored by the manufacturer, found that device-aided referral sensitivity increased from 82.0% to 91.4%, while referral specificity decreased from 44.2% to 32.4% (Ferris et al., 2025). The study used selected cases rather than in-person patient care; all cases involved White patients, and 68% were Fitzpatrick type II or III. It supports a defined adjunctive role, not autonomous diagnosis or claims of real-world outcome benefit.

Nevisense is a PMA device (P150046) for dermatologists to obtain additional information when considering biopsy of a lesion with clinical or historical characteristics of melanoma. Its FDA labeling excludes acral skin and other special anatomic sites, and it should not be used for clinically obvious melanoma (FDA, P150046). Its intended use should not be generalized to broad skin-cancer screening or patient-operated imaging.

Decision Rule for Lesion AI

Use a lesion AI result only within the device’s intended population, operator group, lesion criteria, and care pathway. Before adoption, define the confirmatory assessment, referral destination, expected false-positive volume, and monitoring plan. Do not use a negative output to override concerning history, examination, dermoscopy, or clinical follow-up.

For a current device inventory, see AI Tools by Medical Specialty. For a general implementation framework, see Evaluating AI Clinical Decision Support Systems.

Equitable Image-Based AI

Skin tone, disease morphology, image capture, and care setting are inseparable sources of uncertainty in dermatology AI. A claim that an algorithm was tested “across Fitzpatrick types” is insufficient by itself. Fitzpatrick skin type was developed to describe response to ultraviolet exposure, not to measure skin tone, race, ethnicity, or the optical properties of a clinical image.

In a prospective comparative study of skin-tone labeling methods, no assessed approach generalized reliably across in-person assessment, clinical photographs, and dermoscopy images. The study concluded that Fitzpatrick skin type is not a proxy for skin tone (Weir et al., 2025). This does not supply one universal replacement scale. It does establish that image-based dermatology studies should justify their measurement method and report how labels were obtained.

Minimum Evidence for Equity Claims

An adoption dossier should report:

  1. The intended population, clinical task, image modality, acquisition protocol, and reference standard
  2. Disease subtype and anatomic-site representation, including acral and amelanotic lesions when relevant
  3. A prespecified and justified skin-tone measurement approach, rather than a race label or Fitzpatrick type alone
  4. Group-specific sensitivity, specificity, calibration, and uncertainty, with sample sizes and confidence intervals
  5. External validation across sites and capture conditions, followed by post-deployment performance monitoring

The implication is practical rather than rhetorical: a vendor that cannot characterize its data and subgroup performance has not supplied enough evidence for equitable deployment. Broader population-health implications of representation and measurement bias are addressed in the Public Health AI Handbook’s ethics and equity guidance.

Synthetic Images Are Not a Remedy

Generative models can create dermatology-like images, but synthetic imagery should not be treated as a clinical illustration, training datum, or evidence of diversity without expert review and provenance. An experimental assessment of four image generators found limited representation of darker skin tones and low accuracy in depicting the prompted dermatologic condition (Joerg et al., 2025). Synthetic augmentation requires the same scrutiny as any other data intervention: it can reproduce or amplify the errors and gaps of its source data.

Patient-Facing Tools and Teledermatology

Patient-facing dermatology AI has three distinct uses: educational information, asynchronous image submission to a clinician, and automated diagnosis or triage. They should not be evaluated or communicated as interchangeable.

Informational Tools Are Not Self-Triage Tools

A randomized study of US consumers found that an AI-powered informational interface improved users’ ability to name possible skin conditions but did not improve accuracy of the next step in care (Sayres et al., 2026). This supports a narrow role for condition information and question preparation. It does not validate self-diagnosis, cancer exclusion, or treatment recommendations.

Patients should be advised that an app result must not rule out malignancy or defer clinically indicated assessment. Clinicians should document the app result as part of the history when it has influenced care seeking, then assess the lesion on its clinical merits.

Teledermatology Is a Care Model, Not an AI Validation Shortcut

Teledermatology has evidence separate from AI. A 2026 systematic review and meta-analysis reported diagnostic concordance of 76% across skin conditions and 73% for skin cancers; dermoscopy improved skin-cancer concordance from 67% to 80% (Martyin et al., 2026). These findings support teledermatology as an access pathway in suitable settings. They do not establish that an AI classifier added to patient-submitted images improves diagnosis, referral efficiency, equity, or outcomes.

Any AI layer used with teledermatology should therefore be validated using the same image source, workflow, follow-up pathway, and patient population in which it will operate. A model trained on high-quality dermoscopy images should not be assumed to generalize to consumer smartphone photographs.

Longitudinal Inflammatory-Disease Monitoring

The most credible near-term expansion beyond lesion triage is measurement support for chronic inflammatory disease. Image-based systems are being developed to assess severity in atopic dermatitis, psoriasis, and acne, where repeated measures may support follow-up, treatment-response assessment, and research. The clinical question is not whether a model can reproduce a score in a retrospective dataset, but whether it improves decisions, burden, access, or outcomes.

A systematic review of 45 studies and meta-analysis of 19 studies found pooled sensitivity of 80.5% and specificity of 96.2% for AI severity assessment, with performance varying materially by disease and scoring system (Cai et al., 2025). The included evidence was heterogeneous and was not sufficient to establish clinical benefit from automated severity scoring.

In a real-world observational study of patient-uploaded atopic dermatitis photographs, an AI-derived severity score correlated with clinician assessments, providing an early signal for longitudinal measurement rather than autonomous remote management (Okata-Karigane et al., 2025). The study’s single-country setting, image-quality exclusions, and small objective-reference subset limit generalizability.

For psoriasis, acne, and atopic dermatitis, prospective studies should compare AI-supported monitoring with standard care on treatment changes, patient-reported outcomes, clinician burden, access, and performance across skin tones and image conditions. Until then, AI severity scores should complement validated clinical instruments rather than replace them.

Questions to Ask Before Adopting Dermatology AI
  1. What exact clinical decision does the tool support, and what is the reference standard?
  2. Does the validation population, image source, and workflow match local use?
  3. What are the group-specific sensitivity, specificity, calibration, and uncertainty estimates?
  4. What clinical action follows a positive or negative result, and who owns follow-up?
  5. What false-positive volume, specialist capacity, and patient burden should be expected?
  6. What prospective or real-world evidence shows benefit beyond image-level accuracy?
  7. How will drift, subgroup performance, and adverse events be monitored after deployment?

Clinical Bottom Line

Dermatology AI is clinically useful only when the task, users, images, and follow-up pathway are defined. A high score on a selected image dataset does not establish diagnostic safety, equitable performance, or patient benefit. The strongest current use cases are clinician-supervised adjuncts for defined lesion decisions and emerging longitudinal measurement support for inflammatory disease. Patient-facing informational tools and teledermatology can improve access and understanding, but neither should be presented as proof that automated self-triage is safe.

See Clinical AI Safety and Risk Management for common image-model failure modes, Physician AI Liability and Regulatory Compliance for documentation and oversight considerations, and Medical Ethics, Bias, and Health Equity for the clinical equity framework.