Welcome back to the journal review. This month we're working through the May 2026 issue of the Journal of the American Academy of Dermatology, and we've got four pieces on deck — a commentary on artificial intelligence and board exam performance, an original eye-tracking study comparing dermatologists and an AI algorithm on dermoscopy, a brief report on melanoma malpractice litigation, and a follow-up letter on the dermoscopic "plumage sign" in pigmented squamous cell carcinoma in situ. Let's get into it. First up is a letter to the editor, a commentary responding to an earlier meta-analysis by Chen and colleagues that evaluated how well ChatGPT-4 performs on dermatology board-style exam questions. This is not new data — it's a critical appraisal of that meta-analysis, raising three points worth flagging for anyone thinking about where large language models actually stand in our field. The first and probably most clinically relevant point is what the commentary authors call the "visual gap." The original meta-analysis found that across all ChatGPT versions, text-based question performance was noticeably better than performance on image-based questions — roughly two-thirds correct on text versus about half on images. For ChatGPT-4 specifically, that gap was even more dramatic: about eighty percent accuracy on text questions but only around sixty percent on image-based ones. The commentary's point, and it's a fair one, is that dermatology is fundamentally a visual specialty, and a model that can reason well about a written vignette but stumbles on image interpretation is nowhere near ready for autonomous use in the tasks that actually matter to us. The second point is about reproducibility — the original study found a significant correlation between when a given ChatGPT model was accessed and how well it scored, meaning performance drifts over time as these models are updated. The commentary authors note this creates a real problem for benchmarking: any snapshot evaluation of an AI model may already be stale by the time it's published, which complicates comparing results across studies done even months apart. Third, they reiterate a limitation the original authors already acknowledged — board exam performance, however impressive, is not clinical validation. High multiple-choice scores don't capture the uncertainty, nuance, and safety stakes of real patient encounters. Bottom line for practice: nothing here is actionable yet, but it's a useful reminder to stay skeptical of headline AI accuracy numbers, especially for image-heavy tasks, and to recognize that these benchmarks have a shelf life. Next is an original study, and this one's methodologically interesting: a comparison of dermatologists' visual attention and an AI algorithm's decision-making, using eye tracking as the common currency. The clinical problem here is the "black box" issue — AI-based dermoscopic image classifiers can output a diagnosis and a heat map showing which pixels drove that prediction, but we don't actually know whether the algorithm is looking at clinically meaningful structures or just picking up on spurious correlations in the training data. That matters enormously for trust and adoption. So the authors set out to quantify whether an AI algorithm's attention overlaps with expert human attention, using eye tracking as the ground truth for "where humans look." Methodologically, four dermatologists, blinded to diagnosis, evaluated 120 dermoscopic images spanning melanoma, basal cell carcinoma, squamous cell carcinoma, nevi, benign keratoses, and vascular lesions, in equal numbers, while wearing high-precision infrared eye trackers. This generated gaze heat maps — essentially a record of where each dermatologist's eyes lingered and for how long. For the same images, the authors ran a proprietary algorithm called DEXI, short for Dermoscopy EXplainable Intelligence, which is a five-network convolutional neural network ensemble built into a commercial dermoscopy software platform, and pulled its class activation maps, which are the AI's own version of a heat map showing which regions drove its prediction. The clever design choice here is the reference framework: rather than just reporting a raw correlation number, which is hard to interpret in isolation, the authors built in two anchors — an upper bound, which is the correlation between different dermatologists looking at the same image, essentially "how much do humans agree with each other," and a lower bound, a null correlation generated by pairing DEXI's heat map for one image against a dermatologist's eye-tracking map from a completely different image. That null comparison establishes what correlation you'd expect from pure chance alignment. This is a smart way to contextualize an otherwise abstract pixel-correlation statistic, and it's worth noting as a general technique — anchoring a novel metric between a best-case and chance-level reference is exactly how you make a single correlation coefficient clinically interpretable. On to results. The correlation between dermatologists' gaze patterns and DEXI's heat maps came out to about 0.54. Inter-dermatologist correlation — the upper anchor — was about 0.59. And the chance-level null correlation was about 0.43. So DEXI's attention pattern was nearly as similar to a dermatologist's gaze as two dermatologists were to each other, and clearly above what you'd expect from random chance. That's a meaningful signal that the algorithm is, broadly speaking, looking where expert human eyes look. A few secondary findings are worth flagging. Fixation-based heat maps — which capture where the eye actually pauses to process information, as opposed to just where it passes over — showed lower correlation with DEXI than the gaze-based maps did, for both human-to-AI and human-to-human comparisons, and that difference was statistically significant. When the authors tightened the analysis to just the lesion itself, excluding peri-lesional skin, correlations actually dropped for both DEXI and inter-dermatologist comparisons, suggesting that some of the shared "attention" in the main analysis was being driven by how both humans and the algorithm scan the surrounding skin, not just the lesion. Diagnostic accuracy varied across readers — the most experienced dermatologist, with over three decades of experience, had the fewest misdiagnoses at under twenty percent, while the less experienced readers missed closer to a quarter to nearly a third of cases, with the highest miss rates seen in benign keratoses and squamous cell carcinoma. Interestingly, correlation between dermatologist and DEXI attention was actually a bit higher on lesions the dermatologists got wrong than on ones they got right — a subtle finding the authors don't over-interpret, but it hints that when humans are drawn to the same ambiguous features an algorithm is weighing, that ambiguity itself may be diagnostically treacherous for everyone. DEXI misclassified under ten percent of lesions overall, with its errors concentrated in melanomas read as nevi and pigmented benign keratoses read as melanoma or nevi. For limitations, the authors are upfront: sample sizes within each lesion subtype were small, and they didn't have lesion size data, which could plausibly affect gaze patterns independent of diagnosis. I'd add, as my own observation rather than something the authors state explicitly, that this is a single AI algorithm tested against four readers at one institution, so generalizability to other explainable-AI systems or other dermatologist populations is unproven. Practically, this is an interesting proof-of-concept for AI interpretability research rather than anything that changes how you practice tomorrow. It doesn't tell you whether to trust DEXI's diagnosis on a given lesion, and it's not evaluating diagnostic accuracy of AI versus human head-to-head. What it does provide is reassurance that when this particular algorithm flags a region as important, it's usually flagging something a trained dermatologist would also be looking at — which is a meaningful, if incremental, step toward the kind of transparency that would eventually support clinical integration. Third, a brief report — really structured like a research letter — on malpractice litigation specifically around melanoma, using the LexisNexis legal database from 2000 to 2024. The motivating fact, cited up front, is sobering: by age sixty-five, nearly three-quarters of dermatologists will face a malpractice lawsuit at some point, and melanoma is consistently one of the top three drivers of liability claims in the specialty. Prior work has looked at dermatology malpractice broadly, but melanoma litigation cuts across specialties — pathology, primary care, oncology — so the authors wanted a melanoma-specific picture. Methodologically, this was a straightforward retrospective search of the database using the terms "melanoma" and "malpractice," excluding cases where melanoma wasn't central to the lawsuit or where the case was still pending, yielding one hundred sequential cases for descriptive analysis. This is about as simple as study design gets — essentially a case series pulled from public legal records — and the rationale, though not spelled out at length by the authors, is obvious: there's no other way to systematically characterize litigation patterns except by mining the legal record itself, since malpractice claims and outcomes aren't captured in clinical databases. On the numbers: defendants won sixty of the hundred cases, plaintiffs won thirty-four, and six ended in mixed rulings or settlements. Cases clustered in the South and in private practice settings. Among plaintiff victories, the three leading allegations — failure to test or diagnose, false-negative misdiagnosis, and negligent follow-up — each accounted for about one-fifth of cases. Financial payouts were recorded in only four cases, ranging from about two hundred sixty thousand dollars up to four and a quarter million dollars — a wide spread, but a reminder that when melanoma litigation does result in a payout, the numbers can be substantial. Compared to keratinocyte carcinoma litigation, where defendants win about three-quarters of the time, and nail disorder litigation, where defendants win over ninety percent of the time, melanoma cases were tougher for the defense, with close to a forty percent plaintiff success rate. That's also higher than a 2016 Westlaw database review, which found defendants winning about half of melanoma cases versus sixty percent here — though the authors note that over half of defendant wins in their cohort came from procedural or technical defenses, like statute of limitations issues, rather than a finding of no clinical negligence, so the true "clinical merit" defendant win rate is likely lower than the topline number suggests. Perhaps the most striking finding for a Mohs and dermatologic oncology audience: prison medical staff were the single most frequently named defendant group, ahead of dermatologists, which the authors attribute to a quirk of the database — prison-related cases are automatically adjudicated in state or federal court and therefore overrepresented in public legal records. Among plaintiff victories specifically, eighty-five percent of defendants were non-dermatologists, and pathologists were the most common specialty named, underscoring how central histopathologic interpretation is to liability exposure — both false-negative reads driving delayed diagnosis, and false-positive melanoma calls driving lawsuits over unnecessary treatment. Limitations here are inherent to the data source: this only captures publicly filed cases, so confidential settlements are invisible, less severe claims are likely underrepresented, and prison cases are likely overrepresented relative to their true frequency in overall melanoma litigation. The practical takeaway isn't practice-changing in a technical sense, but it is a useful risk-management reminder — document thoroughly around biopsy decisions and follow-up recommendations, since failure to diagnose and negligent follow-up remain the two most litigated failure points, and recognize that pathologic interpretation, not just clinical judgment, is a major locus of liability in this space. Last, a short letter following up on last year's original description of the "plumage sign" in pigmented Bowen's disease, meaning pigmented squamous cell carcinoma in situ. The original article described a novel dermoscopic pattern — curled, tubular brown structures with hyperpigmented tips, arranged like overlapping bird feathers — as a clue for this diagnosis. These authors decided to prospectively look for the sign in their own practice over a four-month stretch. They diagnosed twenty cases of pigmented Bowen's disease during that window, and thirteen of them, about two-thirds, displayed the plumage sign. Every one of those thirteen occurred in older adults with sun-damaged skin, mean age in the early seventies, mostly men, and notably, all thirteen were located on the lower extremities — mostly the leg, one on the thigh. All were histopathologically confirmed, showing the expected full-thickness epidermal dysplasia and atypical mitoses of Bowen disease, along with increased basal melanin in every case. In this small experience, the sign was perfectly specific — all thirteen lesions showing the plumage pattern were confirmed pigmented Bowen's disease, and the authors did not see this pattern in melanocytic lesions, pigmented basal cell carcinomas, or solar lentigines, though they flag a possible overlap with flat seborrheic keratosis, large cell acanthoma, and melanoacanthoma that hasn't been worked out yet. They offer a histogenetic hypothesis — that the sign likely reflects basally pigmented atypical keratinocytes arranged linearly around dermal papillae — but are careful to note this wasn't confirmed with targeted biopsy correlation. This is a small, uncontrolled case series, not a validation study, so the specificity numbers shouldn't be over-read, but the practical takeaway is straightforward and low-cost to adopt: when you're doing dermoscopy on a pigmented lesion on the lower leg of an older, sun-damaged patient, add the plumage sign to your pattern-recognition checklist for pigmented Bowen's disease, while keeping in mind the differential overlap with pigmented benign proliferations that still needs to be sorted out with larger, controlled studies. That wraps up this month's review — a cautionary note on AI benchmarking, a genuinely interesting look under the hood of an explainable AI dermoscopy tool, a sobering snapshot of melanoma litigation risk, and a nice piece of bedside pattern recognition to add to your dermoscopic vocabulary. Thanks for listening, and we'll see you next month.