Welcome back to the journal review. This is our look at the Journal of the American Academy of Dermatology, July 2026 issue, and we've got four pieces worth your attention this month — two letters to the editor engaging with recent AI-in-dermatology publications, a controversies-style opinion piece on the pros of artificial intelligence, and a brief report using a national registry to stress-test the thirty-one gene expression profile test in thin melanomas. Let's get into it. First up is a letter to the editor, essentially a rapid-fire follow-on study responding to a prior exchange in the journal about ChatGPT-4o's performance on dermoscopic images. You'll recall the original piece by Tadros and colleagues, and a response from Ke and colleagues, debating whether ChatGPT-4o's poor dermoscopic accuracy was a fair test given image standardization concerns. This new letter, from Gandhi and colleagues, essentially says: fine, let's standardize the images and test other large language models too. They're building on an earlier study of theirs using seven hundred seventy-five images from the Harvard HAM10000 dataset — a well-curated, pathology-confirmed pigmented lesion database — in which ChatGPT-4o's accuracy was significantly below the greater-than-seventy-percent mean accuracy dermatologists achieve with dermoscopy, and, notably, misclassified malignant melanoma as benign in about half of cases. That's the backdrop. For this letter, they ran the same prompt-based methodology — asking simply for the most likely diagnosis and whether the lesion is benign or malignant — but this time on Claude and on Gemini, using two independent raters per model across four hundred images each, split evenly among melanoma, benign nevi, seborrheic keratosis, and vascular lesions. The rationale for using two raters is worth flagging: it lets them assess interrater reliability, which turns out to be a critical part of the story, not just diagnostic accuracy in isolation. The results are unimpressive across the board. Claude landed at an overall diagnostic accuracy of about 43%, with interrater agreement only around half the time. Gemini came in similarly at about 45% accuracy, but with interrater agreement down near one in five — meaning the same model, fed the same image twice by different reviewers, frequently gave different answers. The difference between Claude and Gemini was not statistically significant, and clinically, that's almost beside the point — both are simply poor, and inconsistently so. Sensitivity and specificity for melanoma versus everything else were unimpressive for both models, with positive predictive values sitting in the thirty-to-fifty percent range — not something you'd want anchoring a triage decision. The authors close by returning to the equity question raised in the original exchange: stratification by Fitzpatrick skin type would be valuable, but a cited scoping review found that among over a hundred fifty dermoscopy studies, only about one in eight even reported phototype, and none included phototype six. Their bottom line is straightforward and worth remembering: this isn't a ChatGPT problem, it's a category problem. Commercially available large language models, whatever the brand, underperform on dermoscopic diagnosis even against a clean, standardized, high-quality dataset — so calls for "better datasets" as the fix likely miss the more fundamental architectural limitations of these models for image-based diagnosis. Nothing here is practice-changing since none of us should be using these tools diagnostically anyway, but it's a useful, well-documented data point to have on hand next time a patient — or a hospital administrator — asks whether an AI chatbot could triage pigmented lesions. Next is a controversies piece — one half of a paired pro-con debate — arguing the "pro" side of artificial intelligence in dermatology, written by Klufas, Zhou, and Grant-Kels. This is an opinion piece, not original data, so there's no methods section to walk through; it's a position paper marshaling existing evidence to make its case, and it's worth knowing what evidence they lean on. Their central argument is that AI's biggest win isn't replacing dermatologists but augmenting non-dermatologists, especially in under-resourced settings. They cite a meta-analysis showing AI assistance in skin cancer diagnosis produced a pooled increase in sensitivity and specificity of roughly six and five points respectively, up from baseline figures in the mid-seventies and low-eighties — a real but modest bump. The more clinically interesting number, though, is that the benefit was concentrated in non-dermatologists, who saw sensitivity and specificity gains in the ten-to-thirteen-point range when using AI assistance — meaningfully larger than what dermatologists themselves gained, which makes sense if you think of these tools as raising the floor rather than the ceiling. They also point to human-AI collaborative full-body skin exam and dermoscopy studies showing improved diagnostic performance and, notably, close to a one-fifth reduction in unnecessary excisions of benign nevi when AI-assisted tools are used collaboratively — a number that, if it holds up in broader practice, would actually matter to procedural burden and healthcare costs. Their framing throughout is collaborative, not autonomous: AI as adjunct, not replacement, with human oversight remaining essential, and a call for more transparent, representative training datasets going forward. There's no new data here for you to act on tomorrow, but it's a useful, well-referenced articulation of the optimistic case, and worth having in mind as the counterpoint to the sobering LLM letter we just covered — the difference being purpose-built, validated convolutional neural network tools used collaboratively versus general-purpose chatbots asked to freelance a diagnosis. Third is another letter to the editor — a response by Klufas, Zhou, and colleagues to a prior letter from Hu and colleagues, continuing a discussion about the downstream effects of AI ambient scribes, particularly on premedical students and medical assistants who currently work as scribes. Again, this is correspondence, not a study, so think of it as an argument being extended rather than data being generated. The authors concede the concern — that AI scribes could erode a training and employment pipeline for premedical students — but pivot to the productivity and burnout case for AI scribes in specialties like dermatology, where documentation burden is a major burnout driver and many clinicians don't have human scribes at all. The one concrete figure worth remembering here: cited data from an ambulatory cohort found AI scribe use associated with an increase in weekly encounters and relative value units, translating to about three thousand dollars in additional annual productivity per provider, with stable billing approval rates. They reasonably note that high-volume, fast-paced specialties like dermatology could see outsized benefit from small per-encounter efficiency gains, though they also flag a real dermatology-specific caveat — scribes often double as chaperones during exams, a function AI scribes cannot fill, so any efficiency gains may be partially offset by needing to staff that role separately. Their overall stance is measured: AI scribes are tools, not replacements for judgment about how documentation and staffing should evolve, and any time or cost savings should be reinvested in the patient and provider experience rather than simply used to justify larger patient panels. Nothing practice-changing to act on today, but a reasonable framework if you're evaluating scribe technology for your own practice. The last piece is a brief report — real registry data this time — looking at whether the thirty-one gene expression profile test, the 31-GEP, actually predicts sentinel lymph node biopsy positivity in AJCC pT1b melanomas, using the SEER-DecisionDx linked registry. This is squarely relevant to your daily shared decision-making conversations, so let's spend some time here. The background: NCCN guidelines already call for discussing sentinel node biopsy with pT1b patients. The 31-GEP has been marketed as a way to refine that conversation — low-risk Class 1A results theoretically allowing patients to reasonably forgo biopsy, while intermediate or high-risk results argue for proceeding. But the evidence has been conflicting: the industry-sponsored DecisionDx impact study supported using the GEP to forgo biopsy, while a more recent Mayo Clinic cohort found the test had limited value in predicting nodal positivity. This registry study is essentially an independent arbiter, using a large, real-world, linked national dataset rather than a single-institution or industry-sponsored cohort — a methodological choice that matters because it reduces referral bias and industry influence, though as we'll see it comes with its own real-world messiness. Methodologically, they pulled histologically confirmed melanoma cases from SEER using standard morphology and topography codes, restricted to pT1b tumors with known Breslow depth and ulceration status — a sensible restriction since those are the variables that define the T1b category and let them confirm they're looking at the right population. This gives them thirteen hundred pT1b patients, average age right around sixty, just over half male, tumors trunk-located in just under half the cohort, with an average Breslow depth of about nine-tenths of a millimeter — squarely mid-range for T1b — and ulceration present in roughly one in eight tumors. Here are the numbers that matter. Just over half the cohort — about 57% — actually underwent sentinel node biopsy, and of those, about 11%, or one in nine, were node-positive — a rate entirely consistent with what you'd expect for pT1b disease generally. But here's the finding that undercuts the GEP's proposed clinical utility: of those eighty node-positive patients, more than four out of five — 82% — had been classified as Class 1A, the ostensibly low-risk category by the GEP. In other words, the test's low-risk call was wrong the vast majority of the time it mattered most. And among patients who didn't undergo biopsy at all, those classified Class 1B or higher had a numerically higher hazard of melanoma-specific death compared to Class 1A patients, but this difference was not statistically significant, so we can't lean on it either way — though the authors are careful to note that this at least fails to support the reassuring claim that Class 1A patients who skip biopsy are provably not harmed. The authors are appropriately blunt about the implication: this data does not support using 31-GEP results to forgo sentinel node biopsy in pT1b melanoma, and it echoes — rather than contradicts — the Mayo Clinic experience, which paradoxically also found higher nodal positivity rates in low-risk GEP cases. They add useful context on the economics, too — Medicare reimbursement runs around seven thousand dollars for the GEP test versus roughly thirty-five hundred for a sentinel node biopsy done in an ambulatory setting, or around six thousand two hundred if done through a hospital outpatient department — figures worth having on hand given they flag this as a potential source of financial conflict of interest pushing test adoption. They also remind us that sentinel node biopsy itself isn't a perfect gold standard, citing a false-negative rate of about 10% in one recent series, so neither test is without error — but the GEP's error, in this dataset, runs in the direction of falsely reassuring the highest-risk group of missed patients. Limitations here are real and the authors own them: SEER-DecisionDx doesn't capture recurrence data, so they can't speak to recurrence-free survival, only overall and melanoma-specific mortality signals, and the survival comparison in patients who skipped biopsy was underpowered. This is registry data, with all the usual concerns about selection into who actually got GEP testing and who got sentinel node biopsy in the first place. But practically speaking, for those of you having this conversation with pT1b patients: this is another data point — now from a large, independent, real-world registry — arguing against using a reassuring low-risk GEP result as grounds to skip sentinel node biopsy. I would not call this fully practice-changing on its own, since it's observational and conflicts with the industry-sponsored dataset, but combined with the Mayo Clinic findings, it should raise your threshold for leaning on GEP results alone to talk a pT1b patient out of biopsy, and it strengthens the case for continuing to frame sentinel node biopsy as the more reliable staging tool until better prospective data settles this discrepancy. That wraps our four pieces for July. The throughline this month is really about calibrating trust in adjunct technologies — where large language models and gene expression profiling both promise to lighten diagnostic and decision-making burden, but where the actual data, when you look closely, argues for real caution before letting either one steer a management decision on its own. Thanks for listening, and we'll see you next issue.