Welcome back to the journal review. This is our December 2025 walkthrough of the Journal of the American Academy of Dermatology, and we've got four pieces worth your time this month — an editor's letter on artificial intelligence in scientific publishing, a brief report stress-testing GPT-4 Vision against dermoscopic images, a national database study on treatment patterns and survival in dermatofibrosarcoma protuberans, and a phase three randomized trial of red light photodynamic therapy for superficial basal cell carcinoma. Let's get into it. First up is a letter from the editor, written by Dirk Elston, titled "Artificial Intelligence: The Good, the Bad, and the Ugly." This is an editorial, so there's no methods or results section to walk through — it's Elston laying out a position, and it's worth hearing in full because it's squarely about how our literature is being produced right now. His core argument is that AI cannot be an author, full stop, because authorship requires the ability to take responsibility for content, and a language model can't do that. He grants that AI has legitimate uses — tightening prose, helping non-native English speakers with readability — but he stresses that every AI-assisted edit needs careful human review, because these tools can shift meaning in ways that are easy to miss. He uses a couple of punchy examples about comma placement changing meaning entirely, which sounds almost cute until you map it onto a methods section or a drug dosing sentence. The heavier concern is intellectual property and the flood of AI-generated or AI-assisted "research" now hitting predatory journals — he cites data on retractions rising sharply, with the causes shifting from honest error toward outright misconduct, plagiarism, and paper mills. His worry for us specifically is mentorship: a resident or junior faculty member under pressure can generate a stack of polished-looking manuscripts in half an hour, and a mentor's name can end up attached to something with fabricated references or unvetted claims before anyone notices. His closing point is essentially a call to action for anyone supervising trainees — that vetting AI use is now a core mentorship responsibility, not an optional add-on, and that journals including JAAD have adopted formal policies on AI disclosure for exactly this reason. There's no data to critique here, just a warning worth internalizing, especially if you're reviewing manuscripts or signing off on trainee work. That sets up nicely for the second piece, which is a brief report putting one of these AI tools through an actual diagnostic test. This is "Comparative Assessment of GPT-4 Vision in Dermoscopic Image Analysis," out of Johns Hopkins. The gap they're addressing: there's been plenty of work on purpose-built dermoscopy AI classifiers, but general-purpose large language models with vision capability — the kind anyone can access through a consumer chatbot — hadn't been systematically benchmarked against dermatologists on dermoscopic images specifically. Methodologically, this is a straightforward comparative accuracy study. They pulled two hundred seventy-three biopsy-proven dermoscopic images from a public Instagram dermoscopy archive, deliberately selecting images posted after GPT-4 Vision's training cutoff — a smart design choice to make sure the model couldn't have simply memorized these exact images during training, which is a real contamination risk with any public-image benchmark. GPT-4 Vision was prompted with chain-of-thought reasoning and structured output, and five human dermatologists independently read the same images, each providing a top-three differential, a confidence rating, and a biopsy recommendation. The mix of pathology was heavy on basal cell carcinoma and melanoma, with smaller numbers of squamous cell carcinoma, seborrheic keratosis, dermatofibroma, verruca, and intradermal nevi. The headline result: human dermatologists were correct on their top diagnosis about eight in ten times, versus roughly half the time for GPT-4 Vision — a substantial and clinically meaningful gap, not just a statistical one. Interestingly, in the melanoma subgroup specifically, GPT-4 Vision actually outperformed the humans on top-diagnosis identification, catching around ninety-six percent versus seventy-nine percent for the dermatologists. But — and this is the important asymmetry — that apparent strength was bought at the cost of massive overcalling. GPT-4 Vision recommended biopsy for literally one hundred percent of seborrheic keratoses and about nine in ten dermatofibromas, compared to roughly a third of each for the human graders. So it wasn't discriminating melanoma well so much as it was calling everything pigmented or irregular "concerning" and recommending biopsy reflexively. Confidence tracked accuracy in both groups — both were less confident when they were wrong — but that self-awareness signal was much more pronounced in the human graders than in the model. The authors' own conclusion is appropriately blunt: GPT-4 Vision is not ready for independent clinical use in dermoscopy. The false-positive burden, especially the reflexive biopsy recommendations for benign lesions, would translate into real patient anxiety and unnecessary procedures if this were let loose unsupervised. The one constructive thread they pull out is that the confidence-accuracy relationship, weak as it is, might be a useful lever for future model refinement. Limitations are the ones you'd expect from a single-institution proof-of-concept — a few hundred images from one public source, a single vision model rather than a head-to-head of multiple LLMs, and a dataset that likely doesn't capture the full range of skin phenotypes. For your practice, the takeaway is not practice-changing in the sense of anything you'd do differently tomorrow, but it is a useful data point to have on hand the next time a patient shows you a chatbot's read of their own dermoscopic photo — the tool is particularly prone to overcalling benign lesions, and that pattern is now documented, not just anecdotal. Now to something with actual management implications: "National Trends in Treatments and Outcomes for Dermatofibrosarcoma Protuberans," a National Cancer Database analysis. The clinical gap here is real — there's been growing evidence that Mohs micrographic surgery outperforms wide local excision for DFSP, but it wasn't clear whether that evidence had actually translated into national practice patterns, or whether the survival advantage would hold up in a much larger, more heavily adjusted dataset than prior single-institution or SEER-based studies. This is a retrospective cohort study using the National Cancer Database Participant User File, capturing DFSP cases diagnosed between 2004 and 2017 — over seven thousand patients, which is the whole point of using a registry like this rather than an institutional series. A database of this scale is really the only way to get adequate statistical power for a tumor this uncommon, and it lets you adjust for a long list of sociodemographic and clinical covariates simultaneously, something no single Mohs practice's case series could do. The tradeoff, which the authors acknowledge, is that the NCDB has no recurrence data and no cause-of-death coding, so this can only speak to overall mortality, not disease-specific mortality or local control. On the results: wide local excision remained the dominant approach, used in about half of patients, while Mohs surgery was used in roughly one in nine. But Mohs utilization rose meaningfully in the more recent era — diagnosis after 2011 was associated with about a fifty percent higher odds of receiving Mohs relative to wide excision — and treatment at an academic center was the single strongest predictor, more than tripling the odds of receiving Mohs. Head and neck location also favored Mohs, which makes intuitive sense given the tissue-sparing rationale. On the flip side, Black race and lack of insurance were both associated with significantly lower odds of receiving Mohs — a disparity pattern that unfortunately echoes what's been documented across other skin cancers. The mortality analysis is where this study earns its keep. In the fully adjusted Cox model, Mohs surgery was associated with about a fifty percent reduction in five-year mortality compared to wide local excision — a hazard ratio around zero point five — and this was statistically significant even after controlling for age, comorbidity burden, tumor grade, insurance status, and facility type. That's a clinically meaningful signal, not just a statistical footnote, especially since prior SEER-based studies had given genuinely mixed results on this exact question, some showing a Mohs survival benefit and others showing none. The authors reasonably attribute their more consistent finding to the larger sample and broader covariate adjustment. Beyond the surgery-type comparison, the other independent mortality predictors were unsurprising but worth noting: age over sixty-five, male sex, meaningful comorbidity burden, prior malignancy, public or no insurance, non-academic treatment setting, and undifferentiated tumor grade all independently predicted worse survival. The authors are careful to point out that these socioeconomic and structural factors — insurance status, academic versus non-academic care — are themselves independent predictors of death, meaning survival here isn't just about which scalpel technique was used; access and care setting matter enormously in parallel. Limitations are the ones inherent to any tumor registry study: no recurrence data at all, no ability to isolate disease-specific survival, and the coarse clinical granularity typical of NCDB coding — margin status beyond the one-centimeter WLE definition isn't available, and there's no central pathology review. So for your practice, here's the honest framing: this is a large, well-adjusted, but still retrospective and registry-based confirmation that Mohs likely does carry a real survival advantage for DFSP, on top of its already well-established tissue-sparing benefit. It's not a randomized trial and never will be for a tumor this rare, but the consistency of the mortality signal here, in the largest NCDB analysis of DFSP to date, is a reasonable piece of ammunition when you're advocating for Mohs as first-line therapy in this disease, particularly when you're pushing back against insurance denials or facility-level barriers to referral. Last up, and the most classically "actionable" piece this month: a phase three randomized, vehicle-controlled, double-blind, multicenter trial of red light photodynamic therapy using ten percent aminolevulinic acid gel — brand name Ameluz, delivered as BF-200 ALA — for superficial basal cell carcinoma. This is a full original clinical trial, so it's worth walking through properly. The background is familiar: surgical excision remains the gold standard for BCC, but it comes with cosmetic tradeoffs, and there's sustained demand for effective noninvasive alternatives, particularly for low-risk superficial lesions. Photodynamic therapy is already guideline-endorsed for this indication, and ten percent ALA gel is already approved in the US for actinic keratosis and in Europe for a broader set of indications including superficial and nodular BCC. What this trial adds methodologically, and the authors are explicit about this, is that prior studies of ALA-PDT for BCC generally relied on clinical assessment alone for clearance; here, the entire lesion was surgically excised twelve weeks after the final PDT cycle and assessed histologically at a central dermatopathology lab, giving a much more rigorous, biopsy-confirmed clearance endpoint rather than relying on visual inspection alone. The design: twenty-one US centers, adults with at least one biopsy-confirmed, treatment-naive superficial BCC, randomized four-to-one to active gel versus vehicle, stratified by whether they had a single lesion or multiple. Participants got up to two PDT cycles, each consisting of two illuminations a week or two apart, with the primary target lesion pre-selected for excision and histology at week twelve. Randomizing at four-to-one rather than one-to-one is a common efficient design choice in a trial like this where the biological plausibility and prior clearance data already strongly favor the active arm — it maximizes safety and secondary-endpoint data on the active treatment while still preserving a valid vehicle comparison for the primary efficacy claim. The statistical plan used a stringent one-sided significance threshold for the primary and key secondary endpoints, tested hierarchically, which is a fairly conservative regulatory-grade approach appropriate for a pivotal trial intended to support an indication expansion. The results were unambiguous. Composite clinical-plus-histological clearance of the main target lesion — the most rigorous endpoint — was about two-thirds with active gel versus under five percent with vehicle, a massive and highly significant difference. Breaking that down, histological clearance alone was roughly three-quarters with active treatment versus about one in five with vehicle, and clinical clearance alone was around eight in ten versus about one in five. All of these differences were highly statistically significant, and given the size of the gap between arms, they're clinically meaningful as well — this isn't a case of statistical significance without clinical relevance. Clearance was consistent regardless of whether the lesion was on the trunk or the extremities, though it did trend downward as lesion size increased, dropping from around ninety percent clearance for the smallest lesions to somewhere in the low seventies for the largest ones — a sensible dose-response pattern that has obvious implications for patient selection. On the patient experience side, close to nine in ten participants treated with active gel rated the cosmetic outcome as good or very good, and the safety profile revealed no new or unexpected adverse events beyond what's already known for this formulation. The authors' own limitations are worth repeating rather than glossing over: relatively few study lesions were located on the face or scalp, so the generalizability of these clearance rates to cosmetically sensitive facial superficial BCCs is less certain, and the sixty-month long-term follow-up — which will matter enormously for durability of clearance and true recurrence rates — was still ongoing at the time of publication, so we don't yet have long-term recurrence data to weigh against surgical outcomes. Practically, this one is closer to practice-relevant than merely interesting, with the caveat that long-term durability is still pending. For appropriately selected low-risk superficial BCCs — good candidates being smaller lesions, non-facial locations, patients who are poor surgical candidates or who strongly prioritize cosmetic outcome — this trial gives you a rigorous, histologically confirmed clearance rate in the mid-seventies to talk through with patients as a real noninvasive alternative, not just an anecdotal one. It doesn't change your approach to high-risk histology, recurrent disease, or anything requiring margin control, and the durability question means you'll still want to counsel patients that long-term recurrence data are pending before treating this as equivalent to surgical clearance. That wraps our four articles for December. To sum up the throughline: caution and human oversight around AI in our literature and in our clinics, a sobering reminder that general-purpose AI still overcalls benign lesions in dermoscopy, reassuring national-level confirmation that Mohs likely improves survival in DFSP on top of its tissue-sparing benefits, and solid new trial-grade evidence for ALA-PDT as a noninvasive option for superficial BCC. Thanks for listening, and we'll see you next month.