Welcome back to this July twenty twenty-six run-through of the Journal of the American Academy of Dermatology. We've got four pieces this month — a health policy piece on sunscreen regulation, a brief report on an EMR documentation tool, a diagnostic accuracy study on large language model bias in melanoma triage, and a letter to the editor pushing back on a cSCC imaging study. Let's get into it. First up is a health policy and practice piece — not original research, more of a regulatory update with clinical commentary — on the long-stalled state of sunscreen innovation in the United States and what's finally changing. The framing here is one you've probably grumbled about yourself: US sunscreens have lagged behind European and Asian formulations for decades. Sunscreens got reclassified as over-the-counter drugs back in the nineteen-seventies, which locked filter approval into drug-level regulatory pathways, and the FDA hasn't updated allowable filter concentrations since nineteen ninety-nine. A whole regulatory effort in the two-thousands — the so-called Time and Extent Application — came and went without adding a single new category-one filter, meaning an ingredient generally recognized as safe and effective that can be marketed under the OTC monograph without a full new drug application. Then in twenty-nineteen, the FDA's proposed rule only accepted zinc oxide and titanium dioxide as GRASE, right as absorption studies were generating public anxiety about chemical filters. Legislative attempts since then — the Sunscreen Innovation Act, provisions under the CARES Act — didn't translate into actual new filter approvals by twenty twenty-five. The news here is that Congress passed the SAFE Sunscreen Standards Act in November twenty twenty-five, directing the FDA to use real-world evidence and non-animal testing to speed up filter evaluation, and just weeks later, in December twenty twenty-five, the FDA proposed adding bemotrizinol to the OTC monograph at concentrations up to six percent. If finalized, this would be the first new category-one UV filter added under this modernized framework. Clinically, bemotrizinol isn't new to the world — it's been used in Europe, Australia, and parts of Asia for over two decades. Its selling points are a large molecular size that limits systemic absorption, dual absorption peaks around 310 and 340 nanometers giving genuine UVB and UVA1 coverage, and high photostability. That UVA1 durability is the real practical gap it fills, since a lot of current US formulations lean on avobenzone, which is UVA-active but photounstable without stabilizers. For your patients with photosensitive dermatoses — cutaneous lupus, polymorphous light eruption — or those chasing pigmentary and photoaging concerns, better UVA1 coverage is clinically relevant, not just a marketing point. It may also give you another option for patients with contact allergy to existing organic filters. The authors' practical takeaway, and I'd agree, is that this is worth flagging to patients proactively once these products hit shelves, but that your core counseling doesn't change — broad-spectrum, adequate SPF, reapplication, and photoprotective behavior remain the fundamentals. The equity angle they raise is worth remembering too: expanded filter access only helps if manufacturers actually produce affordable, cosmetically elegant formulations across skin tones — white cast and texture are still the biggest reasons patients with skin of color abandon sunscreen, and a new filter on paper doesn't fix that unless the finished products address it. So — interesting and something to watch for on your shelves later this year, not something that changes what you tell patients tomorrow. Second is a brief report — a quality improvement or workflow study, not a randomized design — looking at a custom Epic-based documentation module built by a dermatology department at the University of Pennsylvania, and its effect on documentation burden and after-hours EMR use, the so-called pajama time. The background problem is one every one of us lives with: Epic's default dermatology application is intentionally generic, so it typically requires heavy customization to actually fit specialty workflow, and a lot of institutions never get around to building that customization out. This group did — they created "DermModule," which adds disease-specific templated plans for common diagnoses, embeds patient education directly into the note so it's automatically available at the point of care instead of requiring separate handouts, and standardizes procedure note templates with structured dropdowns for rapid documentation. Methodologically, this is a pre-post observational study — they pulled Epic signal data, the built-in analytics Epic generates on clinician EMR use, for seventy-four active providers, comparing a seven-month window before implementation against the same seven-month window a year later after rollout. They used paired t-tests to compare each clinician's own before-and-after averages. The pre-post design here makes practical sense — you can't randomize an EMR module mid-department without contaminating your control group, since everyone's using the same shared system — but the authors are upfront that this leaves them vulnerable to temporal confounding, things like concurrent Epic platform updates, staffing changes, or documentation policy shifts that happened to coincide with rollout and could account for some of the effect independent of the module itself. The results were consistent across every efficiency metric they tracked. Time in notes per day dropped from about forty-seven minutes to about thirty-nine minutes, a meaningful and statistically significant drop. Time in notes per individual appointment fell from just under nine minutes to under eight, also significant. Pajama time — after-hours and weekend EMR use — dropped from about thirty-five minutes to twenty-eight minutes a day, again significant. And importantly, the number of appointments per day actually trended upward slightly, from roughly twelve to twelve-and-a-half, which didn't reach statistical significance but tells you productivity wasn't sacrificed to get these efficiency gains — if anything it moved the right direction. Averaged out, they estimate this saved each clinician somewhere around forty-eight hours of documentation time across the seven-month study window. The takeaway: this is a nicely pragmatic, low-cost intervention story rather than a landmark trial, and the effect sizes — several minutes per day, several minutes per note — are individually modest but add up to real burnout-relevant time savings at scale, especially the pajama-time reduction, which is the metric most tied to physician wellbeing. It's practice-relevant in the sense that if your own institution has an underbuilt dermatology Epic module, this is a concrete argument and template for investing in building one out — genuinely actionable if you have any influence over your department's EMR configuration, though the specific numbers are institution-specific and won't necessarily transfer one-to-one to your own build. Third is a diagnostic accuracy study examining framing bias in ChatGPT when it's asked to assess pigmented lesions — a nicely designed little experiment out of a group in Rome. The core question: large language models are increasingly being poked at for dermatologic image interpretation, and we know they can inherit cognitive biases similar to clinicians — anchoring, suggestibility, framing effects. This group asked specifically whether the narrative framing of a prompt, independent of the image itself, systematically shifts ChatGPT's melanoma risk scoring. Their method was clean and worth understanding because it's a good template for this kind of bias-testing study. They took one hundred dermoscopic images, all histopathology-confirmed — eleven melanomas and eighty-nine dysplastic nevi, all Fitzpatrick skin types one through three — and ran each image through ChatGPT-5 six separate times, once under a neutral baseline prompt asking for a one-to-five nevus-to-melanoma score, and then five more times with identical images but different narrative framing bolted on: a worried patient wanting a same-day appointment, a patient minimizing concern based on prior harmless lesions, an irrelevant detail about a nearby tattoo, an anxious spouse convinced it's melanoma, and a strongly reassuring frame. They used the model as a black box with no fine-tuning, deliberately, because that mirrors how a clinician or patient would actually interact with it off the shelf. They then computed area under the ROC curve, sensitivity, specificity, and positive predictive value for each framing condition, used sign tests with Holm correction to compare score distributions against baseline, and used Spearman correlation to see how consistently lesions were ranked relative to each other across the different frames. The results are the whole point of the paper: framing mattered, a lot. Compared to the neutral baseline, average lesion scores rose significantly under the worried-patient and minimizing-concern frames, and dropped under the strong-reassurance frame — meaning the model was pulled toward higher suspicion scores by anxious framing and toward false reassurance by confident-sounding framing, in both cases regardless of the actual image. Diagnostic discrimination swung substantially depending on framing — the area under the ROC curve ranged from about 0.56 under the strong-reassurance prompt, which is barely better than a coin flip, up to about 0.75 under the irrelevant-tattoo-detail prompt, which was actually the best-performing condition, better even than the neutral baseline's 0.72. Sensitivity stayed reasonably high across most frames, generally in the eighty-to-ninety percent range, but specificity was consistently poor, dropping as low as the mid-thirties percent under the reassurance and anxious-spouse frames, and positive predictive value hovered in the mid-teens to high-teens percent throughout — reflecting the low melanoma prevalence in the sample as well as the model's tendency to over-call. Scores across the different prompts were only moderately correlated with each other, in the range authors describe as 0.54 to 0.79 — meaning the model didn't completely reorder its risk ranking of lesions, but shifted absolute scores enough that the same lesion could cross a clinical decision threshold purely based on how the question was phrased. The authors' conclusion is appropriately restrained: this is a single-center pilot using one model version on a modest and imbalanced image set with only eleven true melanomas, so the point estimates themselves shouldn't be over-read, and there's no external validation cohort. But the qualitative finding — that identical images generate materially different risk outputs purely from narrative framing — is a legitimate and important caution. For you, this isn't practice-changing in the sense of altering how you manage patients today, since nobody's suggesting you replace your own judgment with ChatGPT scoring. But it is directly relevant if your patients start showing up having already run their own mole photos through a chatbot with some anxious or reassuring framing baked into how they asked the question — you should assume the output they got may not reflect the image alone, and it's a useful, concrete data point for anyone on a P and T or AI-governance committee evaluating LLM tools for patient-facing triage. Last is a letter to the editor, so no methods or results section of its own to walk through — this is a critical commentary responding to a previously published retrospective cohort study by Wei and colleagues on radiologic imaging in high-risk cutaneous squamous cell carcinoma, which had reported that imaging turned up unexpected findings in nearly half of high-risk cSCC cases and changed management in about forty-seven percent of patients. The letter writers, a group from Qingdao, argue that this headline finding is more fragile than it looks, and their critique is worth internalizing given how often "imaging changed management" numbers get quoted uncritically. Their first point is selection bias — since imaging in the original cohort was clinically driven rather than protocol-mandated, patients who got imaged were the ones already generating clinical suspicion, so of course that group had higher rates of nodal or distant disease; the imaging didn't necessarily create the finding, the pre-existing suspicion did, and the "management changed" figure may really just be measuring disease severity rather than the independent value of imaging itself. Their second point is that the original study never adjusted for confounders like tumor size, depth, or location, and notably the difference in five-year disease-related outcomes between imaged and non-imaged patients — thirty-six percent versus twenty-five percent — didn't reach statistical significance, nor did overall survival differ meaningfully between groups. Without multivariable or propensity-based adjustment, you can't distinguish "imaging improves outcomes" from "imaging just identifies patients who were already going to do worse." Their third point is more practical: the original recommendation to image "the local site and nodal basins" doesn't specify modality, and while CT and PET-CT dominated that cohort, ultrasound — increasingly the workhorse for nodal surveillance — was barely used, though the letter writers reasonably note the original study's two-thousand-ten to twenty-twelve enrollment period predates ultrasound's broader adoption for this indication. There's no new data here, so the takeaway is purely interpretive: treat the original study's forty-seven percent management-change figure as hypothesis-generating rather than as grounds for routine imaging in every high-risk cSCC, until someone does the prospective, confounder-adjusted, modality-specific version of this question. For your own practice, the sensible posture the letter implicitly supports is what most of us already do — reserve imaging for cases with genuine clinical concern for nodal or deep invasion rather than treating it as a blanket reflex for every BWH T2b or T3 tumor, pending better evidence. That wraps this month's four. Bemotrizinol's on its way onto US shelves and worth pre-briefing patients on, a homegrown Epic module shows a workable template for clawing back documentation time, ChatGPT's melanoma scoring is more swayed by how a question is asked than we'd like, and the cSCC imaging literature still needs a properly adjusted prospective study before "image everything high-risk" becomes real guidance. Thanks for listening, and I'll see you next month.