A curious problem has begun appearing in my dermatopathology practice.
I routinely send clinicians histopathology reports accompanied by good-quality representative microscopic images. Increasingly, some of these images are being uploaded into ChatGPT and the diagnosis questioned on the basis of what it says.
One recent case was particularly instructive.
Clinically, there was good reason to suspect a nerve-related process. A prominent cord-like structure extended along a limb, and Hansen’s disease was among the clinical considerations. The question was therefore quite reasonable: could the biopsied structure represent a thickened nerve?
The histology, however, did not show peripheral nerve. It showed fibromuscular tissue.
Figure 1. Histology of the structure clinically thought to represent a thickened nerve. (A) Low-power architecture showing a broad spindle-cell proliferation rather than nerve fascicles. (B–C) Intersecting and whorled smooth-muscle bundles with intervening collagen and clefting. (D) Bland spindle cells with eosinophilic cytoplasm and elongated blunt-ended nuclei, confirming the morphological appearance of smooth muscle. No peripheral nerve architecture is identified.
This was not a subtle distinction between two closely related diseases. There was a more basic question.
Was this nerve at all?
The images were uploaded to ChatGPT with the clinical premise that this could be a thickened nerve in suspected Hansen’s disease. Its answer was remarkably confident.
It described a “multifascicular nerve trunk” with “pronounced expansion by dense collagen.” The supposed abnormalities included “perineurial thickening,” “endoneurial fibrosis” and “microfasciculation,” leading to a diagnosis of severe chronic fibrosing neuropathy.
Hansen’s disease had already been placed in the prompt. ChatGPT then took the interpretation further and argued that chronic or end-stage neural leprosy was a leading consideration. It cited published literature as support for fibrosis and microfasciculation in leprous nerves and suggested Fite-Faraco staining and M. leprae PCR.
It even proposed wording for the pathology report.
The answer sounded convincing. It was detailed, internally coherent and written in the language of a specialist opinion, complete with references and suggestions for further investigation.
Except that the biopsy was not nerve.
When I pointed this out, ChatGPT immediately agreed. The same microscopic appearances were now explained differently: the supposed nerve fascicles were muscle bundles cut in different planes, and there was no convincing peripheral nerve architecture.
Nothing in the images had changed.
The discussion then became more interesting. Could this be a leiomyoma? Tendon was raised as another possibility, and then a thickened muscular cord. Each new proposition produced another reconsideration and another plausible pathological explanation.
That, in my view, was the important part. This was more than an incorrect diagnosis. It was diagnostic instability presented with considerable confidence.
When the prompt becomes part of the pathology
Clinical information is indispensable in dermatopathology. But it should inform the interpretation; it cannot be allowed to determine what structures are present on the slide.
If the biopsy does not contain nerve, a strong clinical suspicion of neural disease does not turn fibromuscular tissue into nerve. One has to consider whether the intended structure was sampled, whether the palpable cord represents something else, or whether further investigation is required.
That is ordinary clinicopathological correlation.
Conversational AI introduces an unusual problem because the clinical hypothesis is embedded directly within the question. In this case, “thickened nerve in suspected Hansen’s disease” appears to have strongly influenced the first interpretation, to the extent that structures resembling neither conventional nerve fascicles nor normal peripheral nerve architecture were described using the vocabulary of neuropathology.
Yet the opposite problem appeared as soon as the premise was challenged. The model was remarkably willing to accommodate the new suggestion too.
A genuine second opinion should sometimes resist the person asking for it.
A dermatopathologist looking at these sections should be capable of saying, quite simply: “No. Whatever the clinical suspicion, I do not think this is nerve.”
There is nothing sophisticated about that sentence, but the ability to reject a leading premise is part of diagnostic expertise. A system that repeatedly accommodates the latest suggestion may produce an excellent discussion while still being a poor second opinion.
The dangerous answer is the plausible one
Much of the discussion around medical AI has concentrated on hallucination. Obvious nonsense is certainly a problem, although in practice it is often the easier problem to recognise.
What worries me more is the wrong answer that sounds exactly like a specialist answer.
Here, the terminology was appropriate and the reasoning sounded sophisticated. References were supplied. The model went on to recommend ancillary investigations and then produced language that could have been copied into a pathology report.
For a clinician who does not routinely interpret histopathology, this can look very much like an independent expert opinion. There is no obvious warning label within the prose itself telling the reader that the entire interpretation began with a basic morphological error.
And fluency makes the problem worse. A poorly written wrong answer invites suspicion; an articulate explanation, particularly one containing the expected terminology and references, may do the opposite.
ChatGPT is not a second dermatopathologist
None of this means that ChatGPT has no useful role in pathology. It can be helpful for teaching and terminology, and for discussing differential diagnoses. Purpose-built computational pathology systems may also prove extremely valuable for particular validated tasks.
But a general-purpose multimodal chatbot examining selected photomicrographs is not another specialist reviewing the slides. I think this distinction is becoming important now that clinicians can upload an image within seconds and receive something that reads like a formal pathological opinion.
If ChatGPT raises a possibility that conflicts with a pathology report, ask the reporting pathologist.
If the discrepancy persists, the specimen can be reviewed and additional sections or ancillary studies obtained. Where appropriate, re-biopsy the lesion or ask another dermatopathologist to review it. That is how a genuine second opinion is obtained.
What concerned me in this case was therefore not simply that ChatGPT got the diagnosis wrong. Pathologists get diagnoses wrong too. The more peculiar feature was that the model could be confidently wrong, accept a correction with equal confidence, and then continue producing different convincing explanations as the proposition changed.
Technical language can give an answer the appearance of diagnostic reasoning without providing the resistance and judgement that diagnosis sometimes requires. For now, when a chatbot interpretation conflicts with the pathology report, I would regard that conflict as a reason to go back to the slides and the pathologist—not as a reason to keep prompting until the model arrives at a story that sounds satisfactory.





