A conversation about clinical intuition, pattern recognition and what machines still cannot see
In one of my previous Substack posts, I described how ChatGPT correctly diagnosed an extremely rare dermatopathological condition after being shown the histopathology images.
This does not mean that ChatGPT has suddenly become a dermatopathologist. It gets several cases wrong, sometimes quite badly. What interested me was that, in this particular case, it recognised a highly unusual pattern that many human observers might have missed.
The post led to an interesting discussion in a dermatology WhatsApp group with a senior colleague, a renowned dermatologist in India.
He made the point that clinical diagnosis is not simply about fitting symptoms and signs into predefined criteria. Experienced clinicians sometimes “smell a rat”.
A genital ulcer may somehow not look like an STI. A dermatitis may not behave like ordinary contact dermatitis. Something about a presentation may make one suspect tuberculosis, Hansen disease, a drug eruption, connective tissue disease or dermatitis artefacta.
That feeling may prompt the clinician to ask one specific question, order a particular test or reconsider an apparently obvious diagnosis.
He called this clinical intuition.
I agree that this phenomenon exists. My only disagreement is over what we call it.
Is it really intuition?
The word “intuition” can make it sound as though an experienced clinician possesses some mysterious faculty that younger doctors do not.
In most clinical situations, that is probably not what is happening.
Years of practice expose us to thousands of combinations of morphology, colour, texture, distribution, symptoms, behaviour and context. Gradually, these patterns become so deeply ingrained that the mind processes them without consciously listing every clue.
The conclusion appears first. The explanation may come later.
Sometimes the explanation does not come at all.
Medical education literature describes this as experience-based, non-analytical reasoning or pattern recognition. The expert rapidly recognises similarities to previously encountered cases, often before beginning a deliberate step-by-step analysis.[1]
This does not mean that no reasoning occurred. It means that much of the reasoning happened automatically, below conscious awareness.
What we call clinical intuition may often be diagnostic reasoning that has become too fast and automatic for us to notice each intermediate step.
I remember a boss in the UK who, while assessing pigmented lesions, would occasionally say:
“I don’t know why, but this smells of melanoma.”
He obviously was not literally smelling melanoma. Nor was he using some extrasensory faculty.
He was probably noticing a combination of subtle clues that years of experience had taught him to recognise. Perhaps one colour looked slightly out of place. Perhaps the symmetry was disturbed in a way that was difficult to define. Perhaps the border, surface or overall architecture did not fit his internal picture of a benign lesion.
He was certainly an excellent diagnostician.
However, when I asked him to explain exactly what he had noticed, he sometimes could not.
That made him a very good diagnostician, but perhaps not always an equally good teacher.
When expertise becomes mystique
There is a problem when experienced clinicians repeatedly describe their decisions as a “feeling”.
It can create an unnecessary aura around them. Juniors begin to think that the senior clinician possesses some special diagnostic sense that they themselves lack.
The boss has a feeling. The trainee does not. The trainee therefore feels inadequate.
But the senior clinician’s feeling is usually not magic. It is accumulated experience compressed into rapid pattern recognition.
A good teacher should at least try to unpack it.
What looked unusual?
Which feature did not fit?
Why was the obvious diagnosis unsatisfactory?
What prompted that particular investigation?
What alternative was being considered?
An expert may not always be able to retrieve every subconscious step. That is understandable. But “intuition” should not become a convenient substitute for explanation.
Teaching does not merely mean announcing the correct diagnosis. It means making at least some part of the reasoning transferable.
Could an LLM like CHATGPT acquire this kind of intuition?
This brings us back to the original question.
Could an LLM (Large Language Model) eventually reproduce this so-called clinical intuition?
Perhaps, ironically, yes.
If intuition is the subconscious integration of numerous subtle clues learnt from previous cases, then an LLM trained on enough reliable examples might eventually recognise similar combinations.
It too may learn to “smell a rat”.
Image-based AI is already remarkably good at recognising patterns within defined datasets. Research in digital pathology and skin-lesion classification has shown considerable promise, although many studies still have problems involving dataset quality, selection bias, limited external validation and artificial test settings.[2,3]
Recognising a pattern in a selected image is also not the same as independently handling a real patient, or even a complete biopsy case.
Still, the underlying principle is not entirely different.
The experienced clinician has learnt from thousands of previous cases. The model has also learnt from previous examples. Both may recognise a familiar configuration before breaking it into individual components.
The machine may even appear to have one advantage. We can ask it to list the features it noticed, explain why it preferred one diagnosis and tell us what alternatives it considered.
But this advantage needs an important qualification.
An LLM can produce a convincing explanation even when that explanation does not faithfully represent how it reached the answer. It may identify the correct diagnosis first and then construct a plausible justification afterwards.
In other words, an articulate explanation is not necessarily a genuine audit trail.
This matters because fluency can create an illusion of transparency. A model that explains itself confidently may appear more accountable than a clinician who says, “I just have a feeling.” Yet both may be offering a post hoc account rather than revealing the true process that produced the conclusion.
We should therefore judge an AI explanation not merely by how sensible it sounds, but by whether the cited features are actually present, discriminating and reproducibly linked to the diagnosis.
If a machine can consistently do that, then what appears to be intuition may simply be very fast diagnostic reasoning.
A slide is not a patient
This is where the difference between dermatopathology and clinical dermatology becomes important.
A histopathology slide is complex, but the problem is relatively contained. The model sees tissue architecture, cells, inflammatory patterns, staining characteristics and their spatial relationships.
A slide does not arrive tired, anxious, jaundiced, embarrassed or evasive.
A patient does.
Within seconds of someone entering the consultation room, a clinician takes in far more than the lesion being presented.
The human mind is almost like an organism with multiple tentacles. While one tentacle examines the rash, others are simultaneously taking in the patient’s build, complexion, posture, movements, expression, voice, behaviour and interaction with accompanying relatives.
The patient may look fatigued, anorexic or cachectic. There may be mild jaundice, a prominent abdomen and rounded cheeks in someone with alcoholic liver disease. There may be proximal weakness or other Cushingoid features.
Subtle striae or facial telangiectasia may suggest prolonged topical steroid use, even before the patient admits to applying a potent steroid cream for acne.
We do not consciously decide to assess each of these features one by one. The mind scans the patient as a whole, often within a few seconds.
This complete visual, behavioural and emotional impression is what I might loosely call the patient’s “aura”.
I do not mean an aura in any supernatural sense. I mean the total impression carried by the patient into the room.
Someone may say that the disease is not troubling them, while their expression suggests fear. A patient may avoid eye contact when asked about medication use. A parent may answer every question before the adolescent is allowed to speak. Someone may appear embarrassed, defensive or strangely indifferent to a potentially serious diagnosis.
These are clinical clues too.
Empathy influences how we interpret them. It helps us recognise when a patient is frightened, ashamed, confused or withholding something because they fear being judged. It affects when we pause, when we ask a question differently and when we decide not to press immediately.
An LLM can generate empathic-sounding language. That is not quite the same as sharing the emotional space of a consultation or recognising what a particular patient is struggling to say.
More importantly, the model generally sees only the information it is given.
If it receives a close-up photograph of a rash, the rest of the patient is outside its field of vision. It cannot notice jaundice, cachexia, rounded cheeks, striae, anxiety or an unusual interaction with an accompanying relative unless someone captures those details and supplies them.
This may change.
Future multimodal systems could be connected to multiple high-definition cameras, microphones and other sensors. They might analyse the lesion, the entire skin surface, posture, gait, speech and facial expression simultaneously.
Technically, machines may eventually gather much more of the information that clinicians now absorb almost automatically.
But reproducing the richness of an actual consultation is a far harder problem than interpreting a selected image.
Why histopathology may be more approachable for AI
Histopathology is not easy. Nor is it independent of clinical information.
The patient’s age, biopsy site, duration of the lesion, clinical differential diagnosis, biopsy technique and treatment history can completely alter how a slide is interpreted.
A poorly sampled biopsy can mislead both the human pathologist and the machine.
However, the information within the slide itself is more bounded. The slide does not hesitate, minimise symptoms or conceal its treatment history out of embarrassment.
It simply presents its morphology.
This may partly explain why AI can sometimes perform impressively in histopathology. The task is difficult, but the visual field is defined.
Clinical medicine is less tidy. The relevant clue may not be in the photograph, the written history or the answer to the question that was asked.
It may be elsewhere in the room.
So will AI ever develop intuition?
I suspect that AI will eventually reproduce much of what we currently call intuition, particularly in bounded, pattern-heavy fields such as radiology, dermatology and histopathology.
If human intuition is largely the subconscious recognition of subtle patterns learnt through repeated exposure, there is no obvious reason why machines cannot learn some version of it.
The larger challenge is giving the machine access to all the information that the clinician naturally gathers.
It must not merely inspect the lesion. It must learn where else to look.
It must recognise which clues matter, which are distractions and which missing questions need to be asked.
It must also know when its pattern recognition is unreliable and when slower, more analytical reasoning is required.
The irony is that if an LLM does develop something resembling intuition, we will immediately demand that it explain itself.
Once it identifies the clues, lists the alternatives and reliably shows how it reached its conclusion, perhaps we will no longer call it intuition.
We will simply call it reasoning.
And if ChatGPT one day says, “I don’t know why, but this smells of melanoma,” perhaps we will have come full circle.
For now, the human clinician still has an important advantage.
Not because we possess supernatural intuition, but because we see the patient, not merely the data.
References
Norman G, Young M, Brooks L. Non-analytical models of clinical reasoning: the role of experience. Medical Education. 2007;41:1140–1145.
McGenity C, Clarke EL, Jennings C, et al. Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy. NPJ Digital Medicine. 2024;7:114.
Haggenmüller S, Maron RC, Hekler A, et al. Skin cancer classification via convolutional neural networks: systematic review of studies involving human experts. European Journal of Cancer. 2021;156:202–216.





