What large language models actually do, and what professionals need to know about it.
There is a version of the conversation about AI in professional settings that treats scepticism as a failure of imagination, as though those who ask hard questions about AI reliability are simply behind the curve, waiting to be convinced by better demonstrations. This piece argues the opposite. Informed professional scepticism about AI is not a position to grow out of; it is a professional competency.
This is not an argument against AI use in professional settings. It is an argument for understanding what AI systems actually do, mechanically, not metaphorically, before using them to produce outputs that will influence professional decisions, client communications, legal documents, or clinical assessments.
What a large language model actually does
The term ‘artificial intelligence’ is, in the context of current language models, significantly misleading. It implies reasoning, understanding, and cognition that these systems do not possess. What a large language model does is substantially simpler and substantially stranger than what the word ‘intelligence’ suggests: it predicts, token by token, the most probable continuation of the text presented to it, based on statistical patterns learned from an enormous corpus of training data.
This is a remarkable technical achievement. The outputs of state-of-the-art language models are often fluent, contextually appropriate, and superficially impressive. The problem is that fluency and accuracy are not the same thing. A system that generates text by predicting what words are likely to follow what other words, based on patterns in data, will produce authoritative-sounding text whether or not it is correct. It has no mechanism for distinguishing truth from falsehood. The 2023 case in which New York lawyers submitted a brief citing six fabricated cases, all generated by ChatGPT, none of them real, is the most widely reported illustration of this failure mode. It will not be the last.
The calibration problem
The specific challenge AI presents to professional users is not that it is frequently wrong, human professionals are also frequently wrong. It is that the confidence of AI outputs does not correlate reliably with their accuracy. Research published in Nature and elsewhere has examined calibration across a range of factual and reasoning tasks, consistently finding that language models express confidence in incorrect outputs at rates that preclude using expressed confidence as a reliability signal. For professional users, this creates a specific and underappreciated risk: when they use an AI output without adequate verification, they inherit the AI’s lack of calibration.

The EU AI Act and the UK’s regulatory position
The regulatory landscape for AI is evolving faster than most professional training has been able to absorb. The EU AI Act, which came into force in 2024 and began applying from 2025, introduces a risk-based classification system for AI applications. High-risk categories include AI used in employment, education, essential services, law enforcement, and administration of justice. The United Kingdom has taken a different approach: rather than a single horizontal Act, the government has adopted a sector-led model in which existing regulators apply existing frameworks to AI within their domains. The NCSC’s guidance on AI security, the ICO’s guidance on AI and data protection, and sector-specific guidance from the FCA, CQC, and SRA are all relevant depending on professional context.
The professional accountability question
Regulatory frameworks aside, there is a professional accountability question that no AI tool resolves. When a legal professional uses AI to draft an advice note that contains a material error, the professional remains responsible for the error. The Solicitors Regulation Authority has begun addressing this directly. Its warning notice on the use of AI makes clear that solicitors are responsible for verifying AI-generated content before relying on it, and that the SRA’s standards, including obligations around competence, honesty, and client care, apply equally to AI-assisted work.
What informed professional use looks like
The applications where current AI is genuinely strong are those that involve language manipulation rather than factual accuracy or professional judgement: drafting, summarising, reformatting, translating, restructuring. An AI that produces a draft that requires professional review and verification is functioning appropriately. An AI that produces a final output that bypasses professional review is not.
The applications where current AI is genuinely weak, and where the risk of uncritical reliance is highest, are those that require accurate factual knowledge, reliable probabilistic reasoning, and the exercise of professional judgement. Legal research, clinical diagnosis, risk assessment, and financial advice are all domains where the failure modes of current AI systems are precisely the failure modes that professional standards exist to prevent.
The practical framework for any professional considering AI use is straightforward, even if the application is not. Identify the specific claim or output that will influence a professional decision. Assess how verifiable that claim is. Verify it, against authoritative sources, not against other AI outputs. Document the verification. The burden is on the professional who relies on AI output to demonstrate that they have exercised the judgement that their professional status requires.
Further reading
EU AI Act (2024). Full text and guidance.
ICO. Guidance on AI and data protection.
NCSC. Guidelines for secure AI system development.
SRA. Using AI in legal practice: warning notice.
Bender, E. et al. (2021). On the Dangers of Stochastic Parrots. Proceedings of FAccT 2021.