
[dropca]E[/dropcap]very month, NHS England publishes a tidy set of numbers about general practice. Thirty million appointments. Neat breakdowns by face-to-face, telephone, online. Trend lines. Practice-level dashboards. It looks like a well-measured system.1
It isn’t.
The publication carries a quiet caveat that almost no-one reads. The data, in NHS England’s own words, “…does not show the totality of GP activity/workload” and “…does not assess the complexity of activity.“1 The publisher of the national workload metric is telling us the metric is incomplete.
That is not a small concession. This is the number that drives contract negotiations, workforce planning, commissioning, and the political conversation about general practice. We’re managing a workforce crisis with a yardstick its own producers admit cannot do the job.
What we’re missing
We’re managing a workforce crisis with a yardstick its own producers admit cannot do the job.
The shape of GP work has changed in ways the appointment count can’t see. Whittaker and colleagues’ BJGP analysis of UK primary care across 2005 to 2019 found that direct consulting rates fell modestly while indirect and administrative work nearly tripled. Admin now accounts for around 30% of total GP workload.2 The proportion of consultations involving patients with three or more chronic conditions rose from under 10% to over 16%.
A consultation involving multimorbidity, polypharmacy, mental health, social vulnerability, and a safeguarding concern within fifteen minutes is not the mathematical equivalent of a routine medication review. But to the appointment system, both are one tick in the box.
And the indirect work sits entirely outside the count. Results, referrals, prescription queries, letters, supervision, and the meeting that should have been an email are all invisible. The BMA recommends a 3:1 clinical-to-admin ratio.3 Most of us recognise that as theoretical. The clinical always wins. The admin happens in the gaps that don’t officially exist.
Why this matters for safety and system finances
The BMA’s 2024 Safe Working handbook recommends a maximum of 25 patient contacts per GP per day, 15-minute appointments, and no more than three hours of patient-facing time per session.4 The 2024 RCGP GP Voice survey found that 76% of GPs believe patient safety is actively compromised by their workload.5 Hodkinson and colleagues’ BMJ meta-analysis of 170 studies and over 230,000 physicians found that burnt-out clinicians have roughly twice the odds of being involved in patient-safety incidents.6
The 25-contact ceiling may be adjusted when sessions are dominated by complex cases. In the executive summary, on page 2 of the BMA’s 2024 Safe Working handbook, the exact sentence reads: “It recommends a maximum of 25 patient consultations daily to prevent overload and ensure that each patient receives adequate time and attention. This can be adjusted based on the complexity of cases and individual circumstances within the practice“. The BMA’s own phrasing is that the limit “can be adjusted based on the complexity of cases,” rather than that it should reduce. Without a shared way of measuring complexity, that adjustment is left to the individual clinician, applied inconsistently, or not applied at all.7
But there is a financial argument here too, and commissioners should pay attention to it. When we fail to measure complexity, we fail to price it. An unmeasured “high complexity” psychosocial consultation doesn’t just represent a burden on the GP; it represents cost-avoidance for the wider system. When an exhausted GP breaches an invisible safe-working limit and stops absorbing that unresourced complexity, the patient defaults to A&E, crisis teams, or acute admission. The evidence is clear that better access to primary care is directly associated with fewer emergency admissions. [3] GP workload tracking isn’t just about defending the clinician. It’s about quantifying the invisible safety net that prevents secondary care collapse.
Workload isn’t one number
The simplest model of workload is intuitive: time multiplied by complexity. The trouble is that complexity isn’t one thing. Working clinicians know this, and the case-mix literature confirms it. What looks like a single variable breaks down into at least six partially overlapping dimensions:
- Clinical complexity: the number, severity, and interaction of active medical problems.
- Psychosocial complexity: domestic abuse, safeguarding, suicidal ideation, family crisis.
- Diagnostic uncertainty: the undifferentiated symptom that may be nothing or may be serious.
- Cognitive switching cost: moving rapidly between unrelated problems and modalities.
- Interruption frequency: duty calls, urgent queries, and results requiring action mid-consultation.
- Medico-legal and emotional risk: the consultation requiring careful documentation because of a possible complaint.
Consider two patients. Patient A has stable diabetes, COPD, and chronic kidney disease, requiring a polypharmacy review. Patient B presents with domestic abuse, acute suicidal ideation, and a safeguarding concern about her son. A disease-count algorithm would score Patient A as more complex. Every GP reading this knows Patient B is, by an order of magnitude, the more demanding consultation. A workload model that can’t distinguish the two has reproduced rather than solved the problem.
Can we measure this reliably and will it be gamed?
Even with the dimensions properly named, a hard question remains. Can clinicians rate them consistently?
Commissioners will argue that subjective complexity scoring is open to gaming, assuming an exhausted workforce will naturally inflate complexity to protect their time. That’s a valid risk. But we should ask whether the current alternative is any less distorted. Right now, a 10-minute multimorbidity consultation and a 10-minute sore throat are mathematically identical. A system that accepts a known margin of error in complexity-adjusted tracking is vastly safer, and more honest, than a system that insists complexity doesn’t exist at all.
Any honest framework has to name the dimensions being rated, offer brief operational definitions, provide anchored examples to align judgement across clinicians, and quantify the residual variation rather than pretend it isn’t there. It also has to be used carefully for individual reflection and aggregate analysis, not for inter-clinician performance management. Two GPs scoring the same consultation won’t mean exactly the same thing by “moderate.” That’s a feature to be managed, not a reason to abandon the attempt.
A starting point: the Contractual Time Equivalent
The institutional desire for this data already exists. The NHS Lothian Primary Care Workload Toolkit explicitly asked practices to stop presenting anecdotes and to start capturing un-resourced admin pressure through structured recording.9 The BMA itself has endorsed it as a model. But asking a burned-out workforce to manually tally every Docman letter on a spreadsheet is a doomed methodology. We need a frictionless, digital translation of what the LMC toolkits are trying to do.
I propose that as a starting point we adopt a Contractual Time Equivalent (CTE), a unit of workload built from time and a clinician-assigned complexity weight. In its simplest form, each activity carries a base time and a three-tier multiplier (Low, Moderate, High). A clinical session is treated as a finite budget, such as 210 minutes, with explicit clinical, admin, and CPD allocations. As consultations and tasks are logged, the budget runs down. Where logged clinical work exceeds the clinical allocation, the framework names the displaced admin as an ‘administrative deficit.’ This is work that currently goes, silently, into unpaid time.
CTE isn’t a finished instrument. It’s a hypothesis-generating model. Its value lies not in any claim to measurement precision, but in forcing an explicit conversation about what GP workload actually consists of. It would need staged empirical validation through a Delphi process with GPs of mixed experience, feasibility and inter-rater reliability testing, and construct validity comparison against external measures of fatigue and overtime. That’s achievable for a clinician-researcher on modest funding.
What I’m not claiming
I’m not arguing that a personal work-logging app should drive contract negotiations. It shouldn’t, until it’s been validated against independent benchmarks. I’m not arguing that complexity scoring should be used to compare one GP’s output to another’s because the variation in judgement makes that misleading. And I’m not arguing that better measurement will fix the GP workforce crisis on its own. It won’t.
I propose that as a starting point we adopt a Contractual Time Equivalent (CTE), a unit of workload built from time and a clinician-assigned complexity weight.
What I’m arguing is more modest. The metric we rely on is, by its own admission, incomplete. The BMA’s safe-working guidance is an undefended ceiling without a shared way to measure when complexity ought to lower it. The dimensions that most consume our days, such as cognitive load, interruption frequency, and medico-legal burden, have no place in any national dataset.
That’s a real gap. Filling it requires validated workload instruments, multi-site studies, and independent governance. It requires research infrastructure. And it requires political will.
A small claim, honestly held
I’ve started logging my own work not because I think my numbers are the truth, but because paying attention has already changed how I work. That itself is interesting, and it complicates any claim that self-measurement is neutral observation. The Hawthorne effect (defined as people modifying their behaviour in response to being observed)10 is real in this instance, and I’ve noticed it in myself.
But the broader point survives. What we count shapes what we see. And what we’re currently counting isn’t enough to keep patients safe, to keep clinicians in the profession, or to support an honest conversation about what general practice actually involves.
We should keep asking the question.
References
- https://digital.nhs.uk/data-and-information/publications/statistical/appointments-in-general-practice/may-2026 [accessed 29/6/26]
- Whittaker W, Hadi M, Sutton M, et al. Trends in clinical workload in UK primary care 2005–2019: a retrospective cohort study. Br J Gen Pract 2024;74(747):e659. DOI: 10.3399/BJGP.2023.0527
- https://www.bma.org.uk/pay-and-contracts/job-planning/job-planning-process/salaried-gp-model-job-planning-templates-and-guidance [accessed 29/6/26]
- https://www.bma.org.uk/advice-and-support/gp-practices/managing-workload/safe-working-in-general-practice/daily-working-contacts [accessed 29/6/26]
- https://www.rcgp.org.uk/News/GP-workloads-threaten-patient-safety [accessed 29/6/26]
- Hodkinson A, Zhou A, Johnson J, et al. Associations of physician burnout with career engagement and quality of patient care: systematic review and meta-analysis. BMJ 2022;378:e070442. DOI: 10.1136/bmj-2022-070442
- https://www.bma.org.uk/media/a31pbfv1/bma-gp-safe-working-guidance-handbook-england.pdf [accessed 29/6/26]
- Cowling TE, Cecil EV, Soljak MA, et al. Access to primary care and visits to emergency departments in England: cross-sectional, population-based analysis. PLoS ONE 2013;8(6):e66699. DOI: 10.1371/journal.pone.0066699
- NHS Lothian Quality Improvement. Primary Care Workload Toolkit. Available from: www.qilothian.scot.nhs.uk/pc-toolkit-workload [accessed 29/6/26]
- https://catalogofbias.org/biases/hawthorne-effect/ [accessed 29/6/26]