Electronic monitoring chain in England and Wales from order to response, showing a 95% attempt threshold, 62% fitted and an incomplete outcome record.
Case Study

Electronic Monitoring UK: What a 95% Tag-Fitting Attempt Target Actually Measured

This electronic monitoring UK case study focuses specifically on the service operating across England and Wales. Serco met a 95% timeliness threshold for electronic-monitoring installation attempts. Yet in the same reporting month, only 62% of the adults it visited were successfully fitted with a tag after the standard two attempts.

At first glance, those facts seem incompatible. How could the contractor responsible for electronic monitoring across England and Wales meet the contractual threshold while the fitting result remained so much lower?

The answer is not that one number was true and the other false. The 95% was the threshold for making two installation attempts on time. The 62% described whether the adults visited during those attempts actually ended up with a tag fitted.

They observed two different moments in a longer process.

That distinction opens this case. The electronic monitoring service did recover from a difficult transition. A large visit backlog was cleared, systems were migrated, and the contractor later met its contractual targets. But those achievements did not, by themselves, show whether every person who should have been monitored was fitted, actively communicating, and followed through to a recorded response when a potential breach occurred.

To understand what happened, I need to separate the stages rather than collapse them into one judgment.

The Recovery Was Real

When Serco and Allied Universal took over the electronic monitoring service for England and Wales on May 1, 2024, the transition did not go smoothly. The service had to manage fitting, checking, and removing devices while moving data and operations onto new systems. Full operational status arrived eleven months later than planned.

The National Audit Office review shows how quickly the strain became visible. Outstanding visits rose from roughly 1,200 to around 7,000 before falling below the tolerable level of about 400 after Serco and HM Prison and Probation Service concentrated resources on clearing the backlog.

That fall matters. It is evidence of genuine operational recovery, not a cosmetic improvement in reporting.

But the performance regime also changed during the transition. For a period, HMPPS formally assessed Serco against only four of the 14 contractual performance indicators, and timely tag fitting was not among the four. The full set was restored later through an improvement plan. By the end of that process, Serco was meeting all contractual targets.

This gives us two facts that should remain together. The early service was under severe pressure, and the later service was operating more reliably against the measures it had been given.

The analytical question is narrower: what did meeting those measures prove?

What the Electronic Monitoring Service’s 95% Target Actually Measured

The fitting target applied to adult orders. It required two installation attempts to be made by midnight on the day after an order was received in at least 95% of cases.

Serco met that threshold. In the same reporting month, it successfully fitted tags on 62% of the adults it visited within those two attempts.

I would not discard the contractual target. It tells us something important: the service was making the required attempts on time. It shows that a defined part of the process had become more dependable.

I also would not let that result answer a different question. An attempt can happen on time without ending in a successful fit. The person may not be at the address, may refuse the installation, may be in custody, or the order or equipment may present another problem. Some of those conditions sit outside the contractor’s direct control.

The target therefore measured timely activity. The fitting rate measured the immediate result of that activity. Neither figure cancels the other.

The apparent contradiction appears only when “an attempt was made on time” is treated as equivalent to “a tag was successfully fitted.” Once the two stages are separated, the figures become compatible—and more useful.

This is the first interpretive lesson in the case. A valid performance measure becomes an invalid basis for inference when it is used as evidence about a stage it does not observe.

Who Counted as Monitored

Successful fitting was still not the end of the chain.

An order had to be usable. A fitting visit had to occur. The required equipment had to be installed. That equipment then had to communicate with the monitoring system. If it generated an alert, information had to move to the appropriate authority, and that authority had to decide what action to take.

Different measures—and different organizations—covered different parts of that sequence, so the meaning of the evidence changed as it moved through the system.

This becomes especially important when reading the headline caseload. The Ministry of Justice’s official electronic monitoring statistics defined the caseload as people who had been assigned equipment. That was a legitimate population to report, but it did not include everyone with an active electronic monitoring order.

The Ministry’s later monitoring-status annex made the structure clearer. Excluding a small specialist group, it identified 27,761 people with the required equipment installed. Within that group, 1,596 were classified as not actively communicating because no signal had been received during the previous 14 days.

That second number sits inside the first. It is not an additional population.

The annex separately identified 5,043 people who had an active order but did not have the minimum required equipment installed. Because the headline caseload counted people assigned with equipment, this no-equipment group sat outside it by definition.

I would treat the 14-day rule with the same restraint. The Ministry described it as a reporting classification, not the point at which operational action began. It also does not explain why a device fell silent: non-communication could reflect non-compliance, equipment damage or fault, or a loss of power.

These categories can be shown beside one another, but they should not be compressed into a single definitive count of “unmonitored people.” One group had no equipment. Another had equipment that had not communicated within the reporting window. The reasons, operational status, and required response could differ.

The headline caseload was not false. It answered a narrower question: how many people had been assigned equipment? The later annex answered additional questions that the headline number could not.

Where the Chain Became Harder to See

The farther the process moved from installation, the harder it became to state the overall outcome with confidence.

Serco could record whether visits happened and equipment communicated. Police, probation, courts, and prisons controlled other parts of the system. Alerts could become breach packs; breach packs could be reviewed; and officials could decide whether further action was appropriate. No single performance measure observed that whole route.

The National Audit Office reported that only around a fifth to a quarter of breach packs had an outcome recorded in the data it examined, while HMPPS could not state a system-wide count of everyone who should have been monitored but was not. I would not treat that absence as proof of the worst-case interpretation. It means the public evidence stopped before the full outcome could be established.

This is an important boundary. The audit described public-protection risk, not a demonstrated increase in harm; HMPPS reported no identified rise in serious further offences linked to service failures. Uncertainty about the monitoring chain should not be turned into evidence of events the record does not document.

The data itself was also still being repaired. Before publishing the new monitoring-status categories, the Ministry carried out data-quality work. Expired orders that still appeared active were closed, including some that still had equipment assigned. The official caseload subsequently fell by about 3%.

That movement was not a clean measure of operational improvement. Some of it reflected records being corrected or closed rather than a corresponding change in the number of people requiring monitoring.

The distinction matters because a revised number can change what is visible without proving that the underlying service changed by the same amount. Data cleansing can improve the accuracy of the picture. It cannot, by itself, tell us that the real-world outcome improved.

What This Case Reveals About Performance Measures

The UK electronic monitoring case is tempting to summarize as a story about misleading KPIs. I think that would be too easy—and less accurate.

The measures were not meaningless. Each one did useful work—inside its own boundary.

The problem began when one stage was treated as proof of the next. Timely attempts did not establish successful fitting. Successful fitting did not establish active communication. Active communication did not establish that every alert produced a recorded and appropriate response.

This is also why the case is not best understood as a contractor story. Serco’s performance mattered, and it improved. But the outcome depended on a system that crossed suppliers, HMPPS, courts, prisons, police, and probation. A measure attached to one actor could not establish the performance of every actor downstream.

In Data-Driven Decision Making Begins Before the Decision, I argued that interpretation begins before a finished number reaches the person expected to act on it. This case shows one of those earlier decisions very clearly: deciding where a measure begins and ends.

Before I use a performance number as evidence, I would ask four questions:

  • Who is included in the number?
  • Which stage of the process does it observe?
  • What outcome does it actually measure?
  • Where does the evidence stop?

Those questions do not weaken performance reporting. They make it harder for a valid measure to be stretched into an unsupported conclusion.

What Can We Responsibly Carry Forward?

The non-obvious finding in this case is not that the electronic monitoring service failed to improve. It did improve. Nor is it that the published figures contradicted one another. They did not.

The stronger conclusion is that operational recovery and incomplete outcome visibility coexisted.

The service became better at completing the activities its contract measured. Later reporting also made previously separate monitoring states more visible. Yet the public record still did not join every stage into one end-to-end account of who should have been monitored, who was actively monitored, and what followed when the system detected a possible breach.

That is a bounded lesson, but a transferable one. In any multi-stage system, a measure can be accurate and still support only a narrow conclusion. Our job is not to distrust the number. It is to identify the part of reality the number can see—and refuse to make it speak for the parts it cannot.

Related Analysis