Why buses run early, and what punctuality really measures
About a quarter of the buses we track arrive ahead of time. That is not a fleet of speeding drivers — it is what happens when a timetable has to be honest about a journey whose length varies by twenty minutes depending on traffic.
Our bus punctuality league reports a lower figure than the one most operators publish, and it reports a surprising number of early arrivals. Both are real, and both come from decisions about measurement rather than from the buses.
Why early is so common
A bus timetable has to work on a bad day as well as a good one. The scheduled running time between two points is set with enough slack that the route stays viable in traffic, which means that on a clear run the bus simply gets there sooner. The slack is usually concentrated in the final leg, where being early costs the operator least — a bus arriving early at its destination inconveniences nobody, whereas a bus running early at an intermediate stop leaves passengers behind.
Measure arrival at the destination, therefore, and a large share of perfectly well-run journeys land ahead of time. In our sweeps that is around a quarter of them. It is a property of how the timetable is written, not a fault.
Why our figure is lower than the operator's
The published measure counts departures at timing points along the route. We count arrival at the destination. These reward different things.
A journey that leaves every timing point within tolerance and then loses ten minutes on the last stretch scores well on the first measure and badly on ours. A journey that starts late, recovers, and arrives on time scores the reverse. Neither is dishonest; they answer different questions. The regulated measure asks whether the service is operating to plan. Ours asks whether you got where you were going when you were told you would.
We use the second because it is the passenger's question, and we state the difference everywhere we publish a number, because a figure that looks comparable to an official one and is not is worse than no figure.
The mistake that produced 21.9%
The first version of this league reported an on-time rate of 21.9%, which is implausible on its face — and being implausible is what saved it, because it prompted us to look rather than publish.
The error was counting position readings instead of journeys. The live feed reports each vehicle every few seconds. A bus stuck in one jam therefore contributed hundreds of “late” observations, while a bus running cleanly contributed a similar number of readings spread over its whole route. The figure was measuring how long buses spent being late, not how many of them were.
A journey is now counted exactly once, judged at its arrival — which mirrors how we treat trains, and makes the two networks' figures mean the same kind of thing.
“Nearly there” is not there
A subtler version of the same error: it is tempting to count a journey as arrived when the vehicle enters the final leg of its route. That is wrong, and it flatters the numbers, because the last leg is exactly where the padding sits and exactly where urban traffic bites.
We require the vehicle to come within 250 metres of the final calling point before the journey counts as complete. Anything less and a bus that entered the last mile and then sat in a queue for eight minutes would be recorded as punctual.
Working out which journey a bus is on
None of the above is possible without knowing which scheduled trip a moving vehicle is actually running, and the feed does not say. It gives a position, a line and an operator; the timetable gives the trips. Matching them is guesswork that has to be done carefully.
The naive approach — pick the trip on that line which should be running now — fails immediately on a frequent route. On a ten-minute headway, all six buses visible on one line resolve to the same trip, and five of them get somebody else's schedule. We match by position instead, and require the best candidate to beat the runner-up by at least five minutes. Where nothing wins clearly, the journey is not counted at all: an uncounted journey is a small loss, whereas a journey scored against the wrong timetable is a wrong number.
Direction needs care too. Comparing the origin and destination stop codes between feed and timetable does not work, because the two frequently name different stands of the same interchange. We compare positions within a kilometre instead, which tolerates the naming difference while still telling the two directions of a route apart.
What the figures do and do not cover
- Sampled areas only. The sweep covers a set of urban areas rather than the whole country, so an operator's figure here reflects where we watch, not everywhere it runs.
- Only journeys that could be matched confidently. Ambiguous ones are discarded rather than guessed.
- Early counts as not on time, and is reported separately rather than folded in with late. Lumping them together would hide the single most interesting thing in the data.
- It is our measurement, not the regulator's, and should not be quoted as an official statistic.
Where this comes from
Vehicle positions are the Department for Transport's Bus Open Data Service SIRI-VM feed; timetables are the same service's GTFS export; stop and operator reference data are NaPTAN and the national operator codes list. All are described with their licences on where the data comes from. The league itself is at bus punctuality, and the rail equivalent of this guide covers why the train figures have the same shape of caveat.