Methodology / Glossary

Glossary

Every term and metric used across the site, with plain-language definitions that explain not just what a term means, but why the convention exists. Each entry has an anchor so score cards and verdict lines can deep-link directly to it.

Band

One of five colour-coded categories that group scores: STRANDED (0–29), POOR (30–49), PATCHY (50–69), DECENT (70–84), or GOOD (85–100). Bands make it possible to talk about "a STRANDED suburb" without quoting a precise score that will shift with each data refresh.

Why this convention: A score of 34 and a score of 41 describe materially the same experience (long waits, limited coverage), but the numbers invite false precision. Bands group scores into meaningfully different tiers so the story is about the category, not the decimal place.

Viable route

A route whose best-scoring stop in the area has a final_score above 50 (the PATCHY/DECENT boundary). Below that threshold, the service is too infrequent, limited, or unreliable to count as a genuine travel option.

Why this convention: Without a viability floor, a bus that comes once every 90 minutes would count as "an option" for diversity bonus purposes, inflating the score of an address whose real choices are drive or wait. The 50-point line marks where a service becomes worth planning around.

Reach

The share of a suburb's population living within 800m of a viable route. When reach drops below 40%, the verdict prioritises this fact over any mode-split or frequency story — describing what the served minority experiences would be misleading for a suburb where most people cannot walk to anything useful.

Why this convention: A suburb might have excellent buses along one corridor and nothing elsewhere. Stop-level data would describe the excellent buses and miss the story. Reach catches the gap: "most of this suburb cannot walk to a stop worth using" is a fundamentally different verdict from "the buses that exist are infrequent".

Best route

The highest-scoring viable route within walking distance of a point or cell centre. The address-level formula weights this best option at 70% of the final score, making it deliberately dominant — an address score should answer "what is the best service I can actually walk to", not "what is the average of everything nearby". The best route figure on a suburb page is deliberately not population-weighted, since "your best option" is a different question from "typical access here".

Why this convention: Averaging all nearby stops would let a mediocre tram stop dilute the effect of an excellent train station 200m away. The best-route approach rewards genuine anchor stations and corridors, which is the correct signal for an address-level reader asking "how good is my best option".

Headway

The time between consecutive services on the same route in the same direction (e.g. a bus every 15 minutes). What riders actually experience is the headway, but the score uses average wait — half the headway — as its input, because the expected wait for a randomly arriving rider is half the gap. See also: average wait.

Why this convention: Headway is what most people mean when they ask "how often does the bus come". Scoring on average wait rather than headway is a mathematical convenience (each represents the same underlying data, scaled by 0.5), and the verdict labels convert back to headway for readability.

Average wait (half-gap convention)

Half the headway, used as the input to the headway score curve. If services run every 30 minutes, the average wait for a randomly arriving passenger is 15 minutes. The score curve treats 15 minutes as its input, not 30.

Why this convention: The half-gap convention is standard in transit planning and avoids double-counting the gap between two services. It does not make the score more or less generous — the curve constants were calibrated against the half-gap input, not the full headway.

Representative day

A date chosen to represent typical service for scoring. Rather than picking one date from the feed (which could land inside a school-holiday or service-change window), the processor samples five near-term Wednesdays and five near-term Saturdays, taking the per-stop median trip count across the sample. Departures and headway are populated from whichever sampled date's count is closest to that median.

Why this convention: A single date can be anomalous (a public holiday, a planned shutdown, a future timetable change that zeroes out a route). Sampling multiple dates and taking the median ensures a stop only drops out if it is genuinely inactive on most dates, not because of one exception window.

Grid cell

A 250m x 250m square used to score suburb-level access. Each cell's centre is scored like a single address (best viable route plus diversity bonus), and the suburb's headline score is the population-weighted average of its cells.

Why this convention: Averaging stops directly gives stop density — a corridor with a tram stop every 200m — too much influence over the suburb score. Gridding ensures the score reflects what the typical resident can actually walk to, not how many stop signs exist in the suburb.

Dasymetric weighting

A population-modeling technique that distributes a known total (ABS Estimated Resident Population for an SA2) across smaller areas (mesh blocks) proportional to a correlated variable (G-NAF residential address counts). Unlike a uniform-density assumption, dasymetric weighting puts population only where addresses exist — commercial districts, parks, and industrial zones carry zero weight.

Why this convention: Uniform density would give a park the same population weight as a housing estate, silently dragging the suburb score toward the unpopulated areas. Dasymetric weighting is more honest about where people actually live, and more accurate than the 2021 Census snapshot for growth corridors.

Mesh block

The smallest geographic unit published by the ABS, typically 30–60 dwellings. Mesh blocks are the building block for all higher-level ABS geographies (SA1, SA2, etc.) and the finest resolution at which dwelling counts are publicly available.

Why this convention: SA2s are suburb-scale — using them directly would lose within-suburb gradients (one side of a suburb has a train station, the other does not). Mesh blocks let us distribute population at a scale where we can meaningfully say "this cell is residential, that cell is parkland".

SA2 (Statistical Area 2)

A medium-sized statistical geography published by the ABS, typically representing a suburb or group of related suburbs. SA2s are the level at which ABS Estimated Resident Population (ERP) totals are published annually.

Why this convention: ERP is the most current official population estimate, but it only exists at SA2 level. To get finer resolution, we distribute the SA2 total across its mesh blocks using G-NAF address counts — the dasymetric model described above.

ERP (Estimated Resident Population)

The ABS's official annual population estimate for each SA2, adjusted from Census counts for births, deaths, and migration. More current than the 5-yearly Census but only available at SA2 level.

Why this convention: Using ERP replaces the 2021 Census 5-year lag with an annual refresh. The trade-off is SA2-level granularity, which we then distribute dasymetrically to mesh blocks via G-NAF.

G-NAF (Geocoded National Address File)

Australia's official address register, maintained by Geoscape and published by the ABS. Every titled residential property in Victoria has a G-NAF record with coordinates and a mesh-block assignment. Monthly Core releases are available via data.gov.au.

Why this convention: G-NAF provides up-to-date address locations and their mesh-block assignments, enabling the dasymetric population model. Its quarterly refresh cadence captures new estates faster than the 5-yearly Census.

Legacy-mean fallback

A simpler suburb-scoring method used when a suburb has no real polygon boundary or zero populated grid cells: the plain arithmetic mean of every scored stop's final_score. Unlike the grid method, legacy-mean can let a single high-scoring station or dense stop spacing inflate the suburb's published number.

Why this convention: Every suburb should have a published score even when it falls outside the metro polygon scope. The legacy-mean gives a number — honestly labelled — rather than leaving the page blank. The methodology page and suburb detail page both state which method produced the number.

Trunk anchor

A stop with a final_score of 70 or above (the DECENT threshold), used as the target for feeder-bus network-plan proposals. Anchors represent "transit that already works" — a train station or high-frequency bus corridor that a proposed feeder route should connect residents to.

Why this convention: The network plan is about connecting people to existing quality transit, not rebuilding the whole system. The 70-point threshold is the same "decent" boundary used everywhere on the site, not a separate bar invented for this feature.

Canonical route

A merged representation of a single public route number, combining all its GTFS route_ids (direction variants, branches, night-network versions) into one entity for scoring and page display. Train lines key on route_long_name; other modes key on (normalised route_short_name, mode).

Why this convention: The "901" should have one page and one score, regardless of how many internal route_ids the GTFS feed uses for its variants. Without canonicalisation, a route with 8 direction variants would dominate the route league table with 8 separate entries.

School_special

A flag for routes whose primary purpose is school transport, identified by name patterns, low trip counts (under 6 weekday trips), or few active days (under 4). School_special routes get a page but are excluded from league tables so rankings reflect genuine public transport, not the school run.

Why this convention: A route that operates 2 return trips at 8am and 3pm weekdays only is a school service, not general public transport. Including it in the worst-20 bus list would hand critics an easy rebuttal: "you're calling a school bus the worst route in Melbourne".

Circuity ratio

A measure of route directness: the length of the route's representative shape divided by the straight-line distance between its termini. A ratio of 1.0 means the route is perfectly direct; 2.0 means it travels twice as far as the crow flies.

Why this convention: A bus that travels 15km to cover 5km of straight-line distance is not competitive with a car regardless of frequency. The circuity ratio captures this inefficiency as a continuous score, penalising routes above 1.2 and scoring near-zero above 2.5.

Intermodal bonus

Up to 20 additional points added to a stop's coverage key when different transit modes interconnect within 150m — a train station with a bus interchange scores higher than an isolated train station.

Why this convention: Physical interchanges multiply the usefulness of a stop: a rider who can alight a train and board a bus has access to a much larger network than a rider whose stop serves only one mode. The bonus is discounted when the connecting service itself scores under 70, so a poorly-served interchange does not inflate the score.

Free-flow

Driving time under ideal conditions — no traffic, no delays, green lights all the way. Free-flow times represent the best possible driving experience, which is the best case for the car and the worst case for our argument that public transport should be prioritised.

Why this convention: Until peak-congestion data is available (blocked on a working DTP traffic-data API key), free-flow is the honest default. Publishing free-flow times with a stated limitation is more defensible than inventing a congestion multiplier.

Gravity decay / BETA

A weighting function used in the car-competitiveness formula: closer destinations get higher weight, and the weight decays exponentially with distance. The decay rate is controlled by BETA, currently set to 0.05 per minute — a commonly-cited moderate work-trip decay rate from the accessibility literature.

Why this convention: Without gravity decay, a suburb's car-competitiveness ratio would be equally influenced by a destination 5 minutes away and one 90 minutes away. Gravity decay ensures the ratio reflects the destinations a real resident would actually travel to. BETA is not yet calibrated against real Melbourne journey-to-work data — flagged as a placeholder.

Cumulative accessibility

A measure of how many jobs a resident can reach within a fixed time threshold by public transport. Currently set at 45 minutes, computed from the same OSM + GTFS travel-time matrix used for car competitiveness.

Why this convention: Cumulative accessibility answers a different question from the car-competitiveness ratio: not "how much longer than driving does it take", but "how many opportunities does the network actually connect you to within a reasonable commute".

Snapshot / vintage

A dated record of which GTFS build and data versions produced a given set of scores. Every score card, share image, and page footer carries the data vintage and methodology version, making any screenshot fully citable on its own.

Why this convention: PTV publishes new GTFS feeds regularly; scores change between builds. Without a snapshot label, a screenshot from March and one from July would be indistinguishable but possibly different. The vintage line turns every score into a citable fact about a specific point in time.

Methodology version

A semver identifier (currently 1.0.0) that changes when the scoring formulas, constants, or aggregation methods change. Major = band or list membership change; minor = new measure or coverage expansion; patch = fix restoring intended behaviour.

Why this convention: The previous date-based scheme ("2026.07") could not distinguish a data refresh from a formula change. Semver tells a reader immediately whether a score moved because the network changed (same methodology version, new data) or because the ruler changed (new methodology version). Every changelog entry carries its methodology version so the timeline is fully citable.