Methodology

How the numbers are built, written for an actuary

Every figure on this page traces to a source file in the pipeline that produced it. Nothing here is invented to fill a gap — where the evidence is thin or missing, that is stated directly, not smoothed over. Coverage is 590+ models across 60+ makes today (38,000+ have real curves in the pipeline) and grows as we surface more real inspection data and add sources — each new source folded in with the same measured-vs-modelled discipline below.

TECHNICAL TRACK ↓ACTUARIAL TRACK ↓

Two tracks, one page. The technical track covers the data sources, the pipeline, and its honest limits. The actuarial track covers the frequency × severity pricing math, written for a fellow underwriter — denser on purpose.

Real-world evidence in
MOT inspectionsshape
FCA disclosureslevel
NHTSA recalls / complaintsengine · xm
Workshop pricingseverity
OEM warranty · EV studiesEV · early
Credibility-weighted engine
Frequency × severity, resolved per make / model / year / fuel / mileage.
Thin cells blend toward the market average — nothing overstated.
Cross-market signals kept structurally separate, never merged in.
Honest output
A–E reliability grademeasured
Expected annual costest.
5-year component outlookper part
Confidence band ±always

Every output on the right traces back to a named source on the left. The rest of this page is the detail of each step — read the track you need.

Technical track

Data sources, pipeline mechanics, and the honest limitations.

Sources

  • DVSA anonymised MOT tests and results (2023 + 2024, 109,074,256 raw test records) — data.gov.uk dataset. Licensed under the Open Government Licence v3.0. Provides the failure-frequency SHAPE — cleaned to 84.2M kept test events after class-4 filtering, retest dedup, and mileage/age sanity bounds.
  • FCA general insurance value measures (2022–2024, Extended warranty - motor, Stand-alone) — regulated claims disclosures, cross-referenced against SEC 10-K filings for two OEMs. Provides the absolute claims LEVEL anchor: 26.4% frequency / £844.61 severity, 3-year exposure-weighted.
  • OEM warranty accrual disclosures (Toyota, Honda, Ford, Hyundai, Volkswagen, via Warranty Week’s aggregation of 10-K/annual-report data) — informs the age-0–2 modelled hazard shape where no MOT data exists (first UK MOT test is required at age 3).

Credibility approach

Failure curves are credibility-weighted using a two-level Bühlmann-style blend: a model+fuel+age+mileage cell blends toward its make (weighted by the make’s own exposure), which blends toward the market. Credibility Z = n / (n + K), K = 1,000 supporting tests for 50% credibility — tuned against a reliability sanity check (known-reliable and known-problem models ranking as expected at a fixed age/mileage reference point) during pipeline validation, not chosen arbitrarily.

Every score exposes a match_level (1 = exact cell match through 5 = market-average fallback) and supporting_tests, so you can see exactly how much real data backs any given number, not just the number itself.

MOT-scope caveat

MOT data is a GB roadworthiness test, not a claims or repair record. It captures whether a component failed a specific set of inspection criteria at a point in time — not manufacturer recalls, technical service bulletins (TSBs), or defects that don’t affect roadworthiness (e.g. some infotainment or comfort-feature failures). It also reflects the GB vehicle parc and driving conditions, which may not transfer directly to APAC markets or climates.

MOT-derived numbers should be read as roadworthiness-failure risk specifically, which correlates with but is not identical to total cost of ownership — the cross-market signal below is what fills the sealed-engine/transmission/battery gap this leaves.

Cross-validated across national datasets

The most direct answer to that caveat: we audit the signal against a second, fully independent national inspection regime — the Netherlands’ RDW APK (~13M inspections) — and check whether it agrees with the UK DVSA MOT (~84.2M tests). It does, in two independent ways.

0.73
correlation of model reliability (Pearson r)
0.77
rank agreement (Spearman ρ)
82%
agree on direction vs market average
283
models compared across both regimes
1 · The same physical signature

Share of a fuel type’s inspection defects, computed identically for both regimes. The absolute mix differs by country, but the EV signature is reproduced independently: EVs wear tyres far faster (weight + torque + regen), use their friction brakes far less (regenerative braking), and — having no exhaust — record essentially zero exhaust defects.

defect shareUK petrolUK electricNL petrolNL electric
Tyres10.3%40.5%24.9%43.2%
Brakes13.7%6.2%12.1%6.5%
Exhaust9.1%0.1%4.0%0.0%
2 · The same model rankings

Each model’s age-standardised failure rate relative to its own market (1.0 = average). Across 283 shared models the two regimes correlate at r = 0.73 — the same cars are reliable, and the same cars are troublesome, in both countries.

Reliable in both
Porsche 911UK 0.26 · NL 0.47
Porsche BoxsterUK 0.36 · NL 0.46
Porsche PanameraUK 0.40 · NL 0.70
Porsche MacanUK 0.41 · NL 0.77
Audi A8UK 0.42 · NL 0.71
Porsche CayenneUK 0.49 · NL 0.79
Troublesome in both
Ford Transit CustomUK 1.56 · NL 1.34
Fiat TipoUK 1.56 · NL 1.19
Citroen BerlingoUK 1.54 · NL 1.12
Peugeot PartnerUK 1.51 · NL 1.11
Dacia SanderoUK 1.51 · NL 1.09
Citroen C4 CactusUK 1.50 · NL 1.20

Where they diverge, the reason is legible rather than random: the largest gaps are almost all large commercial vans (Sprinter, Ducato, Crafter) — the duty-cycle confound we already flag for light commercials — plus Tesla Model S, which the Dutch data rates worse, echoing Germany’s TÜV Report. An independent regime reproducing both the physics and the rankings is what lets us treat the signal as regime-independent — and defensible as a basis for APAC, not a single-country artifact.

Reading the signals

Every scored vehicle carries a coverage_source — how it’s covered, not just whether:

  • mot_direct — real UK MOT data for this exact model.
  • platform_proxy — no UK sale, but a documented shared-platform/engine match with a UK-sold sibling covers the whole vehicle (e.g. Toyota Vios scored via the UK-market Yaris, which shares its XP150 platform and 1.5L engine). Confidence band widens.
  • component_proxy — no whole-vehicle match, but a genuine engine/transmission lineage match exists (e.g. Perodua Bezza’s Toyota-Daihatsu group engine family). Only that component is scored; reliability_grade is withheld, not estimated from a partial signal.
  • insufficient_data — no match at any level. The vehicle stays selectable — never hidden — with a sibling_note where one exists (e.g. Proton X50 is a rebadged Geely Coolray, which has no UK presence but real sales data in the Philippines, Russia and South Africa).

The cross-market (“xm”) signal: MOT inspection cannot see inside a sealed engine, transmission, or EV battery pack. For those specifically, this build adds a second, independent signal from NHTSA complaints and recalls — filtered to engine/transmission/HV-battery categories, credibility-weighted the same way the MOT curves are, expressed as a defect_index relative to market average (1.0 = average). It is never merged into the MOT-derived relativity number — shown separately, styled violet in the console and outlook table, tagged signal: "xm" in the API. NHTSA only covers US-sold vehicles, so this signal has its own, different coverage gaps from the MOT data — a model can be well-covered on one signal and absent on the other.

The age/mileage axis, and carry-forward projection: for warranty risk, the real drivers are vehicle age and mileage — model year is mostly a proxy for age. A 6-year-old car at 90,000km has a similar mechanical risk profile whether it’s a 2019 observed in 2025 or a 2017 observed in 2023, provided it’s the same platform generation. So a recent model year isn’t “no data” — it’s “read this model’s real age/mileage curve at its current position,” the same treatment already used for the age-0–2 modelled segment. The risk lives in pricing a forward term: if a model’s current platform generation hasn’t reached a given age yet in the real MOT data, that portion of the curve is projected — holding the model’s own relativity to market constant at its newest real observed age, and letting the market’s own well-populated age-curve shape carry the trajectory forward. Projected cohorts are tagged projection: "carry_forward" (vs. "observed"), dashed the same way the modelled segment is, and widen the confidence band by 5% per year of projection distance (capped at 25%). Carry-forward only activates within a single, continuous platform generation — a real generation or powertrain change (e.g. the 2025 Proton X50 facelift’s switch from a 3- to 4-cylinder engine, or the Vauxhall Corsa E→F switch from a GM to a PSA platform) blocks it, so an older generation’s real data is never silently presented as if it described the current one.

For the very newest model years, this creates a real inversion worth naming directly: MOT data for a car first tested in 2026 won’t exist in volume for years, but a NHTSA complaint or recall lands within months of a defect emerging. For those vehicles, the cross-market signal is genuinely the more current one — the opposite of the usual case, and worth reading that way rather than assuming roadworthiness data is always ahead.

EV battery data: OEM-stated warranty terms (years/km/minimum capacity %) exist for every researched EV, sourced from the manufacturer directly. Real-world degradation curves — degradation_data: "fleet_empirical" — exist only for Tesla (a 2023 peer-reviewed study) and Nissan Leaf (a New Zealand study, 283 vehicles, which found the 24kWh and 30kWh packs degrade at meaningfully different rates and are kept as separate curves, not blended). Every other EV — including BYD and other recent Chinese entrants — shows "warranty_terms_only": the OEM floor is real and shown, but no fleet degradation curve is fabricated or borrowed from another manufacturer’s chemistry. HV-battery defect density reuses the same NHTSA-derived cross-market signal as engine/transmission, kept separate from the degradation estimate — density and degradation rate are different questions.

Modelled vs measured — the newer tiers

To give every vehicle a signal on every dimension, some values are modelled, not measured. These are always tagged distinctly and never merged into real matched data, and they never move the MOT-derived reliability grade:

  • Engine & transmission frequency is estimated — MOT never inspects these internals. A low research baseline that rises with age is scaled by the model’s cross-market NHTSA defect index (real US complaint/recall data where it exists). A known-bad transmission like the Ford Focus PowerShift (~9× market) gets a much higher estimate than a reliable one. Repair cost is estimated Singapore workshop pricing. Rows carry an EST tag.
  • Engine defect index tiers: nhtsa_matched (real US data) > engine_proxy (borrowed from a documented US-sold engine sibling, e.g. Skoda Octavia → VW Golf) > segment_modelled (the segment-average, an explicit placeholder for genuine dead ends).
  • EV battery: fleet_empirical (real per-model study) > chemistry_modelled (a published LFP/NMC/NCA chemistry-class fade curve — real chemistry, class-level curve, not per-model telemetry) > warranty_terms_only.
  • Projection to 2026: researched_generation (a verified platform-generation boundary, 10 models) > default_relativity_hold (projecting past a model’s own newest observed age with no generation research — a wider band, since we can’t confirm the platform hasn’t silently changed). Projected catalog years are marked “· projected”.

Expanded coverage here does not mean “gaps closed” — it means gaps filled with disclosed, lower-confidence estimates. That is a different claim, and the confidence band and tier tag say which one you’re looking at.

Actuarial track

How a premium is built, written for a fellow underwriter.

Frequency × severity, and the FCA anchor

A term’s burn cost = frequency × severity, and the single premium grosses that up to your target loss ratio. Each side is sourced separately:

  • Frequency = the market claim frequency (from FCA-regulated warranty disclosures) × the vehicle’s own MOT-derived relativity × the covered-component share(narrower cover → fewer claims trigger) × an age-trajectory factor that chains the vehicle’s hazard forward year by year over the term.
  • Severity — the FCA-disclosed average claim amount sets the absolute level(converted GBP→SGD at a documented rate); the estimated severity priors only reshape it across the covered components/segments. The better-sourced number sets the scale; the weaker one only the shape.
  • Term pricing — 1/3/5-year burn costs are computed simultaneously by chaining the annual hazards. The headline “expected annual repair cost” is the same engine at a 5-year average (so it doesn’t collapse for a brand-new car whose year 1 is ~zero). Engine/transmission add an estimated burn scaled by the cross-market defect index; an EV’s HV battery adds a severity-only contribution discounted by the OEM-warranty overlap — no fabricated frequency, since no real EV-battery claims history exists.
Worked example · 2018 Toyota Corolla · 1-year term · Singapore
Annual failure frequency (this cell, credibility-blended)0.264
× covered-component share of claims× 0.71
= expected claims in the term0.187
× average severity per claim (estimated, SGD)× S$1,650
= term burn costS$309
÷ target loss ratio (65%) → indicative single premiumS$475

Illustrative, using the same engine the console runs. Every line carries the confidence band; severity is a research estimate, flagged as such.

UK→APAC adjustment

Pinned at 1.0

The absolute claims LEVEL anchor is derived from UK (FCA) disclosures. A UK→APAC adjustment factor exists in the pricing formula but is explicitly pinned at 1.0 — i.e. currently applying no adjustment — pending real judgment input on how UK claims severity translates to APAC repair markets. This is a deliberate placeholder, not a calibrated estimate; treat any output as UK-anchored until this is resolved.

Country-level severity localisation currently uses a placeholder parts/labour index (SG 1.00, MY 0.72, TH 0.66, PH 0.60, ID 0.61) carried over from the design prototype — not a real cost-of-repair index. Component repair costs themselves (below) are a simple research pass, not audited data.

Severity priors — the weakest link, stated plainly

Component repair costs are a simple web-research estimate: 13 component base costs (from noisy, incomplete SG/UK repair-cost search results) × 12 segment multipliers (a documented judgment call, not data). Every cost figure derived from this table — typical repair cost, term burn, single premium — is labelled severity_priors: "estimated" in the API and marked with an est. flag in the console. This is materially weaker evidence than the frequency SHAPE (MOT-derived) or LEVEL (FCA-anchored) sides of this product, and should not be presented to a counterparty as audited.

Confidence band mechanics

Every score carries a confidence band, built additively from named drivers (not a single opaque number):

  • Base ±15%
  • +10% per fallback level beyond an exact cell match (see match_level above)
  • +15% if fewer than 500 real tests support the estimate
  • +20% flat, always, because severity is currently a research estimate (see above)

The confidence_band.drivers field in every API response lists exactly which of these applied.

Full OpenAPI reference: /docs.