Diagnostics
Brazil · GRU
Curves moving together = infra fault. Quality drop on flat latency = app-side.
Rebuffering and origin latency rise together on 24 August, peak together, and recover together. The edge is the cause, not the client.
Detected chains
CDN POP degradation → regional QoE → regional cancellation
Edge health at GRU explains regional quality better than any client-side factor.
GRU cache hit ratio
CLOUDWATCH81.6%from 94.1%no lagRegional rebuffer ratio
NPAW2.44%from 1.31%no lagSão Paulo cancellations
REDSHFT4.02%from 3.84%+4d lag
CaveatCloudWatch joins at request level, not user level, so the first link is a cohort-level association by design. It can never be attributed to an individual user.
QoE degradation → engagement drop → cancellation
When rebuffering rises in a cohort, minutes per session falls within two days and cancellations rise within four.
Rebuffer ratio
NPAW2.81%from 1.34%no lagMinutes per session
FIREBASE18.4from 20.8+2d lagCancellation rate
REDSHFT4.14%from 3.90%+4d lag
CaveatThe lag structure is consistent across four prior edge incidents, which is why the detector assigns high confidence. It remains an association: the confirming check is whether cancellations fall back to baseline within a week of GRU recovering.
App release → crash rate → churn
The 3.4.1 rollout raised crashes on Android TV within 36 hours, and D7 retention followed.
3.4.1 rollout share
FIREBASE42%from 0%no lagCrash-free sessions
FIREBASE98.71%from 99.30%+2d lagD7 retention
FIREBASE38.2%from 60.6%+7d lag
CaveatRollout share and crash rate move together by construction. The load-bearing claim is the retention link, and it rests on a 7-day lag with 46k users — enough to act on, not enough to close the case.
Payment failure → suspension → continued play attempts → ticket
Suspended users keep trying to watch. Every attempt fails silently, and a support ticket follows within a day.
Payment failure rate
REDSHFT8.40%from 2.10%no lagUsers entering SUSPENDED
AWG1,840from 410no lagPlay attempts while SUSPENDED
NPAW6,204from 180no lagTickets raised
ZENDESK94from 8+1d lag
CaveatThe mechanism here is not in dispute — it is a state machine, not a statistical association. The finding is that nobody was told: 6,204 failed attempts produced no alert until this tool existed.
Entitlement mismatch: AWG says ACTIVE, client behaves as unentitled
412 users are billed and blocked at once. Neither system alone shows a problem.
Users reported ACTIVE
AWG412from 412no lagMFT-401 error volume
NPAW2,884from 38no lagMinutes watched
NPAW0from 198no lag
CaveatThis is a contradiction, not a correlation. AWG's view and the client's behaviour cannot both be right, and the tool's only claim is that they disagree.
Automatic cohort decomposition
- cdn_pop · GRU88.0%341,200 users
- os_version · Android 1484.2%612,400 users
- tenant · revolut-br79.6%402,100 users
- app_version · 3.4.171.4%1,284,110 users
- device_model · Moto G8422.8%62,100 users
- country · BR91.2%534,700 users
- plan · Partner bundle · monthly68.1%1,142,800 users
- title · Costa Norte18.4%284,100 users
Cache hit ratio, same windows
- Brazil GRUlow 81.6%now 94.1%