28 Aug 14:20 UTC

Diagnostics

Showing Praia Telecom BR only — readings below are that tenant's sessions, not the whole portfolio

Brazil · GRU

Sony infrar = 0.91
Rebuffer ratioOrigin latency
6.8%5.4%4.0%2.6%1.2%
927ms727ms528ms328ms128ms
15 Aug17 Aug19 Aug21 Aug23 Aug25 Aug27 Aug

Curves moving together = infra fault. Quality drop on flat latency = app-side.

Rebuffering and origin latency rise together on 24 August, peak together, and recover together. The edge is the cause, not the client.

Detected chains

5 chains

CDN POP degradation → regional QoE → regional cancellation

r = 0.83341,200 usersF-2291

Edge health at GRU explains regional quality better than any client-side factor.

  1. GRU cache hit ratio

    CLOUDWATCH
    81.6%from 94.1%
    no lag
  2. Regional rebuffer ratio

    NPAW
    2.44%from 1.31%
    no lag
  3. São Paulo cancellations

    REDSHFT
    4.02%from 3.84%
    +4d lag

CaveatCloudWatch joins at request level, not user level, so the first link is a cohort-level association by design. It can never be attributed to an individual user.

QoE degradation → engagement drop → cancellation

r = 0.87214,800 usersF-2291

When rebuffering rises in a cohort, minutes per session falls within two days and cancellations rise within four.

  1. Rebuffer ratio

    NPAW
    2.81%from 1.34%
    no lag
  2. Minutes per session

    FIREBASE
    18.4from 20.8
    +2d lag
  3. Cancellation rate

    REDSHFT
    4.14%from 3.90%
    +4d lag

CaveatThe lag structure is consistent across four prior edge incidents, which is why the detector assigns high confidence. It remains an association: the confirming check is whether cancellations fall back to baseline within a week of GRU recovering.

App release → crash rate → churn

r = 0.7946,120 usersF-2288

The 3.4.1 rollout raised crashes on Android TV within 36 hours, and D7 retention followed.

  1. 3.4.1 rollout share

    FIREBASE
    42%from 0%
    no lag
  2. Crash-free sessions

    FIREBASE
    98.71%from 99.30%
    +2d lag
  3. D7 retention

    FIREBASE
    38.2%from 60.6%
    +7d lag

CaveatRollout share and crash rate move together by construction. The load-bearing claim is the retention link, and it rests on a 7-day lag with 46k users — enough to act on, not enough to close the case.

Payment failure → suspension → continued play attempts → ticket

r = 0.941,840 usersF-2271

Suspended users keep trying to watch. Every attempt fails silently, and a support ticket follows within a day.

  1. Payment failure rate

    REDSHFT
    8.40%from 2.10%
    no lag
  2. Users entering SUSPENDED

    AWG
    1,840from 410
    no lag
  3. Play attempts while SUSPENDED

    NPAW
    6,204from 180
    no lag
  4. Tickets raised

    ZENDESK
    94from 8
    +1d lag

CaveatThe mechanism here is not in dispute — it is a state machine, not a statistical association. The finding is that nobody was told: 6,204 failed attempts produced no alert until this tool existed.

Entitlement mismatch: AWG says ACTIVE, client behaves as unentitled

r = 0.99412 usersF-2294

412 users are billed and blocked at once. Neither system alone shows a problem.

  1. Users reported ACTIVE

    AWG
    412from 412
    no lag
  2. MFT-401 error volume

    NPAW
    2,884from 38
    no lag
  3. Minutes watched

    NPAW
    0from 198
    no lag

CaveatThis is a contradiction, not a correlation. AWG's view and the client's behaviour cannot both be right, and the tool's only claim is that they disagree.

Automatic cohort decomposition

  • cdn_pop · GRU88.0%
    341,200 users
  • os_version · Android 1484.2%
    612,400 users
  • tenant · revolut-br79.6%
    402,100 users
  • app_version · 3.4.171.4%
    1,284,110 users
  • device_model · Moto G8422.8%
    62,100 users
  • country · BR91.2%
    534,700 users
  • plan · Partner bundle · monthly68.1%
    1,142,800 users
  • title · Costa Norte18.4%
    284,100 users

Cache hit ratio, same windows

CloudWatch · cohort-level only
  • Brazil GRUlow 81.6%now 94.1%