Case 03 · Global payments rail
Real-time fraud graph
A continuously learning entity graph that scores every hop in the money path — blocking an estimated $1.1B in attack surface year-over-year while holding false positives under 0.08% at 18ms p99.
Context
A global payments rail faced adaptive fraud rings that mutated faster than rule packs could be written. Device farms, synthetic identity, mule networks, and friendly fraud blended into traffic that looked legitimate on any single feature. Marketing promotions — flash sales, referral bursts — created legitimate spikes that old rules treated as attacks, burning conversion and brand trust.
The problem, precisely
Pointwise classifiers missed coordinated rings. Pure graph approaches were too slow for authorization paths. The business needed entity-level reasoning inside a hard latency envelope, with kill switches that risk and compliance could operate without waiting on a model release.
What we built
A streaming graph neural system over devices, merchants, accounts, and intermediaries. Adversarial example harnesses continuously probed the scorer. Policy-as-code encoded jurisdictional and partner constraints. Marketing calendars fed the feature plane so promo-shaped traffic was expected, not punished.
Architecture highlights
- Streaming GNN with incremental neighborhood aggregation.
- Adversarial red-team harness in CI for every model candidate.
- Policy-as-code kill switches operable by risk on-call.
- Promo / campaign features shared with growth analytics.
- Latency budget enforced with graceful feature degradation.
The hard parts
Graph staleness under burst write loads threatened both latency and recall. We partitioned hot entities, prioritized edges by risk prior, and degraded to a fast linear scorer when neighborhoods could not be materialized in time — with explicit telemetry when that happened.
Organizationally, fraud and growth had been adversaries. We created a joint review where every major campaign had a fraud simulation before launch. That single process saved more false positives than any model tweak that quarter.
Rollout
Shadow scores for six weeks, then advisory, then enforce on high-risk corridors, then global. Holiday peak was the exam. The system held p99 under 18ms with degraded-mode rates inside agreed SLOs.
Outcomes
Estimated $1.1B attack surface blocked year-over-year. False positive rate held under 0.08% on the measured portfolio. Growth teams reported fewer “why was my campaign killed” escalations after promo features landed.
What we’d repeat
Put marketing and fraud on the same feature bus. Budget latency like you budget loss. And assume the attacker reads your release notes.