Problem
Design a system that collects GPU driver crash reports and performance telemetry from millions of end-user machines.
Requirements
Functional:
- Receive crash dumps and perf counters
- Deduplicate and cluster crashes by signature
- Prioritize by impact (users affected)
- Feed fixes back to driver teams
Non-functional:
- Millions of clients, bursty uploads
- Privacy-preserving
- Fast triage of new crash signatures
Discussion points
- Client upload with sampling and back-off
- Crash signature/stack clustering
- Storage and dedup at scale
- Prioritization dashboard
- Privacy and PII scrubbing