The observability bill is a design problem
The telemetry bill for one service passed its compute bill in March. The instinct is to sample harder and log less. We did the opposite: fewer, wider events, tail-based sampling that keeps every error, and a one-week retention tier for the boring 95 percent. The bill went down by two thirds and the incidents got easier to debug, not harder.
11 min readRead more →