Building observability into hybrid networks: beyond ping
Monitoring asks if a path is up. Observability asks why the path is behaving this way—especially when satellite and terrestrial disagree.
Monitoring vs. observability
- Monitoring — did a known condition occur? (up/down, threshold crossed)
- Observability — can we explain novel failure modes from the telemetry we already collect?
You need both. Monitoring catches the failures you predicted. Observability shortens the failures you did not.
A minimal hybrid telemetry set
For each site, aim for:
- Reachability of customer-critical targets (not only the router)
- Underlay identity — which path is active for which traffic class
- Capacity / usage in windows your business understands
- Latency and loss on the paths that matter for interactive apps
- Change events — firmware, policy, failovers, ticketed work
If you cannot answer “which underlay carried the pain?”, you are still guessing.
Vanity metrics vs. decision metrics
| Vanity | Decision |
|---|---|
| Raw interface counter screenshots | Error budget burn for a customer journey |
| “Devices online” alone | Online and meeting path SLO |
| Alert count | Actionable incident count |
Practical starter kit
- Synthetic checks from outside the LAN
- Distinct health for primary vs. backup
- Dashboards organized by customer journey, not by vendor logo
- Post-incident: which signal was missing? add only that