Network Monitoring and Performance Questions
Observing and optimizing network health: monitoring and telemetry, latency and throughput optimization, performance baselining, and capacity monitoring. Covers instrumenting networks for visibility, detecting and diagnosing performance degradation, and tuning for latency-sensitive workloads. The reliability and performance view of the network.
You observe intermittent high RTT spikes across regions that correlate with BGP route updates. Describe diagnostic steps (BGP monitoring, looking-glass, Paris traceroute, netflow correlation), detection/mitigation strategies (RPKI validation, route filtering, BGP communities, prepending), and an automated response plan to reduce tail latency caused by transient routing changes.
Explain sampling in flow telemetry: define sampling ratio (e.g., 1:1000), discuss how sampling affects accuracy for different use-cases (total bandwidth estimation vs heavy-hitter detection), and describe mathematical methods or estimators you would use to correct for sampling (inverse-probability weighting, confidence intervals).
As a staff network engineer with limited budget, prioritize observability investments across device monitoring, flow analytics, synthetic testing, and packet capture. Provide measurable KPIs to evaluate ROI (e.g., MTTR reduction, incident detection rate), a 6-month roadmap with milestones, and a stakeholder engagement plan to get buy-in from app developers, security, and execs.
Design a system for real-time topology discovery and alerting based on link-state changes across multiple data centers using routing protocol telemetry and device state. Include data sources (LSDB/RIB, streaming telemetry), graph model for topology, detection of partitioning and flapping links, impact analysis (affected prefixes/services), and how to surface this to operators.
Define tail latency and explain why percentiles (p50, p95, p99, p999) are more useful than averages for latency-sensitive services. Describe measurement challenges (histogram resolution, aggregation across nodes, sliding windows) and how you would report and visualize tail latency for an SLO-driven service.
Unlock Full Question Bank
Get access to all Network Monitoring and Performance interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.