Distributed Systems Security and Trust Questions

Security problems that exist because a system is distributed: keeping trust state correct while it propagates across many services, clusters and regions. Covers credential, token and certificate revocation under eventual consistency and network partitions; fleet-wide rotation of signing keys, secrets and trust anchors without outages (key rollover and grace windows, canary rotation, recovery from a compromised root); distributed authorization (replicated policy decision points, cached decisions, fail-open versus fail-closed when an auth dependency degrades); propagating caller identity and permissions through service call chains; tamper-evident audit trails across services and regions (hash chains, Merkle proofs, ordering events with imperfect clocks); Byzantine and partially trusted participants; cross-cluster and cross-organization trust federation; securing shared distributed components such as caches and message brokers against injection, replay and cross-tenant access; protecting data in transit across region boundaries; and tenant isolation as a security blast-radius boundary. Steady-state mTLS, service-mesh identity and network segmentation mechanics are covered by zero-trust service-to-service security; single-system cryptography and KMS basics by applied cryptography.

HardTechnical
45 practiced

Design a secure service-to-service authentication and authorization system across multiple clusters and regions. Cover token issuance, short-lived credentials, certificate/key rotation, trust boundaries, least-privilege authorization, fallback when an identity provider is down, and operational practices for key management.

HardSystem Design
45 practiced

Your platform needs a secrets and API key management system for services sitting behind an API gateway. It must support zero-downtime key rotation, immediate revocation, audit logging, and minimal blast radius if a key leaks. Explain the storage model, how keys get distributed to gateways and services, how you would orchestrate rotation, and the emergency revocation flow.

EasyTechnical
46 practiced

Why are tamper-evident and append-only audit logs important for a distributed system? As an SRE, describe how you would design audit logging across services and regions so logs can't be silently altered or deleted, how you'd keep events correctly ordered per entity, and what operational controls you would put in place to preserve logs during an incident.

HardTechnical
34 practiced

You operate a multi-tenant control plane where API keys and control APIs are currently shared across all tenants. Design an architecture that minimizes the blast radius if a single tenant is compromised, and explain how you would migrate off the shared-key model and what that isolation costs you versus what it buys you.

HardSystem Design
44 practiced

Design a distributed policy decision point (PDP) for authorization that must serve 1,000 decisions/sec with a 50ms latency SLA across three regions, while preventing privilege escalation during partitions. Describe how to store and replicate policies, your caching strategy at local PDPs, decision versioning, how to roll back a bad policy quickly, and the operational monitoring and audit logging compliance requires.

Unlock Full Question Bank

Get access to all 12 Distributed Systems Security and Trust interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.