A repeatable way to debug cross-region network failures

When a service works from one network but fails from another, I try not to jump straight to a DNS or CDN change. A small evidence chain makes the failure easier to reproduce and much easier to explain.
1. Capture each layer
Start with the exact hostname, time, client network and resolver. Save the raw output instead of relying on a browser error page.
dig +short A example.com
dig +short AAAA example.com
curl -sv --connect-timeout 10 https://example.com/ -o /dev/null
openssl s_client -connect example.com:443 -servername example.com -brief
These checks separate name resolution, TCP reachability, TLS negotiation and HTTP response handling. A successful ping does not prove that HTTPS is healthy, and an HTTP 200 from one location does not prove that every edge or route is healthy.
2. Compare independent vantage points
Run the same commands from at least two networks or regions. Record the resolved address, connection address, TLS certificate, response code, latency and timestamp. If possible, test both IPv4 and IPv6 explicitly.
A useful pattern is to compare:
the local resolver with a public resolver;

the normal hostname with the origin or a controlled test endpoint;

IPv4 with IPv6;

a direct request with a request through the normal CDN path.

Do not treat a single probe as proof of a global outage. The goal is to find the smallest layer and path that differs.
3. Read the failure pattern
Different answers from resolvers usually point to DNS, delegation, DNSSEC or cache state.

A resolved address with a timeout before TLS points to routing, firewall, congestion or an unavailable edge.

A TLS alert or certificate mismatch points to SNI, certificate deployment or interception.

A clean TLS handshake followed by 4xx/5xx points to HTTP routing, origin health, authentication or application behavior.

Keeping these categories separate prevents a certificate change from being used to “fix” a routing problem, or a DNS change from hiding an unhealthy origin.
4. Keep a small incident timeline
For each observation, record UTC time, vantage point, command, result and the next hypothesis. Include the raw command output and response headers when possible. This makes it possible to compare a recovered path with the failed path instead of relying on memory.
Practical checklist
Confirm the hostname and record type.

Save A/AAAA and resolver results.

Test TCP, TLS and HTTP separately.

Repeat from independent regions and networks.

Compare IPv4/IPv6 and CDN/origin paths.

Change one layer at a time and keep rollback evidence.

Disclosure
I maintain dnspup.com, a public multi-region network diagnostic toolkit. It is included only as a reference for running the same checks; every command above can be run independently.

Sign In or Register to comment.