Mail flow troubleshooting: when messages do not arrive
Messages disappear, end up in quarantine, loop between systems or take twenty minutes instead of seconds. For cases like these I offer short, clearly bounded troubleshooting engagements: find the cause, fix it, prove it. Methodically, with controlled tests instead of configuration guessing.
- NDRs
- Mail Loops
- Delivery Delays
- Quarantine
- Message Trace
- Header Analysis
- Exchange & M365
Typical symptoms
- NDRs and bounces550 errors, "hop count exceeded", relay rejections: the NDR rarely names the actual cause.
- Delayed deliveryMessages arrive, but minutes or hours late; the delay almost always sits at one identifiable hop.
- MisclassifiedInternal mail is treated as external, legitimate mail lands in junk or quarantine, spoofing gets through.
- After changesSince the gateway swap, the certificate rollover or the migration, the mail flow behaves differently than expected.
Approach
- Pin down the symptomAffected routes, time window, message IDs; "mail is not arriving" becomes a verifiable finding.
- Secure the evidenceMessage trace, message tracking, full headers and NDRs, before anyone touches the configuration.
- Test the hypothesis in a controlled wayTargeted test messages over defined routes, one variable per test.
- Fix and proveThe correction, the proof over the same test route, and the documentation of why it happened.
Real-world examples
- Common causes of mail loops: recognising and resolving loops in Exchange Online, Hybrid and gateways.
- Analyzing Exchange mail flow: the queries with example output, from tracking to the SMTP log.
- Testing SMTP on Linux: layer by layer, from the TCP connection to the delivered message.
- Mail header analyzer: received chain with per-hop timing and authentication results, entirely in the browser.