24 Operational Runbooks

Monitoring & Reputation Troubleshooting

Protect service, collect evidence, isolate scope and verify recovery before changing production behavior.

PowerMTA Queue Growing Rapidly

Diagnose whether growth is caused by provider throttling, DNS failures, listener abuse, disk pressure or a sudden traffic change.

Postfix Deferred Queue Increasing

Inspect DSNs, relay responses, transport maps, DNS and resource constraints before forcing retries.

MailWizz Campaign Throughput Dropped

Verify cron execution, delivery-server status, application locks, database load and MTA acceptance.

Microsoft Deferrals Increased

Correlate SMTP response families, source IPs, domains, complaint signals and recent volume changes.

Gmail Temporary Failures Increased

Check authentication, reputation, content shifts, volume distribution and queue age before altering rates.

Yahoo/AOL Delivery Slowed

Separate connection throttling, policy deferrals, DNS issues and campaign-specific failures.

Authentication Pass Rate Dropped

Identify whether SPF, DKIM or DMARC changed, which streams are affected and whether alignment or DNS propagation is responsible.

DKIM Failures After Key Rotation

Check selector publication, private-key deployment, canonicalization, clock timing and stale DNS caches.

SPF Permerror Appeared

Count lookup-causing mechanisms, inspect nested includes, redirects, void lookups and record syntax.

DMARC Reports Show Unknown Sources

Inventory legitimate systems, forwarding paths and abuse before changing policy.

TLS Failures by Provider

Verify certificate chain, hostname, protocol support, cipher compatibility and provider-specific requirements.

Disk Usage Growing Unexpectedly

Locate spool, log, database or temporary-file growth and protect queue data before cleanup.

High CPU on MTA Host

Separate message processing, TLS, DNS, logging, compression and external process load.

High Database Load in MailWizz

Inspect campaign queries, bounce processing, cron overlap, table growth and missing maintenance.

Complaint Rate Increased

Identify source campaign, recipient source, domain, template and timing before pausing only the affected stream.

Hard Bounce Rate Increased

Check list source, domain typos, stale recipients, routing changes and provider-specific rejection patterns.

Soft Bounce Rate Increased

Classify temporary mailbox, policy, reputation, connection and resource failures before retry-policy changes.

Delivery Latency Increased

Measure queue age, connection delays, DNS time, TLS time and provider response latency.

Monitoring Data Missing

Verify exporters, log permissions, time synchronization, service discovery, storage and dashboard queries.

Alert Storm During Provider Incident

Consolidate duplicate alerts, preserve provider scope and keep a single incident timeline.

Metrics Cardinality Exploded

Find high-cardinality labels such as recipient, queue ID or message ID and redesign aggregation.

False Positive Blacklist Alert

Confirm listing source, exact IP, timestamp, delisting status and whether the list is relevant to recipient providers.

PTR or HELO Drift Detected

Validate forward-confirmed reverse DNS, MTA configuration, provisioning records and recent IP changes.

Configuration Change Caused Regression

Compare versions, isolate changed directives, roll back safely and preserve evidence for post-incident review.

Search Trushilla Documentation