A demo webhook handler can receive a POST request, parse JSON, run business logic, and return 200 OK. A payment, provisioning, deployment, or account workflow needs more. The reliability layer around the handler determines what happens after a network timeout, process crash, duplicate delivery, or downstream failure.

1. Verify signatures against the exact request bytes

Many providers sign the raw request body. If middleware parses the JSON and serializes it again before verification, the bytes may change even when the decoded object looks equivalent. Preserve the original bytes, read the provider headers, verify the signature, and only then decode the payload.

import hashlib
import hmac

def verify_github_signature(raw_body: bytes, header: str, secret: bytes) -> bool:
    if not header.startswith("sha256="):
        return False
    expected = "sha256=" + hmac.new(secret, raw_body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, header)

This focused example illustrates GitHub's documented HMAC-SHA256 header shape. Production code must also validate header parsing, payload limits, secret handling, and the exact protocol documented by the provider. Stripe, for example, uses a timestamped signature scheme and recommends freshness checking. Do not invent one generic verifier for incompatible protocols.

2. Treat duplicate delivery as normal network behavior

The sender cannot always know whether an attempt succeeded. Your service may commit a database change and then lose the response. Retrying is reasonable, so the receiver must assume that a successful event can arrive more than once.

Store a provider event identifier—or another deliberately defined deduplication key—behind a database uniqueness constraint. Seeing the same valid event again should not repeat the business effect. If the event is already safely stored, an acknowledged duplicate will often prevent an unnecessary retry loop; follow the provider's response contract.

3. Keep retries durable, bounded, and observable

Blindly retrying inside the request handler is not durable. The process may exit, the request may time out, or every worker may hammer the same failing dependency. A useful lifecycle records the state, attempt count, next eligible attempt, last error, and terminal failure reason.

Exponential backoff gives a failing dependency time to recover. A dead-letter state preserves events that exhaust their retry budget instead of silently dropping them. Retry policy should distinguish failures that may recover from failures that require operator action.

4. Make replay a guarded operator action

Once the downstream problem is fixed, an operator needs a controlled way to replay a failed event. Replay should be auditable, limited to eligible events, and routed through the same idempotent business path as normal delivery. Manually editing a status column can skip invariants and erase the evidence needed to understand the incident.

5. Avoid the “exactly once” promise

Distributed systems have an uncomfortable failure window: a remote side effect may complete just before the process dies, leaving the sender uncertain about the result. No local flag can remove that uncertainty by itself. The practical design is commonly at-least-once delivery combined with idempotent handling, explicit reconciliation, and honest documentation of the remaining limits.

A practical production checklist

  • Preserve raw request bytes for verification.
  • Implement each provider's documented signature and freshness rules.
  • Support controlled secret rotation.
  • Store receipts durably and define duplicate behavior.
  • Use bounded retries with backoff.
  • Keep terminal failures inspectable.
  • Guard and audit replay operations.
  • Redact sensitive payload fields from logs and operator output.
  • Exercise duplicate, stale, malformed, timeout, crash, and terminal-failure paths in tests.

A tested source-code starting point

BuildShelf's Webhook Reliability & Replay Starter Kit for Python packages these paths into a local, inspectable reference implementation: Stripe and GitHub verification adapters, SQLite receipt state, retries, dead letters, replay, CLI operations, deterministic fixtures, and Docker/Compose workflows.

It is source code, not a hosted service. Buyers still own deployment, TLS, backups, access control, business handlers, provider-specific review, and production operations. It does not guarantee exactly-once remote effects.

Source-code starter kit

Skip rebuilding the webhook failure lifecycle.

Review the included features, verified QA signals, limitations, commercial license, and support scope before purchasing.

View the $199 kit

Disclosure: BuildShelf built and sells the linked source-code kit. The engineering guidance in this article stands independently of the product.