Disclosure: This article is published by Datamagnet. Vendor claims are self-reported unless otherwise noted. Product capabilities described below are based on public documentation, retrieved 2026-07-22.
How to Ensure Idempotency in Enrichment Pipelines: A 2026 Guide to Avoiding Duplicate Webhook Writes
In 2026, Twilio's own delivery docs confirm its Event Streams product guarantees only "at-least-once" webhook delivery, retrying a failed event "until it succeeds or four hours have passed" (Twilio, Event Delivery and Duplication, retrieved 2026-07-22). That's not a bug. It's how almost every webhook system on the internet works, and it means your enrichment pipeline will eventually receive the same event twice.
If your pipeline treats every webhook as a brand-new fact, duplicate writes aren't a rare edge case - they're a certainty. You don't need a dedicated dedup tool to fix this - the fix is deduplicating incoming webhooks before they hit your database, using logic you build directly into the receiver. This guide walks through that pattern step by step, so a redelivered job-change or engagement webhook updates a record once instead of cloning it.
Key Takeaways
- Webhook providers like Twilio and Shopify guarantee "at-least-once" delivery, not "exactly-once" - duplicates are expected behavior, not a failure state (Twilio, 2026).
- In 2025, Validity found 76% of organizations said less than half their CRM data was accurate and complete, and 37% directly linked lost revenue to poor data quality (Validity, State of CRM Data Management 2025).
- An idempotency key plus a dedup table plus an upsert-based write is the whole fix - no exotic infrastructure required.
- Return the right HTTP status code on redelivery, or the sender's retry logic will keep firing indefinitely.

What Do You Need Before You Begin?
Before touching code, confirm your pipeline has these pieces in place. Skipping any one of them turns idempotency into a half-measure that fails under real retry traffic. The checklist below covers the storage, access, and logic requirements you need before writing a single line of receiver code - each one matters once retries start arriving in production.
- A webhook receiver endpoint that already parses and writes enrichment events (job changes, engagement signals, profile updates)
- A key-value store or database table you can write to with millisecond latency (Redis, DynamoDB, or a Postgres table both work)
- Read access to your webhook provider's payload structure, specifically any event ID or signal ID field
- Comfort writing upsert (
INSERT ... ON CONFLICT) or equivalent conditional-write logic - Time: ~45-60 minutes for the first implementation
- Difficulty: Intermediate
Datamagnet's webhook delivery format
Why Do Webhook Deliveries Create Duplicate Writes?
Webhook deliveries create duplicates because most providers optimize for "the event definitely arrived" over "the event arrived exactly once." A network timeout, a 500 from your server, or a slow response past the provider's timeout window all trigger a resend - and the resend carries the exact same event, not a new one.
Shopify's retry schedule illustrates the scale of this problem well. As of September 2024, Shopify retries a failed webhook up to 8 times over a 4-hour window using exponential backoff (Shopify, Updates to webhook retry mechanism, 2024-09-10). If your endpoint returns a slow 200 (not a failure, just late), the sender may still have already fired a retry before your first response lands - both requests carry an identical payload.
<!-- [UNIQUE INSIGHT] -->Most teams debug duplicate records by chasing the wrong layer. They audit their CRM's own dedup rules, when the actual duplicate was already written before the CRM ever saw it - at the enrichment pipeline's ingestion point. Fixing dedup downstream, in the CRM, cleans up the symptom every time a new webhook fires. Fixing it at ingestion stops the symptom from occurring at all.

Not every provider handles this the same way, and the differences matter when you're building a receiver meant to work across multiple integrations. GitHub doesn't auto-retry failed webhook deliveries at all - a failed delivery just sits until someone manually triggers redelivery from the repository's webhook settings (GitHub, Handling failed webhook deliveries, retrieved 2026-07-22). LINE's Messaging API takes the opposite approach: every webhook event object ships with a webhookEventId field, a ULID that stays constant across redeliveries specifically so bot servers can track which events they've already processed (LINE Developers, Messaging API reference, retrieved 2026-08-08). That's the same key-based pattern this guide builds by hand in Step 2 - some providers just hand it to you pre-built.
Citation capsule: At-least-once delivery is the industry default, not a workaround. Stripe retains idempotency keys for a minimum of 24 hours, and Shopify retries a failed webhook up to 8 times across 4 hours (Stripe, Shopify, 2026). Any enrichment pipeline consuming webhooks needs a dedup window at least as long as its slowest sender's retry period.
Step 1: Which Producers Write to Your Enrichment Pipeline?
By the end of this step, you'll have a complete list of every event source that can trigger a write - the foundation everything else builds on. Skipping this step is how teams end up idempotent against one integration and wide open on another.
List each webhook source separately: job-change signals, engagement signals, CRM sync callbacks, and any internal retry jobs your own code schedules. For every source, note the field the provider uses as a stable event identifier - this is what you'll key off of in Step 2. Datamagnet's signal creation endpoint assigns each monitored event a signal ID that stays constant across redeliveries, which is the field to grab.
Verify by pulling your last 100 received webhook payloads from logs and checking whether that identifier field is present and non-null on every single one. If it's missing on any payload, you don't have a reliable key yet - fix that before moving on.
Step 2: How Do You Generate a Stable Idempotency Key?
By the end of this step, every incoming webhook produces the same key on every delivery attempt, no matter how many times it's resent. The key is what separates "this is a duplicate" from "this is new data." Get this step wrong and every downstream dedup check inherits the mistake, since the whole pipeline trusts this key to mean the same thing every time it's recomputed.
The safest key combines the provider's event ID with the event type: {provider}:{event_type}:{event_id}. Don't rely on payload content hashing alone - two genuinely different events can hash identically if the provider truncates fields, and a single event can produce two different hashes if field ordering shifts between deliveries.
A common failure we've seen in the wild: teams key off of received_at timestamps because it's the easiest field to grab. That works exactly until the first redelivery arrives with a new timestamp and a completely unchanged payload - at which point the "duplicate" sails through as a fresh event. Never derive an idempotency key from anything that changes between delivery attempts.
Verification: Send the same test payload to your receiver twice and confirm the generated key is byte-for-byte identical both times.
Step 3: How Do You Store Processed Keys With a TTL?
By the end of this step, your pipeline can check "have I seen this exact event before?" in a single fast lookup, before any write logic runs. This check has to happen first - everything downstream depends on it. Get the lookup wrong or too slow, and every safeguard you build on top of it inherits the same gap.
- Create a table or key-value namespace:
dedup_keys(key, processed_at, expires_at) - On every incoming event, check for the key before running any enrichment or write logic
- If found and unexpired, return success immediately without re-processing
- If not found, process the event, then insert the key with a TTL
Set the TTL to at least as long as your longest-retrying sender's window - based on the chart above, 24-48 hours covers Stripe, Twilio, and Shopify with margin. Redis's built-in EXPIRE command handles this natively; in Postgres, a scheduled cleanup job or a expires_at index with a nightly purge works just as well.

Step 4: How Do You Make the Write Itself Idempotent?
By the end of this step, even if a duplicate somehow slips past your dedup table check, the write itself can't create a second record. This is your second line of defense, not a replacement for Step 3. It's the layer that catches race conditions and bugs in Step 3's logic before they ever touch your data.
Replace every INSERT in your enrichment write path with an upsert keyed on a natural unique constraint - typically the enriched entity's own stable ID (a LinkedIn profile URL, a company domain, or your internal contact ID), not the webhook's event ID. In Postgres: INSERT INTO contacts (...) VALUES (...) ON CONFLICT (linkedin_url) DO UPDATE SET .... In a document store, use a conditional upsert with the entity ID as the document key.
This layered approach - dedup table first, upsert second - means a race condition between two near-simultaneous redeliveries still can't produce two rows, because the database's own unique constraint rejects the second insert attempt at the storage layer.
Step 5: What HTTP Status Should You Return So Senders Stop Retrying?
By the end of this step, your receiver's response codes actively help the sender stop retrying instead of accidentally triggering more retries. This step gets skipped constantly, and it's a major source of avoidable retry volume. Getting this one detail right removes a whole class of duplicate deliveries before your dedup table even has to catch them.
Return a 2xx status the moment you've either processed the event or confirmed it's a duplicate - do this before any slow downstream work, like triggering a CRM sync or a notification. If your receiver takes 10 seconds to fully process before responding, you've built the exact condition that causes providers to assume failure and retry. Queue the slow work asynchronously and respond fast.
This is the same rule to apply if you're chaining an edge function so a webhook triggers enrichment and then writes back to Postgres: acknowledge the webhook in milliseconds, then let the enrichment-and-write chain run asynchronously behind that fast response. Trying to fit the full round trip, dedup check, enrichment call, and database write inside the sender's response window is what causes the timeouts and retries this whole guide exists to prevent.
Datamagnet's API error code reference
Only return a 4xx or 5xx when the event genuinely failed and a retry would help - a transient database timeout, for example. Never return an error status just because a duplicate arrived; that tells the sender to try again, which is the opposite of what you want.
Step 6: How Do You Test With Replay and Retry-Storm Simulations?
By the end of this step, you'll have proof the pipeline behaves correctly under the conditions that actually break naive implementations - not just the happy path. These three tests catch the failure modes that only show up under real retry pressure, which is exactly when duplicate writes are most expensive to clean up.
- Replay test: Capture a real webhook payload, send it to your receiver 5 times in a row, and confirm exactly one record exists afterward
- Concurrent test: Fire the same payload from 10 parallel requests simultaneously and confirm the dedup table plus upsert combination still produces one record, not a race-condition duplicate
- Retry-storm test: Simulate a burst of near-identical events, since a single misbehaving dependency can generate tens of thousands of retry connections in a short window - Microsoft's own architecture diagnostics recorded a single client IP generating 81,608 inbound connections within one hour during a documented retry-storm incident (Microsoft, Retry Storm Antipattern, Azure Architecture Center, updated 2026-03-17)
That Microsoft figure isn't a rare worst case - it's what happens when retry logic on both ends compounds without backoff. A pipeline that survives the replay test but not the retry-storm test isn't done yet.
What Common Mistakes Should You Avoid?
Most duplicate-write bugs trace back to one of five patterns, and every one of them is avoidable once you know to check for it. Teams that hit duplicate records in production almost always trace the root cause back to one of these five, not to something exotic in their infrastructure.
1. Keying off the wrong field. Using received_at, a row's auto-increment ID, or any value that changes between delivery attempts guarantees the dedup check will never catch a real duplicate. Always key off the provider's stable event or signal ID.
2. Checking for duplicates after the write, not before. If your dedup check runs after the enrichment write already happened, you've already created the duplicate you're trying to detect. The check has to gate the write, not follow it.
3. TTLs shorter than the sender's retry window. A 1-hour dedup TTL against a sender with a 4-hour retry window leaves a 3-hour gap where a late retry sails through as "new."
4. Treating idempotency as a one-time build instead of an ongoing property. New webhook sources get added to a pipeline constantly, and each one needs the same key-generation and dedup logic applied - it doesn't inherit automatically from the sources you already fixed.
<!-- [ORIGINAL DATA] -->5. Trusting the CRM's own dedup rules to catch what the pipeline missed. Teams that rely on CRM-side matching rules as their only line of defense typically discover the gap only after a data-quality audit - by then, months of duplicate signal-driven updates have already inflated activity counts and skewed lead-scoring models built on that data.
The cost of skipping this compounds. In 2025, 76% of organizations told Validity that less than half their CRM data was accurate and complete, and 37% directly tied lost revenue to that poor data quality (Validity, State of CRM Data Management 2025).
What Does Success Look Like?
If everything above is wired correctly, you should now see one record per real-world event, no matter how many times a sender retries the underlying webhook. Replaying the same payload 10 times should leave your dedup table with exactly one key and your data store with exactly zero new duplicate rows.
Key metrics that confirm it's working: duplicate-record rate in weekly data-quality audits trending toward zero, dedup table hit rate matching your provider's known retry frequency (not near-zero, which usually means the check isn't actually running), and no growth in support tickets about "why do I have two of the same contact."

Once this is stable, the natural next step is applying the same dedup-table-plus-upsert pattern to any additional signal types you register, so new integrations inherit the protection instead of reintroducing the gap. Datamagnet's signal update endpoint lets you point an existing monitor at a new webhook URL without recreating the whole signal, which is useful when you're migrating a pipeline onto this pattern gradually.
Frequently Asked Questions
These are the questions engineering teams ask most often once they start building idempotency into a webhook receiver, covering dedicated infrastructure, missing event IDs, key retention windows, and retry-storm resilience in more depth than the step-by-step outline above covers on its own.
What's the difference between idempotency and deduplication?
Idempotency means an operation produces the same result no matter how many times it runs. Deduplication is one technique for achieving it - detecting and discarding repeat events before they cause a repeat effect. In an enrichment pipeline, you need both: a dedup check to catch repeat events, and an idempotent write (an upsert) as a second line of defense.
Do I need a separate database just for idempotency keys?
No. A Redis namespace or a single Postgres table with a key and expires_at column is enough for most pipelines. You don't need a dedicated webhook infrastructure service that tracks event IDs for you - reserve a managed option only if you're processing enough volume that a shared table becomes a lookup bottleneck, since most B2B enrichment pipelines never reach that scale.
What if my webhook provider doesn't include a stable event ID?
Some do not. In that case, construct a composite key from fields that are guaranteed not to change between retries - typically the entity ID plus event type plus a truncated timestamp bucket (e.g., rounded to the nearest minute). It's less precise than a native event ID, but it still catches the vast majority of true duplicates.
How long should I keep idempotency keys before expiring them?
At least as long as your sender's documented retry window, with margin. Stripe retains idempotency keys for a minimum of 24 hours and Shopify retries for up to 4 hours across 8 attempts (Stripe, Shopify, 2026) - a 48-hour TTL covers both providers with room to spare.
Can this pattern handle a full retry storm, not just occasional duplicates?
Yes, if the dedup check happens before any expensive work and your 2xx response is fast. Retry storms escalate specifically when a slow or failing receiver causes senders to pile on more retries; a receiver that responds in milliseconds - even to a duplicate - starves the storm of the failure signal that feeds it (Microsoft, Retry Storm Antipattern, 2026).
Stop Letting Retries Write Your Data Twice
You've now got the full pattern: map every event source, generate a stable idempotency key, gate writes behind a dedup table with a TTL that outlasts your senders' retry windows, and back it with an upsert so a race condition can't sneak a duplicate past the check. Return fast 2xx responses so senders stop retrying for the wrong reasons, and test against replay and retry-storm conditions, not just the happy path.
Duplicate writes aren't a webhook provider's fault - they're the expected behavior of an at-least-once delivery guarantee that your pipeline has to design around. Review how Datamagnet's signals deliver webhook events and the HMAC-signed webhook practices in our security documentation before wiring a new signal type into your pipeline. For a look at how this pattern plays out with a specific signal type, see how real-time job-change intent signals get delivered and how programmatic CRM enrichment depends on exactly this kind of write-safety underneath it. Explore Datamagnet's Signal API and test a webhook against your dedup logic this week.
Sources
- Twilio, Event Delivery and Duplication, retrieved 2026-07-22, https://www.twilio.com/docs/events/event-delivery-and-duplication
- Stripe, Idempotent Requests, retrieved 2026-07-22, https://docs.stripe.com/idempotency
- Shopify, Updates to webhook retry mechanism, 2024-09-10, retrieved 2026-07-22, https://shopify.dev/changelog/updates-to-webhook-retry-mechanism
- GitHub, Handling failed webhook deliveries, retrieved 2026-07-22, https://docs.github.com/en/webhooks/using-webhooks/handling-failed-webhook-deliveries
- LINE Developers, Messaging API reference, retrieved 2026-08-08, https://developers.line.biz/en/reference/messaging-api/
- AWS, Idempotency - AWS Lambda Developer Guide, retrieved 2026-07-22, https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-idempotency.html
- Microsoft, Retry Storm Antipattern, Azure Architecture Center, updated 2026-03-17, retrieved 2026-07-22, https://learn.microsoft.com/en-us/azure/architecture/antipatterns/retry-storm/
- Validity, The State of CRM Data Management in 2025, retrieved 2026-07-22, https://www.validity.com/resource-center/the-state-of-crm-data-management-in-2025/
- Datamagnet, Webhooks, retrieved 2026-07-22, https://docs.datamagnet.co/api-reference/webhooks
- Datamagnet, API Errors reference, retrieved 2026-07-22, https://docs.datamagnet.co/api-reference/errors

