Why a separate guide on best practices?
We answer the basic question "what is an API integration" in our guide to what an API integration is. This page is for the next stage: you now know you're going to build an integration, or already manage one, and want to know which patterns make the difference between an integration that works in the demo and one that holds up under production pressure, with vendor outages, peak loads and a GDPR audit due next quarter.
Many teams start an integration with the mindset "it's just an HTTP call, we'll get it working". That holds for the prototype. It holds less well once the first retry storm hits a vendor's rate limit, or when it turns out the same payment confirmation was processed twice because a webhook was repeated. The best practices below are the defensive layer that separates production quality from demo quality. If you first want to know what an API integration costs and which architectural choices affect that, start there; for the trade-off between custom and platform, see API versus integration platform.
The sections run from "outside in": first patterns at the network edge (authentication, retries, webhooks), then those concerning your own architecture (observability, gateways, BFF), and finally the governance layer (versioning, secrets, GDPR).
Authentication: OAuth2, JWT and mTLS
Authentication is the part of an integration where a lot of costly time disappears. Not because it is conceptually complicated, but because the assumptions of two systems about how a caller identifies itself rarely align well. Three patterns you will encounter in practice, often in combination.
OAuth2 for customer and SaaS flows
OAuth2 is the de facto standard for scenarios in which a user grants a third party permission to access data on their behalf. The client_credentials grant is for server-to-server communication, authorization_code for user flows involving a redirect. Establish in advance which scopes you request, how long refresh tokens live, and what your plan is if a token expires during an overnight batch. Many teams only discover in production that a refresh token becomes invalid after a long period of inactivity, and by then they have no automated recovery path.
JWT for service-to-service
JSON Web Tokens with asymmetric signing suit scenarios where you manage the identity domain yourself: one microservice calls another with a token signed with a private key, and the recipient validates it with the public key. Short expiry, limited scopes per token, key rotation as a fixed routine, and an unambiguous audclaim so that tokens cannot be reused against another endpoint.
mTLS for B2B and enterprise integrations
Integrations with banks, insurers, healthcare providers or government bodies often use mutual TLS: both sides of the connection present a certificate. It raises the bar for an attacker considerably, but introduces new operational complexity, as certificates expire, and always at the worst possible moment. Build certificate monitoring in as a first-class concern, and treat renewal as a process that goes through the test environment first, not as manual work in production.
What you explicitly want to avoid
API keys in URL parameters (logged by anyone with access to access logs), basic auth without TLS, shared service accounts that nobody can any longer link to a user, and custom authentication protocols. For something as fundamental as authentication, it rarely pays to invent your own recipe; use proven patterns, and invest your time instead in observability around authentication failures.
Idempotency keys: the pattern that saves the most grief
An operation is called idempotent if performing it twice produces the same result as performing it once. For read operations (GET) this holds naturally. For write operations (POST, PUT, PATCH) you must design for it deliberately. The mechanism: with each write operation the caller generates a unique Idempotency-Key (a UUID will do) and sends it along as a header. The server temporarily stores this key together with the result of the first successful call. If the same key arrives again within a window, for example because a network timeout triggered a retry while the original request did succeed, the server returns the stored result without executing the operation again.
POST /v1/payments
Idempotency-Key: 9f4c2b0e-13aa-4f3a-8e6b-a5e0c3b1d8c2
{ "amount": 2500, "currency": "EUR", "order_id": "ord_8731" }
Key design decisions: how long do you keep keys, what do you do when the same key arrives with a different body (a 409 is cleaner than silently ignoring it), and how do you document in your API docs which endpoints accept this header? Are you building an API that others will consume? Don't make idempotency optional — make it the default, and validate that the header is present on all non-GET actions.
Retry strategy: exponential backoff with jitter
Networks fail, vendors are busy, deploys run long. A well-built integration retries — but not like a parrot. Three principles that work together:
- Exponential backoff: after the first failure, wait for example 1 second, then 2, then 4, then 8. You give the other system breathing room instead of overwhelming it.
- Jitter on top: without randomised variation in the wait time, you get a "thundering herd" — ten clients retrying at exactly the same moment after a vendor outage, causing a second, self-inflicted outage.
- A retry budget: define a maximum number of attempts or a time budget, and then park the request in a dead-letter queue. A retry loop that runs for hours without anyone noticing is exactly what you want to avoid.
def retry_with_backoff(call, max_attempts=6):
delay = 1.0
for attempt in range(max_attempts):
try: return call()
except RetryableError:
if attempt == max_attempts - 1: raise
time.sleep(delay * (0.5 + random.random()))
delay *= 2
Not every error is retryable. Retrying a 401 or 400 solves nothing, as those are permanent errors. A 429, a 502/503/504 or a network timeout are classic retryable cases. Encode this distinction explicitly. For the broader approach to robust integrations, see our service smart API integrations.
Circuit breakers: stop knocking when the door is shut
When a dependency fails completely, carrying on with retries can be worse than stopping. A circuit breaker monitors an endpoint's success rate over a rolling window. If it drops below a threshold, the breaker "opens": all subsequent calls fail immediately, without putting any further load on the vendor. After a cooling-off period, the breaker tries a half-open state with a few calibration calls; if those succeed, it closes again.
The benefit is twofold: your own system stays responsive, and you help the vendor recover by no longer adding pressure. Libraries such as Resilience4j (Java), Polly (.NET) or a custom wrapper around the HTTP client make the pattern easy to implement. Be careful with thresholds that are too low in low-volume situations: if an endpoint receives only a few calls, a single failure produces a high failure rate quickly and the breaker opens too readily. Tune the window and threshold to your traffic.
Respecting rate limits — and spreading the load proactively
Almost every modern API has rate limits. Handling them well comes down to two things: responding to the signals the vendor gives, and proactively spreading the work out.
- Honour limit headers: with a 429, or sometimes even with a 200, vendors send
X-RateLimit-Remaining,X-RateLimit-ResetandRetry-After. Build your retry policy so that it first followsRetry-Afterand only falls back to exponential backoff when that header is absent. - Spread batches with a token bucket: for large import jobs, a naïve loop is a guarantee of 429 storms. A token bucket in your API client limits the outbound rate to what the vendor tolerates. Batches that look "too slow" on a single thread often finish sooner in the end than a burst that is continually being cut off.
- Negotiate where needed: many SaaS vendors will raise rate limits on reasonable request for enterprise contracts. For an initial historical sync, it's better to ask for an increase in advance than to let your production integration hit the wall.
Webhooks versus polling: choose deliberately per stream
To retrieve changes you have roughly two options: polling or receiving a webhook. For rare events that need to be processed quickly (payments, fraud alerts, disputes), webhooks are better — polling wastes capacity. Polling wins when you want replay control, want to set the window yourself, or when the vendor does not offer reliable webhook delivery. An hourly incremental pull on a updated_atWebhook-polling is often more robust than a webhook you can never be sure has arrived.
What a production-grade webhook handler does
- Verify signatures: every serious vendor signs webhooks with HMAC. Verify before you even parse the body, or you leave yourself open to spoofing.
- Respond quickly, process asynchronously: reply with a 2xx within seconds, put the payload on a queue, and process it in a separate worker. A slow handler triggers retries on the vendor's side.
- Process idempotently: use the vendor's event ID as the idempotency key, so that the same event has only one effect.
- Keep unprocessable events: if processing fails, move the event to a dead-letter queue with enough context for manual or automated recovery.
Most robust integrations combine both: webhooks for near-real-time signals, and periodic polling reconciliation to detect missed events. Anyone who relies only on webhooks will sooner or later discover that they have missed something.
Eventual consistency: embrace it, or pay the price
Two systems synchronising via APIs are never truly equal in real time. There is always a delay: milliseconds in the best case, minutes or longer if batch processing is involved. Denying that delay leads to bugs that manifest exclusively under high load.
Two patterns help. First: avoid questions like "can you retrieve this record immediately after creating it?". Replace them with event-driven flows in which the receiving system itself waits for confirmation that the message has been processed. Second: handle state conflicts explicitly. Last-write-wins is an acceptable strategy if you choose it deliberately; problematic if it creeps in unnoticed. For records with multiple sources, it is better to work with version vectors, CRDTs, or a leading-system-per-field agreement.
What definitely helps: consistently store timestamps in UTC, do not rely on wall-clock ordering between distributed systems, and monitor critical state with checksums or reconciliation jobs that periodically compare system A and B.
Observability: correlation IDs, structured logging, traces
When a fault spans two or three systems, your ability to diagnose it depends almost entirely on the observability you built in during development. Adding it afterwards is technically possible but painful.
- Correlation IDs across the entire chain: generate a unique
X-Request-ID(ortraceparent) at the start of each flow and pass it on to every downstream call. Without this pattern, integration debugging becomes a reading exercise through twenty different logs, searching for overlapping timestamps. - Structured logging: log in JSON with consistent field names —
level,service,request_id,action,resource,duration_ms,outcome. Plain-text logs are unusable once you need to aggregate them across services. - Distributed tracing: OpenTelemetry is the standard. Spans around each API call give a visual picture of where the time in a request goes, often much faster than logging alone.
- Metrics that matter: success ratio per endpoint (not averaged across everything), P95/P99 latency instead of averages, queue depth for async flows, and dead-letter-queue volume as a business metric.
If you cannot establish within five minutes whether a specific transaction succeeded, what happened along the way, and which system caused the failure, your observability is not yet production-ready.
Schema versioning: on both sides of the integration
A schema is not a monument; it evolves. The difference between an integration that adapts smoothly and one that has to be "patched up" every year lies in how you handle versioning from day one.
On the vendor side: pin your integration to a specific API version. A client that says "v1" and implicitly gets the latest minor version can break along the way without anyone having made a release. Read changelog announcements and plan deprecations as regular work.
On your own API side: version explicitly — in the URL (/v1/orders), in a header (Accept: application/vnd.appfront.v1+json) or via a query parameter. Make backwards-incompatible changes only under a new version, and give old versions a documented sunset date.
Deliberately additive versus breaking: adding fields is almost always safe, provided consumers are built to tolerate schema changes. Removing fields, changing types or making validation stricter is always breaking and belongs in a version bump — to the consumer, a field that has silently become stricter is hard to tell apart from a bug.
Contract testing: the same language on both sides
A unit test verifies that your code does what you intend. A contract test verifies that your integration does what the other party expects, and vice versa. Tools like Pact were built for this: the consumer describes which calls it makes and which responses it expects; the provider verifies that its actual behaviour matches. This pays off most when you hold both the consumer and provider roles within the same organisation. For integrations with external vendors, integration tests against a sandbox usually suffice, with regular re-runs to catch regressions on the vendor's side. Combine mocked tests in CI with scheduled sandbox runs for genuine verification.
Error handling and dead-letter queues
What do you do with a message that keeps failing despite retries? "Retry indefinitely" is wrong, "discard it" is worse. The right answer is a dead-letter queue (DLQ): a separate holding area for messages your processing cannot handle, with enough context to be repaired manually or automatically. Categorise errors accordingly: transient (network glitch, vendor outage) goes into retry with backoff, permanent (4xx except 429) goes straight to the DLQ without retries, unknown (5xx, vague timeouts) is retried until the budget is exhausted and then sent to the DLQ.
A DLQ should hold the original payload, all failed attempts with error details, the timestamp, the correlation ID and, crucially, a mechanism to replay the message into the main flow once it has been corrected. A DLQ without a replay path becomes a digital graveyard. A DLQ that suddenly grows is always a signal that someone needs to look at it within the hour — define thresholds and alerts, and treat emptying the DLQ as routine work rather than crisis response.
Secrets management and security
API keys, OAuth client secrets, certificate private keys, database credentials — an integration layer is a factory full of secrets. Four rules:
- A central secret store: no secrets in code, no secrets in environment files that end up in Git. Use HashiCorp Vault, AWS Secrets Manager / KMS, Azure Key Vault, GCP Secret Manager or similar — with short-lived credentials and an audit log.
- Rotation as a rule, not an incident: automate rotation rather than doing it by hand after a leak. Refresh-token rotation for OAuth, vendor-specific flows for API keys, monitoring plus automatic renewal for certificates.
- Least privilege: integration accounts should only have what they need. A sync account for product data has no rights to customer data. A breach in one integration then does not become a breach across your entire systems landscape.
- Keep secrets out of logs: a serialiser that places the entire request, including the Authorization header, into an error message leaks secrets into your observability stack. Build redaction in as a first-class concern.
For the wider security context, see our information security policy for the stance we take on projects where this matters.
API gateway and backend-for-frontend
Once you have more than a handful of integrations, it pays to centralise access patterns. An API gateway sits between consumers and your back ends and handles cross-cutting concerns: authentication and token validation, rate limiting per client, request logging, caching, transformation. Examples include Kong, Apigee, AWS API Gateway, Tyk, or a custom build around Envoy. You don't rebuild this logic in every microservice, but make sure the gateway itself doesn't become a critical single point of failure.
A backend-for-frontend (BFF) is a thin layer built for one specific type of frontend (web, mobile, partner portal) that aggregates data from multiple backends. It avoids complex orchestration in the frontend and stops your microservices from filling up with frontend-specific endpoints. For frontend-heavy products where you consolidate several system landscapes, this is almost always worth it. Our page on integrations within web development covers this in more depth.
GraphQL versus REST: it depends on your use case
REST often wins for well-defined resources with clear CRUD semantics, caching at the edge (REST is HTTP-native, GraphQL much less so), public APIs for a broad audience that doesn't want to learn about schemas, and scenarios where you want to design idempotency and versioning explicitly per endpoint.
GraphQL often wins for frontend-heavy applications that aggregate data from many sources and suffer from over- or under-fetching, rapidly evolving client needs, and internal APIs where consumers and providers sit in the same team. Watch out for pitfalls: N+1 queries call for dataloaders, query complexity must be limited through cost analysis or allowlists, and caching is no longer an out-of-the-box HTTP cache. The pragmatic choice in many organisations: REST for public integration APIs, GraphQL for the BFF layer between internal frontends and backends.
The GDPR layer: data minimisation, EU residency, redaction, audit trails
An API integration that carries personal data falls under the GDPR. This isn't a coat of paint applied after the build, but a set of design requirements that run through the architecture.
- Data minimisation in payloads: send only what the recipient actually needs. A marketing tool doesn't need a citizen service number, and a logistics partner doesn't need a date of birth. Define an explicit projection of the source record for each integration, rather than passing the whole blob along.
- EU residency of endpoints: for GDPR-sensitive data, the physical location of the processor matters. Check whether each vendor offers EU hosting, choose it deliberately, and document it in your record of processing activities.
- PII redaction in logs: a log line containing
email=jansen@example.com order=1234 status=failedis a GDPR incident waiting to happen as soon as that log stream is accessed. Locate PII fields, redact or hash them, and treat error bodies with the same care as payload bodies. - Audit trails: immutable, append-only, with enough context to trace individual actions, and a retention period that matches your record of processing activities. For healthcare, NEN 7510 applies as well; for finance, DORA.
The mistakes we keep seeing
Finally, a collection of anti-patterns we regularly encounter in practice. Consider it a checklist rather than a judgement.
- Non-idempotent write operations without a safety net: a payment confirmation processed twice, or an order arriving in duplicate. Almost always preventable with idempotency keys.
- Retry storms: a vendor has three minutes of downtime, your 200 nodes all retry at once, and now the vendor has hours of downtime. Jitter, circuit breakers and retry budgets are there for good reason, not for show.
- Hidden coupling: system B relies on system A always responding quickly, without this being written down in any contract. When A's latency shifts, B breaks, invisibly, until it does.
- Versioning deadlocks: you need to move to a new API version, the vendor has removed the old one, consumers are still on the old version, and nobody feels responsible for the migration. Plan version transitions in advance with a sunset date.
- Missing observability until the first outage: you only notice an integration has broken when a customer calls. By the time that phone rings, you should already have had alerts.
- Security through obscurity: an endpoint that "nobody can find because the URL isn't documented". Scanners will find it quickly enough. Authentication and authorisation are non-negotiable, internal ones included.
- Lock-in to vendor-specific features: a handy proprietary feature, and a year later you want to migrate, but that feature doesn't exist anywhere else. Keep track in your architecture documentation of which vendor assumptions you accept and what exiting would cost.
An API integration is not a standalone HTTP-POST — it is a set of agreements between two systems that must keep working together under production pressure. The best practices above are those agreements expressed in a defensive form. None of them is exotic, and almost none of them is expensive. The difference between integrations that hold up and those that keep falling over lies in which ones you take seriously enough to build in before go-live.