Every integration project reaches the same moment: the third-party API that worked perfectly in staging starts returning 429 Too Many Requests in production. Most tutorials about a Node.js retry API call exponential backoff pattern stop at a ten line setTimeout loop. That is not enough for real integration work, because a naive retry loop is exactly what turns a small rate limit warning into a full outage.
This guide covers the defensive part that usually gets skipped: soft rate limits enforced on our side, exponential backoff with jitter, respecting the Retry-After header, idempotency rules, and a shared request queue. All examples are production shaped code you can drop into a Node.js 22 or 24 service, using native fetch, Axios interceptors and Bottleneck.
Why a plain retry loop makes things worse
Imagine 40 concurrent jobs hitting a partner API that allows 100 requests per minute. The limit is reached, the API answers 429, and every job retries after a fixed 1 second delay. All 40 retries land in the same millisecond window. The API answers 429 again, and now you have a self inflicted thundering herd: constant traffic, zero successful calls, and often a temporary IP ban.
A resilient integration layer needs four things working together:
- Exponential backoff so the pressure on the remote API decreases over time.
- Jitter so concurrent clients spread out instead of synchronising.
- A soft rate limit (client side throttle) so you rarely hit 429 in the first place.
- Retry policy rules so you never retry something that is not safe to repeat.

Exponential backoff in one formula
The core formula is simple:
delay = min(maxDelay, baseDelay * 2 ** attempt)
With baseDelay = 300ms and maxDelay = 8000ms, the growth curve looks like this once full jitter is applied:
| Attempt | Exponential ceiling | Actual wait with full jitter | Worst case cumulative wait |
|---|---|---|---|
| 0 | 300 ms | 0 to 300 ms | 0.3 s |
| 1 | 600 ms | 0 to 600 ms | 0.9 s |
| 2 | 1 200 ms | 0 to 1 200 ms | 2.1 s |
| 3 | 2 400 ms | 0 to 2 400 ms | 4.5 s |
| 4 | 4 800 ms | 0 to 4 800 ms | 9.3 s |
| 5 | 8 000 ms (capped) | 0 to 8 000 ms | 17.3 s |
Always cap the delay. Without maxDelay, attempt 10 waits more than five minutes, which will time out your HTTP handler or your job runner long before the retry fires.
Which jitter strategy should you use?
| Strategy | Formula | When to use it |
|---|---|---|
| No jitter | base * 2 ** n |
Single process, single request. Avoid in anything concurrent. |
| Full jitter | random(0, base * 2 ** n) |
Best default. Maximum spread, lowest collision rate. |
| Equal jitter | d/2 + random(0, d/2) |
When you want a guaranteed minimum pause between attempts. |
| Decorrelated jitter | min(cap, random(base, prev * 3)) |
Very high fan out (hundreds of workers hitting one endpoint). |
Step 1: decide what is actually retryable
Retrying the wrong thing creates duplicate orders, duplicate emails and duplicate charges. Use this matrix as your policy baseline. p-retry vs async-retry vs exponential-backoff 2026 is a useful companion to this.
| Response | Retry? | Notes |
|---|---|---|
| 408, 425 | Yes | Timeout or too early, transient by definition. |
| 429 | Yes | Read Retry-After first, backoff only as a fallback. |
| 500, 502, 503, 504 | Yes | Server side, usually transient. |
| 400, 404, 409, 422 | No | Your payload is wrong. Retrying wastes quota. |
| 401, 403 | Once | Only after a token refresh, never in a loop. |
| ECONNRESET, ETIMEDOUT, EAI_AGAIN | Yes | Network layer failures with no response object. |
| POST without idempotency key | Careful | Only retry if the endpoint documents idempotency. |

Step 2: the shared backoff utility
Put the maths in one module so the fetch layer, the Axios layer and the queue layer all agree.
// lib/backoff.js
const RETRYABLE_STATUS = new Set([408, 425, 429, 500, 502, 503, 504]);
const RETRYABLE_CODES = new Set([
'ECONNRESET', 'ECONNREFUSED', 'ETIMEDOUT', 'EAI_AGAIN', 'EPIPE', 'ECONNABORTED'
]);
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
// Full jitter: random point between 0 and the exponential ceiling
function backoffDelay(attempt, { baseDelay = 300, maxDelay = 8000 } = {}) {
const ceiling = Math.min(maxDelay, baseDelay * 2 ** attempt);
return Math.round(Math.random() * ceiling);
}
// Retry-After can be seconds ("30") or an HTTP date
function retryAfterMs(headerValue, { hardCap = 60000 } = {}) {
if (!headerValue) return null;
const seconds = Number(headerValue);
if (Number.isFinite(seconds)) return Math.min(seconds * 1000, hardCap);
const timestamp = Date.parse(headerValue);
if (!Number.isFinite(timestamp)) return null;
return Math.min(Math.max(0, timestamp - Date.now()), hardCap);
}
module.exports = { RETRYABLE_STATUS, RETRYABLE_CODES, sleep, backoffDelay, retryAfterMs };
Two details that most examples miss:
Retry-Afterwins over your own formula. The server knows its own window better than your exponent does.- Cap the server hint too. Some APIs answer
Retry-After: 3600. Blocking a worker for an hour is not a retry, it is a hang. Cap it, then push the job back to your queue instead.
Step 3: retry with native fetch (no dependencies)
Node.js ships with fetch and AbortSignal, so a dependency free version is short. Note the timeout: a retry loop without a per attempt timeout can hang forever on a stalled socket.
// lib/fetchWithRetry.js
const { RETRYABLE_STATUS, sleep, backoffDelay, retryAfterMs } = require('./backoff');
async function fetchWithRetry(url, init = {}, options = {}) {
const { retries = 4, baseDelay = 300, maxDelay = 8000, timeout = 10000 } = options;
let lastError;
for (let attempt = 0; attempt <= retries; attempt++) {
try {
const response = await fetch(url, {
...init,
signal: AbortSignal.timeout(timeout)
});
const isLast = attempt === retries;
if (!RETRYABLE_STATUS.has(response.status) || isLast) return response;
const hint = retryAfterMs(response.headers.get('retry-after'));
const wait = hint ?? backoffDelay(attempt, { baseDelay, maxDelay });
console.warn('[retry] http', {
url, status: response.status, attempt: attempt + 1, waitMs: wait
});
// Drain the body so the socket returns to the pool
await response.body?.cancel();
await sleep(wait);
} catch (error) {
lastError = error;
if (attempt === retries) throw error;
const wait = backoffDelay(attempt, { baseDelay, maxDelay });
console.warn('[retry] network', { url, code: error.code, attempt: attempt + 1, waitMs: wait });
await sleep(wait);
}
}
throw lastError;
}
module.exports = { fetchWithRetry };

Step 4: the Axios interceptor version
If your codebase already uses Axios, the cleanest place for retry logic is a response error interceptor. It keeps every call site clean and applies the policy globally. A longer analysis is available for anyone who wants it.
// lib/httpClient.js
const axios = require('axios');
const {
RETRYABLE_STATUS, RETRYABLE_CODES, sleep, backoffDelay, retryAfterMs
} = require('./backoff');
const SAFE_METHODS = new Set(['get', 'head', 'options', 'put', 'delete']);
const client = axios.create({
baseURL: process.env.PARTNER_API_URL,
timeout: 10000,
headers: { Accept: 'application/json' }
});
client.interceptors.response.use(null, async (error) => {
const config = error.config;
if (!config) return Promise.reject(error);
const policy = { retries: 4, baseDelay: 300, maxDelay: 8000, ...(config.retryPolicy || {}) };
config.retryState = config.retryState || { attempt: 0 };
const status = error.response?.status;
const method = (config.method || 'get').toLowerCase();
const transient = status
? RETRYABLE_STATUS.has(status)
: RETRYABLE_CODES.has(error.code);
const idempotent = SAFE_METHODS.has(method)
|| Boolean(config.headers?.['Idempotency-Key'])
|| config.forceRetry === true;
if (!transient || !idempotent || config.retryState.attempt >= policy.retries) {
return Promise.reject(error);
}
const hint = retryAfterMs(error.response?.headers?.['retry-after']);
const wait = hint ?? backoffDelay(config.retryState.attempt, policy);
config.retryState.attempt += 1;
if (status === 429) tripCooldown(wait); // see the queue section below
console.warn('[axios-retry]', {
url: config.url,
method,
status: status || error.code,
attempt: config.retryState.attempt,
waitMs: wait,
serverHint: hint !== null
});
await sleep(wait);
return client.request(config);
});
module.exports = { client };
Interceptor gotchas
- Never mutate a shared config object. Store the counter on the request config (
config.retryState), which is per call. - Streams cannot be replayed. If
config.datais a readable stream or a form built from a file handle, the retry will send an empty body. Rebuild the payload with a factory function instead. - Axios timeouts surface as
ECONNABORTED(andETIMEDOUTin some adapters), so include both in the retryable code list. - Keep the total budget in mind. 4 retries with a 10 s timeout each can occupy a worker for almost a minute.
Step 5: the missing piece, a soft rate limit with Bottleneck
Backoff is reactive. It only helps after you have already been rejected. A soft rate limit is proactive: you throttle yourself slightly below the documented quota so 429 becomes a rare event rather than a daily routine.
Bottleneck is still the most practical tool for this in Node.js because it gives you a reservoir (a quota window), concurrency control and a minimum spacing between jobs.
// lib/limiter.js
const Bottleneck = require('bottleneck');
// Partner quota: 100 requests / minute. We run at 80 % of it.
const limiter = new Bottleneck({
reservoir: 80,
reservoirRefreshAmount: 80,
reservoirRefreshInterval: 60 * 1000, // must be a multiple of 250
maxConcurrent: 5,
minTime: 150, // ~6.6 req/s ceiling in bursts
highWater: 1000,
strategy: Bottleneck.strategy.OVERFLOW
});
// Global cooldown gate: any 429 pauses every queued job
let cooldown = Promise.resolve();
function tripCooldown(ms) {
const until = Date.now() + ms;
cooldown = cooldown.then(() => {
const remaining = until - Date.now();
return remaining > 0 ? new Promise((r) => setTimeout(r, remaining)) : null;
});
}
function schedule(fn, opts = {}) {
return limiter.schedule(opts, async () => {
await cooldown;
return fn();
});
}
limiter.on('depleted', () => console.warn('[limiter] reservoir empty, queueing'));
limiter.on('dropped', (dropped) => console.error('[limiter] job dropped', dropped.args));
module.exports = { limiter, schedule, tripCooldown };
Usage becomes boringly simple, which is the goal:
const { client } = require('./httpClient');
const { schedule } = require('./limiter');
const getInvoice = (id) =>
schedule(() => client.get(`/invoices/${id}`), { priority: 5 });
const syncInvoices = (ids) => Promise.all(ids.map(getInvoice));
With this in place, Promise.all over 500 ids no longer fires 500 simultaneous requests. Bottleneck releases them at your chosen pace, and the interceptor handles the rare 429 that slips through.
Mapping API docs to Bottleneck options
| What the API docs say | Bottleneck option | Example value |
|---|---|---|
| 100 requests per minute | reservoir + reservoirRefreshInterval |
80 / 60000 |
| Max 10 concurrent connections | maxConcurrent |
5 |
| No more than 5 req/s | minTime |
250 |
| Quota shared across all your servers | datastore: 'ioredis' + id |
clustered limiter |
| Bulk endpoints preferred | priority and batching |
1 to 9 |
Running more than one instance? An in memory limiter per pod means your effective rate is limit x pods. Use the clustered mode so the quota is enforced globally:
const limiter = new Bottleneck({
id: 'partner-api',
datastore: 'ioredis',
clearDatastore: false,
clientOptions: { host: process.env.REDIS_HOST, port: 6379 },
reservoir: 80,
reservoirRefreshAmount: 80,
reservoirRefreshInterval: 60 * 1000,
maxConcurrent: 5,
minTime: 150
});
Alternative: let Bottleneck own the retries
Bottleneck can also drive the backoff itself through its failed event, which keeps the retry inside the queue accounting:
const { backoffDelay, retryAfterMs } = require('./backoff');
limiter.on('failed', async (error, jobInfo) => {
const status = error.response?.status;
const retryable = status === 429 || status >= 500;
if (!retryable || jobInfo.retryCount >= 3) return; // undefined = give up
const hint = retryAfterMs(error.response?.headers?.['retry-after']);
return hint ?? backoffDelay(jobInfo.retryCount, { baseDelay: 500, maxDelay: 10000 });
});
limiter.on('retry', (message, jobInfo) =>
console.warn('[limiter] retrying', { retryCount: jobInfo.retryCount })
);
Pick one owner for retries, not two. If both the Axios interceptor and the Bottleneck failed handler retry, you get 4 x 3 = 12 attempts and a very confused quota. Our recommendation: interceptor for HTTP level retries, limiter for pacing only.

Step 6: stop retrying when the API is genuinely down
Backoff assumes the failure is transient. When a provider has a real outage, retries just burn CPU and delay your error handling. Add a circuit breaker on top:
const CircuitBreaker = require('opossum');
const { schedule } = require('./limiter');
const { client } = require('./httpClient');
const breaker = new CircuitBreaker(
(id) => schedule(() => client.get(`/invoices/${id}`)),
{
timeout: 15000,
errorThresholdPercentage: 50,
resetTimeout: 30000,
volumeThreshold: 10
}
);
breaker.fallback(() => ({ data: null, degraded: true }));
breaker.on('open', () => console.error('[breaker] partner API circuit OPEN'));
The full defensive stack, from the outside in:
- Circuit breaker: is this dependency healthy at all?
- Limiter / queue: am I allowed to send a request right now?
- Retry interceptor: this specific attempt failed transiently, backoff and repeat.
- Per attempt timeout: no attempt may hang forever.
Observability: what to log on every retry
Retries that are invisible are retries you cannot tune. Log a structured line per attempt with: target URL, method, status or error code, attempt number, computed wait, whether Retry-After was present, and a correlation id. Then track these metrics:
http_client_retries_totalby host and statushttp_client_429_totalby host (your soft limit tuning signal)limiter_queue_depthandlimiter_wait_secondshttp_client_give_up_total(retries exhausted, the real incident metric)
If 429_total stays above zero during normal traffic, your reservoir is set too high. Lower it by 10 % and observe again.

Testing your backoff without waiting 17 seconds
Use nock to script the failures and fake timers to skip the waiting:
const nock = require('nock');
const { client } = require('../lib/httpClient');
jest.useFakeTimers({ doNotFake: ['nextTick'] });
test('retries a 429 and honours Retry-After', async () => {
nock('https://api.partner.test')
.get('/invoices/42').reply(429, {}, { 'retry-after': '2' })
.get('/invoices/42').reply(200, { id: 42 });
const promise = client.get('/invoices/42');
await jest.advanceTimersByTimeAsync(2000);
const response = await promise;
expect(response.status).toBe(200);
});
Also write a test that asserts no retry happens on a 422. That is the test that protects you from duplicate side effects six months from now.
Production checklist
- Exponential backoff with a hard
maxDelaycap: yes - Full jitter applied to every delay: yes
Retry-Afterhonoured and capped: yes- Retry allow list by status code and error code, not a catch all: yes
- Non idempotent writes protected by an idempotency key: yes
- Per attempt timeout plus a total time budget: yes
- Client side soft rate limit below the documented quota: yes
- Shared limiter across instances when running multiple pods: yes
- Only one component owns retries: yes
- Structured retry logs and give up metrics: yes
- Circuit breaker or dead letter queue for sustained outages: yes
FAQ
How many retries should I configure in Node.js?
For a user facing HTTP request, 2 to 3 attempts with a 300 ms base and a 3 s cap keeps the total under a typical request budget. For background jobs, 5 attempts with an 8 to 30 s cap is reasonable, and after that the job should return to a queue with a much longer delay rather than looping in memory.
Is exponential backoff enough to avoid 429 errors?
No. Backoff only reacts after rejection. To largely avoid 429 you need a client side throttle (a soft rate limit) sized below the provider quota, plus batching where the API offers bulk endpoints. Backoff is the safety net, not the plan.
Should I use a library or write my own retry helper?
Libraries such as exponential-backoff, p-retry or axios-retry are fine for the pure delay maths. Write your own thin wrapper when you need custom rules: Retry-After parsing, idempotency checks, cooldown gates shared with a limiter, and structured logging. The 40 lines in this article are usually easier to reason about than tuning a generic library through options.
Can I retry POST requests safely?
Only when the operation is idempotent. Either the API supports an Idempotency-Key header (most modern payment and billing APIs do), or you can deduplicate server side with a natural key. Without one of those, a retried POST after a timeout may create a duplicate record because the first request could have succeeded before the connection dropped.
What is jitter and why is it mandatory in concurrent code?
Jitter is randomness added to the delay. Without it, all clients that failed at the same moment retry at the same moment, so the load pattern stays spiky and the API keeps rejecting. Full jitter (a random value between 0 and the exponential ceiling) spreads retries across the whole window and measurably increases success rates.
Where should the retry logic live: HTTP client, service layer or queue?
Transport level concerns (network errors, 429, 5xx, timeouts) belong in the HTTP client through an interceptor. Business level recovery (this sync failed entirely, reschedule it in 15 minutes, notify the operator) belongs in the job queue. Mixing the two produces retry storms that are very hard to debug.
Need this hardened in your own integration?
Defensive integration work is invisible when it succeeds and extremely visible when it does not. At Box Software we build and audit Node.js integration layers: retry policies, rate limit budgeting, idempotency, queueing and observability for third-party APIs. If your service is currently living with occasional 429 storms or silent duplicate writes, get in touch and we will review your client layer.
