+1 (726) 227-3745

Rate Limiting a MEAN API That Actually Holds: Express 5, Redis, Per-Tenant Quotas and Angular 22 Backoff

Every MEAN API we are asked to rescue has the same hole: there is no limit on how often anyone can call it. In 2026 that is no longer a theoretical risk. Credential-stuffing bots hammer /api/login, scrapers and AI crawlers pull entire collections one paginated request at a time, and a single enthusiastic customer script can saturate the Node event loop for everybody else. This tutorial adds a layered defense to an Express 5 + MongoDB 8 + Angular 22 application: a Redis-backed sliding-window limiter, stricter rules on authentication endpoints, per-tenant quotas, correct 429 semantics, and an Angular client that backs off instead of making things worse.

Versions used: Node 24 LTS, Express 5, Mongoose 8, MongoDB 8.0, Redis 7, Angular 22.

Decide what you are protecting against

Three different problems get lumped together as "rate limiting," and they need different controls:

  1. Abuse of credentials — brute force and credential stuffing on login, password reset, and token refresh. Needs tight, per-identifier limits with exponential lockout.
  2. Volumetric load — one caller consuming capacity meant for everyone. Needs a global per-IP or per-key limit plus concurrency control.
  3. Commercial quota — a paying tenant exceeding the plan they bought. Needs a counter tied to the tenant, measured per month or per day, and a clear error that tells them to upgrade.

Build them as separate middleware. Mixing them into one limiter produces a number that is simultaneously too strict for paying customers and too loose for attackers.

Why in-process counters are not enough

express-rate-limit with its default memory store is fine for a single process on a laptop. The moment you run two containers, or node --run cluster, each replica keeps its own counter and your "100 requests per minute" becomes 100 times the number of replicas. Worse, memory counters reset on every deploy, which attackers notice.

Use a shared store. Redis is the usual choice and the one below.

npm install express-rate-limit rate-limit-redis redis

A sliding-window limiter on Redis

src/limits.js:

import { rateLimit, ipKeyGenerator } from 'express-rate-limit';
import { RedisStore } from 'rate-limit-redis';
import { createClient } from 'redis';

export const redis = createClient({ url: process.env.REDIS_URL ?? 'redis://localhost:6379' });
redis.on('error', (err) => console.error('redis', err));
await redis.connect();

const store = (prefix) =>
  new RedisStore({
    prefix,
    sendCommand: (...args) => redis.sendCommand(args),
  });

// Identify the caller: API key, then authenticated user, then IP.
// ipKeyGenerator normalises IPv6 to a /64 so a single client cannot
// rotate through a huge address range for free.
const callerKey = (req) =>
  req.get('x-api-key') ?? req.user?.sub ?? ipKeyGenerator(req.ip);

export const globalLimiter = rateLimit({
  windowMs: 60_000,
  limit: 300,
  standardHeaders: 'draft-8',
  legacyHeaders: false,
  keyGenerator: callerKey,
  store: store('rl:global:'),
  message: { error: 'too_many_requests', detail: 'Slow down and retry shortly.' },
});

export const writeLimiter = rateLimit({
  windowMs: 60_000,
  limit: 60,
  standardHeaders: 'draft-8',
  legacyHeaders: false,
  keyGenerator: callerKey,
  store: store('rl:write:'),
});

export const authLimiter = rateLimit({
  windowMs: 15 * 60_000,
  limit: 10,
  // Successful logins should not burn the budget of a legitimate user.
  skipSuccessfulRequests: true,
  standardHeaders: 'draft-8',
  legacyHeaders: false,
  keyGenerator: (req) => `${ipKeyGenerator(req.ip)}:${String(req.body?.email ?? '').toLowerCase()}`,
  store: store('rl:auth:'),
});

Two details that matter more than the numbers:

  • standardHeaders: 'draft-8' emits the IETF RateLimit and RateLimit-Policy headers. Clients, including your own Angular app, can read remaining budget instead of guessing.
  • The auth key combines IP and email. Keying on email alone lets an attacker lock a real user out of their own account; keying on IP alone lets a botnet spread attempts across thousands of addresses. Both together is the usual compromise, and you should pair it with an account-level counter that triggers a step-up challenge rather than a hard lock.

Wire it into Express 5

Order matters. Cheap rejections go first, and the limiter must see the real client IP.

import express from 'express';
import { globalLimiter, writeLimiter, authLimiter } from './limits.js';

export const app = express();

// Only if you actually run behind a known proxy or load balancer.
// Never set this to `true` blindly: it lets anyone spoof X-Forwarded-For
// and defeat every IP-based limit you just wrote.
app.set('trust proxy', Number(process.env.TRUSTED_PROXY_HOPS ?? 1));

app.get('/healthz', (req, res) => res.json({ ok: true }));   // before the limiter

app.use(globalLimiter);
app.use(express.json({ limit: '100kb' }));                   // body size is a limit too

app.post('/api/login', authLimiter, loginHandler);
app.post('/api/password-reset', authLimiter, resetHandler);

app.use(['/api'], (req, res, next) =>
  ['POST', 'PUT', 'PATCH', 'DELETE'].includes(req.method) ? writeLimiter(req, res, next) : next(),
);

Health checks sit above the limiter so an attack does not make your orchestrator think the service is dead and restart it mid-incident. The express.json({ limit }) call is part of the same story: a 50 MB JSON body costs you CPU before any limiter opinion is formed.

Per-tenant quotas in MongoDB

Plan quotas are business state, not infrastructure state, so keep them next to the tenant. An atomic $inc with an upsert gives you a correct counter without a transaction:

import { Schema, model } from 'mongoose';

const usageSchema = new Schema({
  tenant: { type: Schema.Types.ObjectId, ref: 'Tenant', required: true },
  period: { type: String, required: true },   // '2026-09'
  calls:  { type: Number, default: 0 },
}, { timestamps: true });

usageSchema.index({ tenant: 1, period: 1 }, { unique: true });
export const Usage = model('Usage', usageSchema);

export function quota(planLimit) {
  return async function quotaMiddleware(req, res, next) {
    const period = new Date().toISOString().slice(0, 7);
    const limit = planLimit(req.tenant);
    const usage = await Usage.findOneAndUpdate(
      { tenant: req.tenant._id, period },
      { $inc: { calls: 1 } },
      { new: true, upsert: true },
    );
    res.set('X-Quota-Limit', String(limit));
    res.set('X-Quota-Remaining', String(Math.max(0, limit - usage.calls)));
    if (usage.calls > limit) {
      return res.status(429).json({
        error: 'quota_exceeded',
        detail: `Plan allows ${limit} calls in ${period}.`,
        upgrade: 'https://example.com/billing',
      });
    }
    next();
  };
}

Give the Usage collection a TTL index on createdAt if you do not need the history, and read the counter from a cached tenant document rather than joining on every request.

Shed load before the event loop melts

Rate limits are per caller. They do not help when ten thousand legitimate callers arrive at once. For that you want the server to notice it is behind and say so immediately:

import { monitorEventLoopDelay } from 'node:perf_hooks';

const histogram = monitorEventLoopDelay({ resolution: 10 });
histogram.enable();
setInterval(() => histogram.reset(), 5_000).unref();

app.use((req, res, next) => {
  const p99ms = histogram.percentile(99) / 1e6;
  if (p99ms > 250 && req.method === 'GET' && !req.path.startsWith('/healthz')) {
    res.set('Retry-After', '5');
    return res.status(503).json({ error: 'overloaded' });
  }
  next();
});

A fast 503 is kinder than a queued request that times out thirty seconds later holding a MongoDB connection the whole time. Pair it with a bounded Mongoose pool (maxPoolSize) so a traffic spike cannot open connections until the database refuses them.

Do not forget the crawlers

A large share of 2026 "attack" traffic is just automated agents reading everything you expose. For a MEAN app:

  • Serve a robots.txt that disallows /api/, and expect it to be ignored by the worst offenders.
  • Require an API key for list endpoints, even a free self-service one, so anonymous bulk reads are impossible.
  • Cap pagination server-side (limit = Math.min(Number(req.query.limit) || 25, 100)) — unbounded ?limit=100000 is how a single request becomes an outage.
  • Put static and cacheable GET responses behind a CDN so the Node process never sees repeat traffic. Our post on caching a MEAN app properly covers the ETag and Redis side of this.

Handle 429 properly in Angular 22

A client that retries instantly turns a rate limit into a self-inflicted denial of service. An interceptor that respects Retry-After with jittered backoff fixes that for the whole app:

import { HttpInterceptorFn, HttpErrorResponse } from '@angular/common/http';
import { throwError, timer } from 'rxjs';
import { retry } from 'rxjs/operators';

const delayFor = (error: HttpErrorResponse, attempt: number) => {
  const header = Number(error.headers.get('Retry-After'));
  const base = Number.isFinite(header) && header > 0 ? header * 1000 : 2 ** attempt * 500;
  return Math.min(base, 20_000) * (0.5 + Math.random() / 2);   // full-ish jitter
};

export const backoffInterceptor: HttpInterceptorFn = (req, next) =>
  next(req).pipe(
    retry({
      count: 3,
      delay: (error, attempt) => {
        const retryable = error instanceof HttpErrorResponse && [429, 503].includes(error.status);
        const idempotent = ['GET', 'HEAD', 'PUT', 'DELETE'].includes(req.method);
        if (!retryable || !(idempotent || req.headers.has('Idempotency-Key'))) {
          return throwError(() => error);
        }
        return timer(delayFor(error as HttpErrorResponse, attempt));
      },
    }),
  );

Note the idempotency check: silently replaying a POST that may have already succeeded is how duplicate orders happen. If you want retryable writes, send an Idempotency-Key and deduplicate server-side — the pattern in our idempotent webhook handling post applies to any write endpoint.

Surface the remaining budget in the UI with a signal fed from the RateLimit header, so power users see "18 of 300 requests left this minute" rather than an unexplained error.

Prove it works

Limits that were never tested are configuration, not protection. A minimal integration test:

import test from 'node:test';
import assert from 'node:assert/strict';
import request from 'supertest';
import { app } from '../src/app.js';

test('login is limited after 10 failures', async () => {
  const body = { email: 'victim@example.com', password: 'wrong-password' };
  for (let i = 0; i < 10; i++) await request(app).post('/api/login').send(body).expect(401);
  const res = await request(app).post('/api/login').send(body).expect(429);
  assert.ok(res.headers['ratelimit']);
});

Flush the Redis prefix between tests, then run a short k6 or autocannon burst against a staging deployment and watch three numbers: the 429 rate, the event loop p99, and MongoDB connection count. If the 429 rate is zero under a deliberate flood, your limiter is not wired into the path you think it is.

A sane starting configuration

SurfaceWindowLimitKey
All /api requests1 min300API key → user → IP /64
Writes1 min60same
Login / reset15 min10 failuresIP + email
Signup1 hour5IP /64
Export / report endpoints1 hour10user

Start generous, log every rejection with the key and route, and tighten once you have a week of real percentiles. The goal is for legitimate users never to see a 429 while an attacker sees nothing else.


If your MEAN API is currently running with no limits, or you have discovered them the hard way during an incident, our performance tuning and MEAN Stack consulting teams do this work as a short, fixed-scope engagement. Get in touch with your traffic numbers and we will tell you where the first failure is.