Skip to content
DocspackagesDocumentation

@lunora/ratelimit

Token-bucket, fixed-window, and sliding-window rate limiting as procedure middleware.

PackagesRatelimit

@lunora/ratelimit enforces named rate limits over a pluggable store. You build one RateLimiter per app from a config map, then either attach the rateLimit middleware to a procedure's .use(...) chain or call limiter.limit(...) directly. On rejection the middleware throws a structural LunoraError that the runtime maps to 429 Too Many Requests (or 403 Forbidden for a deny-list hit), carrying data.retryAfterMs — the retry hint the client and every SDK read off a TOO_MANY_REQUESTS envelope (protocol/README.md §4.3).

import type { RateLimitConfigMap } from "@lunora/ratelimit";
import { dbRateLimit } from "@lunora/ratelimit";

import { mutation } from "./_generated/server";

// `satisfies` is load-bearing: a bare `const` widens each `kind` to `string`,
// and the config map needs the literal to pick the algorithm.
const config = {
    login: { kind: "fixed window", period: 60_000, rate: 5 },
    send: { kind: "token bucket", period: 1_000, rate: 10, capacity: 20 },
} satisfies RateLimitConfigMap;

// `key` returns `string | undefined` (omit it for a global limit), while
// `ctx.auth.userId` is `string | null` — collapse the null yourself.
export const send = mutation.use(dbRateLimit(config, "send", { key: (ctx) => ctx.auth.userId ?? "anonymous" })).mutation(async ({ ctx }) => {
    // …
});

A RateLimiter built with no explicit store falls back to createMemoryStore — a per-Durable-Object-instance in-memory Map. That's correct for a limit that only ever needs to hold within one DO instance, but it stops enforcing the configured rate the moment the instance is sharded, replicated (.global()), or recreated after eviction. dbRateLimit above sidesteps the question by building the limiter per call, backed by a Lunora table via ctx.db — see Pluggable stores. Construction without an explicit store now warns once so the choice is visible instead of silent.

Like row-level security and data masking, it rides the .use(...) chain and is opt-in per procedure: a bare query/mutation is never rate-limited.

Algorithms

Each named limit declares a kind. Pick by the burst behavior you want:

kindBehaviorPick it when
"token bucket"Tokens refill continuously at rate / period per ms up to capacity; a fresh key starts full, so a burst is allowed.You want smooth throughput that tolerates short bursts (API calls).
"fixed window"rate tokens granted at the start of each window aligned to start + n * period. With capacity > rate, unused tokens roll over and a fresh key starts at capacity.You want a hard cap per discrete window (e.g. 5 logins per minute).
"sliding window"A weighted estimate blending the current and previous window's counts; always caps at rate per period.You want fixed-window's simplicity without its boundary-burst spike.
const limiter = new RateLimiter({
    config: {
        api: { kind: "token bucket", period: 1_000, rate: 10 },
        login: { kind: "fixed window", period: 60_000, rate: 5 },
        search: { kind: "sliding window", period: 10_000, rate: 30 },
    },
});

Limit config

Each entry in the config map is a RateLimitConfig:

FieldTypeDefaultNotes
kindRateLimitKind—"token bucket", "fixed window", or "sliding window". Required.
ratenumber—Tokens granted per period. Required; must be a positive number.
periodnumber—Window/refill period in milliseconds. Required; must be a positive number.
capacitynumberrateRollover ceiling. Caps a token-bucket burst; for fixed windows enables cross-window rollover. Ignored by sliding windows.
startnumber0Phase offset in epoch ms for windowed algorithms; windows align to start + n * period. Ignored by token buckets.

The constructor validates every config up front: a non-positive period or rate, or a negative capacity throws at construction rather than corrupting accounting later.

Consuming a limit

limiter.limit(name, args) consumes capacity and returns a RateLimitStatus ({ ok, reason?, retryAfter }). limiter.check(name, args) peeks without consuming. limiter.reset(name, { key }) clears accounting for a pair (e.g. on a successful login).

import { RateLimiter, RateLimitError, createDbStore } from "@lunora/ratelimit";

import { action } from "./_generated/server";

// An action, not a mutation: see the callout below.
export const login = action.action(async ({ ctx, args }) => {
    const limiter = new RateLimiter({ config, store: createDbStore({ db: ctx.db }) });
    const status = await limiter.limit("login", { key: args.email });

    if (!status.ok) {
        // Stop here. `limit` has already consumed capacity; carrying on would
        // authenticate the attempt the limiter just rejected.
        // status.retryAfter is milliseconds until the request would succeed.
        // status.reason is "rate" or "deny".
        throw new RateLimitError(status);
    }

    await authenticate(args);

    // Only AFTER a successful login — resetting on the way past a failed one
    // hands the attacker a fresh budget on every wrong password.
    await limiter.reset("login", { key: args.email });
});

DB-store consumption commits with the mutation. A mutation's ctx.db writes ride its storage transaction, and a throw rolls all of them back — including the unit limit() just consumed. Written as a mutation, the example above charges nothing for a wrong password (authenticate throws), so only successful logins count and an attacker gets an unlimited number of guesses. Consume from an action when failed attempts must count: its writes commit on their own, so the unit stays spent whether or not the handler throws. The same applies to .use(dbRateLimit(...)) on a mutation whose handler rejects. A mutation that must keep the charge can return a failure value instead of throwing.

Per-call RateLimitArgs:

OptionTypeDefaultNotes
keystring—Sub-key isolating the limit (per user/team/IP). Omit for a global limit.
countnumber1Units to consume. Must be a positive integer.
reservebooleanfalsePermit now and reserve future capacity (the stored value goes negative); retryAfter reports when the debt clears. A count above capacity is never permitted — see below.
throwsbooleanfalseThrow RateLimitError instead of returning a failing status.

A count larger than the limit's capacity is caller misuse, not a rate rejection. It can never succeed no matter how long you wait, so it does not return { ok: false } and does not honour throws: false — it throws an INTERNAL LunoraError (a 500), with or without reserve. rateLimit(...) middleware re-throws it as-is rather than masking it behind the fail-open/fail-closed policy, which is reserved for genuine store outages. Clamp the count, or raise the limit's capacity.

Procedure middleware

rateLimit(limiter, name, options) returns a Middleware you attach with .use(...). The first argument is a LimiterResolver: either a fixed RateLimiter or a (ctx) => RateLimiter that binds a context-derived limiter at call time (handy for an ORM-backed store, see Stores).

import { RateLimiter, rateLimit, createDbStore } from "@lunora/ratelimit";

import { mutation } from "./_generated/server";

const config = { send: { kind: "token bucket", period: 1_000, rate: 10 } } as const;

export const send = mutation
    .use(
        rateLimit((ctx) => new RateLimiter({ config, store: createDbStore({ db: ctx.db }) }), "send", {
            key: (ctx) => ctx.auth.userId ?? "anonymous",
        }),
    )
    .mutation(async ({ ctx, args }) => {
        // …
    });

For the common DB-backed case, dbRateLimit(config, name, options) is shorthand for the resolver above; it builds the per-call RateLimiter + createDbStore for you:

import { dbRateLimit } from "@lunora/ratelimit";

import { mutation } from "./_generated/server";

const config = { send: { kind: "token bucket", period: 1_000, rate: 10 } } satisfies RateLimitConfigMap;

export const send = mutation.use(dbRateLimit(config, "send", { key: (ctx) => ctx.auth.userId ?? "anonymous" })).mutation(async ({ ctx, args }) => {
    // …
});

Pass options.store to point at a non-default backing table/index/key column; the rest of options matches rateLimit's.

RateLimitMiddlewareOptions:

OptionTypeDefaultNotes
key(ctx) => string | undefined—Sub-key derived from context. Omit for a global limit. A resolver that returns undefined throws INTERNAL rather than silently pooling every keyless caller into one bucket.
countnumber1Units to consume per call.
messagestring—Override the error message thrown on rejection.
failOpenbooleanfalseBehavior when the limiter itself throws (store unavailable, etc). See below.

Failure policy: the middleware fails closed by default: if resolving or invoking the limiter throws, it logs via console.error and rejects with 503. This is the safer default for security-sensitive limits (auth, account creation). Set failOpen: true only when degraded availability is preferable to refusal, because a failing limiter then admits every request.

Deny list

A denyList passed to the constructor short-circuits before any token accounting: a matching key is denied with reason: "deny" and retryAfter: Infinity (the middleware maps this to 403 Forbidden).

const limiter = new RateLimiter({
    config: { api: { kind: "token bucket", period: 1_000, rate: 10 } },
    denyList: ["198.51.100.7", "abuser@example.com"],
});

The list is consulted as-is. If you also pass a normalize function (applied to every incoming key for case-folding, trimming, or canonicalizing IPs/emails), both the normalized and raw form are checked, but normalize your deny-list entries up front to be safe.

Pluggable stores

Persistence is a RateLimitStore (get/set/delete, sync or async):

FactoryBackingUse when
createMemoryStore()An in-process Map. The default when no store is given.Single-DO/process limits where state needn't survive eviction.
createSqlStore({ sql })state.storage.sql (workerd SqlStorage, also node:sqlite).Durable per-DO state that survives hibernation, and you have raw SQL.
createDbStore({ db })A Lunora table via ctx.db (the ORM writer on a mutation/action).Durable per-DO limits inside a procedure, where there is no raw SQL. Inside a mutation the consumed unit commits with the handler and a throw returns it.
createReadOnlyDbStore({ db })The same table through a query context's reader ctx.db.Showing remaining quota from a query — getValue/check only; limit/reset throw.

createReadOnlyDbStore exists so a remaining-quota display does not have to cast a query's ctx.db to the writer type. getValue and check project the stored value forward to the current clock, which is a pure read; only limit/reset write, and those reject with a message naming createDbStore as the fix.

createSqlStore creates its table (default _lunora_rate_limits) if missing. createDbStore expects a table you declare in your schema with a key column and its index (defaults: table rateLimits, index by_key, key column key):

import { defineTable, v } from "lunorash/server";

export const rateLimits = defineTable({
    key: v.string(),
    ts: v.number(),
    value: v.number(),
    prev: v.optional(v.number()),
}).index("by_key", ["key"]);

Inside a Durable Object the input gate serializes the limiter's read-modify-write against storage, so both SQL- and ORM-backed stores are atomic against concurrent calls without an explicit transaction. Driving createSqlStore outside a DO, you own serialization yourself.

Token budgets

A request limit is the wrong shape for an LLM: one call can cost a hundred tokens or a hundred thousand, so "20 calls a minute" bounds nothing that matters. tokenBudget binds one named limit whose unit is model tokens (or cents) and spends it in arrears, which is how tokens actually work — the cost is only known once the call returns.

import { tokenBudget } from "@lunora/ratelimit";

const budget = tokenBudget(limiter, "tokens");
const allowed = await budget.check(userId);

if (!allowed.ok) {
    throw new LunoraError("RATE_LIMITED", `token budget exhausted; retry in ${String(allowed.retryAfter)}ms`);
}

const { text, usage } = await generateText({ model: ctx.ai.model(), prompt });

await budget.record(userId, usage?.totalTokens ?? 0);

return text;

check runs before the call and consumes nothing: a tenant whose bucket is already empty (or negative, from an expensive previous call) is turned away. record runs after it — always, including when the call threw, since a failed generation that consumed input tokens still has to be paid for. The charge is a reservation, so the bucket may go negative: the tokens are already spent, and the next call is refused until the debt refills away. A charge larger than the limit's own ceiling is clamped to it rather than rejected, so the one oversized call the budget exists to catch cannot escape it by being too big to record.

Plugin form

ratelimitPlugin(limiter) packages the limiter as a Plugin that exposes the resolved RateLimiter under ctx.api.ratelimit, so a handler can call limit()/check()/reset() programmatically instead of (or alongside) the enforcing middleware. Install its middleware with one .use(...):

import { RateLimiter, ratelimitPlugin } from "@lunora/ratelimit";

import { mutation } from "./_generated/server";

const limiter = new RateLimiter({
    config: { send: { kind: "token bucket", period: 60_000, rate: 5, capacity: 5 } },
});

export const send = mutation.use(ratelimitPlugin(limiter).middleware!).mutation(async ({ ctx }) => {
    const status = await ctx.api.ratelimit.limit("send", { key: ctx.auth.userId ?? "anonymous" });

    if (!status.ok) {
        throw new Error("slow down");
    }
    // …
});

The plugin ships no schema extension (persistence is whatever store the limiter was built with), so it is middleware-only and skipped by the plugin schema fold.