@lunora/ratelimit enforces named rate limits over a pluggable store. You build
one RateLimiter per app from a config map, then either attach the rateLimit
middleware to a procedure's .use(...) chain or call limiter.limit(...)
directly. On rejection the middleware throws a structural LunoraError that the
runtime maps to 429 Too Many Requests (or 403 Forbidden for a deny-list hit),
carrying data.retryAfterMs — the retry hint the client and every SDK read off
a TOO_MANY_REQUESTS envelope (protocol/README.md §4.3).
import type { RateLimitConfigMap } from "@lunora/ratelimit";
import { dbRateLimit } from "@lunora/ratelimit";
import { mutation } from "./_generated/server";
// `satisfies` is load-bearing: a bare `const` widens each `kind` to `string`,
// and the config map needs the literal to pick the algorithm.
const config = {
login: { kind: "fixed window", period: 60_000, rate: 5 },
send: { kind: "token bucket", period: 1_000, rate: 10, capacity: 20 },
} satisfies RateLimitConfigMap;
// `key` returns `string | undefined` (omit it for a global limit), while
// `ctx.auth.userId` is `string | null` — collapse the null yourself.
export const send = mutation.use(dbRateLimit(config, "send", { key: (ctx) => ctx.auth.userId ?? "anonymous" })).mutation(async ({ ctx }) => {
// …
});A RateLimiter built with no explicit store falls back to createMemoryStore — a per-Durable-Object-instance in-memory Map. That's correct for a limit
that only ever needs to hold within one DO instance, but it stops enforcing the configured rate the moment the instance is sharded, replicated
(.global()), or recreated after eviction. dbRateLimit above sidesteps the question by building the limiter per call, backed by a Lunora table via
ctx.db — see Pluggable stores. Construction without an explicit store now warns once so the choice is visible instead of silent.
Like row-level security and data masking,
it rides the .use(...) chain and is opt-in per procedure: a bare
query/mutation is never rate-limited.
Algorithms
Each named limit declares a kind. Pick by the burst behavior you want:
kind | Behavior | Pick it when |
|---|---|---|
"token bucket" | Tokens refill continuously at rate / period per ms up to capacity; a fresh key starts full, so a burst is allowed. | You want smooth throughput that tolerates short bursts (API calls). |
"fixed window" | rate tokens granted at the start of each window aligned to start + n * period. With capacity > rate, unused tokens roll over and a fresh key starts at capacity. | You want a hard cap per discrete window (e.g. 5 logins per minute). |
"sliding window" | A weighted estimate blending the current and previous window's counts; always caps at rate per period. | You want fixed-window's simplicity without its boundary-burst spike. |
const limiter = new RateLimiter({
config: {
api: { kind: "token bucket", period: 1_000, rate: 10 },
login: { kind: "fixed window", period: 60_000, rate: 5 },
search: { kind: "sliding window", period: 10_000, rate: 30 },
},
});Limit config
Each entry in the config map is a RateLimitConfig:
| Field | Type | Default | Notes |
|---|---|---|---|
kind | RateLimitKind | — | "token bucket", "fixed window", or "sliding window". Required. |
rate | number | — | Tokens granted per period. Required; must be a positive number. |
period | number | — | Window/refill period in milliseconds. Required; must be a positive number. |
capacity | number | rate | Rollover ceiling. Caps a token-bucket burst; for fixed windows enables cross-window rollover. Ignored by sliding windows. |
start | number | 0 | Phase offset in epoch ms for windowed algorithms; windows align to start + n * period. Ignored by token buckets. |
The constructor validates every config up front: a non-positive period or rate, or a negative capacity throws at construction rather than corrupting
accounting later.
Consuming a limit
limiter.limit(name, args) consumes capacity and returns a RateLimitStatus
({ ok, reason?, retryAfter }). limiter.check(name, args) peeks without
consuming. limiter.reset(name, { key }) clears accounting for a pair (e.g. on a
successful login).
import { RateLimiter, RateLimitError, createDbStore } from "@lunora/ratelimit";
import { action } from "./_generated/server";
// An action, not a mutation: see the callout below.
export const login = action.action(async ({ ctx, args }) => {
const limiter = new RateLimiter({ config, store: createDbStore({ db: ctx.db }) });
const status = await limiter.limit("login", { key: args.email });
if (!status.ok) {
// Stop here. `limit` has already consumed capacity; carrying on would
// authenticate the attempt the limiter just rejected.
// status.retryAfter is milliseconds until the request would succeed.
// status.reason is "rate" or "deny".
throw new RateLimitError(status);
}
await authenticate(args);
// Only AFTER a successful login — resetting on the way past a failed one
// hands the attacker a fresh budget on every wrong password.
await limiter.reset("login", { key: args.email });
});DB-store consumption commits with the mutation. A mutation's ctx.db writes ride its storage transaction, and a throw rolls all of them back —
including the unit limit() just consumed. Written as a mutation, the example above charges nothing for a wrong password (authenticate throws), so only
successful logins count and an attacker gets an unlimited number of guesses. Consume from an action when failed attempts must count: its writes commit
on their own, so the unit stays spent whether or not the handler throws. The same applies to .use(dbRateLimit(...)) on a mutation whose handler rejects. A
mutation that must keep the charge can return a failure value instead of throwing.
Per-call RateLimitArgs:
| Option | Type | Default | Notes |
|---|---|---|---|
key | string | — | Sub-key isolating the limit (per user/team/IP). Omit for a global limit. |
count | number | 1 | Units to consume. Must be a positive integer. |
reserve | boolean | false | Permit now and reserve future capacity (the stored value goes negative); retryAfter reports when the debt clears. A count above capacity is never permitted — see below. |
throws | boolean | false | Throw RateLimitError instead of returning a failing status. |
A count larger than the limit's capacity is caller misuse, not a rate rejection. It can never succeed no matter how long you wait, so it does not
return { ok: false } and does not honour throws: false — it throws an INTERNAL LunoraError (a 500), with or without reserve. rateLimit(...)
middleware re-throws it as-is rather than masking it behind the fail-open/fail-closed policy, which is reserved for genuine store outages. Clamp the
count, or raise the limit's capacity.
Procedure middleware
rateLimit(limiter, name, options) returns a Middleware you attach with
.use(...). The first argument is a LimiterResolver: either a fixed
RateLimiter or a (ctx) => RateLimiter that binds a context-derived limiter at
call time (handy for an ORM-backed store, see Stores).
import { RateLimiter, rateLimit, createDbStore } from "@lunora/ratelimit";
import { mutation } from "./_generated/server";
const config = { send: { kind: "token bucket", period: 1_000, rate: 10 } } as const;
export const send = mutation
.use(
rateLimit((ctx) => new RateLimiter({ config, store: createDbStore({ db: ctx.db }) }), "send", {
key: (ctx) => ctx.auth.userId ?? "anonymous",
}),
)
.mutation(async ({ ctx, args }) => {
// …
});For the common DB-backed case, dbRateLimit(config, name, options) is shorthand
for the resolver above; it builds the per-call RateLimiter + createDbStore
for you:
import { dbRateLimit } from "@lunora/ratelimit";
import { mutation } from "./_generated/server";
const config = { send: { kind: "token bucket", period: 1_000, rate: 10 } } satisfies RateLimitConfigMap;
export const send = mutation.use(dbRateLimit(config, "send", { key: (ctx) => ctx.auth.userId ?? "anonymous" })).mutation(async ({ ctx, args }) => {
// …
});Pass options.store to point at a non-default backing table/index/key column;
the rest of options matches rateLimit's.
RateLimitMiddlewareOptions:
| Option | Type | Default | Notes |
|---|---|---|---|
key | (ctx) => string | undefined | — | Sub-key derived from context. Omit for a global limit. A resolver that returns undefined throws INTERNAL rather than silently pooling every keyless caller into one bucket. |
count | number | 1 | Units to consume per call. |
message | string | — | Override the error message thrown on rejection. |
failOpen | boolean | false | Behavior when the limiter itself throws (store unavailable, etc). See below. |
Failure policy: the middleware fails closed by default: if resolving or invoking the limiter throws, it logs via console.error and rejects with
503. This is the safer default for security-sensitive limits (auth, account creation). Set failOpen: true only when degraded availability is preferable
to refusal, because a failing limiter then admits every request.
Deny list
A denyList passed to the constructor short-circuits before any token
accounting: a matching key is denied with reason: "deny" and
retryAfter: Infinity (the middleware maps this to 403 Forbidden).
const limiter = new RateLimiter({
config: { api: { kind: "token bucket", period: 1_000, rate: 10 } },
denyList: ["198.51.100.7", "abuser@example.com"],
});The list is consulted as-is. If you also pass a normalize function (applied to
every incoming key for case-folding, trimming, or canonicalizing IPs/emails),
both the normalized and raw form are checked, but normalize your deny-list
entries up front to be safe.
Pluggable stores
Persistence is a RateLimitStore (get/set/delete, sync or async):
| Factory | Backing | Use when |
|---|---|---|
createMemoryStore() | An in-process Map. The default when no store is given. | Single-DO/process limits where state needn't survive eviction. |
createSqlStore({ sql }) | state.storage.sql (workerd SqlStorage, also node:sqlite). | Durable per-DO state that survives hibernation, and you have raw SQL. |
createDbStore({ db }) | A Lunora table via ctx.db (the ORM writer on a mutation/action). | Durable per-DO limits inside a procedure, where there is no raw SQL. Inside a mutation the consumed unit commits with the handler and a throw returns it. |
createReadOnlyDbStore({ db }) | The same table through a query context's reader ctx.db. | Showing remaining quota from a query — getValue/check only; limit/reset throw. |
createReadOnlyDbStore exists so a remaining-quota display does not have to cast
a query's ctx.db to the writer type. getValue and check project the stored
value forward to the current clock, which is a pure read; only limit/reset
write, and those reject with a message naming createDbStore as the fix.
createSqlStore creates its table (default _lunora_rate_limits) if missing.
createDbStore expects a table you declare in your schema with a key column and
its index (defaults: table rateLimits, index by_key, key column key):
import { defineTable, v } from "lunorash/server";
export const rateLimits = defineTable({
key: v.string(),
ts: v.number(),
value: v.number(),
prev: v.optional(v.number()),
}).index("by_key", ["key"]);Inside a Durable Object the input gate serializes the limiter's read-modify-write against storage, so both SQL- and ORM-backed stores are atomic against
concurrent calls without an explicit transaction. Driving createSqlStore outside a DO, you own serialization yourself.
Token budgets
A request limit is the wrong shape for an LLM: one call can cost a hundred
tokens or a hundred thousand, so "20 calls a minute" bounds nothing that
matters. tokenBudget binds one named limit whose unit is model tokens (or
cents) and spends it in arrears, which is how tokens actually work — the
cost is only known once the call returns.
import { tokenBudget } from "@lunora/ratelimit";
const budget = tokenBudget(limiter, "tokens");
const allowed = await budget.check(userId);
if (!allowed.ok) {
throw new LunoraError("RATE_LIMITED", `token budget exhausted; retry in ${String(allowed.retryAfter)}ms`);
}
const { text, usage } = await generateText({ model: ctx.ai.model(), prompt });
await budget.record(userId, usage?.totalTokens ?? 0);
return text;check runs before the call and consumes nothing: a tenant whose bucket is
already empty (or negative, from an expensive previous call) is turned away.
record runs after it — always, including when the call threw, since a
failed generation that consumed input tokens still has to be paid for. The
charge is a reservation, so the bucket may go negative: the tokens are already
spent, and the next call is refused until the debt refills away. A charge larger
than the limit's own ceiling is clamped to it rather than rejected, so the one
oversized call the budget exists to catch cannot escape it by being too big to
record.
Plugin form
ratelimitPlugin(limiter) packages the limiter as a Plugin that
exposes the resolved RateLimiter under ctx.api.ratelimit, so a handler can
call limit()/check()/reset() programmatically instead of (or alongside) the
enforcing middleware. Install its middleware with one .use(...):
import { RateLimiter, ratelimitPlugin } from "@lunora/ratelimit";
import { mutation } from "./_generated/server";
const limiter = new RateLimiter({
config: { send: { kind: "token bucket", period: 60_000, rate: 5, capacity: 5 } },
});
export const send = mutation.use(ratelimitPlugin(limiter).middleware!).mutation(async ({ ctx }) => {
const status = await ctx.api.ratelimit.limit("send", { key: ctx.auth.userId ?? "anonymous" });
if (!status.ok) {
throw new Error("slow down");
}
// …
});The plugin ships no schema extension (persistence is whatever store the limiter was built with), so it is middleware-only and skipped by the plugin schema fold.