Skip to content
DocspackagesDocumentation

@lunora/mcp

Model Context Protocol servers for Lunora — one exposing a deployment to AI agents, one exposing the framework's documentation.

PackagesMcp

@lunora/mcp ships two Model Context Protocol surfaces.

The deployment server (this package's main entry) exposes a deployed Lunora app to AI agents. It registers tools for introspecting a deployment (lunora_list_functions, lunora_list_tables, lunora_get_function_schema) and invoking its functions (lunora_run_query, lunora_run_mutation, lunora_run_action), each backed by @lunora/client over HTTP RPC. It needs an admin token.

The documentation server (@lunora/mcp/docs) exposes the framework's docs, so an agent writing Lunora code can look up the real API instead of inventing one. It reads published documentation only, with no credentials and no writes, which is why it can be hosted unauthenticated; Lunora runs it at https://lunora.sh/mcp.

Most users never install this package. lunora mcp install wires both servers into whichever editor they use. See AI coding agents.

pnpm add @lunora/mcp

The server is transport-agnostic. The shipped lunora-mcp binary speaks JSON-RPC over stdio (the transport MCP clients use when they spawn a process), or you can build a server with createLunoraMcpServer and connect any transport yourself.

The lunora-mcp binary

MCP clients spawn the lunora-mcp binary and talk to it over stdio. Configuration comes from the environment, so the spawn config stays a plain { command, env }:

  • LUNORA_URL (required): base URL of the deployed Worker.
  • LUNORA_ADMIN_TOKEN (required): the deployment's admin bearer, sent on every RPC. It cannot be scoped down — lunora_list_functions, lunora_list_tables and the allowlist precheck that runs before every lunora_run_* call all hit admin-gated /_lunora/admin/* routes, so a least-privilege token returns ADMIN_FORBIDDEN on the first tool call. The binary refuses to start without it rather than advertising tools that can only 403.
  • LUNORA_MCP_ALLOW_WRITES (optional): set to 1/true/yes/on to expose the mutation/action run tools. Default: read-only (writes disabled).
  • LUNORA_MCP_ALLOW_OBSERVABILITY (optional): set to 1/true/yes/on to expose the five lunora_get_* observability tools. Default: disabled. They are read-only, but what they return — production log lines, request metadata, grouped error messages — is user data that lands in the model's context and therefore at its provider, so holding the admin bearer (which every tool needs) is not by itself consent to ship it.
  • LUNORA_MCP_ALLOW_DATA_READS (optional): set to 1/true/yes/on to expose lunora_find_related. Default: disabled. It returns raw table rows read through the deployment's admin writer, so RLS policies and column masks do not apply to what it hands the model — and it returns everything reachable within depth hops of the start row, not just that row. Separate from the observability gate on purpose: log lines and table rows are different data classes, and enabling one must not silently enable the other.
  • LUNORA_MCP_ALLOW_AGENTS (optional): set to 1/true/yes/on to expose the agent tools. Default: agents disabled.
  • LUNORA_MCP_AGENTS (optional): a ;-separated list of name:description pairs selecting which agents to expose (see Expose an agent).
{
    "mcpServers": {
        "lunora": {
            "command": "lunora-mcp",
            "env": {
                "LUNORA_URL": "https://app.example.workers.dev",
                "LUNORA_ADMIN_TOKEN": "...",
            },
        },
    },
}

The binary exits non-zero if LUNORA_URL is missing or the transport fails to connect, so the spawning client surfaces a startup failure immediately.

Exposed tools

Each tool maps onto a method LunoraClient already provides. toolDefinitions(allowWrites, allowObservability, allowDataReads) returns the advertised tool surface for a given set of gates (the tiers are also exported on their own as READ_ONLY_TOOL_DEFINITIONS, OBSERVABILITY_TOOL_DEFINITIONS, ROW_READ_TOOL_DEFINITIONS and WRITE_TOOL_DEFINITIONS); callTool dispatches a call against a client.

ToolInputWhat it does
lunora_list_functionsnoneLists the deployment's public functions (queries, mutations, actions) with their kinds.
lunora_list_tablesnoneLists the deployment's .global() tables with their row counts.
lunora_get_function_schemafunctionPathReturns one function's argument descriptors and kind, so a caller can build valid args.
lunora_run_queryfunctionPath, args?, shardKey?Runs a query and returns its result. Read-only.
lunora_explain_errorcode?, message?Explains a Lunora error from the static catalog: status, title, hint, any matched solution, and a link to the reference. No deployment needed.
lunora_run_mutationfunctionPath, args?, shardKey?, confirmed?, actionDigest?, idempotencyKey?Proposes, then (on a second confirming call) runs a mutation. Writes data. Requires writes enabled.
lunora_run_actionfunctionPath, args?, shardKey?, confirmed?, actionDigest?, idempotencyKey?Proposes, then (on a second confirming call) runs an action. May call external services. Requires writes enabled.
agent_<name>prompt, threadKey?, title?Starts a durable @lunora/agent run and awaits its answer. One tool per exposed agent. Requires agents enabled.
lunora_agent_statusthreadKeyPolls a running agent by threadKey and returns its answer once finished. Requires agents enabled.
lunora_get_logslimit?, level?, shardKey?Recent ctx.log.* lines from the deployment. Requires observability enabled.
lunora_get_issueslimit?, status?, shardKey?Grouped error Issues with their fingerprints and messages. Requires observability enabled.
lunora_get_advisorieslimit?, shardKey?Advisor findings for the running deployment. Requires observability enabled.
lunora_get_query_insightslimit?, range?, shardKey?Slow/frequent query insights over a recent window. Requires observability enabled.
lunora_get_migration_statusshardKey?Applied and pending migrations for a shard. Requires observability enabled.
lunora_find_relatedtable, id, depth?, direction?, edges?, limit?, cursor?, shardKey?Follows the schema's foreign keys out of one row and returns what it is connected to. Reads through the ADMIN writer, so RLS policies and column masks do NOT apply. Requires data reads enabled.

The default surface is read-only and non-privileged: only the introspection tools, lunora_run_query and lunora_explain_error are exposed. Three separate gates hold the rest back, and each of them both hides the tools from the advertised list and refuses them at dispatch. lunora_run_mutation and lunora_run_action need LUNORA_MCP_ALLOW_WRITES (or allowWrites: true). The five lunora_get_* tools need LUNORA_MCP_ALLOW_OBSERVABILITY (or allowObservability: true) — they are read-only, but they return production log lines, request metadata and grouped error messages, all of which land in the model's context and therefore at its provider. lunora_find_related has its own gate, LUNORA_MCP_ALLOW_DATA_READS (or allowDataReads: true), deliberately NOT folded into the observability one: it returns raw table rows read through the deployment's admin writer with RLS policies and column masks bypassed, and enabling log reading for debugging must not silently also hand over every row. Those gates decide what the server may reach at all: the token must be the deployment's admin bearer (every tool reads admin-gated routes), so it cannot be scoped down to enforce anything from the credential side. Past the write gate, each individual write additionally needs its own confirmation. Guard who can reach the server process too.

The three run-tools share one base input schema:

  • functionPath (required, string): a function reference, e.g. "messages:send".
  • args (object): the arguments object passed to the function. Omitted (or null) means an empty bag, so an all-optional function runs with its defaults, and a JSON-stringified object (which models commonly emit) is parsed. Anything else — an array, a number, a boolean, a string that isn't a JSON object — is rejected with a BAD_REQUEST isError result naming the actual mistake, not coerced to {}.
  • shardKey (string): an optional shard key when the function is .shardBy()-partitioned.

The two write tools take three more fields — confirmed, actionDigest and idempotencyKey — which drive the confirmation handshake below.

Results are returned as MCP content text: the function's return value JSON, or the JSON null literal for a void-returning mutation/action. Unknown tools and thrown errors come back as isError tool results (not rejections), so the calling model sees the failure as tool output.

Write confirmation

allowWrites decides whether this server may write at all. It says nothing about whether a particular write was reviewed — and that is the gap that matters most for lunora_run_action, the tool that can send mail, charge a card, or call a third-party API. The destructiveHint annotation is a UI hint, not a gate.

So past that gate, lunora_run_mutation and lunora_run_action each take two calls. The first executes nothing and returns the proposal:

{
    "status": "action_required",
    "actionDigest": "1789129912052.0ZR2…",
    "expiresAt": "2026-09-11T14:41:52.052Z",
    "proposedAction": {
        "tool": "lunora_run_mutation",
        "kind": "mutation",
        "functionPath": "messages:send",
        "args": { "roomId": "r1", "text": "hi" },
    },
    "nextStep": "Show proposedAction to a human. To execute, call …",
}

The client renders proposedAction for a human, then calls the same tool again — before expiresAt — with the identical functionPath, args, shardKey and idempotencyKey, plus confirmed: true and that actionDigest. Only that second call writes.

The confirmation is bound to the exact action a human saw, for ten minutes. actionDigest is <expiresAt>.<signature>, the signature an HMAC over a canonical, sorted-key encoding of the tool name, function path, arguments, shard key, idempotency key and that deadline, keyed by the deployment's own identity — so re-serializing args in a different key order still verifies, while changing the target, any argument, or the shard key produces a different digest and the confirmation is refused with nothing written. A digest minted against another deployment never verifies, a lunora_run_mutation digest can never confirm a lunora_run_action, and an expired one is refused rather than silently re-proposed.

The digest carries its own proof rather than naming a stored record, because there is no store to name: createMcpFetchHandler serves statelessly — a fresh MCP server per HTTP request, no session, no cross-request state — and the confirming request need not even reach the instance that issued the proposal. Any instance holding the same deployment URL and admin bearer recomputes and verifies the same digest; nothing without that bearer can mint one. The deadline rides inside the digest for the same reason: there is nowhere else to keep it.

What the handshake does not do

It binds intent, not human presence.

A verified digest proves the call about to run is exactly the call that was proposed, on this deployment, inside its window. It does not prove a human saw it, and nothing server-side can: an MCP server has no channel to a person, and MCP puts the human-in-the-loop at the host — the client is what renders a tool call for approval. A client that asks nobody can send the digest it was just handed straight back with confirmed: true and the write runs.

That is why the write surface is off by default and refused at dispatch as well as omitted from ListTools. Setting allowWrites is the operator's statement that the client on the other end does the asking; the handshake is a client-UI affordance and an audit record of what was proposed, not a gate against the model.

Its scope is also deployment-wide, not principal-bound: the signing key is the deployment URL plus the admin bearer, with nothing identifying a user in it. On an OAuth-fronted server every principal shares that bearer, so inside the ten-minute window any principal holding write scope can confirm another's identical proposal. Binding a confirmation to a person would mean folding the verified sub claim into the key, which this package does not do today.

What idempotencyKey does and does not guarantee

idempotencyKey is an optional caller-chosen token folded into the digest.

It guarantees: a client that timed out can resubmit the confirmation it already holds, for as long as that digest is inside its window, instead of asking for a second human review — and a deliberately-repeated identical write sent under a new key gets its own digest, so it cannot ride the first review.

It does not guarantee deduplication. This server keeps no state between requests, so it cannot remember that a call already ran and cannot replay an earlier result; it also never forwards the key to your function, which never declared it as an argument. A resubmitted confirmed call therefore executes again. If the underlying write must happen at most once, make the function itself idempotent — for example by storing a caller-supplied key in a row with a unique index and short-circuiting on a repeat.

An agent discovers the surface before it calls anything:

  1. lunora_list_functions: discover the available paths and their kinds.
  2. lunora_get_function_schema: fetch the argument descriptors for one path.
  3. lunora_run_query / lunora_run_mutation / lunora_run_action: call the function with a well-formed args object. For the two write tools this is two calls — propose, have a human review proposedAction, then resubmit with confirmed: true and the actionDigest.

lunora_get_function_schema returns { path, kind, args }, where kind is "query", "mutation", or "action" and args is an array of argument descriptors (name, kind, optional, and optionally element or table). It returns an isError result if the path doesn't exist.

Building a server programmatically

createLunoraMcpServer returns a transport-agnostic MCP Server. Pass a url (and optional token), and the tools dispatch against a LunoraClient built from them:

import { createLunoraMcpServer } from "@lunora/mcp";

const server = createLunoraMcpServer({ url: "https://app.example.workers.dev", token: "..." });

await server.connect(myTransport);

For the common stdio case, connectStdio builds the server and connects it over a StdioServerTransport in one step:

import { connectStdio } from "@lunora/mcp";

await connectStdio({ url: process.env.LUNORA_URL, token: process.env.LUNORA_ADMIN_TOKEN });

LunoraMcpServerOptions accepts either a url with a token (plus an optional fetch) or a pre-built client; the latter is the injection point for tests. You must supply one or the other; passing neither — or a url with no token — throws.

Auth and admin gating

The token (sourced from LUNORA_ADMIN_TOKEN for the binary) is set as the client's auth token and sent as a bearer token on every RPC the tools make. It has to be the admin bearer: introspection (lunora_list_functions, lunora_list_tables) and the allowlist precheck in front of every run tool go through admin-gated /_lunora/admin/* routes.

So the read-only guarantee does not come from the token's scope. It comes from allowWrites (LUNORA_MCP_ALLOW_WRITES) defaulting off, which omits the write tools from tools/list and refuses them at dispatch. With writes on, the confirmation handshake is the second layer: a mutation or action executes only on a call carrying a digest issued for those exact arguments, so "this server may write" and "this write was reviewed" stay separate questions. Beyond that, gating is enforced by your deployment: the tools call ordinary queries, mutations, and actions, so whatever auth those functions require applies unchanged. Treat the MCP server process itself as the trust boundary and control who can reach it — createAuthedMcpFetchHandler is the supported way to expose it beyond a local stdio process. Run-tools open no WebSocket (every call is plain HTTP RPC), so the server is safe to run as a short-lived stdio process.

Serving MCP over OAuth

createMcpFetchHandler serves anyone who can reach the URL. That is the right trade for a stdio binary on your laptop, and the wrong one for a public endpoint: the tools carry the deployment's admin bearer, so the network path is the authorization.

createAuthedMcpFetchHandler mounts the same server behind an OAuth 2.1 gate — the flow MCP clients already know how to walk, discovered through the RFC 9728 protected-resource metadata that better-auth's mcp() plugin serves. The authorization server is a @lunora/auth instance running that plugin; the resource server is the handler:

// lunora/auth.ts
import { createAuth } from "@lunora/auth";
import { jwt, mcp } from "@lunora/auth/plugins";

export const auth = createAuth({
    database: env.DB,
    secret: env.AUTH_SECRET,
    plugins: [
        jwt(),
        mcp({
            consentPage: "/consent",
            loginPage: "/login",
            // The canonical protected-resource id (RFC 8707 / RFC 9728). Issued
            // tokens are audience-bound to it. HTTPS, no query or fragment.
            resource: "https://app.example.workers.dev/mcp",
        }),
    ],
});
// The MCP route.
import { requireMcpAuth } from "@lunora/auth/plugins";
import { createAuthedMcpFetchHandler, mcpTokenScopes } from "@lunora/mcp";

import { auth } from "./auth";

export const handleMcp = createAuthedMcpFetchHandler({
    protect: (handler) => requireMcpAuth(auth, handler, { requiredScopes: ["lunora:read"] }),
    server: (claims) => ({
        // Writes need a scope a read-only token does not carry.
        allowWrites: mcpTokenScopes(claims).has("lunora:write"),
        token: env.LUNORA_ADMIN_TOKEN,
        url: env.LUNORA_URL,
    }),
});

An unauthenticated request never reaches the MCP server at all — it gets a 401 carrying the WWW-Authenticate challenge that starts the client's authorization flow, and no LunoraClient holding the admin bearer is ever constructed for it.

Scoping tools to the token

server takes the verified token claims, not just a fixed options object. A gate that only answers yes/no gives every authorized agent the same capabilities, which throws away the scopes the token was issued with. Resolving allowWrites per request is what lets one endpoint serve a read-only agent and a read-write one: for the former the write tools are omitted from tools/list and refused at dispatch.

mcpTokenScopes(claims) parses the space-delimited scope claim (RFC 6749 §3.3) into a Set. A malformed or absent claim yields an empty set rather than throwing, so a scope check denies instead of surfacing a 500 a client might retry.

protect is a lambda rather than an auth instance because @lunora/mcp does not depend on better-auth — and the same seam takes either better-auth entry point. Use requireMcpAuth(auth, handler, opts) when the resource server shares a deployment with the authorization server, or createMcpProtectedRequestHandler(verifyOptions, handler) when it does not (a separate resource server holds verification config, not an auth instance).

Dynamic client registration is off unless you turn it on. better-auth's mcp() plugin expects you to compose it with @better-auth/cimd for the Client ID Metadata Document flow that MCP 2026-07-28 pins, or to enable the OAuth provider's registration options explicitly. @lunora/auth does not re-export cimd: its transport contract requires resolve-once DNS validation and connection pinning, which Cloudflare Workers cannot provide — so on Workers, register MCP clients explicitly instead.

Connecting an MCP client

Any MCP client that can spawn a stdio server works. With the mcpServers config above, an agent like Claude can call lunora_list_functions to discover the deployment's surface, then lunora_run_query / lunora_run_mutation / lunora_run_action to invoke it, passing functionPath (e.g. "messages:list"), an args object, and an optional shardKey. The two write tools answer the first call with an action_required proposal; the client shows it to the user and calls again with confirmed: true and the returned actionDigest — see write confirmation.

Resources and annotations

Both servers implement MCP tools. The documentation server additionally exposes every page as an MCP resource (lunora-docs:/docs/…, text/markdown), so a client can list and attach a page on the user's behalf, before the model knows what to search for.

Every tool carries annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint, title), so a client can badge the read-only surface and prompt before a write. These are hints for presentation; the guarantee itself is still enforced at dispatch — by the allowWrites / allowObservability gates, and for the two write tools by the confirmation handshake.

The documentation server

@lunora/mcp/docs is an independent surface with three tools:

ToolDescription
lunora_search_docsSearch the docs; returns matching pages and sections with their URLs.
lunora_get_docReturn one page in full, as Markdown.
lunora_list_docsList every page with its title and description.

It touches no user data and needs no token, so it is safe to expose publicly. Nothing in the subpath imports @lunora/client or a Node built-in, so it runs unchanged on Workers, Netlify/Vercel functions, Deno, and Bun.

The tools read a DocsIndex, and two implementations satisfy that contract. A docs site wires up its own in-process search index and mounts the server as a route:

import { createDocsMcpFetchHandler } from "@lunora/mcp/docs";

const handle = createDocsMcpFetchHandler({ index: myDocsIndex });

Anything else (the CLI's lunora mcp serve, a script) reads a published site over its /api/search, /llms.mdx/*, and /llms.txt endpoints:

import { createDocsMcpServer, createRemoteDocsIndex } from "@lunora/mcp/docs";

const server = createDocsMcpServer({ index: createRemoteDocsIndex({ baseUrl: "https://lunora.sh" }) });

Because both backends map results through the same helper, a model sees identical hits whichever one answered.

Hosting it safely

createDocsMcpFetchHandler screens each request before the transport sees it, because this surface is meant to be public and unauthenticated:

  • Bodies are capped: 128 KiB by default, overridable with maxRequestBytes.
  • JSON-RPC batches are refused. The stateless transport buffers a whole batch's replies into one response body, so a single small request carrying thousands of tools/call messages would amplify into hundreds of megabytes out, with no initialize and no session to rate-limit against. A documentation client gains nothing from batching.
  • lunora_search_docs bounds its query, and lunora_list_docs caps how many pages it serialises in one call.

Composing a local server

createLocalMcpServer / connectLocalStdio assemble the docs tools, the deployment tools, and any extra tools a host supplies into one stdio server. This is what lunora mcp serve runs, and it is why the CLI depends only on this package rather than on the protocol SDK.

import { connectLocalStdio } from "@lunora/mcp";

await connectLocalStdio({
    deployment: () => readMyDevServer(),
    docs: { baseUrl: "https://lunora.sh" },
    extraTools: myLocalTools,
});

deployment accepts a resolver, consulted on every tool call. An editor spawns its MCP servers when the project opens, usually before a dev server is running, and keeps them alive across every restart, so a URL captured once at startup would be absent for the whole first session and stale after the first restart. The deployment tools are advertised either way (clients cache the tool list); calling one with nothing running returns an actionable error.

The observability tools are the one exception. Their gate is a snapshot taken when the tool list is built, so a session that started before lunora dev was running does not advertise them — and because clients cache the list, they stay absent for that session even after the dev server comes up. Restart the MCP server (or the editor) once the dev server is running to get them.

Expose an agent

A deployment's durable @lunora/agent agents can be fronted as MCP tools. Like writes, this is opt-in and fail-closed: starting an agent run is a side effect, so the agent tools are hidden from the advertised list and refused at dispatch unless you enable them. @lunora/mcp takes no dependency on @lunora/agent; it calls the agent's public agents:agentRun mutation over HTTP RPC like any other function.

Two opt-ins are needed, and they sit on opposite sides of the RPC:

On the agent, mark it runnable from outside the deployment:

export const support = defineAgent({ name: "support", publicRun: true /* … */ });

agents:agentRun throws FORBIDDEN for any agent whose handle is not publicRun: true — starting a durable run is a side effect, so the agent's author allows it, not the MCP server. The env vars below cannot grant it: an agent exposed here without the flag is advertised in tools/list and fails on its first call.

On this server, the env vars (or the matching createLunoraMcpServer options):

  • LUNORA_MCP_ALLOW_AGENTS: 1/true/yes/on to expose the agent tools.
  • LUNORA_MCP_AGENTS: a ;-separated list of name:description pairs selecting which agents to expose, e.g. "support:Support questions;billing:Billing help".
  • LUNORA_MCP_AGENT_TIMEOUT_MS (optional): wall-clock budget a single agent_<name> call awaits before returning a pending result to poll.
{
    "mcpServers": {
        "lunora": {
            "command": "lunora-mcp",
            "env": {
                "LUNORA_URL": "https://app.example.workers.dev",
                "LUNORA_ADMIN_TOKEN": "...",
                "LUNORA_MCP_ALLOW_AGENTS": "1",
                "LUNORA_MCP_AGENTS": "support:Support questions;billing:Billing help",
            },
        },
    },
}

Each exposed agent gets an agent_<name> tool taking:

  • prompt (required, string): the task or message for the agent.
  • threadKey (string): reuse to continue a conversation; omit to start a new thread. The tool returns the threadKey it used either way.
  • title (string): an optional thread title, applied on the first run only.

The tool starts a durable run and awaits it up to the timeout budget. If the run outlasts the budget, the tool returns a pending result ({ status: "running", threadKey, runId, hint }) instead of hanging. Feed that threadKey to the generic lunora_agent_status tool to poll for the final answer once the run finishes.

Agent runs are owner-scoped to the identity the configured token resolves to — the deployment's admin identity, since that is the only token every tool works with. Every thread this server starts therefore shares one owner; run a separate MCP server per deployment to keep threads isolated. If the token resolves to no identity, threads are not owner-isolated at all.