Skip to content
DocsconceptsDocumentation

Vector search

Typed Vectorize indexes on ctx.vectors — automatic write sync, similarity search, and RAG with embeddings.

Last updated:

@lunora/bindings/vectors is the Cloudflare Vectorize adapter. You declare a typed vector index alongside a regular table, the adapter keeps it in sync on every write, and you run similarity search from any function handler via ctx.vectors. Together with @lunora/ai embeddings, that's everything you need for retrieval-augmented generation (RAG).

pnpm add @lunora/bindings

When your schema declares at least one vector index, codegen imports @lunora/bindings/vectors into the generated server and wires a typed ctx.vectors onto your query, mutation, and action contexts.

Declaring an index

A vector index always names a source of text to embed. The common case is to embed one string column: chain .vectorize off defineTable. The index is the logical name (it must match a vectorize binding in wrangler.jsonc); embed is your own embedder, and model names the model it calls.

// lunora/schema.ts
import { defineSchema, defineTable, v } from "lunorash/server";
import { embed } from "../app/embed"; // your own embedder

export default defineSchema({
    docs: defineTable({
        title: v.string(),
        body: v.string(),
        workspaceId: v.id("workspaces"),
    })
        .shardBy("workspaceId")
        .vectorize("body", {
            index: "docs-body",
            dimensions: 1024,
            metric: "cosine",
            embed,
            model: "@cf/baai/bge-m3", // the model `embed` calls; changing it re-embeds on the next backfill
            metadata: ["title", "workspaceId"], // mirrored into Vectorize metadata for filtering
        }),
});

When the embedded text is derived from multiple columns, use the standalone defineVectorIndex(...) form with a source.select(row) projection. See the package reference for the full shape.

Automatic write sync

You never upsert vectors by hand for declared indexes. On every insert, update, or delete to a table that sources an index, the adapter embeds the source and upserts under the row's id (or removes it on delete). A soft-deleted row (.softDelete()) has no vector: the soft delete removes it, a later patch or replace on the hidden row does not bring it back, and restore() re-embeds the row's current text.

Write sync only sees writes made after the index exists. Rows already in the table when you add .vectorize() (or change an index's field, dimensions, metric, metadata or model, or the table's .softDelete() field) are indexed by the backfillVectors admin operation; see Indexing existing rows.

A .global() table cannot declare a vector index. Its writes go to D1 or Hyperdrive, not through the shard write path the sync runs on, so the index would stay empty. defineSchema and lunora codegen both reject the combination.

The sync runs after the mutation's transaction commits, not inside it. Vectorize is outside the shard's storage and cannot be rolled back, so running it inline would leave a vector for a row a later throw took away. Two consequences worth knowing:

  • A mutation that rolls back syncs nothing. No vector is written for a row that never existed, and no vector is removed for a delete that did not happen.
  • A sync failure does not fail the mutation. The rows are already committed, so reporting the write as failed would push callers into retrying a non-idempotent insert. The failure is logged server-side and the row is left indexed in some indexes and not others. Upserts and deletes are keyed by row id, so re-running the write converges — but nothing tells the caller to, and nothing retries on its own. If a stale index is a problem for your app, re-run the write from a scheduled job rather than expecting an error.

Writes to the same row do not race each other: one mutation's sync completes before the next mutation's begins, so the last commit is the one the index ends up reflecting.

Indexing existing rows

Run the backfillVectors admin operation on each shard after deploying a new vector index over a table that already has rows. Target a shard with --shard:

lunora run '__lunora_admin__:backfillVectors' --args '{"maxPages":5}'
lunora run '__lunora_admin__:backfillVectors' --args '{"maxPages":5}' --shard workspace:acme

Each call embeds at most maxPages pages of 50 rows (one page when you omit it, at most 20). Every row is a remote embed plus a Vectorize write, so a large table cannot be embedded in one request. The result reports { done, pages, rows, failed, failedIds }. Progress is stored in the shard, so repeat the call until done is true.

Two kinds of failure are handled differently:

  • A row that can never be indexed is counted in failed, its id listed in failedIds (the first 20), and skipped. Examples: a non-string source value, text the embedding model refuses, metadata Vectorize rejects. The live write sync only logs these too. A row whose source text is empty or only whitespace is not a failure: it has nothing to embed, so its vector is removed, as when the field is cleared.
  • A page that fails as a whole, because the embedder or Vectorize is unreachable, answers 503 with the progress made so far. That page is retried by the next call. When the error carries a status, it decides: a 5xx, 408 or 429 (or a timeout) is the service and is retried for as long as it lasts, and a 4xx is the row. Without one, a page that fails whole on three calls in a row is written off: its rows are counted in failed and the walk moves on, so one bad page cannot stop every table after it.

Only one backfillVectors call runs per shard at a time. A second call made while one is running answers 409; repeat it once the first returns.

The backfill records the index configuration it built from. When you change an index's field, dimensions, metric, metadata or model, add or rename an index, or change the table's .softDelete() field, the next call re-embeds the table from the top. Existing vectors are overwritten in place, never cleared first. Nothing else is compared, so any other change needs restart (below).

embed is a function, and Lunora cannot tell which model it calls: its source text changes with every rebuild of your bundle, so it is not compared. The optional model string is what makes a model swap visible. Set it to the model embed calls and change the two together. model is only compared, never parsed, so any string works, and changing it on purpose re-embeds the table. Adding model to an index that had none re-embeds that table once, because the backfill has no record of which model built the stored vectors. An index without model keeps its recorded configuration and is not re-embedded.

Pass {"restart":true} once to re-embed everything when the check cannot see the change:

  • a different model behind embed with no model declared, or with model left unchanged;
  • a change to a defineVectorIndex select or metadata function;
  • rows written through the admin writer, such as a bulk import with lunora import. That path skips write sync, so those rows have no vectors until a backfill walks them.

A backfill never overwrites a newer write, from a mutation or an action. Each page is read inside the shard's single-writer gate and synced in commit order with live writes. A row patched or deleted while its page is embedding ends up with the patched vector, or none.

That ordering has a cost while a backfill runs. A page holds the shard's write-sync queue for as long as it takes to embed its rows (a few seconds with a typical remote embedder), and a mutation or action that commits meanwhile waits for its own sync behind it before it returns. Run large backfills with a small maxPages in quiet periods if that latency matters.

Tenant isolation. Vectorize indexes are account-global and shared by every shard. In a multi-tenant / sharded app you must scope writes and queries with a namespace (the shard/tenant key). Without it, one tenant's vectors are queryable by another.

ctx.vectors

The function context exposes the bridged search surface. The read half is available everywhere; the mutating half is gated by context kind:

MethodQueryCtxMutationCtx / ActionCtxUse
query(index, input)✓✓Similarity search.
getByIds(index, ids)✓✓Fetch stored vectors by id.
upsert(index, input)—✓Upsert one vector, after the mutation commits.
upsertNow(index, input)—✓Upsert one vector immediately.
deleteByIds(index, ids)—✓Remove vectors by id.

A query gets a read-only VectorSearchReader (search + fetch); mutations and actions get the full VectorSearch (also upsert + delete). This matches the reactivity model: a query may not write, so search lives naturally in a reactive query.

upsert follows the same rule as the automatic sync above — it is held until the mutation's transaction commits and dropped if the mutation throws, so a hand-written upsert cannot outlive the row it describes. upsertNow writes straight through; reach for it only when you need the vector visible before the mutation ends and you accept that a rollback will strand it. In an action there is no transaction to wait for, so the two behave identically.

deleteByIds is not deferred: it writes immediately, in both contexts.

query takes either a precomputed vector or an input plus an embed function. filter and namespace scope the search.

topK is capped at 100 for id/score-only queries, and at 50 once the query also asks for the heavy payload (returnValues: true or returnMetadata: "all"). Over either ceiling the call is refused locally with a message naming which one applied, rather than being silently capped at the far side.

// lunora/searchDocs.ts
import { query, v } from "@/lunora/_generated/server";
import { embed } from "../app/embed";

export const searchDocs = query.input({ q: v.string(), workspaceId: v.id("workspaces") }).query(async ({ ctx, args: { q, workspaceId } }) => {
    const { matches } = await ctx.vectors.query("docs-body", {
        input: q,
        embed,
        topK: 10,
        namespace: workspaceId,
        filter: { workspaceId },
    });

    return matches.map((m) => ({ id: m.id, score: m.score, ...m.metadata }));
});

Each match carries id, score, and (by default) the index's indexed metadata. You must pass either vector, or both input and embed.

RAG with @lunora/ai

The retrieval half of RAG is ctx.vectors.query; the embedding half is an external, non-deterministic call, so it lives on an action. Embed the user's question with @lunora/ai, search with ctx.vectors, then feed the retrieved context into a generation call:

// lunora/answer.ts
import { action, v } from "@/lunora/_generated/server";
import { embed, generateText } from "@lunora/ai";

export const answer = action.input({ q: v.string(), workspaceId: v.id("workspaces") }).action(async ({ ctx, args: { q, workspaceId } }) => {
    // 1. Embed the question.
    const { embedding } = await embed({
        model: ctx.ai.embeddingModel("@cf/baai/bge-base-en-v1.5"),
        value: q,
    });

    // 2. Retrieve the nearest chunks (scoped to the tenant).
    const { matches } = await ctx.vectors.query("docs-body", {
        vector: embedding,
        topK: 5,
        namespace: workspaceId,
    });

    const context = matches.map((m) => m.metadata?.title).join("\n");

    // 3. Generate an answer grounded in the retrieved context.
    const { text } = await generateText({
        // An `action` is public RPC and generation bills per output token, so
        // cap the completion; without it an anonymous caller can drive
        // arbitrarily long answers in a loop.
        maxOutputTokens: 500,
        model: ctx.ai.model("@cf/meta/llama-3.3-70b-instruct-fp8-fast"),
        prompt: `Answer using only this context:\n${context}\n\nQuestion: ${q}`,
    });

    return text;
});

To index content you embed yourself (rather than the schema's automatic sync), pass the precomputed vector to upsert via its embed thunk:

const { embedding } = await embed({ model: ctx.ai.embeddingModel("@cf/baai/bge-base-en-v1.5"), value: text });
await ctx.vectors.upsert("docs-body", { id, input: text, embed: async () => embedding });

Binding & config wiring

Each declared index needs a matching vectorize binding in wrangler.jsonc whose index_name equals the index name from your schema. @lunora/vite validates this, and a declared index with no binding fails the build.

// wrangler.jsonc
{
    "vectorize": [{ "binding": "DOCS_BODY", "index_name": "docs-body" }],
}

The generated createShardDO(config) factory then takes a vectors thunk mapping each index name to its binding, wired from the worker entry:

import { createShardDO } from "./_generated/shard";

export const ShardDO = createShardDO({
    vectors: (env) => ({ "docs-body": env.DOCS_BODY }),
});

Omit the thunk and ctx.vectors throws a descriptive "no vectors configured" error on first use.

Backing an index with your own Postgres

The thunk's values are typed as VectorizeIndexLike, not as the binding itself — so a Vectorize binding is one implementation, not the only one. An app already on Hyperdrive can serve the same index from pgvector and skip the binding entirely:

import { createHyperdrive, fromPostgresJs } from "@lunora/hyperdrive";
import { createPgVectorIndex } from "@lunora/hyperdrive/global";
import postgres from "postgres";

// Built once — `vectors` runs on every dispatch.
let docsBody: ReturnType<typeof createPgVectorIndex> | undefined;

export const ShardDO = createShardDO({
    vectors: (env) => {
        docsBody ??= createPgVectorIndex({
            client: fromPostgresJs(postgres(createHyperdrive(env.HYPERDRIVE).connectionString)),
            dimensions: 768,
            metric: "cosine",
            name: "docs_body",
        });

        // Map key mirrors the schema's index name; `name` is the SQL table
        // identifier, so a hyphenated index name cannot be reused verbatim.
        return { "docs-body": docsBody };
    },
});

The schema, the write-through sync, and ctx.vectors are unchanged — only where the vectors live moves. See Vector search on your own Postgres for the provisioning details and the three deliberate differences from Vectorize.