Semantic Search and Deep Search: Two Retrieval Layers

Why Context Repo ships both a catalog-level semantic search and a hierarchical chunk navigator, and how to pick the right one for the question your AI agent is actually asking.

Context Repo Team

12 min read

Finding a saved prompt by topic and locating a relevant passage in a large PDF are both retrieval problems. They are not the same retrieval problem.

Context Repo is a context repository for AI agents and humans. Inside it, retrieval splits into two surfaces on purpose: one for finding a relevant artifact, one for finding candidate passages inside that artifact. Both semantic paths use OpenAI text-embedding-3-small and Convex vector search, but prompt embeddings and document chunks live in separate indexes. The tools answer different questions and return different shapes.

Rendering diagram

This article walks through both surfaces, when each one earns its place, and the implementation contracts that shape what each tool can do.

Many retrieval questions in a context repository fit one of two useful shapes:

  1. "Which artifact is relevant here?" This is catalog-level. The unit of return is a prompt, a document, or a collection. An agent uses it when it knows the topic but does not know which item in the workspace contains relevant material.
  2. "What does this artifact say about X?" This is content-level. The unit of return is a chunk inside one document. The agent uses it once it has identified the right document (or a small candidate set) and needs to inspect a specific passage.

find_items answers the first question. deep_search answers the second. They share an embedding model while using separate prompt and document-chunk indexes. They are separate tools with separate inputs because they perform separate jobs.

How does find_items search the catalog?

find_items takes a natural-language query and returns matches across prompts, documents, and collections. The MCP tool signature, simplified:

find_items({
  query: string,
  type?: 'prompts' | 'documents' | 'collections' | 'all', // default: 'all'
  semantic?: boolean,   // default: true
  tags?: string[],      // AND: every tag must match
})

The MCP tool has no limit or collectionId input. It leaves the REST limit at its default of 10, so a call can return up to 10 matches for each selected type.

In semantic mode, prompt matches come from the prompt-embedding index and document matches come from their best-matching section or paragraph chunks. Collection scores are derived from matching member prompts and documents. If that process produces no collection match, the implementation falls back to case-insensitive collection name and description matching.

Returns

The MCP response contains readable Markdown plus structured data grouped into prompts, documents, and collections. Rows use type-specific IDs (promptId, documentId, or collectionId) rather than a generic id. Semantic prompt and document rows include similarity scores; collection scores are derived from matching members or supplied by the name-and-description fallback. Literal rows omit scores.

Snippets are optional, are not query-term-bolded, and are not available for every type:

  • Semantic document matches include the leading portion of the best-matching chunk.
  • Literal prompt body matches and literal document body or preview matches can include a snippet in structured data.
  • Title-only, description-only, and collection matches may have no snippet. Collections have no highlight field.

The agent can use these item-level results to choose what to read next. Unlike deep_search, find_items can return prompts in semantic mode.

When to reach for it

Two specific reasons to use this surface:

  • The agent does not know the title yet. "Find the document about HIPAA compliance" can match a document called "Compliance Notes 2026" because the body mentions HIPAA. Semantic search bridges the title-to-content gap.
  • The agent wants to restrict the result type. Pass type: 'prompts' to retrieve only matching templates. Pass type: 'documents' to retrieve only matching reference material. Pass 'all' (the default) for all three catalog types.

Set semantic: false for case-insensitive literal matching. This path checks prompt titles, descriptions, and content; document titles, bounded previews, and indexed paragraph chunks; and collection names and descriptions. It avoids generating a query embedding. Full document-body coverage depends on the asynchronous chunking and indexing pipeline, so a newly written document may be searchable by title or preview before its deeper chunks are ready.

search_prompts is different: it lists a paginated prompt page first, then filters that page by literal title or description match. It does not search prompt content. Use find_items with semantic: false when literal prompt-body matching matters.

The matching REST surface is GET /v1/search. find_items forwards q, type, semantic, and tags, then returns that endpoint's grouped results.

How does deep_search read inside documents?

deep_search is what makes Context Repo a real retrieval system, not a search bar with a vector index bolted on. Signature:

deep_search({
  query: string,
  documentId?: string,    // search inside one document
  collectionId?: string,  // search inside one collection's documents
  limit?: number,         // default: 10, clamped to 1–50
  sessionId?: string,     // filter chunk IDs returned in earlier calls
})

Scope: documents only. Prompts use a separate single-embedding index and are not reachable through deep_search, deep_read, or deep_expand. Use find_items (or the generic MCP search alias) for semantic prompt discovery. search_prompts provides paginated literal filtering by prompt title or description.

Returns

Each match is chunk-level. The MCP result includes:

  • A current chunkId for follow-up calls. Reprocessing a document can replace its chunk tree, so an old ID can become stale.
  • A leading content preview of roughly 200 characters. Use deep_read for the full chunk.
  • A level whose contract allows "document", "section", or "paragraph". The current ingestion pipeline creates semantic embeddings for section and paragraph chunks; the document root is a navigation anchor.
  • A parentId (or null at the root).
  • siblingIds for the previous and next chunks under the same parent.
  • The owning documentId and documentTitle.
  • A similarity score.

Calling deep_read with a current chunkId returns the full chunk and richer metadata, including chunkIndex, sectionPath, parentChunkId, prevSiblingId, nextSiblingId, optional headingText, and wordCount. If a re-index makes the ID stale, run deep_search again to obtain current IDs.

The current tree is document root → section → paragraph. Every heading-defined section is a direct child of the document root; heading ancestry is retained in the section's sectionPath, not represented as nested section nodes. Content before the first heading and documents without headings can produce paragraphs directly under the root.

When to reach for it

This is a useful shape for retrieval-augmented generation (RAG) inside large documents. An agent can retrieve ranked chunk candidates, inspect promising chunks with deep_read, and navigate only where it needs more context instead of loading the full document by default.

sessionId records returned chunk IDs. Reusing the same live session filters those exact IDs from later searches; it does not deduplicate semantically equivalent text. REST callers mint a session with POST /v1/pd/session and pass it explicitly. When MCP callers omit sessionId, the server creates or reuses a caller-specific auto-session within the current warm server instance. A new instance can start a new auto-session.

How does deep_expand navigate document chunks?

A chunk on its own is often not enough. The matching paragraph might assume context from its parent section ("As mentioned in §3, the policy applies to..."). The agent needs to walk the tree.

deep_expand is the navigator:

deep_expand({
  chunkId: string,
  direction: 'up' | 'down' | 'next' | 'previous' | 'surrounding',
  count?: number,   // for 'surrounding' direction
})

The five directions:

  • up: read the parent chunk. Useful when the matching paragraph needs the surrounding section's setup.
  • down: read the child chunks. Useful when the match landed on a section heading and the agent needs the section body.
  • next and previous: read sibling chunks under the same parent. Useful for continuing a list or a sequential narrative.
  • surrounding: read a window of same-parent siblings around the target. When that window contains only the target, surrounding attempts a bounded walk through children of the parent's previous and next sibling sections, adding available context from those neighbours.

The agent picks the direction based on what it needs. up, next, and previous return zero or one chunk. down can return multiple direct children, and surrounding returns a target-centered window. Each call follows an explicit structural relation; callers can decide whether a larger parent chunk is necessary.

Why two retrieval surfaces, not one

Concretely, here is the failure mode we avoided by splitting them:

  • A single combined endpoint returns "document X scored 0.71, document Y scored 0.69, chunk inside X scored 0.66." The agent has to reason about whether to read X in full or just the matching chunk, and ranking does not really sort that out.
  • Or the endpoint returns only chunks, and the agent loses the catalog view it needs to choose between candidate documents.
  • Or the endpoint returns only documents, and the agent has to pull the entire candidate document just to confirm the match.

By keeping find_items catalog-level and deep_search chunk-level, an agent can run a two-step retrieval pattern that mirrors how humans search:

  1. Find a relevant artifact.
  2. Find a relevant passage within it.

This supports a common two-step RAG workflow and avoids loading a full candidate document unless the task calls for it.

What embedding model and vector store does Context Repo use?

We are pre-launch and have not published benchmark numbers; we will not invent them here. The shape that matters:

  • Vectors are 1536-dimensional, computed with OpenAI text-embedding-3-small. Both Convex vector indexes declare that dimension, so stored and query vectors must match it.
  • Prompts and document chunks use separate native indexes. Semantic prompt search uses one embedding per indexed prompt, built from its title, description, and content. Documents use searchable section and paragraph embeddings in the hierarchical-chunk index. Both indexes live in the same Convex backend as the source rows, so there is no separate vector database to synchronize.
  • Stored-content embeddings are generated asynchronously after ingest or relevant updates. Query embeddings are computed at search time.
  • Each low-level Convex vector-search call requests at most 256 candidates. Public result limits are lower: MCP find_items returns up to 10 matches per selected type, while deep_search defaults to 10 results and clamps the requested limit at 50.
  • The chunk vocabulary has three levels: document, section, and paragraph. The current tree places section nodes directly under the document root and paragraphs under sections or, for preamble and headingless content, directly under the root.

When to use each surface

QuestionSurface
Which document covers our HIPAA policy?find_items (catalog)
What does the HIPAA policy say about audit logs?deep_search (content) inside the doc found above
Show me the paragraph after that onedeep_expand with direction: 'next'
What is the section heading this paragraph lives under?deep_expand with direction: 'up'
Find prompts about code reviewfind_items with type: 'prompts'
Search within documents in one collectiondeep_search with collectionId
Search the literal phrase "error code E2049"find_items with semantic: false

How this connects to the rest of Context Repo

Both surfaces are exposed as MCP tools and as REST endpoints. The MCP server at contextrepo.com/mcp advertises both tools (plus deep_read and deep_expand for chunk navigation). The REST API at /v1 exposes:

  • GET /v1/search for catalog retrieval (matches the find_items tool).
  • POST /v1/pd/search for chunk-level retrieval (matches deep_search).
  • GET /v1/pd/read/{chunkId} for chunk inspection (matches deep_read).
  • POST /v1/pd/expand for hierarchy navigation (matches deep_expand).
  • POST /v1/pd/session to mint a sessionId for cross-call dedup.

The dashboard search bar at contextrepo.com/dashboard/search calls Convex actions and queries directly rather than routing through the public REST endpoints. Semantic mode delegates to the same internal semantic-search action used by REST. Keyword and progressive-disclosure paths also reuse the corresponding internal search logic. These are shared backend capabilities, not a promise that every client will display identical result sets.

The prompt and document-chunk vector indexes remain separate, but both live beside the source data in Convex. There is no second vector database to keep in sync with the row store.

Your questions, answered.

  • `find_items` searches the prompt, document, and collection catalog. It returns grouped item-level matches with type-specific IDs; scores appear in semantic mode, and snippets appear only when available. `deep_search` returns ranked document-chunk candidates with hierarchy metadata so an agent can inspect and navigate relevant passages.