Prompt and Document Management for AI Agents

How a context repository handles the day-to-day mechanics of prompts and documents that AI agents actually consume: version snapshots, stored prompt variables, semantic search, and 75+ file formats over MCP and REST.

Context Repo Team

12 min read

Prompt management and document management sound like two features on a checklist. In practice they are the spine of the product. If the primitives are wrong, nothing built on top of them works. The search ranks the wrong stuff. The agent retrieves stale templates. The version history cannot be trusted to roll back. This article walks through how a context repository handles both, in enough detail to ground a buying decision.

How does prompt versioning work in Context Repo?

A prompt in Context Repo is a piece of text with metadata. The text is the template body. Its current record also carries a title, description, tags, parameters, variable definitions, and a version number.

The template body can use ${variableName} as a readable placeholder convention. Context Repo stores the body as text and can store separate variable definitions; it does not execute the syntax or require placeholders to match those definitions. A prompt can look like this:

You are a senior engineer reviewing ${language} code for ${reviewType}.

The code under review:
${code}

Output requirements:
- Be specific about line numbers
- Cite ${language}-specific best practices
- Skip nitpicks unless ${reviewType} == "pedantic"

Does Context Repo render prompt variables?

No. When an MCP client reads the prompt via read_prompt, it gets the template verbatim. Context Repo ignores request arguments for rendering and does not fill in ${language}, ${reviewType}, or ${code}. If you want substitution, your calling client or workflow must provide and replace those values before sending the resulting text to a model. Whether that happens automatically depends on the client; it is not part of the Context Repo contract.

Versions, change logs, and restore

A new prompt starts with an original snapshot numbered version 0. When its content changes, Context Repo writes the next version row. The REST layer can also version prompt parameter or variable updates, while title, description, and tag changes patch the prompt's current metadata without creating a version. Each snapshot contains:

  • A monotonically incrementing version number.
  • The full template body, parameters, and variable definitions (no diffs-only storage).
  • The timestamp and author.
  • An optional change log message the editor can attach.

The dashboard surfaces recorded content versions as a diff view. Over MCP, get_prompt_versions lists every recorded version with its timestamp, author name when available, change log, and a content preview capped at 200 characters. It does not return each full historical body. restore_prompt_version can still select one of those version IDs and copy the stored snapshot's full content, parameters, and variables into a new version. The prompt's current title, description, tags, visibility, and other unversioned metadata stay in place.

This matters because prompt engineering is iterative. You tune a system prompt across thirty edits to find the version that gets the model to behave. Version 17 was the best, then you broke it on version 24, and you cannot remember exactly what version 17 said. Without history this is a hostage situation. With history it is a restore_prompt_version call.

Content edits and restores append snapshots rather than rewriting earlier ones. Context Repo does not prune versions while the prompt remains, but permanently deleting a prompt also deletes its version rows.

How do you search across prompts?

Prompts are indexed two ways:

  • Paginated filtering via search_prompts fetches a prompt page, then filters that page by case-insensitive title or description match.
  • Semantic or literal catalog search via find_items spans prompts, documents, and collections. Semantic mode compares your query with generated 1536-dimensional embeddings; literal mode can match prompt bodies as well as titles and descriptions.

Both are auth-scoped to the calling user. Context Repo generates query and stored-content embeddings with OpenAI's text-embedding-3-small, and both vector indexes are fixed at 1536 dimensions. Public MCP and REST search inputs do not accept caller-supplied vectors.

What file formats does Context Repo support for documents?

Documents are the larger artifacts. The upload pipeline accepts 75+ file formats through LlamaIndex Cloud: PDFs, Word, Excel, PowerPoint, Markdown, plain text, source code, and most image formats. The /v1/documents/scrape endpoint and the Chrome extension's URL-capture path use Firecrawl to extract rendered webpage content. The extension can also submit text extracted directly from a supported chat page or a selected page element.

A document, once ingested, has:

  • Title and tags for cataloging.
  • Full body text stored verbatim.
  • Embeddings at the chunk level for semantic retrieval.
  • A hierarchical chunk tree for deep_search and deep_expand.
  • Canonical revision history covering title, content, HTML, source URL, metadata, and tags. Any changed canonical field creates one revision; an exact unchanged update is a no-op.

What does the ingest pipeline do?

When you upload a file or scrape a URL, the pipeline does three things in order:

  1. Extract text. PDFs lose their layout, Word docs lose their formatting, code files keep their syntax. The output is a normalized markdown-like representation.
  2. Chunk hierarchically. A document is not a flat string. We split each document into a tree of sections so an agent can read the top-level summary, then drill into a specific section, then drill into a paragraph, without ever loading 200 pages into the model.
  3. Embed searchable chunks. Embeddings live in Convex's native vector index. The dimension is 1536. Each chunk in the current tree has an ID, a parent ID (or null at the root), a position within its parent, and raw text. Reprocessing a document can replace the tree, so callers should not treat chunk IDs as permanent.

How does hierarchical retrieval actually feel?

Two retrieval surfaces, intentionally separate:

  • find_items is catalog-level search. "Find the document about HIPAA compliance." It returns grouped prompt, document, and collection matches with type-specific IDs. Semantic mode includes scores, while snippets appear only when available. Use this when you know roughly what you want and need to identify the artifact.
  • deep_search is content-level search inside one or all documents. Returns ranked, hierarchical chunks with their position in the tree. Use this when you know which document is relevant and you want a passage.

Once a chunk surfaces, deep_expand lets the agent navigate the document:

  • up reads the parent (the section the chunk belongs to).
  • down reads the children (subsections of the chunk).
  • next and previous read sibling chunks under the same parent.
  • surrounding reads a window of N chunks before and after.

The shape is similar to the way a human navigates a book's table of contents. The reason it matters for agents is context-window economics. Pulling a large document into a model's working memory can be expensive, slow, and distracting. Reading a few relevant chunks and expanding only the ones that need more context keeps retrieval bounded.

Visually, deep_search lands inside a document and deep_expand navigates from there:

Rendering diagram

How does Context Repo isolate data between users?

A collection is a named group of prompts and documents. The same prompt can live in multiple collections. The same document can live in multiple collections. Collections do two useful things, and both connect to per-user isolation.

  • Scope search. deep_search accepts a collectionId parameter. Pass it, and the search runs against only the documents in that collection. find_items searches the whole workspace and narrows by type and tags instead.
  • Scope credentials by resource type. API keys use four scopes: prompts.read, prompts.write, documents.read, and documents.write. Collections share the document scopes. Keys cannot be restricted to individual collection IDs, so collections and tags are organizational controls, not security boundaries.

Underneath every collection, prompt, and document, the same hard line holds: read paths scope records to the authenticated user and enforce access before returning data. Interactive session and OAuth tokens carry the signed-in user's authority. API keys resolve to that user but remain limited to their granted resource scopes. There is no collection-scoped API-key boundary: two keys with the same resource scope can see the same eligible resource type.

Can I export everything if I leave?

Yes. The same resource model is available through three doors, but complete bulk export is a REST workflow:

Authentication supports two modes with different authority:

Rate limits use a sliding window: 10 scrapes per minute, 100 mutating API calls per minute, and 120 read-only calls per minute. Responses may include X-RateLimit-* metadata. On 429, honor Retry-After.

Pagination uses opaque cursors. You echo back pagination.cursor from the previous REST response to get the next page; we do not expose internal offsets that could break across schema migrations. Default page size is 20, max 100. Continue while pagination.hasMore is true. Native MCP prompts/list instead returns nextCursor.

For a complete export, use the REST list endpoints and continue each opaque cursor through prompts, documents, collections, and GET /v1/collections/{id}/items. A Bearer token carries the signed-in user's read authority. An API key needs prompts.read and documents.read (or the corresponding write scopes, which imply read). MCP exposes the same core resource types, but get_collection includes at most 50 members, so it is not the complete collection-membership export path.

How to save a versioned prompt and retrieve it from any AI client

Five steps, a few seconds each.

  1. Sign in to Context Repo. Open contextrepo.com and start your 7-day free trial (Pro or Pro Max). The trial includes all 29 MCP tools, prompt and document storage, and Chrome extension capture.
  2. Create a prompt in the dashboard. Open the Prompts page, click New Prompt, paste a template with ${variable} placeholders, add a description, and save. The original snapshot is version 0.
  3. Connect an MCP client. From the MCP Server page, use the one-click install for Cursor, the Claude Desktop instructions, or copy the manual JSON config. Use OAuth where the remote client supports it, or a scoped API key with the stdio bridge or another compatible client.
  4. Retrieve and use the prompt from the client. Use search_prompts with a title or description term, or find_items for semantic or literal prompt-body search. Then call read_prompt with the returned ID. Context Repo returns the template verbatim. Supply any placeholder values explicitly in your client or workflow before model use.
  5. Edit and review history. Edit the prompt content in the dashboard with an optional change-log message. The first content update creates version 1. Call get_prompt_versions to inspect version metadata and short content previews, and use restore_prompt_version to copy an earlier stored snapshot into a new current version.

The same stored primitives are available across connected clients. The authority of each connection still depends on whether it uses an interactive token or a scoped API key.

Where this lands for your team

The mental model we settled on after building this:

  • Prompts are the system prompts and templates you reuse across tools.
  • Documents are the reference material you upload everywhere and lose track of.
  • Collections are the project, client, or context boundary.

If your AI work is fragmented across three or more clients, a context repository pays for itself in the time you stop spending re-uploading PDFs and re-pasting prompts. If your AI work is concentrated in one tool with its native storage working fine, you do not need one yet, and the pricing page will tell you that before you commit.

Your questions, answered.

  • A prompt starts at version 0. A content update creates the next version with an optional change-log entry, while title, description, and tag edits patch metadata in place. `get_prompt_versions` lists recorded versions with change logs and content previews of up to 200 characters; `restore_prompt_version` copies a selected snapshot's full content, parameters, and variables into a new current version without replacing current metadata.