Prompt and Document Management for AI Agents
How a context repository handles the day-to-day mechanics of prompts and documents that AI agents actually consume: version snapshots, stored prompt variables, semantic search, and 75+ file formats over MCP and REST.
Context Repo Team
12 min read
Prompt management and document management sound like two features on a checklist. In practice they are the spine of the product. If the primitives are wrong, nothing built on top of them works. The search ranks the wrong stuff. The agent retrieves stale templates. The version history cannot be trusted to roll back. This article walks through how a context repository handles both, in enough detail to ground a buying decision.
How does prompt versioning work in Context Repo?
A prompt in Context Repo is a piece of text with metadata. The text is the template body. Its current record also carries a title, description, tags, parameters, variable definitions, and a version number.
The template body can use ${variableName} as a readable placeholder convention. Context Repo stores the body as text and can store separate variable definitions; it does not execute the syntax or require placeholders to match those definitions. A prompt can look like this:
You are a senior engineer reviewing ${language} code for ${reviewType}.
The code under review:
${code}
Output requirements:
- Be specific about line numbers
- Cite ${language}-specific best practices
- Skip nitpicks unless ${reviewType} == "pedantic"Does Context Repo render prompt variables?
No. When an MCP client reads the prompt via read_prompt, it gets the template verbatim. Context Repo ignores request arguments for rendering and does not fill in ${language}, ${reviewType}, or ${code}. If you want substitution, your calling client or workflow must provide and replace those values before sending the resulting text to a model. Whether that happens automatically depends on the client; it is not part of the Context Repo contract.
Versions, change logs, and restore
A new prompt starts with an original snapshot numbered version 0. When its content changes, Context Repo writes the next version row. The REST layer can also version prompt parameter or variable updates, while title, description, and tag changes patch the prompt's current metadata without creating a version. Each snapshot contains:
- A monotonically incrementing version number.
- The full template body, parameters, and variable definitions (no diffs-only storage).
- The timestamp and author.
- An optional change log message the editor can attach.
The dashboard surfaces recorded content versions as a diff view. Over MCP, get_prompt_versions lists every recorded version with its timestamp, author name when available, change log, and a content preview capped at 200 characters. It does not return each full historical body. restore_prompt_version can still select one of those version IDs and copy the stored snapshot's full content, parameters, and variables into a new version. The prompt's current title, description, tags, visibility, and other unversioned metadata stay in place.
This matters because prompt engineering is iterative. You tune a system prompt across thirty edits to find the version that gets the model to behave. Version 17 was the best, then you broke it on version 24, and you cannot remember exactly what version 17 said. Without history this is a hostage situation. With history it is a restore_prompt_version call.
Content edits and restores append snapshots rather than rewriting earlier ones. Context Repo does not prune versions while the prompt remains, but permanently deleting a prompt also deletes its version rows.
How do you search across prompts?
Prompts are indexed two ways:
- Paginated filtering via
search_promptsfetches a prompt page, then filters that page by case-insensitive title or description match. - Semantic or literal catalog search via
find_itemsspans prompts, documents, and collections. Semantic mode compares your query with generated 1536-dimensional embeddings; literal mode can match prompt bodies as well as titles and descriptions.
Both are auth-scoped to the calling user. Context Repo generates query and stored-content embeddings with OpenAI's text-embedding-3-small, and both vector indexes are fixed at 1536 dimensions. Public MCP and REST search inputs do not accept caller-supplied vectors.
What file formats does Context Repo support for documents?
Documents are the larger artifacts. The upload pipeline accepts 75+ file formats through LlamaIndex Cloud: PDFs, Word, Excel, PowerPoint, Markdown, plain text, source code, and most image formats. The /v1/documents/scrape endpoint and the Chrome extension's URL-capture path use Firecrawl to extract rendered webpage content. The extension can also submit text extracted directly from a supported chat page or a selected page element.
A document, once ingested, has:
- Title and tags for cataloging.
- Full body text stored verbatim.
- Embeddings at the chunk level for semantic retrieval.
- A hierarchical chunk tree for
deep_searchanddeep_expand. - Canonical revision history covering title, content, HTML, source URL, metadata, and tags. Any changed canonical field creates one revision; an exact unchanged update is a no-op.
What does the ingest pipeline do?
When you upload a file or scrape a URL, the pipeline does three things in order:
- Extract text. PDFs lose their layout, Word docs lose their formatting, code files keep their syntax. The output is a normalized markdown-like representation.
- Chunk hierarchically. A document is not a flat string. We split each document into a tree of sections so an agent can read the top-level summary, then drill into a specific section, then drill into a paragraph, without ever loading 200 pages into the model.
- Embed searchable chunks. Embeddings live in Convex's native vector index. The dimension is 1536. Each chunk in the current tree has an ID, a parent ID (or null at the root), a position within its parent, and raw text. Reprocessing a document can replace the tree, so callers should not treat chunk IDs as permanent.
How does hierarchical retrieval actually feel?
Two retrieval surfaces, intentionally separate:
find_itemsis catalog-level search. "Find the document about HIPAA compliance." It returns grouped prompt, document, and collection matches with type-specific IDs. Semantic mode includes scores, while snippets appear only when available. Use this when you know roughly what you want and need to identify the artifact.deep_searchis content-level search inside one or all documents. Returns ranked, hierarchical chunks with their position in the tree. Use this when you know which document is relevant and you want a passage.
Once a chunk surfaces, deep_expand lets the agent navigate the document:
upreads the parent (the section the chunk belongs to).downreads the children (subsections of the chunk).nextandpreviousread sibling chunks under the same parent.surroundingreads a window of N chunks before and after.
The shape is similar to the way a human navigates a book's table of contents. The reason it matters for agents is context-window economics. Pulling a large document into a model's working memory can be expensive, slow, and distracting. Reading a few relevant chunks and expanding only the ones that need more context keeps retrieval bounded.
Visually, deep_search lands inside a document and deep_expand navigates from there:
Rendering diagram
How does Context Repo isolate data between users?
A collection is a named group of prompts and documents. The same prompt can live in multiple collections. The same document can live in multiple collections. Collections do two useful things, and both connect to per-user isolation.
- Scope search.
deep_searchaccepts acollectionIdparameter. Pass it, and the search runs against only the documents in that collection.find_itemssearches the whole workspace and narrows bytypeandtagsinstead. - Scope credentials by resource type. API keys use four scopes:
prompts.read,prompts.write,documents.read, anddocuments.write. Collections share the document scopes. Keys cannot be restricted to individual collection IDs, so collections and tags are organizational controls, not security boundaries.
Underneath every collection, prompt, and document, the same hard line holds: read paths scope records to the authenticated user and enforce access before returning data. Interactive session and OAuth tokens carry the signed-in user's authority. API keys resolve to that user but remain limited to their granted resource scopes. There is no collection-scoped API-key boundary: two keys with the same resource scope can see the same eligible resource type.
Can I export everything if I leave?
Yes. The same resource model is available through three doors, but complete bulk export is a REST workflow:
- The dashboard at contextrepo.com/dashboard is the UI for humans.
- The hosted MCP server at contextrepo.com/mcp exposes 29 tools over streamable-HTTP transport for AI clients. See How MCP Servers Connect AI Agents to Knowledge Bases for the protocol details.
- The REST API under contextrepo.com/v1 covers 35 logical operations for programmatic clients. OpenAPI 3.1 spec at /openapi.json.
Authentication supports two modes with different authority:
- OAuth 2.1 with PKCE (the modern OAuth flow that lets you log in safely without a shared secret in the URL). The deployment issues tokens via Clerk at
https://clerk.contextrepo.com. Authorization-server metadata lives at/.well-known/oauth-authorization-server(RFC 8414); the MCP and REST resources publish separate RFC 9728 metadata at/.well-known/oauth-protected-resource/mcpand/.well-known/oauth-protected-resource/v1. - Per-user API keys sent as
Authorization: API-Key gm_...headers. Generated from the dashboard with granular permission scopes per key.
Rate limits use a sliding window: 10 scrapes per minute, 100 mutating API calls per minute, and 120 read-only calls per minute. Responses may include X-RateLimit-* metadata. On 429, honor Retry-After.
Pagination uses opaque cursors. You echo back pagination.cursor from the previous REST response to get the next page; we do not expose internal offsets that could break across schema migrations. Default page size is 20, max 100. Continue while pagination.hasMore is true. Native MCP prompts/list instead returns nextCursor.
For a complete export, use the REST list endpoints and continue each opaque cursor through prompts, documents, collections, and GET /v1/collections/{id}/items. A Bearer token carries the signed-in user's read authority. An API key needs prompts.read and documents.read (or the corresponding write scopes, which imply read). MCP exposes the same core resource types, but get_collection includes at most 50 members, so it is not the complete collection-membership export path.
How to save a versioned prompt and retrieve it from any AI client
Five steps, a few seconds each.
- Sign in to Context Repo. Open contextrepo.com and start your 7-day free trial (Pro or Pro Max). The trial includes all 29 MCP tools, prompt and document storage, and Chrome extension capture.
- Create a prompt in the dashboard. Open the Prompts page, click New Prompt, paste a template with
${variable}placeholders, add a description, and save. The original snapshot is version 0. - Connect an MCP client. From the MCP Server page, use the one-click install for Cursor, the Claude Desktop instructions, or copy the manual JSON config. Use OAuth where the remote client supports it, or a scoped API key with the stdio bridge or another compatible client.
- Retrieve and use the prompt from the client. Use
search_promptswith a title or description term, orfind_itemsfor semantic or literal prompt-body search. Then callread_promptwith the returned ID. Context Repo returns the template verbatim. Supply any placeholder values explicitly in your client or workflow before model use. - Edit and review history. Edit the prompt content in the dashboard with an optional change-log message. The first content update creates version 1. Call
get_prompt_versionsto inspect version metadata and short content previews, and userestore_prompt_versionto copy an earlier stored snapshot into a new current version.
The same stored primitives are available across connected clients. The authority of each connection still depends on whether it uses an interactive token or a scoped API key.
Where this lands for your team
The mental model we settled on after building this:
- Prompts are the system prompts and templates you reuse across tools.
- Documents are the reference material you upload everywhere and lose track of.
- Collections are the project, client, or context boundary.
If your AI work is fragmented across three or more clients, a context repository pays for itself in the time you stop spending re-uploading PDFs and re-pasting prompts. If your AI work is concentrated in one tool with its native storage working fine, you do not need one yet, and the pricing page will tell you that before you commit.
Where to read next
- What Is an AI Context Repo for Agents?. The category framing for the whole product line.
- How MCP Servers Connect AI Agents to Knowledge Bases. The protocol layer and the transport and authentication support clients need to connect.
- Semantic Search and Deep Search: Two Retrieval Layers. How retrieval works once your AI is connected.
- Using Context Repo with Claude, Cursor, and ChatGPT. Concrete workflows in each client.
- API reference. Request and response shapes for every endpoint.
- MCP tools reference. Full reference for the 29 hosted tools.