MCP Prompts: What They Are and How Servers Expose Them
MCP prompts are reusable message templates an MCP server exposes through prompts/list and prompts/get. Here is how they work, with the 2026-07-28 changes.
4 min read
MCP sampling lets a server ask the client's model for a completion. The 2026-07-28 MCP spec deprecates it, so here is how it worked and how to migrate.
MCP sampling is a feature that lets an MCP server ask the client's language model to generate a completion, so the server does not need an API key of its own. The Model Context Protocol specification deprecated sampling as of protocol version 2026-07-28. New implementations should not adopt it, and existing implementations should move to calling LLM provider APIs directly.
This post explains how the feature works, what the specification now asks of implementers, and how the Python SDK expresses it, so you can tell when to avoid it.

A server requests a generation while it handles a client request. The specification says servers send an InputRequiredResult containing a sampling/createMessage request, and the client returns its answer inside inputResponses on the retried request. A trimmed version of the spec's example request looks like this:
{
"method": "sampling/createMessage",
"params": {
"messages": [
{ "role": "user", "content": { "type": "text", "text": "What is the capital of France?" } }
],
"systemPrompt": "You are a helpful assistant.",
"maxTokens": 100
}
}The request can also set modelPreferences, which carries model hints and three priorities (costPriority, intelligencePriority and speedPriority), plus temperature and includeContext. The includeContext parameter defaults to "none". The values "thisServer" and "allServers" are deprecated under SEP-2596.
Clients must declare the capability on every request, inside _meta.io.modelcontextprotocol/clientCapabilities. A client that handles tool use declares sampling.tools as well, and servers must not send tool-enabled sampling requests to clients that lack it.
The specification also expects a person to stay in the loop. It says there should always be a human who can deny a sampling request, and that applications should let users review requests, view and edit prompts, and review responses before delivery.
Tools can run inside sampling. The request carries a tools array and an optional toolChoice. The server executes the tool calls the model returns, sends a new request with the results appended, and repeats. The specification suggests capping the number of rounds, and passing toolChoice: {mode: "none"} on the last round to force a final answer. A user message that carries tool results must contain only tool results.
The Python SDK's sampling page shows the feature through a resolver. The resolver returns a Sample object, and the tool receives the model's CreateMessageResult through Resolve:
from typing import Annotated
from mcp.server import MCPServer
from mcp.server.mcpserver import Resolve, Sample
from mcp.types import CreateMessageResult, SamplingMessage, TextContent
mcp = MCPServer("Bookshop")
def draft_blurb(title: str) -> Sample:
prompt = f"Write a one-sentence blurb for the book {title!r}."
return Sample(
[SamplingMessage(role="user", content=TextContent(type="text", text=prompt))],
max_tokens=60,
)
@mcp.tool()
async def blurb(title: str, draft: Annotated[CreateMessageResult, Resolve(draft_blurb)]) -> str:
"""Draft a blurb for a book."""
return draft.content.text if draft.content.type == "text" else "No blurb."If the connected client has not declared the sampling capability, the call fails with a -32021 protocol error. The SDK does not send a request the client cannot answer.
Do not use sampling in new servers. The specification says it remains fully functional, and that it stays in the specification for at least twelve months after the 2026-07-28 revision before it becomes eligible for removal. Existing code that works with clients that declare the capability can keep running for now.
The migration the specification names is to integrate directly with an LLM provider's API. That trade is visible in the design: sampling lets the client keep control over model access, selection and permissions, with no server API keys. Calling a provider from the server gives the server that control, and the server then holds the key.
The Python SDK page also marks roots as deprecated. Its suggested replacement is to pass directories through tool parameters, resource URIs or server configuration.
includeContext values "thisServer" and "allServers" are deprecated under SEP-2596, and they default to "none" when omitted.sampling.tools.MCP sampling is the mechanism by which an MCP server sends a sampling/createMessage request to the client, which generates a completion with its own model and returns the result. The client keeps control over which model runs and whether the request goes ahead.
Yes. The specification marks it deprecated as of protocol version 2026-07-28 under SEP-2577. It remains in the specification for at least twelve months after that revision's release, and existing implementations should migrate to LLM provider APIs.
The specification says the server needs no API key for sampling, because the client supplies the model. A server that migrates to a provider API calls that provider directly and therefore holds a key of its own.
For servers that are already published, browse the MCP servers hub on AI Agents Listing, which lists about 578 MCP servers as of 4 October 2026.
One email a week. New agents, MCP servers and skills, and what is actually getting traction.
MCP prompts are reusable message templates an MCP server exposes through prompts/list and prompts/get. Here is how they work, with the 2026-07-28 changes.
4 min read
A Copenhagen agency founder on logging MCP agent runs, splitting agent time from human review, and turning a run into an invoice line without watching anyone's screen.
5 min read
MCP elicitation lets a server pause a tool call to ask the user for missing input, in form mode or URL mode, before it continues. Here's how it works.
1 min read