MCP's elicitation feature lets a server pause a tool call to ask the human for more input — but the protocol splits that into two modes with very different trust models, one of which exists specifically to stop an account-takeover attack.
A tool's input schema is a contract written in advance. The server declares what it needs — a file path, a search query, a date range — and the client's model fills it in before the call ever happens. That works until the server discovers, mid-call, that it needs something it couldn't have asked for up front: a missing preference, a disambiguating choice between two matching records, or worse, a credential it has no business seeing in the tool arguments at all.
Elicitation is the Model Context Protocol's answer to that gap. It lets a server pause a request it's already processing and ask the human on the other end of the client for more information, then resume once it has an answer. The interesting part isn't that this exists — plenty of systems let a backend ask a frontend for more input. It's that MCP splits elicitation into two modes with deliberately different trust models, and that the second mode exists specifically to prevent a phishing attack the spec authors clearly spent real time thinking through.
Tool schemas are static. requestedSchema for a tool call is fixed at definition time, validated once, and done. That's fine for deterministic utilities — a checksum validator or a unit converter doesn't need to ask follow-up questions. It breaks down for anything that behaves like a real interactive workflow: a booking tool that needs to know which of three matching flights the user meant, a deployment tool that needs an explicit "yes, I understand this is production" before proceeding, or a server that needs an API key for a third-party service it doesn't already have on file.
Before elicitation, servers handled this by either guessing (bad), failing with an error that told the model to ask the human and retry the whole call with more arguments (clunky, and easy for the model to get wrong), or baking in long upfront forms that collected everything that might ever be needed (worse UX, and it still didn't cover the sensitive-data case). Elicitation turns that into a first-class protocol primitive: the server can stop, ask, and resume, within the same logical request.
Mechanically, a server that needs more information returns an InputRequiredResult instead of a final result, carrying one or more elicitation/create requests inside it. The client surfaces each request to the user, collects a response, and retries the original request with the answers attached as inputResponses. The server picks up where it left off — either it has what it needs and returns a real result, or it asks again.
This is a multi-round-trip pattern, not a side channel: the elicitation request is nested inside the response to the very call that triggered it. A client has to opt in explicitly — it declares an elicitation capability, naming which modes it supports, and a server must never send a mode the client hasn't declared:
{
"_meta": {
"io.modelcontextprotocol/clientCapabilities": {
"elicitation": { "form": {}, "url": {} }
}
}
}
An empty elicitation: {} object is backward-compatible shorthand for "form mode only," which tells you something about the order these two modes arrived in the spec — form mode first, URL mode as a later, more careful addition.
Form mode is the mode most people picture: the server describes a small JSON Schema, the client renders a form (or form-like prompt) from it, the user fills it in, and the data comes back as structured content.
{
"method": "elicitation/create",
"params": {
"mode": "form",
"message": "Please provide your contact information",
"requestedSchema": {
"type": "object",
"properties": {
"name": { "type": "string", "description": "Your full name" },
"email": { "type": "string", "format": "email" },
"age": { "type": "number", "minimum": 18 }
},
"required": ["name", "email"]
}
}
}
What's notable is what the schema is not allowed to express. requestedSchema is restricted to flat objects with primitive properties only — strings, numbers, booleans, and single- or multi-select enums, each with an optional default. No nested objects, no arrays of objects, no oneOf/allOf composition beyond the enum pattern. That's a real constraint on a format (JSON Schema) that normally supports arbitrary nesting, and it's intentional: a client has to be able to generate a sane input form from any schema a server throws at it, without running a general-purpose JSON Schema renderer. Capping the shape of the data keeps "render this form" a solvable problem instead of an open-ended one.
The spec is blunt about the one thing form mode must never be used for: "Servers MUST NOT use form mode elicitation to request sensitive information such as passwords, API keys, access tokens, or payment credentials." That's not a style suggestion — it's a MUST NOT, because form-mode content is, by design, visible to the MCP client and therefore to whatever is sitting in the request pipeline between the user and the server: logging middleware, the host application, potentially the LLM's own context window if the client surfaces elicitation content back into the conversation. A password that flows through that pipeline is a password that just became multi-party.
General profile data — a name, an email, a timezone — is fine in form mode, at the server's discretion and subject to the user being able to review and decline it. Secrets are categorically out. That line is exactly why a second mode exists.
URL mode, added in the 2025-11-25 revision of the spec, hands the user a URL instead of a form:
{
"method": "elicitation/create",
"params": {
"mode": "url",
"url": "https://mcp.example.com/ui/set_api_key",
"message": "Please provide your API key to continue."
}
}
The client's job shrinks to almost nothing: show the user where they're being sent, get explicit consent, open the URL in a secure, isolated browser surface, and wait. The actual interaction — typing in an API key, completing an OAuth consent screen, authorizing a payment — happens entirely outside the MCP client, on a page the server controls directly. When the user accepts, the client's response contains no content field at all:
{ "action": "accept" }
Accepting only means the user agreed to go open the link — not that whatever happens on the other end has finished. The original tool call gets retried later (possibly repeatedly), and the server uses its own stored state, correlated through an opaque requestState token, to decide whether the out-of-band step is done yet.
That diagram is the whole design in one picture: form mode's data passes through the client on its way to the server, while URL mode's sensitive data never enters the MCP connection at all — it goes straight from the user's browser to the server's own web page.
| Form mode | URL mode | |
|---|---|---|
| Where data goes | Through the MCP client, as structured content | Directly from browser to server; client never sees it |
| Allowed for secrets? | No — MUST NOT be used for passwords, keys, tokens, payment info | Yes — this is what it's for |
| Schema | Restricted flat JSON Schema (primitives, enums, defaults) | None — just a URL and a message |
| Client's role | Render a form, validate against schema, let user review before sending | Show the domain, get consent, open in an isolated webview, nothing more |
| Completion signal | Immediate — content comes back in the same response | Delayed — accept just means "user agreed to go"; actual completion is out of band |
| Added in spec | 2025-06-18 | 2025-11-25 |
Both modes share the same three-action response shape, which matters because "the user didn't fill out the form" and "the user actively refused" are different signals a server should handle differently:
accept — explicit approval. Form mode includes the submitted content; URL mode omits it (acceptance just means "I'll go look").decline — the user was asked and said no. A server should treat this as a real answer, not a glitch — maybe offer an alternative path instead of re-asking.cancel — the user dismissed the dialog without choosing either way (closed it, hit escape, the client failed to render it). This is ambiguous enough that retry-later is usually the right move.The spec's security section spends unusual effort on one specific scenario, because URL mode creates a structural opening for it: a server hands the client a URL, and nothing stops a malicious user from forwarding that URL to someone else.
The attack goes like this. Alice is a legitimate but malicious user of a benign MCP server. She triggers an elicitation that generates a third-party OAuth authorization URL — say, the server wants to connect to a calendar API on her behalf. Instead of opening the link herself, she sends it to Bob, a different user of the same server, and tricks him into clicking it. Bob completes the OAuth consent screen, believing he's authorizing his own connection. The third-party authorization server redirects back to the MCP server with a valid grant — but the server, if it isn't careful, has no way to know the person who finished the flow wasn't the person who started it. If it binds the resulting token to Alice's session anyway, Bob has just handed Alice access to his calendar. That's an account takeover, delivered entirely through a feature designed to look like routine consent.
The spec's fix is that the server — not the client, not the protocol itself — is responsible for verifying that the user who opens the elicitation URL is the same user for whom it was generated, typically by routing through an intermediate "connect" page that checks the browser's session cookie against the identity on file before forwarding to the real third-party authorization endpoint. The client's obligations are narrower but still concrete: never pre-fetch the URL, never open it without explicit consent, always show the full domain (so a lookalike subdomain stands out), and render it in something like SFSafariViewController rather than an embedded webview the host application or LLM could inspect.
It's easy to blur elicitation together with other MCP concepts it sits next to:
requestState tokens, say), that state has to be bound to a verified user identity — derived from the sub claim on an authorization token, never from a self-reported string in the elicitation response itself.If you're exposing tools through an MCP server, the elicitation question worth asking per tool is simple: will this tool ever need something from the user that isn't already in its input schema, and if so, is that something a password-shaped thing? If it's a preference, a disambiguation, or a confirmation, form mode's restricted schema is almost certainly expressive enough — resist the urge to reach for nested objects it doesn't support. If it's a credential or anything that authorizes a transaction, form mode is off the table by the spec's own rules, and you need a URL-mode "connect" page that checks the requesting user's session before it ever talks to the third party. Most purely deterministic utility servers — the kind that validate a checksum or convert a unit and return byte-exact output — will never need either mode at all, which is itself a useful signal that a server might be over-engineered if it's reaching for elicitation on every call.
Elicitation is a small protocol surface with an outsized amount of security reasoning packed into it, precisely because "ask the human for something" sounds harmless until the something is a credential and the human on the other end isn't the one who asked.