Bezpečnostní a architektonický audit · září 2026Security and architecture audit · September 2026

Co jsme našli, co jsme opravili a co žádné nastavení neodstraní.What we found, what we fixed, and what no configuration removes.

Interní audit s pomocí AI: šestnáct nezávislých revizí po oblastech (kryptografie, OAuth, poštovní vrstva, prompt injection, oprávnění nástrojů, HTTP, provoz, ochrana dat, závislosti, testy, tvrzení na webu) a druhý, nezávislý průchod jiným modelem. Není to penetrační test třetí strany. Všechny vážné nálezy byly opraveny před zveřejněním (verze 0.4.1); přijaté kompromisy jsou popsané i s důvodem.An internal, AI-assisted audit: sixteen independent reviews by area (cryptography, OAuth, the mail layer, prompt injection, tool authorization, HTTP, operations, data protection, dependencies, tests, public claims) and a second, independent pass by a different model family. It is not a third-party penetration test. Every serious finding was fixed before publication (version 0.4.1); the accepted trade-offs are listed with their reasons.

Shrnutí česky

Hlavní závěr: hesla ke schránkám se šifrují ve vašem prohlížeči a server je má rozbalená jen v paměti, když vyřizuje váš dotaz (nejdéle 15 minut po posledním). OpenAI ani Anthropic je nikdy nedostanou, vidí ale výsledky každého nástroje. Token je pro ně nečitelný, přesto je to živý klíč k vaší poště, chraňte ho jako heslo.

Opraveno před zveřejněním: nástroj pro odeslání konceptu šel zneužít k trvalému smazání libovolné zprávy; názvy příloh a hlavičky mohly podvrhnout rámování obsahu pro asistenta; detekce skrytého textu měla mezery; asistent četl textovou variantu zprávy místo té, kterou vidíte vy; skryté kopie unikaly příjemcům; přílohy se tiše ořezávaly; token šel nasměrovat na vnitřní síť serveru a další.

Zůstává záměrně: přístupový token platí 30 dní a nejde odvolat jednotlivě (okamžité odvolání je smazání hesla pro aplikace u poskytovatele); limity pokusů platí na instanci; provozovatel serveru technicky vidí vaše nastavení během dotazu. To poslední je hlavní důvod, proč firmě doporučujeme vlastní nasazení (část 7): vlastní klíč, vlastní vypínač, vlastní hosting a region, žádná závislost na provozovateli.

Celý dokument je anglicky níže. Zdrojová verze: GitHub.

Version audited: 0.4.0 (2026-09-05). Fixes shipped as 0.4.1 and 0.4.2 (2026-09-06). Published at https://mailmcp.ai/audit.

Sections 13 to 18 are addenda written for releases 0.8.0 (Outlook and Microsoft 365 over Microsoft Graph), 0.10.0 (triage and bulk clean-up), 0.11.0 (follow-ups, snooze, templates, confirmed sends and unsubscribe), 0.12.0 (the clickable inbox), 0.13.0 (protection levels) and 0.9.1 (usage statistics and MCP request counting; never released on its own, shipped with 0.13.0). They were not part of the reviews described in section 1; they document the new components to the same standard so the tables in sections 3 and 9 stay true.

1. What this is, and what it is not

This is an internal, AI-assisted audit commissioned by the author of mailmcp. Sixteen independent review agents (Claude Fable 5.1 for cryptography, OAuth, the mail layer and the threat model; Claude Opus 5 and Claude Sonnet 5 for the remaining areas) each examined one part of the system read-only, with file and line evidence, and several of them reproduced their findings by executing the real code against a local mail server. A second, independent review by a different model family (OpenAI Codex with GPT-6 Astra) followed the same brief on the code after the first round of fixes; its findings, and what they changed, are in section 12.

It is not a third-party penetration test and nobody outside the project signed it. It is published so that you can read what was found, what was fixed, what was deliberately left as is, and what no configuration can remove. Every High finding from both reviews was fixed before publication (releases 0.4.1 and 0.4.2, all with regression tests); the Medium findings were fixed or are listed as accepted trade-offs with the reason. Fixes are claimed against commits in the repository, and section 12 records one case where the second reviewer caught a fix that had been described but not actually shipped.

2. The system in one page

mailmcp connects an AI assistant (ChatGPT, Claude, Cursor, VS Code, Gemini CLI) to your mailboxes over IMAP and SMTP through the Model Context Protocol. Since 0.8.0 Outlook.com and Microsoft 365 mailboxes take a second path instead: a Microsoft sign-in and Microsoft Graph over HTTPS, with no password anywhere (section 13). There are three ways to run it.

  • Shared server (mailmcp.ai or any server run by someone else). The server holds one master key. You create a token in your browser: the browser encrypts your mailbox settings with a random key K (AES-256-GCM), the server seals K under its master key and hands it back, and the token is mmt1.<sealed K>.<encrypted settings>. Whoever holds only the token cannot decrypt it. Whoever holds only the master key has no data. The server combines the two while it serves your request, decrypts your settings in memory, opens IMAP/SMTP to your provider, and keeps nothing on disk. Decrypted settings stay in memory for at most 15 minutes after your last request.
  • Your own deployment (Vercel one-click, Docker, a VPS). The same code, your own master key. Your organization is the operator.
  • Local (Claude Desktop extension, Claude Code). No server, no OAuth, no links; settings live on your machine.

Everything else derives from the master key: OAuth codes and tokens, dynamically registered client ids, one-hour attachment download links, one-hour upload links. The server is stateless: no database, no session store, no revocation list. Capabilities (read, draft, send, modify, trash, and since 0.11 unsubscribe and host-only sending) are stored inside the encrypted settings per mailbox and enforced on every call; tools that no mailbox permits are not offered to the assistant at all. Sending is off by default and, when on, limited to an allowlist of addresses or domains.

3. Who sees what

Parties and the three ways of running mailmcp. A = shared server run by the vendor, B = your own deployment, C = local.

PartyA: vendor shared serverB: your own deploymentC: local
YouEverything about your mailboxes. You hold the token (a live credential) and the edit password.Same.Plaintext settings on your machine.
Other users of the same serverNothing of yours. Runtimes are keyed by a hash of the token; no code path shares them. They compete only for capacity (throttles, cache slots).Same; colleagues. Registration can be closed with an invite code.Not applicable.
AI vendor (OpenAI, Anthropic, …)Every tool result: subjects, addresses, message text, names of attachments, text attachments, download links. Holds your token or an encrypted OAuth token, neither of which it can decrypt. Never sees passwords.Identical. Your own server changes nothing about what the model sees.Same, and binary attachments as base64 when asked.
Server operatorThe vendor. Holds the master key; during your requests has your decrypted settings, passwords and mail in memory. Could change the code at any time.Your organization, under your change control.You.
Microsoft (Outlook mailboxes only)The mailmcp application signing in to your mailbox and every Graph call it makes. It is your mail provider, so it sees the mail anyway; what is new is that it can see and revoke this application's access per account.Identical, under the vendor's app registration unless the operator sets MAILMCP_MS_CLIENT_ID.Identical (the sign-in itself happens on mailmcp.ai).
Hosting provider (Vercel)Environment variables (master key), TLS termination, function memory, request logs with IPs and URLs (including one-hour link references), region.Same, under your account, region and audit log; none if you run Docker on your own hardware.None.
Someone who steals your token or OAuth tokenReads and, if enabled, sends mail at the token's capabilities until you delete the app password at your provider or the operator rotates the master key. For an Outlook mailbox there is no password to delete: the credential inside the token is a Microsoft refresh token, so you revoke the application at Microsoft (account.live.com/consent/Manage, or My Apps for work accounts). Changing the mailbox password does not help. Outside mailmcp that refresh token carries the consented Microsoft scopes, which on a mailbox with drafts, sending or labels enabled are Mail.ReadWrite and Mail.Send. Access tokens live 30 days, refresh tokens 90 days; nothing revokes them individually.Same, and your operator can rotate the master key as a kill switch.Not applicable (a leaked local config is plaintext passwords).
Someone who gets a download or upload linkDownloads that one attachment, or uploads files into your mailmcp-uploads folder, for one hour.Same.Not applicable.
A malicious senderGets text in front of the model. Hidden text is stripped, headers and file names are neutralized, bodies are marked as third-party data; the model may still follow a visible instruction, within the allowlist and rate limit, with no permanent deletion.Identical.Identical.
A newsletter sender's server (0.11, only when the owner asks to unsubscribe and enabled it)The IP address of the server, the time, and whatever identifies the owner in the one-click link itself.Same, with your server's address.Same, with your machine's address.
The authorIs the operator of A.Ships the prebuilt bundle; you decide when to update; the licence check is offline and there is no telemetry beyond the daily usage-statistics ping described in section 18 (server copies on by default, desktop copies opt-in; opt-out MAILMCP_NO_STATS=1, no request at all with MAILMCP_NO_UPDATE_CHECK=1).Same as B.

What is where:

AssetABC
Mailbox app passwordsServer memory, at most 15 min after the last use; never at restSame, on your serverAt rest on your machine
Microsoft refresh token (Outlook mailboxes)Encrypted in the browser into your token like a password; in server memory while serving you, at most 15 min after the last request; relayed once from Microsoft through the server to your browser at sign-in; never at restSame, on your serverAt rest on your machine (the sign-in itself is done on mailmcp.ai)
Key K of your settingsSent once over TLS at token creation, then only sealed inside the tokenSameAt rest as MAILMCP_KEY
Master keyVendor's Vercel environmentYour environmentSame value as K
Your tokenYour AI client's config or keychain; inside encrypted OAuth tokens and linksSameNot applicable
Mail content and addressesAI vendor, server memory, hosting provider memory and logsAI vendor, your serverAI vendor, your machine
Attachments (binary)Browser to server directly; the AI vendor sees a link unless you ask for the contentsSameBase64 through the AI vendor
Edit passwordOnly a PBKDF2 hash (600 000 iterations, browser-side), sealed in the tokenSameNot applicable
Licence keySigned and readable; carries the buyer's name and an e-mail fingerprint (not the address)Your environmentLocal
Follow-up, snooze, bulk-journal, send-claim and draft records (0.11)Signed messages in the folder mailmcp-state of your own mailbox; never on the serverSameSame
Snoozed mail and templates (0.11)Your own mailbox: mailmcp-snoozed, mailmcp-templatesSameSame
State secret (policy.state_secret, 0.11)Inside your encrypted settings, sealed in the tokenSameIn the local configuration
Clickable inbox (0.12)The page is part of the server code; its data (the draft card, account labels and capabilities) travels in tool results to the AI vendor's app like any other result; the app's widget state (ChatGPT) may keep which drafts were sent from a card (account, id, fingerprint); nothing on the serverSameSame

4. Findings

Sixteen reviews produced 6 High, about 30 Medium and about 40 Low or informational findings, with overlaps; the second-model review added 5 High and 15 Medium, again with overlaps. All High and most Medium findings were fixed in 0.4.1 and 0.4.2 before this document was published. The remaining Medium findings are accepted trade-offs and are listed with the reason.

4.1 Fixed in 0.4.1 and 0.4.2

AreaSeverityFindingFix
ToolsHighsend_draft accepted any folder and any message, sent it verbatim and then permanently removed it, with only the send capability. A prompt-injected assistant could re-send and destroy any message. Reproduced against a local mail server.Only messages flagged as drafts in the Drafts folder qualify; a different folder is refused; read and send capabilities required.
ToolsHighpolicy.attachments = "metadata" only hid one tool; download links were still minted and mailbox attachments could still be re-sent.The setting now disables download links, attachment re-sending and forwarding of attachments.
ToolsHighforward_message read message bodies and attachments from mailboxes marked as not readable (read: false, send: true).Read capability required. list_uploads gained a capability check too; forward_message is available to draft-only configurations as well.
ToolsHighConsuming a staged upload deleted the source message even from a mailbox the token could only read, and any message moved into the staging folder qualified.Uploads are consumed only from mailboxes the token may write to, and only messages mailmcp staged itself (marked with its own header) are ever removed.
ToolsMediummodify_message could move a message to Trash with the modify capability alone, and could move mail into the staging folder.A move to Trash (by folder or Gmail label) requires the delete capability; the staging folder is reserved.
ToolsMediumThe hourly send limit lived inside the cached runtime, so eviction after idle time or cache pressure reset it within the hour.Limiters are kept per token outside the runtime cache.
ToolsMediumsend_draft sent stored drafts regardless of the attachment policy and without a size bound.Drafts with attachments are refused under a metadata-only policy; drafts are bounded like every other transfer and never sent incomplete.
Prompt injectionHighAttachment file names were inserted into the assistant's text outside the untrusted wrapper and could forge its delimiters; subjects and sender names could do the same on one line. Reproduced.One sanitizer for every sender-controlled field: control and invisible characters removed, wrapper delimiters neutralized, length bounded. Applied to names, subjects, senders, recipients, forwarded headers, upload listings.
Prompt injectionHighHidden-text stripping missed <style> class rules, entity-encoded styles, CSS comments, self-closing hidden elements, fonts of 2 px, near-white colours, zero-size boxes. Reproduced.Style-block rules are applied by class and id; styles are entity-decoded and comment-stripped before matching; thresholds widened; self-closing non-void elements treated as open; presentational zero sizes honoured.
Prompt injectionHighThe plain-text part of a message was preferred over HTML, so a sender could show the human one message and the assistant another.The HTML part is read when present (the same one the user sees); plain text is the fallback.
CryptographyMediumA sealed key was not bound to its encrypted settings: an old seal (no expiry, old edit password) could be paired with a newer settings blob of the same key. Reproduced.The seal carries a hash of the settings blob and is refused with any other blob; regenerating a token uses a fresh key by default.
CryptographyMediumA cached runtime outlived its token's expiry while the client stayed active.Expiry is checked on every cache hit.
CryptographyLowThe master key was not validated in token mode; PBKDF2 iterations had no ceiling; token text accepted non-canonical base64.Master key must decode to 32 bytes; iterations 100 000 to 2 000 000; strict base64url.
MailMediumBcc recipients were disclosed to every recipient: the raw message carried the Bcc header onto the wire. Reproduced.The Bcc header is stripped from the wire copy.
MailMediumOversized bodies and attachments were silently truncated by the IMAP library and delivered corrupted.Downloads ask for one byte over the limit and refuse anything larger.
MailLowA recipient string with two angle-bracket addresses displayed one address and delivered to another.Exactly one angle address allowed.
MailLowForwarded headers were not sanitized; forwarded text was unbounded.Sanitized and bounded.
MailMediumAttached messages and containers (message/rfc822) were split into inner parts.Recorded as one attachment.
HTTPHighAnyone could mint a token pointing IMAP or SMTP at a private address, using a public server to probe internal networks; a name-based check alone could be bypassed with a public hostname that resolves to a private address.User tokens may not name loopback, link-local, private or CGNAT hosts on public servers (opt-in for intranet mail servers); mail hosts are resolved before connecting, every answer is vetted and the connection is pinned to the vetted address with TLS still verifying the hostname.
HTTPHighRequest bodies were buffered before any size check; a chunked body without Content-Length bypassed the declared limit entirely.Limits are enforced on the bytes actually read (streaming body limit): 1 MB on every route that parses a body, 50 MB on uploads, plus the per-token policy; uploads must declare Content-Length; at most 10 files per request; container content types stored as opaque files; on Vercel the upload cap is stated as 4 MB.
HTTPMediumThrottles keyed on X-Real-IP and friends even when no proxy was trusted, so a client could rotate headers.Proxy headers are trusted only on Vercel or with MAILMCP_TRUST_PROXY=1; otherwise the socket address is used.
HTTPMediumThe sign-in and unseal throttles shared one budget; unseal attempts with a correct password were never counted although each costs a PBKDF2 verification.Separate unseal budget; every attempt counts.
HTTPLowX-Forwarded-Proto was reflected unvalidated; /start (where the master key is generated) and /api/* lacked Cache-Control: no-store; checkout links were not scheme-checked.Fixed.
OAuthLowAn authorization code issued on one host could be exchanged on another host sharing the key.Codes are bound to the host that issued them.
OAuthLowDynamic registration accepted redirect URIs with fragments or credentials.Refused.
LicensingMediumThe signed licence key carried the buyer's e-mail address in readable form.The key carries the name and a 12-character fingerprint of the e-mail, not the address.
LicensingMediumThe Claude Desktop extension still honoured a test-only public-key override that the other bundles had compiled out.Compiled out of the extension as well.
LicensingHigh (vendor)A licence key for an unrelated Lemon Squeezy product whose name contained "unlimited" or "personal" would have been accepted and signed.Only the configured variant ids are accepted, product names are never trusted, and an optional store id pin refuses keys from other stores.
LicensingLowThe Lemon Squeezy webhook had no replay guard; the first version of the guard marked an event as handled before processing succeeded.Per-instance guard on event id, recorded only after success.
OperationsMediumThe Docker image and Vercel function shipped source maps; the Node entry had no crash handlers; an SMTP transport was not closed on failure; release tags were force-pushed; the customer changelog was generated from internal commit messages; the source Docker build broke when the audit page was added to the build.Source maps off; handlers added; finally around SMTP; tags and release assets are immutable unless explicitly re-tagged; curated CHANGELOG.md; Docker build fixed and verified.
Setup pageMediumLoading a token back into the form to add a mailbox silently reset policy fields the form does not show (a send limit of 0 became 10; attachment and token-lifetime settings reverted to defaults).Fields not shown in the form survive the round trip unchanged; 0 stays 0.
Prompt injectionMediumHidden-class rules were lost when rendering link text; a hidden image still contributed its alt text; white text in rgb(100%,100%,100%) form passed; an out-of-range HTML entity threw; deep nesting cost quadratic time; the untrusted wrapper did not remove invisible Unicode from text attachments.All fixed; nesting is bounded and dropped depth is counted, not scanned.
Local filesLowLocal attachments were checked and then opened separately, and special files were not refused.The file is opened once, inspected through the same descriptor, must be a regular file, and is read up to the checked size.
OperationsLowSending quotas, connection state and other limits could be raised without bound by a user's own token; hosting logs saw the credential-bearing link paths in Referer headers.Schema maxima independent of the token (mailboxes per token, bytes, rates); Referrer-Policy is no-referrer.
Public pagesMediumThe home page in single-owner mode listed the owner's e-mail addresses and permissions to anonymous visitors; /api/claim echoed the buyer's name and e-mail; error strings from the licence check and the webhook were echoed.Removed or made generic. (Fixed on 2026-09-05, before this audit.)
Public pagesLowSeveral sentences overstated the guarantees: "nothing in memory after the request" (15 minutes is the truth), "not even the operator sees passwords" (the operator's process decrypts them), "files never travel through the chat" (text attachments and explicit requests do), "a crafted e-mail cannot order it around" (mitigation, not immunity), and "security audit" without saying it was internal.Reworded on the home page, README and guide; this document is linked from the home page.

4.2 Accepted trade-offs (not changed, with the reason)

FindingWhy it stays
Access tokens live 30 days and refresh tokens 90 days; nothing revokes a single token; /revoke accepts and ignores.The design is stateless (no database), and ChatGPT does not refresh tokens proactively, so short access tokens produce daily "connection expired" prompts. The immediate revocation path is deleting the app password at your provider; the operator's kill switch is rotating the master key, which logs everyone out.
Refresh tokens are not rotated and a stolen refresh token can be renewed until the underlying user token stops working.Same reason: rotation without a revocation store gives no security benefit, and the design has no store. Deleting the app password at the provider ends every chain at once; on your own deployment rotating the master key does too. A bounded chain age is on the roadmap.
Throttles, the authorization-code replay guard and the send rate limit are per process. On Vercel each instance keeps its own counters.Stateless by design. They are honest best-effort limits, not global guarantees; the allowlist is the hard control on sending.
Anyone can create a token on a shared server unless the operator sets an invite code.That is what a shared server is for. Companies set MAILMCP_INVITE_CODE; the vendor's server is open so that people can try the product. Since 0.4.2 such tokens can only reach public mail hosts.
Two concurrent send_draft calls for the same draft can both send it; SMTP cannot promise exactly-once delivery.Known; the tool reports what was accepted. A per-draft lock is on the roadmap.
Generating your deployment configuration on someone else's setup page means trusting that page's code with the initial credentials.Inherent. Companies should generate their configuration on their own deployment's /setup (or locally with the CLI), which is what the guide recommends.
The operator's process decrypts your settings in memory while serving you.Unavoidable: the server must open IMAP with your password. This is the main reason to run your own deployment (section 7).
The AI vendor sees every tool result.That is the product. Passwords are the only thing kept from it.
Upload links are bearer credentials valid for one hour with unlimited uses; someone holding one can fill your upload folder or plant a file.Stateless links cannot count uses. The assistant is told that uploads may come from anyone with the link and to confirm with the user which file to attach; files are deleted once attached.
Bearer-header clients (Claude Code, Cursor, Gemini CLI) keep the raw token in a config file.Their design. Prefer OAuth in ChatGPT and claude.ai; keep the token in a password manager.
Drafts can be addressed to anyone (no allowlist on create_draft).Drafts are reviewed and sent by you from your mail client. Sending them through the assistant (send_draft) is allowlisted.
Personal-tier "one person" and the licence check itself are enforced by contract, not by code; anyone controlling the server can bypass the licence.The Sendy model. Offline verification and no telemetry beyond the daily usage-statistics ping described in section 18 were chosen over enforcement; the ping carries no licence key or tier.
Text attachments enter the chat; inline: true embeds binaries; in local mode everything is embedded.Requested behaviour, bounded at 2 MB, stated on the pages.
Google Fonts is loaded from Google's servers on every page.Convenience; Google receives visitor IPs. Self-hosting the font is on the roadmap.
The vendor's instance loads Vercel Web Analytics (MAILMCP_ANALYTICS_SCRIPT, served by the deployment itself).Page views, referrer, country, browser and device class as aggregate statistics; cookieless, visitors are a per-day hash, no identities. Only the deployment that sets the variable reports, and only to its own Vercel project. A self-hosted copy loads no analytics unless its operator opts in, and never reports to the vendor.

4.3 Open recommendations (roadmap)

  • Per-response CSP nonces instead of script-src 'unsafe-inline' (defence in depth; no injection point was found).
  • Publish SHA-256 sums with every release; sign them.
  • A bounded refresh-token chain age; an optional deny list for access tokens on long-running Node deployments.
  • Move CI into .github/workflows (blocked by a token scope) and run pnpm audit there; one transitive build-time dependency (tmp, via the extension packer) has two advisories that never reach the runtime bundle.
  • An explicit Vercel function duration and a streaming /files response.
  • Tests for a mailbox with read off and send on, for the rate limiter end to end, and for cache eviction.
  • A privacy notice and imprint on the vendor's site.

5. Properties verified as true

Each of these was checked in the code and, where marked, executed.

  • Mailbox passwords are encrypted in the browser with AES-256-GCM before anything leaves it; the server receives only the 32-byte key at token creation, never the settings. (executed)
  • A token alone does not yield passwords: unsealing needs the token and the edit password, which no assistant ever sees; the edit password is hashed with PBKDF2-SHA256 at 600 000 iterations in the browser and the server refuses fewer than 100 000. (executed)
  • The server has no database and writes no user data to disk; configuration comes from environment variables or a startup file; OAuth artefacts are self-contained encrypted tokens. Nothing in the logs contains tokens, passwords or message text.
  • Every key derived from the master key has its own purpose (seal, code, access, refresh, file, sign); ciphertexts cannot be replayed across layers; IVs are random per encryption; there is no padding oracle. (executed)
  • PKCE S256 is mandatory; authorization codes live 3 minutes and are bound to client, redirect URI, PKCE challenge, host and user token; redirect URIs match exactly; error responses never redirect; the login form has a SameSite=Strict CSRF cookie compared in constant time. (executed)
  • Client metadata documents are fetched with SSRF hardening: HTTPS only, port 443, no IP literals, private and loopback ranges refused on every DNS answer, the connection pinned to the vetted address, no redirects, 64 KB cap. (executed)
  • Recipient strings that could expand into several SMTP recipients are refused; the SMTP envelope is built only from validated bare addresses; an empty allowlist refuses to send; domain rules do not match look-alike domains. (executed)
  • Certificate validation is never disabled; every provider preset uses TLS or STARTTLS; IMAP commands cannot be injected through folder names, search strings or part ids.
  • Permanent deletion of mail is not implemented: trash moves to the Trash folder, moves to Junk/Spam need the delete capability, and moves to Outlook's Recoverable Items (which Outlook purges) are refused. mailmcp removes only messages it wrote itself (staged uploads and superseded signatures, identified by their marker header) and a draft it has just sent, and only with UID EXPUNGE of exactly those messages, after the server accepted the \Deleted flag; a server without UIDPLUS keeps them (tombstoned uploads and signatures, the draft untouched). On Gmail an expunged message leaves the folder but stays in All Mail (Gmail's default expunge behaviour). A server with neither MOVE nor UIDPLUS refuses moves instead of risking other messages. (since 0.13.0; built as a 0.9 hotfix that was never released on its own)
  • Runtimes are isolated per token; a token can never reach another token's mailboxes, uploads or cache entries. (executed, end to end)
  • Local file attachments (Claude Desktop) are confined to policy.attachment_dirs with real-path resolution and a separator-aware prefix check. (executed)
  • The licence is verified offline with Ed25519 against an embedded public key; the licence check never contacts the author; there is no telemetry beyond the daily usage-statistics ping described in section 18 (opt-out MAILMCP_NO_STATS=1, no request at all with MAILMCP_NO_UPDATE_CHECK=1).
  • The prebuilt distribution and the Claude Desktop extension contain no secrets, source maps or local paths; no secret has ever been committed to the repository. (executed)
  • Security headers on every response: HSTS, nosniff, X-Frame-Options DENY, CSP with frame-ancestors 'none', base-uri 'none', form-action limited to the server and the registered client, Referrer-Policy strict-origin-when-cross-origin, no-store on pages that show or take secrets. (verified live)

6. Residual risks no configuration removes

  1. The AI vendor sees every tool result and keeps it under its own retention policy.
  2. The operator's process holds your decrypted settings in memory while serving you and for up to 15 minutes after. "Stores nothing" is a property of this code, not a cryptographic guarantee; a modified server could store anything.
  3. Stateless credentials cannot be revoked individually. Delete the app password at your provider to end access at once.
  4. A token issued with broad permissions stays valid after you issue a narrower one.
  5. Visible instructions inside e-mails still reach the model. Sanitization removes hidden text and prevents the framing from being forged; it cannot stop a model from complying with text it can read. The allowlist, the hourly limit and the absence of permanent deletion bound the damage.
  6. Throttles and replay guards are per instance on serverless hosting.
  7. One-hour links are bearer credentials that appear in hosting logs and browser history.
  8. Plain-text IMAP/SMTP (tls: none) can be selected deliberately; presets never do.
  9. Whatever the assistant reads can leave through the assistant itself: its reply, or any other tool the user has enabled in ChatGPT or Claude. The permissions bound what mailmcp will do after a successful injection (no sending outside the allowlist, no permanent deletion); they cannot bound what the model does with content it has already read. That is the connector's edge, not a setting: connect only mailboxes whose content the assistant may see.

7. Why your own deployment is the best option for a company

  1. You own the kill switch. Every credential on the server derives from the master key. Rotating it invalidates every token, OAuth token and link at once. On the shared server that lever belongs to the vendor.
  2. You remove the unverifiable honest-operator assumption. During every request the operator's process holds your users' passwords and mail in memory. Nothing in the protocol lets a user verify what code a remote operator runs. On your deployment the operator is you, under your change control.
  3. The blast radius is yours alone. The vendor's project also holds the licence signing key and the tokens of everyone who tried the free server. An incident there is everyone's incident.
  4. You control registration and offboarding. MAILMCP_INVITE_CODE closes token creation; your Google Workspace or Microsoft 365 admin can revoke app passwords per user, which is the real revocation path in this stateless design.
  5. You choose the hosting party and region. Vercel sees the master key and plaintext requests after TLS termination. You pick the region, enable deployment protection and audit logs, or run the Docker image on your own hardware where no third party sees memory. That is what a DPO will ask for: a named processor under your DPA.
  6. No dependency on the vendor's uptime or business continuity. The licence verifies offline; the Deploy button clones the distribution into your GitHub; the running code is yours even if the vendor disappears. Tokens on a shared server die with it.
  7. You review updates. Updates reach you only when you pull a tag. The source code is not provided (EULA 1.1); the statutory rights of a lawful user under § 66 of the Czech Copyright Act stay untouched.
  8. You tune the policy. Owner configuration sets token lifetimes, rate limits, attachment limits and can pin owner mailboxes.

What you give up: the Unlimited licence (€149, one-time) plus hosting (Vercel Hobby is non-commercial, so Pro or a small VPS), and someone who watches the distribution repository and redeploys; security fixes do not arrive by themselves.

When the shared server is reasonable: trying the product, and individuals whose concern is "I do not want OpenAI or Anthropic holding my Google account" and who trust a small Czech vendor about as much as their hosting. Use read and draft only, set an edit password, and prefer OAuth over pasting the token into config files.

When local is best: one person, one machine, Claude Desktop or Claude Code. No server, no operator, no links. The trade: passwords sit on the machine, and binary attachments go through the AI vendor as base64. Not available for ChatGPT or claude.ai on the web.

8. Notes for data protection officers

  • Data categories touched: mailbox credentials (encrypted in the browser, decrypted in server memory per request), addresses and names of correspondents, message text, attachment names and, on request, contents, uploaded outgoing files (staged in the user's own mailbox), IP addresses (in-memory throttles and hosting logs), OAuth tokens (stateless, encrypted), licence buyer name and e-mail fingerprint.
  • Roles: on the vendor's shared server the vendor is a processor and Vercel a sub-processor; on your own deployment your organization processes for itself and the author receives only the daily update check and usage-statistics ping (section 18); the AI vendor and the mail provider are separate controllers in every option.
  • Sub-processors and recipients: the AI vendor (every tool result), Vercel (execution, logs, memory), the mail provider (IMAP/SMTP sign-ins from the server's egress address), Google Fonts (visitor IP on page loads), Vercel Web Analytics on the vendor's instance only (aggregate page views, cookieless), Stripe (purchases; merchant of record since 0.6.0, previously Lemon Squeezy), GitHub (downloads), Upstash (the vendor's usage-statistics counts since 0.9.1, EU region Frankfurt; section 18).
  • Data subject rights: there is no account and no stored profile. Erasure is deleting the app password and discarding the token; since 0.11 mailmcp also leaves records, folders, labels and categories in the owner's own mailbox, which the owner removes as described in /docs, "Removing what mailmcp keeps in your mailbox". Content already sent to the AI vendor is governed by that vendor's terms.
  • What to document: records of processing naming the AI vendor, the host and the mail provider; Article 28 agreements with Vercel and with the AI vendor (zero-retention or enterprise terms change the analysis); a DPIA, since correspondence contains third parties' data and may contain special categories; a legal basis for correspondents' data; employment-law groundwork for staff mailboxes; user instructions about prompt injection and about what the AI vendor sees.
  • Statements you can rely on are in section 5; the licence text contains no data-protection terms and does not replace a processing agreement.

9. Limits and defaults

SettingDefaultNotes
Message body per read8 000 characterspolicy.max_body_chars, 500 to 200 000
Search results50policy.max_results, up to 200
Attachment read into the chat2 MBpolicy.max_attachment_bytes
Attachment download link25 MBpolicy.max_download_bytes; link valid one hour
Outgoing attachments per message20 MBpolicy.max_upload_bytes; on Vercel uploads through a link are limited to 4 MB per request by the platform
Sends per hour per mailbox10 (Free, enforced), 50 (subscription, enforced, one quota per seat), the configuration's value on licensed serverspolicy.send_rate_per_hour lowered to the tier ceiling, per process
Recipients per field50to, cc, bcc
Access token / refresh token30 days / 90 daysserver.access_token_ttl_seconds, server.refresh_token_ttl_seconds
Authorization code3 minutesreplay guard per process
Sign-in attemptsowner password: 5 per address and 30 overall per 15 minutes; refused user tokens: 10 per address per 15 minutes, no overall cap (since 0.8.1)per process
Token creation60 per address per hourper process
Decrypted settings in memory15 minutes idle, 100 runtimes per processRuntimeCache
IMAP connection idle60 secondsone connection per mailbox per process
Outlook mailboxes per token3refresh tokens are large and the token travels in a request header
Microsoft sign-in validityabout 90 days from the sign-inshown on /setup and as reauth_by; rolled automatically on every OAuth refresh (section 13)
PDF text extraction5 MB, 50 pages, 20 000 characters, 10 s, one at a timeget_attachment; larger or scanned PDFs keep the link
Undo of a bulk batch (0.11)7 days, paid plansfrom when the batch ran; journals housekept after that
Bulk, unsubscribe and mode-B confirmations (0.10, 0.11)10 minutessealed refs; mode B single use claimed in the mailbox, claims kept 1 day
Unsubscribe one-click request (0.11)20 per hour per mailbox, 5 s total, 8 KB responseper process; off by default (capabilities.unsubscribe)
Draft card text (0.12)20 000 characterscut on a code-point boundary; the owner's own text only
Bulk preview from the page (0.12)re-run after 9 minutesconfirmations live 10 minutes
Mailboxes per licenceFree 2 per token (tokens created before 0.7.0 keep 5), subscription 10 per token on mailmcp.ai, Personal 5 per token or configuration, Unlimited no capFree (no key) adds a "Sent with mailmcp.ai" signature to every composed message (0.5.0)

0.7.0: the free tier caps new tokens at 2 mailboxes (seals carry iat; older seals keep 5). The seal may carry a licence key (lic) verified per request with its expiry. Nothing new is stored.

0.9.0 (mailmcp.ai subscription; like section 13, this paragraph was written after the reviews in section 1 and was not part of them): subscription keys (tier hosted: the Stripe subscription id, the seat number and an expiry of the paid period's end plus 30 days, no personal data) are sealed in the token and verified offline whenever the server builds a runtime for the token (its first request on an instance, and again after an idle runtime was dropped). Only a server started with MAILMCP_HOSTED=1 honours them, and distribution builds never do (the switch is compiled out); a subscription key never licenses a server of its own. Keys are issued only for invoice periods Stripe reports as paid; /claim?session_id= shows the keys of the first period only, later keys arrive by e-mail, and the e-mail re-send is throttled per IP address and per purchase address. /api/unseal now also returns the sealed key to whoever holds the token and its edit password; /api/relicense replaces the key with the token alone and gains no mail access: it answers with the new token only, never with the mailbox settings, so it gives nothing the token did not already give. Free-tier tokens are held to 10 sends per hour per mailbox by the server, whatever their configuration says, and subscription tokens to 50, shared by all tokens made with the same seat's key. Revocation works through the MAILMCP_REVOKED denylist (whole subscriptions or single seats), which takes effect on redeploy (instances of the previous deployment, and the runtimes they cached, keep the old licence until they end); this closes finding 7 of the 2026-09-23 security review for the subscription (a refunded or disputed subscription stops working), while a refunded one-time Personal or Unlimited key still works offline and can still be claimed, so for those the finding stays open. Finding 5 of that review (the send-limiter map was cleared at 5 000 entries, resetting every quota) is fixed: the oldest entry is evicted instead. Nothing new is stored.

10. What the test suite proves

88 automated tests: unit tests without network for cryptography, sanitization, policies, licensing, the purchase flow (Lemon Squeezy at the time of the audit, Stripe since 0.6.0) and the full OAuth flow; end-to-end tests against a real IMAP/SMTP server (GreenMail in Docker) for stdio and HTTP transports, shared multi-user mode, attachments in and out, uploads, forwarding, drafts and the ChatGPT connector contract. They prove the request-shaped attacks are handled: recipient smuggling, CSRF, PKCE and replay, SSRF against metadata and loopback addresses, tenant isolation, licence caps enforced server-side against client-built tokens, hidden-text stripping, delimiter forging through file names, path traversal for local attachments, tampered links. They do not prove behaviour under sustained load or across several serverless instances.

11. How the pages were corrected

The reviews compared every claim on the public pages with the code. Accurate: encryption in the browser, no database, offline licence verification, allowlist-only sending, no permanent deletion, one payment, 60-day refund (14 days since 0.6.1), tier limits. Corrected: the memory window (15 minutes, not "during the request"), what the operator can see, which attachments enter the chat, what hidden-text stripping guarantees, and the nature of this audit. Claims about the official Gmail and Outlook connectors of ChatGPT and Claude describe those vendors' documentation as of September 2026 and cannot be verified from this code.

12. Independent second-model review

OpenAI Codex (GPT-6 Astra, high reasoning, read-only sandbox) received the same brief and worked from the repository after the first round of fixes. Its report ran to 26 findings. It confirmed the cryptographic construction, the OAuth protections, the recipient and envelope handling and the Bcc stripping, and it found no cross-tenant access path. Its top ten, and what happened to each:

#FindingSeverityOutcome
1Private-network check on mail hosts was string-based; a public hostname resolving to a private address bypassed itHighFixed in 0.4.2: hosts are resolved, every answer vetted, the connection pinned to the vetted address
2Refresh tokens can be renewed indefinitely; /revoke does nothingHighAccepted trade-off of the stateless design, documented in 4.2; app-password deletion or master-key rotation ends every chain
3Consuming a staged upload deleted from a read-only source mailboxHighFixed in 0.4.2
4The 1 MB body limit relied on Content-Length; chunked bodies bypassed itHighFixed in 0.4.2 with a streaming limit
5A foreign Lemon Squeezy product could become a signed licence through the name fallbackHigh (vendor)Fixed in 0.4.2
6Silent truncation of oversized attachments and draftsMediumFixed in 0.4.2. The first round had described this fix, but it had not reached the code: a scripted edit had failed and the change was lost. The second reviewer caught it.
7modify_message could move to Trash without the delete capabilityMediumFixed in 0.4.2
8Remaining hidden-text gaps and inconsistent untrusted handling of structured fieldsMediumFixed in 0.4.2 (the HTML-part preference had likewise been described and not shipped; it is now in the code with tests)
9Send quotas reset with runtime evictionMediumFixed in 0.4.2
10The setup form weakened imported policyMediumFixed in 0.4.2

Also fixed from its list: the source Docker build broken by the audit page, webhook events marked as handled before success, release assets overwritten under an existing tag, the local-file check-then-open race, the missing forward tool for draft-only configurations, invalid HTML entities throwing, and quadratic work on deeply nested HTML. Left as documented trade-offs: per-instance code replay guards, concurrent send_draft duplicates, generating a configuration on a third party's setup page.

Its independent opinion on deployment matched section 7: a company should run its own deployment, provided it maintains it, with restricted egress, conservative capabilities and narrowly scoped mailbox credentials, and it should not treat the split-key design as a substitute for those controls. The sentence "I would not rely on the present split-key marketing or published fixed claims as substitutes for those controls" is quoted here on purpose: the pages were reworded, and this document states fixes against released versions with tests.

Lesson recorded for the project: every fix in this report is tied to a test that reproduces the original problem, and a release is not called fixed until that test passes in the shipped version.

13. Addendum: Microsoft Graph and the sign-in relay (0.8.0)

Written for release 0.8.0, after the reviews in section 1. Outlook.com and Microsoft 365 mailboxes do not use IMAP or SMTP: Microsoft rejects passwords on most accounts. Instead the user signs in at Microsoft and mailmcp holds a refresh token.

Components. src/mail/msauth.ts (device-code and token-endpoint calls, MsTokenSource with an in-memory rotation cache and error mapping), src/mail/graph.ts (GraphAccount, the second implementation of the Mailbox interface, talking to graph.microsoft.com/v1.0), the routes /api/ms/config, /api/ms/start, /api/ms/callback, /api/ms/device and /api/ms/poll next to /api/seal, the sign-in UI in the setup page, and src/mail/ms-reseal.ts (the rolling re-seal on our own OAuth refresh).

What is stored: nothing. As everywhere else in this codebase there is no database and no disk write. The Microsoft refresh token is encrypted in the browser into the user token exactly like a mailbox password; rotated refresh tokens and access tokens exist only in the process that obtained them. At sign-in the server relays the refresh token once from Microsoft to the browser (a same-origin postMessage from the popup callback page, or the answer to a poll) and keeps no copy.

Lifetimes.

ArtefactLifetime
PKCE verifier and state (authorization-code flow)encrypted cookie, httpOnly, Secure, SameSite=Lax, path /api/ms/callback, 10 minutes, single use (the reference carries a jti the server retires as it opens it)
Device-code reference (device flow)JWE with the device_code and the tenant that issued it (consumers for personal accounts, organizations for work accounts; the user picks the account type, because a code from /common cannot be redeemed by personal accounts) inside, TTL = Microsoft's expires_in (15 minutes), single use (each reference carries a jti, and the one it replaces is refused from then on), with a poll counter (240 polls) in the payload
Microsoft access tokenabout one hour, in process memory only
Microsoft refresh tokeninside the user token; about 90 days from the sign-in, re-sealed with a fresh one on every OAuth refresh of a long-lived client. The re-seal reaches only the client that performed that refresh: a token pasted into several clients leaves the others holding the original refresh token, which expires on its own 90-day schedule.
Decrypted configuration holding itthe existing RuntimeCache window: 15 idle minutes

Hardening. The Graph and login hosts are fixed constants, never derived from user input, so an oauth2 mailbox names no host an attacker could choose. Every path segment is encodeURIComponent-ed and every query string is built with URLSearchParams. All three POST routes require content-type: application/json, so a cross-origin form cannot drive them. /api/ms/device is counted in the sign-in throttle (5 per address and 60 overall per 15 minutes) only once the invite gate has passed; /api/ms/start, which asks Microsoft nothing, is limited to 10 per address with no overall cap (since 0.8.1), so a wrong invite code costs the caller's own address one of ten attempts and never the shared budget; /api/ms/poll has a budget of its own (250 per address per 15 minutes, against the roughly 180 polls a real sign-in makes), because the poll counter and the poll spacing live inside the reference and an in-memory replay guard is empty after a cold start; /api/ms/callback is counted at 20 per address, with no overall cap since 0.8.1. As in the rest of the design these counters are per process, so on serverless hosting the real budget is multiplied by the number of live instances. Permanent deletion stays impossible: the only DELETE /me/messages/{id} calls target the uploads folder for uploads and the signature folder for the signature prune, each only for messages carrying the matching X-Mailmcp-Upload or X-Mailmcp-Signature marker header, the same rule the IMAP path uses. Bodies from Graph pass through the same sanitizer and the same untrusted wrapper.

Accepted risks.

RiskWhy it stays
The device-code flow is a phishing primitive: an attacker who can get a victim to type a code at microsoft.com/link or login.microsoft.com/device collects a token. The relay runs on a trusted domain.Since 0.8.1 it runs only under the operator's own Entra app (MAILMCP_MS_CLIENT_ID); under the vendor's app it answers 404, and "Allow public client flows" is off on the vendor's app, so copies of 0.8.0 cannot relay a device code under it either. The vendor app's only redirect URIs are the hosted callbacks (mailmcp.ai and its test deployment); the generic nativeclient URI was removed, so no one can run an authorization-code sign-in under it outside those hosts. It is always behind the invite code, including when MAILMCP_OPEN_SIGNUP=1 opens the rest of token creation, so it is never an open relay. The vendor's server uses the popup flow and never the device flow. Microsoft's own advice is to block device code by Conditional Access (security defaults do so in new tenants); a tenant that does gets a message pointing at the fix: the server operator registers an Entra app of their own and switches to the popup flow (guide, chapter Outlook), or the tenant's admin allows the flow.
Consent at Microsoft is broader than mailmcp's own caps: a mailbox with drafts, sending, labels or Trash enabled consents to Mail.ReadWrite and Mail.Send, because Graph has no finer delegated permission.mailmcp enforces its per-mailbox capabilities on every call, but the refresh token used outside mailmcp carries the consented scopes. Read-only mailboxes are kept to Mail.Read, which is why the scope set is derived from the capabilities rather than requested once for everything. Widening capabilities later requires a new sign-in, and removing capabilities does not shrink the grant already given at Microsoft: the consent stands until the owner signs in again with fewer capabilities ticked or revokes the application at Microsoft. The setup page says so next to the capability boxes of a connected mailbox.
A replayed old refresh token of *our* OAuth (which cannot be revoked in a stateless design) makes the server re-seal a stale Microsoft refresh token, and Microsoft does not revoke a rotated refresh token.The same trade-off as the rest of section 4.2: our access and refresh tokens are stateless and unrevocable, and the kill switches are rotating the master key or revoking the application at Microsoft. Same class: a rotation that succeeds at Microsoft but whose re-seal fails leaves the token with the pre-rotation RT0, which Microsoft keeps valid until its own expiry, so the cost is at worst an earlier re-sign-in.
PDF text extraction parses untrusted attachments in-process (pdf.js through unpdf).Bounded: 5 MB, 50 pages, 20 000 characters, a 10-second budget and one extraction at a time, loaded lazily so it stays out of the cold-start path. The extracted text is wrapped as untrusted like any body, and the pages say it may contain text that is invisible in the rendered document, because no hidden-text stripping exists for PDF.
Self-hosted copies used the vendor's Entra application for sign-in until 0.8.0.Since 0.8.1 a self-hosted copy needs its own registration (MAILMCP_MS_CLIENT_ID) for any Microsoft sign-in; the vendor's app is used on mailmcp.ai only. MAILMCP_MS_DISABLED=1 removes the sign-in and MAILMCP_DISABLE_GRAPH=1 refuses Outlook mailboxes even in tokens already issued.
Since 0.13.0 every copy asks mailmcp.ai once a day (Claude Desktop at each start) whether a newer release exists (it no longer asks GitHub); server copies send the usage statistics with it.Described in section 18. The update link put into model-facing instructions is always built locally from the distribution repository's address; only a strictly newer x.y.z from the answer is used. MAILMCP_NO_UPDATE_CHECK=1 (in Claude Desktop the "Check for updates" setting) turns every request off.

14. Addendum: triage and bulk clean-up (0.10.0)

Written for release 0.10.0, after the reviews in section 1 and not part of them. The release was built and reviewed block by block by AI agents (Claude) against a written plan, with every fix tied to a test; that is the same kind of internal, AI-assisted review as the rest of this document, not a third-party audit.

New tools. triage, awaiting_replies and digest read message headers only (List-Id, List-Post, List-Unsubscribe, Precedence, Auto-Submitted, X-Autoreply, iCalendar parts, Outlook Focused/Other and the last-verb property), never bodies, and never set \Seen (IMAP BODY.PEEK[HEADER.FIELDS …]). The classification reasons are a fixed vocabulary, so no text a sender wrote reaches the model through them; subjects and senders pass through the same header sanitizer as everywhere else. bulk_preview and bulk_apply archive, mark read, label, move or trash many messages at once.

State: still none. 0.10 writes nothing to disk, to a database or to the mailbox. A bulk confirmation and a digest cursor are sealed refs (compact JWE, dir + A256GCM) that travel with the assistant's call: purpose-specific keys derived from the master key (oauth-confirm, oauth-cursor, distinct from the OAuth and file keys), the purpose also as the JWT subject, the principal in its own claim (derived from the token's sealed blob key, so it survives Microsoft re-seals and relicensing, and owner for owner mode and Claude Desktop), and the account. A confirmation never opens as a cursor, a ref sealed for one token never opens for another, and without MAILMCP_KEY the keys are random per process (refs then work only on that process; stdio logs this once).

Approval mode A only. A confirmation binds the operation: the ref carries the account, the action, the label, the destination, the absolute criteria (relative dates are resolved once, at the preview), the tier, and the exact members. bulk_apply also takes the action, the account and the count as plain arguments that must equal the sealed values, so the client's approval dialog shows what is being approved instead of an opaque string; the ref stays the authority. It expires after 10 minutes. It does not prove that a person saw the preview: a model, or a scheduled run, can preview and apply at once. The instructions, the prompts, /llms.txt and /docs tell the assistant to show the preview and ask, and tell the owner to "always allow" bulk_preview (read-only) but never bulk_apply, and to keep bulk_apply out of scheduled tasks. The server cannot tell a scheduled call from an interactive one.

Exact identity. On IMAP the confirmation seals the folder's UIDVALIDITY and the uid set itself (packed; 500 scattered uids fit a 4,000-character ref). On Outlook, immutable ids are too long to seal, so the confirmation carries the query that found the members, the preview instant and 32-bit SHA-256 prefixes of the member ids; bulk_apply re-runs the query bounded by receivedDateTime ≤ preview instant and keeps the candidates whose hash was sealed. The query is narrowed to the time span of the members actually sealed (after the plan cap and after shortening; it ends with the newest member's whole second, because Exchange reports receivedDateTime in whole seconds but filters on the stored milliseconds), and an exact sender is part of Graph's filter. If more candidates match than were sealed, or one sealed hash is matched by two candidates, the whole batch is refused; fewer is fine (the missing ones are reported as gone). One case cannot be detected: a member that left the folder after the preview while a message outside the preview, inside the same span and matching the same criteria, has the same 32-bit hash prefix. That message would be acted on in its place. The odds are about one in 4.3 billion per departed member and candidate in the span (for explicit ids, every message of the folder in that span is a candidate); the action stays recoverable (nothing is deleted), and the message must still pass the per-member checks below. A batch whose confirmation would not fit is shortened at preview time and says so. Apply revalidates every member before touching it: still in the source (IMAP: UIDVALIDITY and the uid; Outlook: the parent folder id), not flagged now when flagged mail was excluded, not read now when only unread mail was selected, not already done. Members are only ever dropped. A message delivered or unflagged after the preview is never touched.

Replay. A used confirmation is refused by the same instance (an in-memory set, like OAuth codes; the text names the batch). It counts as used only once every check has passed, and a refusal before anything changed (a rebuilt folder, a plan change, a transient error) leaves it valid. Another serverless instance does not know it was used, and a replay there is harmless: every member has already left the source (gone) or is already in the target state (already done), and nothing outside the sealed set can be touched. Cross-instance single use needs state and is left to a later release.

Safe mutations. Every IMAP move, including the bulk ones, goes through src/mail/imap-safe.ts (UID MOVE, or UID COPY + verified copy + \Deleted + UID EXPUNGE of exactly those uids); a lint test fails the build if messageMove, messageCopy, messageDelete, exec, a \Deleted store or CLOSE appear anywhere else. A server with neither MOVE nor UIDPLUS refuses archive, move and trash before anything changes; mark read and labels still work. On Outlook every member is one request through the per-mailbox semaphore (three in flight), never $batch; category changes read the @odata.etag and PATCH with If-Match, rereading on 412. Nothing is permanently deleted: trash is a move to Trash (Gmail and Outlook empty it after about 30 days; the preview says so), moves and labels to Junk/Spam need delete (a Gmail label is compared with the account's own Trash and Spam folders from LIST, so [Google Mail]/… and localized names count too), Recoverable Items is refused. On Gmail a bulk move out of All Mail is refused, because Gmail may delete a message that loses its last label there, depending on the account's IMAP expunge setting.

Reserved names. mailmcp-uploads, mailmcp-state, mailmcp-templates, mailmcp-snoozed and mailmcp-live-test (the last path segment, either delimiter, any case) are refused as move destinations and as sources of any change, and mailmcp/… or those names are refused as labels or categories. On Outlook a reserved folder is recognised by name, by id and by the message's parent folder; a reserved folder below a parent is recognised by id only once a folder listing has named it.

Deadlines. Each new tool call runs under one budget from request receipt (50 s over HTTP, under Vercel's 60 s limit; 120 s on stdio). Queue wait, connect and every IMAP command and Graph request share it; an IMAP command still running at expiry is stopped by closing the connection. Reads keep 5 s to answer and report partial coverage; bulk_apply starts a chunk (250 IMAP uids or 25 Outlook messages) only while 12 s remain, and a chunk still in flight at expiry is reported as unknown_outcome ("The mail server did not confirm the last step. Nothing was deleted. Run bulk_preview with the same criteria"), never retried silently. With account: "all", four mailboxes run in parallel, each with its own share of the time; one that times out reports its error and the others answer. Header bytes per call are capped at 4 MB.

Limits. triage: newest 60 per mailbox (at most 200), since at most 30 days back; 50 items listed per bucket, 15 for newsletter, lists and automated (counts cover all). awaiting_replies: Sent 500, inbound 2,000 per IMAP folder or 4,000 Outlook rows; Outlook headers are read for at most 25 candidate replies per call. digest: 500 headers per mailbox per call, cursor valid 30 days, at most 4,096 characters. Bulk: 50 per call on Free, 500 paid, at most 100 explicit uids, 1,000 candidates examined per scan window (continued with scan_from), confirmation valid 10 minutes.

Known limitations.

LimitationWhy it stays
The deadline covers the new tools and get_thread; the 0.9 tools keep their per-operation budget (Graph 50 s, IMAP socket timeout).Retrofitting all 21 is a follow-up (decision 5 of the release plan).
Gmail-specific paths (bulk label, archive by \Inbox removal, All Mail thread search) are tested against GreenMail and Dovecot semantics only; no live Gmail or Seznam account was available at release time.The release notes say so. The live suite covers the Microsoft paths on two test mailboxes (run on both on 2026-09-30: refs, triage buckets, bulk move / mark read / label, a stale If-Match answered with 412, get_thread); the Gmail and Seznam cases are written and wait for test accounts.
Outlook can deliver a message whose receivedDateTime is a few seconds earlier than one already read; a digest cursor at the later instant does not see it.The cursor is exact on IMAP (by uid); on Outlook it is exact at its boundary instant only. The next triage still sees the message.
Replay protection is per instance.Harmless by construction (see Replay).

15. Addendum: follow-ups, snooze, templates, confirmed sends and unsubscribe (0.11.0)

Written for release 0.11.0, after the reviews in section 1 and not part of them. Like 0.10, the release was built and reviewed block by block by AI agents (Claude) against a written plan, with every fix tied to a test; an internal, AI-assisted review, not a third-party audit.

State in the mailbox, never on the server. 0.11 is the first release that keeps state, and all of it is in the owner's own mailbox: follow-up, snooze, bulk-journal, send-claim and draft-provenance records in mailmcp-state, snoozed mail in mailmcp-snoozed, templates in mailmcp-templates. A record is one message whose body is JSON and whose X-Mailmcp-Mac header is an HMAC-SHA256 over the kind, the key, the Message-ID, the mailbox address, created and the exact decoded body. Records are authority only when they verify; an unverified record is listed as such and is never cleared, woken, superseded or deleted unattended.

The state key. policy.state_secret in the configuration, or, when absent, stateSecretOf(configKey) = HKDF-SHA256(configKey, salt "mailmcp-state", info "state-secret"), where the configuration key is the one that decrypts the configuration (the token's sealed blob key; MAILMCP_KEY in owner mode and Claude Desktop). The per-mailbox HMAC key is HKDF-SHA256(secret, salt "mailmcp-state", info "v2\0" + lower-cased address); the account id is not an input, so renaming an account keeps its records and changing its address orphans them (the owner's way back is wake_snoozed orphans=true and discard_unverified=true, which move to INBOX and to Trash, never expunge). The setup page writes state_secret into every new configuration and keeps it on Edit. Without any key (a plaintext developer configuration) bulk clean-up runs as in 0.10 without a journal and says so; setting follow-ups or snoozes, saving templates, provenance and mode B refuse or skip; clearing, waking and listing still run.

Recognition by location. A message is treated as a record only where it sits: an IMAP path of a record folder, a Gmail X-GM-LABELS entry of a record label, a Graph parentFolderId of a record folder; in Trash only when it also carries the record header and a @mailmcp.invalid Message-ID (tombstones). A message delivered to INBOX with a forged X-Mailmcp-State header stays visible in every read tool, and a record can never hide ordinary mail. Records in the record folders are refused by every generic read tool. The Date header is used only as a coarse index for the due query: a client that rewrites it can hide a record from that query, never make one verify. mailmcp removes only its own verified records: on Gmail by moving them to Trash (tombstones), elsewhere by their exact identity (on IMAP through the safe-mutation path). A tombstoned record restored from Trash or Outlook's Recoverable Items verifies again; restoring it needs access to the mailbox, which already allows everything the record could do.

Record-first follow-ups and snooze. Setting writes an intent record (with the prior flag state), then sets the markers and reads them back (a server may answer OK to a STORE and keep nothing), then writes a completion record with the markers actually present, then tombstones older records of the same key in the total order (created, Message-ID). Clearing trusts markers only from a completion and restores the prior flag only if it is still ours; a flag the owner changed afterwards is left alone. Snooze writes its intent before the move and its completion after; a wake finds the message by its identities (Message-ID, X-GM-MSGID, Graph id, UIDVALIDITY + uid), and two concurrent wakes move it once. Snooze moves one message, never its conversation. On Outlook the record keeps the exact due time and clearing compares against Outlook's read-back of the flag; spike S3 found the due time accepted and read back exactly, and what Outlook's clients show for it is checked in E5.

Bulk journal and undo. bulk_apply writes an intent (the exact members resolved at apply) before the first change and a completion after; an interrupted batch is reconciled from the mailbox on the next call or by bulk_preview batch= (read-only: status never writes). Undo (paid, 7 days) acts only on members found by exact identity: the uid within the recorded UIDVALIDITY, the Graph id or the Gmail X-GM-MSGID. A member found only by Message-ID could be the owner's own copy and is left alone and reported. A member the owner changed after the batch is left alone. A flag action (mark read, label, Gmail archive) whose completion had to be reconciled after an interruption is never reversed: mailmcp cannot tell its own change from one the owner made after the crash. If two runs of one batch both proceeded (a completion not yet visible in a lagging Outlook listing), status and undo merge their outcomes per member, so an all-skipped duplicate never hides the run that did the work. On Outlook the journal records a mark-read member as unread without a request (the preview selected unread mail only) and leaves a label's prior state unknown.

Send fingerprints. Draft tools return a fingerprint of the stored draft (IMAP: the appended bytes; Graph: one read of $value after creation), computed over the parsed message: bare lower-cased addresses including Bcc, the decoded subject, SHA-256 of the transfer-decoded text and HTML parts and of every attachment, and the source (reply, forward). send_draft with expect_fingerprint re-reads and refuses a draft that changed ("The draft changed since it was shown; nothing was sent."). On Graph, /send does not honour If-Match (spike S5), so the re-read is the guard; the window between that read and the send remains. Fingerprints are computed per request and never stored. Every send path strips X-Mailmcp-* headers before the wire: mailmcp adds no header to mail it sends, and draft provenance is a record, not a header.

Mode B: host-mediated confirmation (MRTR). On a modern-era request from a client that declares form elicitation and is listed as clientName@protocolVersion (the list ships empty; MAILMCP_ELICITATION_HOSTS can only add entries), a send, a confirmed bulk apply, an undo or an unsubscribe first returns input_required with a server-built dialog and a sealed requestState (purpose key oauth-mrtr, the tool, a hash of the canonical arguments, the fingerprint, the principal, the account, 10 minutes). The SDK's requestState.verify hook checks integrity only (a tampered state is a -32602 protocol error); expiry, arguments, principal and fingerprint are checked by the tool and answered in words. legacyShim: false was chosen deliberately: a legacy-era client never receives input_required (the SDK answers -32603), which matters only on stdio, where the era comes from the envelope. The dialog lists totals and every address (IDN domains in punycode, the subject quoted and labelled as coming from the draft, bidi controls removed); if the full text does not fit 300 characters no dialog is offered and a draft is saved. The retry must reproduce the same message fingerprint and the same dialog text (its hash is sealed too), so a Bcc added after the dialog is refused even where Exchange's stored MIME would not show it, and send_draft sends only against the fingerprint the owner confirmed; a message without a fingerprint gets no dialog. Single use across instances is a send-claim record keyed by the confirmation's jti: written, re-read, and only the earliest claim proceeds, and a tie in the server's order (Graph orders by creation time in whole seconds) makes both refuse, so a confirmation sends at most once or not at all. On Outlook send_draft claims first on the draft itself: a PATCH of a mailmcp property with If-Match on the etag seen at the dialog, writing a value unique to that attempt (Exchange accepts a stale etag when a PATCH writes the value already stored, seen live) and reading it back, so a second instance gets a 412, or finds another value, and sends nothing. Sends without a stored draft (send_message, reply_send, forward_message) keep the record claim and read it again after 1.5 s on Outlook; Graph's listing lag has no bound, so a small window remains there, bounded in practice by the in-process replay guard; a replay on any instance answers "This confirmation was already used; nothing was sent again." Mode B never widens policy: every check the tool makes without it runs first. clientInfo.name is self-reported, so the host list protects against a misbehaving model inside an honest host, not against a client that holds the token and lies.

Host-only sending. A mailbox set to send only after the dialog is stored as send: false plus send_host_only: true, so 0.10 and older copies refuse to send from it instead of sending without the dialog. On clients without mode B its sends become drafts and send_draft leaves the draft. It requires the draft capability.

Unsubscribe trust boundary. mailmcp relies only on an authentication verdict it can attribute to the owner's own provider:

Account classTrusted Authentication-ResultsStatus
Gmail (X-GM-EXT-1)the topmost header, authserv-id exactly mx.google.comtrusted
Outlook / Microsoft 365 (Graph)the header Exchange Online addsmanual-only until spike S9 proves its format and position
Other IMAPnonemanual-only

The one-click POST is sent only when every condition holds on a fresh read of the complete header block (at most 64 KB; a longer one is manual): the trusted header is the topmost one; it has exactly one dkim=pass identifying exactly one signature (header.b as a prefix of at least 8 characters, else header.d + header.s); that signature's h= covers List-Unsubscribe and List-Unsubscribe-Post; each header occurs once; List-Unsubscribe-Post is List-Unsubscribe=One-Click with exactly one HTTPS URI of at most 2,048 characters, without userinfo, fragment, IP literal or a port other than 443; the URI's host is the signer's d= or below it, and d= is not a bare public suffix (a valid signature from attacker.example cannot unsubscribe the owner at victim.example); the message is not in Drafts, Sent or a mailmcp folder (APPENDed mail carries no provider verdict); and the mailbox has the unsubscribe capability, which is off by default. The dry run shows only the host and names the verified signer as the sender. Execute re-evaluates everything and requires the URL, the signature, the Message-ID and the UIDVALIDITY to equal the sealed values. The POST goes to an address vetted and pinned like the CIMD fetch, carries only List-Unsubscribe=One-Click as multipart/form-data, no cookies, follows no redirect (a 3xx is not a confirmation), and runs under one 5-second deadline over DNS, connect, TLS, request and response, reading at most 8 KB; 20 per hour per mailbox. This trusts the provider's verification; mailmcp does not verify DKIM itself.

Recovery mode. An expired mailmcp.ai subscription token with more mailboxes than the Free cap signs in and runs only the recovery set on all its mailboxes (listing accounts, follow-ups and templates, clearing follow-ups, cancelling and waking snoozes, deleting templates, batch status); every other tool answers with the renewal steps. The set is a table checked at registration, so a new tool is refused in recovery mode unless listed.

Known limitations.

LimitationWhy it stays
The invocation deadline covers the tools added or changed in 0.10 and 0.11; the remaining 0.9 tools keep their per-operation budget (Ruling 17).Retrofitting the rest is a follow-up.
Gmail and Seznam paths (labels, records in All Mail, Trash tombstones, the unsubscribe verdict) are tested against GreenMail, Dovecot and fake clients only; no disposable Gmail or Seznam account was available (Ruling 19).The live suite covers the Microsoft paths; the Gmail cases are written and wait for test accounts.
Graph /send ignores If-Match: a draft changed in the moments between the fingerprint re-read and the send is sent as changed.Spike S5 showed no conditional send on Graph; the re-read narrows the window to one request.
Rate limits and the per-instance replay map for mode A refs remain best-effort per instance.Single use that matters (mode B) is claimed in the mailbox.
Mode-B sends on Outlook without a stored draft rely on the record claim, read again after 1.5 s; a listing that lags longer lets two instances both see only their own claim.send_draft uses a compare-and-swap on the draft; the others have no object to compare on, and the in-process replay guard catches a retry on the same instance.
Outlook send_draft posts the stored draft, so X-Mailmcp-* headers someone else put on a draft cannot be stripped; Reply-To is not part of the fingerprint.mailmcp itself adds no such header; the fingerprint's fields follow Ruling 8 (a gap noted for the owner).

16. Addendum: the clickable inbox (MCP Apps, 0.12.0)

Written for release 0.12.0, after the reviews in section 1 and not part of them. Like 0.10 and 0.11, the release was built and reviewed block by block by AI agents (Claude, with a separate security reviewer on the server side and on the widget) against a written plan, with every fix tied to a test; an internal, AI-assisted review, not a third-party audit.

The resource. One MCP Apps page, ui://mailmcp/app.html with MIME type text/html;profile=mcp-app, is registered on every transport (HTTP, stdio, the MCPB, recovery mode included) unless the configuration turns the panels off (0.13.0). It is a generated, committed string (src/http/app-bundle.ts, built deterministically by scripts/build-app.mjs), never served by an HTTP route, so dist copies, Docker and the MCPB serve the same bytes. Its _meta.ui declares empty CSP lists (connectDomains: [], resourceDomains: []) and no permissions; the listing and the content item carry the same metadata. Nine tools point at it with _meta.ui.resourceUri: triage, digest, search_messages, get_thread, create_draft, reply_draft, bulk_preview, bulk_apply, unsubscribe. No tool is app-only and no tool was added: the model sees 33 tools, as in 0.11. Since 0.13.0 five tools point at it (triage, digest, create_draft, reply_draft, bulk_preview); the other four answer as in 0.11, and a configuration with policy.ui: "off" registers neither the page nor any _meta.ui.

The meta CSP. The page's first <head> element is its own Content-Security-Policy: default-src 'none'; script-src 'sha256-…'; style-src 'sha256-…'; img-src data:; font-src data:; connect-src 'none'; frame-src 'none'; worker-src 'none'; manifest-src 'none'; form-action 'none'; base-uri 'none', with the hashes of the one inline script and the one inline style. It is defence in depth under the host's own sandbox and CSP. zod is switched to jitless before anything else loads, so no eval is needed. A test asserts the tag is first, the hashes match and no unsafe- source appears. The browser suite counts every request the frame attempts, including ones the CSP blocks, and requires zero.

No Claude domain; the OpenAI origin only when trustworthy. _meta.ui.domain is never set: Claude derives the sandbox domain from the exact URL the user entered and refuses to render on a mismatch, and the page makes no network calls that would need a stable origin. For ChatGPT, _meta['openai/widgetDomain'] carries the server's own origin only when it is https:, not loopback or private, and comes from MAILMCP_PUBLIC_URL, Vercel or a trusted proxy (MAILMCP_TRUST_PROXY=1); otherwise it is omitted, so a self-hosted server behind an untrusted proxy never advertises an internal address. stdio never sets it. A wrong value can only affect ChatGPT's rendering, never a text result.

Allowed links. The page opens exactly one kind of link, through one function (safeOpen, the only caller of the host's openLink): an attachment download link taken from a get_attachment result, https: (or http: on loopback), with the path /files?r=<16–9000 safe characters> and no other query (0.13: Vercel refuses paths over ~2 KB) or the older /files/<16–4096 safe characters> without a query, no userinfo or fragment, and the origin of the first such link pinned for the session. Everything else is refused and the chip says "Ask the assistant to save it". The Claude directory entry lists https://mailmcp.ai; self-hosted copies get Claude's confirmation prompt for every link.

Mail text never becomes markup or navigation. Every mail-derived string reaches the DOM as a Text node through one helper that removes control characters and the invisible and bidi set of sanitize.ts; mail-derived elements are dir="auto" and isolated, so a subject cannot reorder a neighbouring button label. Links in mail stay text and remote images are never loaded. A lint over the widget sources (with a fixture per rule) forbids innerHTML and every other markup sink, window.open, location, .href =, .src =, srcdoc, creating a, form, iframe, img, link or script elements, non-literal attribute names, on* property handlers, script-dispatched clicks, and any tool call outside the single click gateway.

Model-channel templates. The page writes into the conversation (ui/message) and the model context (ui/update-model-context) only through fixed builders that accept account ids, uids (IMAP numbers or Graph ids of [A-Za-z0-9_=+/-], at most 512 characters), plain folder names, counts and batch ids, and throw on anything else. The one free text is what the owner types into "Ask for changes" (at most 2,000 characters, the field starts empty). No subject, sender, address or mail text can reach either channel (a property test over 1,000 hostile rows). Model-context updates are best effort and never the only way the model learns of an action.

_meta versus structuredContent. Every UI result keeps its 0.11 content text and structuredContent byte for byte (snapshot test on IMAP- and Graph-shaped fakes), plus structuredContent.kind from a fixed vocabulary. Everything else the page needs (the draft card, account labels and capability flags, as_of) goes into the result's _meta['mailmcp/ui'], which hosts document as going to the widget and not the model. Because a host may forward _meta to the model after all, it follows the same rules as content: header strings through sanitizeLine, the card's text inside the untrusted-content markers (the page strips exactly those two lines), and only text mailmcp can attribute to the owner is shown as the owner's: the reply part (or a forward's comment) when the provenance split adds up. The quoted or forwarded message is named by its recorded sender, date and subject, and at most its first 600 characters travel separately, inside the same markers, shown collapsed. Without _meta the page renders what structuredContent already exposes and shows no Send.

The draft card, and refresh before arm. The card is built from the parse of the stored draft, the same parse send_draft fingerprints, so it shows every recipient including Bcc, every attachment and the sending identity. Its Send is disabled whenever the shared pre-check (src/mail/send-precheck.ts, also used by send_draft: read, send, attachments policy, size, recipients, allowlist including Bcc) refuses, when a send of the draft is recorded (already_sent: after every send a sent record keyed by the draft's Message-ID, which send_draft also refuses on, so a draft left in Drafts is never sent twice), when mailmcp's record cannot show the owner's own text or name the quoted message (unreviewable, also from the first card when the configuration has no records key), and when the draft has an HTML part mailmcp did not generate from the text itself (html_differs: the record keeps the hash of generated HTML only; any HTML the assistant supplies, or any HTML changed afterwards, disarms, because the card cannot show HTML). The first click calls get_message, which returns a fresh card for a draft in Drafts (only to clients that declare MCP Apps support, or whose request carries no capabilities, and for drafts within max_download_bytes, the whole result within 60,000 characters), and arms only if the fingerprint, account, uid, recipients and attachments are unchanged; the second click calls send_draft with that fingerprint, which the server checks again. Any other action, a new result, a theme change or Back disarms. A card replayed from an old conversation is never sent from what it showed; a send is shown as "Sent" only on send_draft's success shape, never on an input_required answer. Bulk actions follow the same pattern: the first click previews per (account, folder) group, the second applies with each group's own confirmation, a preview older than 9 minutes (confirmations live 10) is run again, and a preview that returns after the selection changed never arms. Handlers that render results call no tool (lint and a browser test).

The postMessage target. The ext-apps transport checks event.source on what it receives, but sends with the target origin "*": card payloads and tool results go to whatever origin the parent frame has. The page cannot pin the host's origin across hosts, so it relies on the host's sandbox; the payloads are the same data the host already holds from the tool result.

Hosts may skip their own per-tool prompts for widget calls. A host that asks before a model's tool call may not ask before a call from the page. Then the gates are the page's two clicks and the server's checks, which are the same as for any model call.

Mode A plus UI. A call from the page is an ordinary tools/call that the server cannot tell from a model's. Every policy runs exactly as for the model: capabilities, the allowlist including Bcc, the rate limit, the tier, the fingerprint, and host-only sending. _meta.ui.visibility is not used and is never treated as authorisation. The card is not a confirmation: the model can send the same draft within the owner's rules; only a mailbox set to send after the app's dialog (send_host_only) makes a send wait for the owner.

Third-party code in the bundle. The page bundles the App client of @modelcontextprotocol/ext-apps 2.0.3, tree-shaken against the repository's own @modelcontextprotocol/client and core 2.0.0 and zod 4.5.4, so every copy is pinned by our lockfile and covered by pnpm audit (the upstream app-with-deps build, which embeds its own copies, is not used). The page is 653,518 bytes, under a budget of 660,616 (the measured client plus 80 KB for our code) and a cap of 716,800. ext-apps and playwright-core 1.63.0 (the browser test suite) are development dependencies only; neither is a runtime dependency of the server.

Gate outcomes. schedule_send (Outlook send later) and respond_invite (invitation replies) were planned behind feasibility tests. The schedule_send test ran before 0.13.0 and failed (Microsoft 365 cannot cancel a deferred send), so the tool was dropped; Outlook send later stays a scheduled-task recipe. The respond_invite test was not run and the tool is deferred. Neither is registered, and the release notes say so. The owner tested the panels in the Claude app before release, which led to the 0.13.0 narrowing to five tools and the per-token switch; the remaining host behaviour (_meta forwarding, prompts, fonts, the CSP in each host) was not checked host by host, so the /docs host table still says "(to be verified)" and problems are expected from user reports.

In-card editing deferred. A revise_draft tool that edits the draft from the card was designed and deferred to a later release (it is not in 0.13.0). The review of its design found problems a safe edit has to solve first: header injection and encoding in a hand-written splice, recovering the quote from text-only offsets, signature images after a signature change, Exchange re-rendering stored drafts (which would refuse every edit after the first), the IMAP selection after an APPEND, a crash window between writing the new draft and retiring the old one, and concurrent edits. The card is read-only apart from Send; "Ask for changes" asks the assistant for a new draft, and the earlier draft stays in Drafts.

Related changes. create_draft and reply_draft return a uid on IMAP servers without UIDPLUS (a fresh SELECT, then the Message-ID); every draft mailmcp saves gets a provenance record; send_draft refuses an Outlook draft larger than max_download_bytes instead of reading it whole; forward_message with as_draft no longer returns the forwarded body outside the untrusted-content markers.

Known limitations.

LimitationWhy it stays
The "*" postMessage target (above).The ext-apps transport sends that way and the host's origin differs per host; the host's sandbox is the boundary.
The page's contrast is tested against our own colours; a host that supplies its own palette decides the contrast of structural colours.Host variables take precedence by design.
Graph Bcc in the fingerprint, and whether hosts honour cacheHint, are checked live before release, not by the automated suite.They need a real Exchange mailbox and real hosts.
Homograph or IDN look-alike recipients are shown as they are, not flagged.The allowlist is the enforcement; flagging is a follow-up.

Review before release. A release-wide /recheck by Claude agents ran on the 0.12 work (record docs/replan/audit/2026-09-30-134346-recheck-0d6700e.md): no Critical finding; the Important findings (among them hidden HTML passing the card and a double send after a replayed card) were fixed with tests before 0.13.0. No independent second-model pass of this addendum was run.

17. Addendum: protection levels (0.13.0)

Written for release 0.13.0, after the reviews in section 1 and not part of them. Like 0.10 to 0.12, the release was built and reviewed block by block by AI agents (Claude, with a separate security reviewer) against a written plan and specification, with every fix tied to a test; an internal, AI-assisted review, not a third-party audit.

Threat model. A verification email (a one-time code, a sign-in link, a password reset) is a bearer credential delivered to the inbox. An assistant connected to mailmcp holds private data, reads untrusted content and has ways out: the send tools, and the client's own web fetch and link rendering. Published attacks make an assistant read such a code and pass it on. 0.13 removes the private-data leg for that subset of mail: the mailmcp server, not the model, decides per mailbox which verification emails the assistant may see, before any tool result is built. It does not make the rest of the mailbox safe to read (section 6 still applies).

The setting. accounts[i].capabilities.protection: off, basic (codes, sign-in links, recovery), standard (plus security alerts, and redaction of codes and sign-in links in other mail), strict (plus bank and payment mail) or {level: "custom", categories}. It lives only in the sealed configuration; no tool argument, prompt, resource or _meta field can lower it or ask for hidden mail. Parsing is lenient and fails closed: an unknown level or type is read as Strict, and a configuration without the field is Off, reported as set: false by list_accounts. New mailboxes on the setup page start at Standard. Changing the level creates a new token; the old token (and OAuth grants of up to 90 days) keeps the old level, so every page, the CHANGELOG and the setup page pair the change with replacing the token in every assistant. There is no server-wide minimum level. One mailbox entered twice with different levels would let the assistant read through the lower one; the setup page warns about that, and list_accounts shows every level.

The classifier. src/mail/protect.ts and protect-rules.ts are pure functions: English and Czech rules, headers first (subject, sender, the One-Time-Code header), then the body, NFKC plus a confusable fold with an offset map, invisible characters stripped before matching. Categories are a union: a message hidden by any selected category is hidden. At most 256 KB of each text and 4 KB of a subject are analysed; text beyond that is never returned (it is cut with the usual truncation note) and paths that send bytes onward treat a truncated text as unanalysed. Length never hides a message on its own. Meeting passcodes stay visible. A body code counts only next to a verification keyword, or next to the generic word "code"/"kód" (any case) when no qualifier makes it another kind of code (booking, reservation, order, customer, access, error, voucher …); a sign-in word right before or after the keyword ("customer login code", "kód pro přihlášení", "kód pro vstup do internetového bankovnictví", "přístupový kód do aplikace") overrides the qualifier, because a code for signing in is a login code. Amounts with decimals never count. Next to a bank or payment sender the weak words never mean codes: "Potvrzení platby" is banking (Strict only), while a bank's code mail is caught by its verification phrase or the code in its subject. A message is hidden whole by its body only when it comes from a table sender, an automated or role address, or carries a precise or strong phrase in its subject; mail from a person is only redacted. These precision rules came from the release recheck, which measured booking codes, invoice amounts and colleague mail being redacted or hidden at the default level; every changed corpus pin is listed in the release report. The rules version is reported as rules in list_accounts.

The decorator. Every mailbox the runtime hands to a tool is wrapped in ProtectedMailbox (src/mail/protected-mailbox.ts), whatever its level. Its table is typed Record<keyof Mailbox, MethodPolicy>, so a new backend method does not compile until it has a policy. The kinds: pass (no foreign content, each with a written reason: folder lists, record storage, positions, markers, bulk members already selected under protection); byRef (a method naming a message reads that message's header facts first, so the body of header-hidden mail is never downloaded, and at hiding levels mutations also run the body pass); mask (the housekeeping reads used by wake and follow-ups return hidden rows with subject and sender blanked, never dropped, so moving mail back keeps working); rows (triage and digest scans drop hidden rows and count them); body (message, thread, reply context, writing samples and draft inspection: header check, body pass, redaction); keep (search and bulk resolution: the backend applies the predicate before counting, sampling or paging, and stamps the result; an unstamped result has its total discarded and is logged as protect_unstamped). Facts ride on results under a non-enumerable symbol and are answered by position.

Two kinds of hidden. Header-hidden mail is left out of listings and refused by id without its body being read. Body-hidden mail (the headers look ordinary, the body is a verification email) is refused wherever a body is returned or sent, and by the destructive single-message operations (trash, modify, snooze, set a follow-up), which read the body at hiding levels. Plain listings may show a body-hidden row's headers.

Fail closed. At every level except Off: missing facts, a failed header fetch, a message that cannot be found, and an unknown level all count as hidden or unchecked, with one refusal text ("could not find or check this message…"). Unchecked rows are counted separately from hidden ones. The folder a caller names never exempts a message from classification: only the backend says whether a message sits in one of mailmcp's reserved folders (Graph from its parent folder id, IMAP from the folder it opened), and Graph read paths also refuse a named reserved folder the message is not in. The first review found that a by-id call on Outlook naming mailmcp-uploads skipped the check and returned the body unfiltered; it was fixed before release and is covered by a test of every by-reference tool on the real Graph and IMAP backends.

Searches and the oracle. The model controls the query, so any result that depends on hidden or redacted content would let it read a code digit by digit. Plain listings (folder, dates, flags only) skip header-hidden rows, keep the backend's total and report hidden_by_protection. Content queries (any query, from, to, subject, every Gmail raw query) skip header-hidden rows and body-check every remaining candidate, dropping body-hidden rows and, at redacting levels, every row where redaction would fire, whether or not the query touches those words; they report no hidden count. The body check reads the same text the redaction analyses: on IMAP both alternatives (the HTML, sanitised and with hidden text kept, and the plain part, which SEARCH TEXT matches too), each part up to 1 MB; on Graph the whole body. A candidate whose part is longer than 1 MB, or whose text is longer than the 256 KB that is analysed, is left out, since it could not be checked whole. The server's search also matches inside attachments, so the pass reads them too, within bounds (five parts of at most 1 MB each): text parts with their charset, attached mail (its own header and body pass; a hidden one is not looked into) and PDF text layers through the extractor; pictures, sound and video hold no searchable text. A candidate whose body or analysed attachment is hidden or would be redacted is dropped. When some part could not be analysed, the candidate is kept only if every query word shows in its analysed text (subject, sender, analysed body and attachments, and file names the level lets out unchanged), so nothing unanalysed can change the answer. A file name the level would hide or change counts as not analysed, since servers match names too. A query with a negation ("-word", NOT), a wildcard or an operator other than the header, state, label, date and size operators can make the server's own answer depend on an attachment or a name (it excluded or matched there), so such a query keeps only candidates with no attachment at all. Accepted residue: the check compares words as substrings of the analysed text, while servers match whole words or word prefixes; the two can differ on a word, but the difference depends only on analysed text, never on a hidden part. The release recheck found the earlier gap as a digit-by-digit oracle on a code inside an attached verification mail; a test now compares the answers with that mail attached and with a harmless one. The walk stops once offset + limit + 1 rows are visible, as plain listings do; the exact visible total is reported only when the walk ran out of matches, otherwise the total is the end of the page with total_exact: false, so a total never depends on how many bodies were dropped. Every search result from a protected account carries one constant line built from the account id and the setting only, byte-identical whether anything was hidden or not. A test compares whole content-query results between a mailbox with hidden, body-hidden and redaction-fired matches and one without, on the fake backend and on Dovecot. The comparison runs on the fake, on the real IMAP and Graph backends (over test doubles of the servers) and on Dovecot, with up to 100 matches at offsets 0 and 40. Accepted residue, stated here: each call examines at most offset + limit + 1 + 200 rows and reads at most offset + limit + 1 + 40 bodies, and hitting either bound yields total_exact: false; that happens with or without a secret only when more than 40 candidates in the window are dropped, and it can say "many matches were left out" but never what they contain; plain listings bisected by date reveal when hidden mail arrived; get_thread reveals that a thread has hidden members through coverage.hidden_by_protection. Response times can differ with the number of bodies checked; the model does not see latency, so this is accepted.

Redaction and bytes that leave. From Standard up (and Custom with codes, links or recovery), code tokens near a code keyword become [code removed] and sign-in URLs [sign-in link removed] in returned text and in reply quotes (quote HTML is rebuilt from the redacted text when redaction fired or an original link matches). The bytes of mail where redaction fires, or that could not be analysed, never leave: forward_message, send_draft (bound to the draft's fingerprint), attachment references in compose and get_attachment refuse, and /files answers 403. Text-like parts, PDF text layers and nested message/rfc822 parts (three levels deep, deeper refuses) are classified themselves, and an attachment's name is shown, downloaded, forwarded and attached only as the level lets it out ("(hidden by the protection level)" otherwise). A part is classified by the stricter of its declared type, its file name and its first bytes, so an attached mail or a text file sent as application/octet-stream is checked as mail or text; text is decoded with its declared charset, and a NUL left after decoding counts as not analysed. Attachment names are classified like subjects (Graph names an attached item after its subject; Gmail's forward-as-attachment calls it "<subject>.eml"): a hidden name is replaced with "(hidden by the protection level)" in message, thread and card listings, and codes are removed from the others at redacting levels. The text an onward path checks is read up to one character past the analysed 256 KB, so a longer text counts as not analysed; on IMAP a text part longer than the 4 MB read is flagged as cut. A PDF whose text layer could not be read (another extraction running, a timeout, the size or page cap, a parse error) does not leave either; only a PDF with no text layer at all (a scan) passes, as a binary attachment does. Such refusals say the message "could not be checked" rather than that it holds a code. Text the model composes in send_message and reply_send is not scanned: the owner may dictate a code on purpose. The draft card shows the reason protected and does not send.

Bulk. Header-hidden mail is never selected, sampled or counted by bulk_preview. Body-hidden mail can be moved (never deleted) by a bulk action, because selection is header-only; this is accepted. The confirm ref binds a hash of the protection setting, so an apply after a changed setting is refused. That matters most in owner mode, where the principal is the constant owner: without the hash a ref issued before a configuration change would stay valid. A 0.12 ref, which carries no hash, is accepted only for accounts at Off.

Microsoft Graph. Thread listings request the internet headers; rows without them are classified with known: false, which is more cautious and can make the same message visible in one call and hidden in the next. A search, triage or digest row whose headers were not read (beyond the per-call header fallback) cannot rule out a One-Time-Code header, so at a level that hides codes it is left out and counted as unchecked. The $search window is at most 1,000 rows; a full window with too few visible rows gives total_exact: false.

Download links. A /files/… link stays a bearer URL for one hour: an assistant with web fetch can read what it points at. Links are issued only for mail that passed the checks above, and /docs says so.

Binary attachments. Images and office files attached to visible, unredacted mail are not scanned; a code inside a screenshot passes.

Confusables. NFKC folds full-width and compatibility forms, and a fixed map folds common Cyrillic and Greek look-alike letters to Latin; other look-alike characters are not caught.

Records, Sent rows and Outlook flags. Follow-up and snooze records, awaiting items and Outlook flag rows are read without the decorator; every subject or counterpart a tool takes from them goes through the same header rules first. Rows of list_followups, the digest's due follow-ups and snoozes, and wake_snoozed are also checked against the message where it sits now, header and body (one batched header read, a body read only where the body pass could hide it): a record written at Off cannot name mail hidden by a One-Time-Code header or by its body. A row whose message cannot be found or checked in time is shown as hidden. The draft card keeps the quoted source's One-Time-Code fact in its record. Writing samples (list_templates samples) get the body pass too. Hidden mail still wakes and follow-ups on it still clear; results report it only as a count.

Nothing stored, nothing logged. Classification runs in memory while answering. The only log lines are protect_unchecked (account id, method, error class, count) and protect_unstamped; no subject, sender, code or URL is logged. A leak walk calls every tool in tools/list plus the accounts and app resources with canaries in every field of hidden fixtures and in the body, code and link of body-hidden and redacted ones, and fails on any canary in any serialised result.

Rollback is fail-open. A 0.12 server does not know the field and shows everything again, verification emails included; the CHANGELOG says so. Upgrading again restores the level from the configuration.

Known limitations.

LimitationWhy it stays
The rules catch typical English and Czech verification emails, not all; long-tail services with generic subjects rely on the body pass and on redaction.A deterministic classifier with sourced rules and 36 pinned hard negatives; precision over recall outside the tested corpus.
A sender can hide its own mail from the assistant (a fake "Security alert" is hidden, so the assistant cannot warn about it); a verification-like reply in a real thread makes get_thread refuse when it is the root.Neither reveals real verification mail; the owner still sees the mail in their mail app.
Bulk can move body-hidden mail; binary attachments are not scanned; confusables outside the fold map pass.Stated above.
Rolling back to 0.12 shows everything.Older servers cannot enforce a field they do not know.

Review before release. Two security reviews and a release-wide /recheck by Claude agents ran on the 0.13 work (record docs/replan/audit/2026-10-01-003426-recheck-9feb4b6.md): the recheck found a Critical search oracle (server-side search reading attached mail and query operators), fixed with "with vs. without the hidden part" probes on every backend before release, and its Important findings were fixed with tests. No independent second-model pass of this addendum was run.

18. Addendum: Usage statistics and MCP request counting (0.9.1, shipped in 0.13.0)

Written for release 0.9.1 (developed as 0.9.1, never released on its own, shipped in 0.13.0), after the reviews in section 1 and like section 13 not part of them. Two features, both statistics only. B counts running self-hosted copies: every copy sends at most one small request a day to mailmcp.ai, which is also its update check. A counts MCP requests on the vendor's own server mailmcp.ai as Vercel Web Analytics events. Code: src/stats/ping.ts (sender, ships in every build), src/stats/receiver.ts, src/stats/store.ts, src/http/stats-page.ts and src/stats/mcp-events.ts (vendor only, absent from the distribution, see "Dist boundary"). The public description with the exact request, the switches and the numbers is https://mailmcp.ai/stats.

Data flow (B).

StepWhat happensRetention
PayloadPOST https://mailmcp.ai/api/ping, user-agent: mailmcp-ping, body exactly {"s":1,"v":"0.13","r":"docker","d":"…"}: schema, version as major.minor, runtime (vercel, docker, node, claude-desktop, stdio) and the digest. No licence key, tier, mode, address, URL, project or machine id, configuration, mailbox count or content. With statistics off the copy sends GET /api/version without a body instead. Both answer {"latest":"x.y.z"}; the update notice uses only a strictly newer x.y.z, and its link is built locally. 2-second timeout, every error swallowed, never delays a response.none on the sender
WhenAt most once per UTC day per process, and statistics at most once per UTC day per instance: copies outside Vercel keep the day of their last ping in ~/.mailmcp/stats-sent-<runtime>-<id> (the id is a short HMAC of a fixed label under the instance's statistics secret, so other instances on the machine keep their own), so a Claude Desktop launch later that day sends only the body-less version check. Vercel: on the first request of the day (waitUntil; statistics from production deployments only, previews send the version check). Docker/Node: 1–60 minutes after start, then at a uniformly random time of each later day (no interval anchored to the start time). stdio: at start (waits at most 1 second), then at a random time of each later day. Nothing is sent between 23:50 and 00:10 UTC.
Digest with local saltsecret = HKDF-SHA256(ikm = local salt, salt = MAILMCP_KEY bytes or empty, info "mailmcp-stats-v1"); d = the first 16 bytes of HMAC-SHA256(secret, message) as base64url (22 characters), where the message joins mailmcp-stats, v1, the UTC month YYYY-MM and the runtime with vertical bars. The local salt is VERCEL_PROJECT_ID on Vercel, elsewhere /etc/machine-id or /var/lib/dbus/machine-id, else a random 32-byte salt written once to ~/.mailmcp/stats-salt (mode 0600). Where that directory is not writable, and on Vercel without a project id, the salt is random per process: never a public value such as the production domain, and never host facts such as the host name, user or MAC addresses. It never leaves the process; the vendor cannot compute d even though /setup on mailmcp.ai can know a MAILMCP_KEY. d is stable for one calendar month and unrelated to the next month's.
Receiver checksBody at most 512 bytes (413), keys exactly s v r d with enumerated values (400, nothing stored). The /stats sample digest is answered and never stored. In-memory throttle of 120 pings per network prefix (/64 or the IPv4 address) and 20 000 per instance per hour, keyed by a SHA-256 of a per-process random salt and the prefix (429). A version above the receiver's own major.minor, or not canonical, is recorded as other.throttle: one hour, memory only
Hashed HLL elementsFor each of the 3 dimensions (all, v:<minor>, r:<runtime>) and 2 periods (day, month), the element is the first 16 characters of the base64url SHA-256 of <period>:<dim>, a vertical bar and d. Redis (Upstash) never receives the raw d, and no two HyperLogLogs share an element. There is no key that combines two dimensions.
Lua scriptOne EVALSHA per ping (published verbatim on /stats): INCR pings:<day>, and only while the count is within the daily budget (MAILMCP_STATS_DAILY_BUDGET, default 860: a month stays under the Upstash free plan's 400,000 commands even if every command inside the script is billed, 15 per ping as measured on the test deployment) PFADD and EXPIRE of the six HLLs. After the budget the ping is answered and not counted, the day is marked capped, and each instance stops touching the store for the rest of the day. A store error or quota refusal still answers 200 {latest}; a store call is given up after 1.5 seconds, before the sender's 2-second timeout, so a slowly recorded ping is not sent again.HLLs: day 48 h, month 40 days
Periods and finalizeOnce per instance and UTC day, the last two finished days and the previous month (whenever its HLLs still exist) are counted with PFCOUNT, written with HSETNX as one JSON field of the days / months hash, and their HLLs deleted. A day's HLLs live 48 hours from its last ping, so a day is lost only if nothing reaches the receiver on the following day.aggregates: days 400 days, months 25 months (pruned at finalize)
Publication/stats: daily and monthly totals exact, breakdowns by version and runtime rounded to 5 ("< 5" for 1–4); cached 10 minutes; any query string redirects to /stats.

Data flow (A). Only on mailmcp.ai with MAILMCP_MCP_EVENTS set, and only from a Vercel production deployment. Counted: initialize, tools/list and tools/call messages of a POST to /mcp (of a batch at most the first 10) and requests refused before MCP (401, unlicensed 503; at most 60 per instance and hour). Not counted: notifications, ping, server/discover and subscriptions/listen (the newer stateless protocol's first requests, sent instead of initialize: such clients appear only through tools/list and tools/call; owner ruling for 0.13.0), resources/*, prompts/*, calls of unknown tools; a body without a declared length (chunked) or declared larger than 64 KB is not inspected (its tools/call messages are still counted from the tool result). The event name is mcp <method>; the properties are the tool name if it is one of the server's own tools (- for other methods) and "<client family> <outcome>", where the family comes from a fixed list (ChatGPT, Claude, Claude Code, Cursor, other) and the outcome is ok, error or refused. The event goes to <public origin>/_vercel/insights/event with a fixed user agent and ts rounded down to the hour; no client IP, cookie, client user agent, token or content is forwarded. The raw client name and user agent are logged only with MAILMCP_STATS_DEBUG=1 on a deployment that also sets MAILMCP_STATS_NAMESPACE (the test project).

Can an installation be identified? (amended threat review)

Who / whatCould it identify an installation?Why not / residual
PayloadNoThree enumerated values and a digest; no install id, host, URL, licence or tier.
Across monthsNod is an HMAC of the month under a secret derived from local-only entropy; months cannot be joined.
Within a monthA pseudonym, not an identityThe days of one installation share d for that month (needed for monthly-active counts). It carries no other data, never reaches Redis in raw form and ends at month-end.
Vendor with Redis accessMarginal counts onlyHashed elements in HLLs until counted (≤ 48 h / 40 days), then numbers per single dimension. A count of 1 for a rare version or runtime says one such copy ran, not which; public breakdowns are rounded.
Vendor serving /setupNot today/setup on mailmcp.ai can know a MAILMCP_KEY, which is why the local salt, which the browser never sees, is the HKDF input key. See residual risks.
Vercel request logsThe sender's IP address, time and pathAs for any request to mailmcp.ai: 1 day on the Pro plan. Our code never stores the address; the throttle holds only a salted hash of the prefix. The rollout sets IP Address Visibility to hidden in the team settings, so log views and drains show no address.
Country / geographyNot collectedVercel's country header and every other geo header are never read (a test greps src/stats; the build guard fails a bundle that names the header).
Network observerThat a host talks to mailmcp.ai dailyTLS, as for any update check.
MCP events (A)The hour, method, tool, client family and outcomeNo IP, cookie, user agent, token or content; within Vercel's log retention an event could be linked to the request log entry of the same hour.

Legitimate-interest assessment (Art. 6(1)(f) GDPR).

  • *Purpose.* Know how many copies run, which versions and runtimes to support, when an old version can be retired, and which MCP features and clients are used on mailmcp.ai, to plan support and security fixes.
  • *Necessity.* There is no account or licence server that could answer this; the licence check is offline by design. A daily request with enumerated fields is the least data that yields daily- and monthly-active counts. Alternatives rejected: a server-side hash of the IP address (processes the address for statistics, collapses copies behind NAT), exact sets of digests (would keep a monthly pseudonym per copy for a month), tier or geography fields (can single out a buyer).
  • *Balancing.* No identifier; marginal counts only; the raw digest never reaches storage; retention bounded (hashes 48 h / 40 days, aggregates 400 days / 25 months); what is sent and the totals are public; the request is easy to see (one start-up log line) and to switch off. Server operators run infrastructure and reasonably expect an update check; the ping is disclosed in the changelog before upgrading, in /privacy, the README and the EULA draft.
  • *Safeguards.* MAILMCP_NO_STATS=1, DO_NOT_TRACK=1, MAILMCP_NO_UPDATE_CHECK=1 (no request at all; "Check for updates" in Claude Desktop), each accepting 1, true, yes or on; CI and source builds never send, Vercel preview deployments send only the version check; the switch is the Art. 21 objection.
  • *Desktop copies opt in.* Claude Desktop and other stdio copies run on a person's computer, where Art. 5(3) ePrivacy (§ 89 of Czech Act 127/2005) reads broadly. Their statistics are off unless the user ticks "Anonymous usage statistics" or sets MAILMCP_STATS=1; an empty or unexpanded setting counts as off, so upgrades stay off. They keep the body-less update check.

Dist boundary. The receiver, the store, /stats and the MCP events are loaded by dynamic import() behind an inline process.env.MAILMCP_DIST !== '1' check in createApp, which esbuild folds away in the dist and MCPB builds. Each vendor module carries a marker string (the prefix is deliberately not quoted here, because this audit is itself bundled into the distribution); scripts/vendor-guard.mjs, called from build-dist.mjs and build-mcpb.mjs, fails the build if any bundle contains one, if a bundle lacks the sender or names the geo header, if the MCPB manifest's statistics setting is not off by default or its "Check for updates" setting is missing, or if the Dockerfile lacks MAILMCP_RUNTIME=docker. A dist copy answers 404 on /api/ping, /api/version and /stats even with MAILMCP_STATS_RECEIVER=1 (checked on the built dist/node.js).

Residual risks.

RiskWhy it stays
Vercel keeps request logs (IP address, time, path) for 1 day on the Pro plan.Any request to mailmcp.ai is logged by the platform; the receiver stays on mailmcp.ai (owner decision). Our code never stores the address; the rollout hides IP addresses in the team's log views and drains.
The days of one installation within a month share one pseudonym.Required for monthly-active counts; it carries nothing else and ends at month-end.
The /setup JavaScript is served by the vendor and could in theory read local values and send them.It does not: the local salt is read only in the server process (VERCEL_PROJECT_ID, machine id, the salt file), never by the browser. A change would show in the published bundle.
Copies without a machine id keep a salt file in their home directory.It is random (32 bytes), readable only by the copy's user and never sent; deleting it only makes the copy count once more that month. No host name, user name or MAC address is hashed.
A recreated container without /etc/machine-id and without a persistent home directory writes a new salt and counts again that day and month.Accepted upper-bound error (plan decision 13); a volume on /home/node/.mailmcp keeps one salt across recreates.
Pings are unauthenticated and can be forged.Bounded by the per-prefix throttle, the Vercel Firewall rate limit on /api/ping, the Redis-side daily budget (the day is marked capped) and the version allowlist. Counts are labelled self-reported on /stats.
The Web Analytics insights endpoint is used without Vercel's package and is not documented as an API.Pinned by a test and checked with exact counts on the test deployment; if it changes, events stop arriving and nothing else breaks.
MCP events carry the hour and could be linked to a request log entry within Vercel's 1-day log retention.Stated in /privacy; the event itself has no IP, cookie, user agent or token.