Deploying an MCP server to production means moving from a local stdio process on your laptop to a stable, remotely accessible service that handles multiple clients reliably. The shift introduces requirements that local development sidesteps: persistent availability, authentication, TLS, monitoring, and graceful handling of errors that would just crash a local process. This guide covers the deployment architecture decisions, the practical steps for getting a server running on cloud infrastructure, and the operational practices that keep it reliable.
Choosing Between Managed and Self-Hosted
The first production decision is whether to deploy on your own infrastructure or use a managed platform. Self-hosted on a cloud VM (EC2, Compute Engine, DigitalOcean Droplet) gives full control over the environment, no vendor constraints on what the server can access, and typically lower per-instance cost for high-throughput servers. Managed serverless platforms (AWS Lambda, Google Cloud Functions, Cloudflare Workers) handle availability and scaling automatically but impose execution time limits that conflict with long-running MCP tool calls — most serverless platforms have 15–30 second maximum execution times, which is insufficient for tool calls that make slow external API requests or process large files. Container-based managed services (Google Cloud Run, AWS App Runner, Fly.io) are the practical middle ground: containerised deployment with automatic scaling, no execution time limits, and less operational overhead than self-managed VMs. For most production MCP servers, Cloud Run or Fly.io is the recommended starting point.
Containerising Your MCP Server
Containerisation is the foundation of a production MCP deployment. A Dockerfile for a Python MCP server installs dependencies, copies the server code, exposes the HTTP port, and sets the start command. The container should be stateless — all persistent state stored externally in a database or object storage, not in the container filesystem. Environment variables inject configuration at runtime: credentials, database URLs, API keys. The container image should be as small as practical — use slim base images (python:3.12-slim rather than python:3.12), install only required dependencies, and exclude development tools. Build the image in CI, push to a container registry (Google Artifact Registry, ECR, Docker Hub), and deploy from the registry rather than building on the production host. This reproducibility — the same image runs in staging and production — eliminates “works on my machine” deployment failures.
Deploying to Cloud Run
Google Cloud Run is a strong default for production MCP SSE servers. It runs containerised workloads with automatic HTTPS (including a free managed TLS certificate), scales to zero when idle (no cost when not in use), and scales up quickly under load. The deployment process: build and push the container image to Google Artifact Registry, run gcloud run deploy with the image URL and configuration flags (region, memory, CPU, environment variables), and Cloud Run generates a public HTTPS URL for the service. Secrets (credentials, API keys) should be stored in Google Secret Manager and injected as environment variables at deploy time rather than hardcoded in the container configuration. Cloud Run’s minimum instances setting (set to 1 for servers where cold start latency is unacceptable) prevents the few-second delay on first connection after a period of inactivity.
Figure 1 — Production MCP server deployment: architecture overview
Health Checks and Availability
Production MCP servers need a health check endpoint — a simple HTTP GET that returns 200 when the server is running correctly. Cloud Run, App Runner, and similar platforms use this endpoint to determine whether to send traffic to an instance. A useful health check verifies not just that the server process is alive but that its critical dependencies are accessible: a database-backed server should verify it can connect to the database; an API-backed server should verify its credentials are valid. Failing health checks that route traffic away from unhealthy instances prevent clients from receiving cryptic connection errors when the server’s dependencies are down. For SSE servers, the health check endpoint is separate from the SSE endpoint — it is a plain HTTP GET, not a persistent connection.
Logging and Observability
Production MCP servers should emit structured logs for every tool call: timestamp, tool name, arguments (sanitised to remove sensitive values), execution time, success or error. Structured logs (JSON format) integrate with cloud logging services (Cloud Logging, CloudWatch, Datadog) and enable querying for specific patterns — all tool calls that took longer than 5 seconds, all errors from a specific client. Beyond logs, emit metrics: request count, error rate, latency percentiles, and any domain-specific metrics (e.g. database query count per tool call). Cloud Run’s built-in metrics cover request volume and latency; application-level metrics require instrumentation in the server code using a metrics library (Prometheus client, OpenTelemetry). Set up alerts on error rate spikes and latency p99 degradation so you are notified of problems before users report them.
Versioning and Rollback
MCP servers evolve — tools are added, modified, or removed. Clients that depend on specific tools need migration paths when those tools change. Use semantic versioning for your server (expose the version via the MCP server info response) and maintain backward compatibility within a major version: adding new tools is non-breaking; modifying or removing existing tools is breaking and requires a major version increment. Cloud Run’s traffic splitting enables gradual rollouts: deploy a new version and route 5% of traffic to it while keeping 95% on the previous version, observe error rates, and gradually shift traffic when confident. This eliminates the binary “all old / all new” deploy risk for changes that affect production clients. Maintain the ability to roll back to the previous image version with a single command — for Cloud Run, this is gcloud run services update-traffic --to-revisions=previous-revision=100.
Cost Management
Cloud Run bills per request and per CPU/memory second during request processing, with no charge when idle (scale-to-zero). For MCP servers with bursty, low-volume traffic (a team of 5 developers using it throughout the day), Cloud Run’s cost is typically under $10/month including free tier credits. For high-volume servers processing thousands of tool calls per hour, a dedicated VM or container cluster may be cheaper. Monitor actual usage and cost in the first month after deployment and set budget alerts to catch unexpected traffic spikes. One practical optimisation: keep tool call responses concise — returning large JSON payloads increases response time, memory usage, and network cost simultaneously. Return only the data the model actually needs to complete its task, not everything your system has available.
Production MCP deployment is straightforward when the architecture is right: a stateless containerised server behind a managed HTTPS endpoint, with OAuth authentication, structured logging, health checks, and a rollback path. Cloud Run or Fly.io covers most use cases with minimal operational overhead. The investment in production readiness — containerisation, auth, monitoring — pays back quickly in the form of a reliable service rather than one that needs manual restarts and generates support tickets from confused users. Get the foundation right and the server runs itself.
Connection Limits and Rate Limiting
SSE servers maintain persistent HTTP connections — one per connected client. A server handling 20 simultaneous users maintains 20 open connections, each occupying memory and file descriptors. Configure your container with sufficient resources: 256MB RAM handles about 50 simultaneous SSE connections for a typical server; 512MB handles about 150. Set connection timeouts to reclaim resources from idle connections (clients that are connected but not actively making tool calls): a 5-minute idle timeout is a reasonable starting point. Rate limiting protects against abuse — a misconfigured agent that calls a tool in a tight loop can generate thousands of requests per minute. Implement rate limiting at the connection level (maximum requests per minute per client) or use a gateway like Kong or Nginx to enforce rate limits before requests reach the server. Log rate limit events so you can distinguish between a runaway agent (which needs to be fixed) and a legitimate high-throughput use case (which needs a higher limit).
Staging Environment and CI/CD
A staging environment that mirrors production configuration — same container, same secrets structure, same auth setup — is essential for validating changes before they reach production users. Deploy every code change to staging first, run automated integration tests that call each tool and verify responses, and only promote to production after tests pass. For CI/CD, a typical pipeline: on pull request, build the container image and run unit tests; on merge to main, deploy to staging and run integration tests; on tag, deploy to production with traffic splitting at 10% initially. GitHub Actions, GitLab CI, and Cloud Build all integrate with Cloud Run deployments. The integration tests for an MCP server should cover the full tool call cycle — not just unit testing the tool implementations in isolation but actually calling the tools through the MCP protocol to verify the full request-response cycle works correctly.