Runpod
Runpod is an on-demand GPU cloud for AI teams. It rents dedicated Pods, autoscaling serverless inference endpoints and multi-node clusters across 30+ GPU models and 31 regions, billed by the second with no contract.
What is Runpod?
Runpod is an on-demand GPU cloud that positions itself as the AI developer cloud, covering the whole lifecycle from experiment to production on a single account. It is operated by Runpod, Inc., a US company based in Moorestown, New Jersey, with a remote-first team spread across the United States, Canada, Europe and India and a presence in San Francisco. The company reports more than one million developers and announced $120 million in annual recurring revenue in January 2026.
The platform rents compute in three shapes. Pods are dedicated GPU instances giving direct control over the container, drivers, storage and runtime, suited to development, training, fine-tuning and long-running jobs. Serverless exposes containerised models as autoscaling GPU endpoints behind an API, moving from zero to hundreds of workers on demand, with FlashBoot cold starts advertised under 200 milliseconds and no charge while idle. Clusters handle distributed training and large batch work, reaching 200+ simultaneous GPUs over InfiniBand on demand and up to 10,000 under reservation. Two further layers sit on top: the Hub, offering deployable open-source templates and models, and Public Endpoints, pre-deployed models such as Qwen3, Granite 4, Flux, Sora 2, Wan and Whisper, callable by API with no infrastructure at all.
The catalogue spans over 30 GPU models, from a 24 GB RTX A5000 to a 288 GB HBM3e B300, taking in H100, H200, B200, A100, L40S and RTX 4090. Runpod advertises 31 regions across the US, Europe, Asia and Australia, in Community Cloud and network-isolated Secure Cloud tiers. Billing is metered per second from worker start to full stop, with no egress fees, no minimums and no contract.
Compliance is documented: SOC 2 Type 2 is complete, with SOC 3, HIPAA and GDPR material held in a trust centre. Certifications apply at platform level, and data centre coverage varies by region.
What it does
- Rent a dedicated cloud GPU billed by the second and running in under 30 seconds
- Deploy an autoscaling inference endpoint that drops to zero workers when idle
- Train or fine-tune models on multi-GPU instances with persistent storage
- Run distributed multi-node clusters of 200+ simultaneous GPUs over InfiniBand
- Call pre-deployed models through an OpenAI-compatible API with no infrastructure setup
- Host persistent AI agents sharing memory and model weights on network volumes
- Deploy open-source templates and models in one click from the Runpod Hub
When to use Runpod / When not to
A quick filter to help you decide if Runpod is the right fit.
When to use Runpod
- Machine learning engineers who need H100, A100 or B200 capacity for training and fine-tuning runs without signing a contract or waiting on a provisioning queue
- Teams serving bursty inference traffic, where serverless workers scale to zero between requests and idle capacity costs nothing
- AI startups and independent developers watching unit economics, who benefit from per-second metering, published per-GPU rates and no egress fees
- Platform and MLOps engineers who want full control of the container, drivers and runtime rather than an abstracted managed service
- Research groups and labs running image, video or audio generation at scale, as in the Civitai case study of 800,000 LoRAs trained monthly on 500+ concurrent GPUs
When not to use Runpod
- Anyone looking for a ready-made AI application: Runpod sells raw infrastructure, and you must bring a container and your own code
- Users who want to try a product before paying, since the account is prepaid with credits and deploying a Pod requires at least one hour of GPU balance
- Organisations that need guaranteed data residency out of the box, because compliance coverage depends on the data centre selected and must be configured deliberately
- Fault-intolerant production jobs placed on spot instances, which can be evicted whenever demand spikes
- Teams expecting a mobile app or a no-code console: the product is web, API and CLI only, and the interface is English only
How to use Runpod
A typical end-to-end flow, from setup to results.
- Create an account on the Runpod console; no sales call or contract is required to start
- Fund the account with prepaid credits, keeping at least one hour of GPU balance, which is the minimum needed to deploy a Pod
- For a dedicated instance, open Pods then Deploy, pick a GPU model and a region, attach a container image and launch; the instance is live in under 30 seconds
- For serverless inference, write a handler function containing your application logic
- Package the handler in a Docker image and push it to a container registry such as Docker Hub or ECR
- Deploy the serverless endpoint from the console or the REST API, then start sending requests to it
- For the quickest path, use Public Endpoints instead: set parameters in the console playground and copy the generated code from the API tab
- Connect external tools by pointing any client that accepts a custom base URL at the OpenAI-compatible endpoint with a Runpod API key
- Reach running Pods through the automatic HTTP proxy, direct TCP ports, SSH or a remote IDE
- Manage costs with spending limits and auto-pay; billing starts when the workload runs and stops when it is terminated
Pros & Cons
Pros
- Per-second metering with no egress fees, no minimums and no contract on self-serve usage
- Very fast availability: a GPU instance running in under 30 seconds and serverless cold starts advertised under 200 milliseconds
- Broad, current GPU catalogue reaching a 288 GB B300, with no provisioning queue or sales call
- One account covers the entire path from prototype to enterprise agreement without replatforming
- Genuine pricing transparency: Pod and cluster rates are published in the clear, GPU by GPU
- Documented compliance posture: completed SOC 2 Type 2, a publicly available DPA, a named sub-processor list and a designated Article 27 representative
- Full container control and OpenAI-compatible endpoints, so PyTorch, TensorFlow, JAX, n8n, CrewAI or LangChain all work without adapters
Cons
- No free plan and no free trial: the account is prepaid and deploying requires an hour of GPU credit up front
- Serverless prices do not render on the pricing page without JavaScript; only the $0.58 to $9.98 per hour range is stated in plain text elsewhere
- Several cluster tiers and all reserved capacity are quote-only, with no public price
- Compliance coverage varies by data centre, region, provider and deployment model, so residency must be configured and verified rather than assumed
- Compliance documents such as the SOC 2 and SOC 3 reports require an approved access request through the trust centre
- Spot instances can be evicted when demand spikes, and network volumes keep billing while Pods are stopped
- Technical entry barrier: a container, a Docker image and code are required, with no no-code path and an English-only interface
Pricing & Plans
There is no permanent free plan and no free trial. Runpod operates a prepaid, pay-as-you-go account in US dollars, where credits are drawn down in real time and deploying a Pod requires a balance covering at least one hour of the chosen configuration. The lowest published entry point is an RTX A5000 at $0.27 per hour, confirmed both in the pricing table and in the Pods FAQ, with an L4 at $0.49, an A40 at $0.44 and an RTX 3090 at $0.50. Larger cards rise to $1.39 for an A100 PCIe, $2.89 for an H100 PCIe and $7.89 for a B300. Rates are displayed hourly but metered per second, with no egress fees or minimums. Storage is charged separately from $0.05 per GB per month. Prices were recorded on 24 August 2026 and Runpod states openly that it moves prices to keep GPUs available.
- Pods
- on-demand dedicated GPU instances billed per second
- available in Community Cloud and network-isolated Secure Cloud tiers
- in Reserved (guaranteed) or Spot (interruptible
- cheaper) mode
- Serverless
- autoscaling inference endpoints charged per second
- with flex workers that scale to zero and pre-warmed active workers
- ranging from $0.58 to $9.98 per hour
- Clusters on demand
- multi-node compute with published rates for H200 SXM at $4.31 per hour and A100 SXM at $1.79 per hour
- other GPUs quote-only
- Reserved Clusters
- committed capacity over 1
- 3
- 6
- 12 months or longer
- entirely quote-only
- Public Endpoints
- pre-deployed models billed per request
- per second or per 1
- 000 characters with no infrastructure to configure
- Enterprise
- reserved capacity at committed-use rates with a contractual 99.99% uptime SLA
- SSO
- role-based access control
- consolidated post-paid invoicing and migration support
- Storage
- billed separately from $0.05 to $0.20 per GB per month depending on tier and whether the volume is running or idle
Data, GDPR & hosting
A consolidated view of how Runpod handles your data.
GDPR overview
GDPR implementation is concrete and documented rather than merely asserted. Runpod announced on 6 February 2026 that it had been independently audited as meeting GDPR standards, and it publishes a full data processing agreement that customers can sign and return. That agreement incorporates the EU standard contractual clauses of decision 2021/914, the UK addendum B1.0 and Brazilian clauses, and imposes a purpose limitation on Runpod and its sub-processors. Attachment 2 names nine sub-processors, including AWS, Google Cloud, Stripe, HubSpot and Cloudflare. An Article 27 representative is designated for the EU and the UK, Prighter Group, reachable through a dedicated portal, and a data protection officer is contactable at privacy@runpod.io. Rights requests go to dsar@runpod.io. The minimum age is 18. The trust centre lists SOC 2 Type 2, SOC 3, HIPAA and GDPR material, though downloads require approved access.
Who owns the data?
The terms of service state plainly that Runpod asserts no ownership over customer content and that the customer retains full ownership of it, together with any associated intellectual property rights. The customer grants Runpod a non-exclusive, worldwide, royalty-free licence limited to two uses: operating and providing the service, and using the content in aggregated and anonymised form to improve the service and Runpod's related products. Under the published data processing agreement, Runpod acts as processor and the customer as controller. Personal data is deleted irretrievably or returned on the customer's request unless law requires retention, and a shared-responsibility model leaves the customer accountable for protecting data within its own perimeter.
Reuse rights
Customers keep their content and may reuse, move or delete it freely without asking Runpod for permission, since Runpod claims no ownership and the licence it holds is limited to running the service. Runpod's own reuse is narrower but not trivial: the terms permit aggregated and anonymised use of customer content to improve the platform and related products. The data processing agreement adds a purpose limitation, stating that neither Runpod nor its sub-processors may process customer personal data for their own purposes, and listing the categories concerned as identity, communications and IT usage data. The service also generates performance data automatically, meaning logs, telemetry and technical metrics. The privacy policy, effective 7 August 2025, separately covers marketing use, message personalisation and targeted advertising, each with an opt-out. The terms were last updated on 24 March 2026.
Data retention & training
Hosting summary
Runpod advertises 31 regions spread, in its own words, across the US, Europe, Asia and Australia, with the customer choosing a region at deployment. The only nominative list published is a documentation table of 17 data centres, covering Canada, the Czech Republic, France, the Netherlands, Romania, Sweden, Iceland, Australia and nine US locations in California, Georgia, Illinois, Kansas, North Carolina, Texas and Washington. That table describes the global networking scope rather than the whole fleet, so it is verified but probably incomplete. The data processing agreement states the services run on AWS infrastructure and on Secure Cloud data centres, platforms certified to ISO 27001 and SOC 2 Type 2, with encryption in transit and at rest. Documentation explicitly invites customers to pick a data centre on data residency grounds, and serverless deployments can be restricted to data centres meeting a given standard such as HIPAA. Transfers outside the EU are covered by the 2021/914 standard contractual clauses, the UK addendum and, for Brazil, ANPD clauses.
Where Runpod works
Country-level availability.
Not available in
Things to keep in mind
Risks and trade-offs to weigh before adopting Runpod.
- Compliance is certified at platform level but coverage varies by data centre, region, provider and deployment model, so it is on the customer to filter data centres rather than assume residency is handled
- The terms reserve to Runpod a licence to use customer content in aggregated and anonymised form to improve the service, which is easy to skim past when accepting terms quickly
- Nothing is published about whether models are trained on customer data, and no opt-out exists, so a team with sensitive workloads has no documented position to rely on
- Costs can run away quietly: metering by the second feels harmless, but forgotten Pods, idle network volumes and auto-pay together turn inattention into a bill, and spending limits must be set deliberately
- Data loss is a real operational risk rather than a theoretical one, since spot instances can be evicted mid-job and network volumes may be terminated irrecoverably if the balance stays at zero
- Published prices are explicitly mobile, as Runpod states its policy is to move prices to keep GPUs available, so any cost model built on today's rates should be revisited
- The terms include an arbitration clause with a class-action waiver, opt-out available within thirty days, and exclude embargoed countries; the minimum age is 18
Setup & Integrations
Technical difficulty
Difficulty depends on which product you pick, and it ranges from trivial to demanding. Public Endpoints require no infrastructure at all: set parameters in the playground and copy an OpenAI-compatible API call. A Pod is only slightly harder, since a supplied base image with a common ML stack can be running in under thirty seconds. Serverless is the demanding path, requiring a handler function, a Dockerfile, an image pushed to a registry and a deployed endpoint. Overall this is a developer product: comfort with Docker and the command line is assumed, and there is no no-code route.
Deployment
Integrations
Supported languages
Behind Runpod
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement Runpod.
Frequently asked questions
What exactly does Runpod provide?
Is there a free plan or a free trial?
What is the cheapest GPU available, and how is it billed?
Which GPUs and frameworks are supported?
Where is data hosted, and can I choose the region?
How does Runpod handle GDPR?
Does Runpod train models on customer data?
What certifications does Runpod hold?
What are the main caveats before committing?
Should you pick Runpod?
Runpod is a mature, unusually well-documented piece of AI infrastructure, and it knows exactly who it is for. Its central strength is the ratio of granularity to price: rates published GPU by GPU, metering by the second rather than the hour, no egress fees, no minimums and no contract on self-serve usage. A team can put an H100 to work in under thirty seconds without a sales call, then keep the same account, the same containers and the same endpoints all the way to a reserved enterprise agreement with a contractual 99.99% uptime SLA. That continuity is rarer than it sounds, and it removes the replatforming step that usually punishes early success.
The compliance posture is also more substantial than the marketing-page average: a completed SOC 2 Type 2, a data processing agreement published in full rather than promised on request, nine named sub-processors, standard contractual clauses and a designated Article 27 representative. Two qualifications matter. First, that compliance is delivered at platform level, while data centre coverage varies by region and deployment model, so residency is something you configure and verify rather than inherit. Second, the underlying reports require an approved access request.
The main reservation is commercial rather than technical. There is no free plan and no free trial, so evaluation begins by funding a prepaid account. Add spot eviction, network volumes that keep billing while Pods are stopped, quote-only reserved capacity, and a company that states plainly it moves prices to keep GPUs available, and the picture is of a platform that rewards teams who monitor their own consumption.
For machine learning engineers, MLOps teams, researchers and startups who can build a container, it is a strong and honestly priced option. For anyone expecting a finished AI product, it is the wrong layer of the stack.
- Choosing a selection results in a full page refresh.
- Opens in a new window.