
ZeroGPU
ZeroGPU is an edge inference cloud that runs small specialised and open-weight models behind an OpenAI-compatible API. It targets high-volume repeatable workloads such as classification, moderation, extraction and embeddings, with vendor-claimed cost savings of 50-70%.
What is ZeroGPU?
ZeroGPU describes itself as the edge inference cloud. Rather than sending every request to a centralised datacentre GPU, it routes inference to whichever compute is most efficient at that moment: consumer edge devices, edge servers, or the cloud as a fallback. It is built by ZeroGPU, Inc., a Texas corporation with offices in Austin, founded and led by Maddy Arvapally, whose background spans GoPro streaming, Replay and robotics, supported by four named advisors. Two families of models sit behind the API. The proprietary ZLM models, or ZeroGPU Language Models, run from 86M to 149M parameters and handle production tasks. Alongside them, a hosted open-weight catalogue covers Qwen, DeepSeek, Kimi, GLM, GPT-OSS, Llama, LiquidAI, GLiNER and DeBERTa. In total the catalogue lists 18 models, from 22.7M to 753B parameters, with context windows from 100 to 1,048,576 tokens. Access runs through an OpenAI-compatible endpoint at api.zerogpu.ai/v1, with official SDKs for Go, JavaScript, Python, Ruby and Rust, so an existing OpenAI client can often be redirected by changing its base URL. Ready-made integrations include LangChain, an MCP server, a Claude Code plugin, a Claude skill, Hermes Agent and OpenClaw. The vendor claims more than 100,000 edge devices, 50-70% lower inference cost and up to ten times faster responses, while noting that performance varies by workload, model, and configuration. It also publishes a field result: 1.04 million production requests served across 5,078 consumer devices in more than 20 countries over 63 hours, unsupervised. Edge capacity is announced in eleven cities, among them San Francisco, London, Frankfurt, Dubai, Mumbai, Tokyo and Sydney. The platform works in both directions. On the supply side, an SDK embeds into Android apps, games, Telegram mini-apps and Chrome extensions, turning idle devices into paid inference nodes under battery, thermal and Wi-Fi guardrails. A separate storefront lets autonomous AI agents buy capacity through x402, MPP or card payment without any human sign-up. The footer carries a Patents Pending notice and a 2026 copyright.
What it does
- Route repeatable inference workloads to small models through an OpenAI-compatible API
- Classify content against the IAB 1.0 and 2.2 taxonomies, audience segments and intent
- Moderate text across 13 OpenAI categories, returning a verdict with calibrated scores
- Detect and mask personally identifiable information across 40+ entity types in six native languages
- Summarise, extract and translate text, and generate embeddings
- Run inference on edge devices with automatic cloud fallback
- Compare estimated monthly token costs against OpenAI, Anthropic and Google
When to use ZeroGPU / When not to
A quick filter to help you decide if ZeroGPU is the right fit.
When to use ZeroGPU
- Engineering teams sending millions of repeatable classification, moderation or extraction calls to a frontier model
- Adtech and audience platforms that need IAB classification and intent signals in real time
- Publishers and content platforms running moderation, brand safety and categorisation at scale
- Developers already building on the OpenAI SDK, who can switch by changing the base URL
- Teams shipping autonomous AI agents that buy their own inference, without a human sign-up
When not to use ZeroGPU
- Teams looking for complex reasoning, which the vendor itself says belongs to frontier models
- Non-developers, as there is no consumer interface beyond an API and a technical dashboard
- Organisations bound to strict European hosting, since the databases sit in the United States
- Anyone under 18, who is excluded by both the terms of service and the privacy policy
- Workloads built on sensitive data such as health or biometric records, which need prior written approval
How to use ZeroGPU
A typical end-to-end flow, from setup to results.
- Create an account on the ZeroGPU platform through the Start building button
- Collect the three values you need: an API key, an optional project ID and a model name
- Pick a model in the Model Catalog according to the task and its price per million tokens
- Point an existing OpenAI client at api.zerogpu.ai/v1, or call POST /v1/responses directly
- Send the x-api-key and x-project-id headers, with model and input (or messages) in the body
- Or start from a ready-made integration: LangChain, MCP Server, Claude Code Plugin, Claude Skill, Hermes Agent or OpenClaw
- Track usage, latency and cost per request and per model from the dashboard
- For an autonomous agent, buy a plan on the agent storefront and call through the proxy, with no human account
Pros & Cons
Pros
- Very low prices on the specialised models, from $0.02 per million input tokens
- OpenAI compatibility, so integration requires no application rewrite
- Benchmarks published with figures rather than promises alone, such as 0.899 against 0.853 F1 on moderation
- Mixed catalogue spanning nano models and open-weight models up to 753B parameters
- A documented customer case with before and after latency figures (Dappier)
- A storefront for autonomous agents, several payment rails and a $0.01 minimum purchase
- Thorough technical documentation with code examples in six languages
Cons
- No pricing page on the main site: prices live in the documentation and on the agent storefront
- No readable free offer for human customers; the "free allowance" mentioned for agents is contradicted by includedUnits = 0 in the storefront manifest
- Databases in the United States, no European hosting option and no Article 27 representative
- No published certification, neither SOC 2 nor ISO 27001
- A very young company: domain registered in September 2025, first archived capture in January 2026
- Inference may run on third-party consumer devices, which means trusting the attestation and verification layer
- Specialised models do not replace a frontier model on complex reasoning, as the vendor states itself
Pricing & Plans
ZeroGPU publishes neither a free plan nor a free trial. Pricing is usage-based, billed per token and bought as credits. The floor of the catalogue is $0.02 per million input tokens and $0.05 per million output tokens, applying to the ZLM, GLiNER, DeBERTa, LFM2.5 and llama-3.1-8b-instruct-fast models. The top of the catalogue, glm-5.2, is priced at $1.10 per million input tokens and $3.50 per million output tokens. Intermediate examples include gpt-oss-120b at $0.15 and $0.60, deepseek-v4-flash at $0.20 and $0.40, qwen3-30b-a3b-fp8 at $0.05 and $0.30, and t5-small at $0.05 and $0.40. Embeddings are billed at $0.50 per million input tokens, with output not charged. On the agent storefront the minimum purchase is $0.01, or $0.50 when paying by card, with no subscription. No monthly or per-seat price is published. All amounts are quoted in US dollars.
- the single named plan on the agent storefront
- with credit billing
- no interval and a $0 base price
- an account with organisations
- projects and API keys
- billed on usage
- with no named tier published
- a monetisation programme for edge operators
- subject to approval
- with no pricing tier
Data, GDPR & hosting
A consolidated view of how ZeroGPU handles your data.
GDPR overview
The privacy policy, last revised on 28 February 2026, carries a dedicated section on European Union, United Kingdom and Swiss data subject rights. It enumerates access, rectification, erasure, withdrawal of consent, portability, objection, restriction of processing and the right to lodge a complaint with a supervisory authority, with a link to the EDPB. Rights are exercised through legal@zerogpu.ai. ZeroGPU presents itself as a controller for the data processed through the service, and as a joint controller alongside the generative AI providers it calls. Several gaps remain visible: no Article 27 representative in the European Union is named, no data protection officer is appointed, and transfers to the United States rest on user consent rather than on standard contractual clauses. ZeroGPU never claims GDPR compliance for itself: the phrase "GDPR, HIPAA, and CCPA compliant" appears only on the PII redaction use-case card.
Who owns the data?
The terms of service state that ZeroGPU and its licensors own all right, title and interest in the Services, excluding your Data, which therefore remains the customer's. The company does claim ownership of Aggregated Data, on the condition that it neither identifies the customer nor its users, and it undertakes not to attempt to re-identify any de-identified data. ZeroGPU describes itself as a data controller when it passes content on to third-party AI providers, and as a data processor when it handles the personal data of its customers' own end users. No standard data processing agreement is published; the terms only foresee executing one where the law requires it.
Reuse rights
Because submitted Data stays the customer's own under the ownership clause, nothing in the terms makes reuse of that data, or of the model output, conditional on ZeroGPU's permission. On the vendor's side, data serves to deliver the service, routing and running inference workloads, while de-identified and aggregated data may be used to operate, maintain and improve ZeroGPU's services and other products. Requests may be passed to third-party generative AI providers, with OpenAI and Anthropic named in the privacy policy; ZeroGPU states it does not authorise those providers to train any generative AI model on personal information. Trackers in use include Google Analytics, Google Tag Manager, Brevo, Cloudflare, Amazon and Supabase. On the supply side, the device SDK collects no personal data, requests no access to photos, contacts, files, camera, microphone or location, and does not retain payloads once a job has run.
Data retention & training
Hosting summary
The privacy policy states plainly that ZeroGPU's databases are currently located in the United States. It adds that data may be stored and processed in the user's own region, in the United States, and in any country where ZeroGPU or its service providers operate facilities. Named providers for storage and security include Cloudflare, Grafana, Brevo, Lovable and Google Cloud Platform. The company claims that all data is encrypted during transfer and while at rest. Inference itself is a separate matter: it can execute on third-party edge devices, and the vendor states that payloads are not retained on the device once the job has run. The website sits behind Cloudflare's anycast network, so the country attached to the web server's IP address says nothing about where the data actually lives. No European hosting region is offered, and no region beyond the United States is named.
Things to keep in mind
Risks and trade-offs to weigh before adopting ZeroGPU.
- Prices sit outside the main site and the agent storefront explicitly asks callers not to cache them, so a budget can drift without warning
- The performance figures are the vendor's own, published with the caveat that performance varies by workload, model and configuration
- Requests may travel to third-party generative AI providers and execute on third-party consumer devices
- Databases are in the United States, users outside it are asked to consent explicitly to that transfer, and ZeroGPU states it makes no claim that its Services are appropriate or lawful for use outside the United States
- Sensitive data such as health or biometric records may not be submitted without prior written approval
- The terms impose binding arbitration, waive class actions and apply Texas law in Travis County
- No client-side training opt-out is documented, and the company has barely a year of public existence behind it
Setup & Integrations
Technical difficulty
Low for developers, and out of reach for anyone else. There is no non-technical interface: you need an API key, an optional project ID and a model name, then a single HTTP call. Because the endpoint is OpenAI-compatible, changing a base URL is often enough, and official SDKs cover Go, JavaScript, Python, Ruby and Rust, with LangChain and MCP integrations available. The vendor's FAQ says most teams integrate in under an hour. The real effort is not the plumbing but the judgement: choosing the right model per task and calibrating the routing remain the team's responsibility.
Deployment
Integrations
Supported languages
Behind ZeroGPU
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement ZeroGPU.
Frequently asked questions
What is edge AI inference?
What is a small language model?
Does ZeroGPU replace frontier LLMs?
Do I have to change my code?
How much does ZeroGPU cost?
Which workloads suit ZeroGPU?
Where is the data hosted?
Is my data used to train models?
Is a data processing agreement available?
Is there a minimum age?
Should you pick ZeroGPU?
ZeroGPU answers a real and expensive problem: the cost of the repeatable inference calls that sit inside most AI applications. Its proposition is clear and quantified, with small specialised models and hosted open-weight models run at the edge and billed from $0.02 per million input tokens. Because the API is OpenAI-compatible, the friction of trying it is unusually low, since in many cases a new base URL and a model name are enough. What sets the company apart at this stage is that it shows its work. Benchmarks are published model by model, with figures such as 0.899 against 0.853 F1 on moderation, and a customer case with Dappier reports latency falling from roughly 1,800-2,000 ms to under 100 ms. That is rare for a company this young, and it makes the claims testable rather than merely asserted. The reservations are just as concrete. ZeroGPU, Inc. is very new: the domain was registered in September 2025 and the first archived capture dates from January 2026. No security certification is published, neither SOC 2 nor ISO 27001. The databases sit in the United States, with no European hosting option and no Article 27 representative. Prices live in the documentation and on the agent storefront rather than on the main site, and there is no readable free plan or free trial, the free allowance mentioned for agents being contradicted by the storefront's own manifest. The headline performance figures are the vendor's own, and carry its caveat that performance varies by workload, model and configuration. The sensible approach is therefore empirical: run a representative workload through the catalogue, measure quality, latency and cost against your current provider, and keep frontier models for the reasoning ZeroGPU itself says it was not built for.
- Choosing a selection results in a full page refresh.
- Opens in a new window.