Featherless logo
Inference Hosting · Llm Providers

Featherless

Featherless is a serverless inference platform that puts more than 40,000 open-weight models behind a single OpenAI-compatible API key. Flat monthly subscriptions replace per-token billing, and the vendor stores neither prompts nor completions.

Active GDPR compliant Subscription API available 13+ Verified by Guidaio
Overview

What is Featherless?

Featherless is a serverless inference platform for open-weight AI models, published by Recursal AI, Inc. under the Featherless name and run out of San Francisco. Its bet is breadth rather than curation: instead of a short list of frontier models, it keeps an unusually large catalogue permanently online. The homepage advertises more than 40,000 models; the company page claims over 30,000 served on behalf of more than 6,000 model creators, across 19 countries of operation.

Access runs through a single API key on an interface deliberately made compatible with OpenAI's. A client written against the OpenAI SDK works after changing the base URL and the key — no rewrite, no separate library. The documented endpoints cover chat completions, plain completions, model listing, tokenisation and plan status, with tool calling, vision and embeddings described in their own guides. The catalogue is organised by intent rather than by vendor: Most Popular, Trending, Top Reasoning, Top Productivity, Top RP & Creative Writing, Top Small Models and Top Language Specific. Families served include DeepSeek, Qwen, Llama, Mistral, Gemma, GLM, Kimi, MiniMax, GPT-OSS and RWKV, alongside community fine-tunes and the uncensored or abliterated variants that mainstream providers rarely host.

The commercial model is the second differentiator. Featherless bills a flat monthly subscription with unlimited requests and a fixed number of concurrent units rather than metering every token, though credit-based and per-request plans exist for teams that prefer usage billing. A managed agent marketplace lets non-developers launch pre-configured applications on cloud sandboxes with inference already wired in, and a dedicated-GPU tier reserves H100, B200, B300, MI325X or RTX Pro 6000 capacity together with the engineering team that operates it.

The company's roots are in research: its founders co-lead RWKV, a Linux Foundation project on attention-free architectures. It launched as Recursal.AI, took the Featherless name in 2024, and raised a $20M Series A in 2026 co-led by AMD Ventures and Airbus Ventures. On privacy, the vendor states in three separate places that it stores neither prompts nor completions.

What it does

  • Call more than 40,000 open-weight models through one API key
  • Point an existing OpenAI client at Featherless by changing the base URL and the key
  • Run reasoning, coding, role-play, vision and embedding models from the same account
  • Launch a pre-configured AI agent on a managed cloud sandbox in seconds
  • Reserve dedicated H100, B200, B300, MI325X or RTX Pro 6000 capacity with an operations team included
  • Estimate the savings against a current provider with the public cost calculator
  • Send unlimited monthly requests within a fixed number of concurrent units
Audience

When to use Featherless / When not to

A quick filter to help you decide if Featherless is the right fit.

When to use Featherless

  • Developers who want to try many open-weight models without provisioning or paying for a single GPU
  • Engineering teams tired of unpredictable per-token bills and looking for a fixed monthly cost
  • Writers and role-play communities who need uncensored, abliterated or community fine-tuned models that mainstream providers do not host
  • Builders of coding agents who want Qwen3-Coder, DeepSeek-V4, MiniMax or GLM behind an OpenAI-compatible endpoint
  • Researchers and hobbyists hunting niche models — language-specific checkpoints, sub-2B models, experimental architectures

When not to use Featherless

  • Anyone who needs proprietary frontier models: no GPT, no Claude, no Gemini — the catalogue is open weights only
  • Buyers who want to evaluate before paying, since there is neither a free tier nor a free trial
  • Teams routing production API traffic through the $25 entry plan, which the terms restrict to interactive human use
  • Organisations that require published security certifications such as SOC 2, ISO 27001 or HIPAA
  • Users looking for a mobile app, a browser extension or a ready-made chat product rather than an inference layer
Get started

How to use Featherless

A typical end-to-end flow, from setup to results.

  1. Create an account on the sign-up page — access to models requires a paid subscription, so pick a plan first
  2. Choose a tier: Chat at $25 a month for interactive use, developer at $50 a month for production workloads
  3. Generate an API key from the API keys section of your account
  4. Browse the model library and filter by family, capability, modality or style to pick a checkpoint
  5. Point your existing OpenAI client at https://api.featherless.ai/v1 and swap in the Featherless key
  6. Set the HTTP-Referer and X-Title headers so the vendor can identify your application and support it faster
  7. Follow the quickstart guide or the open-source cookbook on GitHub for a first working call
  8. For a no-code route, open the agents section of your account and launch a pre-built app from the marketplace
  9. Connect a third-party client instead of writing code: Cursor, Aider, Cline, n8n, Dify, LangChain, SillyTavern and others have dedicated guides
  10. For dedicated capacity, book a call with the team — GPU reservations are quoted, not self-served
Quick read

Pros & Cons

Pros

  • Catalogue breadth with no real equivalent among managed providers, including niche and community models
  • No infrastructure to run: no provisioning, no servers, no GPU hours to reconcile
  • Predictable flat monthly cost with unlimited requests instead of a token meter
  • Drop-in OpenAI compatibility makes migration a one-line configuration change
  • Prompts and completions are not stored, and the commitment is repeated in three separate documents
  • A full data processing addendum with Standard Contractual Clauses and fourteen named sub-processors
  • Research pedigree — the founders co-lead RWKV — and a $20M Series A backing the infrastructure

Cons

  • No free tier and no trial: evaluating the service costs $25 a month up front
  • No proprietary frontier models — open weights only, which rules out some benchmarks and workloads
  • The entry plan is contractually limited to interactive human use, excluding API traffic and benchmarking
  • Privacy policy and terms have not been revised since 10 June 2024, while the DPA is dated August 2026
  • The terms still describe the service as beta and reserve the right to de-list previously working models
  • No published security certification (SOC 2, ISO 27001, HIPAA) and no Article 27 EU representative
  • Support runs through email and Discord only: no contact page, no phone, no ticketing portal
Pricing

Pricing & Plans

There is no free plan and no free trial. The lowest paid entry point is USD 25 per month for the Chat tier, billed monthly and cancellable at any time, with the current month non-refundable. The developer tier costs USD 50 per month, credit-based plans start at USD 25, and dedicated GPU capacity is quoted individually under annual contract.

Plan 1
  • Chat — $25/month — context up to 32K
  • 4 concurrent units
  • unlimited tokens
  • interactive human use only (listed as Featherless Premium in the documentation)
Plan 3
  • Featherless Token-Based Business — from $25/month — credit-based monthly billing
  • 8 concurrent units
  • 256K context
  • one agent sandbox
Plan 4
  • Feather Per-Request — from $25/month — prepaid credits
  • pay per successful request on model and token usage
  • no model size limit
  • 100 concurrent units
Plan 5
  • business / Featherless GPU-Based Business — custom quote — dedicated H100
  • MI325
  • B200 and B300 GPUs
  • engineering team included
  • burst and failover to public cloud
  • annual contracts and volume pricing
Special offers — A public cost calculator estimates savings against your current provider based on spend and token volume · A promotional campaign highlights running GLM 5.2 on AMD hardware exclusively through Featherless, with its own landing page · The company claims to have sponsored more than 3,000 builders · No student, non-profit or annual-commitment discount is published, and no promotional code is offered
Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how Featherless handles your data.

GDPR overview

GDPR is never mentioned on any HTML page of the site — not in the privacy policy, not in the terms, not in the documentation. The commitment lives entirely in a data processing addendum published at the /legal/dpa route, which is linked from nowhere and redirects to a public PDF last updated on 7 August 2026. That document is substantial: it covers the GDPR, the UK GDPR, the e-Privacy Directive, the CCPA and thirteen further US state privacy laws, incorporates the European Commission's Standard Contractual Clauses (Decision 2021/914/EU), designates Ireland as competent supervisory authority, names a Data Privacy Officer reachable at dpo@featherless.ai, and lists fourteen sub-processors by name and address. No Article 27 EU representative is designated. The privacy policy, unchanged since 10 June 2024, simply states that personal data is not transferred internationally.

Who owns the data?

The terms of service are unusually clear on ownership. Everything a user uploads, publishes or displays on the platform, together with every output returned by a model, is defined as "Your Content" and remains the property of the user. Featherless keeps all rights to the platform itself — its software, technology and processes — and transfers no intellectual property to the customer. Crucially, it does not license the third-party models it hosts: the user alone is responsible for complying with each model's own licence and applicable law. Under the data processing addendum the customer acts as controller and Featherless as processor, which places the customer in charge of the personal data flowing through the service.

Reuse rights

Users may reuse their own content and model outputs without asking Featherless for permission: the terms grant ownership of "Your Content" to the user and reserve no licence over it for the vendor. Two limits apply and neither comes from Featherless itself. First, the licence of the underlying open-weight model governs what may be done with its outputs, and enforcing it is the user's responsibility. Second, the terms forbid presenting content generated on the platform as human-written. On the vendor's side, reuse is narrow by design: prompts and completions are neither collected nor stored, so only account data and aggregate usage metrics — request counts, input and output token volumes, and possibly sampler settings — are processed, for billing, account management and model discoverability.

Data retention & training

Retention summary
Prompts and completions are never retained: API requests are processed in real time and are not stored on the servers. What is kept is limited to account data — email address and payment details — plus usage metrics such as request counts and input and output token volumes, and possibly sampler settings used to guide new users. Session cookies are used. Under the data processing addendum, covered data is processed for the duration of the contract and for 30 days afterwards, unless the law requires otherwise; on expiry or termination, and at the customer's request, data is securely returned or destroyed. Where a legal obligation forces retention, the data is isolated and shielded from further processing until deletion becomes possible. Users can request deletion of their personal data by emailing support.
Trains on customer data
No
Subprocessors disclosed
Yes
DPA available
Yes
GDPR contact

Hosting summary

Featherless publishes hosting locations only for its dedicated GPU offering: infrastructure is available in the US, the EU and Southeast Asia, the customer chooses the region to match user location and data residency requirements, and multi-region deployments are supported. Nothing equivalent is stated for the shared serverless service, so its physical location is undocumented. The privacy policy adds that personal data is not transferred internationally. The infrastructure sub-processors named in the DPA give an indirect picture: Amazon Web Services, DigitalOcean, Hydra Host, TensorWave and Vercel for cloud hosting, Skylab Services Pte. Ltd. in Singapore, Cloudflare for non-persistent network transmission, MongoDB for account data persistence and Elasticsearch for operational metadata. Dedicated workloads can be isolated at VPC level, with prompts, completions and model weights kept out of any shared environment. The public site is served behind Cloudflare.

Hosting regions
USEUSoutheast Asia
Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting Featherless.

  • The $25 Chat plan is contractually restricted to interactive human use; reselling, app or API traffic, background automation and benchmarking can lead to cancellation without refund
  • The terms, unchanged since June 2024, still describe the service as beta and allow previously working models to be de-listed
  • The privacy policy and the data processing addendum are two years apart and are not aligned; the DPA is not linked from any page on the site
  • Uncensored and abliterated models are a headline feature, and the stated minimum age is 13 — a combination that deserves supervision in any shared or family context
  • Running third-party open-weight models makes the user, not Featherless, responsible for each model's licence and for lawful use
  • Flat pricing with unlimited requests can encourage volume over judgement: cheap tokens are still tokens someone has to read and verify
  • No published security certification and no Article 27 EU representative, which will slow or block procurement in regulated organisations
Setup

Setup & Integrations

Technical difficulty

Low for a developer. Create an account, subscribe, generate an API key and change the base URL of an existing OpenAI client — a few minutes, with no infrastructure to provision. A quickstart guide and an open-source cookbook cover the first call. Non-developers have a no-code route through the agent marketplace, where a pre-configured app launches on a managed sandbox in seconds. The real difficulty is not technical but editorial: choosing sensibly among more than 40,000 models, and sizing concurrent units for the workload. Dedicated GPU capacity is not self-served and requires a sales conversation.

Deployment

Web appAPI

Integrations

Hugging Face OpenAI LangChain LlamaIndex LiteLLM N8n Dify Cursor Aider Cline Roo Code Open WebUI SillyTavern WyvernChat HammerAI Venus AI JanitorAI Anime.gf Typing Mind OpenLIT Stripe

Supported languages

English
Company

Behind Featherless

Company name
Recursal AI, Inc.
Founded
31/05/2024
Country of origin
🇺🇸 United States
Headquarters
2261 Market Street STE 5498, San Francisco, CA 94114, United States
UBO
INFORMATION_NOT_FOUND
UBO country
INFORMATION_NOT_FOUND
Domain registrar country
🇺🇸 United States
Support contact

Fundraising

$20M Series A announced in 2026, co-led by AMD Ventures and Airbus Ventures
Participation from BMW i Ventures, Kickstart Ventures, Panache Ventures and Wavemaker Ventures
Stated purpose: scaling open-source AI infrastructure independent of any single vendor
The company was founded in 2023 as Recursal.AI and renamed Featherless in 2024

Social

Official links

Resources

All the official URLs gathered for verification and reference.

Compare

Alternatives

Tools that compete with or complement Featherless.

T Together AIR ReplicateF Fireworks AIO OpenRouterR RunPodH Hugging Face InferenceA AWS BedrockA Anthropic
FAQ

Frequently asked questions

What exactly does Featherless do?
It is a serverless LLM inference platform. You call open-weight models — more than 40,000 of them, drawn from the Hugging Face ecosystem — through a single API key, without provisioning or operating any GPU yourself.
Is there a free plan or a free trial?
Neither is advertised. The FAQ states that you create an account and subscribe to a plan to get started, so the lowest entry point is $25 per month.
How much does it cost?
The Chat plan is $25 per month, the developer plan $50 per month, and credit-based plans start at $25. Dedicated GPU capacity is quoted individually. Per-request billing publishes prices per million tokens by model class.
Can I reuse my existing OpenAI code?
Yes. The API is deliberately OpenAI-compatible: point the base URL at api.featherless.ai/v1 and substitute a Featherless key for your OpenAI key. No other change is required.
Are my prompts and conversations logged?
No. The privacy policy states that prompts and completions are neither collected nor stored, and the documentation confirms that API requests are processed in real time and not kept on the servers.
Which models can I run?
The stated goal is to serve every open-weight model on Hugging Face. Llama 2 and 3, Mistral, Qwen, DeepSeek, Gemma, GLM, Kimi, MiniMax, GPT-OSS and RWKV are all served, alongside vision, multimodal and embedding models.
How do I get a missing model added?
Business customers deploy models themselves from their dashboard. Users on individual plans request additions through the Featherless Discord community.
Is a data processing agreement available?
Yes. A global DPA last updated on 7 August 2026 is published at the /legal/dpa route. It incorporates the Standard Contractual Clauses and lists fourteen sub-processors by name, address and processing purpose.
Where is the infrastructure hosted, and what is the minimum age?
Dedicated GPU infrastructure is available in the US, the EU and Southeast Asia, with the region left to the customer; no location is published for the shared serverless offering. Use of the service by anyone under 13 is prohibited.
Conclusion

Should you pick Featherless?

Featherless answers a question most inference providers avoid: what if the catalogue were the product? Where competitors curate a dozen frontier models, Featherless keeps tens of thousands of open-weight checkpoints online — including the community fine-tunes, language-specific models and uncensored variants that no one else will host — and sells access to all of them for a flat monthly fee. For a developer who wants to test broadly, or a role-play and creative-writing community that lives on niche models, that combination is genuinely hard to replicate.

The engineering story is credible. The founders co-lead RWKV, a Linux Foundation project, and a $20M Series A led by AMD Ventures and Airbus Ventures funds the infrastructure. The API is OpenAI-compatible in the useful sense — change two lines and existing code runs. The privacy posture is stronger than most: prompts and completions are not stored, and a serious data processing addendum with Standard Contractual Clauses and fourteen named sub-processors sits behind the service.

The reservations are equally concrete. There is no free tier and no trial, so evaluation costs $25 before you learn anything. The $25 plan is contractually barred from carrying API traffic, which is a trap worth reading twice. The privacy policy and terms have not been touched since June 2024 and still call the service beta, while the DPA is dated August 2026 — and that DPA is linked from nowhere on the site, which is an odd way to treat the best document the company has written. No security certification is published.

In short: an excellent fit for teams that need model variety and cost predictability, and a poor one for buyers who need frontier proprietary models, a free evaluation path, or a compliance file they can hand to procurement without further questions.