Featherless
Featherless is a serverless inference platform that puts more than 40,000 open-weight models behind a single OpenAI-compatible API key. Flat monthly subscriptions replace per-token billing, and the vendor stores neither prompts nor completions.
What is Featherless?
Featherless is a serverless inference platform for open-weight AI models, published by Recursal AI, Inc. under the Featherless name and run out of San Francisco. Its bet is breadth rather than curation: instead of a short list of frontier models, it keeps an unusually large catalogue permanently online. The homepage advertises more than 40,000 models; the company page claims over 30,000 served on behalf of more than 6,000 model creators, across 19 countries of operation.
Access runs through a single API key on an interface deliberately made compatible with OpenAI's. A client written against the OpenAI SDK works after changing the base URL and the key — no rewrite, no separate library. The documented endpoints cover chat completions, plain completions, model listing, tokenisation and plan status, with tool calling, vision and embeddings described in their own guides. The catalogue is organised by intent rather than by vendor: Most Popular, Trending, Top Reasoning, Top Productivity, Top RP & Creative Writing, Top Small Models and Top Language Specific. Families served include DeepSeek, Qwen, Llama, Mistral, Gemma, GLM, Kimi, MiniMax, GPT-OSS and RWKV, alongside community fine-tunes and the uncensored or abliterated variants that mainstream providers rarely host.
The commercial model is the second differentiator. Featherless bills a flat monthly subscription with unlimited requests and a fixed number of concurrent units rather than metering every token, though credit-based and per-request plans exist for teams that prefer usage billing. A managed agent marketplace lets non-developers launch pre-configured applications on cloud sandboxes with inference already wired in, and a dedicated-GPU tier reserves H100, B200, B300, MI325X or RTX Pro 6000 capacity together with the engineering team that operates it.
The company's roots are in research: its founders co-lead RWKV, a Linux Foundation project on attention-free architectures. It launched as Recursal.AI, took the Featherless name in 2024, and raised a $20M Series A in 2026 co-led by AMD Ventures and Airbus Ventures. On privacy, the vendor states in three separate places that it stores neither prompts nor completions.
What it does
- Call more than 40,000 open-weight models through one API key
- Point an existing OpenAI client at Featherless by changing the base URL and the key
- Run reasoning, coding, role-play, vision and embedding models from the same account
- Launch a pre-configured AI agent on a managed cloud sandbox in seconds
- Reserve dedicated H100, B200, B300, MI325X or RTX Pro 6000 capacity with an operations team included
- Estimate the savings against a current provider with the public cost calculator
- Send unlimited monthly requests within a fixed number of concurrent units
When to use Featherless / When not to
A quick filter to help you decide if Featherless is the right fit.
When to use Featherless
- Developers who want to try many open-weight models without provisioning or paying for a single GPU
- Engineering teams tired of unpredictable per-token bills and looking for a fixed monthly cost
- Writers and role-play communities who need uncensored, abliterated or community fine-tuned models that mainstream providers do not host
- Builders of coding agents who want Qwen3-Coder, DeepSeek-V4, MiniMax or GLM behind an OpenAI-compatible endpoint
- Researchers and hobbyists hunting niche models — language-specific checkpoints, sub-2B models, experimental architectures
When not to use Featherless
- Anyone who needs proprietary frontier models: no GPT, no Claude, no Gemini — the catalogue is open weights only
- Buyers who want to evaluate before paying, since there is neither a free tier nor a free trial
- Teams routing production API traffic through the $25 entry plan, which the terms restrict to interactive human use
- Organisations that require published security certifications such as SOC 2, ISO 27001 or HIPAA
- Users looking for a mobile app, a browser extension or a ready-made chat product rather than an inference layer
How to use Featherless
A typical end-to-end flow, from setup to results.
- Create an account on the sign-up page — access to models requires a paid subscription, so pick a plan first
- Choose a tier: Chat at $25 a month for interactive use, developer at $50 a month for production workloads
- Generate an API key from the API keys section of your account
- Browse the model library and filter by family, capability, modality or style to pick a checkpoint
- Point your existing OpenAI client at https://api.featherless.ai/v1 and swap in the Featherless key
- Set the HTTP-Referer and X-Title headers so the vendor can identify your application and support it faster
- Follow the quickstart guide or the open-source cookbook on GitHub for a first working call
- For a no-code route, open the agents section of your account and launch a pre-built app from the marketplace
- Connect a third-party client instead of writing code: Cursor, Aider, Cline, n8n, Dify, LangChain, SillyTavern and others have dedicated guides
- For dedicated capacity, book a call with the team — GPU reservations are quoted, not self-served
Pros & Cons
Pros
- Catalogue breadth with no real equivalent among managed providers, including niche and community models
- No infrastructure to run: no provisioning, no servers, no GPU hours to reconcile
- Predictable flat monthly cost with unlimited requests instead of a token meter
- Drop-in OpenAI compatibility makes migration a one-line configuration change
- Prompts and completions are not stored, and the commitment is repeated in three separate documents
- A full data processing addendum with Standard Contractual Clauses and fourteen named sub-processors
- Research pedigree — the founders co-lead RWKV — and a $20M Series A backing the infrastructure
Cons
- No free tier and no trial: evaluating the service costs $25 a month up front
- No proprietary frontier models — open weights only, which rules out some benchmarks and workloads
- The entry plan is contractually limited to interactive human use, excluding API traffic and benchmarking
- Privacy policy and terms have not been revised since 10 June 2024, while the DPA is dated August 2026
- The terms still describe the service as beta and reserve the right to de-list previously working models
- No published security certification (SOC 2, ISO 27001, HIPAA) and no Article 27 EU representative
- Support runs through email and Discord only: no contact page, no phone, no ticketing portal
Pricing & Plans
There is no free plan and no free trial. The lowest paid entry point is USD 25 per month for the Chat tier, billed monthly and cancellable at any time, with the current month non-refundable. The developer tier costs USD 50 per month, credit-based plans start at USD 25, and dedicated GPU capacity is quoted individually under annual contract.
- Chat — $25/month — context up to 32K
- 4 concurrent units
- unlimited tokens
- interactive human use only (listed as Featherless Premium in the documentation)
- developer — $50/month — context up to 256K
- one agent environment included
- fastest response times
- unused credits roll over
- billed per token
- Featherless Token-Based Business — from $25/month — credit-based monthly billing
- 8 concurrent units
- 256K context
- one agent sandbox
- Feather Per-Request — from $25/month — prepaid credits
- pay per successful request on model and token usage
- no model size limit
- 100 concurrent units
- business / Featherless GPU-Based Business — custom quote — dedicated H100
- MI325
- B200 and B300 GPUs
- engineering team included
- burst and failover to public cloud
- annual contracts and volume pricing
Data, GDPR & hosting
A consolidated view of how Featherless handles your data.
GDPR overview
GDPR is never mentioned on any HTML page of the site — not in the privacy policy, not in the terms, not in the documentation. The commitment lives entirely in a data processing addendum published at the /legal/dpa route, which is linked from nowhere and redirects to a public PDF last updated on 7 August 2026. That document is substantial: it covers the GDPR, the UK GDPR, the e-Privacy Directive, the CCPA and thirteen further US state privacy laws, incorporates the European Commission's Standard Contractual Clauses (Decision 2021/914/EU), designates Ireland as competent supervisory authority, names a Data Privacy Officer reachable at dpo@featherless.ai, and lists fourteen sub-processors by name and address. No Article 27 EU representative is designated. The privacy policy, unchanged since 10 June 2024, simply states that personal data is not transferred internationally.
Who owns the data?
The terms of service are unusually clear on ownership. Everything a user uploads, publishes or displays on the platform, together with every output returned by a model, is defined as "Your Content" and remains the property of the user. Featherless keeps all rights to the platform itself — its software, technology and processes — and transfers no intellectual property to the customer. Crucially, it does not license the third-party models it hosts: the user alone is responsible for complying with each model's own licence and applicable law. Under the data processing addendum the customer acts as controller and Featherless as processor, which places the customer in charge of the personal data flowing through the service.
Reuse rights
Users may reuse their own content and model outputs without asking Featherless for permission: the terms grant ownership of "Your Content" to the user and reserve no licence over it for the vendor. Two limits apply and neither comes from Featherless itself. First, the licence of the underlying open-weight model governs what may be done with its outputs, and enforcing it is the user's responsibility. Second, the terms forbid presenting content generated on the platform as human-written. On the vendor's side, reuse is narrow by design: prompts and completions are neither collected nor stored, so only account data and aggregate usage metrics — request counts, input and output token volumes, and possibly sampler settings — are processed, for billing, account management and model discoverability.
Data retention & training
Hosting summary
Featherless publishes hosting locations only for its dedicated GPU offering: infrastructure is available in the US, the EU and Southeast Asia, the customer chooses the region to match user location and data residency requirements, and multi-region deployments are supported. Nothing equivalent is stated for the shared serverless service, so its physical location is undocumented. The privacy policy adds that personal data is not transferred internationally. The infrastructure sub-processors named in the DPA give an indirect picture: Amazon Web Services, DigitalOcean, Hydra Host, TensorWave and Vercel for cloud hosting, Skylab Services Pte. Ltd. in Singapore, Cloudflare for non-persistent network transmission, MongoDB for account data persistence and Elasticsearch for operational metadata. Dedicated workloads can be isolated at VPC level, with prompts, completions and model weights kept out of any shared environment. The public site is served behind Cloudflare.
Things to keep in mind
Risks and trade-offs to weigh before adopting Featherless.
- The $25 Chat plan is contractually restricted to interactive human use; reselling, app or API traffic, background automation and benchmarking can lead to cancellation without refund
- The terms, unchanged since June 2024, still describe the service as beta and allow previously working models to be de-listed
- The privacy policy and the data processing addendum are two years apart and are not aligned; the DPA is not linked from any page on the site
- Uncensored and abliterated models are a headline feature, and the stated minimum age is 13 — a combination that deserves supervision in any shared or family context
- Running third-party open-weight models makes the user, not Featherless, responsible for each model's licence and for lawful use
- Flat pricing with unlimited requests can encourage volume over judgement: cheap tokens are still tokens someone has to read and verify
- No published security certification and no Article 27 EU representative, which will slow or block procurement in regulated organisations
Setup & Integrations
Technical difficulty
Low for a developer. Create an account, subscribe, generate an API key and change the base URL of an existing OpenAI client — a few minutes, with no infrastructure to provision. A quickstart guide and an open-source cookbook cover the first call. Non-developers have a no-code route through the agent marketplace, where a pre-configured app launches on a managed sandbox in seconds. The real difficulty is not technical but editorial: choosing sensibly among more than 40,000 models, and sizing concurrent units for the workload. Dedicated GPU capacity is not self-served and requires a sales conversation.
Deployment
Integrations
Supported languages
Behind Featherless
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement Featherless.
Frequently asked questions
What exactly does Featherless do?
Is there a free plan or a free trial?
How much does it cost?
Can I reuse my existing OpenAI code?
Are my prompts and conversations logged?
Which models can I run?
How do I get a missing model added?
Is a data processing agreement available?
Where is the infrastructure hosted, and what is the minimum age?
Should you pick Featherless?
Featherless answers a question most inference providers avoid: what if the catalogue were the product? Where competitors curate a dozen frontier models, Featherless keeps tens of thousands of open-weight checkpoints online — including the community fine-tunes, language-specific models and uncensored variants that no one else will host — and sells access to all of them for a flat monthly fee. For a developer who wants to test broadly, or a role-play and creative-writing community that lives on niche models, that combination is genuinely hard to replicate.
The engineering story is credible. The founders co-lead RWKV, a Linux Foundation project, and a $20M Series A led by AMD Ventures and Airbus Ventures funds the infrastructure. The API is OpenAI-compatible in the useful sense — change two lines and existing code runs. The privacy posture is stronger than most: prompts and completions are not stored, and a serious data processing addendum with Standard Contractual Clauses and fourteen named sub-processors sits behind the service.
The reservations are equally concrete. There is no free tier and no trial, so evaluation costs $25 before you learn anything. The $25 plan is contractually barred from carrying API traffic, which is a trap worth reading twice. The privacy policy and terms have not been touched since June 2024 and still call the service beta, while the DPA is dated August 2026 — and that DPA is linked from nowhere on the site, which is an odd way to treat the best document the company has written. No security certification is published.
In short: an excellent fit for teams that need model variety and cost predictability, and a poor one for buyers who need frontier proprietary models, a free evaluation path, or a compliance file they can hand to procurement without further questions.
- Choosing a selection results in a full page refresh.
- Opens in a new window.