fal logo
Api Tools · Inference Hosting

fal

fal is a generative media platform built for developers: a single unified API reaches over 1,000 image, video, audio, speech and 3D models, billed by usage, with a free entry tier and infrastructure for deploying your own models.

Active Free plan Usage Based API available 18+ Verified by Guidaio
Overview

What is fal?

fal is a generative media platform aimed squarely at developers. It is published by Features & Labels, Inc., a San Francisco company founded in 2021 by Burkay Gur and Gorkem Yurtseven, both previously at Coinbase and Amazon. The original ambition was to scale compute for Python; the pivot towards AI-generated media came at the end of 2022, once diffusion models proved compelling but painfully slow.

The platform is organised around three product lines. Model APIs expose a catalogue of more than 1,000 endpoints behind one interface, so moving from FLUX to Nano Banana, Kling, Veo, Seedance, Seedream, Wan, LTX, PixVerse, Qwen, Ideogram or Whisper means changing a single string rather than onboarding a new provider. Models come from Black Forest Labs, Google, OpenAI, xAI, Alibaba, Kling, ByteDance, ElevenLabs and others, and fal advertises day-zero availability when a new model ships. Serverless lets teams deploy their own or fine-tuned models on the same engine that runs the public catalogue, scaling from zero to thousands of GPUs with a multi-layer cache that blunts cold starts. Compute rents dedicated GPU instances, H100 SXM singly or eight at a time over InfiniBand, with full SSH access for training and other sustained workloads.

Speed is the stated differentiator: fal claims proprietary CUDA kernels, optimised attention layers and inference-engine tuning, alongside 99.99% historical uptime and billions of requests a day. Calls can be synchronous, queued asynchronously, streamed, or held open over WebSocket for real-time work. Python and JavaScript SDKs sit beside plain cURL, a genmedia command-line tool and an MCP server.

Around that core sit a Sandbox for side-by-side model comparison with cost estimates, no-code Workflows and an Agent, free browser tools such as background removal and upscaling, and a Platform API covering model metadata, pricing, usage, logs and metrics. Enterprise contracts add single sign-on, private endpoints, reserved capacity, 24/7 priority support and SOC 2. A dedicated Trust and Safety function, led by a former TikTok executive, runs automated content moderation across the platform.

What it does

  • Call over 1,000 generative media models through a single unified API
  • Generate images, video, audio, music, speech and 3D assets
  • Compare several models side by side on the same prompt before committing
  • Deploy your own or fine-tuned models on autoscaling GPU infrastructure
  • Rent dedicated GPU instances with full SSH access for training
  • Train LoRA adapters and fine-tuned variants from the platform
  • Control retention and access rights of generated files through HTTP headers
Audience

When to use fal / When not to

A quick filter to help you decide if fal is the right fit.

When to use fal

  • Product engineers embedding image or video generation into an existing application
  • Developers who want to compare and switch between competing models without integrating each provider separately
  • Startups that need to scale from zero to thousands of GPUs without running their own infrastructure
  • Machine learning teams deploying their own or fine-tuned models on managed, autoscaling GPU infrastructure
  • Latency-sensitive workloads where raw inference speed is the deciding factor

When not to use fal

  • Non-technical creators looking for a ready-made editing interface rather than an API
  • Teams that need European data residency, since personal data is processed on United States servers
  • Organisations requiring a documented opt-out from model training outside an enterprise contract
  • Projects with a fixed, capped budget, because billing is strictly usage-based
  • Anyone under 18, who is not permitted to use the service
Get started

How to use fal

A typical end-to-end flow, from setup to results.

  1. Create an account on fal.ai, optionally signing in with GitHub
  2. Generate an API key from the dashboard
  3. Browse the marketplace and pick a model that fits the task
  4. Install the Python or JavaScript client, or call the endpoint directly with cURL
  5. Send a first request, passing the model identifier and a prompt
  6. Use the Sandbox to run the same prompt across several models and compare output and cost
  7. Route browser traffic through a server-side proxy so the API key is never exposed
  8. Set retention and access-control headers on any request whose output must not stay public
  9. To host your own model, write a fal.App class declaring machine type, setup and endpoints
  10. Test it with fal run, publish with fal deploy, then tune the concurrency parameters
Quick read

Pros & Cons

Pros

  • A single integration replaces separate contracts and SDKs for hundreds of competing models
  • Day-zero availability of newly released models
  • Inference speed is the company's core investment, backed by proprietary CUDA optimisations
  • Cold starts are not billed on Model APIs, and server errors are never charged
  • The Sandbox allows genuine comparison of models before committing to one
  • Retention and access rights are controllable per request through HTTP headers
  • Dense, versioned documentation, including an llms.txt index aimed at agents

Cons

  • Strictly usage-based billing makes total cost hard to cap in advance
  • Generated media URLs are public by default until they expire
  • New accounts start at only two concurrent requests, capped at forty without a contract
  • The account locks automatically once the credit balance falls below its threshold
  • Purchased credits expire 365 days after purchase
  • Data is hosted in the United States, with no documented European residency option
  • No training opt-out is documented outside an enterprise contract, and no named subprocessor list is published
Pricing

Pricing & Plans

A permanent free entry tier is available, and fal grants 10 USD of credits when a payment method is added. Beyond that, billing is strictly usage-based rather than subscription-based. The lowest published unit price is 0.02 USD per megapixel of generated image, on the Qwen model; per-image models start at 0.03 USD and video models at 0.05 USD per second. Dedicated GPU capacity starts at 1.10 USD per hour for an RTX PRO 6000 and 1.89 USD per hour for an H100. All prices are quoted in United States dollars.

Plan 1
  • Free entry tier
  • with 10 USD of credits granted when a payment method is added
Plan 3
  • Serverless
  • billed per second of execution
  • scaling automatically from zero
Plan 4
  • Compute
  • dedicated GPU instances billed at a fixed hourly rate from 1.10 USD per hour
Plan 5
  • Enterprise
  • a contractual tier adding single sign-on
  • private endpoints
  • analytics
  • 24/7 priority support
  • reserved capacity
  • SLA
  • SOC 2 and invoice-based billing
Special offers — 10 USD of credits granted when a payment method is added to the account · Grants programme announced on the site · Startup programme announced on the site · Volume discounts negotiable with the sales team for higher-volume customers · Reduced GPU hourly rates advertised as as low as, under conditions that are not published
Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how fal handles your data.

GDPR overview

fal addresses the GDPR without ever claiming compliance in so many words. The privacy policy, last updated on 22 July 2026, carries a dedicated section for people located in Europe. It offers the right to access personal information in a portable format, to request deletion and to request correction, each exercised by emailing support@fal.ai. It acknowledges the right to lodge a complaint with the data protection authority of the country of residence and provides contact details for European authorities. International transfers are covered by appropriate safeguards including contractual clauses, and the Global Privacy Control signal is honoured. Two gaps matter: no Article 27 representative in the European Union is named anywhere in the legal pages, and no named list of subprocessors is published, only categories of service providers.

Who owns the data?

Under the terms of service, the customer owns and retains all right, title and interest in their Customer Input. In exchange, fal receives a non-exclusive, non-sublicensable, royalty-free licence to reproduce, store, display, adapt, translate, modify, create derivative works from and otherwise process that input, for the sole purpose of delivering the service. Separately, fal may generate, collect, store, use, transfer and disclose Usage Data to third parties. For enterprise customers, fal acts as a processor on the customer's behalf under a dedicated contract, and that data sits outside the standard privacy policy rather than under it.

Reuse rights

Customers may reuse what they generate without asking fal for permission, but the licence attached to each individual model governs commercial use: models carrying a Commercial use badge may be used in commercial projects, while those marked Research only may not. Checking the badge on the model page before any professional use is therefore necessary. On fal's side, personal information is used to provide, maintain and improve the services, to run analytics and to communicate with users, with legitimate interest cited as the legal basis in Europe. Cookies, pixels and session replay support product improvement, personalisation and marketing. The terms also allow Usage Data to feed the design and development of fal's own products, services and AI models. Enterprise customer data is explicitly carved out: fal states it never trains its large language models on it.

Data retention & training

Retention summary
Generated media and uploaded input files are kept on the fal CDN for at least seven days by default. That period can be changed per request through the X-Fal-Object-Lifecycle-Preference header, which also accepts no expiry at all. The JSON payloads of requests and responses are kept for thirty days, which is what powers the dashboard history; sending X-Fal-Store-IO set to zero prevents them being stored at all. Both payloads and output files can be deleted on demand through the Platform API. Files written to the persistent /data volume are kept indefinitely and managed manually. Expired files are permanently deleted and cannot be recovered, so anything worth keeping must be downloaded first.
Trains on customer data
Unclear
DPA available
Yes
GDPR contact

Hosting summary

fal is established in the United States and states that it processes and stores personal information on servers located in the United States and in other countries, which it does not name. Generated media is served from the fal CDN under the v3b.fal.media domain. The website itself resolves to an anycast address operated by Amazon. Where international transfers are restricted, fal says it puts appropriate safeguards in place, including contractual clauses. No European hosting region is documented, and no data residency option is described outside an enterprise contract, under which fal acts as a processor on the customer's behalf. Organisations with jurisdictional constraints should treat United States processing as the working assumption and raise residency explicitly during contract negotiation.

Hosting countries
🇺🇸 United States
Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting fal.

  • Generated media URLs are public by default, so anything sensitive is exposed to whoever holds the link until it expires
  • Media is deleted after seven days and request payloads after thirty, so anything worth keeping must be downloaded deliberately
  • Model licences differ: using a Research only model in a commercial project is a licensing breach that is easy to commit by accident
  • Usage-based billing rewards experimentation and punishes inattention, and an unattended loop can consume credits quickly
  • Outside an enterprise contract there is no documented way to exclude your data from feeding fal's own model development
  • Automated moderation, built on NCMEC, Thorn, StopNCII and OpenAI's Omni moderation API, can reject legitimate requests with a non-retryable error
  • Delegating creative output to a model catalogue can erode a team's own craft judgement if results are shipped without review
Setup

Setup & Integrations

Technical difficulty

Low for a first call, moderate for production. A working request takes three lines of Python or JavaScript once an account and API key exist. Calling from a browser requires building a server-side proxy so the key is never exposed. Deploying your own model is a further step: a Python class declaring machine type, a setup method and endpoints, then fal run to test and fal deploy to publish, with three concurrency parameters to tune. Free browser tools are usable without any code, but the platform is explicitly aimed at developers.

Deployment

Web appAPI

Integrations

Google Cloud Marketplace
Company

Behind fal

Company name
Features & Labels, Inc.
Founded
26/01/2021
Country of origin
🇺🇸 United States
Headquarters
2261 Market St. Suite 10467, San Francisco, CA 94114
UBO
INFORMATION_NOT_FOUND
UBO country
INFORMATION_NOT_FOUND
Domain registrar country
🇺🇸 United States
Legal contact
Support contact

Fundraising

Seed and Series A totalling 23 million USD, including a 14 million USD Series A led by Kindred Ventures with Andreessen Horowitz, First Round Capital and angel investors including Perplexity founder Aravind Srinivas
Series B of 49 million USD led by Notable Capital and Andreessen Horowitz, with participation from Bessemer Venture Partners, Kindred Ventures and First Round Capital
Series C of 125 million USD announced on 31 July 2025, led by Meritech with new participation from Salesforce Ventures, Shopify Ventures and the Google AI Futures Fund, joined by existing backers Bessemer Venture Partners, Kindred Ventures, Andreessen Horowitz, Notable Capital, First Round Capital, Unusual Ventures and Village Global

Social

Official links

Resources

All the official URLs gathered for verification and reference.

FAQ

Frequently asked questions

How is fal billed?
Strictly by usage, with no compulsory subscription. Model APIs are billed per output unit (per image, per megapixel, per second of video or per video), Serverless per second of execution, and Compute at a fixed hourly GPU rate. A free entry tier exists, and 10 USD of credits is granted when a payment method is added.
Are failed requests charged?
Server errors of HTTP 500 and above are never charged. Client-side errors such as invalid inputs returning HTTP 422 may still be charged if a runner had already spent GPU time on the request before the error was detected.
How long are generated files kept?
Generated media is stored on the fal CDN and available for at least 7 days by default. The retention period can be set per request with the X-Fal-Object-Lifecycle-Preference header, including no expiry at all. Expired files are permanently deleted and cannot be recovered.
Are the URLs of generated files public?
Yes, by default. Anyone holding the URL can open the file until it expires. Access can be restricted by setting an access-control list on the request, making files private to the account, shared with named users, or hidden.
Can the outputs be used commercially?
It depends on the individual model. Most models on fal carry a Commercial use badge and may be used in commercial projects, while models marked Research only are restricted to non-commercial use. The licence is shown on each model page.
Does fal train its models on customer data?
fal states that it never trains its large language models on enterprise customers' data. For standard accounts the position is less clear: the terms of service allow Usage Data to be used to design, develop and offer fal's own products, services and AI models.
Where is the data hosted?
fal is based in the United States and processes and stores personal information on servers located in the United States and other, unnamed countries. No European hosting region is documented, and international transfers rely on contractual safeguards.
Is there a minimum age?
Yes. Users must be 18 or the age of legal majority where they live. Organisations creating a team account are responsible for ensuring their end users meet the same threshold, and parents can contact safety@fal.ai about underage use.
Is there a rate limit?
Every account has a concurrency limit rather than a request quota. New accounts start at 2 concurrent requests, rising automatically to 40 as credits are purchased. Requests beyond the limit are queued. Higher limits require a sales contract.
Can I deploy my own models?
Yes, through fal Serverless. Any Python framework can be used, custom Docker images are supported, and deployments autoscale from zero to thousands of GPUs. Each deployment creates a revision, so rollbacks are immediate.
Conclusion

Should you pick fal?

fal occupies a precise position: it is infrastructure, not a creative tool. Anyone hoping for a polished editor will be in the wrong place, but a team that needs to put image, video, audio or 3D generation into a product will find one of the shortest paths available. The single strongest argument is consolidation: one integration, one key and one billing relationship covering more than a thousand models from a dozen competing labs, with new releases available on day zero. The Sandbox turns model selection into an evidence-based decision rather than a guess, and the fact that fal's own catalogue runs on the very infrastructure it sells is a meaningful alignment.

The reservations are real and worth weighing. Governance of data is uneven: enterprise customers get an explicit promise that their data will never train fal's models, while standard accounts get terms that permit Usage Data to inform the development of fal's own AI models, and no opt-out is documented for them. No named subprocessor list is published, no Article 27 representative is designated in the European Union, and hosting is in the United States with no documented European residency option. Teams under GDPR scrutiny should read the contract carefully before committing.

Operationally, two defaults deserve attention: generated media URLs are public unless an access-control list is set, and files disappear after seven days unless retention is extended. Costs scale with usage and cannot easily be capped. None of this disqualifies the platform; all of it should be configured deliberately rather than discovered later. Backed by 197 million USD across three rounds and growing quickly, fal is a credible long-term supplier for generative media at scale.