Scale AI logo
Data Labeling · Evaluation Benchmarks

Scale AI

Scale AI supplies expert-annotated training data, model evaluations and an enterprise agent platform. It serves foundation model labs, Fortune 500 companies and government agencies, with access through a sales demo or a pay-as-you-go self-serve tier.

Active GDPR compliant Contact Sales API available 18+ Verified by Guidaio
Overview

What is Scale AI?

Scale AI is an AI data and applications provider published by Scale AI, Inc., a Delaware company created on 3 May 2016 and headquartered at 650 Townsend Street, San Francisco. Its stated mission is to build reliable AI systems for the decisions that matter most, and its modern slavery statement describes two business segments: Data Infrastructure and AI Applications.

The first segment centres on the Scale Data Engine, a loop that collects, curates and annotates data, trains a model, evaluates it, then starts again. Scale argues its case on expert quality, annotation budget optimization, elastic volume and data diversity. The Generative AI Data Engine extends the loop to RLHF, data generation, model evaluation, safety and alignment, producing the complex prompt-response pairs that follow pre-training. Around it sit Nucleus for dataset curation and management, SEAL, an in-house frontier benchmarking lab that publishes private-benchmark leaderboards, and Outlier, Scale's proprietary contributor platform.

The second segment is the Scale GenAI Platform, organized in four steps: Connect Data, Build + Execute, Evaluate + Monitor, Learn + Improve. It is deliberately agnostic. It works with the tools, frameworks and models a customer already uses, supports OpenAI, Google, Meta and Mistral among others, and deploys inside the customer's own VPC on AWS, Azure or GCP, so nothing has to be migrated and no supplier has to be swapped out. Agentex and AgentOps are the agentic infrastructure layers behind long-running asynchronous agents and multi-agent coordination, while Dialect encodes an organization's expert judgement as a decision layer. Scale Donovan serves defense and the public sector, turning classified data into actionable intelligence.

The numbers claimed are unusual: 15 billion human decisions used to train models, close to 1 billion USD paid to hundreds of thousands of contributors across more than 150 countries, a 29 billion USD valuation and over 1,000 employees. The home page claims 90% of leading generative model builders are powered by Scale and that 25% of contributors hold a postgraduate degree, and names customers including CDAO, Meta, Mayo Clinic, TIME and British Petroleum. Access runs through a booked demo, a documented public API, or direct sign-up on the dashboard for self-serve work.

What it does

  • Have training data annotated, generated and curated by expert contributors
  • Apply RLHF and alignment work to a foundation model
  • Evaluate and compare models against private benchmarks through SEAL and its leaderboards
  • Run generative red teaming to probe how a model holds up under pressure
  • Connect enterprise sources such as Confluence, SharePoint and Amazon S3 and make them usable by agents
  • Build, deploy and orchestrate long-running agents inside your own VPC
  • Monitor agents in production with evaluation scoring, full traces and semantic monitoring
Audience

When to use Scale AI / When not to

A quick filter to help you decide if Scale AI is the right fit.

When to use Scale AI

  • ML teams that need large volumes of data annotated by domain experts; Scale states that 25% of its contributors hold a postgraduate degree.
  • Foundation model labs running RLHF, alignment, safety work and complex prompt-response generation.
  • Enterprises taking agent projects from proof of concept to production on the Scale GenAI Platform, deployed inside their own VPC.
  • Government, defense and intelligence teams that need FedRAMP High or DoD IL4 accreditation, including CDAO and Donovan users.
  • Regulated sectors such as healthcare, insurance and energy that require governance, traceability and a full audit trail.

When not to use Scale AI

  • Buyers who need published prices before contacting sales: no amount, currency or billing unit appears anywhere on the site.
  • Consumers and private individuals: the offering targets companies and public bodies, and the minimum age is 18.
  • Teams that want a mobile app, a browser extension or a non-English interface: Scale ships an English-language web dashboard and an API, nothing else.
  • Organizations that expect self-service support: assistance runs through a named Engagement Manager or Field Engineer, which presupposes a contract.
  • Teams reading the Self-Serve Data Engine as a production tier: Scale frames it as suited to experimental or research projects.
Get started

How to use Scale AI

A typical end-to-end flow, from setup to results.

  1. Decide which product line fits the need: Data Engine for training data, GenAI Platform for agents, Nucleus for dataset curation, Donovan for defense and intelligence work.
  2. Read the public documentation and the API reference to check key concepts, technical limits and the API compatibility policy.
  3. For an enterprise project, submit the Book a demo form for a one-to-one discovery call; commercial questions go to sales@scale.com.
  4. For a small project, sign up directly on the Scale dashboard: the Studio product for annotation, Nucleus for data management.
  5. Scope the use case with Scale during the discovery call and agree on what the engagement has to deliver.
  6. Connect enterprise data sources such as Confluence, SharePoint or Amazon S3.
  7. Deploy the GenAI Platform inside your own VPC on AWS, Azure or GCP, so no data has to leave your environment.
  8. Build and execute agents, testing and fine-tuning the models you already use rather than rebuilding around new ones.
  9. Evaluate and monitor them in production with evaluation scoring, semantic-layer monitoring and full traces.
  10. Feed human feedback back in as a continuous learning signal, and rely on the named Engagement Manager or Field Engineer for support, 9am to 5pm Pacific, Monday to Friday, excluding US holidays.
Quick read

Pros & Cons

Pros

  • Market reference point: Scale claims 90% of leading generative model builders as customers, and names CDAO, Meta, Mayo Clinic and TIME among them.
  • Rare accreditation level for an AI vendor: FedRAMP High Authorized and DoD IL4 Provisional Authorization on top of SOC 2 Type II and ISO/IEC 27001:2022.
  • Ownership of customer data, business logic and the AI solutions built for the customer is explicitly guaranteed, with no vendor lock-in.
  • Unusually transparent subprocessor disclosure: 22 named third parties with their addresses, plus five affiliated Scale entities.
  • Real GDPR machinery: Article 27 representatives for the EU and the UK, Standard Contractual Clauses, and a DPA available on request.
  • Model- and cloud-agnostic, deploying inside the customer's own VPC rather than in a Scale-hosted tenancy.
  • A self-serve tier lets a team start without going through sales, backed by operation since 2016 and genuine vertical depth in defense, healthcare, insurance, automotive and energy.

Cons

  • No price is published anywhere: no amount, no currency, no billing unit, so budgeting requires a sales conversation.
  • The site never states whether customer data is used to train models, and documents no training opt-out.
  • No public support address: support runs through a named contact and therefore presupposes a contract.
  • No data residency commitment: the privacy policy provides for transfers to the United States and other countries.
  • Legal documents are aging: the site terms are dated 25 August 2021 and the privacy policy 1 July 2024.
  • No mobile app, no browser extension, and no documented interface language other than English.
  • Governance in motion, with three chief executives cited on the site in under two years, and a self-serve tier presented as experimental or research-oriented rather than production-ready.
Pricing

Pricing & Plans

Scale AI publishes no price. Neither an amount, nor a currency, nor a billing unit appears on the pricing page or anywhere else on the site, so no entry price can be quoted here. Two commercial tiers are described. Enterprise is priced on request and reached through the Book a demo form. The Self-Serve Data Engine is billed pay-as-you-go by credit card. The self-serve tier carries a starter allocation at no cost that is capped by volume rather than by time: the first 1,000 labeling units and the first 10,000 uploaded and curated images. Once those volumes are used up, consumption becomes chargeable. That allocation is neither a permanent free plan nor a time-limited free trial. No free trial and no refund policy are published.

Enterprise
  • priced on request
  • described by Scale as ideal for strategic AI initiatives. Enterprise-grade quality and SLAs
  • access to both the Data Engine and the Enterprise GenAI Platform
  • and dedicated customer operations support. Reached through the Book a demo form
  • no amount published.
Self-Serve Data Engine, Data Annotation option
  • bring your own workforce
  • with the first 1
  • 000 labeling units at no cost. Entry point is the Studio product on the Scale dashboard.
Self-Serve Data Engine, Data Management option
  • upload and curate the first 10
  • 000 images at no cost. Entry point is Nucleus on the Scale dashboard.
Special offers — Self-Serve Data Engine starter allocation at no cost: the first 1,000 labeling units and the first 10,000 uploaded and curated images, capped by volume rather than by time. · No promotion, discount code, annual rebate, or student, non-profit, education or startup pricing is published.
Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how Scale AI handles your data.

GDPR overview

Scale publishes concrete GDPR machinery rather than a compliance slogan. Transfers outside the EEA, the UK and Switzerland rely on the European Commission's Standard Contractual Clauses, and Scale self-certifies with the Department of Commerce under the EU-U.S. Data Privacy Framework, its UK extension and the Swiss-U.S. DPF, with JAMS as independent recourse body, binding arbitration and FTC oversight. A Data Processing Addendum is available on request at privacy@scale.com. Two Article 27 representatives are named: Lionheart Squared (Europe) Limited in Dublin for the EU and Lionheart Squared Limited in Hampshire for the UK. The policy details access, rectification, erasure, objection, restriction and portability rights, consent withdrawal and the right to complain to a supervisory authority, adds a dedicated CCPA section for California residents, and restricts the services to users aged 18 and over. It took effect on 1 July 2024; an earlier version remains online.

Who owns the data?

The GenAI Platform FAQ states that customers keep full ownership of their data, their business logic and any custom AI solutions Scale builds for them, while Scale owns the underlying platform infrastructure and technical frameworks, with no vendor lock-in. The privacy policy splits two roles. For website visitors and users, whose information it calls Scale User Data, Scale AI, Inc. is the controller. When Scale handles Scale Services Data on behalf of a customer, the policy does not apply and the customer is the controller; anyone whose account was created by a customer must send access or deletion requests to that customer's administrator, not to Scale.

Reuse rights

Customers may reuse their own data and the solutions built for them without asking permission: ownership stays with the customer and Scale advertises no vendor lock-in. On Scale's own side, the privacy policy sets differentiated legal bases for the EEA, the UK and Switzerland, and describes sharing with customer administrators, affiliates and service providers covering billing, support, hosting, storage, analytics and marketing, plus disclosure in a business transfer or to authorities. Scale states that it does not and will not sell personal information, and that it processes Scale Services Data only as instructed by its customers. Its subprocessor list is public: 22 named third parties with addresses, plus five affiliated Scale entities. One point is simply absent from the site: nowhere does Scale state whether customer data is used to train models, and no training opt-out is documented. The only opt-out published is the CCPA opt-out covering the sale of personal information, which is a different question. Treat model training as an open contractual point to raise in negotiation, not as settled either way.

Data retention & training

Retention summary
Scale keeps personal information for as long as the purposes set out in its privacy policy require, plus whatever is needed to meet legal obligations, resolve disputes and enforce agreements. It states plainly that actual retention periods vary significantly by data type and by service, and lists the criteria applied: the volume, nature and sensitivity of the data, the risk of harm from unauthorized use or disclosure, whether the purpose could be achieved another way, and applicable legal requirements. No figure in days, months or years is published. Data is deleted once those periods expire, and where full deletion is technically impossible Scale says it puts measures in place to prevent further use. Data belonging to anyone under 18 is deleted as soon as Scale becomes aware of it. Erasure requests go to privacy@scale.com, or to the customer's administrator when the account was created by a customer.
Trains on customer data
Unclear
Subprocessors disclosed
Yes
DPA available
Yes
GDPR contact

Hosting summary

Scale AI commits to no specific hosting country or region. Its privacy policy states that information collected from users may be transferred to, stored and processed by Scale, its affiliates and third parties in countries outside the EEA, the UK and Switzerland, including but not limited to the United States and other countries where data protection rules may offer a lower level of protection. Transfers rely on Standard Contractual Clauses and on self-certification under the EU-U.S., UK and Swiss-U.S. Data Privacy Frameworks. The named infrastructure subprocessors are Amazon Web Services in Seattle and Google LLC, for Google Cloud Platform, in Mountain View. The picture is different for the Scale GenAI Platform: it deploys inside the customer's own VPC on AWS, Azure or GCP, so production data stays where the customer puts it and the hosting jurisdiction becomes the customer's decision rather than Scale's. On the public sector side, FedRAMP High Authorized and DoD IL4 Provisional Authorization cover sovereign hosting requirements. The documentation carries a data hosting section that was not examined during this review, so contractual deployment options may go further than the privacy policy alone suggests.

Hosting countries
🇺🇸 United States
Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting Scale AI.

  • No published price means you cannot budget without going through a salesperson, and no refund policy is published either.
  • The site never states whether customer data is used to train models, and documents no opt-out; treat this as an open contractual question rather than a settled one.
  • There is no data residency commitment: the privacy policy provides for transfers to the United States and other countries where protection may be weaker, and subprocessors include providers outside the EU, one of them in the Philippines.
  • The site terms date from 25 August 2021 and the privacy policy from 1 July 2024; check that what you actually sign is current.
  • The site terms cap liability at 100 USD and apply California law, a limit worth weighing against the value of the data you are entrusting.
  • The no-cost self-serve allocation is capped by volume, not by time: it runs out, it is not a free plan, and the switch to paid usage needs to be budgeted.
  • Outsourcing annotation and evaluation can quietly move domain judgement outside your own team; keep enough in-house expertise to challenge what comes back, especially given the repeated leadership changes of the past two years.
Setup

Setup & Integrations

Technical difficulty

Setup is demanding on the enterprise path, light on the self-serve one. Enterprise starts with a discovery call, use-case scoping and integration work; Scale says it finds the use case, builds the system and owns the outcome, but the GenAI Platform deploys inside the customer's own VPC, so infrastructure and security work stays on the customer's side. Connecting Confluence, SharePoint or Amazon S3 means custom pipelines, not generic connectors. Expect data and platform engineering skills, plus a governance function in regulated sectors. The self-serve route needs none of that: sign up on the dashboard and follow the documented API.

Deployment

Web appAPI

Integrations

Confluence SharePoint Amazon S3 AWS Microsoft Azure Google Cloud Platform OpenAI Meta Mistral
Company

Behind Scale AI

Company name
Scale AI, Inc.
Founded
03/05/2016
Country of origin
🇺🇸 United States
Headquarters
650 Townsend Street, San Francisco, CA 94103, United States
UBO
INFORMATION_NOT_FOUND
UBO country
INFORMATION_NOT_FOUND
Domain registrar country
🇺🇸 United States
Legal contact

Fundraising

May 2024, Series F: 1 billion USD raised at a 13.8 billion USD valuation, marking Meta's first entry into the capital on 21 May 2024.
June 2025: Meta invested 14.3 billion USD for a 49% NON-VOTING minority stake, valuing Scale at 29 billion USD.
Meta is a non-controlling minority investor. It is not the owner, the parent company or the publisher of Scale AI, Inc., which remains a Delaware company created on 3 May 2016 and headquartered in San Francisco; Meta also appears among the customers named on the Scale home page.
Aggregators such as Tracxn and Sacra put the total raised at roughly 15.9 billion USD across nine rounds.
Reported governance consequences: co-founder Alexandr Wang left for Meta, Jason Droege served as interim chief executive, and Francis deSouza was appointed CEO on 30 July 2026, this last point being first-party from the site banner and the about page press review.
Source note: apart from the 29 billion USD valuation, which the Scale about page displays directly, these figures come from off-site research and not from scale.com.

Social

Official links

Resources

All the official URLs gathered for verification and reference.

FAQ

Frequently asked questions

Does Scale AI publish its prices?
No. The pricing page exists and describes two tiers, but it carries no amount, no currency and no billing unit. Enterprise is quoted on request through the Book a demo form, and the Self-Serve Data Engine is billed pay-as-you-go by credit card.
Can I use Scale AI without paying?
Only up to a point. The Self-Serve Data Engine includes a starter allocation at no cost: the first 1,000 labeling units and the first 10,000 uploaded and curated images. It is capped by volume, not by time, so it is neither a permanent free plan nor a time-limited free trial. Once the allocation is consumed, usage is charged.
Is there an API?
Yes. Scale documents a public API with its own reference section, alongside developer documentation covering the GenAI Platform, the GenAI Data Engine, the Automotive Data Engine, Nucleus and data hosting. Work is done through a web dashboard and that API; there is no other client.
Which security certifications does Scale AI hold?
SOC 2 Type II, ISO/IEC 27001:2022, a DoD IL4 Provisional Authorization and FedRAMP High Authorized status. Scale points to its Trust Center at trust.scale.com for the underlying reports, and publishes vulnerabilities@scale.com for security disclosures.
Who owns the data and the intellectual property?
The customer does. Scale states that customers retain full ownership of their data, their business logic and any custom AI solutions built for them, and that there is no vendor lock-in. Scale owns the platform infrastructure and the underlying technical frameworks.
Does Scale AI train its models on customer data?
The site does not say. It states that Scale Services Data is processed as instructed by customers and in line with contractual commitments, but it makes no explicit statement about model training and documents no opt-out from training. The only opt-out published concerns the sale of personal information under the CCPA, which is a separate matter. Raise the point in contract negotiation.
Are Scale AI's subprocessors disclosed?
Yes. A dedicated legal page, last updated on 7 August 2024, lists 22 named third-party subprocessors with their addresses, including AWS, Google Cloud, Microsoft, Cloudflare, Datadog, Snowflake and Zendesk, plus five affiliated subprocessors within the Scale group.
Is a Data Processing Addendum available, and is there a GDPR representative in Europe?
Yes to both. Customers transferring personal information from the EEA, the UK or Switzerland can request a copy of the DPA at privacy@scale.com, and transfers rely on Standard Contractual Clauses. Lionheart Squared (Europe) Limited in Dublin acts as the Article 27 representative for the EU, with a separate UK representative.
Where is the data hosted?
There is no residency commitment. The privacy policy provides for information to be transferred to, stored and processed outside the EEA, the UK and Switzerland, including in the United States and other countries. The GenAI Platform is a different case: it deploys inside the customer's own VPC on AWS, Azure or GCP, so production hosting is the customer's choice.
Is there a mobile app, and who can sign up?
There is no iOS or Android app: Scale is a web product with an API. The services are not intended for anyone under 18, and support is handled by a named Engagement Manager or Field Engineer for Pro and Nucleus customers, from 9am to 5pm Pacific on weekdays excluding US holidays, with no public support address.
Conclusion

Should you pick Scale AI?

Scale AI sits at the industrial end of the AI market: data infrastructure and applied AI systems for foundation model labs, large enterprises and public bodies, rather than a tool an individual picks up on a whim. The strengths are concrete. The company has operated since 2016, holds accreditation few AI vendors reach, with FedRAMP High and DoD IL4 alongside SOC 2 Type II and ISO/IEC 27001:2022, guarantees that customers keep ownership of their data and of the solutions built for them, publishes an unusually detailed subprocessor list, and names GDPR representatives for both the EU and the UK. Deployment inside the customer's own VPC, and an architecture that stays agnostic to models and clouds, remove much of the lock-in that comparable vendors impose.

The reservations are just as concrete. Pricing opacity is total: no amount, no currency, no billing unit anywhere, which makes any comparison impossible before a sales conversation. The site never addresses whether customer data feeds model training, and offers no documented opt-out, which is a gap worth closing in the contract rather than assuming away. The published terms date from 2021 and cap liability at 100 USD, and there is no self-service support path.

The practical verdict follows from that. Scale AI makes sense when a framework agreement is realistic and the work genuinely needs expert annotation, rigorous evaluation or agents running in production under audit. The Self-Serve Data Engine offers a way to test the machinery at small scale before committing, though Scale itself frames that tier as experimental or research-oriented. Anyone looking for a self-service product with transparent list pricing should look elsewhere.