Sieve logo
Data Labeling · Dataset Marketplaces

Sieve

Sieve is a multimodal data lab that sources, filters, indexes and annotates video, audio, image and interaction data, then delivers training-ready datasets, evaluation sets and environments to frontier AI labs under negotiated purchase agreements.

Active Contact Sales No public API 18+ Verified by Guidaio
Overview

What is Sieve?

Sieve is a multimodal data lab: it builds and sells the training corpora that frontier AI labs, Fortune 100 companies and fast-growing AI startups use to teach models how the world looks, sounds and moves. The company describes its own pipeline in five stages, namely source, filter, index, annotate and deliver. It captures and aggregates material across real-world, digital and simulated environments; scores it for semantics, rights, artefacts and task quality; indexes billions of videos, images, audio clips and interaction traces with purpose-built detectors and embeddings; adds dense labels, pairings, temporal alignment, transcripts, action metadata and human QA; then packages training-ready datasets, evaluation sets and environments for secure delivery.

Three families sit on the homepage: high-quality video, meaning coherent scenes with clean motion, composition, physics and storytelling; editing pairs, before-and-after media for controlled generation and editing; and audio-visual data, synchronised video, image, speech, music and sound. Scale is the argument. Sieve says it operates on hundreds of petabytes of video, collects millions of hours of new content every month, and has embedded one billion videos into 41 billion queryable vectors. Supply is built proactively through a contributor platform and data partnerships rather than assembled request by request, which is what lets pre-packaged datasets ship within days.

The approach is deliberately software-first. Quality control runs in layers: contributors are qualified before onboarding, new sources ramp in a controlled way, automated validation runs through capture, upload and processing, and every asset ends with human review whose outputs feed back into Sieve's own QA models. Roughly 75% of the company works in research and engineering.

This is not the product Sieve started with. Until early 2025 it sold an API platform for video understanding and editing, covering auto-resizing, translation, scene detection, object tracking and speaker tracking. In a March 2026 post, co-founder Mokshith Voodarla explained the pivot: customers were using those APIs as annotation systems, and model progress looked likely to make the API products obsolete within one to two years. Downstream work today spans generative media, visual understanding, robotics, world models, computer use and agentic systems. There is no self-serve access.

What it does

  • Supply training-ready multimodal datasets at petabyte scale
  • Source video, audio, image and interaction data from real, digital and simulated environments
  • Filter footage on semantics, rights, artefacts and task quality
  • Index billions of clips with purpose-built detectors and embeddings
  • Annotate with captions, transcripts, object labels, action metadata and temporal alignment
  • Deliver evaluation sets and interactive environments for testing models
  • Run custom collection programmes targeting specific model capabilities
Audience

When to use Sieve / When not to

A quick filter to help you decide if Sieve is the right fit.

When to use Sieve

  • Research teams training multimodal models whose bottleneck is data rather than compute
  • AI labs that need evaluation sets and interactive environments alongside training corpora
  • Robotics, world-model and computer-use teams needing captured real, digital and simulated footage
  • Enterprises with strict rights requirements, where licensing, consent and permissions must be filtered in
  • Buyers operating at petabyte scale, where manual collection and human review are simply not feasible

When not to use Sieve

  • Individuals or small teams wanting a self-serve tool: there is no sign-up, no trial and no published price
  • Developers looking for a video processing API, which Sieve retired along with its documentation
  • Anyone who needs to budget upfront, since every engagement runs through a negotiated purchase agreement
  • European buyers requiring a DPA, an Article 27 representative or a named subprocessor list before signing
  • Teams expecting to browse and download a catalogue themselves, as the Browse Datasets button leads nowhere
Get started

How to use Sieve

A typical end-to-end flow, from setup to results.

  1. Open the contact form on the Sieve site; there is no account to create and nothing to install
  2. Describe what your research team needs, or ask which ready-to-use datasets already exist
  3. Request samples across video, audio, image, computer use or interactive environments
  4. Review those samples against your model's failure modes, distributions and evaluation goals
  5. Scope the dataset or environment with Sieve's team: volume, distributions, metadata, licensing, QA and delivery format
  6. Agree a purchase agreement priced on data volume, task complexity and annotation depth
  7. Take delivery, with pre-packaged datasets arriving within days and custom work following an agreed SLA
  8. Receive the material over encrypted secure transfer, under the retention window you negotiated
Quick read

Pros & Cons

Pros

  • Scale very few suppliers can match: hundreds of petabytes, millions of new hours ingested each month
  • Rights, licensing and consent are handled inside the filtering rather than bolted on afterwards
  • SOC 2 Type 2 controls, end-to-end encryption and negotiable retention on delivery
  • Inventory built in advance means catalogue datasets can ship in days rather than months
  • Evaluation sets and interactive environments are delivered alongside training data
  • Engineering-dense team, around 75% in research and engineering, with stated backgrounds at NVIDIA, Scale AI, Zoox and Niantic
  • Institutional backers named openly: Y Combinator, Swift Ventures, Matrix Partners and AI Grant

Cons

  • No published pricing of any kind, not even an order of magnitude
  • No terms of service: the privacy policy is the only contractual document on the site
  • GDPR is never mentioned, and there is no DPA, no Article 27 representative and no named subprocessor list
  • The privacy policy points to a full policy document that is never linked anywhere
  • A six-page site, two of whose pages render almost nothing without JavaScript
  • No postal address and no phone number; contact is a form plus a single support mailbox
  • Browse Datasets is a dead button, so there is no catalogue to browse
Pricing

Pricing & Plans

Sieve publishes no pricing. There is no permanent free plan and no free trial, and the site has no pricing page at all, the /pricing path returning an error. Access is sold under a purchase agreement whose value is set by data volume, task complexity and the annotation work required, so no entry price and no currency can be quoted. The only cost-free step available is the data sample Sieve provides on request through its contact form, which is presented as an evaluation sample rather than a trial.

Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how Sieve handles your data.

GDPR overview

There is no GDPR implementation to report. The word GDPR appears nowhere on the site, neither in the privacy policy nor on any other page. The only regulatory framework Sieve addresses explicitly is the California Consumer Privacy Act, which gets a dedicated supplemental notice. The policy does list rights that look European in shape, namely access and portability, rectification, erasure, restriction and objection, and withdrawal of consent, but attaches them to no legal basis and no text. Requests go to support@sievedata.com or the contact form. International transfers to the United States or other countries are announced without naming any transfer mechanism. No Article 27 representative is designated, no data protection officer is named, and no Data Processing Agreement is published or offered.

Who owns the data?

Sieve, Inc. declares itself the controller of the personal information covered by its privacy policy, last updated on 16 February 2026. That policy deliberately stops at the contract line: personal information Sieve processes on behalf of its Customers under written agreements falls outside it, and readers are told to contact the relevant Customer directly. Content Providers who upload footage see their material shared with the Customers who purchase the resulting datasets. Sieve states that it does not sell personal information, and its California notice confirms no sale or sharing for behavioural advertising in the preceding twelve months. No terms of service are published anywhere on the site, so ownership of delivered datasets is defined only by the individual purchase agreement.

Reuse rights

Nothing on the site grants a buyer any reuse right by default. Datasets are transferred under a purchase agreement whose scope, including volume, distributions, metadata, licensing, QA and delivery format, is negotiated case by case, so what a customer may do with the material is defined by that contract and not by any public licence. In the other direction Sieve reserves broad internal use of the information it collects: training and improving AI models and algorithms is listed among its declared purposes, and the company says outputs from human QA are used to train its own QA and QC models. Content uploaded by contributors is shared with the Customers who buy the datasets built from it.

Data retention & training

Retention summary
Sieve gives rules but no numbers. Personal information is kept for as long as you use the Services, or for as long as it takes to fulfil the purpose it was collected for. How long that is depends on four stated factors: legal requirements, how sensitive the data is, the risk involved, and whether the same purpose could be achieved another way. No retention period is expressed in days, months or years anywhere on the site, and nothing is said about anonymisation schedules beyond a general mention of de-identifying and aggregating data as a business operation. Deletion is available on request, through support@sievedata.com or the contact form, and Sieve may verify your identity first. Separately, the homepage advertises custom data retention as part of secure delivery to customers, but gives no detail.
Trains on customer data
Yes

Hosting summary

Sieve names exactly one hosting jurisdiction: the United States. Its privacy policy warns that data may be transferred to, processed and stored in the United States or other countries whose data protection laws differ from the user's own, and stops there. There is no list of those other countries, no hosting provider named, and no data residency option offered. Third parties appear only as categories, one of which is hosting. No EU region is mentioned, nothing is said about where dataset material sits during processing, and there is no trust or security page beyond the homepage's claim of end-to-end encryption, custom data retention, secure transfer and SOC 2 Type 2 controls. For completeness, the website itself resolves to an IP address geolocated in the United States, but that describes the marketing site's infrastructure and says nothing about where customer or contributor data lives. A buyer with residency requirements will have to establish them contractually, because the published material does not answer the question.

Hosting countries
🇺🇸 United States
Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting Sieve.

  • The privacy policy lists training and improving AI models among its purposes, so anything you upload may end up shaping a model
  • Content Providers hand over government-issued ID, photographs and precise device geolocation, a heavy identity footprint for contribution work
  • Contributed material is resold to Customers, so anyone appearing in it depends entirely on consent having been collected properly upstream
  • Data Sieve processes for its Customers sits outside the published policy, so a buyer cannot learn from the site what their own contract permits
  • With no terms of service published, liability, indemnity and dataset warranties stay invisible until a contract is on the table
  • Training on bought-in corpora concentrates whatever bias, gaps or rights defects the source material carries, so provenance deserves independent audit rather than trust
  • Buying data instead of collecting it can quietly erode a team's own grasp of its data distribution, which is exactly what debugging a model later depends on
Setup

Setup & Integrations

Technical difficulty

There is nothing to install and no account to open, so setup effort on Sieve's side is effectively zero: the entry point is a contact form. The work sits downstream. Delivery formats, metadata schemas and annotation structures are negotiated per engagement, and what arrives is measured in terabytes or petabytes, so the buyer needs storage, ingestion and validation pipelines already in place. Pre-packaged datasets arrive within days, while custom collections and environments follow an agreed SLA. In practice this is a procurement and data-engineering exercise rather than a software integration.

Deployment

Web app
Company

Behind Sieve

Company name
Sieve, Inc.
Founded
INFORMATION_NOT_FOUND
Country of origin
🇺🇸 United States
UBO
INFORMATION_NOT_FOUND
UBO country
INFORMATION_NOT_FOUND
Domain registrar country
🇺🇸 United States
Support contact

Fundraising

Backed by Y Combinator, Swift Ventures, Matrix Partners (Ilya Sukhar) and AI Grant (Nat Friedman and Daniel Gross), all four named on Sieve's own About page
A seed round of roughly USD 4 million, led by Matrix Partners, was announced in a Sieve blog post that has since been taken offline; the figure now survives only in search indexes and third-party databases
Third-party profiles also reference a later Series A with an undisclosed amount, unconfirmed by Sieve, which publishes no round, no figure and no date on its current site

Social

Official links

Resources

All the official URLs gathered for verification and reference.

FAQ

Frequently asked questions

What does Sieve actually sell?
Multimodal training datasets, evaluation sets and interactive environments, packaged and handed over by secure transfer. It is a data supply business, not a piece of software you log into.
Which data modalities does Sieve cover?
Video, image and audio, the latter covering speech, music and sound, plus interaction traces including computer use. The homepage also highlights before-and-after editing pairs for controlled generation.
How much does it cost?
Sieve publishes no prices. Each engagement is a purchase agreement calibrated on data volume, task complexity and the annotations required, so the figure only emerges through a sales conversation.
Is there a free trial or a free plan?
Neither is advertised. Sieve does provide data samples on request through its contact form, but these are presented as evaluation samples rather than as a trial or a free tier.
Does Sieve offer an API?
Not today. The company sold a video understanding and editing API until early 2025, then withdrew it; the current site documents no API and the old documentation host no longer resolves.
What security guarantees are stated?
End-to-end encryption, custom data retention, secure transfer and SOC 2 Type 2 controls, all claimed on the homepage. No trust or security page expands on them, and no audit report is published.
Does Sieve train models on the data it collects?
Yes. Training and improving AI models and algorithms is listed among the declared purposes in its privacy policy. Data processed on behalf of a Customer under a written agreement is explicitly outside that policy's scope.
Is there a minimum age?
Yes, 18. The privacy policy states that the Services are not directed to children under 18 and that Sieve does not knowingly collect their personal information.
Where is the data hosted?
The policy names the United States, then adds or other countries without listing them. No hosting provider is named, no EU region is offered and no data residency option is described.
Who is behind Sieve and how do you reach them?
Sieve, Inc., backed by Y Combinator, Swift Ventures, Matrix Partners and AI Grant. There is no published postal address or phone number: contact runs through the form on the site, or support@sievedata.com for privacy matters.
Conclusion

Should you pick Sieve?

Sieve is not a tool you sign up for, it is a supplier you negotiate with, and judged on that basis it presents a credible case. The scale it claims, hundreds of petabytes, millions of hours ingested each month, a billion videos embedded into 41 billion vectors, is the kind of inventory that makes fast delivery plausible rather than aspirational. Building supply proactively instead of request by request is a real structural difference from most data vendors. Rights handling sits inside the filtering rather than beside it, SOC 2 Type 2 controls are claimed on delivery, and the team's provenance is stated openly.

The reservations concern transparency rather than capability. Nothing about price is public, which is normal in this market but leaves a buyer unable to size an engagement before speaking to sales. More unusually, Sieve publishes no terms of service at all: a privacy policy is the only contractual document on the site, and it expressly declines to cover the data Sieve handles for its own customers. For a European buyer the gaps are sharper still, since GDPR is never mentioned, no Data Processing Agreement is offered, no Article 27 representative is designated and subprocessors appear only as categories. The site itself is thin: six pages, no address, no phone number, and a Browse Datasets button that leads nowhere.

One point of history matters. Until early 2025 Sieve sold a video understanding and editing API, so any description of the company written before 2026 describes a product that no longer exists. Treat older listings as obsolete rather than incomplete. For research teams whose bottleneck genuinely is data, Sieve is worth a conversation, with the contract rather than the website doing the reassuring.