
Pangeanic
Pangeanic is a Valencia-based multilingual AI data company supplying licensed and bespoke datasets, annotation, RLHF and model alignment, adaptive machine translation and anonymisation, all deployable in private cloud, on-premises or air-gapped environments under the customer's own governance.
What is Pangeanic?
Pangeanic is a Spanish language technology and AI data company, founded in Valencia in 2000 and led by its founder Manuel Herranz. It presents itself as the operating layer for multilingual AI and addresses three audiences: AI labs and model builders, enterprise AI teams, and regulated or public-sector organisations.
The offer rests on five blocks. The first is data: a catalogue of ready-to-license assets covering multilingual text, parallel corpora for machine translation, speech and audio, image, video, multimodal and OCR material, noise datasets and OSINT collections, alongside bespoke collection programmes defined by language, domain, demographic requirements, format, consent model and quality thresholds. The second is AI Data Operations: taxonomy design, labelling schemes, metadata structures and expert review covering entity tagging, intent classification, question-answer pairs, document labelling and preference data. The third is evaluation and alignment: benchmark design, multilingual scoring, failure analysis, regression testing, human feedback and RLHF. The fourth is model customisation, which means selecting, fine-tuning and evaluating smaller task-specific models, adapting terminology and style, and grounding them with retrieval. The fifth is the ECO Intelligence Platform, which puts all of it into production through secure translation, document intelligence, multilingual retrieval, anonymisation, quality estimation and APIs.
What distinguishes the company is the deployment layer. ECO can run in private cloud, on premises or fully air-gapped, which is what Pangeanic means by sovereign AI. Its named technologies include Deep Adaptive AI Translation (formerly PangeaMT), MTQE for machine translation quality estimation, multilingual data masking, and PECAT, the internal platform that manages human annotation and review workflows.
The track record is European and checkable: the company coordinated NTEU, which delivered 552 direct neural translation directions across the 24 official EU languages without pivoting through English, and MAPA, which built multilingual anonymisation resources. Gartner has listed it as a Representative Vendor for data masking and synthetic data and for conversational AI. Offices are declared in Valencia, Madrid, London, Boston, New York, Tokyo, Hong Kong and Shanghai.
Most services are sold on quotation. The exception is machine translation, which has four public subscription tiers alongside a free translation dashboard that anyone can use.
What it does
- License ready-made multilingual datasets or commission a bespoke data collection programme
- Run annotation, metadata engineering and expert human review at production scale
- Build gold-standard evaluation sets and RLHF preference data to align and benchmark models
- Fine-tune and deploy task-specific small language models on the customer's own infrastructure
- Translate documents while preserving layout, with quality estimation routing risky output to human review
- Mask and anonymise sensitive multilingual data before translation, retrieval or AI processing
- Connect translation and multilingual retrieval to internal systems through a secure API
When to use Pangeanic / When not to
A quick filter to help you decide if Pangeanic is the right fit.
When to use Pangeanic
- AI labs and model builders that need legally usable multilingual training, evaluation and preference data across text, speech, image and video
- Enterprise AI teams connecting data preparation, annotation, translation, anonymisation and human review into one production pipeline
- Public administrations and regulated organisations that must keep sensitive documents inside on-premises or air-gapped infrastructure
- Localisation and translation departments processing high document volumes with format preservation and quality estimation
- Freelance translators and small teams looking for an affordable adaptive machine translation subscription with CAT tool integration
When not to use Pangeanic
- Individuals wanting an instant self-service data or annotation product, since everything outside the machine translation subscription goes through a quote
- Buyers who need published pricing for datasets, AI Data Operations or sovereign AI deployments, none of which is priced publicly
- Organisations that require a signed data processing agreement or a published subprocessor list before onboarding a vendor
- Teams that must guarantee their content is never reused for model improvement, since no training opt-out is documented
- Mobile-first users, as there is no iOS or Android application and access is limited to the web, the API and CAT tools
How to use Pangeanic
A typical end-to-end flow, from setup to results.
- Try the free translation dashboard first, to judge output quality without signing anything
- Decide which route fits: the self-service machine translation subscription, or a scoped data or deployment project
- For the subscription, pick one of the four tiers on the pricing page and subscribe online
- For everything else, contact an AI architect through the contact form and describe the languages, domain and volumes involved
- Agree the scope: dataset licensing, bespoke collection, annotation, evaluation, model customisation or deployment
- Connect the translation API to your CMS or workflow, remembering that API access starts at the Professional tier
- Plug in your CAT tool if you work with SDL, Memsource or memoQ
- Let Pangeanic customise the engine during the implementation phase, using your previously translated content and terminology
- Process documents one by one or in batch, including PDF, Word, PowerPoint, Excel and scanned files, over an encrypted channel
- Choose the deployment model: managed private cloud through ECO, on premises, or air-gapped for the strictest environments
Pros & Cons
Pros
- Twenty-six years of continuous work on multilingual data, backed by European projects that can be checked, including NTEU and its 552 direct translation directions across the 24 official EU languages
- The whole chain under one roof: data sourcing, annotation, human feedback, alignment, evaluation, language technology and deployment
- Air-gapped and on-premises deployment is genuinely offered, which remains rare among multilingual AI vendors
- ISO 9001:2015/Amd 1:2024, ISO/IEC 27001:2022/Amd 1:2024 and ISO 18587:2017 are displayed, with ISO 17100 cited in the terms
- Verifiable public references, including Barcelona Supercomputing Center, the Spanish Tax Agency, the EFE news agency and DoD Iron Bank through Veritone
- Machine translation pricing is published and starts low, at 9.90 euros per month
- A permanently free translation dashboard lets anyone judge quality before committing to anything
Cons
- No public pricing for the core business, since datasets, AI Data Operations and sovereign AI all require a quote
- The terms give the translator ownership of the translated work and leave the client with no rights over derived parallel corpora
- No training opt-out is documented, while the terms explicitly allow submitted content to improve Pangeanic's or third-party systems
- No data processing agreement is published or offered, and no subprocessor list exists
- No hosting country or region is named for the services Pangeanic runs itself
- No dedicated support address: only a generic mailbox and regional office addresses are published
- The legal identity wavers between two spellings across the site's own pages, and the homepage footer's legal links are inert JavaScript placeholders
Pricing & Plans
A permanent free option exists: the public translation dashboard and the public LLM portal may be used free of charge within what the terms call reasonable limits. The lowest paid entry point is the Standard machine translation subscription at 9.90 euros per month, or 118 euros per year as a single payment, with VAT excluded. Datasets, AI Data Operations and sovereign AI deployments are not priced publicly and are quoted on request.
- translation dashboard and publicly hosted fine-tuned LLM
- free of charge within reasonable limits
- 1 concurrent Deep Adaptive model
- generic MT
- 100 pages per month (about 150
- 000 characters)
- 2 users
- CAT tool integration
- up to 1
- 000
- 1 concurrent Deep Adaptive model
- 400 pages per month (about 600
- 000 characters)
- 50 users
- API access
- CAT tool integration
- technical support
- 15 concurrent Deep Adaptive models
- generic MT on other language pairs
- 2
- 000 pages per month (about 3
- 000
- 000 characters)
- 100 users
- API access
- dedicated machine learning team
- pages
- language pairs
- users and administrators on request
- no restrictions on users
- storage or bandwidth
- custom APIs and integrations
- custom LQA
- VAT excluded
- unlimited language pairs
- one page counted as 1
- 500 characters
- and extra pages billed individually
Data, GDPR & hosting
A consolidated view of how Pangeanic handles your data.
GDPR overview
Pangeanic publishes a dedicated GDPR page structured around controller, purpose, legitimacy, recipients, rights and data origin. The controller is named as Pangeanic B. I. Europa S. L., VAT B97017461, at Av. de les Corts Valencianes 26, bloque 5, oficina 107, 46015 Valencia, with info@pangeanic.com given as the data protection contact. Legal bases are stated: contract performance for clients, consent for prospecting, and the supplier contract and non-disclosure agreement for providers. Access, rectification, erasure, restriction and objection rights are all described, and the legal notice points to Spanish Organic Law 3/2018. The terms add a mutual GDPR compliance clause. Being established in Spain, the company needs no Article 27 representative. Two gaps remain: no data processing agreement is published or offered, and no subprocessor list exists beyond the single tool named on the page.
Who owns the data?
Pangeanic's terms are unusually vendor-favourable on ownership. Once a translation is delivered, the translator retains all intellectual property rights in the translated work, including copyright, patent and trademark rights, unless something else is agreed in writing. The client grants Pangeanic a non-exclusive, irrevocable, worldwide licence to use, reproduce, modify and distribute the original work in order to deliver the service. A separate clause lets Pangeanic reuse both the original and the translated work for data augmentation and parallel corpus creation, and states that the client owns no intellectual property in the corpora produced that way. For personal data specifically, Pangeanic undertakes to process it only on the client's instructions.
Reuse rights
The client cannot freely reuse everything it receives. Because the translator keeps the intellectual property in the delivered translation, the client's right to exploit that output is bounded by the contract, and any use inconsistent with those ownership rights requires written agreement. In the other direction, Pangeanic reserves wide reuse rights over what is submitted to it: documents uploaded to the free public translation portal and conversations held with the publicly hosted fine-tuned LLM may be used to improve Pangeanic's systems or third-party systems, and client material may feed data augmentation and parallel corpus creation. Personal data is handled internally and passed to translation providers when a project requires it, with XTRF named as the project and supplier management system. No opt-out from this reuse is documented anywhere on the site.
Data retention & training
Hosting summary
The site never names a country or a region where Pangeanic hosts customer data. What it offers instead is control over where processing happens: the ECO Intelligence Platform can run in private cloud, on the customer's own premises, or fully air-gapped, and these are presented as options for organisations with strict security and data-residency requirements. The jurisdiction that is clearly established is corporate rather than technical. The company is registered in Valencia, Spain, operates under Spanish law, and disputes go before the courts of Valencia. Being established in the European Union, it falls directly under the GDPR without needing an Article 27 representative. The public website itself sits behind an anycast address, which says nothing about where customer data lives. Anyone with a data residency requirement should therefore treat hosting location as a contractual question to settle during scoping, not as something the website already answers.
Things to keep in mind
Risks and trade-offs to weigh before adopting Pangeanic.
- Ownership of the translated work stays with the translator unless you negotiate otherwise in writing, so read that clause before sending anything you intend to own
- Content you submit can be reused for data augmentation and parallel corpus creation, and nothing on the site describes a way to exclude it
- The free public portals are most tempting when you are in a hurry, which is exactly when confidential material gets pasted into them; the terms say those uploads may improve Pangeanic's or third-party systems
- The absence of a published data processing agreement and of a subprocessor list makes a compliance review harder than it should be, and therefore easy to postpone
- Quality estimation routes risky output to human review, but relying on it without actually staffing that review turns a safety net into a rubber stamp
- Two different legal names appear across the company's own pages, which complicates contracting and due diligence
- Late payment carries a 50 euro charge after 60 days, another 50 euros after 90 days, and interest at 8 percent above the European Central Bank base rate
Setup & Integrations
Technical difficulty
Effort varies enormously. The free translation dashboard needs nothing at all, and subscribing to a machine translation tier is an online purchase. Connecting the API is a modest developer task: text can be sent straight from a CMS without a login, and SDL, Memsource and memoQ plug in directly. Beyond that the difficulty rises sharply. Engine customisation happens during an implementation phase using your previously translated content, and bespoke data collection, model fine-tuning or on-premises and air-gapped deployment are engineering projects run jointly with Pangeanic's teams rather than configurations you complete alone.
Deployment
Integrations
Supported languages
Behind Pangeanic
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Frequently asked questions
What does Pangeanic actually sell?
Can I use it without talking to a salesperson?
How much does it cost to start?
Is there a free option?
Is there an API?
Can it be deployed on our own infrastructure?
Will my content be used to train models?
How long is data kept?
Which certifications does the company hold?
Which CAT tools are supported?
Should you pick Pangeanic?
Pangeanic is not a piece of software you sign up for; it is a data and services partner with a product layer on top. Twenty-six years spent building multilingual corpora, coordinating European projects such as NTEU and MAPA, and delivering machine translation to public administrations give it something few AI data vendors can claim: a documented lineage rather than a launch announcement. The combination it sells - sourcing, annotation, human feedback, alignment, evaluation and then deployment under the customer's own governance - is genuinely rare, and the air-gapped option will matter to defence, tax, health and legal organisations that cannot send documents anywhere.
The reservations are contractual rather than technical. The terms hand intellectual property in the translated work to the translator, leave the client with no rights over the parallel corpora derived from its own content, and allow submitted material to improve Pangeanic's systems or third parties' systems with no documented way to opt out. No data processing agreement is published, no subprocessor list exists beyond a single named tool, and no hosting country is stated for the services Pangeanic runs itself. Anyone in a regulated sector should negotiate these points explicitly rather than accept the published terms as they stand.
Pricing is split in two. The machine translation subscription is inexpensive, public and easy to test through the free dashboard. Everything that constitutes the real business - datasets, AI Data Operations, sovereign deployments - is quoted privately, which makes budgeting impossible from the website alone.
The realistic audience is an organisation with a genuine multilingual AI programme or a sovereignty constraint, not an individual looking for a translation tool. For that audience, Pangeanic is a serious candidate, provided the contract is read closely.
- Choosing a selection results in a full page refresh.
- Opens in a new window.