Folklore logo
Academic Research · Agents Orchestration Frameworks

Folklore

Clinical genomics software that classifies germline variants from a VCF under the ACMG/AMP rules, adds HPO phenotype matching, literature evidence and prioritisation, then hands a traceable case record to a qualified geneticist for review and sign-off.

Active GDPR compliant Free plan Freemium API available 18+ Verified by Guidaio
Overview

What is Folklore?

Folklore is clinical genomics software from Helena Bioinformatics, a Bulgarian company based in Sofia. It takes a single-sample VCF 4.1 or 4.2 file from a gene panel, exome or whole genome and turns it into a reviewable case record: annotated variants, an ACMG/AMP classification with every applied criterion exposed, phenotype-ranked genes, scored literature and a report a qualified specialist signs.

The pipeline runs in six stages. Variants are quality-filtered, with clinically significant ClinVar entries preserved regardless of score. Annotation goes through Ensembl VEP release 113 from a local offline cache, joined to locally stored reference data including gnomAD v4.1, ClinVar, dbNSFP, HPO, ClinGen and precomputed SpliceAI. The vendor is precise about scale here: 16 production databases sit in the core enrichment stage, while 45 reference databases and curated sources is the total across every module. Classification then applies the 2015 ACMG/AMP guidelines with all 28 evidence criteria in a Bayesian point framework. Folklore states plainly that classification is strictly rule-based and that no AI model determines pathogenicity; material relating to a version 4 of the standard is described as beta or draft, and no conformance with a final version 4 is claimed. PP3 and BP4 rest on BayesDel_noAF with ClinGen SVI-calibrated thresholds, with a separate guarded SpliceAI path. Scores such as SIFT, AlphaMissense, MetaSVM, DANN, PhyloP and GERP are displayed for clinical context and carry no ACMG vote of their own.

Around that core sit modules for mitochondrial DNA under MMDWG 2020, structural and copy-number variants under Riggs 2020, family and trio inheritance analysis, cohort analytics, phenotype matching and clinical screening. Phenotype matching orders candidates by HPO semantic similarity; it does not rewrite a classification and neither score establishes a diagnosis. An AI assistant answers questions about the case and shows the SQL it generates, but cannot alter a classification.

Private genomic processing runs on dedicated hardware in Helsinki, Finland, with no outbound network access from the analysis pipeline. Alongside the paid platform, Folklore publishes a free read-only MCP endpoint that classifies one public GRCh38 variant without any account. The vendor states the product is Research Use Only and not CE-marked.

What it does

  • Classify germline variants from a VCF under the ACMG/AMP rules, with every applied criterion listed
  • Rank genes and variants against the patient's HPO phenotype profile using semantic similarity
  • Retrieve and score biomedical literature with traceable PMID, DOI and extracted context
  • Prioritise findings into clinical tiers using age, sex, ancestry, family history and testing indication
  • Interpret mitochondrial variants under MMDWG and copy-number changes under the ClinGen/ACMG Riggs framework
  • Generate a tiered clinical interpretation report in PDF or DOCX with full evidence attribution
  • Query one public GRCh38 variant for free through a read-only MCP endpoint, with no account or API key
Audience

When to use Folklore / When not to

A quick filter to help you decide if Folklore is the right fit.

When to use Folklore

  • Clinical genetics laboratories running roughly 50 to 500 inherited-disease exome or genome cases a year
  • Clinical geneticists who want evidence gathered, cross-referenced and documented before their own review begins
  • Rare-disease, newborn-screening and carrier-screening teams needing phenotype-aware prioritisation
  • Bioinformaticians and research groups analysing trios, families or whole cohorts under a locked classifier version
  • Developers and AI agents that need a free, read-only variant-classification endpoint with no account or API key

When not to use Folklore

  • Patients or members of the public: the service is explicitly not offered directly to patients for diagnosis or treatment decisions
  • Somatic and tumour work: tumour-normal analysis and AMP/ASCO/CAP classification are outside the product scope
  • Laboratories needing repeat-expansion disorders, which are neither detected nor interpreted
  • Teams looking for upstream variant or structural-variant calling: Folklore interprets a VCF, it does not produce one
  • Any workflow requiring a CE-marked in vitro diagnostic device, since the vendor states the product is Research Use Only
Get started

How to use Folklore

A typical end-to-end flow, from setup to results.

  1. Try the product first for free: search a single GRCh38 variant on the home page by coordinate, HGVS, SPDI or rsID
  2. Or connect an AI client to the public MCP endpoint over Streamable HTTP, with no account, API key or OAuth flow
  3. For laboratory use, create an account and pick an Individual, Laboratory or Distributor profile
  4. Expect Helena to ask for evidence of professional qualification or organisational identity before clinical features are enabled
  5. Have an authorised representative accept or sign the data processing agreement before any patient data is submitted
  6. Upload a pseudonymised single-sample VCF 4.1 or 4.2 file, splitting multi-sample files beforehand
  7. Set the quality filters for read depth, genotype quality and allelic balance
  8. Enter the patient's HPO terms directly, by HP:ID, or by AI-assisted extraction from a referral letter
  9. Choose a built-in or custom gene panel and one of the six screening modes, then run the analysis
  10. Review the fired criteria and evidence, reclassify where your judgement differs, and generate the signed report
Quick read

Pros & Cons

Pros

  • Classification is deterministic and auditable: each result carries its criteria, its sources and the classifier version
  • No AI model sits inside the class calculation, and the vendor separates generated text from the decision explicitly
  • Private genomic processing runs on dedicated EU hardware in Helsinki, outside multi-tenant cloud and outside claimed US jurisdiction
  • The variant pipeline has no outbound network route by design, enforced at firewall and container level rather than in the application
  • Public documentation is exceptionally dense, with 131 canonical pages, a per-version changelog and an explicit limitations page
  • Validation is published and citable: four cohorts covering 197 retrospective cases, with public reports carrying OSF DOIs
  • A free read-only machine interface needs no account or API key, and the adapter is open source under Apache-2.0

Cons

  • No price is published anywhere: neither the product site nor the company site exposes a pricing page or any amount
  • The vendor states the product is Research Use Only and not CE-marked, so it cannot serve as an in vitro diagnostic device
  • Helena holds no certification of its own; the ISO 27001 cited in the impact assessment belongs to Hetzner, its hosting subprocessor
  • The published cohort 4 result is 67.0% full plus clinical agreement over 100 cases, on the vendor's own figures
  • All three legal documents carry a Candidate for legal review banner, so the vendor itself flags them as unvalidated
  • The documentation contradicts itself on whether the clinical AI assistant may call an external language model API
  • Scope is strictly germline: no somatic variants, no repeat expansions, no upstream variant or structural-variant calling
Pricing

Pricing & Plans

A permanent free plan exists. The terms of use state that version 1.0 supports a free plan without a payment card, and that paid charges, taxes, renewal, cancellation and refund terms apply only if a paid offer is separately selected, with its ordering and billing terms presented before purchase. The public variant search and the read-only MCP endpoint are free and require no account. No paid price is published. Neither the product site nor the company site exposes a pricing page, and no amount in any currency appears in the vendor's own complete public text corpus. To investors, Helena describes the commercial model as B2B software with usage-based and volume access, and a route running from an own-data pilot to per-case purchase, prepaid volume packs and then an annual commitment or private instance. Pricing enquiries are directed to the company's general contact address.

Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how Folklore handles your data.

GDPR overview

GDPR treatment is documented in unusual detail. A privacy notice, terms of use and data processing agreement are published in full, all version 1.0 effective 14 August 2026, alongside a public impact assessment summary (version 1.1, March 2026) carried out under Article 35 because genetic data is Article 9 special-category data. The laboratory is controller and chooses the Article 9(2)(a) or 9(2)(h) basis; Helena is processor. Private genomic processing stays inside the EEA, with standard contractual clauses reserved for any future transfer. Annex 3 of the agreement authorises a single subprocessor, Hetzner, for private customer data. The Bulgarian supervisory authority is named, and privacy@helena.bio is given as the data-protection contact, with an explicit caveat that this is not a formal DPO designation. All three legal documents carry a Candidate for legal review banner.

Who owns the data?

The customer keeps ownership. The terms state that you retain your rights in the data, files and materials you submit, and grant Helena Bioinformatics only the limited rights needed to host, copy, transmit, analyse and process that content in order to provide, secure, support and maintain the service and to meet legal obligations. For genomic case data the laboratory is the data controller and Helena acts strictly as processor on documented instructions; Helena is controller only for account, security, contact and billing records. Patient data must be pseudonymised before upload. At the end of the agreement the controller chooses return or deletion. Separately, Helena claims a perpetual right over feedback you volunteer, without identifying you or disclosing customer data.

Reuse rights

Customer data is processed only on the controller's documented instructions, for delivering, securing, supporting and maintaining the service. Results belong to the laboratory and may be used in its own professional or research workflow, subject to third-party database licences, attribution requirements and patient rights, and no permission request to the vendor is described for that reuse. Two secondary uses are disclosed and deserve attention. The impact assessment states that de-identified data may be processed for internal research and development, algorithm validation and platform improvement under a separate data use agreement, with Helena then acting as an independent controller on a legitimate-interest basis. And the privacy notice names OpenAI as processing bounded content from the public assistant, with provider storage disabled in Helena's configured request. Nothing on the site states either that customer data trains AI models or that it does not; the classifier itself is rule-based, not trained. No documented mechanism lets a customer opt out of the research and development use.

Data retention & training

Retention summary
Original VCF files are deleted automatically once processing completes, and raw sequencing data is never accepted. Parsed variant data and the structured results, including ACMG evidence, phenotype scores, screening tiers and literature, are kept for the period agreed with the laboratory so that a case can be revisited without re-upload, and are removed when the session is deleted or the agreement ends. Phenotype and clinical context go with the session. Account details are deleted within 30 days of termination or on request, and IP addresses and session logs are purged after at most 12 months. Deletion on request is completed within 30 days and certified in writing, and every deletion event is logged. Processing metadata is kept for the audit trail and holds no genomic data. The impact assessment separately mentions a 90-day default retention period.
Trains on customer data
Unclear
Subprocessors disclosed
Yes
DPA available
Yes
GDPR contact

Hosting summary

Private genomic processing and storage run on dedicated hardware physically located in Helsinki, Finland, identified as a Hetzner server, rather than on multi-tenant cloud infrastructure. The vendor gives four reasons: physical isolation from other customers, a fixed and verifiable data location, no exposure to US disclosure laws because the hosting provider is a European company, and exclusive root-level access for Helena so that no hosting employee reaches the operating system, storage or network. Annex 3 of the data processing agreement authorises Hetzner alone to receive private customer data, and no public website, analytics, email, payment or public-AI provider is authorised for it. Data travels over TLS 1.3, and no genomic or personal data is transferred outside the European Economic Area; standard contractual clauses are reserved for any future transfer. Vercel serves the public websites and front ends and may see network and usage metadata, but not private genomic content. Note one documented inconsistency: encryption at rest is given as AES-256 on the infrastructure page and as AES-128 at application level plus full-disk encryption in the impact assessment.

Hosting countries
🇫🇮 Finland
Hosting regions
EU
Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting Folklore.

  • Research Use Only and not CE-marked: do not deploy it as an in vitro diagnostic device or as the basis of a regulated report
  • The only certification named on the site, ISO 27001, is held by the hosting subprocessor Hetzner and must not be read as Helena's own
  • The documentation contradicts itself on external AI: the impact assessment and FAQ say all inference is local, while the no-external-calls page says the clinical assistant may use an external language model API
  • Retention is stated twice and differently: a 90-day default in the impact assessment, a duration agreed with the laboratory on the retention page
  • Automation bias is the real clinical risk: the vendor insists the geneticist decides, and a 67.0% agreement rate on its largest cohort shows why that insistence matters
  • The public MCP endpoint accepts public variant data only; sending patient identifiers, phenotype or case context to it would be a disclosure incident
  • Displayed predictor scores such as SIFT, AlphaMissense, PhyloP and GERP are context only and carry no ACMG vote, so do not read them as criteria
Setup

Setup & Integrations

Technical difficulty

Two very different levels. Public access is effortless: the variant search needs nothing, and connecting an AI client to the MCP endpoint takes one line of configuration over Streamable HTTP, with no account, API key or OAuth flow. Laboratory access is contractual rather than technical: you create an account, choose an Individual or Laboratory profile, may be asked to evidence professional qualification, and must have an authorised representative accept the data processing agreement before any patient data is uploaded. Nothing is installed: the product is web-based and a conforming pseudonymised VCF is uploaded over TLS.

Deployment

Web appAPI

Integrations

ChatGPT Claude Cursor Visual Studio Code Gemini CLI Perplexity Grok Windsurf Codex Stripe

Supported languages

English
Company

Behind Folklore

Company name
Helena Bioinformatics EOOD
Founded
13/08/2026
Country of origin
🇧🇬 Bulgaria
Headquarters
14 Tsar Ivan Asen II Street, floor 1, apartment 1, 1142 Sofia, Bulgaria
UBO
INFORMATION_NOT_FOUND
UBO country
INFORMATION_NOT_FOUND
Domain registrar country
🇺🇸 United States
Legal contact

Social

Official links

Resources

All the official URLs gathered for verification and reference.

FAQ

Frequently asked questions

What file formats and genome builds does Folklore accept?
Single-sample VCF 4.1 or 4.2 files, plain text or bgzipped; multi-sample files must be split first, and FASTQ or BAM are not accepted. GRCh38 is the primary build. GRCh37 files are accepted and automatically lifted over with CrossMap, while hg18 and earlier builds are unsupported.
Does an AI model decide whether a variant is pathogenic?
No. The vendor states that classification is strictly rule-based and that no AI model determines pathogenicity. The AI assistant can explain why criteria fired and discuss the evidence, but it cannot change a classification. Only the reviewing geneticist can override an automated result.
Is a Folklore result a diagnosis?
No. The vendor describes it as clinical decision support. Every automated classification, screening priority and generated interpretation must be independently reviewed and validated by a qualified clinical professional before it is used in patient care.
Where is patient data stored?
On EU infrastructure in Helsinki, Finland, on a dedicated server rather than a multi-tenant cloud. The vendor states that no genomic or personal data is transferred outside the European Economic Area, and that the variant processing pipeline has no outbound network access.
What does Folklore cost?
No paid price is published. The terms state that version 1.0 supports a free plan without a payment card, and the public variant search and MCP endpoint are free with no account. Paid access is described to investors as per-case purchase, prepaid volumes or an annual commitment, and pricing enquiries go through the company contact address.
Is there an API?
Yes. Folklore publishes a free, read-only MCP endpoint that requires no account, API key or OAuth flow, with four tools and five workflow prompts, a connector guide, an official MCP registry identity and an Apache-2.0 adapter on GitHub. It accepts public variant-level queries only, never patient or case data.
How many reference databases does Folklore use?
The vendor distinguishes two counts: 16 production classification and reference databases in the core enrichment stage, and 45 reference databases and curated sources in total across all modules. Smaller figures quoted for individual subsystems are not platform-wide totals.
What is the regulatory status?
The company states plainly that the product is Research Use Only and not CE-marked. An IVDR conformity assessment submission is described as anticipated for the third quarter of 2027 or later, conditional on funding and organisational growth.
Which variant types are outside the scope?
Somatic and tumour variants are not supported, and repeat expansions are neither detected nor interpreted. The structural and copy-number module interprets calls already present in the VCF and does not replace upstream structural-variant calling.
Who is behind Folklore?
Helena Bioinformatics EOOD, registered at 14 Tsar Ivan Asen II Street in Sofia, Bulgaria. Folklore is one of four products alongside Undertone, Noodle and Evidence. The company describes itself as seed stage and founder-financed, without institutional investment.
Conclusion

Should you pick Folklore?

Folklore is one of the more transparent products a clinical genetics laboratory is likely to evaluate. The vendor publishes 131 canonical pages, a per-version changelog, an explicit limitations page and four validation cohorts with citable DOIs, and it goes as far as publishing reading rules that correct over-interpretation of its own figures. The engineering position is coherent: classification is rule-based, no model touches the class calculation, private genomic processing sits on dedicated hardware in Finland, and the analysis pipeline has no outbound network route.

The reservations are equally clear, and most of them come from the vendor's own documents. The product is Research Use Only and not CE-marked; an IVDR submission is a 2027 intention conditional on funding. Helena holds no certification of its own, and the ISO 27001 that appears in its impact assessment belongs to its hosting provider. The published agreement rate on the largest cohort is 67.0%. All three legal documents are labelled as candidates for legal review, and the documentation contradicts itself on whether the clinical assistant may call an external language model, and on the default retention period. The company is seed stage, founder-financed and currently seeking investment.

There is no published price, so any budgeting exercise starts with a sales conversation. The compensation is that the free public endpoint lets a laboratory judge classification quality on known variants before committing to anything at all. For a team that reads documentation closely and intends to keep a qualified specialist in the loop, that is an unusually honest starting point. For anyone needing a certified diagnostic device today, it is not the right tool.