Folklore
Clinical genomics software that classifies germline variants from a VCF under the ACMG/AMP rules, adds HPO phenotype matching, literature evidence and prioritisation, then hands a traceable case record to a qualified geneticist for review and sign-off.
What is Folklore?
Folklore is clinical genomics software from Helena Bioinformatics, a Bulgarian company based in Sofia. It takes a single-sample VCF 4.1 or 4.2 file from a gene panel, exome or whole genome and turns it into a reviewable case record: annotated variants, an ACMG/AMP classification with every applied criterion exposed, phenotype-ranked genes, scored literature and a report a qualified specialist signs.
The pipeline runs in six stages. Variants are quality-filtered, with clinically significant ClinVar entries preserved regardless of score. Annotation goes through Ensembl VEP release 113 from a local offline cache, joined to locally stored reference data including gnomAD v4.1, ClinVar, dbNSFP, HPO, ClinGen and precomputed SpliceAI. The vendor is precise about scale here: 16 production databases sit in the core enrichment stage, while 45 reference databases and curated sources is the total across every module. Classification then applies the 2015 ACMG/AMP guidelines with all 28 evidence criteria in a Bayesian point framework. Folklore states plainly that classification is strictly rule-based and that no AI model determines pathogenicity; material relating to a version 4 of the standard is described as beta or draft, and no conformance with a final version 4 is claimed. PP3 and BP4 rest on BayesDel_noAF with ClinGen SVI-calibrated thresholds, with a separate guarded SpliceAI path. Scores such as SIFT, AlphaMissense, MetaSVM, DANN, PhyloP and GERP are displayed for clinical context and carry no ACMG vote of their own.
Around that core sit modules for mitochondrial DNA under MMDWG 2020, structural and copy-number variants under Riggs 2020, family and trio inheritance analysis, cohort analytics, phenotype matching and clinical screening. Phenotype matching orders candidates by HPO semantic similarity; it does not rewrite a classification and neither score establishes a diagnosis. An AI assistant answers questions about the case and shows the SQL it generates, but cannot alter a classification.
Private genomic processing runs on dedicated hardware in Helsinki, Finland, with no outbound network access from the analysis pipeline. Alongside the paid platform, Folklore publishes a free read-only MCP endpoint that classifies one public GRCh38 variant without any account. The vendor states the product is Research Use Only and not CE-marked.
What it does
- Classify germline variants from a VCF under the ACMG/AMP rules, with every applied criterion listed
- Rank genes and variants against the patient's HPO phenotype profile using semantic similarity
- Retrieve and score biomedical literature with traceable PMID, DOI and extracted context
- Prioritise findings into clinical tiers using age, sex, ancestry, family history and testing indication
- Interpret mitochondrial variants under MMDWG and copy-number changes under the ClinGen/ACMG Riggs framework
- Generate a tiered clinical interpretation report in PDF or DOCX with full evidence attribution
- Query one public GRCh38 variant for free through a read-only MCP endpoint, with no account or API key
When to use Folklore / When not to
A quick filter to help you decide if Folklore is the right fit.
When to use Folklore
- Clinical genetics laboratories running roughly 50 to 500 inherited-disease exome or genome cases a year
- Clinical geneticists who want evidence gathered, cross-referenced and documented before their own review begins
- Rare-disease, newborn-screening and carrier-screening teams needing phenotype-aware prioritisation
- Bioinformaticians and research groups analysing trios, families or whole cohorts under a locked classifier version
- Developers and AI agents that need a free, read-only variant-classification endpoint with no account or API key
When not to use Folklore
- Patients or members of the public: the service is explicitly not offered directly to patients for diagnosis or treatment decisions
- Somatic and tumour work: tumour-normal analysis and AMP/ASCO/CAP classification are outside the product scope
- Laboratories needing repeat-expansion disorders, which are neither detected nor interpreted
- Teams looking for upstream variant or structural-variant calling: Folklore interprets a VCF, it does not produce one
- Any workflow requiring a CE-marked in vitro diagnostic device, since the vendor states the product is Research Use Only
How to use Folklore
A typical end-to-end flow, from setup to results.
- Try the product first for free: search a single GRCh38 variant on the home page by coordinate, HGVS, SPDI or rsID
- Or connect an AI client to the public MCP endpoint over Streamable HTTP, with no account, API key or OAuth flow
- For laboratory use, create an account and pick an Individual, Laboratory or Distributor profile
- Expect Helena to ask for evidence of professional qualification or organisational identity before clinical features are enabled
- Have an authorised representative accept or sign the data processing agreement before any patient data is submitted
- Upload a pseudonymised single-sample VCF 4.1 or 4.2 file, splitting multi-sample files beforehand
- Set the quality filters for read depth, genotype quality and allelic balance
- Enter the patient's HPO terms directly, by HP:ID, or by AI-assisted extraction from a referral letter
- Choose a built-in or custom gene panel and one of the six screening modes, then run the analysis
- Review the fired criteria and evidence, reclassify where your judgement differs, and generate the signed report
Pros & Cons
Pros
- Classification is deterministic and auditable: each result carries its criteria, its sources and the classifier version
- No AI model sits inside the class calculation, and the vendor separates generated text from the decision explicitly
- Private genomic processing runs on dedicated EU hardware in Helsinki, outside multi-tenant cloud and outside claimed US jurisdiction
- The variant pipeline has no outbound network route by design, enforced at firewall and container level rather than in the application
- Public documentation is exceptionally dense, with 131 canonical pages, a per-version changelog and an explicit limitations page
- Validation is published and citable: four cohorts covering 197 retrospective cases, with public reports carrying OSF DOIs
- A free read-only machine interface needs no account or API key, and the adapter is open source under Apache-2.0
Cons
- No price is published anywhere: neither the product site nor the company site exposes a pricing page or any amount
- The vendor states the product is Research Use Only and not CE-marked, so it cannot serve as an in vitro diagnostic device
- Helena holds no certification of its own; the ISO 27001 cited in the impact assessment belongs to Hetzner, its hosting subprocessor
- The published cohort 4 result is 67.0% full plus clinical agreement over 100 cases, on the vendor's own figures
- All three legal documents carry a Candidate for legal review banner, so the vendor itself flags them as unvalidated
- The documentation contradicts itself on whether the clinical AI assistant may call an external language model API
- Scope is strictly germline: no somatic variants, no repeat expansions, no upstream variant or structural-variant calling
Pricing & Plans
A permanent free plan exists. The terms of use state that version 1.0 supports a free plan without a payment card, and that paid charges, taxes, renewal, cancellation and refund terms apply only if a paid offer is separately selected, with its ordering and billing terms presented before purchase. The public variant search and the read-only MCP endpoint are free and require no account. No paid price is published. Neither the product site nor the company site exposes a pricing page, and no amount in any currency appears in the vendor's own complete public text corpus. To investors, Helena describes the commercial model as B2B software with usage-based and volume access, and a route running from an own-data pilot to per-case purchase, prepaid volume packs and then an annual commitment or private instance. Pricing enquiries are directed to the company's general contact address.
Data, GDPR & hosting
A consolidated view of how Folklore handles your data.
GDPR overview
GDPR treatment is documented in unusual detail. A privacy notice, terms of use and data processing agreement are published in full, all version 1.0 effective 14 August 2026, alongside a public impact assessment summary (version 1.1, March 2026) carried out under Article 35 because genetic data is Article 9 special-category data. The laboratory is controller and chooses the Article 9(2)(a) or 9(2)(h) basis; Helena is processor. Private genomic processing stays inside the EEA, with standard contractual clauses reserved for any future transfer. Annex 3 of the agreement authorises a single subprocessor, Hetzner, for private customer data. The Bulgarian supervisory authority is named, and privacy@helena.bio is given as the data-protection contact, with an explicit caveat that this is not a formal DPO designation. All three legal documents carry a Candidate for legal review banner.
Who owns the data?
The customer keeps ownership. The terms state that you retain your rights in the data, files and materials you submit, and grant Helena Bioinformatics only the limited rights needed to host, copy, transmit, analyse and process that content in order to provide, secure, support and maintain the service and to meet legal obligations. For genomic case data the laboratory is the data controller and Helena acts strictly as processor on documented instructions; Helena is controller only for account, security, contact and billing records. Patient data must be pseudonymised before upload. At the end of the agreement the controller chooses return or deletion. Separately, Helena claims a perpetual right over feedback you volunteer, without identifying you or disclosing customer data.
Reuse rights
Customer data is processed only on the controller's documented instructions, for delivering, securing, supporting and maintaining the service. Results belong to the laboratory and may be used in its own professional or research workflow, subject to third-party database licences, attribution requirements and patient rights, and no permission request to the vendor is described for that reuse. Two secondary uses are disclosed and deserve attention. The impact assessment states that de-identified data may be processed for internal research and development, algorithm validation and platform improvement under a separate data use agreement, with Helena then acting as an independent controller on a legitimate-interest basis. And the privacy notice names OpenAI as processing bounded content from the public assistant, with provider storage disabled in Helena's configured request. Nothing on the site states either that customer data trains AI models or that it does not; the classifier itself is rule-based, not trained. No documented mechanism lets a customer opt out of the research and development use.
Data retention & training
Hosting summary
Private genomic processing and storage run on dedicated hardware physically located in Helsinki, Finland, identified as a Hetzner server, rather than on multi-tenant cloud infrastructure. The vendor gives four reasons: physical isolation from other customers, a fixed and verifiable data location, no exposure to US disclosure laws because the hosting provider is a European company, and exclusive root-level access for Helena so that no hosting employee reaches the operating system, storage or network. Annex 3 of the data processing agreement authorises Hetzner alone to receive private customer data, and no public website, analytics, email, payment or public-AI provider is authorised for it. Data travels over TLS 1.3, and no genomic or personal data is transferred outside the European Economic Area; standard contractual clauses are reserved for any future transfer. Vercel serves the public websites and front ends and may see network and usage metadata, but not private genomic content. Note one documented inconsistency: encryption at rest is given as AES-256 on the infrastructure page and as AES-128 at application level plus full-disk encryption in the impact assessment.
Things to keep in mind
Risks and trade-offs to weigh before adopting Folklore.
- Research Use Only and not CE-marked: do not deploy it as an in vitro diagnostic device or as the basis of a regulated report
- The only certification named on the site, ISO 27001, is held by the hosting subprocessor Hetzner and must not be read as Helena's own
- The documentation contradicts itself on external AI: the impact assessment and FAQ say all inference is local, while the no-external-calls page says the clinical assistant may use an external language model API
- Retention is stated twice and differently: a 90-day default in the impact assessment, a duration agreed with the laboratory on the retention page
- Automation bias is the real clinical risk: the vendor insists the geneticist decides, and a 67.0% agreement rate on its largest cohort shows why that insistence matters
- The public MCP endpoint accepts public variant data only; sending patient identifiers, phenotype or case context to it would be a disclosure incident
- Displayed predictor scores such as SIFT, AlphaMissense, PhyloP and GERP are context only and carry no ACMG vote, so do not read them as criteria
Setup & Integrations
Technical difficulty
Two very different levels. Public access is effortless: the variant search needs nothing, and connecting an AI client to the MCP endpoint takes one line of configuration over Streamable HTTP, with no account, API key or OAuth flow. Laboratory access is contractual rather than technical: you create an account, choose an Individual or Laboratory profile, may be asked to evidence professional qualification, and must have an authorised representative accept the data processing agreement before any patient data is uploaded. Nothing is installed: the product is web-based and a conforming pseudonymised VCF is uploaded over TLS.
Deployment
Integrations
Supported languages
Behind Folklore
Social
Resources
All the official URLs gathered for verification and reference.
Frequently asked questions
What file formats and genome builds does Folklore accept?
Does an AI model decide whether a variant is pathogenic?
Is a Folklore result a diagnosis?
Where is patient data stored?
What does Folklore cost?
Is there an API?
How many reference databases does Folklore use?
What is the regulatory status?
Which variant types are outside the scope?
Who is behind Folklore?
Should you pick Folklore?
Folklore is one of the more transparent products a clinical genetics laboratory is likely to evaluate. The vendor publishes 131 canonical pages, a per-version changelog, an explicit limitations page and four validation cohorts with citable DOIs, and it goes as far as publishing reading rules that correct over-interpretation of its own figures. The engineering position is coherent: classification is rule-based, no model touches the class calculation, private genomic processing sits on dedicated hardware in Finland, and the analysis pipeline has no outbound network route.
The reservations are equally clear, and most of them come from the vendor's own documents. The product is Research Use Only and not CE-marked; an IVDR submission is a 2027 intention conditional on funding. Helena holds no certification of its own, and the ISO 27001 that appears in its impact assessment belongs to its hosting provider. The published agreement rate on the largest cohort is 67.0%. All three legal documents are labelled as candidates for legal review, and the documentation contradicts itself on whether the clinical assistant may call an external language model, and on the default retention period. The company is seed stage, founder-financed and currently seeking investment.
There is no published price, so any budgeting exercise starts with a sales conversation. The compensation is that the free public endpoint lets a laboratory judge classification quality on known variants before committing to anything at all. For a team that reads documentation closely and intends to keep a qualified specialist in the loop, that is an unusually honest starting point. For anyone needing a certified diagnostic device today, it is not the right tool.
- Choosing a selection results in a full page refresh.
- Opens in a new window.