BenchGen logo
Evaluation Benchmarks · Synthetic Data

BenchGen

BenchGen is a simulation and benchmarking platform that runs AI agents through digital-twin environments, scoring every decision step, capturing full trajectories and turning failed runs into reinforcement-learning data for teams shipping agents into production.

Active GDPR compliant Free plan Freemium API available 18+ Verified by Guidaio
Overview

What is BenchGen?

BenchGen is benchmarking and simulation infrastructure for AI agents, published by Benchgen, Inc. Instead of scoring a single model answer, it builds what the company calls digital-twin companies: sandboxed replicas of the systems an agent will actually touch - APIs, databases, workflow engines and communication channels - and lets the agent operate inside them freely. Every tool call, retrieval step and intermediate decision is recorded as a trajectory, scored step by step and kept as an auditable log. The site puts the intent bluntly: most evaluation tools test prompts, BenchGen tests agents.

The platform comes in three parts. Eval registers a model - uploaded, imported from Hugging Face, or simply reached through any OpenAI-compatible endpoint such as OpenRouter or Mistral - and runs it against an evaluation environment, returning task completion rates, per-step accuracy, located failure modes and comparisons across model versions or prompt variants. Train takes the failing cases, exported as a labelled dataset, fine-tunes a model with LoRA, merges the adapter and serves it back. Agents, branded Agentspace, builds agents from Topics and Actions and wires them to external services. The documentation calls the resulting cycle Simulate, Act, Capture, Measure, Generate, Improve.

A second product line, BenchGen for Hermes Agent, is a command-line tool installed with a single pip command. It reads trajectories already sitting in a local Hermes directory, without instrumentation or API keys, and returns a quality score split into tool-call accuracy, goal completion, error recovery, skill coverage and memory utilisation. That line is in early access, announced for the third quarter of 2026.

Deployment covers cloud, on-premise, air-gapped and sovereign environments: the platform installs on a Kubernetes cluster through one Helm release, or from the DT Edge Platform interface. A public catalogue of more than a hundred benchmarks - SWE-bench, MMLU-Pro, GPQA Diamond, TerminalBench, ARC-AGI, plus in-house sets such as FinArena - and a leaderboard ranking over seventy language models are readable without an account. BenchGen claims 1,400 agents improved, over two million trajectories captured and 750 reinforcement-learning environments, and cites deployments in NATO-member defence organisations and on national GPU infrastructure in Turkiye.

What it does

  • Benchmark an AI agent inside a simulated replica of the systems it will actually operate
  • Capture and score the full decision trajectory, step by step, with auditable logs
  • Locate failure modes and see where in the workflow they occur
  • Export failing cases as a labelled dataset and fine-tune a model with LoRA
  • Evaluate any OpenAI-compatible endpoint or any public Hugging Face model
  • Compare model versions, providers and prompt variants against the same environments
  • Run the whole platform in your own cloud, on-premise or air-gapped infrastructure
Audience

When to use BenchGen / When not to

A quick filter to help you decide if BenchGen is the right fit.

When to use BenchGen

  • Machine-learning and MLOps teams that must prove agent reliability before a production rollout
  • Mission-critical and regulated organisations - defence, energy, finance, public sector - that need auditable, deterministic evaluation
  • Platform teams running sovereign, on-premise or air-gapped infrastructure, since the product installs on Kubernetes through a Helm release
  • Applied research teams that want to recycle agent runs into reinforcement-learning datasets and LoRA fine-tuning
  • Hermes Agent developers who want to score the trajectories already sitting on their machine with a single command

When not to use BenchGen

  • Anyone who only needs to test a single prompt: the product is built around multi-step agent behaviour, not one-shot output quality
  • Non-technical buyers looking for a self-serve checkout, since the site publishes its price grid only inside its llms.txt file and shows a Contact us button everywhere else
  • Teams that need a documented public API today: the page labelled API reference is still an unmodified Mintlify sample about a plant store
  • Users expecting a mobile app or a browser extension, neither of which exists
  • Procurement teams that require a signed data processing agreement and a published subprocessor list before starting, as neither is available on the site
Get started

How to use BenchGen

A typical end-to-end flow, from setup to results.

  1. Try the free Skill Checker, which needs no account and no installation, or browse the public leaderboard
  2. Create an account on the platform at app.benchgen.com, or install BenchGen yourself with a Helm release on a Kubernetes cluster
  3. Spin up an isolated runtime on cloud or on-premise GPU and CPU servers, choosing an air-gapped environment if security requires it
  4. Connect the enterprise data sources - CRM, ERP, databases, data warehouse, support tickets, APIs - that make the simulation realistic
  5. Register the model to evaluate: upload an archive, import it from Hugging Face, or point BenchGen at any OpenAI-compatible endpoint with its URL and key
  6. Pick an evaluation environment from the public catalogue, or build your own as a zip bundle driven by a competition.yaml file
  7. Launch the benchmark and follow the run in real time, with usage, tokens, latency and spend monitored per model
  8. Read the report: completion rate, per-step accuracy, located failure modes, comparisons between versions, and auditable logs of every action
  9. Export the failing cases as a labelled dataset, fine-tune with LoRA, merge the adapter and put the improved model back into service
  10. For a Hermes agent, install the command-line tool with pip, run a scan of the local trajectories, open the report and export the best runs for fine-tuning
Quick read

Pros & Cons

Pros

  • Evaluation covers the whole decision path rather than the final answer, so a correct result reached through a broken reasoning chain is still visible
  • The loop is closed inside one product: evaluate, export the failures, fine-tune with LoRA, redeploy
  • Open by design - Hugging Face imports and any OpenAI-compatible endpoint, rather than a closed model roster
  • Self-hosting is genuinely documented, with a Helm chart and a full values reference, and air-gapped deployment is a stated design goal
  • A permanent free tier of 50 benchmark runs per month makes it possible to judge the tool before committing
  • The public benchmark catalogue and leaderboard are open to read without an account
  • The site names its competitors and compares itself to them point by point instead of arguing in the abstract

Cons

  • The price grid is published only inside the llms.txt file: there is no pricing page, and the site still shows a Contact us button everywhere
  • The page labelled API reference is the untouched Mintlify sample about a plant store, so the advertised REST API has no real documentation
  • Several pages linked from the footer and from llms.txt simply do not exist and quietly serve the homepage instead
  • Blog, changelog and glossary are empty shells with no statically rendered content, leaving no verifiable release history
  • The Hermes line is in early access and announced for the third quarter of 2026, so part of what the site describes is not shipping yet
  • The SOC 2, ISO 27001 and HIPAA badges carry no linked report, certificate or trust page
  • The company is very young - domain registered in March 2026, no Wayback Machine capture - with no postal address and no contact page
Pricing

Pricing & Plans

BenchGen offers a permanent free plan. The Starter tier costs nothing and covers 50 benchmark runs per month across 5 evaluation environments; the lowest paid tier, Pro, is priced at 49.00 USD per month for 2,000 runs and unlimited environments; the Enterprise tier is quoted on request. Prospective customers should note that this grid appears only in the site's llms.txt file, dated 29 June 2026, and not on a pricing page: the homepage carries a Contact us call to action and structured data declaring a price of zero. The figures should therefore be confirmed in writing before any commitment. The Skill Checker tool remains free and requires no account.

Plan 1
Starter
  • free - 50 benchmark runs per month
  • 5 evaluation environments
  • trajectory capture and export
  • community support
  • basic failure analysis
Plan 3
Enterprise
  • custom pricing - unlimited benchmark runs
  • custom environment builds
  • full audit trail and compliance
  • dedicated infrastructure
  • SSO and SAML
  • 24/7 SLA support
Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how BenchGen handles your data.

GDPR overview

The privacy policy, effective 23 March 2026, addresses the GDPR concretely rather than generically. Section 3 sets out legal bases for users in the EEA and the United Kingdom - contract, legitimate interests, consent and legal obligation. Section 9 lists rights of access, rectification, erasure, portability, objection and restriction, withdrawal of consent and marketing opt-out, with a stated response time of 30 days; requests go to contact@benchgen.com. Section 8 mentions standard contractual clauses for transfers where required. What is missing is equally clear: the site never claims GDPR compliance in so many words, names no data protection officer, designates no Article 27 representative in the European Union, publishes no data processing agreement and no subprocessor list. SOC 2, ISO 27001 and HIPAA appear as footer badges without any linked report or certificate.

Who owns the data?

Benchgen, Inc. publishes the service. Under section 4 of the terms of use, content you submit stays yours, but you grant Benchgen a worldwide, non-exclusive, royalty-free licence to use, reproduce, process and display it solely to operate and improve the service, and you warrant that you hold the rights to it. The privacy policy states plainly that Benchgen does not sell personal information, while allowing disclosure to service providers handling hosting, analytics and email, to a buyer in a merger or acquisition, and to authorities under legal obligation. In sovereign deployments the company states that model weights, evaluation data and benchmark results never leave the customer's controlled environment.

Reuse rights

The terms grant the customer the right to use the service and its outputs, and the licence flowing back to Benchgen is limited by its own wording to operating and improving the service. Nothing in the documents collected restricts what a customer does with its own benchmark reports, trajectory exports or fine-tuned adapters, and the platform is explicitly designed so that these artefacts can be exported: failing cases become labelled datasets, trajectories become reinforcement-learning training material. The privacy policy lists the purposes of processing as operating and improving the service, answering requests, sending newsletters, usage analysis, security and fraud prevention, legal obligations and the handling of evaluations. It does not state whether customer content feeds model training, which leaves that question open.

Data retention & training

Retention summary
Benchgen keeps personal information for as long as necessary to fulfil the purposes set out in its privacy policy, unless a longer retention period is required by law. Once data is no longer needed it is securely deleted or anonymised. Users may request deletion at any time by writing to contact@benchgen.com, and the company undertakes to respond within 30 days. No specific retention period is published for any category of data - not for account information, not for benchmark and evaluation records, not for usage logs - so the actual duration is left to the vendor's discretion and would need to be settled contractually by anyone with a defined retention requirement.
Trains on customer data
Unclear
Subprocessors disclosed
No
GDPR contact

Hosting summary

The privacy policy states that Benchgen is based in the United States and that information may be transferred to, stored and processed there or in other countries where its service providers operate, with standard contractual clauses used where required. No hosting provider is named and no cloud region is specified, so the jurisdiction can only be described as declared, not verified. The public website itself resolves to an IP address in Turkiye, on a Turkish network operator - that is the marketing site, and it says nothing about where customer data lives. The picture changes entirely under self-hosting: installed through a Helm release on the customer's own Kubernetes cluster, or in an air-gapped environment, the data never leaves the customer's infrastructure, and the company states that model weights, evaluation data and benchmark results remain inside that controlled environment.

Hosting countries
🇺🇸 United States
Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting BenchGen.

  • Pricing is published in one place only, a file aimed at AI answer engines, while the site says Contact us and its structured data declares a price of zero; confirm any figure in writing before committing
  • The advertised REST API has no real documentation, so do not scope an integration before seeing an actual specification
  • SOC 2, ISO 27001 and HIPAA are shown as badges with no linked report or certificate, and the security and trust pages do not exist
  • Impact figures and named customer results carry no source or date, and should be treated as vendor claims rather than verified outcomes
  • The vendor's position on training its models with customer content is not stated in the privacy policy, which matters if you connect a CRM, ERP or support-ticket history to build the digital twin
  • The terms cap total liability at the amount paid over the previous twelve months and allow suspension or termination at any time without notice, under Delaware law and exclusive Delaware jurisdiction
  • Treating a benchmark score as proof that an agent is safe to deploy is the human failure mode this kind of tool invites: a good score measures behaviour in a simulation, not in your production reality
Setup

Setup & Integrations

Technical difficulty

Effort depends entirely on the entry point. The Skill Checker needs nothing at all: no account, no installation. The Hermes command-line tool sits in the middle - one pip install, then a scan that reads local trajectories with no configuration and no API key. The full platform is demanding: an isolated GPU or CPU runtime, enterprise data sources connected, and evaluation environments imported or built as zip bundles. Self-hosting adds a Kubernetes cluster and a Helm release to parameterise. The audience is explicitly technical.

Deployment

Web appAPI

Integrations

Hugging Face OpenRouter Mistral VS Code Cline OpenClaw Kubernetes Helm DT Edge Platform Hermes Agent

Supported languages

English
Company

Behind BenchGen

Company name
Benchgen, Inc.
Founded
INFORMATION_NOT_FOUND
Country of origin
🇺🇸 United States
UBO
INFORMATION_NOT_FOUND
UBO country
INFORMATION_NOT_FOUND
Domain registrar country
🇺🇸 United States
Legal contact
Support contact

Social

Official links

Resources

All the official URLs gathered for verification and reference.

Compare

Alternatives

Tools that compete with or complement BenchGen.

L LangSmithA Arize PhoenixD DeepEval
FAQ

Frequently asked questions

What does BenchGen actually test?
It tests agents rather than prompts. An agent is run inside a simulated replica of the systems it would touch in production, and every step of its decision path - tool calls, retrievals, intermediate actions - is scored, not just the final output.
Why are standard LLM benchmarks not enough?
Benchmarks such as MMLU or HumanEval score a single response. A model that scores 90 percent on them can still hallucinate a tool call or collapse midway through a multi-step workflow, which is precisely what trajectory-based evaluation is designed to surface.
What does it cost?
Three tiers are published: Starter is free with 50 benchmark runs per month, Pro costs 49 USD per month for 2,000 runs and unlimited environments, and Enterprise is quoted on request. Be aware that this grid appears only in the site's llms.txt file, not on a pricing page.
Is there a free plan or a free trial?
There is a permanent free plan, the Starter tier, and the Skill Checker tool is open without an account. No time-limited free trial is advertised anywhere on the site.
Can BenchGen run on our own infrastructure?
Yes. The platform installs on a Kubernetes cluster through a single Helm release, with a complete values reference published, or through the DT Edge Platform interface. Sovereign, on-premise and air-gapped deployments are stated design goals, and the company says model weights, evaluation data and results never leave the controlled environment.
Which models can be evaluated?
A model already deployed on the platform, a public model pulled from the Hugging Face Hub, or any external OpenAI-compatible endpoint - OpenRouter and Mistral are given as examples - reached by supplying its URL and key.
Does BenchGen have an API?
The site advertises an agent-native REST API and the documentation describes OpenAI-compatible inference endpoints that the platform deploys. However, the page labelled API reference is still the unmodified Mintlify sample about a plant store, so no real API specification is published today.
Can it train models, or only evaluate them?
Both. Failing benchmark cases are exported as a labelled dataset, the Train section fine-tunes a model with LoRA, merges the adapter into the base model and serves it back for inference and re-evaluation.
Which tools does BenchGen compare itself to?
A comparison table on its Hermes page names LangSmith, Arize Phoenix and DeepEval, arguing that generic evaluation tools do not understand Hermes trajectory files, skill files or the distinction between prompt-level and model-level failures.
Is there a minimum age to use the service?
Yes. Section 1 of the terms of use sets the minimum age at 18.
Conclusion

Should you pick BenchGen?

BenchGen addresses a real and well-framed problem: agents that look convincing in a demo and fail quietly in production, with no way to tell the difference in advance. Its answer - simulate the systems the agent will touch, score every step of the decision path rather than the final answer, then recycle the failures into training data - is coherent, and the technical substance behind it holds up. The documentation is substantial, the platform accepts models from Hugging Face and any OpenAI-compatible endpoint rather than locking users into a roster, LoRA fine-tuning sits in the same product as evaluation, and self-hosting on Kubernetes is documented down to the Helm values. The public benchmark catalogue and the leaderboard can be read without an account, which is more openness than this category usually offers.

The reservations are about maturity rather than intent. Pricing is now legible - a free Starter tier, Pro at 49 USD per month, Enterprise on quote - but it lives only in a file written for AI answer engines while the site itself still says Contact us. The page presenting the API is an untouched sample template about a plant store. Pages linked from the footer resolve to the homepage, and the blog, changelog and glossary are empty. The compliance badges carry no evidence, the company is barely six months old by its domain registration, and it publishes no postal address.

The sensible approach follows from that. The free tier exists precisely so that a team can measure the product against its own agents rather than against its marketing, and any commercial or compliance commitment should be confirmed in writing first.