Respan
Respan is an LLM engineering platform that combines a multi-model gateway, tracing, evaluations and prompt management for teams running AI agents in production. A permanent free plan is available, and paid tiers start at $199 per month.
What is Respan?
Respan is an LLM engineering platform for teams running AI applications and agents in production. It was formerly known as Keywords AI, with the rebrand announced on 20 February 2026. Four building blocks share a single data plane: an AI gateway, tracing and observability, evaluations, and prompt management. Because they read the same data, a call routed through the gateway is traced automatically, with no extra instrumentation.
The gateway exposes one OpenAI-compatible endpoint onto a catalog the homepage describes as 1,000+ models, with ordered fallback, response caching, weighted load balancing, bring-your-own-key storage and spend caps. Tracing turns every LLM call, tool run, retrieval and agent turn into a span carrying its input, output, latency, cost, token count and model. Evaluations merge an LLM judge, a deterministic code check and human review into one score, run against datasets sampled from production traffic or imported from CSV, with experiments comparing score distributions. Prompt management adds a collaborative editor, versioning and one-click deployment.
Around those blocks, monitoring dashboards break requests, errors, cost, latency and tokens down by model, key or end user, with threshold monitors firing to email, Slack, Microsoft Teams or a webhook. Autonomous incident detection groups failing spans by error class, provider, endpoint and HTTP status. Six built-in behaviors are detected semantically, covering user frustration, laziness, unsafe content, jailbreak attempts, success and escalation, and custom ones can be defined. A Red Team module runs adversarial campaigns against an agent through a Python adapter, a Chat Completions-compatible endpoint or a hosted ShopBot sandbox.
Instrumentation comes in three forms: Python and TypeScript/JavaScript SDKs, OpenTelemetry OTLP, or the gateway proxy. Two MCP servers on mcp.respan.ai expose the platform and the documentation to Claude Code, Cursor and Codex CLI.
Respan claims more than 80 trillion tokens processed, around 80 million LLM requests a day, over a billion logs and 2 trillion tokens a month, 6.5 million end users and more than 100 customer teams. Named customers include Retell AI, Mem0, Gumloop, Lovable, Finta, Giga and AlphaSense. Behind the product is Keywords AI Inc., a ten-person Y Combinator W24 company based in Alameda, California.
What it does
- Route every LLM call through a single OpenAI-compatible endpoint
- Trace each LLM call, tool run, retrieval and agent turn as a span carrying its input, output, latency, cost and tokens
- Fail over automatically to a backup model when a provider goes down or throttles
- Score outputs with an LLM judge, a deterministic code check or human review
- Set budgets and rate limits per key, per customer or across the whole organization
- Trigger an alert as soon as cost, errors, latency or token volume crosses a threshold
- Compare two prompt or model versions on a test set before shipping the change
When to use Respan / When not to
A quick filter to help you decide if Respan is the right fit.
When to use Respan
- Engineering teams running LLM applications and agents in production, from a solo developer with one agent to a large team with several applications and compliance obligations
- Teams on mixed or framework-free stacks, since the platform is explicitly framework-agnostic
- Teams that would rather have a gateway, observability, evaluations and prompt management in one product than pay for four separate subscriptions
- Organizations under HIPAA or SOC 2 obligations, with a BAA available and SOC 2 Type II reports shared under NDA
- Teams already instrumented with OpenTelemetry, since OTLP ingestion is accepted on every plan including the free one
When not to use Respan
- Teams built entirely on LangChain or LangGraph, for which Respan's own comparison page concedes that LangSmith is the path of least resistance
- Teams that need integrated LangGraph deployment, which is explicitly not included
- Teams that require self-hosting today, since the FAQ states it is not currently supported
- Latency-critical applications where the 50 to 150 ms the gateway adds to each request is unacceptable, although the tracing SDK sits off the request path and avoids it
- Non-technical users and consumer use, since the tool assumes instrumenting code or routing API traffic and offers no iOS or Android app
How to use Respan
A typical end-to-end flow, from setup to results.
- Create an account on platform.respan.ai; the free plan needs no credit card
- Generate an API key from the API keys page in settings, which displays it only once
- Add credits or connect your own provider key from the Integrations page
- For the fastest route, run npx @respan/cli setup, which detects the project language, installs the SDK, instruments the code and creates the API key
- Otherwise pick a starting point: the gateway quickstart to route traffic, the tracing quickstart to instrument an existing application, or the evals quickstart to score quality
- To work without an SDK, point an existing OpenTelemetry exporter at Respan or send data through the JSON API
- Watch the first traces arrive within seconds, and make sure short-lived processes such as scripts or serverless functions do not exit before the SDK flushes its buffer
- Invite the team from the members page and assign the Owner, Admin, Member or Annotator role
- Write system and user prompts with {{variable_name}} placeholders, test them in the playground, then Commit and Deploy
- Call deployed prompts from your application by prompt ID
Pros & Cons
Pros
- Four products in one on a shared data plane: a routed call is traced automatically, and evaluation scores land next to the span that produced them, removing the need for four separate subscriptions
- Framework-agnostic and OpenTelemetry-native, so it works without a proprietary SDK and fits stacks that use no framework at all
- A genuinely usable free plan: 100,000 logs, 1,000 scores, five datasets, unlimited seats and OTLP ingestion included, with no credit card required
- Serious compliance for a young product: SOC 2 Type II, HIPAA with a BAA, GDPR, and ISO 27001 displayed on the homepage
- Native cost control: cost and token tracking on every span, budgets and caps per key, per customer or organization-wide, and caching that removes the cost of repeated calls
- Resilience built into the gateway: model fallback, weighted load balancing, retries and caching, with a 99.9% uptime SLA on Team and 99.99% on Enterprise
- Fast onboarding and agent-readable documentation: npx @respan/cli setup instruments a project automatically, and the docs are exposed as .md pages, llms.txt and two MCP servers
Cons
- Short retention on the lower tiers: 7 days on Free and 30 days on Team, which is tight for investigating incidents after the fact
- Self-hosting is contradictory across the site: the FAQ says it is not currently supported, while the pricing page lists it for Enterprise
- The gateway adds roughly 50 to 150 ms of latency to each request
- The $199 per month headline assumes annual billing with a 17% discount; no month-to-month rate is published
- Compliance and seats are billed on top: HIPAA is a $249 per month add-on, and Team includes only five seats before $15 per additional member
- No subprocessor list is published, and there is no written position anywhere on whether customer data is used to train models
- The terms make all purchases non-refundable and grant a very broad license over contributions published in the Services
Pricing & Plans
Respan operates on a freemium model. The Free plan costs nothing, is permanent and requires no credit card. The first paid tier, Team, is listed at USD 199 per month billed annually, a rate presented with a 17% annual discount; Enterprise pricing is quoted on request. Usage beyond plan allowances is charged at USD 8 per additional 100,000 logs and USD 1 per additional 1,000 scores, additional Team seats at USD 15 per member, and HIPAA compliance is a separate add-on at USD 249 per month. All amounts are in US dollars and processed by Stripe, which accepts Visa, Mastercard, American Express, Discover and PayPal. Subscriptions renew automatically and may be cancelled at any time, with effect at the end of the paid period; purchases are non-refundable.
- the full platform with 100
- 000 logs
- 1
- 000 scores
- 5 datasets
- 2 evaluators
- 5 prompts
- 1 workspace
- everything in Free plus unlimited datasets
- evaluators and prompts
- 10
- 000 scores
- 5 included members and USD 15 per additional member
- 30-day retention
- 8
- 400 requests per minute and 4
- everything in Team plus custom packages and SLAs
- volume discounts
- a dedicated support engineer
- priority support
- a HIPAA BAA
- SAML SSO
- retention management
- PII masking
- HIPAA compliance add-on (USD 249 per month)
- available across the tiers
Data, GDPR & hosting
A consolidated view of how Respan handles your data.
GDPR overview
Respan claims GDPR alignment explicitly. The homepage states that the service is operated under GDPR, and the compliance documentation confirms support for every data subject right: access, rectification, erasure, restriction of processing, portability and objection. The privacy policy lists its legal bases (consent, legal obligation, vital interests, contract performance, legitimate interests) and covers the EEA, the United Kingdom, Switzerland and Canada, with a right to lodge a complaint with the national supervisory authority, or the FDPIC for Switzerland. Customers can execute a DPA covering processing instructions, security measures and breach procedures, and EU data residency is available on request. Rights requests go to security@respan.ai. Two gaps remain: no Article 27 EU representative and no data protection officer is named anywhere, for a US company with no declared European establishment.
Who owns the data?
Under the terms of use, customers keep full ownership of their contributions: Keywords AI Inc. asserts no ownership over that content or the intellectual property rights attached to it. In exchange, publishing content inside the Services grants the company an irrevocable, perpetual, worldwide, transferable and sub-licensable license over it. Section 22 adds that Respan retains transmitted data in order to operate the service, that the customer alone remains responsible for that data, and that Respan disclaims any liability for its loss or corruption, the customer waiving any right of action on that ground. Personal data is deleted or anonymized once no legitimate purpose remains.
Reuse rights
The privacy policy in force since 25 February 2026 sets out what is collected: names, email addresses, account identifiers, passwords, billing addresses and card numbers, with payment data handled and stored entirely by Stripe. Technical data is gathered automatically, including IP address, browser, operating system, language, referring URL, country, usage logs and error reports. Section 6 states that the AI products are delivered through third-party AI service providers, named as Anthropic, Google Cloud AI and OpenAI, and that inputs, outputs and personal data are passed on to them. Google Analytics is used, advertising and remarketing features included, with opt-out available through Google's browser add-on. No sensitive information is processed. Respan declares that it has not sold or shared personal information for business or commercial purposes in the preceding twelve months, and commits not to do so. One silence deserves attention: nowhere on the site does the company state whether customer data is used to train its own models, and no opt-out mechanism is described.
Data retention & training
Hosting summary
Respan runs on Amazon Web Services, with the application hosted on Amazon ECS. PostgreSQL holds persistent data, ClickHouse powers analytics and observability, and Redis carries the event queue processed by Celery workers. Primary data centers are in US East (Virginia) and US West (Oregon); European data residency is available on request, and the architecture documentation states that data never leaves the geographic region a customer specifies. Traffic is encrypted with TLS 1.2 or above, the compliance page citing TLS 1.3, data at rest is encrypted with AES-256 in PostgreSQL and ClickHouse, and API keys are stored as SHA-256 hashes. Access relies on mandatory MFA, least-privilege RBAC, just-in-time administrative access and regular access reviews, with no default employee access to customer data. Continuity targets are a four-hour RTO and a one-hour RPO, backed by automated daily backups replicated across regions, and customers are notified of incidents within 24 hours by a dedicated response team. Operational security includes internal audits, weekly testing, continuous AWS CloudWatch monitoring, code review, vulnerability scanning and penetration testing.
Things to keep in mind
Risks and trade-offs to weigh before adopting Respan.
- The site contradicts itself on self-hosting: the FAQ states it is not currently supported, the pricing page lists Cloud and Self-hosted for Enterprise, and the comparison article says Enterprise tier only
- Nowhere does Respan state in writing whether customer data is used to train models, and no opt-out is offered, which leaves an open question for teams routing sensitive prompts
- No subprocessor list is published, even though third parties process data: AWS, Stripe, Anthropic, Google Cloud AI, OpenAI and Google Analytics
- No Article 27 EU representative and no data protection officer is named, for a US company that claims GDPR compliance; EU data residency exists only on request, with United States hosting by default
- Headline figures vary by page: the model catalog is described as 1,000+, 500+ or 250+ depending on where you look, and encryption in transit is stated as TLS 1.3 on the compliance page but TLS 1.2 or above in the architecture documentation
- ISO 27001 appears among the certifications on the homepage but is absent from the compliance page, which documents only SOC 2, HIPAA and GDPR
- Commercial terms deserve a read: subscriptions renew automatically, all purchases are non-refundable, the cookie policy still carries a 1 January 2025 date and a contact address that differs from the rest of the site, and the respan.ai domain itself was only registered on 13 January 2026
Setup & Integrations
Technical difficulty
Setup is developer-level but short. The quickest route is npx @respan/cli setup, which detects the project language, installs the SDK, instruments the code and creates the API key in a few minutes. Teams already using an OpenAI-compatible SDK simply change the base URL to route through the gateway, and teams already on OpenTelemetry point their existing OTLP exporter at Respan with no proprietary SDK. Nothing has to be provisioned, since the product is fully hosted. One documented pitfall: a process that exits before the SDK flushes its buffer, such as a script or a serverless function, loses its traces.
Deployment
Integrations
Behind Respan
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement Respan.
Frequently asked questions
Are Respan and Keywords AI the same product?
Is Respan a gateway or an observability platform?
What security certification does Respan hold?
How long are logs kept?
Can Respan be self-hosted?
Does Respan support HIPAA and sign a BAA?
Which models and providers can I reach through the gateway?
How much latency does the gateway add?
Is OpenTelemetry supported?
Which SDKs, frameworks and AI tooling are covered?
Should you pick Respan?
Respan is a young product carrying an unusually mature compliance posture. Out of Y Combinator's W24 batch, renamed from Keywords AI in February 2026 and backed by a USD 5M seed round the following month, it already claims more than a hundred customer teams and names several of them.
Its real strength is consolidation. A gateway, tracing, evaluations and prompt management normally mean four vendors and four invoices; Respan puts them on one data plane, so a call routed through the gateway is traced without further work and an evaluation score sits next to the span that produced it. Add OpenTelemetry ingestion on every plan, a free tier with unlimited seats, and spend caps per key, and the case for engineering teams running agents in production is easy to follow.
The reservations are just as concrete. Self-hosting is described as unsupported in the FAQ and offered on Enterprise on the pricing page, a contradiction the site never resolves. Retention is seven days on Free and thirty on Team, which is short for after-the-fact investigation. And nowhere does Respan state in writing whether customer data is used to train models, nor does it publish a subprocessor list, even though inputs and outputs pass through Anthropic, Google Cloud AI and OpenAI.
The audience is clear: technical teams with code to instrument or API traffic to route, not general consumers, and there is no mobile application.
Three points are worth settling with the vendor before committing: the month-to-month price without annual commitment, whether self-hosting is genuinely available today, and the exact scope of the ISO 27001 certification shown on the homepage but absent from the compliance documentation.
- Choosing a selection results in a full page refresh.
- Opens in a new window.