
Autoheal
Autoheal is a multi-agent AI SRE platform for regulated enterprises. It triages alerts, builds evidence-backed root cause hypotheses, coordinates incident response in Slack or Teams, and writes postmortems, all under zero-trust governance with bring-your-own-cloud deployment.
What is Autoheal?
Autoheal is a multi-agent AI platform for site reliability engineering, unveiled on 10 March 2026 by Autoheal AI, Inc., a Delaware corporation. It targets one audience deliberately: SRE and production engineering teams inside regulated enterprises, where an incident has to be resolved quickly and explained afterwards to auditors.
At its centre sits the Production Context Graph, a continuously updated map linking infrastructure, application logic, production tooling and the tribal knowledge normally trapped in chat threads. The graph is built by autonomous exploration of the customer's observability, cloud and code stack, then refined through a reinforcement learning loop. Seven named agents work on top of it: the Curator maintains knowledge, the Triager separates signal from noise, the Hypothesizer builds root cause theories, the Coordinator runs the human side of the response, the Analyzer produces postmortems, the Verifier adversarially challenges the others, and the Tracer records why each decision was taken.
Three claims define the product. First, hallucination-proof reasoning: every step traces back to observable evidence, and a separate LLM-as-judge agent checks findings before they reach an engineer. Second, a zero-trust agentic runtime built on Cedar, the policy engine AWS uses for IAM, letting teams specify exactly what agents may do alone and what needs approval, with every tool call, argument and result logged. Third, deployment choice: alongside SaaS, Autoheal runs inside the customer's own cloud account with no outbound calls, in air-gapped mode, or through self-hosted runners, with LLM inference staying in the customer's VPC and encryption using the customer's own KMS keys.
Forty-two integrations cover the usual production stack — Datadog, Grafana, Prometheus, Sentry, New Relic, AWS, Kubernetes, GitHub, GitLab, Jira, ServiceNow, Slack, Teams and Zoom among them. On-call scheduling and incident response are built in, which is how Autoheal positions itself against the customary stack of PagerDuty or Opsgenie plus FireHydrant or incident.io plus a separate AI bot. Agents hold read-only access, and no customer data is used to train or fine-tune models.
What it does
- Normalise, deduplicate and categorise alerts arriving from several monitoring tools at once
- Gather live diagnostic signals across logs, metrics, traces and deployment history
- Rank evidence-backed root cause hypotheses and challenge them through an adversarial verifier agent
- Draft ready-to-run mitigation scripts, code patches and rollback commands for human approval
- Run incident response inside Slack, Microsoft Teams and Zoom, creating dedicated channels and paging the right responder
- Generate structured 5-Why postmortems with timeline, blast radius and preventive actions
- Open pull requests for accepted preventive fixes and keep the Production Context Graph updated
When to use Autoheal / When not to
A quick filter to help you decide if Autoheal is the right fit.
When to use Autoheal
- Site reliability and platform engineering teams in regulated industries running complex distributed systems
- On-call engineers drowning in duplicate alerts across PagerDuty, Datadog, Grafana and Sentry
- Engineering leaders trying to win back capacity lost to production firefighting and repetitive troubleshooting
- Organisations bound by audit and traceability requirements, where every AI action has to be logged and explainable
- Security-conscious enterprises that need the platform and its LLM inference to stay inside their own cloud perimeter
When not to use Autoheal
- Small teams looking for a self-service sign-up: access starts with a 45-minute sales demo and a negotiated order form
- Anyone shopping on price, since no rates are published anywhere and there is no free plan or trial
- Teams without an instrumented production stack, as the agents reason over existing logs, metrics, traces and deployments
- Non-technical users, given that setup requires admin rights and API credentials for at least one observability tool
- Engineers who want a mobile app for paging, because no iOS or Android application is listed
How to use Autoheal
A typical end-to-end flow, from setup to results.
- Contact support@autoheal.ai to have an account provisioned; there is no self-service sign-up
- Make sure you hold the Admin role in your organisation and have credentials for at least one observability tool
- Log in to your tenant instance, at an address of the form https://tenant.autoheal.ai, using email and password or your company SSO
- Open Integrations in the sidebar and connect your primary monitoring tool first, such as Datadog or Grafana, by supplying the requested API credentials
- Go to Instructions and create your first Skill document, describing a common incident type: symptoms, step-by-step remediation, escalation paths
- Save and publish the Skill so the agents can draw on it, and let the Production Context Graph auto-discover your topology and service dependencies
- Start an investigation from the Investigations panel by describing the incident in plain language
- Alternatively, mention @Autoheal in any Slack channel to launch an investigation and get findings posted back to the thread
- Keep the conversation going with follow-up questions, redirections or a wider scope, since the agent retains the full context
- Review and approve any proposed mitigation before it runs, and configure roles, permissions, alert notifications and Cedar authorisation policies in the admin section
Pros & Cons
Pros
- Agent governance is unusually explicit: a Cedar policy engine, defined authorisation boundaries, mandatory human approval before execution, and every tool call logged with its arguments and result
- Deployment options that are rare in this market — bring-your-own-cloud, air-gapped, self-hosted runners and bring-your-own-key, with LLM inference kept inside the customer's VPC
- A clear privacy stance: no customer data used to train or fine-tune models, read-only agent access, and zero outbound calls in BYOC mode
- Compliance transparency above the market average, with seven named subprocessors listed by purpose and location, a DPA on request, SOC 2 Type II, ISO 27001 and HIPAA on request
- Forty-two integrations spanning observability, infrastructure, code, data stores and collaboration, plus custom integrations built on request
- Genuine consolidation: AI investigation, on-call scheduling and incident response in one platform where the market usually requires three
- Detailed public documentation covering capabilities, security, quickstart, CLI and role-based administration
Cons
- No public pricing at all: the pricing page returns a 404 and buying means a 45-minute demo followed by a negotiated order form
- No free plan and no free trial are advertised, and the quickstart tells you to email support to get an account
- A very young product, unveiled in March 2026 and still described as validated with design partners
- No named customers, no case studies and no published customer outcomes
- No mobile application, despite paging being central and alerts going out by SMS, voice call, email and push
- No API reference page is published, although the API is contractually provided and underpins data portability
- The headline figures on the site — 2M USD per hour of downtime, four-hour MTTR, 73% on-call burnout, two thirds of postmortems skipped — are stated without sources
Pricing & Plans
No pricing is published. Autoheal has no pricing page — the /pricing address returns a 404 — and no amount, currency figure or tier appears anywhere on the site. There is no free plan and no free trial. Access is arranged commercially: section 7.1 of the terms of service states that fees are those set out in the applicable order form, quoted and payable in United States dollars unless otherwise specified, and non-refundable except where expressly stated. Invoices fall due within thirty days, with late payment accruing interest at 1.5% per month, and taxes are charged in addition. Prospective customers therefore have no way of situating the cost before speaking to sales.
Data, GDPR & hosting
A consolidated view of how Autoheal handles your data.
GDPR overview
GDPR compliance is claimed in writing. The security documentation lists GDPR and CCPA under current compliance, and the subprocessors page says the list exists to meet contractual and regulatory requirements including SOC 2 and GDPR. The privacy policy names its lawful bases — contractual necessity, legitimate interests and consent — and sets out access, rectification, deletion, portability and opt-out rights, exercised through security@autoheal.ai. A Data Processing Agreement is available on request, with advance notice of new subprocessors and a thirty-day objection window. The gaps matter too: no Article 27 EU representative is named, no data protection officer is designated, no transfer mechanism such as standard contractual clauses is cited, and the default infrastructure sits in the United States.
Who owns the data?
Customers keep full ownership. Section 6.1 of the terms of service states that the customer retains all right, title and interest in its Customer Data and that Autoheal claims no ownership. Section 6.2 grants Autoheal only a limited, non-exclusive licence to access, use, process and transmit that data as needed to provide, maintain and improve the service. Autoheal and its licensors keep the platform, software and documentation (section 9.1), and any feedback submitted may be reused freely (section 9.2). Data is never sold; it is shared only with seven named subprocessors, all located in the United States, including AWS, Google, Slack and Twilio.
Reuse rights
Customers can retrieve and reuse their own data without asking permission. The privacy policy documents export in JSON or CSV through the API, plus programmatic access to profile and organisational data. Section 6.5 of the terms adds a written-request route: Autoheal supplies a copy of Customer Data in a common electronic format within thirty days, though applicable fees may apply. After termination, section 11.4 keeps the data exportable for a further thirty days before Autoheal may delete it. Nothing in the terms restricts what a customer does with its own exported data. The restrictions run the other way: section 3.2 forbids reverse-engineering the service or using it to build a competing product.
Data retention & training
Hosting summary
Primary infrastructure is in the United States: the privacy policy states it runs on AWS us-east-1, and users outside the US consent to their information being transferred there. AWS hosts all customer data — database, compute, storage, email delivery and AI inference through Bedrock — and all seven listed subprocessors are US-based, including GitHub, Drata, Google, Slack, Twilio and Rippling. AI inference happens inside Autoheal's own AWS environment, and no customer data reaches a third-party model provider. The documentation mentions data residency options without naming alternative regions, and the trust page describes a multi-region architecture across separate geographies for disaster recovery. Customers needing tighter control can take the BYOC route, running the platform in their own cloud account with no outbound calls, or an air-gapped variant. Security measures include AES-256 encryption at rest through AWS RDS and KMS, TLS 1.2 or above in transit, logical multi-tenant isolation, private subnets, AWS Shield and AWS WAF.
Things to keep in mind
Risks and trade-offs to weigh before adopting Autoheal.
- Delegating diagnosis to an AI can erode engineers' own understanding of the system, which is precisely the knowledge the tool exists to preserve
- A wrong but well-argued root cause hypothesis can steer a mitigation in the middle of a crisis, adversarial verification notwithstanding
- Agents generate scripts, patches and rollback commands: an approval step rushed through under pressure turns a safeguard into a formality
- The platform concentrates credentials for observability, cloud, code repositories and databases in one place, creating a very wide access surface
- Read-only access and Cedar authorisation boundaries are settings, so an over-permissive policy quietly cancels the guarantee that was advertised
- Depending on a vendor that launched in March 2026 for a critical on-call function carries obvious continuity risk
- Automatically generated postmortems can be read without being discussed, and the value of a postmortem lies mostly in the discussion
Setup & Integrations
Technical difficulty
High, and openly so. Autoheal is built for SRE, DevOps and platform engineers, not for general users. You need an account provisioned by support, the Admin role, and API or OAuth credentials for at least one observability tool. Beyond the connection work, the Production Context Graph only becomes useful once the team writes Skill documents describing symptoms, remediation steps and escalation paths. Administration adds RBAC roles, alert notifications, Cedar authorisation policies, SSO and MFA. A CLI is documented, and BYOC or air-gapped deployments assume an in-house cloud team.
Deployment
Integrations
Behind Autoheal
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement Autoheal.
Frequently asked questions
What does Autoheal actually do?
Who is it built for?
How much does it cost?
Is there a free plan or a trial?
How can it be deployed?
Will my data be used to train AI models?
Where is the data hosted, and for how long?
What certifications and legal guarantees are offered?
Can the agents act on production by themselves?
Is there an API and a mobile app?
Should you pick Autoheal?
Autoheal is one of the more coherent propositions in the AI operations space: the marketing pages, the technical documentation and the legal texts all tell the same story, which is rarer than it should be. Where most AI tools promise reliability, Autoheal publishes the mechanism — Cedar for authorisation, an adversarial verifier agent, decision traces, a full audit trail. Its deployment options answer a real objection from regulated sectors: running inside the customer's own cloud account, air-gapped, or through self-hosted runners, with LLM inference and encryption keys never leaving the customer's perimeter. Compliance transparency follows the same line, with seven subprocessors named by purpose and location, a DPA on request, SOC 2 Type II and ISO 27001.
The reservations are equally clear. Pricing is completely opaque: no pricing page, no amounts, no free plan or trial, and a negotiated order form at the end of a sales demo. The product is very young, unveiled in March 2026 and still described as validated with design partners, with no named customers, no case studies and no published outcomes. Several claims lean on unsourced figures, no Article 27 EU representative is designated despite the GDPR claim, no postal address appears on the site, and there is no API reference page.
The target is narrow and openly stated. An organisation without an instrumented production stack and an on-call rotation will find nothing here to work with. A regulated enterprise already paying for PagerDuty, an incident coordinator and a separate AI bot, and unable to let production data leave its own cloud, will find a platform designed precisely around those constraints — and should press hard on price and on customer references before committing.
- Choosing a selection results in a full page refresh.
- Opens in a new window.