When your AI voice agent goes down, calls do not get answered. Leads are lost. Appointments are missed. Payments are not collected. Customer complaints pile up. Unlike a web application where a few minutes of downtime is an inconvenience, voice AI downtime directly translates to lost revenue and damaged customer relationships.
Yet many businesses sign AI voice agent contracts without closely examining the service level agreement (SLA). They trust the vendor's marketing page that says "99.9% uptime" without understanding what that number actually means, what happens when the platform does go down, and what recourse they have.
This guide explains what SLAs mean in the context of AI voice agents, what you should expect from your provider, how to negotiate better terms, and how to protect your operations against outages.
Understanding Uptime SLAs
What "99.9% Uptime" Actually Means
Uptime is typically expressed as a percentage of total time in a given period (usually monthly) that the platform is operational. Here is what common SLA levels translate to in allowed downtime:
| SLA Level | Allowed Downtime/Month | Allowed Downtime/Year |
|---|---|---|
| 99.0% | 7 hours 18 minutes | 3 days 15 hours |
| 99.5% | 3 hours 39 minutes | 1 day 19 hours |
| 99.9% | 43 minutes 48 seconds | 8 hours 45 minutes |
| 99.95% | 21 minutes 54 seconds | 4 hours 22 minutes |
| 99.99% | 4 minutes 22 seconds | 52 minutes 35 seconds |
The difference between 99.9% and 99.0% is massive -- 7 hours vs. 44 minutes of monthly downtime. Most enterprise-grade AI voice agent platforms target 99.9% or higher. Anything below 99.5% should raise concerns for production workloads.
What Counts as "Downtime"?
This is where SLAs get complicated. Vendors define downtime differently, and the definition matters as much as the percentage.
Watch for these common exclusions:
- Scheduled maintenance: Most SLAs exclude planned maintenance windows from uptime calculations. A vendor could have 99.9% "unscheduled" uptime but take the platform down for 4 hours every month for maintenance and still claim compliance.
- Partial degradation: If the platform is up but latency is 5x normal, some vendors do not count this as downtime. Your calls are technically completing, but they sound terrible.
- Regional outages: If the platform is down in India but up in the US, the vendor may not count this as a global outage. If your business is in India, this matters.
- Third-party failures: If the outage is caused by a telephony provider, cloud infrastructure provider, or LLM API, some vendors exclude this from their SLA even though it affects your calls.
- API downtime vs. call downtime: The dashboard might be unavailable (API downtime) while calls continue to work. Some vendors count only API downtime, not call processing downtime.
What you should demand in your SLA:
- Downtime is defined from your users' perspective: if calls are failing or quality is unacceptable, it is downtime regardless of the root cause.
- Scheduled maintenance windows are clearly defined, limited (e.g., no more than 2 hours per month), and occur during your lowest-traffic periods.
- Latency degradation beyond an agreed threshold (e.g., average latency exceeding 800ms) counts as a service degradation event.
- Regional uptime is measured for your operating regions, not globally.
The Five SLA Dimensions for Voice AI
Uptime is only one dimension of an AI voice agent SLA. A comprehensive SLA should cover five areas.
Dimension 1: Platform Availability
This is the standard uptime metric discussed above. The platform is either available and processing calls, or it is not.
What to expect:
| Provider Tier | Typical SLA | Financial Remedy |
|---|---|---|
| Startup/Early-stage | 99.0-99.5% | None or best-effort |
| Mid-market | 99.5-99.9% | Service credits (5-15% of monthly bill) |
| Enterprise | 99.9-99.99% | Service credits (10-30%) + escalation procedures |
Dimension 2: Call Quality (Latency)
Platform availability means nothing if the calls sound robotic. Latency -- the delay between the caller finishing speaking and the AI beginning its response -- is the primary quality metric.
What to expect:
| Metric | Good | Acceptable | Unacceptable |
|---|---|---|---|
| Average latency | Under 400ms | 400-700ms | Above 700ms |
| P95 latency | Under 600ms | 600-1000ms | Above 1000ms |
| P99 latency | Under 1000ms | 1000-1500ms | Above 1500ms |
Your SLA should include latency commitments, not just uptime. A platform that is "up" but delivering 2-second latency is effectively unusable for natural conversation.
Edesy maintains average latency of approximately 377ms. Read the detailed latency analysis to understand how this is measured and maintained.
Dimension 3: Call Success Rate
What percentage of attempted calls are completed successfully? Failed calls (dropped connections, SIP errors, audio failures) represent lost business regardless of whether the platform is technically "up."
What to expect:
- Call success rate should be above 98% for a production platform
- SLA should define what constitutes a "failed call" (connection failure, audio quality below threshold, call dropped before completion)
- Vendor should provide transparency on failure rates by category
Dimension 4: Support Response Time
When things go wrong, how quickly does the vendor respond?
Industry standard support SLAs:
| Severity | Definition | Response Time Target | Resolution Time Target |
|---|---|---|---|
| Critical (P1) | Platform down, all calls failing | 15-30 minutes | 4 hours |
| High (P2) | Significant degradation, many calls affected | 1-2 hours | 8 hours |
| Medium (P3) | Minor issue, some calls affected | 4-8 hours | 24-48 hours |
| Low (P4) | Feature request, minor bug, no call impact | 24-48 hours | Best effort |
What to look for:
- Named support contacts, not just a generic support email
- After-hours support for critical issues (voice AI does not stop at 5 PM)
- Escalation procedures documented in the contract
- Communication commitments during outages (status updates every 30 minutes for P1 events)
Dimension 5: Data Durability
Call recordings, transcripts, and campaign data are business-critical and often legally required. The SLA should address:
- Durability guarantee: What is the probability of data loss? (Standard for cloud storage is 99.999999999% or "eleven nines")
- Retention compliance: Does the platform maintain data for your required retention period?
- Backup and recovery: How often is data backed up? What is the recovery point objective (RPO) and recovery time objective (RTO)?
- Export guarantee: Can you export your data at any time, even during a contract dispute?
Financial Remedies and Service Credits
An SLA without financial consequences is a marketing document, not a contract. Here is how service credits typically work.
Standard Service Credit Structure
| Uptime Achieved | Service Credit |
|---|---|
| 99.9%+ | No credit (SLA met) |
| 99.5% - 99.9% | 10% of monthly bill |
| 99.0% - 99.5% | 20% of monthly bill |
| Below 99.0% | 30% of monthly bill |
What to Negotiate
- Higher credit percentages: Standard credits of 10-30% often do not cover the actual business impact of downtime. Negotiate higher percentages, especially for P1 events.
- Automatic credits: Credits should be applied automatically, not require you to file a claim and prove the outage.
- Termination rights: If the vendor consistently misses SLA targets (e.g., 3 months in a rolling 12-month period), you should have the right to terminate the contract without penalty.
- Exclusion of SLA credits from minimum commitments: If you have a volume commitment, SLA credits should reduce the bill without counting against your committed spend.
The Limitation of Service Credits
Be realistic: even generous service credits rarely compensate for the actual business impact of downtime. If your AI voice agent handles INR 50 lakh worth of calls per month and goes down for 4 hours during peak time, a 10% service credit of INR 5 lakh does not cover the lost business.
Service credits are a mechanism to keep the vendor accountable, not to make you whole. Your primary protection against downtime is the vendor's architecture and operational practices, which is why the technical due diligence matters more than the SLA percentages.
Protecting Your Operations
Build Redundancy
Do not rely on a single point of failure for your voice operations:
- Failover telephony: If the AI platform goes down, route calls to a simple IVR or voicemail system that captures caller information for callback
- Human backup: Maintain a small team of human agents who can handle calls during outages
- Multi-region deployment: If the vendor offers multi-region hosting, deploy across regions for geographic redundancy
- Secondary vendor: For mission-critical operations, maintain a secondary AI voice agent platform that can be activated during extended outages
Monitor Proactively
Do not rely solely on the vendor's status page. Implement your own monitoring:
- Synthetic testing: Schedule automated test calls at regular intervals (every 15-30 minutes) to verify the platform is operational
- Latency monitoring: Track actual call latency from your infrastructure, not just what the vendor reports
- Alert thresholds: Set alerts for latency increases, call failure rate increases, and integration failures before they become outages
- Independent status tracking: Log your own uptime data to verify the vendor's SLA reporting
Incident Communication Plan
When an outage occurs, your team needs to know immediately and respond according to a plan:
- Detection: Automated monitoring detects the issue
- Verification: Confirm the outage is on the vendor side (not your network or integration)
- Activation: Switch to backup systems (IVR, human agents, secondary vendor)
- Communication: Notify affected internal teams and, if appropriate, customers
- Tracking: Open a support ticket with the vendor, documenting the start time and impact
- Recovery: Once the vendor confirms resolution, test before switching back to primary
- Post-mortem: Review the incident with the vendor, document root cause and preventive measures
Questions to Ask Your Vendor
Before signing, ask these specific questions about reliability and SLA:
- What is your average uptime over the past 12 months? (Ask for actual data, not the SLA target)
- How many P1 incidents have you had in the past 12 months, and what was the average resolution time?
- Where are your data centers located, and what is the failover architecture?
- Do you have a public status page with historical incident data?
- What happens to in-progress calls when the platform goes down? Are they dropped, or do they complete?
- How do you handle LLM provider outages? (If OpenAI goes down, does your platform still work?)
- What is your scheduled maintenance window, and how much notice do you provide?
- Can you provide references from customers who have experienced an outage and can speak to the resolution process?
- What latency guarantees are included in the SLA, not just uptime?
- How are service credits calculated and applied?
For a comprehensive vendor evaluation framework, see the AI Voice Agent RFP Template: 50 Questions to Ask Vendors.
What Enterprise Buyers Should Expect in 2026
The AI voice agent market has matured significantly. Here is what enterprise buyers should consider baseline expectations:
- Uptime: 99.9% or better, with financial penalties for non-compliance
- Latency: Sub-500ms average, sub-800ms P95
- Call success rate: Above 98%
- Support: 24/7 for critical issues, with 30-minute response time
- Status page: Public, with real-time status and historical incident data
- Data residency: Options for India, US, EU, and Middle East
- Redundancy: Multi-region hosting with automatic failover
- Transparency: Willing to share actual uptime data, not just SLA targets
Any vendor that cannot meet these baseline expectations is not ready for enterprise production workloads.
Conclusion
The SLA is not just a legal formality -- it is a window into how the vendor thinks about reliability and customer impact. Vendors that offer vague SLAs with no financial consequences are telling you something about their confidence in their platform. Vendors that offer specific, measurable commitments with real financial penalties are telling you something different.
Invest the time to negotiate a comprehensive SLA before signing. And invest equally in your own operational resilience -- monitoring, failover, and incident response -- because no SLA prevents outages. It only determines what happens after they occur.
Explore Edesy's platform reliability and architecture or contact the team to discuss enterprise SLA terms for your deployment.