ElevenAgents Expands From Voice Calls to Multichannel Service
SMS, Telegram, Intercom, and Freshdesk widen the same-agent model, while channel-specific testing, disclosure, permissions, and auditability become the real deployment test.
By WhatAI Editorial Team ·
ElevenLabs is extending ElevenAgents beyond the voice call. Its July 28, 2026 product announcement added SMS and Telegram responses plus ticket handling in Intercom and Freshdesk. Those channels join phone, web, Zendesk, Slack, and WhatsApp in a redesigned dashboard. The promise is simple: define one agent with its knowledge, tools, procedures, and guardrails, then deploy it wherever customers ask for help. The operational reality is more demanding. A useful phone agent, a safe SMS responder, and a reliable ticketing agent can share a policy brain, but they should not behave as if the channels are interchangeable.
The expansion shows how far ElevenLabs has moved from its original reputation as a realistic text-to-speech generator. The company now sells three connected product families. ElevenCreative covers voice, transcription, dubbing, music, sound effects, Studio, Flows, and multimodal production. ElevenAPI exposes models and workflows to developers. ElevenAgents combines speech, language models, knowledge, tools, procedures, analytics, and customer channels. For buyers, the comparison is no longer just whether one generated voice sounds more human than another. The harder question is whether the platform can run a governed customer interaction from first message to final action.
The new channels are relevant because customers do not stay inside one surface. A person may start with a web chat, reply to an SMS, call when the issue becomes urgent, and later respond to an email ticket. A business does not want four conflicting bots with different order data, policies, and handoff rules. ElevenLabs says a team can define the agent once and then tune behavior for each channel. Its announcement also highlights channel-scoped simulations, which let teams test the tuned behavior against evaluation criteria before updates go live.
### One agent does not mean one behavior
The same knowledge and policy can serve every channel, but the presentation rules should change. A voice agent needs short turns, low latency, interruption handling, pronunciation control, and a clear verbal disclosure. SMS needs compact text, safe links, identity limits, and a plan for messages that arrive hours later. A ticketing agent can provide a longer structured answer, but it must preserve the thread, quote the right customer, avoid duplicating actions, and write into a durable business record.
ElevenLabs' channel behavior tools address part of that difference by letting teams tune response style per surface. The important work remains with the operator. Each channel needs its own maximum response length, supported formats, authentication method, timeout rule, escalation path, and confirmation language. A refund procedure that is acceptable on a verified call may be unsafe in an unauthenticated Telegram conversation. A voice apology that sounds empathetic may look overly verbose in a support ticket.
The safest design is a shared policy core with channel-specific wrappers. The core should contain the approved knowledge, procedure steps, tool permissions, and prohibited decisions. The wrapper should control what the agent can reveal, how it verifies the user, which tools it can call, how it formats the answer, and when it transfers to a person. Teams should test the same scenario on every live channel and compare not only the answer, but the action and audit trail.
### Alpha channels need production discipline
ElevenLabs labeled Telegram, Intercom, and Freshdesk integrations as Alpha in the July announcement. Alpha is not a cosmetic label. It should change the deployment plan. A team should expect missing controls, changing interfaces, edge cases, limited support, and behavior that differs from a mature channel. High-volume or regulated workflows should not move into an Alpha integration because a demo succeeded.
A controlled pilot should start with low-risk queries and a narrow group of users. The agent can answer order-status questions from a test environment before it receives refund authority. Ticket drafts can remain in review before the system posts automatically. The team should record duplicate actions, lost context, incorrect customer matching, formatting failures, latency, channel outages, and transfers that never reach a human. A rollback should be possible without rewriting the entire agent.
Versioning matters as well. When a shared agent changes, a phone workflow and an Alpha ticket integration may not fail in the same way. Channel-scoped simulations should become a release gate. A team can maintain a small set of tests for identity, disclosure, unsupported promises, refunds, payment information, prompt injection, abusive language, silence, interruption, attachment handling, tool failure, and human escalation. A release should fail if any channel crosses its risk threshold.
### Procedures and tools create the action risk
ElevenAgents Procedures let teams define how an agent should complete repeatable tasks. Workflows can branch according to the conversation, and tools can retrieve or update business data. That structure is more controllable than a vague system prompt, but it also turns a wrong interpretation into a real action. An agent might disclose an order to the wrong person, submit a duplicate refund, change a reservation, or write an inaccurate summary into a customer record.
The control point is authorization, not eloquence. Before every consequential tool call, the system should know which user has been verified, which account and record are in scope, what exact change will occur, whether the action is reversible, and whether policy requires a human. A user asking a question should not automatically authorize the agent to act. The agent should preview the target and effect, then obtain confirmation in the same channel or transfer to an approved process.
Multimodality increases both usefulness and exposure. ElevenLabs says agents can work with images, files, and audio, while post-call webhooks can deliver transcripts, analysis, metadata, or full audio to another endpoint. Operators should decide which file types are accepted, scan untrusted inputs, limit prompt-injection paths, redact unnecessary sensitive data, secure webhook destinations, and avoid sending full audio when a smaller event payload is enough.
### Disclosure is a product requirement
ElevenLabs' documentation requires notice immediately before an ElevenAgents interaction. The notice must explain that the user is interacting with AI rather than a person and that the conversation is being recorded and may be shared with ElevenLabs and third-party language-model providers. For a voice call, that can be a verbal message. For web and messaging, it can be a visible screen, banner, or other notice that appears before use.
This should not be reduced to a rushed sentence that the user cannot understand. The disclosure should match the actual workflow, identify the business responsible for the agent, link to the relevant privacy information, and explain any recording or data sharing that matters. It should also work on every channel. A disclosure shown on a website does not necessarily cover a later SMS conversation or an inbound phone call. Local consent, recording, marketing, accessibility, and sector rules can add obligations beyond the platform terms.
Voice identity needs an equally firm boundary. Instant Voice Cloning requires permission from the speaker. Professional Voice Cloning is more restrictive: ElevenLabs says users can create a Professional clone only of their own verified voice, even if another person has consented. Teams that need a celebrity, employee, performer, or character voice should use the correct licensing and product route rather than trying to bypass identity verification. Voice Design and licensed Voice Library voices can provide alternatives without copying an identifiable person.
### Provenance is improving, but it is not proof
ElevenLabs began adding Google's SynthID watermark to text-to-speech generations by Free users in June and said it would expand coverage across its audio generations. It also launched a free Audio Detector. The watermark is designed to remain detectable after common transformations such as compression, clipping, speed changes, or metadata removal. This is meaningful infrastructure because platform metadata is easy to strip when a file is reposted.
A detected watermark can help attribute audio to ElevenLabs. It cannot tell a listener whether the voice owner consented, whether the statement is true, whether a customer was deceived, whether the file was edited, or whether the use complies with law and platform rules. A negative result also does not prove that audio is human. Watermarking should sit beside visible disclosure, account-level traceability, content credentials, permission records, moderation, and a public reporting process.
The distinction is especially important for agents. A customer may hear a realistic voice during a live interaction rather than a published file. The organization needs to disclose the AI at the start, log the selected voice and agent version, preserve the relevant tool actions, and make it possible to investigate a disputed call or message. Provenance should follow the interaction, not only the exported audio.
### Pricing must be modeled by surface
The current self-serve subscription ladder is Free, Starter at $6 monthly, Creator at $22, Pro at $99, Scale at $299, and Business at $990, with custom Enterprise pricing. Creative and API products draw from shared monthly credits. Paid unused credits can roll over for up to two months, capped at two times the plan's monthly quota, while the subscription remains active and is not downgraded.
ElevenAgents uses a different unit. It is billed by call minutes, with 15 included minutes on Free, 75 on Starter, 275 on Creator, 1,238 on Pro, 3,738 on Scale, and 12,375 on Business. Current additional call minutes are listed at $0.08, burst minutes beyond concurrency at $0.16, and text messages at $0.003. Language-model and telephony providers are charged separately at cost. This means a team cannot estimate agent spend from the creative credit balance alone.
A useful forecast separates voice generation, transcription, dubbing, music, Studio or Flows generation, agent hosting, text messages, telephony, external language models, concurrency bursts, and human review. The cost of a resolved customer issue is more informative than the advertised number of included minutes. A cheap call that creates a duplicate refund is not cheap. A more expensive call that verifies identity, resolves the issue, and preserves a clean audit trail may be the better system.
### What teams should test before rollout
Start with a narrow task such as order status or appointment scheduling. Write the procedure, list the approved tools, and define prohibited actions. Add the required AI and recording notice before interaction. Create one evaluation set that covers correct requests, ambiguous identity, policy exceptions, prompt injection, emotional callers, silence, noise, duplicate messages, tool outages, and handoff failure. Then run it across each intended channel.
Review the agent's answer, tool calls, data exposure, timing, formatting, and escalation. Test whether SMS messages remain understandable without voice context, whether a ticket preserves the right thread, and whether a phone transfer carries the necessary summary without oversharing. Keep Alpha channels in a monitored pilot until the failure rate and controls meet the organization's standard.
ElevenLabs' multichannel expansion is valuable because customers should not have to restart every conversation when they switch surfaces. The platform's advantage will not be proven by placing one agent in the largest number of channels. It will be proven when the agent adapts to each channel, acts only with appropriate authority, discloses what it is, and leaves a record that a human can understand and correct.
ElevenLabs now spans much more than text to speech. ElevenCreative combines voices, transcription, dubbing, music, sound effects, Studio, Flows, image, and video tools. ElevenAgents deploys governed voice and text agents across customer channels. ElevenAPI exposes the underlying capabilities to developers.
Pricing, voice tools, agents, rights, and limits
The current self-serve ladder runs from Free to Starter at $6 monthly, Creator at $22, Pro at $99, Scale at $299, and Business at $990. Creative and API tools share credits, while ElevenAgents uses separate call-minute pricing with concurrency, text-message, external language-model, and telephony costs.
How to use ElevenLabs safely and cost-effectively
ElevenLabs is a strong fit when voice quality, multilingual production, dubbing, creative audio, and real-time agents need to live in one ecosystem. Buyers should test the exact model and channel, document voice and content rights, review commercial terms, disclose AI and recording, limit agent permissions, and measure approved output after corrections rather than headline minutes.
About ElevenLabs
ElevenLabs is an AI audio and conversational-agent platform organized around ElevenCreative, ElevenAgents, and ElevenAPI. Creators can generate expressive speech, design or clone approved voices, transcribe recordings, dub performances, create music and sound effects, and assemble audio or video projects in Studio and Flows. Businesses can deploy agents across phone, web, WhatsApp, Slack, Zendesk, SMS, Telegram, Intercom, and Freshdesk, with knowledge bases, tools, procedures, workflows, testing, analytics, and channel-specific behavior. Developers can embed speech, transcription, dubbing, music, sound, and agent capabilities through APIs and SDKs. The platform is powerful but requires explicit voice rights, AI and recording disclosure, human review, cost controls, safe tool permissions, and careful handling of customer data.
Use Cases
Key Features
- ✓ ElevenCreative browser workspace for creating, editing, and localizing audio and video
- ✓ ElevenAgents platform for real-time voice and text agents
- ✓ ElevenAPI for speech, transcription, dubbing, music, sound, and agent integrations
- ✓ Eleven v3 expressive text to speech with dialogue and inline audio tags
- ✓ Text to speech across more than 70 languages with model-dependent coverage
- ✓ Flash and Turbo speech models for low-latency applications
- ✓ Streaming speech generation for real-time playback
- ✓ Instant Voice Cloning from a short approved sample
- ✓ Professional Voice Cloning from extended high-quality recordings
- ✓ Voice Design for creating synthetic voices from descriptions
- ✓ Voice Remixing for changing delivery, cadence, tone, and accent
- ✓ Searchable Voice Library with default and community voices
- ✓ Voice Actor Payouts for eligible shared Professional Voice Clones
- ✓ Voice Changer for transferring delivery into another approved voice
- ✓ Voice Isolator for removing noise, reverb, and background sound
- ✓ Scribe speech to text for recorded audio and video
- ✓ Scribe Realtime for low-latency live transcription
- ✓ Speaker diarization, timestamps, audio-event labels, and keyterm prompting
- ✓ Dubbing v2 for more than 90 languages while preserving speaker delivery
- ✓ Dubbing Studio for transcript, translation, speaker, and timing review
- ✓ Music v2 for songs, instrumentals, vocals, arrangement, and editing
- ✓ Music References for style guidance from approved uploaded tracks
- ✓ Music Finetunes and consistent vocals for eligible workflows
- ✓ AI Sound Effects and the ElevenMusic Sounds library
- ✓ Studio for audiobooks, podcasts, voiceovers, captions, and long-form projects
- ✓ Character Casting that detects manuscript characters and proposes voices
- ✓ Pronunciation dictionaries and project-wide delivery controls
- ✓ Flows visual canvas for chaining audio, image, video, lip-sync, and editing models
- ✓ Flows Agent that can build and run multimodal creative pipelines from conversation
- ✓ Assist mode for approval before expensive Flows generations
- ✓ Ads Engine for generating and localizing advertising creative
- ✓ Agent knowledge bases, retrieval, tools, procedures, workflows, and guardrails
- ✓ Agent deployment over phone, web, WhatsApp, Slack, Zendesk, SMS, Telegram, Intercom, and Freshdesk
- ✓ Channel-specific agent behavior and simulation testing
- ✓ Agent transcripts, analytics, Spotlight insights, sentiment, and topic discovery
- ✓ Real-time monitoring, OpenTelemetry traces, and post-call webhooks
- ✓ Web, mobile, and server SDKs plus REST and WebSocket APIs
- ✓ SynthID watermarking rollout and a free ElevenLabs Audio Detector
- ✓ Startup Grants for qualifying agent builders
- ✓ Enterprise security, privacy, residency, and deployment options where contracted
Pricing
Free
$0
- • 10,000 creative and API credits per month
- • Approximately 10 minutes of standard text to speech
- • 15 ElevenAgents call minutes
- • Four concurrent agent calls
- • Text to speech, speech to text, music, sound, and dubbing trials
- • Voice Design and limited Voice Library access
- • API access to eligible endpoints
- • No commercial license
- • Unused Free credits do not roll over
Starter
$6/month
- • 30,000 creative and API credits per month
- • Approximately 30 minutes of standard text to speech
- • 75 ElevenAgents call minutes
- • Six concurrent agent calls
- • Commercial license for eligible non-beta services
- • Instant Voice Cloning
- • Dubbing Studio and expanded creative tools
- • Text-message support for agents
- • Annual equivalent of $5 per month
Creator
$22/month
- • 121,000 creative and API credits per month
- • Approximately 121 minutes of standard text to speech
- • 275 ElevenAgents call minutes
- • Ten concurrent agent calls
- • Professional Voice Cloning
- • Additional creative credits and agent minutes available
- • First-month promotion may reduce the price to $11
- • Annual equivalent of about $18.33 per month
Pro
$99/month
- • 600,000 creative and API credits per month
- • Approximately 600 minutes of standard text to speech
- • 1,238 ElevenAgents call minutes
- • Twenty concurrent agent calls
- • 44.1 kHz PCM output through the API
- • 192 kbps audio through Studio and API workflows
- • Additional-credit and pay-as-you-go options
- • Annual equivalent of $82.50 per month
Scale
$299/month
- • 1.8 million creative and API credits per month
- • Approximately 1,800 minutes of standard text to speech
- • 3,738 ElevenAgents call minutes
- • Thirty concurrent agent calls
- • Three workspace seats
- • Team collaboration
- • Three Professional Voice Clones
- • Annual equivalent of about $249.17 per month
Business
$990/month
- • 6 million creative and API credits per month
- • Approximately 6,000 minutes of standard text to speech
- • 12,375 ElevenAgents call minutes
- • Forty concurrent agent calls
- • Ten workspace seats
- • Ten Professional Voice Clones
- • Low-latency text to speech from about $0.05 per minute
- • Annual equivalent of $825 per month
Enterprise
Custom
- • Custom credits, call volume, concurrency, seats, and voices
- • Custom DPA and service-level terms
- • BAAs for qualifying HIPAA customers
- • Custom SSO and administration
- • Elevated concurrency limits
- • Regional, local, and zero-retention options where contracted
- • Managed dubbing through Productions
- • Priority support and volume discounts
Pricing varies by plan and region — see current pricing.
Plan features change — last updated: 2026-08-16.
Details
Tags
ElevenLabs Community Discussions
Explore community discussions. Ask and answer questions on ElevenLabs to grow and learn together.
ElevenLabs Showcase
ElevenLabs — Frequently Asked Questions
What is ElevenLabs?
ElevenLabs is an AI audio and conversational-agent platform. ElevenCreative covers speech, voices, transcription, dubbing, music, sound, Studio, and Flows. ElevenAgents covers customer-facing voice and text agents. ElevenAPI lets developers embed those capabilities in software.
Does ElevenLabs have a free plan?
Yes. The current Free plan includes 10,000 creative and API credits, about 10 minutes of standard text to speech, and 15 ElevenAgents call minutes. It does not include a commercial license, and unused Free credits do not roll over.
How much does ElevenLabs cost?
As verified on August 16, 2026, Starter is $6 monthly, Creator $22, Pro $99, Scale $299, and Business $990. Enterprise is custom. Annual plans charge the equivalent of ten monthly payments, and taxes, extra usage, telephony, and external language models can add cost.
How do ElevenLabs credits work?
Creative and API products share a monthly credit pool, but each product and model consumes credits at a different rate. Paid unused credits can roll over for up to two months, capped at two times the monthly quota, while the subscription remains active and is not downgraded.
How is ElevenAgents billed?
ElevenAgents is billed by call minutes rather than the shared creative credit pool. Plans include 15 to 12,375 call minutes and 4 to 40 concurrent calls. Additional minutes currently cost $0.08, burst minutes cost $0.16, text messages cost $0.003, and external providers are extra.
Can I use ElevenLabs output commercially?
The Free plan has no commercial license. Eligible paid-plan output can be used commercially under the applicable service terms if the user owns the necessary rights. Beta Services cannot be used commercially or in production under the current Beta Services Addendum.
What is the difference between Instant and Professional Voice Cloning?
Instant Voice Cloning creates a fast clone from a short approved sample. Professional Voice Cloning uses 30 to 180 minutes of clean audio for stronger consistency and verification, but ElevenLabs permits a user to create a Professional clone only of their own voice.
Can I clone someone else's voice with permission?
Instant Voice Cloning requires permission from the voice owner. Professional Voice Cloning is more restrictive: ElevenLabs says users may clone only their own verified voice, even when another person has consented. Use a licensed Voice Library voice or Voice Design instead.
What is Eleven v3?
Eleven v3 is ElevenLabs' generally available expressive text-to-speech model. It supports more than 70 languages, multi-speaker dialogue, and audio tags for performance cues such as whispering, laughter, or emotion. Real-time use may fit Flash or Turbo models better.
What is Dubbing v2?
Dubbing v2 translates and recreates speech across more than 90 languages and accents while preserving more of the source speaker's identity, tone, emotion, and delivery. Native-language review remains necessary before publication.
What can ElevenLabs Studio and Flows do?
Studio supports long-form voice, audiobooks, podcasts, captions, corrections, and project assembly. Flows is a node-based canvas that chains ElevenLabs audio with image, video, lip-sync, and other models. Flows Agent can build and run a pipeline from natural-language instructions.
Which channels does ElevenAgents support?
ElevenAgents supports channels including phone, web, WhatsApp, Slack, Zendesk, SMS, Telegram, Intercom, and Freshdesk. Availability varies, and Telegram, Intercom, and Freshdesk were still labeled Alpha in the July 2026 product announcement.
Do ElevenAgents need to identify themselves as AI?
Yes. ElevenLabs' disclosure documentation requires notice immediately before interaction that the user is dealing with AI and that the conversation is being recorded and may be shared with ElevenLabs and third-party language-model providers. Organizations remain responsible for legal compliance.
Does ElevenLabs use customer data to improve models?
ElevenLabs says it uses certain submitted data to improve audio models, but users can disable the Improve the models for everyone setting for new data. Enterprise customer data is not used for training by default, subject to service delivery and contract terms.
What is the ElevenLabs Audio Detector?
It is a free tool designed to detect SynthID watermarks in audio generated by ElevenLabs. The watermark can survive common transformations such as compression or trimming, but a positive result does not verify the truth, consent, safety, or legality of the audio.
Sources & References
- Official ElevenLabs platform ↗
- Official ElevenCreative and API pricing ↗
- Official ElevenAgents pricing ↗
- Official ElevenLabs documentation overview ↗
- Official ElevenCreative overview ↗
- Official text-to-speech product guide ↗
- Official Instant Voice Cloning guide ↗
- Official Professional Voice Cloning guide ↗
- Professional Voice Clone identity restriction ↗
- Commercial use and publication rules ↗
- ElevenLabs model-improvement data controls ↗
- Eleven v3 product and model overview ↗
- Official ElevenLabs documentation changelog ↗
- Character Casting in Audiobooks announcement ↗
- Flows Agent announcement and assist mode ↗
- New channels in ElevenAgents announcement ↗
- ElevenAgents disclosure requirements ↗
- ElevenAgents workflow documentation ↗
- SynthID watermarking and Audio Detector announcement ↗
- ElevenLabs safety and reporting resources ↗
- Official ElevenLabs affiliate program ↗
- ElevenLabs Creator Affiliate Program terms ↗
Try ElevenLabs
Visit the official website to get started with ElevenLabs today.
Visit ElevenLabs →