AI / ML
Amazon Polly
Amazon Polly converts text and supported SSML into synthesized speech using a catalog of languages, voices, and model engines, with real-time audio and asynchronous speech-task APIs.
Explore pricing models, common use cases, infrastructure support, and the AWS services that commonly work with Amazon Polly.
Amazon Polly pricing and cost programs
Pricing model: Text-to-speech usage
- On-Demand
- Available
- Reserved Instances or reserved capacity
- Not applicable
- Savings Plans
- Not applicable
- Spot
- Not applicable
Billing dimensions: Characters synthesized · Voice engine · Speech marks
Programs and modes: Standard · Neural · Long-form · Generative voices
Voice engine and character count determine charges.
Free Tier: Available — verify current offers
Pricing reviewed 2026-07-25. Reviewed against the linked official AWS pricing page. Recheck regional rates and program terms before purchase.
Official AWS sources reviewed 2026-07-21.
Why implement Amazon Polly?
- Generates speech without recording studios or customer-managed speech-synthesis infrastructure.
- Offers multiple voices, languages, audio formats, speech marks, and model-dependent engines for interactive and batch experiences.
- Integrates with applications, contact centers, conversational bots, S3, and content-delivery workflows.
How to implement Amazon Polly
- Select a supported language, voice, engine, audio format, sample rate, pronunciation and accessibility requirements using representative listening tests.
- Validate and escape text or SSML, split requests within current quotas, use a least-privilege role, and choose direct synthesis or an asynchronous S3 task for long-form content.
- Cache immutable output when licensing and freshness permit, version source text and voice settings, and monitor failures, latency, throttling, cost, and user accessibility feedback.
Amazon Polly best practices
- Use SSML deliberately and validate it before production; test names, abbreviations, dates, numbers, pronunciation, pacing, and language-specific behavior with native listeners.
- Do not present synthetic speech as a real person's recording or use it deceptively; disclose automation where context calls for it and provide accessible alternatives.
- Protect sensitive source text and generated audio, grant narrow S3 access, reuse cached output to reduce latency and cost, and design retries around per-operation quotas.
Amazon Polly use cases and server impact
- Accessible narration and screen content
- Voice responses for bots and contact centers
- Automated audio articles and announcements
Replaces speech-synthesis servers and many recording workflows; it does not translate text, and content accuracy, disclosure, pronunciation review, storage, and delivery remain yours.
Official implementation resources
Commonly paired AWS services
- Amazon Simple Storage Service — Object storage
- AWS Lambda — Run code without servers
- Amazon API Gateway — Managed APIs
- Amazon Lex — Conversational AI bots
- Amazon Connect — Cloud contact center
- Amazon CloudWatch — Metrics & logs