All Services

AI / ML

Amazon Polly

Amazon Polly converts text and supported SSML into synthesized speech using a catalog of languages, voices, and model engines, with real-time audio and asynchronous speech-task APIs.

Explore pricing models, common use cases, infrastructure support, and the AWS services that commonly work with Amazon Polly.

Amazon Polly pricing and cost programs

Pricing model: Text-to-speech usage

On-Demand
Available
Reserved Instances or reserved capacity
Not applicable
Savings Plans
Not applicable
Spot
Not applicable

Billing dimensions: Characters synthesized · Voice engine · Speech marks

Programs and modes: Standard · Neural · Long-form · Generative voices

Voice engine and character count determine charges.

Free Tier: Available — verify current offers

Pricing reviewed 2026-07-25. Reviewed against the linked official AWS pricing page. Recheck regional rates and program terms before purchase.

Official AWS pricing

Official AWS sources reviewed 2026-07-21.

Why implement Amazon Polly?

  • Generates speech without recording studios or customer-managed speech-synthesis infrastructure.
  • Offers multiple voices, languages, audio formats, speech marks, and model-dependent engines for interactive and batch experiences.
  • Integrates with applications, contact centers, conversational bots, S3, and content-delivery workflows.

How to implement Amazon Polly

  1. Select a supported language, voice, engine, audio format, sample rate, pronunciation and accessibility requirements using representative listening tests.
  2. Validate and escape text or SSML, split requests within current quotas, use a least-privilege role, and choose direct synthesis or an asynchronous S3 task for long-form content.
  3. Cache immutable output when licensing and freshness permit, version source text and voice settings, and monitor failures, latency, throttling, cost, and user accessibility feedback.

Amazon Polly best practices

  • Use SSML deliberately and validate it before production; test names, abbreviations, dates, numbers, pronunciation, pacing, and language-specific behavior with native listeners.
  • Do not present synthetic speech as a real person's recording or use it deceptively; disclose automation where context calls for it and provide accessible alternatives.
  • Protect sensitive source text and generated audio, grant narrow S3 access, reuse cached output to reduce latency and cost, and design retries around per-operation quotas.

Amazon Polly use cases and server impact

  • Accessible narration and screen content
  • Voice responses for bots and contact centers
  • Automated audio articles and announcements

Replaces speech-synthesis servers and many recording workflows; it does not translate text, and content accuracy, disclosure, pronunciation review, storage, and delivery remain yours.

Official implementation resources

Commonly paired AWS services