Claude 3.5 Sonnet costs $3 per million input tokens and $15 per million output tokens. This model provides a 200K token context window. Anthropic retired the two previous Claude 3.5 Sonnet snapshots, claude-3-5-sonnet-20240620 and claude-3-5-sonnet-20241022, from its first-party API on October 28, 2025. Organizations using Amazon Bedrock to access Claude 3.5 Sonnet operate under specific constraints. AWS places these organizations in the Start tier. These users do not move between usage tiers automatically. To increase limits, users must contact an Anthropic account representative or Anthropic support.
| Feature | Specification |
|---|---|
| Input Token Price | $3 per million |
| Output Token Price | $15 per million |
| Context Window | 200K tokens |
| Prompt Caching | Supported on Amazon Bedrock |
| Max Output (via Batch API) | 300K tokens |
Automating development with Claude Code
Claude Code operates directly in a terminal or in IDEs like VS Code and Jetbrains. It uses Claude Sonnet 4 to write code and fix bugs across multiple files. The tool can search through git history, resolve merge conflicts, and create commits or PRs. It also works with AWS CLI, Terraform, and k8s. Developers can connect Claude Code to external tools and data sources through the Model Context Protocol.
Amazon Bedrock prompt caching provides performance benefits for these agentic applications. When a user enables prompt caching, the application inserts cache checkpoint markers at specific points in prompts. Amazon Bedrock creates cache checkpoints that save the model state after processing the preceding text. Subsequent requests that reuse the same prefix load the cached state instead of recomputing. This process reduces response times and lowers input token costs.
For complex codebases, prompt caching reduces token costs for repeated interactions with the same files. In one test, using prompt caching for a simple task resulted in lower resource consumption than running the same task without it. To test this, developers set the environment variable DISABLE_PROMPT_CACHING to enable the feature. To disable it, developers set the variable to true. Does the reduction in latency justify the management of cache checkpoint markers?
Managing API limits and costs
The API enforces service-configured limits at the organization level. Users can also set user-configurable limits for organization workspaces. Two types of limits exist: spend limits and rate limits. Spend limits set a maximum monthly cost an organization incurs for API usage. Rate limits set the maximum number of API requests an organization makes over a defined period.
The API uses the token bucket algorithm for rate limiting. This means capacity replenishes up to a maximum limit rather than resetting at fixed intervals. Rate limits for the Messages API measure requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). For most Claude models, only input tokens after the last cache breakpoint count toward ITPM. Cached input tokens do not count toward ITPM for most models. Instead, they bill at the cache read rate, which is a fraction of the base input price.
If an organization reaches its spend cap, API usage pauses until 00:00 UTC on the first day of the next month. Using the API during this pause returns an HTTP 429 error with the code enforced_spend_limit_reached. You can view your current limits in the Claude Console.
Implementation for enterprise workflows
Claude 3.5 Sonnet performs tasks like code translation and advanced image analysis. It converts code between languages, such as translating Python to Java, while preserving original logic. The model also transcribes text from imperfect images. It extracts information from documents containing printed text, handwritten notes, and custom logos.
You can use Claude 3.5 Sonnet for context-sensitive customer support and multi-step workflow orchestration. For developers, the model solves coding problems with sophisticated reasoning. In an internal agentic coding evaluation, Claude 3.5 Sonnet solved 64% of problems. This outperformed Claude 3 Opus, which solved 38% in the same evaluation.
Deployment on AWS requires planning for security and governance. Organizations should use AWS IAM Identity Center to govern identity and access. This verifies that only authorized developers access Claude Code. Developers can use temporary, role-based credentials through this method.
The model’s audio capabilities are non-existent. Anthropic provides no native audio input, speech output, or real-time voice in its API. Enterprise voice stacks using Claude require third-party speech-to-text and speech-to-text components.
The following table compares the current Claude model lineup.
| Model | Latency | Pricing (Input/Output per MTok) | Context Window |
|---|---|---|---|
| Claude Fable 5.1 | Slower | $10 / $50 | 1M tokens |
| Claude Opus 5 | Moderate | $5 / $25 | 1M tokens |
| Claude Sonnet 5 | Fast | $2 / $10 | 1M tokens |
| Claude Haiku 4.5 | Fastest | $1 / $5 | 200K tokens |
Organizations must review service quotas and set appropriate Token Per Minute (TPM) and Request Per Minute (RPM) values based on active developer counts. For 200 developers, a request for 20,000 TPM per developer would total 4 million total TPM. Use the /cost command in Claude Code to monitor resource consumption and API processing time.

Leave a Reply