Skip to content

Grok 4.6 Lands on Amazon Bedrock With Guardrails and Cross-Region Pricing Tradeoffs

Short answer

xAI's Grok 4.6 launched on Amazon Bedrock on August 18, 2026, xAI's second model on the platform. It adds a 500K token context window, four reasoning effort levels, Bedrock Guardrails, invocation logging and cross-region inference — features that shape how B2B teams build and price agentic workflows on AWS.

What this means for operators

For a 10-200 person company running sales or support agents on AWS Bedrock, this is a real build decision, not just a new model to try. Guardrails, invocation logging and prompt caching now work on the bedrock-runtime endpoint, which is the one to use if you need a policy boundary around an unattended agent and an audit trail of what it actually did. But structured JSON output and server-side tool use stay on bedrock-mantle only, so a workflow depending on strict schema output has to pick that endpoint instead. Pricing adds another lever: routing through the cheaper Global cross-region profile versus the US-only Geo profile, or dropping to the Flex service tier at half the standard rate for non-urgent batch work, changes the unit economics of an agent more than the model choice itself.

xAI's Grok 4.6 is now available in Amazon Bedrock, the AWS Machine Learning Blog reports, launching on the platform on August 18, 2026 as xAI's second model there after Grok 4.3. The model offers a 500K token context window and four configurable reasoning effort levels — low, medium, high, and a new xhigh — set through Bedrock's API parameters rather than a dedicated reasoning field.

Grok 4.6 widens its footprint compared with the earlier Grok launch: it runs on both the bedrock-mantle (OpenAI-compatible) and bedrock-runtime endpoints, and it supports the Converse API alongside Chat Completions and Responses, including streaming. Feature support differs by endpoint. bedrock-mantle covers client-side tool calling, structured outputs, prompt caching, response streaming and abuse detection. bedrock-runtime covers reasoning, prompt caching, response streaming, invocation logs and now, new for this launch, Amazon Bedrock Guardrails — content filters, denied topics, PII redaction and word policies evaluated against both prompt and response. Structured outputs and server-side tool use are not available on bedrock-runtime.

Grok 4.6 is not offered for in-Region inference on bedrock-runtime; requests instead route through cross-Region inference profiles. A Geo profile keeps traffic within US Regions for data residency, while a Global profile routes worldwide across a longer list spanning the US, Canada, Europe, Asia Pacific, the Middle East, Africa and South America. Global is also cheaper: $2.00 per million input tokens against $2.20 for Geo. On bedrock-mantle, in-Region inference is available only in US West (Oregon).

Pricing includes three service tiers built as multipliers on standard rates: Priority at 1.75x for faster processing, and Flex at 0.5x for work that isn't time-sensitive. xAI lists Grok 4.6 pricing starting at $2 per million input tokens and $6 per million output tokens, with a faster variant at double that. Cached input is billed at roughly a quarter of the standard input rate, which the post notes matters for agents that resend a large system prompt or document on every turn.

Reported benchmark figures for Grok 4.6 High, published by xAI on August 12, 2026, include an AA Intelligence Index of 61, a GDPVal-AA v2 score of 1753, and 69.9% on CursorBench v3.2, among others cited in the post.

Source: AWS Machine Learning Blog

Next step

Visibility Analyzer

This is what the Visibility Analyzer measures on a real site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.