Skip to content

Infrastructure & pricing, read from an operations desk

Everything we have published under Infrastructure & pricing, read from an operations desk: what it changes for a B2B company of 10-200 people.

  1. Latest

    Amazon Bedrock Adds Prompt Caching, Cutting AI Support Tool Costs Sharply

    AWS added prompt caching to Amazon Bedrock, letting applications reuse previously processed prompt segments (system instructions, long context, examples) instead of reprocessing them on every call. AWS reports up to 90% lower input token costs and up to 85% lower latency for supported models, directly cutting the operating cost of AI agents and chat tools built on Bedrock.

    What changes for operatorsIf your support chatbot, internal knowledge assistant, or sales copilot runs on Bedrock and sends the same system prompt or product documentation with every request — which most retrieval-augmented tools do — this reduces your per-query cost and speeds up response times without any change to the model itself. Teams running high-volume support automation (think hundreds or thousands of tickets a day) should see the caching applied automatically or configure it explicitly, since AWS notes it works best when a large portion of the prompt — like a knowledge base excerpt or tool definitions — stays identical across calls. For a 10-200 person company already paying per-token for AI support or sales workflows on Bedrock, this is a direct cost lever worth checking this quarter, not a future consideration.

  1. AWS SageMaker Adds Smarter Routing to Cut Self-Hosted LLM Latency

    Most 10-200 person B2B companies call a hosted API like OpenAI or Anthropic and this change does not touch them directly. But if your ops or support automation runs a self-hosted or fine-tuned model behind SageMaker — common when handling sensitive customer data, ticket histories, or proprietary sales scripts that need to stay in your own VPC — this routing update is a free latency and cost reduction. Support bots and agent-assist tools that reuse the same system prompt across thousands of tickets per day will see faster first-token response and lower GPU spend simply by upgrading to the new routing strategy, with no changes to the prompts or application logic themselves.

  1. AWS Shows How to Cut RAG Token Costs on Bedrock by Trimming Irrelevant Context

    If your support bot, sales assistant, or internal knowledge search runs on a retrieval-augmented pipeline through Bedrock (or a similar architecture), the token bill scales with how much irrelevant context gets stuffed into every prompt — long documents, boilerplate, and near-duplicate passages you retrieve 'just in case.' Query-aware compression addresses that by filtering retrieved chunks against the actual question before they reach the model, which is the same lever that determines whether a 20-person support team's AI assistant costs $200 or $2,000 a month at scale. Teams already running RAG in production should treat this as a concrete cost-reduction checklist item, not a future upgrade — it requires no model swap, only a compression step inserted into the existing retrieval-to-generation pipeline.

  1. AWS Adds Cross-Region Routing for GPT-5.6 on Bedrock

    If your support bot, lead-qualification agent, or ops automation calls GPT-5.6 through Amazon Bedrock, this removes a real operational headache: capacity crunches in a single region that cause dropped or delayed responses during peak hours. Instead of writing and maintaining your own retry-and-failover logic across regions, Bedrock now handles that routing for you, which means fewer 3am pages when a customer-facing AI workflow starts throttling. Teams running lean ops (10-200 people) rarely have spare engineering time to build resilience infrastructure themselves, so this is a case where the cloud provider absorbing that complexity is a direct, if modest, win for uptime of any AI-driven sales or support pipeline built on Bedrock.

  1. NVIDIA's Fast, Cheap Nemotron Model Lands on AWS SageMaker

    If your ops or engineering team already runs on AWS, this matters less as a "new AI model" story and more as a procurement and latency story: one-click deployment inside SageMaker JumpStart cuts the integration overhead of adding a fast, lower-cost model to sales chatbots, support triage, or internal workflow automation. For a 10-200 person company, that's the difference between a two-week engineering sprint and an afternoon's work testing whether a lighter model handles ticket routing or lead qualification well enough to replace a pricier one. The catch: this is an AWS-specific convenience, not a universal capability shift — if you're not on AWS, or you don't yet have infrastructure to A/B test model swaps safely, there's nothing to act on here today beyond noting the option exists.

Next step

Free AI Diagnostic

Fifteen minutes, no email required. It maps where your work actually goes and ranks what is worth automating first.

Start the free diagnostic

Starts immediately in the browser.

Fee
Free
Length
15 minutes

You keep the ranked list of candidates either way.