Skip to content

Infrastructure & pricing, read from an operations desk

Everything we have published under Infrastructure & pricing, read from an operations desk: what it changes for a B2B company of 10-200 people.

  1. Latest

    Salesforce Folds AI Agents Into Standard CRM Pricing Tiers

    Salesforce announced a simplified set of 2026 Editions that bundle its Agentforce AI agents and Data Cloud into standard CRM tiers instead of selling them as separate add-ons. For B2B companies running sales or support on Salesforce, this changes what's included by default in a renewal and what previously cost extra now doesn't — or gets pushed into a higher tier.

    What changes for operatorsIf your 10-200 person company runs Sales Cloud, Service Cloud, or both, this repackaging directly affects your next contract renewal: features you may have been paying for as add-ons (or skipping because of cost) could now be bundled into your existing tier, or your current tier could be discontinued and replaced with a pricier one that includes AI agents you didn't ask for. Before renewing, get your account team to map your current add-on spend against the new Edition structure — the bundling can work in your favor if you were already paying for Agentforce or Data Cloud separately, but it can also force an upgrade if the new baseline tier no longer matches what you're actually using. Either way, this is a billing and packaging event, not evidence that agentic AI is now

  1. AWS Shows How to Cut RAG Token Costs on Bedrock by Trimming Irrelevant Context

    If your support bot, sales assistant, or internal knowledge search runs on a retrieval-augmented pipeline through Bedrock (or a similar architecture), the token bill scales with how much irrelevant context gets stuffed into every prompt — long documents, boilerplate, and near-duplicate passages you retrieve 'just in case.' Query-aware compression addresses that by filtering retrieved chunks against the actual question before they reach the model, which is the same lever that determines whether a 20-person support team's AI assistant costs $200 or $2,000 a month at scale. Teams already running RAG in production should treat this as a concrete cost-reduction checklist item, not a future upgrade — it requires no model swap, only a compression step inserted into the existing retrieval-to-generation pipeline.

  1. Stripe Buys OpenRouter: What It Means for Teams Routing AI Traffic Through It

    If your sales or support automation uses OpenRouter to switch between GPT, Claude, Gemini or open models based on cost or uptime, you now depend on a piece of infrastructure owned by a payments company rather than an independent neutral router — worth checking whether pricing tiers, rate limits or SLA terms shift in the next few quarters, and whether Stripe pushes usage-based billing changes that affect your per-request costs. Teams with a single point of failure on OpenRouter for model orchestration should confirm they can fall back to direct provider APIs if terms change, and treat this as a prompt to audit vendor concentration risk in their AI stack rather than a reason to migrate immediately.

  2. Liquid AI Ships LFM2.5-DSpark, Claims Up to 3.2x Faster Inference

    For a B2B company running an AI chat agent, ticket triage bot, or sales qualification assistant on a small, self-hosted or edge-deployed model, a 3.2x inference speedup translates into lower latency per response and fewer GPU-hours per conversation — meaning either cheaper hosting bills at the same volume, or the ability to run a more capable model at the same cost. Teams currently constrained by response-time SLAs in live chat or voice support, where every second of model latency shows up as customer wait time, get the most immediate benefit; teams using hosted API models from major vendors won't see any change until (or unless) those vendors adopt similar techniques.

  1. AWS Adds Cross-Region Routing for GPT-5.6 on Bedrock

    If your support bot, lead-qualification agent, or ops automation calls GPT-5.6 through Amazon Bedrock, this removes a real operational headache: capacity crunches in a single region that cause dropped or delayed responses during peak hours. Instead of writing and maintaining your own retry-and-failover logic across regions, Bedrock now handles that routing for you, which means fewer 3am pages when a customer-facing AI workflow starts throttling. Teams running lean ops (10-200 people) rarely have spare engineering time to build resilience infrastructure themselves, so this is a case where the cloud provider absorbing that complexity is a direct, if modest, win for uptime of any AI-driven sales or support pipeline built on Bedrock.

  1. Researchers Squeeze 33 Points of GPU Utilization Out of Existing Hardware — By Reordering Jobs

    Most 10-200 person B2B companies don't run their own GPU clusters, so this isn't a direct action item — but it's a useful data point when a vendor tells you that scaling an AI feature requires a costly infrastructure upgrade. If job scheduling alone can unlock 33 points of utilization on the same hardware, ask any provider quoting you for "more compute" whether they've actually optimized what they have first. The angle here is procurement leverage and vendor scrutiny, not internal ops change — most readers won't touch a scheduler themselves, but they will pay for one indirectly through inference or fine-tuning costs.

  2. Solar Eclipse Briefly Dipped Internet Traffic Across Iceland, Spain, and Portugal

    For a 10-200 person B2B company, this event carries no operational implication worth acting on. It is not a security incident, capacity risk, or infrastructure failure — it is a predictable, brief dip in regional consumer browsing behavior tied to a natural phenomenon. Unless your customer base is heavily concentrated in Reykjavik, Madrid, or Lisbon and your business depends on real-time traffic during a two-hour window on eclipse day, there is nothing here to change in your support staffing, uptime monitoring, or automation workflows. It's worth noting mainly as an example of how cleanly network telemetry can capture human behavior at scale — useful context if you ever need to explain an unexplained traffic anomaly to a client.

  3. NVIDIA's Fast, Cheap Nemotron Model Lands on AWS SageMaker

    If your ops or engineering team already runs on AWS, this matters less as a "new AI model" story and more as a procurement and latency story: one-click deployment inside SageMaker JumpStart cuts the integration overhead of adding a fast, lower-cost model to sales chatbots, support triage, or internal workflow automation. For a 10-200 person company, that's the difference between a two-week engineering sprint and an afternoon's work testing whether a lighter model handles ticket routing or lead qualification well enough to replace a pricier one. The catch: this is an AWS-specific convenience, not a universal capability shift — if you're not on AWS, or you don't yet have infrastructure to A/B test model swaps safely, there's nothing to act on here today beyond noting the option exists.

  1. OpenAI Pitches Texas Governor on Data Center Buildout, Citing "Responsible" Infrastructure

    For a 10-200 person B2B company running sales and support automation on top of OpenAI's models, this letter itself changes nothing operationally today — it's a policy communication about physical infrastructure siting in one state, not a product, pricing, or API change. The indirect relevance is worth noting: continued data center buildout in Texas and similar states is part of the capacity expansion that underpins model availability and, eventually, cost trends for API access. Operators should treat this as background context rather than an action item — there's no new tool, quota, or rate limit to react to. The one thing worth watching, unconfirmed for now, is whether state-level infrastructure agreements like this start showing up in vendor communications about regional data residency or latency-optimized endpoints, which would matter more directly for compliance-sensitive support and ops workflows.

Next step

Free AI Diagnostic

Fifteen minutes, no email required. It maps where your work actually goes and ranks what is worth automating first.

Start the free diagnostic

Starts immediately in the browser.

Fee
Free
Length
15 minutes

You keep the ranked list of candidates either way.