Skip to content

Models & capabilities, read from an operations desk

Everything we have published under Models & capabilities, read from an operations desk: what it changes for a B2B company of 10-200 people.

  1. Latest

    Claude Sonnet 5.5 lands on AWS Bedrock as the cheaper execution model to pair with Opus 5.5

    AWS made Claude Sonnet 5.5 available on Amazon Bedrock and Claude Platform on AWS. It costs less per task and runs faster than Sonnet 5 on well-scoped work like coding tasks, SQL generation, and document drafting, while staying inside AWS's IAM, CloudTrail, and Guardrails controls. It's designed to pair with Opus 5.5, which handles judgment-heavy work.

    What changes for operators — For a company already running Bedrock-based coding assistants, alert triage, or document drafting, this is a straight cost-and-latency lever, not a new capability to evaluate from scratch: routing well-defined, high-volume tasks (SQL generation, first response to alerts, spreadsheet edits, routine document work) to Sonnet 5.5 while reserving Opus 5.5 for release debugging, security review, or contract redlining should lower per-task spend without touching the existing IAM, CloudTrail, or Guardrails setup. Teams with fixed monthly AI budgets for IDE coding agents or support-ticket triage get the clearest immediate win, since Sonnet 5.5 is pitched specifically at continuous or at-scale workloads with a fixed spend cap.

  1. H Company Ships Holo4, Open-Weight Agents That Work Any Software Interface

    For a 10-200 person B2B company, the practical draw is not the benchmark score but the interface flexibility: most internal tools, legacy CRMs, and vendor portals have no usable API, forcing manual clicking today. A model that can fall back to GUI control when an API is missing, then switch to API calls where one exists, is directly applicable to automating order entry, ticket triage, or data reconciliation across tools that were never built to be automated. The open weights (BF16, FP8, NVFP4, GGUF) mean a technical team could self-host rather than pay per-call frontier pricing, though the reported OSWorld 2.0 gap versus Opus 5.5 (61.7% vs 81.8%) signals this is a cost-performance tradeoff, not a like-for-like replacement, and should be piloted on a specific workflow before being trusted with production tasks.

  2. Microsoft Foundry Adds Two Cheaper GPT-6 Tiers for Production Agents

    For a company running support or ops agents on Azure, this is a direct cost lever: instead of sending every ticket-classification or data-extraction step through an expensive reasoning model, teams can now put GPT-6 Luna behind routine, high-volume steps and reserve Sol or Astra for the parts of a workflow that actually need judgment. At Global Standard short-context pricing, Luna runs $0.10 per million input tokens and $0.50 output versus Sol's $2/$10 and Astra's $10/$50, which changes the math on cost-per-task for anyone running thousands of routine agent calls a day. The practical move is to map an existing agent workflow, identify which steps are classification/routing/summarization versus multi-step reasoning, and split the model assignment accordingly rather than defaulting to one model for everything.

  1. TypeSafe AI ships Jev, a cheap decision layer built to replace LLM calls in routing and triage

    For a support or ops team that currently pays an LLM to tag ticket urgency, screen resumes or decide which model handles a request, Jev is a much cheaper intermediate step: it returns a structured score or choice with a confidence value instead of free text, so you can auto-route high-confidence cases and send low-confidence ones to a human. The catch is the 67.8% accuracy figure and the lack of any explanation for how it reached a decision, so it's a triage layer, not a replacement for judgment on anything with real downside if it's wrong — and adversarial inputs (a crafted support message, a doctored CV) can still hijack its reasoning without strong guardrails around it.

  1. Together AI ships a $17 recipe for custom support-ticket classifiers

    For a support team fielding a few hundred tickets a day, this is a concrete blueprint for a first-line triage bot: feed in the customer message plus your existing intent list, get back a single label in under a token's worth of output, and route accordingly. The training cost (~$17) and time (~25 minutes) make it cheap enough to iterate on your own historical tickets rather than rely on a general-purpose LLM call for every classification, and the fixed temperature=0, max_tokens=8 setup makes outputs predictable enough to feed directly into a routing rule or CRM webhook without a human review step for high-confidence cases.

  1. Anthropic's Opus 5.5 lands on AWS Bedrock with lower cost per task

    For a 10-200 person B2B company running agentic coding assistants or document-heavy knowledge work on Bedrock, Opus 5.5 offers a direct lever: lower average cost per task and clearer step-by-step reporting during long-running sessions, which matters for anyone reviewing agent output before it reaches a customer or a contract. The catch is the new safety classifiers, which refuse more requests than earlier Opus versions in areas like biology and cybersecurity — teams with support or ops bots that occasionally touch security-adjacent topics (vulnerability triage, incident response drafting) should test refusal behavior before rolling this model into production, since a blocked response mid-workflow is worse than a slightly more expensive one.

  2. TypeSafe AI's Jev Turns Classification Into a Cheap API Call

    For a B2B team running lead scoring, support ticket triage, spam filtering or search relevance ranking, Jev's format maps directly onto those workflows: feed it a customer record or ticket text plus a set of yes/no or scored questions, and get back structured confidence numbers cheaply enough to run on every record rather than a sampled subset. The catch is that Jev gives no explanation for its scores, so anything touching hiring, credit, or other high-stakes decisions needs structured evals before deployment, and even lower-stakes uses like ticket prioritization should be spot-checked for skewed outputs.

  1. Anthropic's Claude Opus 5.5 lands in Microsoft Foundry with cheaper tokens and clearer agent reporting

    For a 10-200 person B2B company running agents on Microsoft Foundry to draft reports, triage support tickets, or refactor internal tooling, the practical change is twofold: lower cache and token costs make longer agent sessions cheaper to run in production, and the model's new habit of surfacing what it did, what it found, and where it needs input reduces the amount of manual review needed before trusting an agent's output. Adaptive thinking also removes a configuration step - teams no longer need to hand-tune reasoning budgets per task, which matters for lean ops teams without dedicated AI engineers.

  2. Google Ships Two New TTS Models Built for Production Voice Agents

    For a support or sales team considering a voice agent, this changes what's technically possible today: the Gemini API is live now for developers to build custom-branded voices, multilingual dubbing, or conversational agents with realistic pacing and backchanneling, and Flash-Lite is explicitly pitched for high-volume, cost-efficient voice agents. The catch is that Gemini Enterprise access is still "coming soon" via API, so a 10-200 person company without in-house developers will likely need to wait or work through a partner platform (Agora, LiveKit, Pipecat, Vercel are named integrators) rather than get this through an enterprise console today.

  1. Salesforce Builds a Reasoning Model Trained on Enterprise Workflows, Not General Knowledge

    For a company running sales or support through Agentforce, the practical change is consistency: a model trained specifically on qualifying leads, routing cases and scheduling follow-ups should apply the same rule to the hundredth ticket as the first, rather than reasoning it out differently each time the way a general-purpose model does. The more concrete win is the failure mode Salesforce says it targeted directly — when the right tool isn't available, Koa is trained to say so and hand off to a human rather than call a similar tool or confirm an action that never happened, which is exactly the kind of silent error that erodes trust in automated support and sales workflows. Teams already on Agentforce piloting Koa in service, sales or commerce should watch whether that discipline holds up outside Salesforce's own benchmarks before routing high-stakes cases (refunds, compliance-adjacent qualification) through it unsupervised.

  1. Google Adds Reasoning to Its Real-Time Voice AI, Opening the Door to Smarter Phone Agents

    For a 10-200 person B2B company running phone support or a sales qualification line through voice AI, this closes the biggest gap in current voice agents: handling anything beyond a single-turn lookup. A support call that requires checking an order status, then applying a conditional refund rule, then confirming with the customer previously needed a handoff to a human or a scripted decision tree. Extended Thinking lets the agent reason through that sequence live, on the call, which means fewer escalations and shorter average handle time for the tier-one queue. Teams evaluating or already running voice bots for inbound support or outbound qualification should treat this as the point to re-test latency and accuracy on their actual call scripts — reasoning modes typically add processing time, so the tradeoff between depth and response speed needs to be measured before rolling it into a live queue, not assumed.

  1. Claude 5.1 Lands on Amazon Bedrock, Widening Model Choice for AWS-Based Ops Teams

    If your support ticketing, sales-enablement, or internal copilots already call Claude through Bedrock, this is a low-friction upgrade: change the model ID in your existing integration rather than re-platforming. Before flipping the switch on a production workflow — a support triage bot, a CRM summarizer, a contract-review assistant — run the new version against a sample of real tickets or deals and compare output quality, latency and per-call cost side by side with the model you're currently paying for. Anthropic and AWS have not published independently verified benchmark deltas for this release as of writing, so treat any capability claims as unconfirmed until you've tested against your own data. Companies not yet on Bedrock gain another reason to consolidate model access through AWS if they're already paying for EC2, S3 or other AWS services, since it simplifies billing and IAM permissions compared to managing a separate Anthropic API key.

  1. Google Tightens Developer Controls on Gemini Omni Flash Model

    If your support bot, lead-qualification agent, or internal ops tool runs on Gemini Flash, this update matters because tighter control over output structure and behavior typically reduces the post-processing and validation layer you'd otherwise build to catch inconsistent responses. A 10-200 person B2B company running a Flash-based automation can potentially simplify its prompt engineering and reduce error-handling code, but only after testing the new controls against existing production prompts — assume nothing works identically until verified in a staging environment.

  2. New Open Encoder Model Adds Multilingual Image-Text Search to RAG Pipelines

    If your support or sales team searches across product manuals, screenshots, or tickets in more than one language, you likely run separate embedding models for text and images today, which adds latency and integration overhead. A single multilingual, multimodal encoder like NeoMME could let you consolidate that into one retrieval pipeline — useful for support teams handling attachments (screenshots, scanned invoices, product photos) alongside text queries in different languages. Before switching, confirm NeoMME's retrieval accuracy on your actual document types against your current encoder; open weights mean you can test this in a staging environment without vendor lock-in, but benchmarks from the source blog have not been independently verified.

  1. Google Ships Gemini 3.7 Flash, a Faster Model for High-Volume Automation Tasks

    If your support or sales stack routes high-volume, low-complexity tasks — first-response drafting, ticket classification, inbound lead scoring — through a Flash-tier Gemini model, this release is worth a benchmark test before you assume it's a straight upgrade. Flash models are chosen specifically for cost and speed rather than peak reasoning, so the real question for a 10-200 person company is whether 3.7 Flash cuts per-ticket or per-call cost at the same accuracy, not whether it's smarter. Anyone with existing automations wired to a previous Flash version should re-run their eval set against 3.7 before switching in production, since silent regressions in tone or accuracy are common even in point releases.

  1. Liquid AI Ships a Compact Vision Model That Runs Without Cloud APIs

    For a 10-200 person company handling support tickets with photo attachments, processing scanned invoices, or verifying shipment/damage images, this model type means that work can run on local or on-prem hardware instead of a per-call cloud vision API — cutting marginal cost to near zero and removing the need to send customer images to a third-party service. Teams building internal tools for receipt/invoice OCR, quality-control photo review, or ID verification in onboarding flows get a smaller, cheaper model to self-host behind existing infrastructure, which matters if data residency or per-transaction API cost has been a blocker to automating those steps. It does not replace larger cloud vision models for complex reasoning over images, but it closes the gap for high-volume, simple visual classification and extraction tasks that make up most support and back-office image workloads.

  1. OpenAI publishes GPT-5.6 builder guide with new tool-calling and context specs

    If you have an AI agent handling inbound support tickets, qualifying leads, or triaging ops requests, the guide's tool-calling recommendations matter more than the model's raw benchmark scores: unreliable function calls mean an agent that silently fails to update a CRM record or escalate a ticket, and nobody notices until a customer complains. Teams running 10-200 person operations should re-test any GPT-5.6-based agent against their actual tool schemas (not just chat prompts) before treating it as a drop-in upgrade, and check whether prompt or workflow changes recommended in the guide require updating existing automation logic to avoid regressions in accuracy or latency.

  1. OpenAI Ships GPT-5.6, Pitches It as Cheaper Per Task Than GPT-5

    If your sales or support automation runs on OpenAI's API — lead qualification bots, ticket triage, call summarization, CRM enrichment — this release is worth a look purely on cost grounds. Price-performance improvements in a new model version typically translate into lower per-call spend or faster throughput at the same spend, which matters when you're running thousands of automated interactions a month. The practical move is not to rush to adopt GPT-5.6 blindly, but to have whoever manages your model calls (in-house or your automation vendor) benchmark it against your current model on your actual prompts — support macros, sales scripts, whatever you've built — before switching. Model upgrades sometimes shift output tone or formatting slightly, which can break brittle prompt chains or downstream parsing. Treat this as a scheduled maintenance item: check cost, check quality, then migrate if it holds up.

  2. OpenAI Lays Out Vision for "Abundant Intelligence," Light on Product Specifics

    For a 10-200 person B2B company, this particular post changes nothing operationally this week — there is no new model, API, price, or SDK to evaluate. What it does signal is direction: OpenAI is publicly framing its roadmap around making high-quality AI cheap and ubiquitous, which historically has preceded price drops and capability jumps that make previously uneconomical automation (deeper support triage, multi-step sales research, ops reporting) suddenly viable. The sensible operator response is not to build anything new today, but to keep a running list of manual, judgment-heavy workflows currently deemed "too expensive to automate" — because the cost curve behind this kind of announcement tends to move faster than internal roadmaps expect.

  3. OpenAI Tunes GPT-5.6 Sol's Behavior, Opens Luna to Free ChatGPT Users

    For a 10-200 person B2B company, this is a low-drama update but worth a note to whoever owns your AI tooling stack: if staff use free-tier ChatGPT for drafting emails, summarizing calls, or triaging support tickets, their default model behavior just changed without any action on your part. That's the real risk with consumer AI tools embedded in business workflows — model updates roll out silently and can shift output tone, accuracy, or refusal patterns overnight. If any part of your sales or support process leans on ChatGPT outputs going to customers unreviewed, this is a good prompt to spot-check recent outputs against what you were getting last week, and to confirm whether your team is on a paid tier where model versioning is more predictable.

  1. Hugging Face's Mid-2026 Model Report: Open Models Now Match Closed Ones on Most Business Tasks

    If you're running sales, support or ops automation on a closed-model API today, this matters because it changes your leverage. Open models that perform close to parity mean you can credibly threaten to switch, negotiate pricing with incumbent vendors, or run sensitive workflows — like customer data enrichment or internal ticket triage — on self-hosted infrastructure instead of sending it to a third party. It doesn't mean rip-and-replace tomorrow: switching costs, fine-tuning work and integration testing are real. But it means your next vendor renewal conversation should include "what's our open-model fallback" as a genuine line item, not a hypothetical.

  2. OpenAI Previews Ultrafast Mode for GPT-5.6, Promising 14x Faster Responses

    For a B2B company running automated support chat, voice agents, or real-time sales qualification bots, latency is often the difference between a tool people actually use and one they abandon mid-task. A 14x speed claim, if it holds up in production and not just cherry-picked demos, could make agentic workflows — the kind that chain multiple model calls together for a single customer interaction — feel instant rather than sluggish. That matters most for voice-based support and live chat handoffs, where every second of "thinking" time costs trust. The caveat: speed previews from model labs frequently ship with caveats around cost multipliers or reduced context windows, so treat this as a signal to watch, not a reason to re-architect anything yet.

Next step

Free AI Diagnostic

Fifteen minutes, no email required. It maps where your work actually goes and ranks what is worth automating first.

Start the free diagnostic

Starts immediately in the browser.

Fee
Free
Length
15 minutes

You keep the ranked list of candidates either way.