Skip to content

Models & capabilities, read from an operations desk

Everything we have published under Models & capabilities, read from an operations desk: what it changes for a B2B company of 10-200 people.

  1. Latest

    Microsoft Foundry Adds Two Cheaper GPT-6 Tiers for Production Agents

    Microsoft made GPT-6 Sol and GPT-6 Luna generally available in Microsoft Foundry alongside GPT-6 Astra, creating a three-tier model lineup for production agents. Astra handles complex reasoning, Sol covers general production workloads, and Luna is priced for high-volume tasks like extraction, summarization and request routing at a fraction of Astra's cost.

    What changes for operators — For a company running support or ops agents on Azure, this is a direct cost lever: instead of sending every ticket-classification or data-extraction step through an expensive reasoning model, teams can now put GPT-6 Luna behind routine, high-volume steps and reserve Sol or Astra for the parts of a workflow that actually need judgment. At Global Standard short-context pricing, Luna runs $0.10 per million input tokens and $0.50 output versus Sol's $2/$10 and Astra's $10/$50, which changes the math on cost-per-task for anyone running thousands of routine agent calls a day. The practical move is to map an existing agent workflow, identify which steps are classification/routing/summarization versus multi-step reasoning, and split the model assignment accordingly rather than defaulting to one model for everything.

  1. Anthropic's Claude Opus 5.5 lands in Microsoft Foundry with cheaper tokens and clearer agent reporting

    For a 10-200 person B2B company running agents on Microsoft Foundry to draft reports, triage support tickets, or refactor internal tooling, the practical change is twofold: lower cache and token costs make longer agent sessions cheaper to run in production, and the model's new habit of surfacing what it did, what it found, and where it needs input reduces the amount of manual review needed before trusting an agent's output. Adaptive thinking also removes a configuration step - teams no longer need to hand-tune reasoning budgets per task, which matters for lean ops teams without dedicated AI engineers.

Next step

Free AI Diagnostic

Fifteen minutes, no email required. It maps where your work actually goes and ranks what is worth automating first.

Start the free diagnostic

Starts immediately in the browser.

Fee
Free
Length
15 minutes

You keep the ranked list of candidates either way.