Skip to content

Microsoft Foundry Adds Two Cheaper GPT-6 Tiers for Production Agents

Short answer

Microsoft made GPT-6 Sol and GPT-6 Luna generally available in Microsoft Foundry alongside GPT-6 Astra, creating a three-tier model lineup for production agents. Astra handles complex reasoning, Sol covers general production workloads, and Luna is priced for high-volume tasks like extraction, summarization and request routing at a fraction of Astra's cost.

What this means for operators

For a company running support or ops agents on Azure, this is a direct cost lever: instead of sending every ticket-classification or data-extraction step through an expensive reasoning model, teams can now put GPT-6 Luna behind routine, high-volume steps and reserve Sol or Astra for the parts of a workflow that actually need judgment. At Global Standard short-context pricing, Luna runs $0.10 per million input tokens and $0.50 output versus Sol's $2/$10 and Astra's $10/$50, which changes the math on cost-per-task for anyone running thousands of routine agent calls a day. The practical move is to map an existing agent workflow, identify which steps are classification/routing/summarization versus multi-step reasoning, and split the model assignment accordingly rather than defaulting to one model for everything.

Microsoft has expanded its GPT-6 lineup in Microsoft Foundry, making GPT-6 Sol and GPT-6 Luna generally available alongside the existing GPT-6 Astra model.

The three models are positioned for different jobs in an agent pipeline. Astra is recommended for demanding reasoning, software engineering and computer-use tasks. Sol is aimed at general-purpose production agents, coding and multi-step knowledge work. Luna is built for high-volume, lower-complexity tasks: extraction, summarization, request routing and routine customer interactions.

Pricing scales with capability. At Global Standard rates for short context, Astra costs $10.00 per million input tokens and $50.00 per million output tokens; Sol costs $2.00 input / $10.00 output; Luna costs $0.10 input / $0.50 output. Data Zone deployments in the US and EU carry a 10-20% premium over Global pricing. Microsoft is pushing customers to evaluate "cost per task" rather than per-token pricing when choosing which model handles which step.

Deployment options vary by model. Standard deployment is available for all three models across 28 Global regions and US/EU Data Zones. Provisioned Throughput, which reserves capacity, is available for Astra and Sol. Priority Processing, a faster pay-as-you-go lane, is available only for Sol.

Microsoft cites Manus and Wolters Kluwer Tax & Accounting as customers already running agentic workflows on Azure OpenAI models, with Wolters Kluwer's CTO noting the models' step-by-step reasoning suits research and compliance workflows that need to "hold up to scrutiny."

Foundry's evaluation and monitoring tools are meant to help teams decide, with evidence, which model tier fits which workload rather than defaulting to the newest or most capable option for every task.

Source: Azure Blog

Next step

Visibility Analyzer

This is what the Visibility Analyzer measures on a real site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.