Skip to content

OpenAI Launches "Presence," an Embodied Voice Agent for Real-Time Screen and Device Interaction

Short answer

OpenAI has introduced Presence, a new product built to give AI a persistent, embodied "presence" that can see a user's screen, hear audio, and take real-time action across devices via voice. It matters because it moves beyond chat-based assistants toward always-on, context-aware agents that could sit inside daily sales, support and ops workflows — pending confirmation of API access and pricing.

What this means for operators

For a 10-200 person B2B company, the interesting part isn't the demo — it's what happens when a voice-driven, screen-aware agent can be pointed at your CRM, helpdesk queue, or internal dashboards without someone typing a prompt first. If Presence or its underlying capabilities become available via API, the realistic near-term use is narrow: a rep or support agent gets a live assistant that watches a screen during a call and surfaces account history, past tickets, or pricing without switching tabs. That's a workflow change, not a headcount change — treat early access claims with caution until OpenAI publishes actual API terms, latency numbers, and pricing, since "real-time" and "always-on" products are exactly where cost and reliability surprises show up first.

OpenAI has introduced a new product called Presence, described in OpenAI's announcement as a step toward giving AI systems a persistent, embodied "presence" in a user's environment — one that can see a screen, hear audio, and respond or act in real time through voice, rather than requiring a typed prompt and a back-and-forth chat exchange.

Based on the source announcement, Presence is positioned as a shift away from the request-response pattern that has defined most consumer and business AI tools since ChatGPT's launch. Instead of a user opening a chat window, typing a question, and waiting for a text reply, Presence is framed as something that stays active alongside a user's work — observing context on screen and responding conversationally, similar in spirit to OpenAI's existing voice mode but extended toward continuous, situational awareness rather than a single turn-based session.

Specific technical details — which models power Presence, what data it retains, whether it runs locally or in the cloud, and what latency it achieves — were not fully detailed in the source material reviewed for this item and should be treated as unconfirmed until OpenAI publishes documentation or a developer-facing spec. Likewise, availability terms — whether Presence ships as a consumer feature inside ChatGPT, a standalone app, or an API that third-party developers and vendors can build on — are not yet confirmed. Given OpenAI's typical rollout pattern with recent features (staged access, waitlists, or premium-tier gating), operators should not assume immediate general availability.

The announcement continues OpenAI's broader push into agentic, multimodal products following earlier releases in voice, vision, and computer-use capabilities. Presence appears aimed at closing the gap between "AI that answers questions" and "AI that participates in a task as it happens" — a distinction that matters for any workflow involving live screens, live calls, or live decision-making, which describes much of day-to-day sales and support work.

For consultancies and internal teams evaluating this, the practical takeaway is to wait for confirmed API access and pricing before building roadmap commitments around Presence specifically. The concept — an agent with real-time visual and audio context — is directionally consistent with where automation tooling for sales and support has been heading: fewer manual lookups, less tab-switching, and assistants that carry context forward automatically. Whether Presence itself becomes the vehicle for that in a 10-200 person company's stack, or whether it remains a consumer-facing showcase that later trickles into enterprise tooling, is not yet established and should be tracked as OpenAI releases further technical and commercial details.

Source: OpenAI