OpenAI was among the first major companies to present its own web navigation AI agent, Operator, and now that other technology companies are rolling out their own, the developer has once again taken the lead with its latest major launch: ChatGPT agent.
This is a unified agency system that combines the capabilities of Operator and the advanced reasoning of Deep Research. In other words: a much more intelligent conversational AI agent, capable of autonomously navigating the web to perform tasks on your behalf, executing code, connecting to your applications, and creating presentations and spreadsheets.
ChatGPT can now do work for you using its own computer. Introducing ChatGPT agent—a unified agentic system combining Operator’s action-taking remote browser, deep research’s web synthesis, and ChatGPT’s conversational strengths. pic.twitter.com/7uN2Nc6nBQ — OpenAI (@OpenAI) July 17, 2025
ChatGPT can now do work for you using its own computer.
Introducing ChatGPT agent—a unified agentic system combining Operator’s action-taking remote browser, deep research’s web synthesis, and ChatGPT’s conversational strengths. pic.twitter.com/7uN2Nc6nBQ
— OpenAI (@OpenAI) July 17, 2025
ChatGPT agent is already available to all users of the ChatGPT Pro plan. OpenAI has confirmed that this capability will be extended to the Plus and Team plans starting Monday, July 21st, while the Enterprise and Edu plans will have to wait a few more weeks.
OpenAI introduced Operator at the end of January 2025 and Deep Research at the beginning of February. These two tools were extremely powerful, but the developer ultimately concluded that, if combined, they would be even more so. That is how ChatGPT agent came into being.
“Operator couldn’t dive deep into analysis or write detailed reports, and deep research couldn’t interact with websites to refine results or access content requiring user authentication”, explains OpenAI on its blog.
In this way, ChatGPT agent is built upon the power of three of OpenAI’s most significant technologies. On one hand, there is Operator, which enables it to interact with websites; on another, Deep Research, which grants the agent the capability to conduct advanced research, analyzing and synthesizing complex information; and finally, ChatGPT itself, providing conversational fluency for interaction.
The result is an AI agent capable of navigating the web, reasoning about the best steps to take in order to fulfill the assigned task. Thanks to this, the ChatGPT agent can do everything from checking your calendar and informing you about upcoming meetings, to researching a recipe and purchasing the necessary ingredients, or performing benchmarking on your direct competitors and presenting the findings in a presentation.
“ChatGPT carries out these tasks using its own virtual computer, fluidly shifting between reasoning and action to handle complex workflows from start to finish, all based on your instructions”, states OpenAI.
Of course, this does not mean you shall surrender complete control of the web, your applications, and your data to ChatGPT agent. Each time the agent needs to perform sensitive actions, such as logging in, it will request your permission first. In addition, you will always retain absolute control, with the ability to interrupt the agent’s activity at any time.
To use ChatGPT agent, you must select the “agent mode” in the dropdown tools menu located in the text box where you enter your prompts. You may do this at any point during your conversation with ChatGPT; you only need to describe the task you wish it to perform.
As the agent works, there will be a screen narration indicating the steps it is taking. You can interrupt or pause its activity at any time, taking back control of the browser. “ChatGPT agent is designed for iterative, collaborative workflows, far more interactive and flexible than previous models. Likewise, ChatGPT itself may proactively seek additional details from you when needed to ensure the task remains aligned with your goals”.
While active, ChatGPT agent can utilize ChatGPT connectors to access applications such as Gmail or Github. It will also be able to log in to any website through its own browser. “Giving ChatGPT these different avenues for accessing and interacting with web information means it can choose the optimal path to most efficiently perform tasks.
For instance, it can gather information about your calendar through an API, efficiently reason over large amounts of text using the text-based browser, while also having the ability to interact visually with websites designed primarily for humans”.
If you have the ChatGPT app installed, once the agent has completed the task, you will receive a notification.
This product also introduces new risks due to its remarkable capabilities. The fact that it can carry out tasks online for you, work with your data, or log in to platforms introduces a series of potential dangers that OpenAI has made efforts to mitigate.
“We’ve strengthened the robust controls from Operator’s research preview and added safeguards for challenges such as handling sensitive information on the live web, broader user reach, and (limited) terminal network access. While these mitigations significantly reduce risk, ChatGPT agent’s expanded tools and broader user reach mean its overall risk profile is higher”, the company explains.
Anticipating possible threats, OpenAI has implemented mitigations within the model enabling it to require explicit user confirmation for sensitive actions (logins, purchases…), request active supervision over activities such as sending emails, and reject tasks considered high-risk (for example, bank transfers).
In addition, the following additional controls have been implemented:
The announcement of ChatGPT agent came just one day after Google launched its new Search feature that allows an AI agent to make phone calls to local businesses on your behalf to obtain information or make reservations.
It is clear that technology giants are moving towards a future where their tools serve as a direct bridge to outcomes, relieving us of preliminary processes. Would you like to book a restaurant? Google calls and makes the reservation for you; you only need to go and enjoy the meal. Need to organize a trip to Rome and buy plane tickets and reserve a hotel? A web navigation AI agent will research flight combinations and hotel locations.
Of course, this is the objective, but before reaching that point there must be an adoption phase during which people not only become accustomed to this new paradigm, but also gain trust in delegating these tasks to AI. The pace of development by the companies and the guarantees they offer will also influence this.
What is clear is that in just half a year we have witnessed an impressive evolution in web navigation AI agents, and the future offers many possibilities. In addition to Operator, we have also become acquainted with Amazon Nova Act, an AI introduced in April that enables the creation of agents capable of taking actions within a web browser.
Another example is Google’s Project Mariner, a web navigation AI agent capable of observing browser information, interpreting your requests, reasoning, establishing a plan, and carrying it out. This was introduced as a research prototype in December 2024, but was not officially launched in the United States until May 2025 for individuals subscribed to the Google AI Ultra plan.
Photo: OpenAI
Your email address will not be published. Required fields are marked *
Δ