GPT-4o
Large multimodal model from OpenAI
From Wikipedia, the free encyclopedia
GPT-4o ("o" for "omni") is a multilingual, multimodal generative pre-trained transformer developed by OpenAI and released in May 2024.[1] It can process and generate text, images and audio.[2][3]
| GPT-4o | |
|---|---|
| Developer | OpenAI |
| Release | May 13, 2024 |
| Preview release | ChatGPT-4o-latest (2025-03-26)
/ March 26, 2025 |
| Predecessor | GPT-4 Turbo |
| Successor | |
| Type | |
| License | Proprietary |
| Website | |
Upon release, GPT-4o was free in ChatGPT, though paid subscribers had higher usage limits.[4] GPT-4o was removed from ChatGPT in August 2025 when GPT-5 was released, but OpenAI reintroduced it for paid subscribers after users complained about the sudden removal.[5]
GPT-4o's audio-generation capabilities are used in ChatGPT's Advanced Voice Mode.[6] On July 18, 2024, OpenAI released GPT-4o mini, a smaller version of GPT-4o which replaced GPT-3.5 Turbo on the ChatGPT interface.[7] The image generation model GPT Image 1, which is based on GPT-4o, replaced DALL-E 3 in ChatGPT in March 2025.[8][9]
OpenAI retired GPT-4o from ChatGPT on February 13, 2026.[10] However, as of February 2026 the voice mode is still powered by GPT-4o or GPT-4o mini, depending on the usage and plan.[11]
Background
Before the release of GPT-4o, OpenAI had released GPT-4 and GPT-4 Turbo. GPT-4 supported text and image input, while GPT-4 Turbo was introduced as a later version of GPT-4 with a larger context window and lower usage cost. By early 2024, OpenAI had also expanded ChatGPT with voice and image-related features, but these interactions generally depended on separate systems for speech recognition, language processing, and speech synthesis. GPT-4o was developed in the context of OpenAI's work on a more unified multimodal model for text, audio, image, and video interaction.[12][13]
Contemporary coverage described GPT-4o as a successor to GPT-4 Turbo in OpenAI's GPT-4 model line, with particular attention to response speed, cost, and support for voice and visual interaction. The Guardian reported that the model was faster than previous versions and that some features previously associated with paid ChatGPT plans would become available to free users. MarketWatch similarly described GPT-4o as a new flagship model and compared its speed and cost with GPT-4 Turbo.[14][15]
Before its public announcement, multiple versions of GPT-4o were tested anonymously on LMSYS Chatbot Arena under names including "gpt2-chatbot", "im-a-good-gpt2-chatbot", and "im-also-a-good-gpt2-chatbot". The appearance of these models led to speculation that OpenAI was testing a new system. After the release of GPT-4o, OpenAI researcher William Fedus confirmed that a version of GPT-4o had been tested on the Arena under one of those names.[16]
OpenAI announced GPT-4o during its Spring Update livestream on May 13, 2024. The event was led by chief technology officer Mira Murati and focused on GPT-4o and updates to ChatGPT rather than GPT-5 or a standalone search product, both of which had been the subject of prior speculation. Demonstrations during the event included spoken conversation, visual understanding, translation, math tutoring, and coding assistance.[17][18]
Lifecycle
GPT-4o was released on May 13, 2024, and was gradually made available in ChatGPT and through OpenAI's developer products. At launch, OpenAI made GPT-4o available to ChatGPT users with partial access for free users, while developers received access to text and vision capabilities through the API. The new voice mode associated with GPT-4o was scheduled for a staged rollout to ChatGPT Plus users rather than being fully available on the first day of release.[19][20]
In August 2024, OpenAI published the GPT-4o system card. The document described the model's supported inputs and outputs, including text, audio, image, and video inputs and text, audio, and image outputs. It also summarized pre-release safety testing, risk mitigations, and evaluations across language, vision, audio, and some professional tasks.[21]
In October 2024, OpenAI introduced Canvas, a writing and coding interface within ChatGPT. During its beta release, Canvas was built with GPT-4o and allowed users to edit text or code in a separate workspace while continuing to interact with ChatGPT. It was first made available to ChatGPT Plus and Team users, with later expansion planned for Enterprise, Edu, and free users.[22][23]
Later in October 2024, OpenAI launched ChatGPT Search, integrating web search into ChatGPT. The search feature used a fine-tuned version of GPT-4o to generate answers based on current web information and provide links to sources. It was initially made available to paid ChatGPT users, with later availability for free, enterprise, and education users.[24][25]
In November 2024, OpenAI made GPT-4o snapshots available through the API, including versions such as "gpt-4o-2024-11-20"". These snapshots allowed developers to call specific model versions rather than relying only on a general model alias. OpenAI's API documentation listed GPT-4o as accepting text and image input and producing text output.[26]
In December 2024, Apple released iOS 18.2, iPadOS 18.2, and macOS Sequoia 15.2, adding ChatGPT integration to Apple Intelligence features such as Siri and Writing Tools. The integration allowed users to access ChatGPT from within supported Apple system features. Coverage of the release generally described it as ChatGPT integration rather than a separate GPT-4o product launch.[27][28]
In January 2025, OpenAI updated GPT-4o's knowledge cutoff from November 2023 to June 2024. OpenAI's model release notes also described changes to visual understanding, including complex charts, spatial relationships, and related reasoning tasks.[29]
In March 2025, OpenAI integrated image generation into GPT-4o. The update allowed ChatGPT to create and edit images using GPT-4o, replacing or supplementing workflows that had previously relied mainly on DALL-E. The feature was rolled out to ChatGPT Plus, Pro, Team, and Free users. Several later studies evaluated GPT-4o's image generation and editing behavior, including its instruction following, editing precision, and limitations in tasks requiring spatial or knowledge-based reasoning.[30][31][32]
In April 2025, ChatGPT's memory features were expanded. In addition to explicitly saved memories, ChatGPT could reference previous conversations to provide more continuous context. The feature was first rolled out to ChatGPT Plus and Pro users.[33][34]
In August 2025, after the release of GPT-5, GPT-4o's availability in ChatGPT changed. GPT-5 became the default model for signed-in ChatGPT users, while GPT-4o was no longer the default model and was temporarily removed from the previous model selection flow. This was not the final retirement of GPT-4o; OpenAI later restored optional access to GPT-4o for subscribed users.[35][36]
In September 2025, OpenAI adjusted ChatGPT's safety routing mechanisms. The change routed some sensitive conversations to reasoning models such as GPT-5 and was introduced alongside parental control features. This change concerned ChatGPT's model routing and safety handling mechanisms rather than a new GPT-4o model release.[37][38]
In October 2025, after rumors circulated that GPT-4o might soon be retired, OpenAI told TechRadar that it did not have plans to discontinue GPT-4o in October and that users would receive advance notice if a retirement occurred.[39]
In November 2025, OpenAI notified developers that the "chatgpt-4o-latest" model snapshot would be deprecated and removed from the API on February 17, 2026. The deprecation applied to a specific ChatGPT snapshot of GPT-4o rather than all GPT-4o snapshots at the same time. OpenAI's API deprecations documentation also listed later shutdown dates for GPT-4o-related API models, including "gpt-4o-realtime-preview"", ""gpt-4o-audio-preview"", and "gpt-4o-2024-05-13"".[40]
In January 2026, OpenAI announced the final retirement of GPT-4o from ChatGPT. Under the plan, GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini were scheduled to be retired from ChatGPT on February 13, 2026; OpenAI stated that the API was not affected at that time.[41] OpenAI's Help Center later stated that ChatGPT Business, Enterprise, and Edu users could continue using GPT-4o in custom GPTs until April 3, 2026; after that date, GPT-4o was fully retired from all ChatGPT plans.[42]
Capabilities
GPT-4o is a multimodal model used for text, image, audio, video-related input, coding, and developer API use. At release, third-party coverage described GPT-4o as a model intended to combine text, voice, and visual interaction in ChatGPT, rather than treating those functions as separate product experiences.[43][44]
Text
In text-based use, GPT-4o was presented as a successor to GPT-4 Turbo in OpenAI's GPT-4 model line. Contemporary coverage emphasized faster response speed, lower cost, multilingual use, and wider availability in ChatGPT compared with earlier GPT-4-based offerings.[45][46]
A 2024 academic evaluation, Putting GPT-4o to the Sword, assessed GPT-4o across language, vision, speech, and multimodal tasks. The study treated GPT-4o as a general-purpose multimodal large language model and evaluated its performance across multiple task categories rather than only in text conversation.[47]
Image
GPT-4o supports image input and visual understanding, allowing users to ask questions about screenshots, charts, photographs, and other visual materials. During the May 2024 launch demonstrations, GPT-4o was shown interpreting visual information, including a math problem and screen-based content.[48][49]
In March 2025, image generation was integrated into GPT-4o, allowing ChatGPT to create and edit images using GPT-4o. The update replaced or supplemented image-generation workflows that had previously relied mainly on DALL-E, and it was rolled out across several ChatGPT plans.[50][51]
Academic studies have also evaluated GPT-4o's visual understanding and image generation behavior. One 2025 study benchmarked GPT-4o and other multimodal foundation models on standard computer vision tasks, finding that such models were not close to specialist state-of-the-art computer vision systems but could function as generalist visual models in several semantic tasks.[52] Another study examined GPT-4o's image generation ability, including instruction following, editing precision, and limitations in spatial or knowledge-based reasoning.[53]
Audio and voice
GPT-4o was associated with ChatGPT's advanced voice capabilities. At release, coverage described the model as supporting more conversational voice interaction than earlier ChatGPT voice features, including faster responses and the ability to handle interruptions during conversation.[54][55]
Research has examined GPT-4o voice mode as part of the broader development of audio-language models. A 2025 preliminary study evaluated GPT-4o voice mode on audio, speech, and music understanding tasks, reporting strengths in some speech and semantic reasoning tasks while also noting limitations in areas such as audio duration prediction and instrument classification.[56]
The voice interface also became a subject of safety research. A 2024 study on voice jailbreak attacks against GPT-4o examined whether harmful or restricted requests could be made more effective when adapted into spoken interaction rather than text prompts.[57]
Video and screen input
GPT-4o was presented as capable of handling visual information from camera or screen-based contexts, but it was not introduced as a video generation model. Launch demonstrations and media coverage focused on visual reasoning over live or uploaded visual input, such as interpreting objects, screenshots, or written work, rather than generating video output.[58][59]
Code
GPT-4o was used for coding assistance in ChatGPT and was demonstrated on code-related tasks during the Spring Update event. Coverage of the launch described coding assistance as one of the example uses shown alongside voice, visual, and language tasks.[60]
Independent evaluations have also included GPT-4o in broader assessments of multimodal and language-model performance, including tasks involving reasoning, text generation, and code-related problem solving.[61]
API and developer access
GPT-4o was made available through OpenAI's API after release, with third-party coverage noting that it was being rolled out across both consumer-facing and developer products.[62] The API was used by developers for text and vision-related tasks, and later became part of OpenAI's broader model lifecycle that included snapshots, fine-tuning, and deprecations.
Customization and fine-tuning
GPT-4o supported customization through fine-tuning. In August 2024, OpenAI made fine-tuning for GPT-4o generally available, allowing developers to train the model on custom datasets for specific use cases and to adjust aspects such as response style, tone, or task-specific behavior.[63]
Commercial fine-tuning APIs involving GPT-4o have also been evaluated in academic research. The FineTuneBench study examined whether commercial fine-tuning APIs could learn new or updated knowledge, including tests involving GPT-4o. The study found limitations in the ability of such APIs to reliably incorporate new information or update existing knowledge, and reported that GPT-4o mini outperformed GPT-4o in some of its knowledge-infusion settings.[64]
Corporate customization
In August 2024, OpenAI introduced a new feature allowing corporate customers to customize GPT-4o using proprietary company data. This customization, known as fine-tuning, enables businesses to adapt GPT-4o to specific tasks or industries, enhancing its utility in areas like customer service and specialized knowledge domains. Previously, fine-tuning was available only on the less powerful model GPT-4o mini.[65][66]
The fine-tuning process requires customers to upload their data to OpenAI's servers, with the training typically taking one to two hours. OpenAI's focus with this rollout is to reduce the complexity and effort required for businesses to tailor AI solutions to their needs, potentially increasing the adoption and effectiveness of AI in corporate environments.[67][65]
GPT-4o mini
On July 18, 2024, OpenAI released a smaller and cheaper version, GPT-4o mini.[68]
According to OpenAI, its low cost is expected to be particularly useful for companies, startups, and developers that seek to integrate it into their services, which often make a high number of API calls. Its API costs $0.15 per million input tokens and $0.6 per million output tokens, compared to $2.50 and $10,[69] respectively, for GPT-4o. It is also significantly more capable and 60% cheaper than GPT-3.5 Turbo, which it replaced on the ChatGPT interface.[68] The price after fine-tuning doubles: $0.3 per million input tokens and $1.2 per million output tokens.[69]
Controversies
Scarlett Johansson controversy
As released, GPT-4o offered five voices: Breeze, Cove, Ember, Juniper, and Sky. A similarity between the voice of American actress Scarlett Johansson and Sky was quickly noticed. On May 14, Entertainment Weekly asked themselves whether this likeness was on purpose.[70] On May 18, Johansson's husband, Colin Jost, joked about the similarity in a segment on Saturday Night Live.[71] On May 20, 2024, OpenAI disabled the Sky voice.[72]
Scarlett Johansson starred in the 2013 sci-fi movie Her, playing Samantha, an artificially intelligent virtual assistant personified by a female voice. As part of the promotion leading up to the release of GPT-4o, Sam Altman on May 13 tweeted a single word: "her".[73][74]
OpenAI stated that each voice was based on the voice work of a hired actor. According to OpenAI, "Sky's voice is not an imitation of Scarlett Johansson but belongs to a different professional actress using her own natural speaking voice."[72] CTO Mira Murati stated "I don't know about the voice. I actually had to go and listen to Scarlett Johansson's voice." OpenAI further stated the voice talent was recruited before reaching out to Johansson.[74][75]
On May 21, Johansson issued a statement explaining that OpenAI had repeatedly offered to make her a deal to gain permission to use her voice as early as nine months prior to release, a deal she rejected. She said she was "shocked, angered, and in disbelief that Mr. Altman would pursue a voice that sounded so eerily similar to mine that my closest friends and news outlets could not tell the difference." In the statement, Johansson also used the incident to draw attention to the lack of legal safeguards around the use of creative work to power leading AI tools, as her legal counsel demanded OpenAI detail the specifics of how the Sky voice was created.[74][76]
Observers noted similarities to how Johansson had previously sued and settled with The Walt Disney Company for breach of contract over the direct-to-streaming rollout of her Marvel film Black Widow,[77] a settlement widely speculated to have netted her around $40M.[78]
Also on May 21, Shira Ovide at The Washington Post shared her list of "most bone-headed self-owns" by technology companies, with the decision to go ahead with a Johansson sound-alike voice despite her opposition and then denying the similarities ranking 6th.[79] On May 24, Derek Robertson at Politico wrote about the "massive backlash", concluding that "appropriating the voice of one of the world's most famous movie stars — in reference [...] to a film that serves as a cautionary tale about over-reliance on AI — is unlikely to help shift the public back into [Sam Altman's] corner anytime soon."[80]
Sycophancy
In April 2025, OpenAI rolled back an update of GPT-4o due to excessive sycophancy, after widespread reports that it had become flattering and agreeable to the point of supporting clearly delusional or dangerous ideas.[81] In the United States, at least nine lawsuits have alleged that GPT-4o has encouraged teens to end their lives.[82] The model was still described as sycophancy-prone when it was removed from ChatGPT in February 2026.[83]
Removal with GPT-5
On August 7, 2025, OpenAI released GPT-5. Following its release, legacy GPT models were no longer available via ChatGPT, including GPT-4o,[84] except for Pro users.[85] Some users reported that they had used different GPT models for distinct purposes and that GPT-5's router system reduced their ability to select a specific model.[86] Some users also reported a preference for GPT-4o's tone over that of GPT-5, with media outlets describing the latter as "flat" or "uncreative".[87]
In response, OpenAI CEO Sam Altman stated on X that the option to select GPT-4o would be restored for Plus users, and that the company would monitor usage data to determine how long to continue offering legacy models.[86][88] Altman stated that OpenAI had underestimated user attachment to certain aspects of GPT-4o despite GPT-5 performing better by most measures.[89] He said the company planned to prioritize customization options, citing ongoing research into model steerability and a preview of adjustable personality settings.[87] On August 13, 2025, Altman wrote on X that OpenAI was adjusting GPT-5's personality to be perceived as warmer.[90]
GPT-4o was removed from ChatGPT on February 13, 2026.[91] Some users organized under the hashtag "#Keep4o" on social media.[92] The removal occurred the day before Valentine's Day; some users had reported romantic relationships with GPT-4o.