GPT-4o with improved text, audio and image features is here

GPT-4o (“o” for “omni”) is OpenAI’s latest multimodal Large Language Model (LLM), and brings major advances in generating text, speech, and image content to enable more natural interaction between users and AI. OpenAI claims the new AI model can respond to audio input in just 232 milliseconds, is significantly faster at text responses to non-English prompts, and supports over 50 languages. You can also interrupt the model with new questions or clarifications while it’s speaking.
GPT-4o also features a more powerful human-sounding voice assistant that responds in real time and can observe your surroundings through your device’s camera. You can even tell the assistant to sound happier or revert to a more robotic-sounding voice. You also get real-time translations in over 50 languages and it can act as an accessibility assistant for the visually impaired.
OpenAI demonstrated a long list of GPT-4o’s features in its livestream. You can watch all the demos of the new GPT-4o features on OpenAI’s YouTube channel. GPT-4o will be available for the free ChatGPT users, while ChatGPT Plus subscribers will get five times higher message limits. The new text and image features are already available in the ChatGPT app and on the web. The new language mode will be available as an alpha mode for ChatGPT Plus in the coming weeks.
In related news, OpenAI announced a ChatGPT desktop app for macOS, while a Windows version will be coming later this year. OpenAI also announced its ChatGPT store, which hosts millions of custom chat bots that users can access for free.
