Hello GPT-4o
OpenAI Blog ● Covered by 3 sources
OpenAI released GPT-4o, a multimodal model capable of processing audio, vision, and text simultaneously in real time. The company positioned it as their new flagship model without disclosing specific performance benchmarks or pricing details in the announcement. This enables developers and users to build applications that integrate multiple input types without separate conversion steps between modalities.
Why it matters
We’re announcing GPT-4 Omni, our new flagship model which can reason across audio, vision, and text in real time.