OpenAI recently hosted its annual developer day in San Francisco, showcasing several groundbreaking innovations and updates for its community of developers. The event featured a series of exciting announcements, particularly focusing on advancements in real-time capabilities and model fine-tuning, aimed at enhancing user interaction and customisation of Artificial Intelligence models.
A major highlight of the event was the introduction of a real-time application programming interface (API) for developers. This new feature enables the exchange of spoken-language inputs and outputs during inference operations with OpenAI’s sophisticated production large language model, GPT-4o. This development is anticipated to facilitate more fluid and dynamic conversations between users and AI, opening doors to new applications in various interactive domains.
The pricing model for this cutting-edge feature reflects its complex capabilities. For traditional text input and output, OpenAI charges $5 and $20 per million tokens, respectively. However, voice tokens are costlier, priced at $100 per million audio input tokens and $200 per million audio output tokens. OpenAI provides an overview of these costs, which equates to about $0.06 per minute for audio input and $0.24 per minute for audio output. This cost structure indicates the premium nature of real-time interactions compared to standard API features.
To mitigate these expenses, OpenAI introduced prompt caching, allowing developers to reuse tokens on previously submitted inputs, effectively reducing the price of GPT-4o input text tokens by half. This cost-efficiency measure aids developers in managing their expenditures while leveraging real-time API possibilities.
Another significant announcement was the enhancement of OpenAI’s fine-tuning capabilities through LLM "distillation." This process allows developers to utilise data from larger models to train smaller variants. By capturing the input and output of advanced models like GPT-4o and using these "stored completions" as training data, developers can more easily fine-tune smaller models, such as the GPT-4o mini. This service aims to alleviate the traditionally cumbersome and error-prone processes associated with model distillation.
Furthermore, OpenAI expanded its fine-tuning services by incorporating image fine-tuning. Developers are now able to submit image datasets for refinement, allowing models like GPT-4o to tailor their output to specific tasks or domains. This addition exemplifies how AI can be applied to concrete industry challenges, highlighted by food delivery service Grab’s use of the service to enhance their mapping capabilities. By employing real-world images of street signs, Grab achieved significant accuracy improvements in lane count and speed limit sign localisation, optimising their route mapping operations.
The pricing for image fine-tuning is consistent with standard fine-tuning processes, with developers incurring $3.75 per million input tokens and $15 per million output tokens. When it comes to training new image models, the cost is set at $25 per million tokens, ensuring developers have a clear understanding of usage costs.
Overall, OpenAI's developer day underscored the company’s continuous innovation in AI technologies, offering developers novel tools and methods to refine and enhance their applications. By merging cutting-edge features with user-centric pricing strategies, OpenAI is poised to advance further the field of AI interactions and custom model development.
Source: Noah Wire Services