Cohere Enhances Multilingual AI Capabilities with New Aya Expanse Models
In an ambitious move to bridge linguistic gaps in artificial intelligence, Cohere, an AI-focused company, has unveiled two new models as part of its Aya project. The release includes Aya Expanse 8B and 35B models, now available on the AI platform Hugging Face, promising advancements across 23 different languages. This development is particularly significant as it continues the mission of widening accessibility to AI models capable of supporting a diverse array of languages beyond English.
The unveiling took place following a blog post by Cohere, where the company highlighted the distinct advantages of each model. The Aya Expanse 8B, credited with harbouring eight billion parameters, aims to democratise access to AI breakthroughs by equipping researchers globally with cutting-edge technology. The more robust Aya Expanse 35B model, with 35 billion parameters, offers what Cohere describes as state-of-the-art multilingual capabilities, supporting a wider expanse of linguistic needs.
The Aya project, initiated by Cohere for AI, the company's research arm, was launched last year with the vision of overcoming the language limitations inherent in many foundational models. Earlier this year, the Aya 101 large language model was introduced, a 13-billion-parameter framework that extended support to 101 languages. Alongside, the Aya dataset was released to aid the development and training of models in less commonly supported languages. The new Aya Expanse models build on these foundations, utilising methodologies akin to those employed in creating Aya 101.
Cohere has emphasised the role of sustained research focus on enhancing AI's service to global languages. Innovations forming the current model framework include data arbitrage, a method devised to counter the inefficiencies of using synthetic data, and a novel preference training mechanism that enhances performance and safety while honouring cultural and linguistic diversity.
In benchmark tests, the Aya Expanse models have reportedly outperformed their counterparts from major industry players, including Google, Mistral, and Meta. The 35B model excelled in multilingual assessments, surpassing notably large models such as Gemma 2 27B, Mistral 8x22B, and Meta’s Llama 3.1 70B. Meanwhile, the smaller 8B model also showed superior performance against similar-sized models.
Cohere's innovative data sampling technique, known as data arbitrage, is designed to mitigate the reliance on synthetic data, which often results in subpar outputs when good teacher models are unavailable for many languages, particularly those that are resource-scarce. The company has been enhancing its models by directing preferences toward what it terms "global preferences," acknowledging and integrating the various cultural and linguistic nuances present across the globe. Cohere notes the novelty and importance of applying preference training in a broad multilingual context—a practice not traditionally extended by safety protocols that tend to be Western-centric.
The Aya initiative particularly targets the often-challenging area of creating language models proficient in languages other than English, which is predominant in government, finance, and digital communications sectors. While English data is abundantly available, the challenge exists to find adequate data for other languages and accurately benchmark model performances.
The landscape of language model development continues to see contributions from multiple industry players. For instance, OpenAI recently made available its Multilingual Massive Multitask Language Understanding Dataset on Hugging Face, to foster improved evaluation of language model performance, supporting 14 different languages.
In recent weeks, Cohere has not only been active with the Aya project but has also expanded capabilities in other areas. The company integrated image search functions into Embed 3, a tool used within retrieval augmented generation systems, and enhanced fine-tuning options for its Command R 08-2024 model.
These steps taken by Cohere signify a concerted effort to not only diversify the languages supported by AI models but also to ensure that these technologies are adaptable and sensitive to the nuanced demands of various linguistic and cultural contexts around the world.
Source: Noah Wire Services