Gladia Secures $16 Million in Series A Funding to Pioneer Multilingual AI Transcription and Audio Intelligence
Gladia, a burgeoning player in the AI transcription and audio intelligence sector, has successfully secured $16 million in a Series A funding round. The investment will be channelled into developing an advanced end-to-end audio infrastructure, spearheading the launch of a new real-time engine for audio transcription and analytics. This innovative engine is designed to enhance the capabilities of voice-first platforms by allowing them to deliver more substantial value to users on an international scale through AI technology.
Founded in 2022, Gladia has quickly established itself within the realm of AI-driven audio solutions, now amassing a total funding of $20.3 million. The company's inception was driven by CEO and co-founder Jean-Louis Quéguiner's personal frustration with existing audio transcription services, which were unable to efficiently comprehend his French accent. Quéguiner highlighted the challenge faced by their international team and clientele who frequently switch languages during meetings, underscoring the need for a transcription service capable of handling multiple languages and accents simultaneously.
Most current speech recognition models are predominantly trained on English audio, which can lead to inherent biases. In contrast, Gladia has focused on developing a real-time product that stands out as 'truly multilingual'. Their newly refined engine boasts real-time transcription capabilities in over 100 languages and offers improved support for various accents, with the adaptability to switch languages instantaneously. This engine is particularly distinctive in its ability to derive insights from conversations, such as caller sentiment, pivotal information, and summarised content, all in real time. The process is remarkably swift, generating both transcript and insights in less than a second.
Creating an accurate, low-latency, multilingual engine is inherently complex and demands substantial expertise in language processing, real-time data management, continuous refinement, and maintenance. Real-time models necessitate increased computing power and can face challenges in providing immediate accurate output due to limited context. Nevertheless, Gladia's real-time speech-to-text engine claims a sub-300 millisecond speed without sacrificing accuracy, irrespective of language, geographical factors, or technological frameworks.
Gladia's foray into the market began with the launch of its first asynchronous transcription and audio intelligence API in June 2023, which was built on a proprietary variation of Whisper ASR. The API quickly gained popularity within the enterprise sector, particularly among meeting recorders and note-taking assistants, and has since attracted a user base exceeding 70,000, alongside adoption by over 600 global customers, including Attention, Circleback, Method Financial, Recall, Sana, and VEED.IO.
Looking ahead, Gladia intends to utilise the new investment to further its research and development initiatives and soon release a comprehensive AI toolkit for audio. The company also aims to enrich its product offering with additional selective models, encompassing large language models (LLMs) and retrieval-augmented generation (RAG). With several design partners within the contact-center-as-a-service (CCaaS) sector, Gladia is currently trialling an agent-assist solution powered by its real-time AI engine. Furthermore, the company plans to expand its workforce as it gears up for international growth.
Gladia's latest developments mark a significant milestone in the field of AI transcription and audio intelligence, as it continues to innovate and set new standards for multilingual and real-time audio processing. As the company progresses, its trajectory and potential partnerships will be closely monitored by industry observers.
Source: Noah Wire Services