Austrian Firm Mostly AI Launches Synthetic Text Functionality to Enhance AI Training
In a significant development for artificial intelligence (AI) application in businesses, Austria-based Mostly AI has launched a synthetic text functionality geared towards generating privacy-protected synthetic data. This innovation is set to revolutionise how enterprises train large language models (LLMs) using proprietary customer data.
The growing reliance on generative AI (gen AI) has highlighted the importance of training data, which is often rich in personally identifiable information (PII). This makes it a potential privacy risk. While public data sources have traditionally fed AI models like ChatGPT, there is an increasing need for data tailored to specific business requirements. Mostly AI aims to address these needs by offering a synthetic data solution.
Announced on Tuesday, the new synthetic text functionality allows businesses to upload their proprietary datasets to Mostly AI's platform. These datasets can include emails, support transcripts, and chatbot exchanges—all crucial for understanding customer interactions. The platform then generates synthetic data that preserves the original patterns and diversity without the associated privacy risks.
How It Works
To utilise this technology, companies can upload their proprietary datasets to the Mostly AI platform. These datasets are processed through privacy-protected reusable bundles containing metadata of the original data. After confirming the correct configurations and encoding types, users select from various AI models integrated into the platform, including options from HuggingFace.
The result is a synthetic dataset that matches the statistical properties of the original data, without containing actual PII. The generated data, therefore, maintains the utility of the original data for training LLMs but complies with privacy regulations such as GDPR and CCPA.
According to Mostly AI, the synthetic data not only preserves the statistical fidelity of the original dataset but also augments it by ensuring more diversity. This can be particularly useful for rebalancing datasets to remove biases or generate mock data for software testing. Tobias Hann, CEO of Mostly AI, emphasises that synthetic data offers enhanced quality and potential compared to public data sources, which are increasingly yielding diminishing returns.
Evolving AI Data Requirements
Synthetic data's role becomes more critical as models hit a training plateau due to limited availability of high-quality public data. Mostly AI’s solution aims to circumvent these limitations by making synthetic data as viable and efficient as its real counterpart.
Founded to generate structured synthetic data, Mostly AI has evolved its offerings to include text data, which is often riddled with PII and diversity gaps. This new feature empowers enterprises to clean and utilise their own datasets more effectively, thereby facilitating better model training and innovation.
Performance and Applications
The company claims that training a text classifier on its synthetic data showed a 35% improvement in performance compared to data generated by prompting GPT-4o-mini. Mostly AI generates synthetic text by fine-tuning models with the original structured and text data to ensure the generated text is contextually accurate and useful.
While direct benchmarks comparing Mostly AI’s synthetic text generator to other providers such as Gretel are awaited, the company asserts it has consistently demonstrated superior performance in past comparisons in terms of data quality and privacy.
Future Prospects
As recently highlighted in a Gartner report, the potential of synthetic data in software engineering remains largely untapped but is expected to grow significantly. By 2026, 75% of companies are projected to use generative AI to create synthetic data, a drastic increase from less than 5% today.
Mostly AI’s synthetic text functionality opens new avenues for businesses to leverage AI while adhering to rigorous privacy standards. This development is especially pertinent as enterprises seek more robust and secure methods for AI training, ensuring that the proprietary data they rely on can be used effectively without compromising privacy.
The innovation marks a significant stride forward in the AI sector, particularly for businesses navigating the complex landscape of data privacy and model training. As Mostly AI continues to refine and expand its offerings, the implications for enterprise AI applications and data management could be profound.
Source: Noah Wire Services