Open Source Initiative Launches Open Source AI Definition Amidst Debates

RALEIGH, NC – On October 28, 2024, the Open Source Initiative (OSI) unveiled the inaugural version of the Open Source AI Definition (OSAID 1.0) at the All Things Open conference in Raleigh, North Carolina. This culmination of nearly two years of development represents a pivotal moment in the evolving landscape of artificial intelligence (AI) governance, although the definition has not come without its share of controversy.

The creation of OSAID 1.0 has been marked by vigorous debate and varying levels of acceptance among stakeholders. The OSI acknowledges the definition as a work in progress, a sentiment echoed by its chairman, attorney Carlo Piana. Piana explained the inherent challenges, stating, "Our collective understanding of what AI does and what’s required to modify language models is currently limited. The more we use it, the more we’ll understand."

The primary contention around OSAID lies in its treatment of datasets used for training AI models. This issue has divided opinion into what have been characterised as three distinct groups: pragmatists, idealists, and faux-source business leaders. Mark Collier, COO of the OpenStack Foundation, highlighted a key challenge: "Requiring all raw datasets to be public might seem logical; however, this analogy to source code starts to fall apart. Training data influences models through patterns, while source code provides explicit instructions."

As a compromise, the OSAID allows for "sufficiently detailed information about the data used to train the system" rather than mandating full transparency of the datasets. This approach aims at balancing transparency needs with practical and legal concerns such as copyright and privacy, especially regarding sensitive data like medical records.

Several organisations have endorsed the OSAID, including the Mozilla Foundation, the OpenInfra Foundation, Bloomberg Engineering, and SUSE. Alan Clark, from SUSE's CTO office, applauded these efforts, stating that the definition is crucial given the rapidly evolving AI landscape.

While some academics, like Percy Liang from Stanford University, have shown support for the OSAID, recognising it as a significant step despite data restrictions, the definition has also faced criticism. Idealists object to the inclusion of proprietary data in open-source AI models, arguing it undermines the principles of open source. Tom Callaway from AWS articulated these concerns, expressing that the definition "damages every established understanding of what 'open source' is."

The OSI is aware of these dissenting views and aims to adjust the definition as AI technologies advance. They were driven to define open-source AI as new legislation is emerging in the US and EU without clear definitions, which they feared could lead to companies creating misleading open-source claims for proprietary products.

In response to the ongoing debates, some groups are taking independent actions. Digital Public Goods (DPG) is updating its standard for AI to require open training data, with plans to release a proposal for public comment on GitHub in November.

The matter of defining open-source AI is further complicated by businesses seeking to benefit from the more lenient regulations associated with open-source systems. Companies like Meta and OpenAI might attempt to establish their own definitions to fit their needs, as illustrated by Meta's recent comments on the complexities of defining open-source in the current AI context.

While the OSAID provides a baseline that many organisations will follow, the broader discourse on what constitutes open-source AI is anticipated to continue, reflecting deeper philosophical and practical divisions within the industry. For the everyday AI user, these debates might feel distant, but for businesses and policymakers, the definition of open-source AI remains crucial for both operational and strategic reasons.

Source: Noah Wire Services