The Open Source Initiative (OSI) has unveiled the inaugural version of its open source AI definition, marking a significant development for the technology industry by providing a clear benchmark for what can be classified as Open Source AI. This definition, published today, seeks to establish a standard that facilitates the verification of AI systems' openness, focusing on code, models, and particularly data information.

The announcement underscores a collaboration with Mozilla, a staunch proponent of open source technologies, to advocate for greater transparency within AI systems. Mozilla's engagement is characterised by an emphasis on understanding AI's internal workings to ensure they can be questioned, investigated, and potentially regulated in alignment with open-source principles. Ayah Bdeir, a senior strategic advisor at Mozilla focusing on AI strategy, expressed these views during an interview on the “What the Dev?” podcast by SD Times.

Bdeir highlighted the multifaceted nature of AI systems, noting that they are composites of various elements such as algorithms, code, hardware, and extensive datasets. She pointed out the complexity and opacity that can arise due to the layered data usage in AI – data for training, testing, and fine-tuning models, which she argues provides organisations with a misleading impression that simply open-sourcing the code equates to an open-source AI system.

“There is a distinct difference between AI and traditional open source software,” Bdeir explained. In traditional contexts, components such as the source code, compilers, and licences are clearly demarcated and known for their openness. She stressed that in the AI sphere, multiple interacting components necessitate clarity and honesty in claims of being open source. A truly open-source AI, she argued, should offer the four essential freedoms: to use, study, modify, and share.

The definition introduced by OSI endeavors to resolve ambiguities around what constitutes open-source AI, crafting a definitive checklist to evaluate the authenticity of open-source claims. According to Bdeir, the formulation of this definition incited debate, particularly concerning data information, which is notably controversial due to the legal intricacies and practicalities associated with disclosing complete datasets.

The debate revolves around two primary schools of thought. One insists that for an AI system to be truly open source, its training dataset must be fully open and accessible to allow replication and thorough scrutiny. This perspective posits that without access to the exact data employed, it is impossible to replicate or fully comprehend the AI's functions, thus disqualifying it as open source.

Conversely, the second camp, which includes many involved in hands-on AI development, argues that making data fully available is impractical. Data governance is subject to international legal frameworks and copyright laws, which vary significantly, complicating the clear attribution and legal distribution of datasets.

To navigate this intricate landscape, the OSI has decided against mandating the release of datasets. Instead, they require what is described as "data information," compelling organisations to divulge detailed information about the training data. This requirement is intended to enable a skilled person to recreate a system that is substantially equivalent using similar data, without directly sharing the original dataset itself.

This newly established framework by the OSI represents an earnest effort to delineate the genuine conditions of openness in AI systems, and is expected to provide a concrete, actionable guide for AI developers and industry stakeholders globally.

Source: Noah Wire Services