Based on the standardization required for data utilization, large-scale data must communicate with each other to provide materials researchers with new insights. At this juncture, we encounter the concept of Interoperability. Interoperability refers to the ability of different systems or devices to exchange information and to interpret and use the exchanged information. The efficiency of data-driven materials research can differ drastically depending on whether or not this technology is well-implemented.
The Limits of Traditional Databases: Materials Research Trapped in Data Silos
Historically, the materials research field has widely used Relational Database Management Systems (RDBMS) to manage data systematically. However, as materials research data has become increasingly vast and complex, this traditional approach has begun to reveal clear limitations.
The first obstacle is structural rigidity. RDBMS requires data to be entered according to pre-defined table formats, which cannot keep up with the flexibility of research sites where experimental conditions frequently change or measurement equipment becomes more advanced. The inconvenience of having to change the database’s basic design (schema) every time a new variable arises has acted as a chronic problem hindering research efficiency.
An even more serious issue is the absence of semantic meaning within the data. Information regarding the context in which the numbers recorded in the database were measured, or specifically what they signify, is not included in the data itself. Ultimately, this core context is either recorded separately in other documents or remains only in the individual researcher’s memory. As a result, even if an external system or another researcher retrieves the data, it becomes virtually impossible to automatically understand or interpret the meaning behind it.
These technical limitations ultimately lead to a phenomenon similar to “data islands,” where data from each laboratory remains isolated and unconnected. Data becomes trapped in massive silos, and a cycle of inefficiency repeats where massive manual labor and costs must be invested to integrate and analyze them as a whole.
Strengthening Interoperability through Materials Ontology
Ontology-based knowledge systems are gaining attention as an innovative solution to overcome the limitations of traditional data management. Beyond mere data storage, ontology clearly defines the complex relationships between data and the meanings contained within them in a language that computers can understand. Through this, experimental data goes beyond simple numerical values to include contextual information—such as measurement equipment or conditions—and possesses the intelligence to be automatically interpreted by the system, as demonstrated by the following features:
- Automatic Context Interpretation: By utilizing ontology, context such as “this numerical value is a ‘bandgap’ value at a specific temperature, measured with a certain device” is included within the data itself. This allows the system to interpret the data on its own.
- Flexible Scalability: Because it is not bound by a fixed table structure, the system can expand flexibly while maintaining connectivity with existing systems, even when new research data or complex material properties are added.
- Foundation for AI Research: It provides a machine-learning-friendly environment where AI models can immediately utilize data from different origins (Provenance) without the need for separate pre-processing.
The realization of interoperability is the difference between possessing mere Raw Data or Living, Collaborating Knowledge. True interoperability is achieved when data can “talk” to each other through ontology, moving beyond the limitations of isolated data.
Interoperability Strategies: Leading from Recorded Data to Living Knowledge
While the ontology-based knowledge system described above serves as the theoretical foundation, a specific approach to deploying interoperability is required to apply it to actual research sites and produce results. Moving beyond simply collecting data, practical strategies to make data “find its own way” proceed in the following directions.
The fundamental first step toward interoperability is standardization, which involves unifying data formats and units. If researchers use different formats or inconsistent measurement units, even the most excellent data faces limitations in integrated analysis. Standardization provides the foundation where a single data point generated in a specific field can be utilized in various other fields without any separate conversion process when numerous data items are interconnected.
Meanwhile, from the perspective of data integration required in materials big data, organically connecting numerous data items opens up the possibility for a single piece of data generated in a specific sub-field to be utilized in its original form across entirely different fields. This implies that data must evolve into a universal knowledge asset that can be commonly used in various research contexts, rather than remaining an isolated fragment of information.
While these characteristics may seem similar to the existing concept of “reusability” in FAIR data principles, there are distinct differences in the details. While reusability simply means a state where past data can be manually found and processed by humans, data utilization through interoperability differs in that it aims for machine-centered intelligent connectivity. This allows the system to understand the context of the data on its own and immediately integrate it into other research processes without manual intervention. In other words, the fundamental difference between reusability and interoperability lies in whether the data itself can communicate with other systems and create value, going beyond the mere recycling of information.
Platform-Based Intelligent Integration and Advancement
The concrete means to turn this intelligent connectivity into reality is a platform-based integration strategy. Since the efforts of individual laboratories or institutions are limited in building a vast data ecosystem, it is essential to establish a common platform where data can flow freely.
In the field of materials research, the emergence of common data formats such as OPTIMADE (Open Databases Integration for Materials Design) is one of the most ideal examples of realizing interoperability. OPTIMADE is an open API standard developed to link various materials databases scattered around the world into a single common format. It allows researchers to simultaneously search and extract materials data from multiple sources using a single query language, without having to learn separate access methods for each database. As such, it possesses powerful potential to unite materials databases under one standardized language. This can maximize interoperability in materials research and accelerate the shift toward utilizing global materials knowledge as if it were one massive library.
Technically, the method of submitting data queries and receiving responses using APIs (Application Programming Interface) is a core element that accelerates interoperability. As seen in the case of OPTIMADE, once a standardized API is established, researchers can query multiple data sources simultaneously through the system and receive refined responses in real-time without complex manual tasks. This automated exchange system serves as a powerful incentive for data providers to manage their data according to standards, ultimately creating a virtuous cycle that improves the data quality and connectivity of the entire ecosystem.
Data collected from these various sources can be semantically integrated through the Semantic Web. Going beyond a simple listing of data, the Semantic Web logically connects causal relationships and physical associations between data points. This enables researchers to analyze data derived from different experiments within a single context, leading to insights that can discover unexpected new material properties.
In addition, interoperability can be materialized in various forms through platforms. Representative examples include workflows where raw data generated from experimental equipment is immediately transmitted to a cloud platform and automatically classified, or platforms equipped with visualization-friendly Knowledge Graphs. In this way, platforms serve as a stepping stone for transforming isolated data into collaborative knowledge by supporting interoperability.
Enhancing Data Value Based on Rich Metadata and Expert Insight
For the implementation of interoperability not to be limited to technical connections for exchanging data, it is necessary to be well integrated with a much more vast amount of ‘descriptive data’—that is, metadata—than the data itself.
Since the level of interoperability is proportional to the depth of metadata, Rich Metadata enhances interoperability. For example, rather than just transmitting a numerical value for a property, the entire context leading to that number must be recorded as metadata. Preservation of context means that data finally achieves Machine-readability when minute variables—such as sample purity, synthesis temperature, pressure duration, equipment model name, and noise filtering conditions—are structured as metadata. Furthermore, trust-based integration means that rich metadata provides a basis for evaluating the reliability of data coming from different laboratories. The system can analyze metadata to determine that “Experiment A and Experiment B were performed under comparable conditions,” thereby drastically increasing the accuracy of data integration.
Contribution of Metadata Enhanced by Materials Domain Experts
Ultimately, the quality of metadata is determined by domain experts who possess a deep understanding of the field. It is the expert’s role to define core materials science variables that AI or general data engineers might easily overlook.
- Selection of Core Variables: Materials experts know which core parameters among numerous experimental conditions have a decisive impact on the results. Metadata schemas reflecting their insights enable the construction of high-quality knowledge systems that penetrate the essence of research.
- Formalization of Tacit Knowledge: By elevating experimental know-how (tacit knowledge)—which previously resided only in a researcher’s mind—into a standardized format called metadata, experts provide an environment where subsequent researchers or AI models can utilize data immediately without trial and error.
- Semantic Advancement: Metadata enhanced by experts combines with ontology to enable logical inference between data. For example, by embedding scientific rules such as “this combination of elements is highly likely to have a specific crystal structure” into the metadata system, it enables ‘intelligent exploration’ that goes beyond simple data searching.
When core variables and the tacit knowledge of experiments are formalized into standardized metadata through the insights of materials experts, data gains clear contextual information that machines can immediately understand beyond simple numerical values. This sophisticated metadata serves as a common guideline that helps different research systems interpret the scientific meaning of data identically, completing the practical foundation of interoperability. Furthermore, metadata combined with ontology integrates fragmented information into an organically connected knowledge network by enabling logical inference between systems. Consequently, expert-centered metadata enhancement becomes a key driver for breaking down barriers between isolated data and fostering smooth communication between systems, establishing a truly intelligent research collaboration system.
Examples of Familiar Material Properties: Bandgap and Overpotential
Bandgap is a core property that determines the efficiency of semiconductor or solar cell materials. However, as is common in traditional databases, if only a property value like “1.1 eV” is recorded, that figure alone struggles to achieve interoperability.
- Absence of Context vs. Ontology-based Solution: In the traditional approach, it is difficult to know whether 1.1 eV was measured at room temperature, at cryogenic temperatures, or if it is a DFT calculation result rather than an experimental value. In an ontology-based system, metadata such as “Measurement Temperature: 300K,” “Optical Measurement Method (UV-Vis),” and “Sample State: Thin Film” is semantically linked to this data. Through this, the system can interpret for itself that this data is reliable experimental data suitable for calculating the efficiency of room-temperature silicon solar cells.
- Expert Insight and Interoperability: When materials experts know that doping concentration or crystal structure decisively impacts bandgap measurements, the metadata schemas enhanced by them inevitably include these variables as mandatory items. If this allows machines to determine whether bandgap data produced by different laboratories share the same crystal structure and doping conditions for integrated analysis, these values can be practically, usefully, and immediately utilized for AI training in materials discovery.
Another example is overpotential, which represents the efficiency of electrode reactions in water electrolysis or battery research. This property can be considered data that is extremely sensitive to the experimental environment.
- Overcoming Data Silos and Platform Integration: Overpotential values vary significantly depending on various measurement conditions, such as electrolyte pH, current density, and the type of working electrode. In traditional databases, these conditions were often scattered as unstructured text data, making direct comparison impossible. However, as common platforms like OPTIMADE become active and standardized APIs are utilized, researchers can submit sophisticated queries such as: “Retrieve all catalyst data where the overpotential is below 200 mV at a current density of 10 mA/cm² under alkaline conditions of pH 14.”
- Formalization of Tacit Knowledge and Intelligent Exploration: If experienced electrochemical researchers possess the know-how (tacit knowledge) that IR-compensation determines the practical usability of data, they will select and formalize this compensation status as a core metadata variable. AI models can then automatically filter out uncompensated, inaccurate data for learning. Through this semantic advancement, the system not only accounts for changes in measurement results triggered by such variables but also establishes an intelligent exploration environment capable of inferring causal relationships.
As seen in the previous cases, the true value of interoperability lies in whether data can actually explain its own context. When rich metadata infused with expert insight is combined with ontology technology, fragmented experimental figures can finally “talk” to each other and evolve into living knowledge that provides new inspiration to researchers.
Conclusion: Completing an Intelligent Research Environment through Interoperability
In this way, the implementation of interoperability in materials research is a key process for transforming fragmented “recorded data” into “living knowledge” that can communicate and collaborate on its own. An ontology-based knowledge system, rich metadata infused with materials experts’ insights, and platform-based integration strategies like OPTIMADE break down data silos and complete an intelligent research environment where artificial intelligence can learn immediately. This systematic interoperability goes beyond merely improving research efficiency; it serves as a source of insight that preserves the context between complex materials data and reveals causal relationships, as demonstrated by the cases of bandgap and overpotential. An interoperability ecosystem built on materials data through organic connections between systems will accelerate the era of data-driven discovery of new materials and become the most powerful foundation for solving scientific challenges facing humanity through materials.