A Comprehensive Guide to Metadata in Materials Science Research

When reflecting on the specialized nature and complexity of materials science data, a critical question often arises: “Is our data truly ‘AI-friendly’ enough for effective learning within this domain?”

While many research fields focus heavily on securing a high quantity of data, they often overlook the holistic context—the metadata—such as the environment, equipment, and intent behind the data generation. However, data that is not contextualized through metadata is nothing more than a meaningless sequence of numbers to an AI. This inevitably leads to “Garbage In, Garbage Out” scenarios, where low-quality inputs yield unreliable results.

In this post, we aim to re-examine our perspective on metadata through the lens of NOMAD (Novel Materials Discovery), one of the world’s leading projects for materials data sharing. We will explore why metadata should be treated not merely as supplementary information, but as a core functional element that determines the success of data-driven materials research.

If you wish to maximize research efficiency and uncover the true value of materials data, it is time to turn your attention to “the data behind the data”—metadata.

Redefining the Concept and Role of Metadata

Generally, metadata is defined as “data about data” or “attribute information.” We can find common examples of this in our daily lives, such as library catalog cards or information embedded in digital photo files.

If a book itself is the data, the catalog card recording its author, publication date, and genre is the metadata. Similarly, if a photograph is the data, the EXIF information containing the shooting location, time, and camera settings serves as the metadata. In essence, metadata acts as a guide that helps us easily find, manage, and understand the meaning of specific data within a vast sea of information.

However, in the field of materials science—where understanding complex correlations and the overall context of research data is paramount—metadata must be redefined from a much more rigorous and strategic perspective. According to the philosophy of NOMAD (Novel Materials Discovery), a single dataset can be clearly divided into two sub-elements: ‘data’ and ‘metadata.’ What is particularly noteworthy here is that NOMAD defines the scope of what we commonly call “data” in an extremely restrictive manner.

From NOMAD’s perspective, “data” in its truest sense refers only to the final objective a researcher seeks to obtain: the ‘Target’ data. Every other item within the dataset—whether numerical, textual, or in any other form—is considered metadata. This strict distinction is intended to reveal the intrinsic value that metadata holds within the materials domain.

The essence of metadata in materials research lies not merely in listing information, but in its “explanatory power”—its ability to serve as a descriptor that logically supports the background and causes of the measured results. While the target data shows what was obtained, metadata serves as the collection of contexts that proves why those results occurred.

What is the Target and What is the Context? Understanding Through a Solar Cell Measurement Case

This concept becomes even clearer when applied to the case of solar cell performance measurement. In a dataset obtained through experiments, the key performance indicator, ‘Power Conversion Efficiency (PCE),’ is essentially the target data we aimed to reach.

On the other hand, all other items—such as ambient temperature and humidity during the experiment, the chemical composition of the perovskite used, the thickness of the thin film, and the status of the measurement equipment—serve as metadata. These elements explain why the measured PCE was recorded at that specific value, providing the necessary context to validate the result.

The Core Driver of AI Models: Metadata as a Descriptor and Knowledge System

The fact that metadata possesses such strong “explanatory power” means that, from the perspective of modern machine learning and AI modeling—the heart of contemporary scientific research—metadata functions as a Descriptor. The performance of a predictive model depends less on the target data itself and more on the numerous variables that determine the outcome; in other words, it depends on how rich and sophisticated the metadata is.

Translating this into the “grammar” of machine learning makes the relationship even clearer. The researcher’s objective (the data) becomes the ‘Target’ or ‘Label’ that the model aims to predict, while the metadata acts as the input, serving as the ‘Feature’ that drives the learning process. Ultimately, developing a high-performance materials prediction AI model is a matter of how systematically one can design and acquire high-quality features—the metadata that best explains the properties of the material.

From this perspective, building big data in the materials field is less about simply accumulating target data and more about the effort to discover and store meaningful metadata. While the data itself is a natural byproduct of experiments or calculations, the process of standardizing and recording the metadata—the context that makes data valuable—requires a far more sophisticated strategy.

Therefore, a true materials big data infrastructure needs to be defined as a repository of rich metadata rather than just a massive data storage facility. Such a well-constructed metadata repository serves as a “treasure trove” where numerous researchers can recombine and analyze data from various perspectives according to their specific research objectives, far beyond a single initial purpose.

This is also closely linked to the FAIR principles (Findable, Accessible, Interoperable, Reusable), the global standard for maximizing data value. An environment where data is easily searchable (Findable), reachable (Accessible), compatible across different systems (Interoperable), and ultimately capable of being repurposed (Reusable) can only be realized upon a foundation of a standardized metadata system.

Furthermore, metadata evolves beyond simple records into a knowledge system through its connection with Ontology, which defines the logical relationships between data. When metadata is organically linked via ontology, AI can gain a deeper understanding of the hidden correlations between data. This, in turn, will serve as a powerful foundation for discovering innovative new materials that humanity has yet to find.

Given that metadata functions as a Feature and Descriptor—the core of machine learning—who creates this high-quality metadata, and how? This is a crucial point, and it is here that we must pay attention to the fundamental role of the Materials Domain Expert.

Materials Domain Experts: The Ultimate Architects of Metadata Value

The first step in building a reliable AI model is securing meticulously refined data. No matter how superior an algorithm is, if the input data is inaccurate or lacks context, the results cannot be trusted. To maximize a model’s predictive power, Data Labeling—assigning precise meaning to each piece of data—is essential. This process is, in fact, an intellectual endeavor that demands near-infinite patience and effort.

In response to the growing importance of metadata, the AI industry has recently utilized various data labeling outsourcing platforms to efficiently handle vast amounts of labeling work. Platforms such as Crowdworks, Labelbox, and Outlier provide environments where a large number of workers can participate to process large-scale data quickly. For building general AI training datasets—such as identifying objects in images or converting speech to text—this crowdsourcing strategy can be highly effective.

However, when we enter the specialized field of materials science, the situation changes completely. Tasks such as identifying microstructures in microscopic images, finding defects in complex crystal structures, or structuring the causal relationships of experimental conditions into metadata are beyond the capabilities of general workers. These are highly intellectual tasks that can only be performed by domain experts with deep insight into the field.

Ultimately, the process of designing and labeling metadata is no longer the exclusive domain of computer scientists. Metadata gains its true vitality as a descriptor only when materials experts—who understand the physical and chemical properties of specific materials and can judge which variables critically affect the results—directly participate.

Sophisticated metadata defined by domain experts serves as a milestone, guiding AI models toward the correct learning path. Therefore, future materials research must evolve beyond researchers simply producing data in laboratories; they must project their expertise onto the data in the form of metadata to communicate with AI. This collaboration, combining professional expertise with data science, will be the most powerful key to shortening the “golden time” for new material development.

When a robust data ecosystem is established through the combination of meticulously designed metadata and the insights of domain experts, materials research will evolve into an entirely different dimension.

Capturing Flowing Data: Time-Series Metadata and Future Prediction

From the perspective that metadata can create new value as data that “flows” over time—rather than just being static records—let us assume a scenario where materials research is viewed as a continuous process rather than a one-off event. In this context, metadata must also be managed from a time-series perspective.

Elasticsearch and Kibana serve as incredibly powerful tools for managing and analyzing such time-series data. By utilizing Elasticsearch’s Snapshot feature, the state of data at a specific point in time can be perfectly preserved and revisited whenever necessary. Furthermore, Kibana’s intuitive dashboards allow for the real-time visualization of complex time-series metadata flows.

A particularly noteworthy aspect is tracking the creation and extinction of metadata fields. As research evolves, variables that were once insignificant may emerge as new key descriptors (creation), while others may become obsolete due to technological advancements (extinction). Tracking the lifecycle of these metadata elements provides valuable analytical outputs as a “Research Lineage,” illustrating how research methodologies have evolved over time.

Rich metadata managed in a time-series format acts as both a mirror reflecting past records and a sophisticated simulator for predicting future experimental results. Researchers who master the flow of data—from its creation to its extinction—will be able to make valuable discoveries based on their unique perspectives amidst the overwhelming flood of data.

Summary: Metadata-Driven New Materials Development

In materials research, metadata is far more than just auxiliary information; it is a core descriptor that explains causal relationships and serves as an essential element determining the ultimate value of a dataset. Establishing high-quality metadata requires more than simple data processing—it demands the strategic design and active participation of domain experts who possess deep insights into the field.

When systematically designed metadata is integrated with FAIR principles and ontologies, it maximizes data reusability and evolves into a massive knowledge system optimized for AI learning. Furthermore, tracking metadata variations through a time-series lens enables the mapping of research lineage and functions as a sophisticated simulator for predicting future experimental outcomes. Ultimately, this rigorous metadata-driven approach will prove invaluable across diverse facets of innovative materials development.

Leave a Reply