How to Use EMMO to Build a Materials Domain Ontology

“Anyone building a materials ontology will eventually encounter EMMO.”

Researchers who set out to build an ontology in their own field will eventually arrive at the idea of a top-level ontology while reviewing relevant technologies and examples. In materials science, that path almost inevitably leads to EMMO—the Elementary Multiperspective Material Ontology. Rather than treating materials, processes, properties, measurements, models, and data as isolated entries, EMMO seeks to describe how they exist and connect in the real world through a shared semantic framework.

EMMO does not regard what we recognize as an object as a fixed three-dimensional thing. It views it as a four-dimensional spacetime entity that comes into being, changes over time, and is causally connected to other entities. A material is therefore understood not only through a material name or chemical formula, but also through multiple perspectives involving its physical composition, part–whole relations, manufacturing and transformation processes, measurements, and models.

This philosophy makes it possible to describe the materials world more rigorously, but it can also feel highly abstract and unfamiliar to researchers encountering EMMO for the first time. Opening the documentation often brings them face to face with spacetime, causality, part–whole relations, and the relationship between signs and their referents before they reach familiar materials classifications or property terms. The large number of classes and relations distributed across modules, hierarchies and axioms organized by multiple perspectives, and definitions and IRIs that have changed across versions can make it difficult to know where to begin. If researchers assume that they must first understand the entire philosophy and structure of EMMO before building a materials domain ontology, they may become exhausted before the actual work begins.

But do materials researchers really need to understand all of EMMO from the outset? Views on ontology development differ, but a practical starting point is to examine, step by step, how familiar concepts of materials, processes, structures, properties, and performance may connect to EMMO. This article considers why EMMO feels difficult from the perspective of materials researchers and offers a practical approach to deciding what to examine first and how to map and align familiar materials knowledge structures.

What Does EMMO Provide?

Technically, EMMO is an OWL-based ontology developed to describe entities studied in physics, chemistry, and materials science in a common form and to support semantic interoperability among materials, processes, properties, measurements, models, data, and information. It provides fundamental concepts, relations, and axioms that ontologies from different materials domains can share.

EMMO is not a finished dictionary containing every specialized term used in every materials field. Instead, it provides a semantic framework for distinguishing and connecting foundational concepts that recur across domains: physical entities and their structure, objects and processes, parts and wholes, properties and physical quantities, measurement and modeling, workflows, data, and information. For example, it allows a real sample being measured, the measurement process, the numerical data produced as a result, and the property represented by that data to be distinguished as separate entities while remaining connected within one knowledge structure.

One of the most important features of EMMO is that it does not place an entity along only one axis of classification; it enables the same entity to be viewed from multiple perspectives. An electrode sample can simultaneously be a physical entity made of matter, the result of a manufacturing process, an object participating in a measurement, and the referent of research data. By allowing the same entity to take different roles from physical, processual, and semiotic perspectives, EMMO represents the complex reality of materials research more fully instead of forcing it into a flat classification table.

Using EMMO to build a materials domain ontology therefore does not mean importing every class it contains. The central task is to compare the materials, processes, structures, properties, and performance concepts used in a field with EMMO’s common framework, and then distinguish which concepts should be reused directly, which should be specialized, and which should be semantically aligned. EMMO is less an answer key that defines domain knowledge on the researcher’s behalf than a shared coordinate system that helps different materials knowledge systems describe the same real world consistently.

This Is the Key Point Many of Us Can Relate To: Why Does EMMO Feel Difficult?

EMMO is difficult not simply because it contains many classes and relations. It does not begin with the terminology and classifications familiar to materials researchers; it first asks, at a more fundamental level, what an entity is and how it exists in the real world. EMMO’s four-dimensional and multiperspective approach, together with structural changes across versions, can therefore leave first-time readers unsure where to begin.

A Top-Level Ontology Starts from a Different Place Than a Materials Glossary

A materials domain ontology may define the materials, processes, and properties of a specific field such as batteries, catalysts, or polymers. A top-level ontology instead defines the most fundamental categories and relations that can be shared across many fields. EMMO therefore does not immediately present the specialized terms a materials researcher is looking for; it first asks what those terms mean in the real world.

Suppose a researcher wants to identify different types of heat treatment. A materials researcher will usually think first of specific processes such as sintering or annealing and try to classify them. A top-level ontology such as EMMO takes a broader view: What kind of process is heat treatment? What participates in it? How should the samples before and after the process be distinguished? How are the overall process and its subprocesses related?

The researcher must therefore understand abstract concepts such as entities, processes, parts and wholes, and participation before reaching the materials terms of immediate interest. This difference in starting point is one of the main reasons EMMO feels difficult when it is approached as though it were a materials glossary.

EMMO Views Materials in Four-Dimensional Spacetime and from Multiple Perspectives

This abstraction becomes even clearer in the way EMMO views the physical world. It does not treat what we call an object solely as a fixed three-dimensional thing at a single moment. Instead, it views that object as a four-dimensional spacetime entity that exists and changes across space and time.

An electrode sample, for example, does not remain in exactly the same state after it is produced. It changes through heat treatment or coating, participates in measurement processes, and may undergo structural change or performance degradation during use. In EMMO, the samples before and after manufacturing, the manufacturing process, the measurement process, and their results need not be treated merely as columns in a single data row. They can be represented as entities and processes connected over time.

EMMO also does not confine an entity to one classification axis. The same electrode may be a physical entity composed of matter, the result of a manufacturing process, the object of a measurement, and the referent of specific data. Its meaning and role depend on the perspective from which it is considered.

This multiperspective approach makes it possible to represent complex materials research more faithfully. For researchers accustomed to a conventional tree-shaped taxonomy, however, it can appear as though the same entity repeatedly occurs in several modules and relations. Even after the philosophical perspective becomes clearer, further difficulties arise when researchers begin consulting actual EMMO resources.

Version Changes and Outdated Resources Add Complexity

During its stabilization, EMMO revised its module structure, class definitions, relations, and IRIs several times. Search engines may display older beta or release-candidate documentation alongside current material. Consequently, the same label may appear in different locations, or a class and IRI used in an older example may no longer be easy to find in the current ontology.

EMMO 1.0.0 was the first official release to stabilize the framework from the Top Level through the Reference Level. Version 1.0.1 then corrected issues involving IRI prefixes and catalog files, while later releases continued to refine class and relation definitions and the module structure. A project should therefore pin an exact version—such as 1.0.0, 1.0.1, 1.0.2, or 1.0.3—rather than stating only that it uses “1.0.x.”

Labels or IRIs found in web documents should not be copied directly without verification. First select the EMMO version adopted by the project, and then check the corresponding OWL files, catalog files, and official release information together. An IRI is the unique address by which a system identifies an ontology element, whereas a label is a human-readable name. Similar or identical labels do not by themselves prove that two classes are the same.

The difficulty of EMMO therefore cannot be explained only by its abstract philosophy. A top-level ontology and a materials domain ontology begin from different places; EMMO views an entity across time and from multiple perspectives; and its documentation and structure can differ by version. Rather than trying to understand all of EMMO at once, it is more practical to begin by distinguishing the basic terms that repeatedly appear in its documentation: modules, classes, relations, individuals, entities, and IRIs.

Start by Distinguishing the Terms: Modules, Classes, Entities, and IRIs

When first reading the official EMMO documentation, researchers encounter terms that refer to different levels of the ontology all at once. This can quickly raise questions such as: Is this name a file, a classification category, or actual data? A useful first step is therefore to distinguish the major components of EMMO and understand the basic meaning of each term.

  • Module: A unit of ontology organization that groups related classes and relations by topic.
  • Class: A type whose members satisfy shared conditions. Categories such as Process and Property are classes.
  • Relation/Property: A semantic connection between entities or classes, including relations of parthood, participation, attribution, and signification.
  • Individual: A particular entity registered as actual data, such as a specific sample, heat-treatment experiment, or measurement.
  • Entity: A broad term that may refer to classes, relations, or individuals. In a concrete design document, it is better to specify which kind of entity is intended.
  • IRI: A globally unique identifier for an ontology element. It should be managed separately from the human-readable label.

Distinguishing these basic terms does not mean that every EMMO module and entity must be examined from the beginning. Start by defining the research question or data-use objective, and narrow the scope to concepts directly required by that question. If the goal is to connect a materials sample with its manufacturing process, for example, first examine the major classes and relations involving physical entities, materials, processes, participants, and outcomes. Concepts involving structure, properties and physical quantities, measurements, units, models and simulations, data, and information can then be added step by step. Beginning with a familiar materials-research use case is a practical way to understand EMMO’s large structure.

Mapping and Alignment: Similar Terms with Different Practical Emphases

The terms mapping and alignment are sometimes used with overlapping meanings in the literature and in software tools. Some sources use them almost interchangeably, while others describe mapping as one of the detailed tasks performed during alignment. To make the materials ontology workflow easier to understand, this article uses mapping to mean finding and comparing candidate correspondences, and alignment to mean expressing validated semantic relations in the ontology.

Mapping: Finding EMMO Candidates for Local Concepts

Mapping investigates and records which EMMO concepts may correspond to concepts used in a local data model or PSPP classification. A local concept may be an institution-specific class, but it may also be a database table or column name, a tag, a classification item, or a node type in an existing knowledge graph.

If a local schema contains an item called AnnealingProcess, first determine whether it is simply a process term, an individual heat-treatment process that was actually performed, or a class covering multiple kinds of annealing process. Then follow the relevant EMMO concepts for processes and manufacturing processes to identify the closest semantic candidates. Searching for a similarly named class is not enough; the official definition, parent classes, relations, and scope must also be compared.

The result of mapping is therefore closer to a reviewable candidate list than to a single correct answer. A mapping table may record the local term, candidate IRI, official label, definition, parent class, adopted EMMO version, rationale for the correspondence, and confidence level. When no candidate is an exact match or several candidates remain possible, that uncertainty and the basis for judgment should also be documented. Such records allow domain experts and ontology engineers to evaluate candidates using the same evidence.

Alignment: Expressing Validated Semantic Relations in the Ontology

Alignment uses the candidates identified through mapping to determine the semantic relation between a local concept and an EMMO concept and, when appropriate, to state that relation in the ontology. Mapping asks, “Which EMMO concept could this connect to?” Alignment asks, “Through which relation should these two concepts be connected?”

If a local AnnealingProcess is a kind of general manufacturing process defined in EMMO, the local class may be placed under the relevant EMMO class. This does not mean that every manufacturing process is annealing. It means that every annealing process defined in the local ontology belongs to that broader manufacturing-process category. The approach preserves the specialized local concept while positioning it within EMMO’s shared semantic framework.

When two concepts have been sufficiently verified to have exactly the same meaning and scope, owl:equivalentClass may be considered. It is risky, however, to assert class equivalence merely because names or descriptions look similar. owl:equivalentClass is a strong axiom stating that all individuals of the two classes belong to the same category. Used incorrectly, it can cause a reasoner to classify individuals unexpectedly or reveal logical inconsistencies when combined with other axioms.

When meanings are similar but cannot be shown to be identical, a more conservative approach is appropriate. The local class may be placed under an EMMO class, or the correspondence may be recorded in an annotation or mapping table pending further validation. Keeping a correspondence as a candidate rather than prematurely asserting a strong axiom is itself an important alignment decision.

Alignment is not limited to classes. The meanings of relations must also be compared: participation between a process and a sample, the relation between a property and its bearer, and the relation between a measurement result and a unit. Even when class hierarchies look similar, differences in relation direction or domain and range can cause actual data to be interpreted differently. Alignment results should therefore be tested with representative materials data and competency questions, and an OWL reasoner should be used to check for unexpected classification or logical inconsistency.

In short, mapping is the exploratory step of identifying potential EMMO concepts and documenting the evidence, while alignment is the design step of selecting appropriate relations so that two semantic systems can be used together. This distinction makes it possible to decide systematically whether an EMMO concept should be reused directly, a local concept should be aligned with EMMO, or a new domain concept should be specialized from an EMMO parent.

Three Ways to Use EMMO

The results of mapping and alignment can be incorporated into a materials domain ontology in three main ways: reuse, alignment, and specialization. These approaches are not mutually exclusive. Within a single domain ontology, shared concepts may be reused directly from EMMO, established institutional terms may be aligned with EMMO, and specialized concepts absent from EMMO may be defined as extensions. The appropriate method should be selected concept by concept after examining both its meaning and the existing data structure.

Reuse

When an EMMO class or relation exactly matches the intended domain meaning, its IRI can be used directly. In practice, the required EMMO modules are imported into the ontology, and EMMO classes and relations are used in local axioms or to define the types and relations of actual data.

If EMMO already provides suitable concepts for a common process, physical quantity, measurement, or information entity, there is no need to create another local class with the same meaning. Direct reuse reduces duplicate definitions and improves semantic interoperability with other ontologies and datasets that use the same EMMO concept. Before reuse, however, check more than the class label: confirm the official definition, parent hierarchy, relations, and the adopted EMMO version.

Alignment

When terms and identifiers from an existing system must be preserved, or when a local concept does not have exactly the same meaning as an EMMO concept, the local class can remain separately defined and be connected to an appropriate EMMO parent or corresponding concept. This allows an established data model to retain its identity while creating a point of contact with an external semantic framework.

For example, if a research team has long used a class and IRI named AnnealingProcess, replacing every occurrence with an EMMO IRI may be unnecessary. The existing class can be retained and placed under an appropriate EMMO process class. Class equivalence should be considered only when the meanings are verified to be identical; otherwise, a subclass relation or mapping annotation is safer. Alignment preserves continuity in an existing system while improving the potential for search, querying, and data integration across ontologies.

Specialization

When a materials field requires a concept more specific than those provided by EMMO, a domain-specific subclass can be defined under a general EMMO concept. Examples include a particular electrochemical measurement, catalyst manufacturing process, thin-film sample, or battery degradation process.

Specialization involves more than adding a new class name. Relations and axioms should also define which samples and processes the domain concept connects to, which conditions and components it has, and which properties or measurement results it concerns. Where necessary, validation rules can also specify conditions, units, and required metadata reviewed by domain experts.

The EMMO-based Electrochemistry Domain Ontology (ECHO) is a practical example of this specialization approach. ECHO builds on EMMO’s general semantic framework for physical entities, processes, physical quantities, and measurement, and develops the specialized concepts and relations required in electrochemistry. It shows how a domain ontology can be constructed on a consistent top-level framework even when EMMO does not directly provide every specialized term.

Reuse applies an existing EMMO concept directly, alignment connects an established local concept to EMMO, and specialization defines a new domain concept on the basis of EMMO. In practice, combining all three approaches is often the best way to preserve domain expertise while maintaining a connection to EMMO’s shared semantic framework.

This Is Only the Beginning

Using EMMO does not mean understanding or importing every module and class. It means starting from the questions that materials researchers need to answer and the data structures they already use, determining what each concept refers to, and connecting it step by step to EMMO’s shared semantic framework. Rather than establishing correspondences from familiar names or labels alone, examine the official definition, parent hierarchy, relations, IRI, and adopted version together. Record the results in a mapping table, and select the appropriate approach—reuse, alignment, or specialization—for each concept.

There is no need to cover every field of materials science from the beginning. Start with one familiar use case, such as connecting a materials sample with its manufacturing process, define the necessary classes and relations, and use representative data to test whether queries and interpretations behave as intended. The scope can then expand to measurements and properties, units and provenance, constraints, reasoning, and governance. When domain experts and ontology engineers validate and extend these small cases together, EMMO can move beyond an abstract philosophical framework and become a practical foundation for connecting the meaning of materials data and improving its interoperability and trustworthiness.

Leave a Reply