R&D subventions
Environmental Descriptive and Learning-Based Analytics (ELBA)
Reference:
PID2025-168427NB-C21
Financial institution:
Ministerio de Ciencia, Innovación e Universidades. Convocatoria 2025 Generación de Conocimiento. Cofinanciado UE
Budget amount:
119,875
Euros
Number of researchers:
13
Duration:
Sep 1, 2026 to Feb 28, 2029
Main researchers:
Involved researchers:
Description:
Environmental data are generated mainly through observation and modelling to estimate past, present, and future states of variables of interest. Along with the measurements, provenance and contextual metadata are recorded, with spatial and temporal aspects being essential for their interpretation. The estimates are obtained through sampling over geospatial, vertical, and temporal dimensions. The extent of the phenomena and the resolutions employed determine the volume produced, which continues to grow rapidly thanks to advances in Earth observation and the evolution of storage and processing infrastructures. In addition to size, environmental data show high variability in formats and semantics, as a result of the diversity of sampling strategies and of model outputs. The datasets can take forms such as time series at stations, vertical profiles, transects, trajectories of mobile platforms, static remote-observation grids, and multidimensional grids (2D, 3D, 4D) generated by models. Descriptive and advanced analytics over this diversity present key challenges: harmonising vocabularies, units, ontologies, and geodetic references to achieve interoperability; defining storage and indexing mechanisms capable of handling multiresolution structures and mixed dimensions (dense and sparse); implementing interfaces and services for exploration with low latency, multiscale navigation, and progressive aggregation; and adapting learning methods especially deep learning to data that are heterogeneous in format, semantics, and resolution, handling gaps, noise, uncertainty, and spatiotemporal misalignments. Current solutions only partially address these requirements, often resulting in limited scalability, lack of interoperability, or restricted analytical expressiveness. Based on this diagnosis, the objectives of the project are specified in four main lines: - Develop methods and models that integrate heterogeneous data and metadata under a common semantic framework, with scalable graphs and reasoning algorithms that enable rich queries, inferences, and automatic alignment. - Development of compact data structures for efficient storage and querying of environmental data with multiple levels of detail. These structures will exploit regularities and redundancies in the data present between different resolutions. - Development of spatial, temporal, and statistical data structures that are multidimensional and multiresolution and that support interactive exploration of environmental data, which may present dense and sparse dimensions. The objective will be to achieve low-latency data access during multiresolution navigation and progressive aggregation. These structures will later be redesigned using compact representations to further reduce storage footprints and response times. - Development of efficient training strategies for deep-learning models applied to heterogeneous environmental data that present dense and sparse dimensions. Multimodal architectures will be developed that combine different types of models to address tasks that integrate different types of data and configurations of several resolutions. The project will also explore neural models that operate directly on compact representations, including compact encodings of their own weight matrices to improve efficiency and scalability.





