Catalogue of linguistic resources

Linguistic Resources Catalogue: A federated and sovereign directory designed

 for the publication, discovery and utilisation of linguistic assets, AI models and

 advanced services within the ecosystem.

Part of the Valle de la Lengua Data Space ecosystem

The catalogue is one of the components of the Valle de la Lengua Data Space and serves as a common reference point for identifying resources available within the ecosystem. Its function is not to store data centrally, but to publish structured metadata that enables the discovery of resources distributed across institutional, academic and corporate silos.

Through this infrastructure, usage policies and participants’ digital identities are automatically managed to ensure trust in every transaction.

What resources can you find and share? 

icono

Linguistic data:

Text corpora, speech, specialist terminology and documentary resources.

icono

Models and tools:

NLP models, language technologies, APIs and related services.

icono brujula

Metadata and certification:

Information on origin, quality, conditions of use and traceability.

Access and terms of use

Participation in the catalogue is governed by the Rulebook, a single governance framework that ensures transparency and regulatory compliance (GDPR, Data Act and AI Act).

The catalogue allows for the dynamic negotiation of terms of use:

icono

Absolute sovereignty:

The provider determines who can access their data, how it is used, and for what purposes, through digital policies (ODRL).

icono

Verified identity:

Only entities with trusted credentials can initiate effective exchanges.

icono brujula

Flexible models:

Resources may be offered under subscription models, pay-per-use models or licences specifically for research purposes.

A growing ecosystem

The catalogue is growing steadily. Although it is linguistically based, the strategic roadmap envisages the integration of datasets from other sectors to enhance its value and expand its potential for innovation.

 

Today linguistic.  
Tomorrow multi-sector.  
Always interoperable.