Valle de la Lengua Data Hub
A secure, federated environment for sharing, discovering and using Spanish-language data, with guarantees of data sovereignty, interoperability and traceability.
// WHAT IS IT?
What is the Valle de la Lengua Data Hub ?
The Valle de la Lengua Data Space (EDVAL) is a collaborative environment where organisations can share and use linguistic resources without relinquishing control over their information. The data is not stored in a single location: it remains where it is and is shared in accordance with the conditions set by its owners.
This model facilitates the development of language technologies, digital services and research projects related to Spanish. Furthermore, it adheres to European standards for secure data exchange, ensuring trust and protection throughout the process.
Key components
Within the ‘Key Components’ section, the federated catalogue, the San Millán Spanish Oral Corpus (COE San Millán) and the connector form the cornerstones upon which the system’s reliability and technical utility are built.
Federated catalogue
A smart directory where providers publish metadata for their resources in accordance with international standards such as DCAT-AP. The result? You can find out which language assets are available without moving or centralising the actual data, facilitating collaboration between systems, organisations and sectors, whilst maintaining control at source.
Technology hub
It acts as a guardian of data sovereignty for each participant. This software functions as a secure gateway, separating permission control from the flow of information. Its purpose is to automatically enforce usage policies and licences, ensuring that data is always transmitted directly and in encrypted form.
// HOW DOES IT WORK?
How does the Valle de la Lengua Data Hub work?
Unlike traditional models, where information is stored on a centralised server, the Valle de la Lengua Data Space operates in a federated and decentralised manner. This means that data is not copied to a shared cloud, but remains under the full control of the person who generates it.
An open ecosystem for those who want to innovate with the Spanish language.
Who can take part in the Valle de la Lengua Data Hub ?
The Valle de la Lengua Data Hub is an inclusive ecosystem, open to any organisation wishing to add value or innovate using Spanish-language data, within a framework of federated collaboration.
Key stakeholders in the ecosystem
- Public administrations
- Academic institutions and research centres
- Technology companies and developers
- Cultural and heritage organisations
- Other sector-specific data hubs
Membership is managed through institutional contact and prior validation.
Benefits of participating in the space
Access to top-quality assets:
Validated resources, such as the COE San Millán or specialist terminology.
Sovereignty and total control:
Each participant retains ownership of their intellectual property and decides who can access their data and how.
Legal certainty and ethics:
An operation in line with the AI Act and the European Data Act, ensuring compliance with legal and ethical standards.
Promoting internationalisation:
A solution enabling Spanish-speaking companies to adapt their content for a global audience.
Reducing innovation costs:
It enables solutions to be developed more quickly and at a lower cost.
Institutional support and public funding
EDVAL is an initiative launched by the Government of La Rioja, through the Foundation for the Transformation of La Rioja (FTR) and the National Center for Spanish Language Industries (CNIE), as part of its commitment to the New Language Economy and the development of Spanish in artificial intelligence.
The project is funded by the European Union through NextGenerationEU funds, as part of the Spanish government’s Recovery, Transformation, and Resilience Plan.
Frequently Asked Questions
What is a data space?
A data space is a collaborative ecosystem where organisations share and access data securely, in a controlled manner and in accordance with common rules.
What is a linguistic fact?
Any Spanish-language resource, whether spoken or written (text, audio), together with its metadata and models, that can be used for the development of language technologies and artificial intelligence systems.
Where is the data stored in the Valle de la Lengua Data Space?
Data is not centralised, but is generated under the owner’s control. Data exchange is decentralised and federated, travelling directly from the provider to the consumer via secure technical connectors.
What types of linguistic data can I find and share?
Spoken and written corpora, specialist terminology (health, law, finance), academic knowledge bases, language models (LLMs) and advanced speech and text processing services.
What are the eligibility criteria for taking part?
Obtain a digital identity (DIDs/VCs), accept the Rulebook (common rules), use certified connectors, and comply with European data protection and artificial intelligence regulations.
Do I need to provide any details to take part?
No. You can access the portal and the public catalogue without providing any details. However, the most common options are ‘consumer’ (to make use of existing resources) or ‘provider’ (to share assets), which do require you to provide details. You provide details depending on the role and level of involvement you choose.