The Valle de la Lengua
Data Hub
We are driving the technological future of the Spanish language through a secure, federated infrastructure for sharing, discovering and utilising high-quality linguistic resources
// Valle de la Lengua Data Hub
A challenge to linguistic sovereignty
Spanish is one of the most widely spoken languages in the world, but in the field of Artificial Intelligence it still lags behind English. Many AI models treat it as if it were merely a translation, failing to fully grasp its nuances, regional variations and cultural contexts.
This data hub is the technical solution for bringing together information that is currently scattered. We want to make Spanish a native language in the digital world, enabling machines to understand it in its full depth. We achieve technological autonomy and protect our linguistic richness.
This is made possible by a federated infrastructure that allows for the sharing of high-quality assets whilst always maintaining control at source.
We are driving forward a sovereign AI that understands Spanish in all its diversity.
The data catalogue: the driving force behind the ecosystem
We provide participants with a federated catalogue of high-quality linguistic resources, which are essential for training new technologies. Through this platform, you will be able to access
San Millán Spanish Oral Corpus (COE San Millán)
Millions of words and thousands of hours of audio that reflect how our language is actually spoken.
More information
Specialist terminology
Dictionaries and technical repositories in key fields such as healthcare, engineering and law.
Technology services
Ready-to-use tools, such as natural language processing and automatic speech recognition (ASR), which can be applied in both the public and private sectors.
An ecosystem open to a wide range of profiles
The Data Space is not a closed or static database, but a collaborative network designed to generate shared value among different user groups.
Public authorities
who want to modernise their services and improve the service they provide to the public.
Universities and research centres
which contribute scientific and philological expertise.
Companies and start-ups
who need high-quality data to train artificial intelligence models and compete on a global scale.
Cultural institutions
responsible for preserving our linguistic heritage and promoting it in the digital sphere.
Join the Valle de la Lengua Data Hub
How can you get involved?
Joining this data space is a secure process, based on decentralised digital identity and legal protection. The data always remains at its source: you retain full control over it thanks to sovereign connectors and the shared rules set out in the Rulebook.
You can participate actively according to your needs.
Data provider
Monetising and sharing your resources under your own licences.
Service user
Leveraging assets from the catalogue to create AI products.
Observer
Exploring the portal and the public catalogue before initiating an actual exchange.
News
Keep up to date with technical developments in the infrastructure and the roll-out of new network nodes across the region. Here you will find information on the operational milestones of the Valle de la Lengua Data Space, the launch of flagship use cases, and the latest technical updates to the environment.
La Rioja, the cradle of the Spanish language: from the Monastery of San Millán to digital globalisation
Over a thousand years ago, in the Monastery of San Millán de la Cogolla, in the heart of La Rioja, the first known word in Spanish was written. The Glosas Emilianenses from the 10th century represent the written birth of
EDVAL’s architecture and infrastructure
EDVAL constitutes a strategic infrastructure for Spanish in artificial intelligence, articulated through four complementary technical blocks that guarantee interoperability, linguistic quality, and computational scalabil
EDVAL: The challenge of a generation
95 per cent of the world’s artificial intelligence is trained in English. This means that more than 500 million Spanish speakers are using AI tools that do not fully understand them, and which make mistakes regarding the