AI Calabria for universities, research and artificial intelligence
A territorial linguistic corpus built with community, context and voice.
AI Calabria does not merely collect texts: it links forms of speech to places, preserves pronunciation and distinguishes validated content from proposals still awaiting review.
Why this data is different
AI Calabria preserves what general-purpose models tend to lose
Local varieties have little digital data, inconsistent spelling and a strong oral component. That is why provenance and validation are part of the data, not optional metadata.
Municipality-level granularity
Each form is linked to a real territory and can be compared without erasing differences.
Audio and pronunciation
The voice preserves accent, rhythm and sounds that writing alone cannot fully describe.
Human validation
Guardians review contributions before they are published in the shared heritage.
Usage context
Meaning, intention, register and situation help prevent flat or misleading translations.
Traceability
Source, territory, data status and attribution enable more responsible research.
Progressive development
The corpus grows through new contributions, reviews and community expertise.
Areas for collaboration
An infrastructure for linguistics, speech technologies, NLP and digital humanities
Linguistics and dialectology
Territorial distribution, variation, registers, language contact and generational change.
Speech technologies
Pronunciation, transcription, speech recognition and speech synthesis in low-resource settings.
Language models
Evaluation, retrieval, adaptation and hallucination reduction for local varieties.
Digital heritage
Preservation, accessibility, return to communities and content interoperability.
Theses and laboratories
Focused research questions, territorial analysis and interdisciplinary experiments.
Responsible AI
Consent, licences, attribution, governance and usage limits defined before technological exploitation.
Scientific value grows when method, community and technology work together.
Let’s discuss a collaborationFrequently asked questions
Scientific and technological use of AI Calabria’s heritage
Is AI Calabria already a downloadable dataset?
AI Calabria is building a structured linguistic heritage. Availability, licences, access levels and export formats must respect quality, consent, attribution and project governance.
Which metadata make the data useful for research?
Content can be linked to municipality, type, local form, meaning, context, source, validation status and, when available, audio pronunciation.
Why is human validation important for artificial intelligence?
Validation reduces ambiguity and errors, preserves territorial origin and makes it possible to distinguish initial proposals from content recognised by the community.
Can universities and research centres propose collaborations?
Yes. AI Calabria is open to collaborations consistent with heritage protection, methodological rigour, transparency and public benefit.