Building open AI infrastructure for the languages of Sierra Leone.
SierraLeoneNLP is an open research and developer community focused on building, documenting, evaluating, and sharing Natural Language Processing (NLP), speech, and language AI resources for Sierra Leone.
Our goal is simple:
Make Sierra Leonean languages first-class languages in modern AI.
We believe researchers, developers, students, institutions, and communities should have access to the datasets, models, benchmarks, tools, and research needed to build useful AI systems for Sierra Leonean languages.
Sierra Leone is linguistically diverse, yet many of its languages remain significantly underrepresented in modern AI systems.
SierraLeoneNLP exists to help close this gap by creating an open ecosystem for research and development around Sierra Leonean languages.
We work to:
Our long-term vision is to establish a strong, open, and collaborative Sierra Leonean Language AI ecosystem.
We want a developer anywhere in the world to be able to visit SierraLeoneNLP and find:
for Sierra Leonean languages.
Ultimately, we want Sierra Leonean languages to be represented across the modern AI stack β from datasets and foundational models to speech assistants, translation systems, educational technology, search, accessibility tools, and conversational AI.
SierraLeoneNLP aims to support the linguistic diversity of Sierra Leone.
Our work may include:
The community will expand language coverage as more data, researchers, native speakers, and contributors become involved.
We develop resources for:
We work toward better speech technology for Sierra Leonean languages.
Speech β Text
Projects may include:
Text β Speech
Projects may include:
Speech β Speech
We are interested in both modular and end-to-end speech systems that enable natural interaction with Sierra Leonean languages.
We support research into translation between Sierra Leonean languages and other languages.
Examples include:
Our goal is not simply to produce translation models.
We also want to build the high-quality parallel datasets, evaluation datasets, and benchmarks required to measure meaningful progress.
SierraLeoneNLP supports research into language models capable of understanding and generating Sierra Leonean languages.
This includes:
We encourage researchers to document both successful and unsuccessful experiments where possible so the community can learn collectively.
A major part of SierraLeoneNLP is building reliable benchmarks.
Having a model is not enough.
We need to know:
How well does it actually understand and generate Sierra Leonean languages?
We aim to develop standardized evaluation datasets and benchmarks for:
Evaluation may include both automated metrics and human evaluation by native speakers.
Potential metrics include:
Automated metrics will never be treated as the only measure of language quality.
High-quality data is one of the most important foundations of our work.
SierraLeoneNLP aims to publish and document datasets covering:
Every dataset should provide appropriate documentation describing:
Not every dataset will necessarily be open.
Some datasets may have restrictions because of:
Each dataset should clearly document its access conditions and license.
SierraLeoneNLP follows several principles.
When legally and ethically possible, we encourage open datasets, models, code, documentation, and research.
A large dataset is not automatically a good dataset.
We prioritize:
Sierra Leonean language technology should ultimately be evaluated by people who actually understand and speak the languages.
We encourage researchers to publish:
where possible.
We take privacy, consent, copyright, cultural context, and potential misuse seriously.
SierraLeoneNLP exists to serve the broader research and developer community rather than a single company or individual.
SierraLeoneNLP welcomes contributions from:
You do not need to be an expert to contribute.
You can contribute by:
We aim to maintain a respectful and technically rigorous community.
Contributors should:
SierraLeoneNLP is committed to open research, but openness must be balanced with:
Therefore, projects may have different levels of openness.
A project may be:
Each project should clearly explain its licensing and access conditions.
SierraLeoneNLP may host several types of repositories on Hugging Face.
Examples:
SierraLeoneNLP/krio-text-corpus
SierraLeoneNLP/krio-english-parallel
SierraLeoneNLP/mende-text-corpus
SierraLeoneNLP/temne-speech