Zurück zum Ranking

common-voice/common-voice

TypeScriptcommonvoice.mozilla.org

Common Voice is part of Mozilla's initiative to help teach machines how real people speak.

open-datavoicecrowdsourcinginternet-freedom
Sterne-Wachstum
Sterne
3.5k
Forks
871
Wochenwachstum
Issues
173
1k2k3k
Mai 2017Mai 2020Juni 2023Juli 2026
Artefaktenpmnpm install common-voice
README

Common Voice

This is the web app for Mozilla Common Voice, a platform for collecting speech donations in order to create public domain datasets for training voice recognition-related tools.

Upcoming releases

Type Release Cadence More info
Platform code & sentences Monthly, or as needed Release notes
Dataset Quarterly Dataset metadata

How to contribute

🎉 First off, thanks for taking the time to contribute! This project would not be possible without people like you. 🎉

There are many ways to get involved with Common Voice - you don't have to know how to code to contribute!

  • To add or correct the translation of the web interface, please use the Mozilla localization platform Pontoon. Please note, we do not accept any direct pull requests for changing localization content.
  • For information on how to add or edit sentences to Common Voice, see SENTENCES.md
  • For instructions on setting up a local development environment, see DEVELOPMENT.md
  • For information on how to add a new language to Common Voice, see LANGUAGE.md
  • For information on how to get in contact with existing language communities, see COMMUNITIES.md

For more general guidance on building your own language community using Mozilla voice tools, please refer to the Mozilla Voice Community Playbook.

Discussion

For general discussion (feedback, ideas, random musings), head to our Discourse Category.

For bug reports or specific feature, please use the GitHub issue tracker.

For live chat, join us on Matrix.

Licensing and content source

This repository is released under MPL (Mozilla Public License) 2.0.

The majority of our sentence text in /server/data comes directly from user submissions in our Sentence Collector or they are scraped from Wikipedia using our extractor tool, and are released under a CC0 public domain Creative Commons license.

Any files that follow the pattern europarl-VERSION-LANG.txt (such as europarl-v7-de.txt) were extracted with our thanks from the Europarl Corpus, which features transcripts from proceedings in the European parliament.

Citation

If you use the data in a published academic work we would appreciate if you cite the following article:

  • Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M. and Weber, G. (2020) "Common Voice: A Massively-Multilingual Speech Corpus". Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020). pp. 4211—4215

The BiBTex is:

@inproceedings{commonvoice:2020,
  author = {Ardila, R. and Branson, M. and Davis, K. and Henretty, M. and Kohler, M. and Meyer, J. and Morais, R. and Saunders, L. and Tyers, F. M. and Weber, G.},
  title = {Common Voice: A Massively-Multilingual Speech Corpus},
  booktitle = {Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020)},
  pages = {4211--4215},
  year = 2020
}

Cross Browser Testing

This project is tested with Browserstack

Ähnliche Repositories
public-api-lists/public-api-lists

A curated list of free public APIs — searchable, community-maintained, with a free JSON API.

MIT Licensepublic-apiapi
public-api-lists.github.io/public-api-lists/
15.2k1.6k
jackvale/rectg

Telegram 中文频道、群组与机器人精选索引,结合自动化抓取与人工整理,支持在线搜索与分类浏览。

PythonPyPIApache License 2.0telegramtelegram-bot
rectg.com
9.1k528
ckan/ckan

CKAN is an open-source DMS (data management system) for powering data hubs and data portals. CKAN makes it easy to publish, share and use data. It powers catalog.data.gov, open.canada.ca/data, data.humdata.org among many other sites.

PythonPyPIOtherpythonopen-data
ckan.org
5.1k2.1k
okfn-brasil/serenata-de-amor

🕵 Artificial Intelligence for social control of public administration | **This repository does not receive frequent updates. Check out the README**

PythonPyPIMIT Licensemachine-learningdata-science
serenata.ai/en
4.6k654
hudl/open-data

Free football data from StatsBomb

Otherfootballfootball-data
statsbomb.com/resource-centre/
3.5k959
codyogden/killedbygoogle

Part guillotine, part graveyard for Google's doomed apps, services, and hardware.

TypeScriptnpmMIT Licensereactgoogle
killedbygoogle.com
2.7k411
mdeff/fma

FMA: A Dataset For Music Analysis

Jupyter NotebookMIT Licensedatasetmusic-analysis
arxiv.org/abs/1612.01840
2.6k456
statsbomb/open-data

Free football data from StatsBomb

Otherfootballfootball-data
statsbomb.com/resource-centre/
2.6k795
github/CodeSearchNet

Datasets, tools, and benchmarks for representation learning of code.

Jupyter NotebookMIT Licensedeep-learningnatural-language-processing
arxiv.org/abs/1909.09436
2.4k408
datopian/portaljs

🌀 AI-native framework for building data portals. Scaffold a full portal from a brief and load datasets in minutes with agentic skills — any backend (CKAN, GitHub, Frictionless).

TypeScriptnpmMIT Licensedata-portalckan
portaljs.com
2.3k334
open-thoughts/open-thoughts

Fully open data curation for reasoning models

PythonPyPIApache License 2.0open-datareasoning
open-thoughts.ai
2.3k193
GSA/datagov-wptheme

Data.gov WordPress Theme (obsolete)

JavaScriptnpmOtheropen-datagovernment
data.gov
1.9k406