{
  "id": 491283,
  "title": "External Dataset Thread",
  "url": "/competitions/ben10/discussion/491283",
  "author_name": "Tahsin",
  "post_date": "2024-04-05T09:45:10.662000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The models developed in this competition must be trained on datasets with CC-BY-4.0 license. All new data created to train models must be publicly disclosed after the first phase of the competition.</p>\n<p>Let's dump the available datasets in this thread <br>\nOOD-Speech, <a href=\"https://www.kaggle.com/competitions/bengaliai-speech\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech</a><br>\nSubak.ko, <a href=\"https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO\" target=\"_blank\">https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO</a><br>\nOpenSLR, <a href=\"https://openslr.org/resources.php\" target=\"_blank\">https://openslr.org/resources.php</a><br>\nShrutilipi, <a href=\"https://ai4bharat.iitm.ac.in/shrutilipi/\" target=\"_blank\">https://ai4bharat.iitm.ac.in/shrutilipi/</a><br>\nFleurs, <a href=\"https://huggingface.co/datasets/google/fleurs\" target=\"_blank\">https://huggingface.co/datasets/google/fleurs</a><br>\nMADASR, <a href=\"https://sites.google.com/view/respinasrchallenge2023/dataset\" target=\"_blank\">https://sites.google.com/view/respinasrchallenge2023/dataset</a><br>\nKathbath, <a href=\"https://huggingface.co/datasets/ai4bharat/kathbath\" target=\"_blank\">https://huggingface.co/datasets/ai4bharat/kathbath</a><br>\ncleaned/pseudo data made by Tugstugi, <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110</a></p>\n<p>(edit: new entry)<br>\nShruti, <a href=\"https://cse.iitkgp.ac.in/~pabitra/shruti_corpus.html\" target=\"_blank\">https://cse.iitkgp.ac.in/~pabitra/shruti_corpus.html</a></p>",
  "messages": [
    {
      "id": 2736612,
      "postDate": "2024-04-05T09:45:10.663Z",
      "content": "<p>The models developed in this competition must be trained on datasets with CC-BY-4.0 license. All new data created to train models must be publicly disclosed after the first phase of the competition.</p>\n<p>Let's dump the available datasets in this thread <br>\nOOD-Speech, <a href=\"https://www.kaggle.com/competitions/bengaliai-speech\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech</a><br>\nSubak.ko, <a href=\"https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO\" target=\"_blank\">https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO</a><br>\nOpenSLR, <a href=\"https://openslr.org/resources.php\" target=\"_blank\">https://openslr.org/resources.php</a><br>\nShrutilipi, <a href=\"https://ai4bharat.iitm.ac.in/shrutilipi/\" target=\"_blank\">https://ai4bharat.iitm.ac.in/shrutilipi/</a><br>\nFleurs, <a href=\"https://huggingface.co/datasets/google/fleurs\" target=\"_blank\">https://huggingface.co/datasets/google/fleurs</a><br>\nMADASR, <a href=\"https://sites.google.com/view/respinasrchallenge2023/dataset\" target=\"_blank\">https://sites.google.com/view/respinasrchallenge2023/dataset</a><br>\nKathbath, <a href=\"https://huggingface.co/datasets/ai4bharat/kathbath\" target=\"_blank\">https://huggingface.co/datasets/ai4bharat/kathbath</a><br>\ncleaned/pseudo data made by Tugstugi, <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110</a></p>\n<p>(edit: new entry)<br>\nShruti, <a href=\"https://cse.iitkgp.ac.in/~pabitra/shruti_corpus.html\" target=\"_blank\">https://cse.iitkgp.ac.in/~pabitra/shruti_corpus.html</a></p>",
      "rawMarkdown": "The models developed in this competition must be trained on datasets with CC-BY-4.0 license. All new data created to train models must be publicly disclosed after the first phase of the competition.\n\nLet's dump the available datasets in this thread \nOOD-Speech, https://www.kaggle.com/competitions/bengaliai-speech\nSubak.ko, https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO\nOpenSLR, https://openslr.org/resources.php\nShrutilipi, https://ai4bharat.iitm.ac.in/shrutilipi/\nFleurs, https://huggingface.co/datasets/google/fleurs\nMADASR, https://sites.google.com/view/respinasrchallenge2023/dataset\nKathbath, https://huggingface.co/datasets/ai4bharat/kathbath\ncleaned/pseudo data made by Tugstugi, https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110\n\n(edit: new entry)\nShruti, https://cse.iitkgp.ac.in/~pabitra/shruti_corpus.html",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2736612": "The models developed in this competition must be trained on datasets with CC-BY-4.0 license. All new data created to train models must be publicly disclosed after the first phase of the competition.\n\nLet's dump the available datasets in this thread \nOOD-Speech, https://www.kaggle.com/competitions/bengaliai-speech\nSubak.ko, https://huggingface.co/datasets/SUST-CSE-Speech/SUBAK.KO\nOpenSLR, https://openslr.org/resources.php\nShrutilipi, https://ai4bharat.iitm.ac.in/shrutilipi/\nFleurs, https://huggingface.co/datasets/google/fleurs\nMADASR, https://sites.google.com/view/respinasrchallenge2023/dataset\nKathbath, https://huggingface.co/datasets/ai4bharat/kathbath\ncleaned/pseudo data made by Tugstugi, https://www.kaggle.com/competitions/bengaliai-speech/discussion/448110\n\n(edit: new entry)\nShruti, https://cse.iitkgp.ac.in/~pabitra/shruti_corpus.html"
  }
}