{
  "id": 425823,
  "title": "External Datasets",
  "url": "/competitions/bengaliai-speech/discussion/425823",
  "author_name": "",
  "post_date": "2023-07-20T14:39:02.707902700Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thread of available External Datasets for Bengali ASR<br>\nI Found the folllowing Datasets in the 🤗 Dataset Hub</p>\n<ul>\n<li><a href=\"https://huggingface.co/datasets/google/fleurs/viewer/bn_in/train\" target=\"_blank\">Fleurs Dataset</a></li>\n<li><a href=\"https://huggingface.co/datasets/openslr/viewer\" target=\"_blank\">OpenSLR</a></li>\n<li><a href=\"https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0\" target=\"_blank\">Common Voice</a> This Dataset requires approval to be accesed</li>\n</ul>\n<h2>The first 2 datasets are saved to Kaggle in <a href=\"https://www.kaggle.com/datasets/ksmcg90/bengali-asr-datasets\" target=\"_blank\">bengali-asr-datasets</a></h2>",
  "messages": [
    {
      "id": "2351990",
      "postDate": "07/20/2023 14:39:02",
      "content": "<p>Thread of available External Datasets for Bengali ASR<br>\nI Found the folllowing Datasets in the 🤗 Dataset Hub</p>\n<ul>\n<li><a href=\"https://huggingface.co/datasets/google/fleurs/viewer/bn_in/train\" target=\"_blank\">Fleurs Dataset</a></li>\n<li><a href=\"https://huggingface.co/datasets/openslr/viewer\" target=\"_blank\">OpenSLR</a></li>\n<li><a href=\"https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0\" target=\"_blank\">Common Voice</a> This Dataset requires approval to be accesed</li>\n</ul>\n<h2>The first 2 datasets are saved to Kaggle in <a href=\"https://www.kaggle.com/datasets/ksmcg90/bengali-asr-datasets\" target=\"_blank\">bengali-asr-datasets</a></h2>",
      "rawMarkdown": "Thread of available External Datasets for Bengali ASR\nI Found the folllowing Datasets in the 🤗 Dataset Hub\n* [Fleurs Dataset](https://huggingface.co/datasets/google/fleurs/viewer/bn_in/train)\n* [OpenSLR](https://huggingface.co/datasets/openslr/viewer)\n* [Common Voice](https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0) This Dataset requires approval to be accesed\n\n## The first 2 datasets are saved to Kaggle in [bengali-asr-datasets](https://www.kaggle.com/datasets/ksmcg90/bengali-asr-datasets)",
      "votes": null
    },
    {
      "id": "2359567",
      "postDate": "07/26/2023 08:40:29",
      "content": "<p>New to kaggle competition. Why the data organized the funny way here? Where is audios?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F8f7692aadb7c1ddc623b9be02e9a69a9%2FScreen%20Shot%202023-07-26%20at%201.37.02%20AM.png?generation=1690360860015607&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "New to kaggle competition. Why the data organized the funny way here? Where is audios?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F8f7692aadb7c1ddc623b9be02e9a69a9%2FScreen%20Shot%202023-07-26%20at%201.37.02%20AM.png?generation=1690360860015607&alt=media)",
      "votes": null
    },
    {
      "id": "2359571",
      "postDate": "07/26/2023 08:43:26",
      "content": "<p>The same data download from openSLR  is so different.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F81f2b10a22612f1aeee65eea28a249bf%2FScreen%20Shot%202023-07-26%20at%201.42.20%20AM.png?generation=1690361004677138&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The same data download from openSLR  is so different.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F81f2b10a22612f1aeee65eea28a249bf%2FScreen%20Shot%202023-07-26%20at%201.42.20%20AM.png?generation=1690361004677138&alt=media)",
      "votes": null
    },
    {
      "id": "2363798",
      "postDate": "07/28/2023 23:55:25",
      "content": "<p>It's the same dataset in different format. You can load them using load_dataset(openslr_bn)</p>",
      "rawMarkdown": "It's the same dataset in different format. You can load them using load_dataset(openslr_bn)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2359567,
      "author_name": "robot2020",
      "author_url": "",
      "post_date": "07/26/2023 08:40:29",
      "content": "<p>New to kaggle competition. Why the data organized the funny way here? Where is audios?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F8f7692aadb7c1ddc623b9be02e9a69a9%2FScreen%20Shot%202023-07-26%20at%201.37.02%20AM.png?generation=1690360860015607&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2359571,
          "author_name": "robot2020",
          "author_url": "",
          "post_date": "07/26/2023 08:43:26",
          "content": "<p>The same data download from openSLR  is so different.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F81f2b10a22612f1aeee65eea28a249bf%2FScreen%20Shot%202023-07-26%20at%201.42.20%20AM.png?generation=1690361004677138&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2363798,
          "author_name": "mahfuzulkabirsourav",
          "author_url": "",
          "post_date": "07/28/2023 23:55:25",
          "content": "<p>It's the same dataset in different format. You can load them using load_dataset(openslr_bn)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2351990": "Thread of available External Datasets for Bengali ASR\nI Found the folllowing Datasets in the 🤗 Dataset Hub\n* [Fleurs Dataset](https://huggingface.co/datasets/google/fleurs/viewer/bn_in/train)\n* [OpenSLR](https://huggingface.co/datasets/openslr/viewer)\n* [Common Voice](https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0) This Dataset requires approval to be accesed\n\n## The first 2 datasets are saved to Kaggle in [bengali-asr-datasets](https://www.kaggle.com/datasets/ksmcg90/bengali-asr-datasets)",
    "2359567": "New to kaggle competition. Why the data organized the funny way here? Where is audios?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F8f7692aadb7c1ddc623b9be02e9a69a9%2FScreen%20Shot%202023-07-26%20at%201.37.02%20AM.png?generation=1690360860015607&alt=media)",
    "2359571": "The same data download from openSLR  is so different.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F81f2b10a22612f1aeee65eea28a249bf%2FScreen%20Shot%202023-07-26%20at%201.42.20%20AM.png?generation=1690361004677138&alt=media)",
    "2363798": "It's the same dataset in different format. You can load them using load_dataset(openslr_bn)"
  },
  "source": "meta"
}