{
  "id": 432639,
  "title": "What exactly is meant by out-of-distribution audio recordings as mentioned in the competition overview page?",
  "url": "/competitions/bengaliai-speech/discussion/432639",
  "author_name": "",
  "post_date": "2023-08-18T08:42:44.443039800Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The overview page states that the objective of this competition is to \"recognize Bengali speech from out-of-distribution audio recordings\" . So what exactly is out-of-distribution audio recordings?</p>",
  "messages": [
    {
      "id": "2396420",
      "postDate": "08/18/2023 08:42:44",
      "content": "<p>The overview page states that the objective of this competition is to \"recognize Bengali speech from out-of-distribution audio recordings\" . So what exactly is out-of-distribution audio recordings?</p>",
      "rawMarkdown": "The overview page states that the objective of this competition is to \"recognize Bengali speech from out-of-distribution audio recordings\" . So what exactly is out-of-distribution audio recordings?",
      "votes": null
    },
    {
      "id": "2396704",
      "postDate": "08/18/2023 12:22:51",
      "content": "<p>The \"distribution\" of input data, which consists of audio files in this competition, implies that these inputs share common characteristics. In speech audio files, these characteristics can encompass tonality, background noise, language, speech from movies, poems, and more. If the distribution of the training data significantly diverges from that of the testing data, it indicates that the audio files in the test set lack the similar characteristics present in the training set's audio files. In this case, the test dataset falls outside the distribution of the training dataset, rendering it out of distribution.<br>\nI hope this explanation is helpful.</p>",
      "rawMarkdown": "The \"distribution\" of input data, which consists of audio files in this competition, implies that these inputs share common characteristics. In speech audio files, these characteristics can encompass tonality, background noise, language, speech from movies, poems, and more. If the distribution of the training data significantly diverges from that of the testing data, it indicates that the audio files in the test set lack the similar characteristics present in the training set's audio files. In this case, the test dataset falls outside the distribution of the training dataset, rendering it out of distribution.\nI hope this explanation is helpful.",
      "votes": null
    },
    {
      "id": "2397541",
      "postDate": "08/19/2023 05:37:07",
      "content": "<p>Great question <a href=\"https://www.kaggle.com/abhisekdash37\" target=\"_blank\">@abhisekdash37</a> and great answer by <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a>! Adding some details to Alenic's answer.</p>\n<p>Our training set (MaCro Train) comprises crowdsourced scripted speech, meaning, contributors were asked to read out a piece of text in uncontrolled natural environments.</p>\n<p>There are two parts of our test set, one in-distribution (MaCro Test) and another out-of-distribution (OOD Test). The in-distribution part of the test set also comprises crowdsourced scripted speech collected from natural environments. On the other hand, the out-of-distribution part of the test set has spontaneous speech from acoustic environments which are intentionally chosen such that they are not present in training.</p>\n<p>Due to this mismatch, you'll see in Table 2 of the dataset paper that the MaCro Test WER is significantly lower than the other domains of the test set, especially for Whisper-small.<br>\n<a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">https://arxiv.org/abs/2305.09688</a></p>",
      "rawMarkdown": "Great question @abhisekdash37 and great answer by @alenic! Adding some details to Alenic's answer.\n\nOur training set (MaCro Train) comprises crowdsourced scripted speech, meaning, contributors were asked to read out a piece of text in uncontrolled natural environments.\n\nThere are two parts of our test set, one in-distribution (MaCro Test) and another out-of-distribution (OOD Test). The in-distribution part of the test set also comprises crowdsourced scripted speech collected from natural environments. On the other hand, the out-of-distribution part of the test set has spontaneous speech from acoustic environments which are intentionally chosen such that they are not present in training.\n\nDue to this mismatch, you'll see in Table 2 of the dataset paper that the MaCro Test WER is significantly lower than the other domains of the test set, especially for Whisper-small.\nhttps://arxiv.org/abs/2305.09688",
      "votes": null
    },
    {
      "id": "2400902",
      "postDate": "08/21/2023 09:58:11",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2396704,
      "author_name": "alenic",
      "author_url": "",
      "post_date": "08/18/2023 12:22:51",
      "content": "<p>The \"distribution\" of input data, which consists of audio files in this competition, implies that these inputs share common characteristics. In speech audio files, these characteristics can encompass tonality, background noise, language, speech from movies, poems, and more. If the distribution of the training data significantly diverges from that of the testing data, it indicates that the audio files in the test set lack the similar characteristics present in the training set's audio files. In this case, the test dataset falls outside the distribution of the training dataset, rendering it out of distribution.<br>\nI hope this explanation is helpful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2400902,
          "author_name": "abhisekdash37",
          "author_url": "",
          "post_date": "08/21/2023 09:58:11",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2397541,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "08/19/2023 05:37:07",
      "content": "<p>Great question <a href=\"https://www.kaggle.com/abhisekdash37\" target=\"_blank\">@abhisekdash37</a> and great answer by <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a>! Adding some details to Alenic's answer.</p>\n<p>Our training set (MaCro Train) comprises crowdsourced scripted speech, meaning, contributors were asked to read out a piece of text in uncontrolled natural environments.</p>\n<p>There are two parts of our test set, one in-distribution (MaCro Test) and another out-of-distribution (OOD Test). The in-distribution part of the test set also comprises crowdsourced scripted speech collected from natural environments. On the other hand, the out-of-distribution part of the test set has spontaneous speech from acoustic environments which are intentionally chosen such that they are not present in training.</p>\n<p>Due to this mismatch, you'll see in Table 2 of the dataset paper that the MaCro Test WER is significantly lower than the other domains of the test set, especially for Whisper-small.<br>\n<a href=\"https://arxiv.org/abs/2305.09688\" target=\"_blank\">https://arxiv.org/abs/2305.09688</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2396420": "The overview page states that the objective of this competition is to \"recognize Bengali speech from out-of-distribution audio recordings\" . So what exactly is out-of-distribution audio recordings?",
    "2396704": "The \"distribution\" of input data, which consists of audio files in this competition, implies that these inputs share common characteristics. In speech audio files, these characteristics can encompass tonality, background noise, language, speech from movies, poems, and more. If the distribution of the training data significantly diverges from that of the testing data, it indicates that the audio files in the test set lack the similar characteristics present in the training set's audio files. In this case, the test dataset falls outside the distribution of the training dataset, rendering it out of distribution.\nI hope this explanation is helpful.",
    "2397541": "Great question @abhisekdash37 and great answer by @alenic! Adding some details to Alenic's answer.\n\nOur training set (MaCro Train) comprises crowdsourced scripted speech, meaning, contributors were asked to read out a piece of text in uncontrolled natural environments.\n\nThere are two parts of our test set, one in-distribution (MaCro Test) and another out-of-distribution (OOD Test). The in-distribution part of the test set also comprises crowdsourced scripted speech collected from natural environments. On the other hand, the out-of-distribution part of the test set has spontaneous speech from acoustic environments which are intentionally chosen such that they are not present in training.\n\nDue to this mismatch, you'll see in Table 2 of the dataset paper that the MaCro Test WER is significantly lower than the other domains of the test set, especially for Whisper-small.\nhttps://arxiv.org/abs/2305.09688",
    "2400902": "Thank you :)"
  },
  "source": "meta"
}