{
  "id": 433469,
  "title": "Training Set Metadata, Model Predictions and Data Quality Scores Released",
  "url": "/competitions/bengaliai-speech/discussion/433469",
  "author_name": "Ahmed Imtiaz Humayun",
  "post_date": "2023-08-21T22:39:24.175000",
  "votes": 17,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi All! </p>\n<p>We have released training metadata, google API and yellowking predictions, and NISQA audio quality scores for the training set. <strong>Check out the notebook that explores the metadata here</strong>:<br>\n<a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda</a></p>\n<p>The metadata can be used to subsample the training set for fine-tuning, reduce label noise, match with common voice recording IDs to include newly collected data on the platform or explore metadata from the common voice servers. </p>\n<p><strong>We are releasing:</strong></p>\n<p><em>Data quality assessment scores</em> (using NISQAv2) on the training set <a href=\"https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-nisqa\" target=\"_blank\">https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-nisqa</a> computed using <a href=\"https://github.com/gabrielmittag/NISQA\" target=\"_blank\">https://github.com/gabrielmittag/NISQA</a></p>\n<p><em>Common Voice filename mappings</em> for the training set <a href=\"https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata</a></p>\n<p><em>Yellowking and Google Speech API predictions</em> on the training set <a href=\"https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata</a></p>\n<p><strong>Why do we need to look at training data quality??</strong></p>\n<p>The training data of our dataset was collected via targeted social media crowdsourcing campaigns, using the Mozilla common voice platform. Since the data is crowdsourced, there is always scope for mislabeled data and low recording quality, even after our prelim quality checks.</p>\n<p><strong>Quality Metrics released:</strong><br>\nMOS: Mean Opinion Score<br>\nNOI: Noiseness<br>\nCOL: Coloration<br>\nDIS: Discontinuity<br>\nLOUD: Loudness </p>\n<p><strong>UPDATES:</strong></p>\n<p>The yellowking WERs were updated after normalizing the ground truth and predicted texts by <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a>  here <a href=\"https://www.kaggle.com/code/mbmmurad/bengaliaisr-train-metadata-eda/notebook\" target=\"_blank\">https://www.kaggle.com/code/mbmmurad/bengaliaisr-train-metadata-eda/notebook</a></p>",
  "messages": [
    {
      "id": 2401940,
      "postDate": "2023-08-21T22:39:24.177Z",
      "content": "<p>Hi All! </p>\n<p>We have released training metadata, google API and yellowking predictions, and NISQA audio quality scores for the training set. <strong>Check out the notebook that explores the metadata here</strong>:<br>\n<a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda</a></p>\n<p>The metadata can be used to subsample the training set for fine-tuning, reduce label noise, match with common voice recording IDs to include newly collected data on the platform or explore metadata from the common voice servers. </p>\n<p><strong>We are releasing:</strong></p>\n<p><em>Data quality assessment scores</em> (using NISQAv2) on the training set <a href=\"https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-nisqa\" target=\"_blank\">https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-nisqa</a> computed using <a href=\"https://github.com/gabrielmittag/NISQA\" target=\"_blank\">https://github.com/gabrielmittag/NISQA</a></p>\n<p><em>Common Voice filename mappings</em> for the training set <a href=\"https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata</a></p>\n<p><em>Yellowking and Google Speech API predictions</em> on the training set <a href=\"https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata</a></p>\n<p><strong>Why do we need to look at training data quality??</strong></p>\n<p>The training data of our dataset was collected via targeted social media crowdsourcing campaigns, using the Mozilla common voice platform. Since the data is crowdsourced, there is always scope for mislabeled data and low recording quality, even after our prelim quality checks.</p>\n<p><strong>Quality Metrics released:</strong><br>\nMOS: Mean Opinion Score<br>\nNOI: Noiseness<br>\nCOL: Coloration<br>\nDIS: Discontinuity<br>\nLOUD: Loudness </p>\n<p><strong>UPDATES:</strong></p>\n<p>The yellowking WERs were updated after normalizing the ground truth and predicted texts by <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a>  here <a href=\"https://www.kaggle.com/code/mbmmurad/bengaliaisr-train-metadata-eda/notebook\" target=\"_blank\">https://www.kaggle.com/code/mbmmurad/bengaliaisr-train-metadata-eda/notebook</a></p>",
      "rawMarkdown": "Hi All! \n\nWe have released training metadata, google API and yellowking predictions, and NISQA audio quality scores for the training set. **Check out the notebook that explores the metadata here**:\nhttps://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\n\nThe metadata can be used to subsample the training set for fine-tuning, reduce label noise, match with common voice recording IDs to include newly collected data on the platform or explore metadata from the common voice servers. \n\n**We are releasing:**\n\n*Data quality assessment scores* (using NISQAv2) on the training set https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-nisqa computed using https://github.com/gabrielmittag/NISQA\n\n*Common Voice filename mappings* for the training set https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\n\n*Yellowking and Google Speech API predictions* on the training set https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\n\n**Why do we need to look at training data quality??**\n\nThe training data of our dataset was collected via targeted social media crowdsourcing campaigns, using the Mozilla common voice platform. Since the data is crowdsourced, there is always scope for mislabeled data and low recording quality, even after our prelim quality checks.\n\n**Quality Metrics released:**\nMOS: Mean Opinion Score\nNOI: Noiseness\nCOL: Coloration\nDIS: Discontinuity\nLOUD: Loudness \n\n\n**UPDATES:**\n\nThe yellowking WERs were updated after normalizing the ground truth and predicted texts by @mbmmurad  here https://www.kaggle.com/code/mbmmurad/bengaliaisr-train-metadata-eda/notebook",
      "votes": 17
    },
    {
      "id": 2457511,
      "postDate": "2023-09-27T02:11:20.667Z",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> the <code>NOI: Noiseness</code> and <code>DIS: Discontinuity</code> are also the bigger the better quality? Or the smaller the better? Thanks.</p>",
      "rawMarkdown": "@imtiazprio the `NOI: Noiseness` and `DIS: Discontinuity` are also the bigger the better quality? Or the smaller the better? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 2458731,
          "postDate": "2023-09-27T18:35:21.647Z",
          "content": "<p>Ahh good question. I don't remember exactly right now and am completely stuck with work to be able to check. Could you please run the code here and listen to some recordings with high/low NOI to confirm? It should be clearly noisier for high or low NOI. My hunch is higher NOI means more noisy. or look into the original NIQSA repository for an answer? </p>\n<p><a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda</a></p>",
          "rawMarkdown": "Ahh good question. I don't remember exactly right now and am completely stuck with work to be able to check. Could you please run the code here and listen to some recordings with high/low NOI to confirm? It should be clearly noisier for high or low NOI. My hunch is higher NOI means more noisy. or look into the original NIQSA repository for an answer? \n\nhttps://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda",
          "votes": 1
        }
      ]
    },
    {
      "id": 2410389,
      "postDate": "2023-08-26T21:54:26.877Z",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> <br>\nDoes the test set contain characters that are the same as those in the train and are valid? </p>\n<p>Is it possible for the test data to have characters from other language such as English ?</p>",
      "rawMarkdown": "@imtiazprio \nDoes the test set contain characters that are the same as those in the train and are valid? \n\nIs it possible for the test data to have characters from other language such as English ?",
      "votes": 1,
      "replies": [
        {
          "id": 2410416,
          "postDate": "2023-08-26T23:05:00.837Z",
          "content": "<p>If you mean unicode characters then yes, all characters in test are present in train. No there are only Bengali characters in test. Great questions! </p>",
          "rawMarkdown": "If you mean unicode characters then yes, all characters in test are present in train. No there are only Bengali characters in test. Great questions! ",
          "votes": 3,
          "replies": [
            {
              "id": 2410660,
              "postDate": "2023-08-27T06:30:03.027Z",
              "content": "<p>Instead of utilising the pre-built vocabulary that comes with the pre-trained model, we can construct a vocabulary based on the train and valid data. </p>\n<p>Thanks a lot for sharing this </p>",
              "rawMarkdown": "Instead of utilising the pre-built vocabulary that comes with the pre-trained model, we can construct a vocabulary based on the train and valid data. \n\nThanks a lot for sharing this ",
              "votes": 1
            },
            {
              "id": 2411514,
              "postDate": "2023-08-27T17:05:16.177Z",
              "content": "<p>You could do that, but remember that there are words in test that are not in training!</p>\n<p>Best of luck!!</p>",
              "rawMarkdown": "You could do that, but remember that there are words in test that are not in training!\n\nBest of luck!!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2406424,
      "postDate": "2023-08-24T12:31:08.410Z",
      "content": "<p>Hi. Great reveal. It would be great if you could write a few lines about each quality metric, specifically mean opinion score. Does mean opinion score encapsulate the other 4 metrics? Also which metric among the 5 relates to accuracy of transcription? </p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi. Great reveal. It would be great if you could write a few lines about each quality metric, specifically mean opinion score. Does mean opinion score encapsulate the other 4 metrics? Also which metric among the 5 relates to accuracy of transcription? \n\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 2410404,
          "postDate": "2023-08-26T22:27:50.937Z",
          "content": "<p>Yes mean opinion score encapsulates the rest. Details on the metrics can be found in the papers listed at the bottom of this repo <a href=\"https://github.com/gabrielmittag/NISQA\" target=\"_blank\">https://github.com/gabrielmittag/NISQA</a></p>",
          "rawMarkdown": "Yes mean opinion score encapsulates the rest. Details on the metrics can be found in the papers listed at the bottom of this repo https://github.com/gabrielmittag/NISQA",
          "votes": 3
        },
        {
          "id": 2410415,
          "postDate": "2023-08-26T23:03:36.950Z",
          "content": "<p>None of the metrics relate directly to the accuracy of transcription. One way to find out label error could be looking at the high MOS recordings and thresholding by yellowking WER. Thr logic could be considering Yellowking to be robust and good quality (MOS) recordings with high WER to be mislabeled. </p>",
          "rawMarkdown": "None of the metrics relate directly to the accuracy of transcription. One way to find out label error could be looking at the high MOS recordings and thresholding by yellowking WER. Thr logic could be considering Yellowking to be robust and good quality (MOS) recordings with high WER to be mislabeled. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2431793,
      "postDate": "2023-09-10T12:07:17.437Z",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> Is train dataset labeled text are normlized?   </p>",
      "rawMarkdown": "@imtiazprio Is train dataset labeled text are normlized?   ",
      "votes": -1,
      "replies": [
        {
          "id": 2432239,
          "postDate": "2023-09-10T17:27:00.047Z",
          "content": "<p>Nope. We provide it in the rawest format without any pre-processing.</p>",
          "rawMarkdown": "Nope. We provide it in the rawest format without any pre-processing.",
          "replies": [
            {
              "id": 2432255,
              "postDate": "2023-09-10T17:35:34.743Z",
              "content": "<p>Here Normalization I mean, bnunicodenormalizer normalization need to do on train text before training the model  ?</p>",
              "rawMarkdown": "Here Normalization I mean, bnunicodenormalizer normalization need to do on train text before training the model  ?"
            },
            {
              "id": 2435147,
              "postDate": "2023-09-12T18:16:46.110Z",
              "content": "<p>Yes no normalization was done to the train text. You'd have to do it yourself! </p>",
              "rawMarkdown": "Yes no normalization was done to the train text. You'd have to do it yourself! "
            },
            {
              "id": 2435150,
              "postDate": "2023-09-12T18:18:57.713Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> </p>",
              "rawMarkdown": "Thanks @imtiazprio "
            }
          ]
        }
      ]
    },
    {
      "id": 2431794,
      "postDate": "2023-09-10T12:08:20.057Z",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> Is it necessary to normalized the text data before training AM, LM or it does not affect?</p>",
      "rawMarkdown": "@imtiazprio Is it necessary to normalized the text data before training AM, LM or it does not affect?",
      "replies": [
        {
          "id": 2432241,
          "postDate": "2023-09-10T17:28:39.230Z",
          "content": "<p>Definitely. Without normalization there could be unwanted tokens which are duplicates of the same character(s).</p>",
          "rawMarkdown": "Definitely. Without normalization there could be unwanted tokens which are duplicates of the same character(s)."
        }
      ]
    },
    {
      "id": 2417695,
      "postDate": "2023-08-31T19:10:10.567Z",
      "content": "<p>Hi. Was the ground truths normalised before calculating the CER, WER? The Yellowking model gives normalized outputs, so it'd not be accurate to calculate the WER between the normalized predictions with unnormalized ground truths.<br>\nFrom some analysis of <a href=\"https://www.kaggle.com/mbmmurad/bengaliaisr-train-metadata-eda\" target=\"_blank\">my notebook</a>, it looks like normalizing both the ground truths and predictions gives different WER values. </p>",
      "rawMarkdown": "Hi. Was the ground truths normalised before calculating the CER, WER? The Yellowking model gives normalized outputs, so it'd not be accurate to calculate the WER between the normalized predictions with unnormalized ground truths.\nFrom some analysis of [my notebook](https://www.kaggle.com/mbmmurad/bengaliaisr-train-metadata-eda), it looks like normalizing both the ground truths and predictions gives different WER values. ",
      "replies": [
        {
          "id": 2417869,
          "postDate": "2023-08-31T23:51:14.137Z",
          "content": "<p>Good point, I doubt they were normalized. Looking at the normalized WER should of course be the better thing to do. Ill edit my post and also refer to your notebook in mine.</p>",
          "rawMarkdown": "Good point, I doubt they were normalized. Looking at the normalized WER should of course be the better thing to do. Ill edit my post and also refer to your notebook in mine.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2422759,
      "postDate": "2023-09-04T08:00:48.730Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2457511,
      "author_name": "Joseph Zhou",
      "author_url": "",
      "post_date": "2023-09-27T02:11:20.667000",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> the <code>NOI: Noiseness</code> and <code>DIS: Discontinuity</code> are also the bigger the better quality? Or the smaller the better? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2458731,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-09-27T18:35:21.647000",
          "content": "<p>Ahh good question. I don't remember exactly right now and am completely stuck with work to be able to check. Could you please run the code here and listen to some recordings with high/low NOI to confirm? It should be clearly noisier for high or low NOI. My hunch is higher NOI means more noisy. or look into the original NIQSA repository for an answer? </p>\n<p><a href=\"https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\" target=\"_blank\">https://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2410389,
      "author_name": "Balaji Selvaraj",
      "author_url": "",
      "post_date": "2023-08-26T21:54:26.877000",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> <br>\nDoes the test set contain characters that are the same as those in the train and are valid? </p>\n<p>Is it possible for the test data to have characters from other language such as English ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2410416,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-08-26T23:05:00.837000",
          "content": "<p>If you mean unicode characters then yes, all characters in test are present in train. No there are only Bengali characters in test. Great questions! </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2410660,
              "author_name": "Balaji Selvaraj",
              "author_url": "",
              "post_date": "2023-08-27T06:30:03.027000",
              "content": "<p>Instead of utilising the pre-built vocabulary that comes with the pre-trained model, we can construct a vocabulary based on the train and valid data. </p>\n<p>Thanks a lot for sharing this </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2411514,
              "author_name": "Ahmed Imtiaz Humayun",
              "author_url": "",
              "post_date": "2023-08-27T17:05:16.177000",
              "content": "<p>You could do that, but remember that there are words in test that are not in training!</p>\n<p>Best of luck!!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2406424,
      "author_name": "Abhisek Dash",
      "author_url": "",
      "post_date": "2023-08-24T12:31:08.410000",
      "content": "<p>Hi. Great reveal. It would be great if you could write a few lines about each quality metric, specifically mean opinion score. Does mean opinion score encapsulate the other 4 metrics? Also which metric among the 5 relates to accuracy of transcription? </p>\n<p>Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2410404,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-08-26T22:27:50.937000",
          "content": "<p>Yes mean opinion score encapsulates the rest. Details on the metrics can be found in the papers listed at the bottom of this repo <a href=\"https://github.com/gabrielmittag/NISQA\" target=\"_blank\">https://github.com/gabrielmittag/NISQA</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2410415,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-08-26T23:03:36.950000",
          "content": "<p>None of the metrics relate directly to the accuracy of transcription. One way to find out label error could be looking at the high MOS recordings and thresholding by yellowking WER. Thr logic could be considering Yellowking to be robust and good quality (MOS) recordings with high WER to be mislabeled. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2431793,
      "author_name": "Raj Gothi",
      "author_url": "",
      "post_date": "2023-09-10T12:07:17.437000",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> Is train dataset labeled text are normlized?   </p>",
      "votes": -1,
      "replies": [
        {
          "id": 2432239,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-09-10T17:27:00.047000",
          "content": "<p>Nope. We provide it in the rawest format without any pre-processing.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2432255,
              "author_name": "Raj Gothi",
              "author_url": "",
              "post_date": "2023-09-10T17:35:34.743000",
              "content": "<p>Here Normalization I mean, bnunicodenormalizer normalization need to do on train text before training the model  ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2435147,
              "author_name": "Ahmed Imtiaz Humayun",
              "author_url": "",
              "post_date": "2023-09-12T18:16:46.110000",
              "content": "<p>Yes no normalization was done to the train text. You'd have to do it yourself! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2435150,
              "author_name": "Raj Gothi",
              "author_url": "",
              "post_date": "2023-09-12T18:18:57.713000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2431794,
      "author_name": "Raj Gothi",
      "author_url": "",
      "post_date": "2023-09-10T12:08:20.057000",
      "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> Is it necessary to normalized the text data before training AM, LM or it does not affect?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2432241,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-09-10T17:28:39.230000",
          "content": "<p>Definitely. Without normalization there could be unwanted tokens which are duplicates of the same character(s).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2417695,
      "author_name": "Md Boktiar Mahbub Murad",
      "author_url": "",
      "post_date": "2023-08-31T19:10:10.567000",
      "content": "<p>Hi. Was the ground truths normalised before calculating the CER, WER? The Yellowking model gives normalized outputs, so it'd not be accurate to calculate the WER between the normalized predictions with unnormalized ground truths.<br>\nFrom some analysis of <a href=\"https://www.kaggle.com/mbmmurad/bengaliaisr-train-metadata-eda\" target=\"_blank\">my notebook</a>, it looks like normalizing both the ground truths and predictions gives different WER values. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2417869,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-08-31T23:51:14.137000",
          "content": "<p>Good point, I doubt they were normalized. Looking at the normalized WER should of course be the better thing to do. Ill edit my post and also refer to your notebook in mine.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2422759,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-04T08:00:48.730000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2401940": "Hi All! \n\nWe have released training metadata, google API and yellowking predictions, and NISQA audio quality scores for the training set. **Check out the notebook that explores the metadata here**:\nhttps://www.kaggle.com/code/imtiazprio/listen-to-training-samples-data-quality-eda\n\nThe metadata can be used to subsample the training set for fine-tuning, reduce label noise, match with common voice recording IDs to include newly collected data on the platform or explore metadata from the common voice servers. \n\n**We are releasing:**\n\n*Data quality assessment scores* (using NISQAv2) on the training set https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-nisqa computed using https://github.com/gabrielmittag/NISQA\n\n*Common Voice filename mappings* for the training set https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\n\n*Yellowking and Google Speech API predictions* on the training set https://www.kaggle.com/datasets/imtiazprio/bengaliai-speech-train-metadata\n\n**Why do we need to look at training data quality??**\n\nThe training data of our dataset was collected via targeted social media crowdsourcing campaigns, using the Mozilla common voice platform. Since the data is crowdsourced, there is always scope for mislabeled data and low recording quality, even after our prelim quality checks.\n\n**Quality Metrics released:**\nMOS: Mean Opinion Score\nNOI: Noiseness\nCOL: Coloration\nDIS: Discontinuity\nLOUD: Loudness \n\n\n**UPDATES:**\n\nThe yellowking WERs were updated after normalizing the ground truth and predicted texts by @mbmmurad  here https://www.kaggle.com/code/mbmmurad/bengaliaisr-train-metadata-eda/notebook",
    "2457511": "@imtiazprio the `NOI: Noiseness` and `DIS: Discontinuity` are also the bigger the better quality? Or the smaller the better? Thanks.",
    "2410389": "@imtiazprio \nDoes the test set contain characters that are the same as those in the train and are valid? \n\nIs it possible for the test data to have characters from other language such as English ?",
    "2406424": "Hi. Great reveal. It would be great if you could write a few lines about each quality metric, specifically mean opinion score. Does mean opinion score encapsulate the other 4 metrics? Also which metric among the 5 relates to accuracy of transcription? \n\nThanks",
    "2431793": "@imtiazprio Is train dataset labeled text are normlized?   ",
    "2431794": "@imtiazprio Is it necessary to normalized the text data before training AM, LM or it does not affect?",
    "2417695": "Hi. Was the ground truths normalised before calculating the CER, WER? The Yellowking model gives normalized outputs, so it'd not be accurate to calculate the WER between the normalized predictions with unnormalized ground truths.\nFrom some analysis of [my notebook](https://www.kaggle.com/mbmmurad/bengaliaisr-train-metadata-eda), it looks like normalizing both the ground truths and predictions gives different WER values. ",
    "2422759": ""
  }
}