{
  "id": 426970,
  "title": "Dataset overlaps with Common Voice dataset : Might lead to overfitting?",
  "url": "/competitions/bengaliai-speech/discussion/426970",
  "author_name": "",
  "post_date": "2023-07-25T20:55:52.594747700Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<blockquote>\n  <p>Most of the best-performing notebooks in this competition used Common Voice Dataset for their training and validation. <br>\n  The <a href=\"https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference\" target=\"_blank\">Yellowking wav2vec2 model</a> used 45k+36k = 81k samples from the <a href=\"https://huggingface.co/datasets/bengaliAI/cvbn\" target=\"_blank\">CV dataset</a> for their <a href=\"https://www.kaggle.com/code/sameen53/yellowking-dlsprint-training\" target=\"_blank\">training</a>. <br>\n  The <a href=\"https://huggingface.co/bangla-speech-processing/BanglaASR\" target=\"_blank\">whisper model</a> was also trained using the Mozilla common voice dataset. </p>\n</blockquote>\n<p>Now we might think of finetuning this model checkpoints for getting better results. But here remains a problem!</p>\n<p>This dataset actually contains most of the <strong>data samples present in the common voice dataset</strong>.</p>\n<blockquote>\n  <p>In <a href=\"https://www.kaggle.com/mbmmurad/dataset-overlaps-with-commonvoice-11-bn\" target=\"_blank\">this notebook</a> I have explored both the datasets and found out 285113 sentences among the 950k+ sentences in this competition data were actually present in the common voice data. </p>\n</blockquote>\n<p>Now sentences might be present in both datasets since we can have multiple audios for the same sentence. But the problem is this dataset contains the exact same audio samples from the common voice dataset samples. <br>\nSome of them are in the train split, and some of them are in the <strong>validation split</strong>. So if you decide to finetune the aforementioned checkpoints and you decide to train it with a subsample of the actual set (for example 100k audios) chances are the model has already seen some of the audios in the training set and in the validation set, thus it may lead to overfitting.</p>\n<blockquote>\n  <p>Workaround :</p>\n  <ol>\n  <li>Train from scratch following the way these models were trained. Not from these checkpoints.</li>\n  <li>Remove the samples present in both datasets, especially those that are present in the validation set.</li>\n  </ol>\n</blockquote>\n<p>Please do share your thoughts/comments! </p>",
  "messages": [
    {
      "id": "2358932",
      "postDate": "07/25/2023 20:55:52",
      "content": "<blockquote>\n  <p>Most of the best-performing notebooks in this competition used Common Voice Dataset for their training and validation. <br>\n  The <a href=\"https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference\" target=\"_blank\">Yellowking wav2vec2 model</a> used 45k+36k = 81k samples from the <a href=\"https://huggingface.co/datasets/bengaliAI/cvbn\" target=\"_blank\">CV dataset</a> for their <a href=\"https://www.kaggle.com/code/sameen53/yellowking-dlsprint-training\" target=\"_blank\">training</a>. <br>\n  The <a href=\"https://huggingface.co/bangla-speech-processing/BanglaASR\" target=\"_blank\">whisper model</a> was also trained using the Mozilla common voice dataset. </p>\n</blockquote>\n<p>Now we might think of finetuning this model checkpoints for getting better results. But here remains a problem!</p>\n<p>This dataset actually contains most of the <strong>data samples present in the common voice dataset</strong>.</p>\n<blockquote>\n  <p>In <a href=\"https://www.kaggle.com/mbmmurad/dataset-overlaps-with-commonvoice-11-bn\" target=\"_blank\">this notebook</a> I have explored both the datasets and found out 285113 sentences among the 950k+ sentences in this competition data were actually present in the common voice data. </p>\n</blockquote>\n<p>Now sentences might be present in both datasets since we can have multiple audios for the same sentence. But the problem is this dataset contains the exact same audio samples from the common voice dataset samples. <br>\nSome of them are in the train split, and some of them are in the <strong>validation split</strong>. So if you decide to finetune the aforementioned checkpoints and you decide to train it with a subsample of the actual set (for example 100k audios) chances are the model has already seen some of the audios in the training set and in the validation set, thus it may lead to overfitting.</p>\n<blockquote>\n  <p>Workaround :</p>\n  <ol>\n  <li>Train from scratch following the way these models were trained. Not from these checkpoints.</li>\n  <li>Remove the samples present in both datasets, especially those that are present in the validation set.</li>\n  </ol>\n</blockquote>\n<p>Please do share your thoughts/comments! </p>",
      "rawMarkdown": ">Most of the best-performing notebooks in this competition used Common Voice Dataset for their training and validation. \n>The [Yellowking wav2vec2 model](https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference) used 45k+36k = 81k samples from the [CV dataset](https://huggingface.co/datasets/bengaliAI/cvbn) for their [training](https://www.kaggle.com/code/sameen53/yellowking-dlsprint-training). \nThe [whisper model](https://huggingface.co/bangla-speech-processing/BanglaASR) was also trained using the Mozilla common voice dataset. \n\nNow we might think of finetuning this model checkpoints for getting better results. But here remains a problem!\n\nThis dataset actually contains most of the **data samples present in the common voice dataset**.\n\n> In [this notebook](https://www.kaggle.com/mbmmurad/dataset-overlaps-with-commonvoice-11-bn) I have explored both the datasets and found out 285113 sentences among the 950k+ sentences in this competition data were actually present in the common voice data. \n\nNow sentences might be present in both datasets since we can have multiple audios for the same sentence. But the problem is this dataset contains the exact same audio samples from the common voice dataset samples. \nSome of them are in the train split, and some of them are in the **validation split**. So if you decide to finetune the aforementioned checkpoints and you decide to train it with a subsample of the actual set (for example 100k audios) chances are the model has already seen some of the audios in the training set and in the validation set, thus it may lead to overfitting.\n\n>Workaround :\n1. Train from scratch following the way these models were trained. Not from these checkpoints.\n2. Remove the samples present in both datasets, especially those that are present in the validation set.\n\nPlease do share your thoughts/comments!",
      "votes": null
    },
    {
      "id": "2360115",
      "postDate": "07/26/2023 15:03:15",
      "content": "<p>my suggestion is to read the datset paper carefully.</p>\n<p>note that for OOD test data, there is OOV (out of vocab).<br>\nThis makes me wonder if the kaggle dataset i sufficient at all.</p>\n<p>(e.g. you can check the word histogram of train mps and exmples wave)</p>\n<p>i think mopst people will train LM word prediction for correcting CTC charcter-wise prediction.<br>\ni think we need another train set for LM</p>",
      "rawMarkdown": "my suggestion is to read the datset paper carefully.\n\nnote that for OOD test data, there is OOV (out of vocab).\nThis makes me wonder if the kaggle dataset i sufficient at all.\n\n(e.g. you can check the word histogram of train mps and exmples wave)\n\ni think mopst people will train LM word prediction for correcting CTC charcter-wise prediction.\ni think we need another train set for LM",
      "votes": null
    },
    {
      "id": "2362276",
      "postDate": "07/28/2023 00:48:19",
      "content": "<p>Yeah, there are some larger text corpora that can be used for LM training, like the one they <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">used here</a>. But this would lead to slower inference tho. Another point is for OOD vocabs, no one knows how large vocab is sufficient. And even LM can't solve some issues like handling named enitites/similar words</p>",
      "rawMarkdown": "Yeah, there are some larger text corpora that can be used for LM training, like the one they [used here](https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali). But this would lead to slower inference tho. Another point is for OOD vocabs, no one knows how large vocab is sufficient. And even LM can't solve some issues like handling named enitites/similar words",
      "votes": null
    },
    {
      "id": "2362745",
      "postDate": "07/28/2023 09:20:59",
      "content": "<p>wtf, how come an independent dataset contains data from others? I don't think they mention this in their paper. I mean, is this even morally correct? Hope hosts can clarify on this. </p>",
      "rawMarkdown": "wtf, how come an independent dataset contains data from others? I don't think they mention this in their paper. I mean, is this even morally correct? Hope hosts can clarify on this.",
      "votes": null
    },
    {
      "id": "2379120",
      "postDate": "08/08/2023 02:34:03",
      "content": "<p>Yes its clearly mentioned in the paper! We crowdsourced data on the common voice platform through online campaigns and then curated it automatically+manually to get the training set of this dataset. If you dig deeper you'll see that most of the public models are partially trained on common voice, not the 1500+ hrs.</p>",
      "rawMarkdown": "Yes its clearly mentioned in the paper! We crowdsourced data on the common voice platform through online campaigns and then curated it automatically+manually to get the training set of this dataset. If you dig deeper you'll see that most of the public models are partially trained on common voice, not the 1500+ hrs.",
      "votes": null
    },
    {
      "id": "2379186",
      "postDate": "08/08/2023 04:16:40",
      "content": "<p>Thanks for the clarification. I initially thought that common voice is a not-updating dataset. Now I understand that it's also a platform and updating over time. So the data collected form  campaigns has already been included in Common Voice Corpus 12.0 12/15/2022, right?</p>",
      "rawMarkdown": "Thanks for the clarification. I initially thought that common voice is a not-updating dataset. Now I understand that it's also a platform and updating over time. So the data collected form  campaigns has already been included in Common Voice Corpus 12.0 12/15/2022, right?",
      "votes": null
    },
    {
      "id": "2379235",
      "postDate": "08/08/2023 04:56:52",
      "content": "<p>Yes! We have included stuff from upto Common Voice 13 (not all of it) in this training set. We should have a list of common voice -&gt; training mapping that we can share. </p>",
      "rawMarkdown": "Yes! We have included stuff from upto Common Voice 13 (not all of it) in this training set. We should have a list of common voice -> training mapping that we can share.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2360115,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/26/2023 15:03:15",
      "content": "<p>my suggestion is to read the datset paper carefully.</p>\n<p>note that for OOD test data, there is OOV (out of vocab).<br>\nThis makes me wonder if the kaggle dataset i sufficient at all.</p>\n<p>(e.g. you can check the word histogram of train mps and exmples wave)</p>\n<p>i think mopst people will train LM word prediction for correcting CTC charcter-wise prediction.<br>\ni think we need another train set for LM</p>",
      "votes": null,
      "replies": [
        {
          "id": 2362276,
          "author_name": "mbmmurad",
          "author_url": "",
          "post_date": "07/28/2023 00:48:19",
          "content": "<p>Yeah, there are some larger text corpora that can be used for LM training, like the one they <a href=\"https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali\" target=\"_blank\">used here</a>. But this would lead to slower inference tho. Another point is for OOD vocabs, no one knows how large vocab is sufficient. And even LM can't solve some issues like handling named enitites/similar words</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2362745,
      "author_name": "aphysict",
      "author_url": "",
      "post_date": "07/28/2023 09:20:59",
      "content": "<p>wtf, how come an independent dataset contains data from others? I don't think they mention this in their paper. I mean, is this even morally correct? Hope hosts can clarify on this. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2379120,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "08/08/2023 02:34:03",
      "content": "<p>Yes its clearly mentioned in the paper! We crowdsourced data on the common voice platform through online campaigns and then curated it automatically+manually to get the training set of this dataset. If you dig deeper you'll see that most of the public models are partially trained on common voice, not the 1500+ hrs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2379186,
          "author_name": "aphysict",
          "author_url": "",
          "post_date": "08/08/2023 04:16:40",
          "content": "<p>Thanks for the clarification. I initially thought that common voice is a not-updating dataset. Now I understand that it's also a platform and updating over time. So the data collected form  campaigns has already been included in Common Voice Corpus 12.0 12/15/2022, right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2379235,
              "author_name": "imtiazprio",
              "author_url": "",
              "post_date": "08/08/2023 04:56:52",
              "content": "<p>Yes! We have included stuff from upto Common Voice 13 (not all of it) in this training set. We should have a list of common voice -&gt; training mapping that we can share. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2358932": ">Most of the best-performing notebooks in this competition used Common Voice Dataset for their training and validation. \n>The [Yellowking wav2vec2 model](https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference) used 45k+36k = 81k samples from the [CV dataset](https://huggingface.co/datasets/bengaliAI/cvbn) for their [training](https://www.kaggle.com/code/sameen53/yellowking-dlsprint-training). \nThe [whisper model](https://huggingface.co/bangla-speech-processing/BanglaASR) was also trained using the Mozilla common voice dataset. \n\nNow we might think of finetuning this model checkpoints for getting better results. But here remains a problem!\n\nThis dataset actually contains most of the **data samples present in the common voice dataset**.\n\n> In [this notebook](https://www.kaggle.com/mbmmurad/dataset-overlaps-with-commonvoice-11-bn) I have explored both the datasets and found out 285113 sentences among the 950k+ sentences in this competition data were actually present in the common voice data. \n\nNow sentences might be present in both datasets since we can have multiple audios for the same sentence. But the problem is this dataset contains the exact same audio samples from the common voice dataset samples. \nSome of them are in the train split, and some of them are in the **validation split**. So if you decide to finetune the aforementioned checkpoints and you decide to train it with a subsample of the actual set (for example 100k audios) chances are the model has already seen some of the audios in the training set and in the validation set, thus it may lead to overfitting.\n\n>Workaround :\n1. Train from scratch following the way these models were trained. Not from these checkpoints.\n2. Remove the samples present in both datasets, especially those that are present in the validation set.\n\nPlease do share your thoughts/comments!",
    "2360115": "my suggestion is to read the datset paper carefully.\n\nnote that for OOD test data, there is OOV (out of vocab).\nThis makes me wonder if the kaggle dataset i sufficient at all.\n\n(e.g. you can check the word histogram of train mps and exmples wave)\n\ni think mopst people will train LM word prediction for correcting CTC charcter-wise prediction.\ni think we need another train set for LM",
    "2362276": "Yeah, there are some larger text corpora that can be used for LM training, like the one they [used here](https://huggingface.co/arijitx/wav2vec2-xls-r-300m-bengali). But this would lead to slower inference tho. Another point is for OOD vocabs, no one knows how large vocab is sufficient. And even LM can't solve some issues like handling named enitites/similar words",
    "2362745": "wtf, how come an independent dataset contains data from others? I don't think they mention this in their paper. I mean, is this even morally correct? Hope hosts can clarify on this.",
    "2379120": "Yes its clearly mentioned in the paper! We crowdsourced data on the common voice platform through online campaigns and then curated it automatically+manually to get the training set of this dataset. If you dig deeper you'll see that most of the public models are partially trained on common voice, not the 1500+ hrs.",
    "2379186": "Thanks for the clarification. I initially thought that common voice is a not-updating dataset. Now I understand that it's also a platform and updating over time. So the data collected form  campaigns has already been included in Common Voice Corpus 12.0 12/15/2022, right?",
    "2379235": "Yes! We have included stuff from upto Common Voice 13 (not all of it) in this training set. We should have a list of common voice -> training mapping that we can share."
  },
  "source": "meta"
}