{
  "id": 426605,
  "title": "CV vs LB - how is it?",
  "url": "/competitions/bengaliai-speech/discussion/426605",
  "author_name": "",
  "post_date": "2023-07-24T10:45:12.358391300Z",
  "votes": 8,
  "comment_count": 19,
  "views": 0,
  "content": "<p>How is your CV compare to LB?<br>\nDo you use the \"default\" train/val split?</p>",
  "messages": [
    {
      "id": "2356722",
      "postDate": "07/24/2023 10:45:12",
      "content": "<p>How is your CV compare to LB?<br>\nDo you use the \"default\" train/val split?</p>",
      "rawMarkdown": "How is your CV compare to LB?\nDo you use the \"default\" train/val split?",
      "votes": null
    },
    {
      "id": "2356949",
      "postDate": "07/24/2023 12:57:58",
      "content": "<p>see<br>\n<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425496#2356341\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425496#2356341</a></p>",
      "rawMarkdown": "see\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/425496#2356341",
      "votes": null
    },
    {
      "id": "2356961",
      "postDate": "07/24/2023 13:11:32",
      "content": "<p>it may be more importnat to check this</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F390e29106b1d53147e9a0868da08bfa7%2FSelection_999(2801).png?generation=1690204289638959&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "it may be more importnat to check this\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F390e29106b1d53147e9a0868da08bfa7%2FSelection_999(2801).png?generation=1690204289638959&alt=media)",
      "votes": null
    },
    {
      "id": "2357881",
      "postDate": "07/25/2023 07:25:13",
      "content": "<p>That's neat! Thank you!<br>\nSo someone has transcribed the given examples I guess?</p>",
      "rawMarkdown": "That's neat! Thank you!\nSo someone has transcribed the given examples I guess?",
      "votes": null
    },
    {
      "id": "2357964",
      "postDate": "07/25/2023 08:22:37",
      "content": "<p>Yeah, I have done it <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">here</a>. It can be used as a pseudo-label, but don't completely rely on the numerical results generated using these transcriptions.</p>",
      "rawMarkdown": "Yeah, I have done it [here](https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932). It can be used as a pseudo-label, but don't completely rely on the numerical results generated using these transcriptions.",
      "votes": null
    },
    {
      "id": "2357979",
      "postDate": "07/25/2023 08:38:10",
      "content": "<p>Of course, it's very low number of transcripts, but it's much much better than nothing, it can tell you if something is very off.<br>\nSo thank you very much for this!:)</p>",
      "rawMarkdown": "Of course, it's very low number of transcripts, but it's much much better than nothing, it can tell you if something is very off.\nSo thank you very much for this!:)",
      "votes": null
    },
    {
      "id": "2357987",
      "postDate": "07/25/2023 08:45:58",
      "content": "<p>Well, thank you too. Hope this helps!</p>",
      "rawMarkdown": "Well, thank you too. Hope this helps!",
      "votes": null
    },
    {
      "id": "2358147",
      "postDate": "07/25/2023 10:10:40",
      "content": "<p>anyone has good results (better than yellow-king model) using this kaggle training set?</p>",
      "rawMarkdown": "anyone has good results (better than yellow-king model) using this kaggle training set?",
      "votes": null
    },
    {
      "id": "2358470",
      "postDate": "07/25/2023 14:23:21",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  How long does it take for you to train an epoch?</p>",
      "rawMarkdown": "hengck23  How long does it take for you to train an epoch?",
      "votes": null
    },
    {
      "id": "2360893",
      "postDate": "07/27/2023 04:45:40",
      "content": "<p>Since this competition is aim to test mdel's out-of-distribution performance, I believe we should use GroupKFold, however, there's no domain column in the dataset</p>",
      "rawMarkdown": "Since this competition is aim to test mdel's out-of-distribution performance, I believe we should use GroupKFold, however, there's no domain column in the dataset",
      "votes": null
    },
    {
      "id": "2361110",
      "postDate": "07/27/2023 07:29:13",
      "content": "<p>I've only listened a couple of the samples, they all seem like someone reading a text or something like that.<br>\nIt would be good to see though, how well CV correlates with LB.</p>",
      "rawMarkdown": "I've only listened a couple of the samples, they all seem like someone reading a text or something like that.\nIt would be good to see though, how well CV correlates with LB.",
      "votes": null
    },
    {
      "id": "2361114",
      "postDate": "07/27/2023 07:33:45",
      "content": "<p>Yeah. They collected most of the audios using this process, where given a sentence, a person would read this out. <br>\nThese are in the train fold.</p>\n<p>The test set contains audios from TV channels, songs, slangs etc.</p>",
      "rawMarkdown": "Yeah. They collected most of the audios using this process, where given a sentence, a person would read this out. \nThese are in the train fold.\n\nThe test set contains audios from TV channels, songs, slangs etc.",
      "votes": null
    },
    {
      "id": "2361341",
      "postDate": "07/27/2023 10:53:39",
      "content": "<p>Sounds like that we need to collect the validation set according to the examples by ourselves:)</p>",
      "rawMarkdown": "Sounds like that we need to collect the validation set according to the examples by ourselves:)",
      "votes": null
    },
    {
      "id": "2361587",
      "postDate": "07/27/2023 13:10:13",
      "content": "<blockquote>\n  <blockquote>\n    <p>Sounds like that we need to collect the validation set according to the examples by ourselves:)</p>\n  </blockquote>\n</blockquote>\n<p>This is what i recommend too.<br>\n(but collecting data for ASR is time consuming)</p>\n<p>You should also read papers that  discuss about out-of-domain ASR.<br>\ne.g. </p>\n<ul>\n<li>how to get good results if you use or do not use out-of-domain in trainining.</li>\n<li>how to estimate/analyse results for  out-of-domain evaluation.</li>\n</ul>",
      "rawMarkdown": ">>Sounds like that we need to collect the validation set according to the examples by ourselves:)\n\nThis is what i recommend too.\n(but collecting data for ASR is time consuming)\n\nYou should also read papers that  discuss about out-of-domain ASR.\ne.g. \n- how to get good results if you use or do not use out-of-domain in trainining.\n- how to estimate/analyse results for  out-of-domain evaluation.",
      "votes": null
    },
    {
      "id": "2362665",
      "postDate": "07/28/2023 08:21:26",
      "content": "<p>this is from the dataset paper:</p>\n<p>We machine-validate the un-validate test dataset using Google ASR model and another model (wav2vec 2.0 [13]) and exclude blank recordings. We manually validated samples that had at least one full word difference between wav2vec 2.0 and ground truth. After this we finalize the machine-validated data.</p>\n<hr>\n<p>wav2vec 2.0 model[13] refers to yellow-king model.</p>\n<p>so if your model consistently perform better  wav2vec 2.0 model[13] and Google ASR , you can be sure that there will be no shakeup in hidden private test set (even though it is OOD). THis is the trick in winning in this competition</p>",
      "rawMarkdown": "this is from the dataset paper:\n\nWe machine-validate the un-validate test dataset using Google ASR model and another model (wav2vec 2.0 [13]) and exclude blank recordings. We manually validated samples that had at least one full word difference between wav2vec 2.0 and ground truth. After this we finalize the machine-validated data.\n\n---\n\nwav2vec 2.0 model[13] refers to yellow-king model.\n\nso if your model consistently perform better  wav2vec 2.0 model[13] and Google ASR , you can be sure that there will be no shakeup in hidden private test set (even though it is OOD). THis is the trick in winning in this competition",
      "votes": null
    },
    {
      "id": "2362759",
      "postDate": "07/28/2023 09:39:57",
      "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Please refer this article, I find out this useful</p>",
      "rawMarkdown": "nofreewill Please refer this article, I find out this useful",
      "votes": null
    },
    {
      "id": "2363849",
      "postDate": "07/29/2023 02:21:51",
      "content": "<p>Thanks for sharing. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, let me check the resources that you mentioned</p>",
      "rawMarkdown": "Thanks for sharing. @hengck23, let me check the resources that you mentioned",
      "votes": null
    },
    {
      "id": "2363854",
      "postDate": "07/29/2023 02:30:36",
      "content": "<p>I finetuned the YellowKing-model with 30k new data and 10k validation.<br>\nCV : 0.561 on the new 10k validation set.<br>\nLB : 0.487 </p>",
      "rawMarkdown": "I finetuned the YellowKing-model with 30k new data and 10k validation.\nCV : 0.561 on the new 10k validation set.\nLB : 0.487",
      "votes": null
    },
    {
      "id": "2363994",
      "postDate": "07/29/2023 04:17:44",
      "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> <br>\nThanks for the information.</p>\n<p>I have been studying wave2vec paper,<br>\n<a href=\"https://arxiv.org/abs/2006.11477\" target=\"_blank\">https://arxiv.org/abs/2006.11477</a></p>\n<p>i think Table 9 can be a good guide on performance versus amount of data.<br>\n(but do note that there is language difference, bn verus en and domain difference)<br>\n\"Table 9: WER on the Librispeech dev/test sets when training on the Libri-light low-resource labeled<br>\ndata setups (cf. Table 1).\"</p>\n<p>for domain experiments, i recommend this: <a href=\"https://arxiv.org/abs/2104.01027\" target=\"_blank\">https://arxiv.org/abs/2104.01027</a></p>",
      "rawMarkdown": "mbmmurad \nThanks for the information.\n\nI have been studying wave2vec paper,\nhttps://arxiv.org/abs/2006.11477\n\ni think Table 9 can be a good guide on performance versus amount of data.\n(but do note that there is language difference, bn verus en and domain difference)\n\"Table 9: WER on the Librispeech dev/test sets when training on the Libri-light low-resource labeled\ndata setups (cf. Table 1).\"\n\nfor domain experiments, i recommend this: https://arxiv.org/abs/2104.01027",
      "votes": null
    },
    {
      "id": "2403238",
      "postDate": "08/22/2023 15:15:24",
      "content": "<p>2023-08-22:<br>\nCV: 0.284<br>\nLB: 0.470<br>\nStill need lots of work.</p>",
      "rawMarkdown": "2023-08-22:\nCV: 0.284\nLB: 0.470\nStill need lots of work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2356949,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2023 12:57:58",
      "content": "<p>see<br>\n<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425496#2356341\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425496#2356341</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2356961,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/24/2023 13:11:32",
          "content": "<p>it may be more importnat to check this</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F390e29106b1d53147e9a0868da08bfa7%2FSelection_999(2801).png?generation=1690204289638959&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2357881,
              "author_name": "nofreewill",
              "author_url": "",
              "post_date": "07/25/2023 07:25:13",
              "content": "<p>That's neat! Thank you!<br>\nSo someone has transcribed the given examples I guess?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2357964,
                  "author_name": "mbmmurad",
                  "author_url": "",
                  "post_date": "07/25/2023 08:22:37",
                  "content": "<p>Yeah, I have done it <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">here</a>. It can be used as a pseudo-label, but don't completely rely on the numerical results generated using these transcriptions.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2357979,
                      "author_name": "nofreewill",
                      "author_url": "",
                      "post_date": "07/25/2023 08:38:10",
                      "content": "<p>Of course, it's very low number of transcripts, but it's much much better than nothing, it can tell you if something is very off.<br>\nSo thank you very much for this!:)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2357987,
                          "author_name": "mbmmurad",
                          "author_url": "",
                          "post_date": "07/25/2023 08:45:58",
                          "content": "<p>Well, thank you too. Hope this helps!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2358147,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2023 10:10:40",
      "content": "<p>anyone has good results (better than yellow-king model) using this kaggle training set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2358470,
      "author_name": "arunodhayan",
      "author_url": "",
      "post_date": "07/25/2023 14:23:21",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  How long does it take for you to train an epoch?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2360893,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "07/27/2023 04:45:40",
      "content": "<p>Since this competition is aim to test mdel's out-of-distribution performance, I believe we should use GroupKFold, however, there's no domain column in the dataset</p>",
      "votes": null,
      "replies": [
        {
          "id": 2361110,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "07/27/2023 07:29:13",
          "content": "<p>I've only listened a couple of the samples, they all seem like someone reading a text or something like that.<br>\nIt would be good to see though, how well CV correlates with LB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2361114,
              "author_name": "mbmmurad",
              "author_url": "",
              "post_date": "07/27/2023 07:33:45",
              "content": "<p>Yeah. They collected most of the audios using this process, where given a sentence, a person would read this out. <br>\nThese are in the train fold.</p>\n<p>The test set contains audios from TV channels, songs, slangs etc.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2361341,
                  "author_name": "wuwenmin",
                  "author_url": "",
                  "post_date": "07/27/2023 10:53:39",
                  "content": "<p>Sounds like that we need to collect the validation set according to the examples by ourselves:)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2361587,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "07/27/2023 13:10:13",
                      "content": "<blockquote>\n  <blockquote>\n    <p>Sounds like that we need to collect the validation set according to the examples by ourselves:)</p>\n  </blockquote>\n</blockquote>\n<p>This is what i recommend too.<br>\n(but collecting data for ASR is time consuming)</p>\n<p>You should also read papers that  discuss about out-of-domain ASR.<br>\ne.g. </p>\n<ul>\n<li>how to get good results if you use or do not use out-of-domain in trainining.</li>\n<li>how to estimate/analyse results for  out-of-domain evaluation.</li>\n</ul>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2362665,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "07/28/2023 08:21:26",
                          "content": "<p>this is from the dataset paper:</p>\n<p>We machine-validate the un-validate test dataset using Google ASR model and another model (wav2vec 2.0 [13]) and exclude blank recordings. We manually validated samples that had at least one full word difference between wav2vec 2.0 and ground truth. After this we finalize the machine-validated data.</p>\n<hr>\n<p>wav2vec 2.0 model[13] refers to yellow-king model.</p>\n<p>so if your model consistently perform better  wav2vec 2.0 model[13] and Google ASR , you can be sure that there will be no shakeup in hidden private test set (even though it is OOD). THis is the trick in winning in this competition</p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 2363849,
                          "author_name": "wuwenmin",
                          "author_url": "",
                          "post_date": "07/29/2023 02:21:51",
                          "content": "<p>Thanks for sharing. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, let me check the resources that you mentioned</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2362759,
      "author_name": "dhirajmwagh1111",
      "author_url": "",
      "post_date": "07/28/2023 09:39:57",
      "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Please refer this article, I find out this useful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2363854,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "07/29/2023 02:30:36",
      "content": "<p>I finetuned the YellowKing-model with 30k new data and 10k validation.<br>\nCV : 0.561 on the new 10k validation set.<br>\nLB : 0.487 </p>",
      "votes": null,
      "replies": [
        {
          "id": 2363994,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/29/2023 04:17:44",
          "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> <br>\nThanks for the information.</p>\n<p>I have been studying wave2vec paper,<br>\n<a href=\"https://arxiv.org/abs/2006.11477\" target=\"_blank\">https://arxiv.org/abs/2006.11477</a></p>\n<p>i think Table 9 can be a good guide on performance versus amount of data.<br>\n(but do note that there is language difference, bn verus en and domain difference)<br>\n\"Table 9: WER on the Librispeech dev/test sets when training on the Libri-light low-resource labeled<br>\ndata setups (cf. Table 1).\"</p>\n<p>for domain experiments, i recommend this: <a href=\"https://arxiv.org/abs/2104.01027\" target=\"_blank\">https://arxiv.org/abs/2104.01027</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2403238,
      "author_name": "berserker408",
      "author_url": "",
      "post_date": "08/22/2023 15:15:24",
      "content": "<p>2023-08-22:<br>\nCV: 0.284<br>\nLB: 0.470<br>\nStill need lots of work.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2356722": "How is your CV compare to LB?\nDo you use the \"default\" train/val split?",
    "2356949": "see\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/425496#2356341",
    "2356961": "it may be more importnat to check this\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F390e29106b1d53147e9a0868da08bfa7%2FSelection_999(2801).png?generation=1690204289638959&alt=media)",
    "2357881": "That's neat! Thank you!\nSo someone has transcribed the given examples I guess?",
    "2357964": "Yeah, I have done it [here](https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932). It can be used as a pseudo-label, but don't completely rely on the numerical results generated using these transcriptions.",
    "2357979": "Of course, it's very low number of transcripts, but it's much much better than nothing, it can tell you if something is very off.\nSo thank you very much for this!:)",
    "2357987": "Well, thank you too. Hope this helps!",
    "2358147": "anyone has good results (better than yellow-king model) using this kaggle training set?",
    "2358470": "hengck23  How long does it take for you to train an epoch?",
    "2360893": "Since this competition is aim to test mdel's out-of-distribution performance, I believe we should use GroupKFold, however, there's no domain column in the dataset",
    "2361110": "I've only listened a couple of the samples, they all seem like someone reading a text or something like that.\nIt would be good to see though, how well CV correlates with LB.",
    "2361114": "Yeah. They collected most of the audios using this process, where given a sentence, a person would read this out. \nThese are in the train fold.\n\nThe test set contains audios from TV channels, songs, slangs etc.",
    "2361341": "Sounds like that we need to collect the validation set according to the examples by ourselves:)",
    "2361587": ">>Sounds like that we need to collect the validation set according to the examples by ourselves:)\n\nThis is what i recommend too.\n(but collecting data for ASR is time consuming)\n\nYou should also read papers that  discuss about out-of-domain ASR.\ne.g. \n- how to get good results if you use or do not use out-of-domain in trainining.\n- how to estimate/analyse results for  out-of-domain evaluation.",
    "2362665": "this is from the dataset paper:\n\nWe machine-validate the un-validate test dataset using Google ASR model and another model (wav2vec 2.0 [13]) and exclude blank recordings. We manually validated samples that had at least one full word difference between wav2vec 2.0 and ground truth. After this we finalize the machine-validated data.\n\n---\n\nwav2vec 2.0 model[13] refers to yellow-king model.\n\nso if your model consistently perform better  wav2vec 2.0 model[13] and Google ASR , you can be sure that there will be no shakeup in hidden private test set (even though it is OOD). THis is the trick in winning in this competition",
    "2362759": "nofreewill Please refer this article, I find out this useful",
    "2363849": "Thanks for sharing. @hengck23, let me check the resources that you mentioned",
    "2363854": "I finetuned the YellowKing-model with 30k new data and 10k validation.\nCV : 0.561 on the new 10k validation set.\nLB : 0.487",
    "2363994": "mbmmurad \nThanks for the information.\n\nI have been studying wave2vec paper,\nhttps://arxiv.org/abs/2006.11477\n\ni think Table 9 can be a good guide on performance versus amount of data.\n(but do note that there is language difference, bn verus en and domain difference)\n\"Table 9: WER on the Librispeech dev/test sets when training on the Libri-light low-resource labeled\ndata setups (cf. Table 1).\"\n\nfor domain experiments, i recommend this: https://arxiv.org/abs/2104.01027",
    "2403238": "2023-08-22:\nCV: 0.284\nLB: 0.470\nStill need lots of work."
  },
  "source": "meta"
}