{
  "id": 425942,
  "title": "help! why submission score error???",
  "url": "/competitions/bengaliai-speech/discussion/425942",
  "author_name": "",
  "post_date": "2023-07-21T05:12:43.357016400Z",
  "votes": 4,
  "comment_count": 25,
  "views": 0,
  "content": "<p>This is the notebook<br>\n<a href=\"https://www.kaggle.com/hengck23/help-why-submission-score-error\" target=\"_blank\">https://www.kaggle.com/hengck23/help-why-submission-score-error</a><br>\nusing public huggingface/ai4bharat/indicwav2vec_v1_bengali model<br>\n(this model has good local CV of 0.39)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19ab286f9fc5c62bfac6fb18cdce5ffd%2FSelection_999(2762).png?generation=1689916005019800&amp;alt=media\" alt=\"\"><br>\n\" submission score error\" occurs when the evaluation script at the server uses your submission csv cannot complete due to error.<br>\ni find it strange:</p>\n<ul>\n<li>my notebook can submit correctly if i can to another model</li>\n<li>i have made the following check:</li>\n</ul>\n<pre><code>assert (sample_submission_df==submit_df)\n  to string\nsubmit_df = submit_df(lambda x: x  (x)==str  ) \n</code></pre>\n<ul>\n<li>my code is ok, with jiwer.wer at local machine, tested over 20k kaggle mp3s.</li>\n</ul>\n<h2>Can anyone suggest what is wrong? Thanks!</h2>\n<p>submission score error :<br>\n\"Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips\"</p>",
  "messages": [
    {
      "id": "2352546",
      "postDate": "07/21/2023 05:12:43",
      "content": "<p>This is the notebook<br>\n<a href=\"https://www.kaggle.com/hengck23/help-why-submission-score-error\" target=\"_blank\">https://www.kaggle.com/hengck23/help-why-submission-score-error</a><br>\nusing public huggingface/ai4bharat/indicwav2vec_v1_bengali model<br>\n(this model has good local CV of 0.39)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19ab286f9fc5c62bfac6fb18cdce5ffd%2FSelection_999(2762).png?generation=1689916005019800&amp;alt=media\" alt=\"\"><br>\n\" submission score error\" occurs when the evaluation script at the server uses your submission csv cannot complete due to error.<br>\ni find it strange:</p>\n<ul>\n<li>my notebook can submit correctly if i can to another model</li>\n<li>i have made the following check:</li>\n</ul>\n<pre><code>assert (sample_submission_df==submit_df)\n  to string\nsubmit_df = submit_df(lambda x: x  (x)==str  ) \n</code></pre>\n<ul>\n<li>my code is ok, with jiwer.wer at local machine, tested over 20k kaggle mp3s.</li>\n</ul>\n<h2>Can anyone suggest what is wrong? Thanks!</h2>\n<p>submission score error :<br>\n\"Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips\"</p>",
      "rawMarkdown": "This is the notebook\nhttps://www.kaggle.com/hengck23/help-why-submission-score-error\n\nusing public huggingface/ai4bharat/indicwav2vec_v1_bengali model\n(this model has good local CV of 0.39)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19ab286f9fc5c62bfac6fb18cdce5ffd%2FSelection_999(2762).png?generation=1689916005019800&alt=media)\n\n\n\" submission score error\" occurs when the evaluation script at the server uses your submission csv cannot complete due to error.\n\ni find it strange:\n- my notebook can submit correctly if i can to another model\n- i have made the following check:\n```\nassert (sample_submission_df['id']==submit_df['id'])\n\n#force all to string\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(lambda x: x if type(x)==str else '') \n```\n\n- my code is ok, with jiwer.wer at local machine, tested over 20k kaggle mp3s.\n\nCan anyone suggest what is wrong? Thanks!\n\n----\n\nsubmission score error :\n\"Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips\"",
      "votes": null
    },
    {
      "id": "2352727",
      "postDate": "07/21/2023 08:16:51",
      "content": "<p>Have same issue(</p>",
      "rawMarkdown": "Have same issue(",
      "votes": null
    },
    {
      "id": "2352900",
      "postDate": "07/21/2023 10:33:06",
      "content": "<p></p>",
      "rawMarkdown": "~~  ```\nif mode=='submit':\n         mp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n         valid_df = pd.read_csv('/kaggle/input/bengaliai-speech/sample_submission.csv') \n```\n\n\nYou are loading sample_submission.csv and then infering on the ids of this dataframe. It contains only 3 values. But the actual test set has around 8k.\n\nWhat you need to do is iterate over the files of the 'test_mp3s' directory, store the filenames and prediction, and then create a new dataframe using these. \nPlease refer to the inference part of [this notebook ](https://www.kaggle.com/code/mbmmurad/lb-0-641-nemo-conformer-baseline-w-o-internet-in)~~",
      "votes": null
    },
    {
      "id": "2352945",
      "postDate": "07/21/2023 10:59:27",
      "content": "<p>Pseudo-code of a correct way of submission :</p>\n<pre><code>filenames = \npredictions = \nfiles = os()\nbase_path = \n\n file  files:\n    filenames(file())\n     = (base_path+file)  function to load the  file\n    text = (audio)   with your model\n     (text)==:\n          text = \n\n    predictions(text)\n\ndf = pd({:filenames,:predictions})\ndf(,index=False)\n</code></pre>\n<p>if you use a normalizer than make sure to check the texts after normalizing.</p>\n<pre><code>def check():\n     ()==:\n         = \n     \ndf. = df..apply(lambda x:normalizer(x)))\ndf. = df..apply(lambda x:check(x)))\n</code></pre>",
      "rawMarkdown": "Pseudo-code of a correct way of submission :\n```\nfilenames = []\npredictions = []\nfiles = os.listdir('/kaggle/input/bengaliai-speech/test_mp3s')\nbase_path = \"/kaggle/input/bengaliai-speech/test_mp3s/\"\n\nfor file in files:\n    filenames.append(file.split(\".\")[0])\n    audio = load_audio(base_path+file) #Some function to load the audio file\n    text = model(audio)  #Infer with your model\n    if len(text)==0:\n          text = \",\"\n\n    predictions.append(text)\n\ndf = pd.DataFrame({'id':filenames,'sentence':predictions})\ndf.to_csv('submission.csv',index=False)\n\n```\n\nif you use a normalizer than make sure to check the texts after normalizing.\n```\ndef check(sentence):\n    if len(sentence)==0:\n        sentence = \",\"\n    return sentence\ndf.sentence = df.sentence.apply(lambda x:normalizer(x)))\ndf.sentence = df.sentence.apply(lambda x:check(x)))\n```",
      "votes": null
    },
    {
      "id": "2353004",
      "postDate": "07/21/2023 11:46:54",
      "content": "<p>i think both test folder and sample_submission.csv are replaced during submission.<br>\nMy code work for other model, so i think it is not the reason.</p>\n<p>besides, if i were only processing three files, the submission time should be very short</p>",
      "rawMarkdown": "i think both test folder and sample\\_submission.csv are replaced during submission.\nMy code work for other model, so i think it is not the reason.\n\nbesides, if i were only processing three files, the submission time should be very short",
      "votes": null
    },
    {
      "id": "2353253",
      "postDate": "07/21/2023 14:34:54",
      "content": "<p>the latest version still course error:</p>\n<h1>use file list from test folder</h1>\n<pre><code> mode==:\n    mp3_dir = \n    glob_file = glob()\n     = ([f[(mp3_dir)+:-]  f  glob_file])\n    valid_df = pd.DataFrame({:,:})\n</code></pre>\n<h1>check each prediction can be process by jiwer</h1>\n<pre><code>def check_text():\n    t =  \n    \n    :\n        jiwer.wer(t, [])\n    except:\n         =  \n     \n\n\nsubmit_df.loc[:,] = submit_df..apply(check_text)\n</code></pre>",
      "rawMarkdown": "the latest version still course error:\n\n\n#use file list from test folder\n```\nif mode=='submit':\n    mp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n    glob_file = glob(f'{mp3_dir}/*.mp3')\n    id = sorted([f[len(mp3_dir)+1:-4] for f in glob_file])\n    valid_df = pd.DataFrame({'id':id,'sentence':''})\n```\n\n\n#check each prediction can be process by jiwer\n```\ndef check_text(sentence):\n    t = 'িনি এবং ও এই করে া। তার' #fake truth\n    #print(sentence)\n    try:\n        jiwer.wer(t, [sentence])\n    except:\n        sentence = '' #t\n    return sentence\n\n\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(check_text)\n```",
      "votes": null
    },
    {
      "id": "2353304",
      "postDate": "07/21/2023 15:20:04",
      "content": "<h1>Solved LB 0.52</h1>\n<p>I successfully submitted using your notebook. The main problem is :</p>\n<blockquote>\n  <p>The model is generating some empty predictions. It seems like you cant't submit empty strings in the dataframe. I replaced the empty strings with \",\" and it got submitted. </p>\n</blockquote>\n<pre><code>filenames = \npredictions = \nfiles = os()\nbase_path = \n\n file  files:\n    filenames(file())\n     = (base_path+file)  function to load the  file\n    text = (audio)   with your model\n     (text)==:\n          text = \n\n    predictions(text)\n\ndf = pd({:filenames,:predictions})\ndf(,index=False)\n</code></pre>\n<p>Refer to <a href=\"https://www.kaggle.com/mbmmurad/fork-of-fork-of-local-nemo-baseline-conformer\" target=\"_blank\">this notebook</a></p>",
      "rawMarkdown": "# Solved LB 0.52\n\nI successfully submitted using your notebook. The main problem is :\n> The model is generating some empty predictions. It seems like you cant't submit empty strings in the dataframe. I replaced the empty strings with \",\" and it got submitted. \n\n``` \nfilenames = []\npredictions = []\nfiles = os.listdir('/kaggle/input/bengaliai-speech/test_mp3s')\nbase_path = \"/kaggle/input/bengaliai-speech/test_mp3s/\"\n\nfor file in files:\n    filenames.append(file.split(\".\")[0])\n    audio = load_audio(base_path+file) #Some function to load the audio file\n    text = model(audio)  #Infer with your model\n    if len(text)==0:\n          text = \",\"\n    \n    predictions.append(text)\n\ndf = pd.DataFrame({'id':filenames,'sentence':predictions})\ndf.to_csv('submission.csv',index=False)\n```\n\nRefer to [this notebook](https://www.kaggle.com/mbmmurad/fork-of-fork-of-local-nemo-baseline-conformer)",
      "votes": null
    },
    {
      "id": "2353306",
      "postDate": "07/21/2023 15:21:12",
      "content": "<p>You're right. My mistake. Your pipeline worked. Just needed to handle the empty strings.</p>",
      "rawMarkdown": "You're right. My mistake. Your pipeline worked. Just needed to handle the empty strings.",
      "votes": null
    },
    {
      "id": "2353310",
      "postDate": "07/21/2023 15:22:15",
      "content": "<p>Just change to <code>sentence = \",\"</code> . </p>",
      "rawMarkdown": "Just change to ``` sentence = \",\"``` .",
      "votes": null
    },
    {
      "id": "2353348",
      "postDate": "07/21/2023 15:43:29",
      "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> </p>\n<p>Thanks a lot!!!</p>",
      "rawMarkdown": "mbmmurad \n\nThanks a lot!!!",
      "votes": null
    },
    {
      "id": "2353352",
      "postDate": "07/21/2023 15:45:00",
      "content": "<p>by the way, this is wave2vec2-CTC without LM</p>",
      "rawMarkdown": "by the way, this is wave2vec2-CTC without LM",
      "votes": null
    },
    {
      "id": "2353382",
      "postDate": "07/21/2023 16:02:55",
      "content": "<p>more information about the model can be found here:<br>\n<a href=\"https://github.com/AI4Bharat/IndicWav2Vec/tree/main\" target=\"_blank\">https://github.com/AI4Bharat/IndicWav2Vec/tree/main</a></p>\n<p>i think the model is trained with author's own dataset.<br>\nHence we can conclude? OOD cause adroup of about 0.12<br>\n(since kaggle valid split has WER = 0.40, public LB 0.52)</p>\n<hr>\n<p>IndicWav2Vec paper:<br>\n<a href=\"https://arxiv.org/pdf/2111.03945.pdf\" target=\"_blank\">https://arxiv.org/pdf/2111.03945.pdf</a></p>\n<p>\"Apart from YouTube, we also curated content<br>\nfrom newsonair3 which is a radio news channel<br>\nrun by the Govt. of India and broadcasts news in<br>\nmultiple Indian languages\"</p>\n<hr>\n<p>there is CTC-LM model which gives better results but you need to convert from fairseq to hf yourself,<br>\nfollowing instructions from githut repo</p>\n<hr>\n<p>\"3.5 Rescoring<br>\nNote that in the above decoding process, we use an<br>\nn-gram KenLM language model (Heafield, 2011).<br>\nHowever, recent advances in language modeling<br>\nhave shown that transformer based language models perform well. To get the best of both worlds,<br>\nwe optionally use an external transformer based<br>\nlanguage model to rescore the n-best hypothesis<br>\nproduced above (Synnaeve et al., 2019). Specifically, for each output in the n-best list, we compute a new score by combining the decoder logprobability, \"</p>\n<p>the rescoring transformer can be used for out ensemble of  different models</p>",
      "rawMarkdown": "more information about the model can be found here:\nhttps://github.com/AI4Bharat/IndicWav2Vec/tree/main\n\ni think the model is trained with author's own dataset.\nHence we can conclude? OOD cause adroup of about 0.12\n(since kaggle valid split has WER = 0.40, public LB 0.52)\n\n\n---\nIndicWav2Vec paper:\nhttps://arxiv.org/pdf/2111.03945.pdf\n\n\"Apart from YouTube, we also curated content\nfrom newsonair3 which is a radio news channel\nrun by the Govt. of India and broadcasts news in\nmultiple Indian languages\"\n\n---\nthere is CTC-LM model which gives better results but you need to convert from fairseq to hf yourself,\nfollowing instructions from githut repo\n\n\n---\n\n\"3.5 Rescoring\nNote that in the above decoding process, we use an\nn-gram KenLM language model (Heafield, 2011).\nHowever, recent advances in language modeling\nhave shown that transformer based language models perform well. To get the best of both worlds,\nwe optionally use an external transformer based\nlanguage model to rescore the n-best hypothesis\nproduced above (Synnaeve et al., 2019). Specifically, for each output in the n-best list, we compute a new score by combining the decoder logprobability, \"\n\nthe rescoring transformer can be used for out ensemble of  different models",
      "votes": null
    },
    {
      "id": "2353407",
      "postDate": "07/21/2023 16:17:50",
      "content": "<p>Yes the model was trained on their own dataset. The model without LM scored WER 0.166 on their evaluation set and with LM performed 0.136.</p>\n<p>So I guess the model with LM might score better than this one. Agreed with the drop for the OOD part! But a thing to consider is the valid split contains 30k files, with comparatively shorter sentence lengths. The public LB is evaluated on around 4k samples. So it will be interesting to see how we make a decision from the LB-CV difference.</p>",
      "rawMarkdown": "Yes the model was trained on their own dataset. The model without LM scored WER 0.166 on their evaluation set and with LM performed 0.136.\n\nSo I guess the model with LM might score better than this one. Agreed with the drop for the OOD part! But a thing to consider is the valid split contains 30k files, with comparatively shorter sentence lengths. The public LB is evaluated on around 4k samples. So it will be interesting to see how we make a decision from the LB-CV difference.",
      "votes": null
    },
    {
      "id": "2353451",
      "postDate": "07/21/2023 17:03:21",
      "content": "<p>considering that indian language is close to bengali, this model is a good starting source. they clamied their training data is public … i am trying to find them now</p>",
      "rawMarkdown": "considering that indian language is close to bengali, this model is a good starting source. they clamied their training data is public ... i am trying to find them now",
      "votes": null
    },
    {
      "id": "2353470",
      "postDate": "07/21/2023 17:11:55",
      "content": "<p>Seems so. Good luck!</p>",
      "rawMarkdown": "Seems so. Good luck!",
      "votes": null
    },
    {
      "id": "2353493",
      "postDate": "07/21/2023 17:26:22",
      "content": "<p>maybe here!!!<br>\n<a href=\"https://ai4bharat.iitm.ac.in/dhwani\" target=\"_blank\">https://ai4bharat.iitm.ac.in/dhwani</a></p>\n<p>Dataset Format<br>\nThe audio files present in separate folders.<br>\nFor YouTubeThe audio filenames are named YouTube-ids and for Newsonair, the contatination of region name, bulletin timing makes the filename.</p>\n<hr>\n<p>more here:<br>\n<a href=\"https://ai4bharat.iitm.ac.in/datasets\" target=\"_blank\">https://ai4bharat.iitm.ac.in/datasets</a></p>\n<p>conformer model demo<br>\n<a href=\"https://models.ai4bharat.org/#/asr/conformer\" target=\"_blank\">https://models.ai4bharat.org/#/asr/conformer</a></p>",
      "rawMarkdown": "maybe here!!!\nhttps://ai4bharat.iitm.ac.in/dhwani\n\nDataset Format\nThe audio files present in separate folders.\nFor YouTubeThe audio filenames are named YouTube-ids and for Newsonair, the contatination of region name, bulletin timing makes the filename.\n\n---\n\nmore here:\nhttps://ai4bharat.iitm.ac.in/datasets\n\n\nconformer model demo\nhttps://models.ai4bharat.org/#/asr/conformer",
      "votes": null
    },
    {
      "id": "2362376",
      "postDate": "07/28/2023 04:01:56",
      "content": "<p>Have the same issue. Is here a feasible solution for this error provided?<br>\nIs that </p>\n<pre><code> ():\n    t =  \n    \n    :\n        jiwer.wer(t, [sentence])\n    :\n        sentence =  \n     sentence\n\n\nsubmit_df.loc[:,] = submit_df.sentence.apply(check_text)\n</code></pre>",
      "rawMarkdown": "Have the same issue. Is here a feasible solution for this error provided?\nIs that \n```py\ndef check_text(sentence):\n    t = 'িনি এবং ও এই করে া। তার' #fake truth\n    #print(sentence)\n    try:\n        jiwer.wer(t, [sentence])\n    except:\n        sentence = ',' #t\n    return sentence\n\n\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(check_text)\n```",
      "votes": null
    },
    {
      "id": "2362384",
      "postDate": "07/28/2023 04:06:32",
      "content": "<pre><code>def check():\n     ()==:\n         = \n     \ndf. = df..apply(lambda x: check(x))\n</code></pre>\n<p>This should work. The problem is your model can sometimes generate empty strings, and when you convert this empty strings to a dataframe, it appears as nan. <br>\nSo we're just handling the empty strings by this function, replacing them with a character so that no empty strings remain in the predictions </p>",
      "rawMarkdown": "~~~\ndef check(sentence):\n    if len(sentence)==0:\n        sentence = \",\"\n    return sentence\ndf.sentence = df.sentence.apply(lambda x: check(x))\n~~~\nThis should work. The problem is your model can sometimes generate empty strings, and when you convert this empty strings to a dataframe, it appears as nan. \nSo we're just handling the empty strings by this function, replacing them with a character so that no empty strings remain in the predictions",
      "votes": null
    },
    {
      "id": "2362390",
      "postDate": "07/28/2023 04:12:28",
      "content": "<p>Is this scenario referring to that the whole prediction for this audio is empty, or there are spaces in the predicted sentences?</p>",
      "rawMarkdown": "Is this scenario referring to that the whole prediction for this audio is empty, or there are spaces in the predicted sentences?",
      "votes": null
    },
    {
      "id": "2362394",
      "postDate": "07/28/2023 04:22:33",
      "content": "<p>For the whole audio. There are some clips in the training/validation set that are 1-2 sec long, consisting of only 2-3 words. Now if the voice Isn't loud enough/audio contain more external noise, the model might fail to even extract a single word/character, thus returning an empty string for the whole audio.</p>",
      "rawMarkdown": "For the whole audio. There are some clips in the training/validation set that are 1-2 sec long, consisting of only 2-3 words. Now if the voice Isn't loud enough/audio contain more external noise, the model might fail to even extract a single word/character, thus returning an empty string for the whole audio.",
      "votes": null
    },
    {
      "id": "2362398",
      "postDate": "07/28/2023 04:27:20",
      "content": "<p>When I submit my notebook and prediction, will the model also predict for audios in training and validating dataset, or some private dataset unseen so that the problem could occur? (I'm somewhat novice, not understanding the competition mechanism)</p>",
      "rawMarkdown": "When I submit my notebook and prediction, will the model also predict for audios in training and validating dataset, or some private dataset unseen so that the problem could occur? (I'm somewhat novice, not understanding the competition mechanism)",
      "votes": null
    },
    {
      "id": "2362403",
      "postDate": "07/28/2023 04:37:05",
      "content": "<p>That's completely fine. You only need to predict for the 'test_mp3s' folder. This folder contains unseen audios (OOD audios), you can only see three of them now. But when you submit your notebook, your notebook will have access to all of the unseen audios( around 8000) </p>\n<p>You can have a look at <a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model\" target=\"_blank\">my notebook</a> or the notebook pinned in the code section to see how it works!</p>",
      "rawMarkdown": "That's completely fine. You only need to predict for the 'test_mp3s' folder. This folder contains unseen audios (OOD audios), you can only see three of them now. But when you submit your notebook, your notebook will have access to all of the unseen audios( around 8000) \n\nYou can have a look at [my notebook](https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model) or the notebook pinned in the code section to see how it works!",
      "votes": null
    },
    {
      "id": "2362410",
      "postDate": "07/28/2023 04:46:17",
      "content": "<p>OK. Thanks a lot for your detailed explanation!🥰</p>",
      "rawMarkdown": "OK. Thanks a lot for your detailed explanation!🥰",
      "votes": null
    },
    {
      "id": "2362559",
      "postDate": "07/28/2023 06:57:12",
      "content": "<p>I'm newly in compitition i know simple basic Python,pandas and numpy knowledge so can you guide me this project?</p>",
      "rawMarkdown": "I'm newly in compitition i know simple basic Python,pandas and numpy knowledge so can you guide me this project?",
      "votes": null
    },
    {
      "id": "2362737",
      "postDate": "07/28/2023 09:17:47",
      "content": "<p>This may help (about this competition): <a href=\"https://www.kaggle.com/code/ayushs9020/understanding-the-competition-bengalai\" target=\"_blank\">https://www.kaggle.com/code/ayushs9020/understanding-the-competition-bengalai</a><br>\nAnd some lectures like Stanford CS230n may help!</p>",
      "rawMarkdown": "This may help (about this competition): https://www.kaggle.com/code/ayushs9020/understanding-the-competition-bengalai\nAnd some lectures like Stanford CS230n may help!",
      "votes": null
    },
    {
      "id": "2445056",
      "postDate": "09/18/2023 15:37:19",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> could you give some pointer or example how rescoring transformer would work for ensembling model ? </p>",
      "rawMarkdown": "hengck23 could you give some pointer or example how rescoring transformer would work for ensembling model ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2352727,
      "author_name": "dimonovez",
      "author_url": "",
      "post_date": "07/21/2023 08:16:51",
      "content": "<p>Have same issue(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2352900,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "07/21/2023 10:33:06",
      "content": "<p></p>",
      "votes": null,
      "replies": [
        {
          "id": 2353004,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/21/2023 11:46:54",
          "content": "<p>i think both test folder and sample_submission.csv are replaced during submission.<br>\nMy code work for other model, so i think it is not the reason.</p>\n<p>besides, if i were only processing three files, the submission time should be very short</p>",
          "votes": null,
          "replies": [
            {
              "id": 2353306,
              "author_name": "mbmmurad",
              "author_url": "",
              "post_date": "07/21/2023 15:21:12",
              "content": "<p>You're right. My mistake. Your pipeline worked. Just needed to handle the empty strings.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2352945,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "07/21/2023 10:59:27",
      "content": "<p>Pseudo-code of a correct way of submission :</p>\n<pre><code>filenames = \npredictions = \nfiles = os()\nbase_path = \n\n file  files:\n    filenames(file())\n     = (base_path+file)  function to load the  file\n    text = (audio)   with your model\n     (text)==:\n          text = \n\n    predictions(text)\n\ndf = pd({:filenames,:predictions})\ndf(,index=False)\n</code></pre>\n<p>if you use a normalizer than make sure to check the texts after normalizing.</p>\n<pre><code>def check():\n     ()==:\n         = \n     \ndf. = df..apply(lambda x:normalizer(x)))\ndf. = df..apply(lambda x:check(x)))\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2353253,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/21/2023 14:34:54",
      "content": "<p>the latest version still course error:</p>\n<h1>use file list from test folder</h1>\n<pre><code> mode==:\n    mp3_dir = \n    glob_file = glob()\n     = ([f[(mp3_dir)+:-]  f  glob_file])\n    valid_df = pd.DataFrame({:,:})\n</code></pre>\n<h1>check each prediction can be process by jiwer</h1>\n<pre><code>def check_text():\n    t =  \n    \n    :\n        jiwer.wer(t, [])\n    except:\n         =  \n     \n\n\nsubmit_df.loc[:,] = submit_df..apply(check_text)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2353310,
          "author_name": "mbmmurad",
          "author_url": "",
          "post_date": "07/21/2023 15:22:15",
          "content": "<p>Just change to <code>sentence = \",\"</code> . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2353304,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "07/21/2023 15:20:04",
      "content": "<h1>Solved LB 0.52</h1>\n<p>I successfully submitted using your notebook. The main problem is :</p>\n<blockquote>\n  <p>The model is generating some empty predictions. It seems like you cant't submit empty strings in the dataframe. I replaced the empty strings with \",\" and it got submitted. </p>\n</blockquote>\n<pre><code>filenames = \npredictions = \nfiles = os()\nbase_path = \n\n file  files:\n    filenames(file())\n     = (base_path+file)  function to load the  file\n    text = (audio)   with your model\n     (text)==:\n          text = \n\n    predictions(text)\n\ndf = pd({:filenames,:predictions})\ndf(,index=False)\n</code></pre>\n<p>Refer to <a href=\"https://www.kaggle.com/mbmmurad/fork-of-fork-of-local-nemo-baseline-conformer\" target=\"_blank\">this notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2353348,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/21/2023 15:43:29",
          "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> </p>\n<p>Thanks a lot!!!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2353352,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "07/21/2023 15:45:00",
              "content": "<p>by the way, this is wave2vec2-CTC without LM</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2353382,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "07/21/2023 16:02:55",
                  "content": "<p>more information about the model can be found here:<br>\n<a href=\"https://github.com/AI4Bharat/IndicWav2Vec/tree/main\" target=\"_blank\">https://github.com/AI4Bharat/IndicWav2Vec/tree/main</a></p>\n<p>i think the model is trained with author's own dataset.<br>\nHence we can conclude? OOD cause adroup of about 0.12<br>\n(since kaggle valid split has WER = 0.40, public LB 0.52)</p>\n<hr>\n<p>IndicWav2Vec paper:<br>\n<a href=\"https://arxiv.org/pdf/2111.03945.pdf\" target=\"_blank\">https://arxiv.org/pdf/2111.03945.pdf</a></p>\n<p>\"Apart from YouTube, we also curated content<br>\nfrom newsonair3 which is a radio news channel<br>\nrun by the Govt. of India and broadcasts news in<br>\nmultiple Indian languages\"</p>\n<hr>\n<p>there is CTC-LM model which gives better results but you need to convert from fairseq to hf yourself,<br>\nfollowing instructions from githut repo</p>\n<hr>\n<p>\"3.5 Rescoring<br>\nNote that in the above decoding process, we use an<br>\nn-gram KenLM language model (Heafield, 2011).<br>\nHowever, recent advances in language modeling<br>\nhave shown that transformer based language models perform well. To get the best of both worlds,<br>\nwe optionally use an external transformer based<br>\nlanguage model to rescore the n-best hypothesis<br>\nproduced above (Synnaeve et al., 2019). Specifically, for each output in the n-best list, we compute a new score by combining the decoder logprobability, \"</p>\n<p>the rescoring transformer can be used for out ensemble of  different models</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2353407,
                      "author_name": "mbmmurad",
                      "author_url": "",
                      "post_date": "07/21/2023 16:17:50",
                      "content": "<p>Yes the model was trained on their own dataset. The model without LM scored WER 0.166 on their evaluation set and with LM performed 0.136.</p>\n<p>So I guess the model with LM might score better than this one. Agreed with the drop for the OOD part! But a thing to consider is the valid split contains 30k files, with comparatively shorter sentence lengths. The public LB is evaluated on around 4k samples. So it will be interesting to see how we make a decision from the LB-CV difference.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2353451,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "07/21/2023 17:03:21",
                          "content": "<p>considering that indian language is close to bengali, this model is a good starting source. they clamied their training data is public … i am trying to find them now</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2353470,
                              "author_name": "mbmmurad",
                              "author_url": "",
                              "post_date": "07/21/2023 17:11:55",
                              "content": "<p>Seems so. Good luck!</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2353493,
                                  "author_name": "hengck23",
                                  "author_url": "",
                                  "post_date": "07/21/2023 17:26:22",
                                  "content": "<p>maybe here!!!<br>\n<a href=\"https://ai4bharat.iitm.ac.in/dhwani\" target=\"_blank\">https://ai4bharat.iitm.ac.in/dhwani</a></p>\n<p>Dataset Format<br>\nThe audio files present in separate folders.<br>\nFor YouTubeThe audio filenames are named YouTube-ids and for Newsonair, the contatination of region name, bulletin timing makes the filename.</p>\n<hr>\n<p>more here:<br>\n<a href=\"https://ai4bharat.iitm.ac.in/datasets\" target=\"_blank\">https://ai4bharat.iitm.ac.in/datasets</a></p>\n<p>conformer model demo<br>\n<a href=\"https://models.ai4bharat.org/#/asr/conformer\" target=\"_blank\">https://models.ai4bharat.org/#/asr/conformer</a></p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    },
                    {
                      "id": 2445056,
                      "author_name": "nyleve",
                      "author_url": "",
                      "post_date": "09/18/2023 15:37:19",
                      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> could you give some pointer or example how rescoring transformer would work for ensembling model ? </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2362376,
      "author_name": "nisshokuitsuki",
      "author_url": "",
      "post_date": "07/28/2023 04:01:56",
      "content": "<p>Have the same issue. Is here a feasible solution for this error provided?<br>\nIs that </p>\n<pre><code> ():\n    t =  \n    \n    :\n        jiwer.wer(t, [sentence])\n    :\n        sentence =  \n     sentence\n\n\nsubmit_df.loc[:,] = submit_df.sentence.apply(check_text)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2362384,
          "author_name": "mbmmurad",
          "author_url": "",
          "post_date": "07/28/2023 04:06:32",
          "content": "<pre><code>def check():\n     ()==:\n         = \n     \ndf. = df..apply(lambda x: check(x))\n</code></pre>\n<p>This should work. The problem is your model can sometimes generate empty strings, and when you convert this empty strings to a dataframe, it appears as nan. <br>\nSo we're just handling the empty strings by this function, replacing them with a character so that no empty strings remain in the predictions </p>",
          "votes": null,
          "replies": [
            {
              "id": 2362390,
              "author_name": "nisshokuitsuki",
              "author_url": "",
              "post_date": "07/28/2023 04:12:28",
              "content": "<p>Is this scenario referring to that the whole prediction for this audio is empty, or there are spaces in the predicted sentences?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2362394,
                  "author_name": "mbmmurad",
                  "author_url": "",
                  "post_date": "07/28/2023 04:22:33",
                  "content": "<p>For the whole audio. There are some clips in the training/validation set that are 1-2 sec long, consisting of only 2-3 words. Now if the voice Isn't loud enough/audio contain more external noise, the model might fail to even extract a single word/character, thus returning an empty string for the whole audio.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2362398,
                      "author_name": "nisshokuitsuki",
                      "author_url": "",
                      "post_date": "07/28/2023 04:27:20",
                      "content": "<p>When I submit my notebook and prediction, will the model also predict for audios in training and validating dataset, or some private dataset unseen so that the problem could occur? (I'm somewhat novice, not understanding the competition mechanism)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2362403,
                          "author_name": "mbmmurad",
                          "author_url": "",
                          "post_date": "07/28/2023 04:37:05",
                          "content": "<p>That's completely fine. You only need to predict for the 'test_mp3s' folder. This folder contains unseen audios (OOD audios), you can only see three of them now. But when you submit your notebook, your notebook will have access to all of the unseen audios( around 8000) </p>\n<p>You can have a look at <a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model\" target=\"_blank\">my notebook</a> or the notebook pinned in the code section to see how it works!</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2362410,
                              "author_name": "nisshokuitsuki",
                              "author_url": "",
                              "post_date": "07/28/2023 04:46:17",
                              "content": "<p>OK. Thanks a lot for your detailed explanation!🥰</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    },
                    {
                      "id": 2362559,
                      "author_name": "sonalisul",
                      "author_url": "",
                      "post_date": "07/28/2023 06:57:12",
                      "content": "<p>I'm newly in compitition i know simple basic Python,pandas and numpy knowledge so can you guide me this project?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2362737,
                          "author_name": "nisshokuitsuki",
                          "author_url": "",
                          "post_date": "07/28/2023 09:17:47",
                          "content": "<p>This may help (about this competition): <a href=\"https://www.kaggle.com/code/ayushs9020/understanding-the-competition-bengalai\" target=\"_blank\">https://www.kaggle.com/code/ayushs9020/understanding-the-competition-bengalai</a><br>\nAnd some lectures like Stanford CS230n may help!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2352546": "This is the notebook\nhttps://www.kaggle.com/hengck23/help-why-submission-score-error\n\nusing public huggingface/ai4bharat/indicwav2vec_v1_bengali model\n(this model has good local CV of 0.39)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19ab286f9fc5c62bfac6fb18cdce5ffd%2FSelection_999(2762).png?generation=1689916005019800&alt=media)\n\n\n\" submission score error\" occurs when the evaluation script at the server uses your submission csv cannot complete due to error.\n\ni find it strange:\n- my notebook can submit correctly if i can to another model\n- i have made the following check:\n```\nassert (sample_submission_df['id']==submit_df['id'])\n\n#force all to string\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(lambda x: x if type(x)==str else '') \n```\n\n- my code is ok, with jiwer.wer at local machine, tested over 20k kaggle mp3s.\n\nCan anyone suggest what is wrong? Thanks!\n\n----\n\nsubmission score error :\n\"Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips\"",
    "2352727": "Have same issue(",
    "2352900": "~~  ```\nif mode=='submit':\n         mp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n         valid_df = pd.read_csv('/kaggle/input/bengaliai-speech/sample_submission.csv') \n```\n\n\nYou are loading sample_submission.csv and then infering on the ids of this dataframe. It contains only 3 values. But the actual test set has around 8k.\n\nWhat you need to do is iterate over the files of the 'test_mp3s' directory, store the filenames and prediction, and then create a new dataframe using these. \nPlease refer to the inference part of [this notebook ](https://www.kaggle.com/code/mbmmurad/lb-0-641-nemo-conformer-baseline-w-o-internet-in)~~",
    "2352945": "Pseudo-code of a correct way of submission :\n```\nfilenames = []\npredictions = []\nfiles = os.listdir('/kaggle/input/bengaliai-speech/test_mp3s')\nbase_path = \"/kaggle/input/bengaliai-speech/test_mp3s/\"\n\nfor file in files:\n    filenames.append(file.split(\".\")[0])\n    audio = load_audio(base_path+file) #Some function to load the audio file\n    text = model(audio)  #Infer with your model\n    if len(text)==0:\n          text = \",\"\n\n    predictions.append(text)\n\ndf = pd.DataFrame({'id':filenames,'sentence':predictions})\ndf.to_csv('submission.csv',index=False)\n\n```\n\nif you use a normalizer than make sure to check the texts after normalizing.\n```\ndef check(sentence):\n    if len(sentence)==0:\n        sentence = \",\"\n    return sentence\ndf.sentence = df.sentence.apply(lambda x:normalizer(x)))\ndf.sentence = df.sentence.apply(lambda x:check(x)))\n```",
    "2353004": "i think both test folder and sample\\_submission.csv are replaced during submission.\nMy code work for other model, so i think it is not the reason.\n\nbesides, if i were only processing three files, the submission time should be very short",
    "2353253": "the latest version still course error:\n\n\n#use file list from test folder\n```\nif mode=='submit':\n    mp3_dir = f'/kaggle/input/bengaliai-speech/test_mp3s'\n    glob_file = glob(f'{mp3_dir}/*.mp3')\n    id = sorted([f[len(mp3_dir)+1:-4] for f in glob_file])\n    valid_df = pd.DataFrame({'id':id,'sentence':''})\n```\n\n\n#check each prediction can be process by jiwer\n```\ndef check_text(sentence):\n    t = 'িনি এবং ও এই করে া। তার' #fake truth\n    #print(sentence)\n    try:\n        jiwer.wer(t, [sentence])\n    except:\n        sentence = '' #t\n    return sentence\n\n\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(check_text)\n```",
    "2353304": "# Solved LB 0.52\n\nI successfully submitted using your notebook. The main problem is :\n> The model is generating some empty predictions. It seems like you cant't submit empty strings in the dataframe. I replaced the empty strings with \",\" and it got submitted. \n\n``` \nfilenames = []\npredictions = []\nfiles = os.listdir('/kaggle/input/bengaliai-speech/test_mp3s')\nbase_path = \"/kaggle/input/bengaliai-speech/test_mp3s/\"\n\nfor file in files:\n    filenames.append(file.split(\".\")[0])\n    audio = load_audio(base_path+file) #Some function to load the audio file\n    text = model(audio)  #Infer with your model\n    if len(text)==0:\n          text = \",\"\n    \n    predictions.append(text)\n\ndf = pd.DataFrame({'id':filenames,'sentence':predictions})\ndf.to_csv('submission.csv',index=False)\n```\n\nRefer to [this notebook](https://www.kaggle.com/mbmmurad/fork-of-fork-of-local-nemo-baseline-conformer)",
    "2353306": "You're right. My mistake. Your pipeline worked. Just needed to handle the empty strings.",
    "2353310": "Just change to ``` sentence = \",\"``` .",
    "2353348": "mbmmurad \n\nThanks a lot!!!",
    "2353352": "by the way, this is wave2vec2-CTC without LM",
    "2353382": "more information about the model can be found here:\nhttps://github.com/AI4Bharat/IndicWav2Vec/tree/main\n\ni think the model is trained with author's own dataset.\nHence we can conclude? OOD cause adroup of about 0.12\n(since kaggle valid split has WER = 0.40, public LB 0.52)\n\n\n---\nIndicWav2Vec paper:\nhttps://arxiv.org/pdf/2111.03945.pdf\n\n\"Apart from YouTube, we also curated content\nfrom newsonair3 which is a radio news channel\nrun by the Govt. of India and broadcasts news in\nmultiple Indian languages\"\n\n---\nthere is CTC-LM model which gives better results but you need to convert from fairseq to hf yourself,\nfollowing instructions from githut repo\n\n\n---\n\n\"3.5 Rescoring\nNote that in the above decoding process, we use an\nn-gram KenLM language model (Heafield, 2011).\nHowever, recent advances in language modeling\nhave shown that transformer based language models perform well. To get the best of both worlds,\nwe optionally use an external transformer based\nlanguage model to rescore the n-best hypothesis\nproduced above (Synnaeve et al., 2019). Specifically, for each output in the n-best list, we compute a new score by combining the decoder logprobability, \"\n\nthe rescoring transformer can be used for out ensemble of  different models",
    "2353407": "Yes the model was trained on their own dataset. The model without LM scored WER 0.166 on their evaluation set and with LM performed 0.136.\n\nSo I guess the model with LM might score better than this one. Agreed with the drop for the OOD part! But a thing to consider is the valid split contains 30k files, with comparatively shorter sentence lengths. The public LB is evaluated on around 4k samples. So it will be interesting to see how we make a decision from the LB-CV difference.",
    "2353451": "considering that indian language is close to bengali, this model is a good starting source. they clamied their training data is public ... i am trying to find them now",
    "2353470": "Seems so. Good luck!",
    "2353493": "maybe here!!!\nhttps://ai4bharat.iitm.ac.in/dhwani\n\nDataset Format\nThe audio files present in separate folders.\nFor YouTubeThe audio filenames are named YouTube-ids and for Newsonair, the contatination of region name, bulletin timing makes the filename.\n\n---\n\nmore here:\nhttps://ai4bharat.iitm.ac.in/datasets\n\n\nconformer model demo\nhttps://models.ai4bharat.org/#/asr/conformer",
    "2362376": "Have the same issue. Is here a feasible solution for this error provided?\nIs that \n```py\ndef check_text(sentence):\n    t = 'িনি এবং ও এই করে া। তার' #fake truth\n    #print(sentence)\n    try:\n        jiwer.wer(t, [sentence])\n    except:\n        sentence = ',' #t\n    return sentence\n\n\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(check_text)\n```",
    "2362384": "~~~\ndef check(sentence):\n    if len(sentence)==0:\n        sentence = \",\"\n    return sentence\ndf.sentence = df.sentence.apply(lambda x: check(x))\n~~~\nThis should work. The problem is your model can sometimes generate empty strings, and when you convert this empty strings to a dataframe, it appears as nan. \nSo we're just handling the empty strings by this function, replacing them with a character so that no empty strings remain in the predictions",
    "2362390": "Is this scenario referring to that the whole prediction for this audio is empty, or there are spaces in the predicted sentences?",
    "2362394": "For the whole audio. There are some clips in the training/validation set that are 1-2 sec long, consisting of only 2-3 words. Now if the voice Isn't loud enough/audio contain more external noise, the model might fail to even extract a single word/character, thus returning an empty string for the whole audio.",
    "2362398": "When I submit my notebook and prediction, will the model also predict for audios in training and validating dataset, or some private dataset unseen so that the problem could occur? (I'm somewhat novice, not understanding the competition mechanism)",
    "2362403": "That's completely fine. You only need to predict for the 'test_mp3s' folder. This folder contains unseen audios (OOD audios), you can only see three of them now. But when you submit your notebook, your notebook will have access to all of the unseen audios( around 8000) \n\nYou can have a look at [my notebook](https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model) or the notebook pinned in the code section to see how it works!",
    "2362410": "OK. Thanks a lot for your detailed explanation!🥰",
    "2362559": "I'm newly in compitition i know simple basic Python,pandas and numpy knowledge so can you guide me this project?",
    "2362737": "This may help (about this competition): https://www.kaggle.com/code/ayushs9020/understanding-the-competition-bengalai\nAnd some lectures like Stanford CS230n may help!",
    "2445056": "hengck23 could you give some pointer or example how rescoring transformer would work for ensembling model ?"
  },
  "source": "meta"
}