{
  "id": 399506,
  "title": "Submission Scoring Error",
  "url": "/competitions/birdclef-2023/discussion/399506",
  "author_name": "",
  "post_date": "2023-04-04T10:56:04.782742500Z",
  "votes": 5,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Repeatedly, I am getting the error </p>\n<p>Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.</p>\n<p>I have correct the data many times still the same, any help much appreciated.</p>\n<p>My notebook is shared, apologies if its not following any standards.</p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023</a></p>\n<p>After a couple of attempts its fixed:</p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook</a></p>",
  "messages": [
    {
      "id": "2208865",
      "postDate": "04/04/2023 10:56:04",
      "content": "<p>Repeatedly, I am getting the error </p>\n<p>Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.</p>\n<p>I have correct the data many times still the same, any help much appreciated.</p>\n<p>My notebook is shared, apologies if its not following any standards.</p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023</a></p>\n<p>After a couple of attempts its fixed:</p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook</a></p>",
      "rawMarkdown": "Repeatedly, I am getting the error \n\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.\n\nI have correct the data many times still the same, any help much appreciated.\n\nMy notebook is shared, apologies if its not following any standards.\n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023\n\nAfter a couple of attempts its fixed:\n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook",
      "votes": null
    },
    {
      "id": "2209017",
      "postDate": "04/04/2023 12:19:35",
      "content": "<p>Hi man, double check your dataset generator, probably the issue it's there, also check your submission CSV file, it must have id's separated by 5 (for 5 seconds), eg, soundscape_29201_5, soundscape_29201_10, soundscape_29201_15, etc. And from soundscape_29201_5 to soundscape_29201_600, with this shape 120 × 265.<br>\nPerhaps the following inference notebooks can help you:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook\" target=\"_blank\">https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-pretraining-is-all-you-need-infer\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/birdclef23-pretraining-is-all-you-need-infer</a></li>\n</ul>",
      "rawMarkdown": "Hi man, double check your dataset generator, probably the issue it's there, also check your submission CSV file, it must have id's separated by 5 (for 5 seconds), eg, soundscape_29201_5, soundscape_29201_10, soundscape_29201_15, etc. And from soundscape_29201_5 to soundscape_29201_600, with this shape 120 × 265.\nPerhaps the following inference notebooks can help you:\n\n- https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook\n- https://www.kaggle.com/code/awsaf49/birdclef23-pretraining-is-all-you-need-infer",
      "votes": null
    },
    {
      "id": "2209237",
      "postDate": "04/04/2023 15:03:45",
      "content": "<p>So I'm having the same problem and I think there are some issues with the scoring code. I have the following piece of code which generates <code>submission.csv</code> and scoring of this file fails. However, if I generate exactly the same CSV file by modifying the example notebook, I get a score of 0.71. I can see that the files are identical by downloading them and diffing them locally.</p>\n<p>Any ideas?</p>",
      "rawMarkdown": "So I'm having the same problem and I think there are some issues with the scoring code. I have the following piece of code which generates `submission.csv` and scoring of this file fails. However, if I generate exactly the same CSV file by modifying the example notebook, I get a score of 0.71. I can see that the files are identical by downloading them and diffing them locally.\n\nAny ideas?",
      "votes": null
    },
    {
      "id": "2209534",
      "postDate": "04/04/2023 18:28:40",
      "content": "<p>Your notebook is just generating a static file for a single visible test soundscape which wouldn't work for the hidden test that is made available to the notebook during submission scoring. Hope this helps :)</p>",
      "rawMarkdown": "Your notebook is just generating a static file for a single visible test soundscape which wouldn't work for the hidden test that is made available to the notebook during submission scoring. Hope this helps :)",
      "votes": null
    },
    {
      "id": "2210076",
      "postDate": "04/05/2023 05:31:38",
      "content": "<p>But the example notebook (<a href=\"https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models\" target=\"_blank\">https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models</a>) does exactly the same, it generates a submission.csv with 3 rows.</p>\n<p>The competition description makes it sound like all that matters is the generated CSV file. But this is clearly not true as the same CSV file sometimes gets processed correctly and sometimes does not.</p>",
      "rawMarkdown": "But the example notebook (https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models) does exactly the same, it generates a submission.csv with 3 rows.\n\nThe competition description makes it sound like all that matters is the generated CSV file. But this is clearly not true as the same CSV file sometimes gets processed correctly and sometimes does not.",
      "votes": null
    },
    {
      "id": "2210096",
      "postDate": "04/05/2023 06:07:21",
      "content": "<p>In the screenshot that you attached, you are doing this: ```return (f'soundscape_29201…'), so essentially, you are hardcoding for just one example in the visible test. You need to make this general so all the samples in the hidden test set have the appropriate name.</p>",
      "rawMarkdown": "In the screenshot that you attached, you are doing this: ```return (f'soundscape_29201...'), so essentially, you are hardcoding for just one example in the visible test. You need to make this general so all the samples in the hidden test set have the appropriate name.",
      "votes": null
    },
    {
      "id": "2210174",
      "postDate": "04/05/2023 07:34:44",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2654254%2Fdba10da9c5f214b625ba0da0f61dece1%2FR1005_BirdCLEF2023.png?generation=1680679854812334&amp;alt=media\" alt=\"\"></p>\n<p>in the snapshot please find my dataframe before I dump into submission.csv and below is the code to convert my predictions (0/1) into df column values.</p>\n<p><em>One-hot encode the labels, assuming that the labels are integers</em><br>\none_hot_labels = np.eye(num_labels, dtype = int)[predictions] </p>\n<p><em>Create a DataFrame with the one-hot encoded labels, assuming that the column names are incorrect</em><br>\nlabel_columns = [f'label_{i}' for i in range(num_labels)]<br>\npred_df = pd.DataFrame(data=one_hot_labels, columns=label_columns)</p>\n<p><em>Create a mapping dictionary between the column names and the labels</em><br>\nmapping = {col_name: birds[label_idx] for label_idx, col_name in enumerate(pred_df.columns)}</p>\n<p><em>Map the column names to the correct labels</em><br>\npred_df = pred_df.rename(columns=mapping)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2654254%2Fdba10da9c5f214b625ba0da0f61dece1%2FR1005_BirdCLEF2023.png?generation=1680679854812334&alt=media)\n\nin the snapshot please find my dataframe before I dump into submission.csv and below is the code to convert my predictions (0/1) into df column values.\n\n*One-hot encode the labels, assuming that the labels are integers*\none_hot_labels = np.eye(num_labels, dtype = int)[predictions] \n\n*Create a DataFrame with the one-hot encoded labels, assuming that the column names are incorrect*\nlabel_columns = [f'label_{i}' for i in range(num_labels)]\npred_df = pd.DataFrame(data=one_hot_labels, columns=label_columns)\n\n*Create a mapping dictionary between the column names and the labels*\nmapping = {col_name: birds[label_idx] for label_idx, col_name in enumerate(pred_df.columns)}\n\n*Map the column names to the correct labels*\npred_df = pred_df.rename(columns=mapping)",
      "votes": null
    },
    {
      "id": "2210727",
      "postDate": "04/05/2023 15:24:19",
      "content": "<p>Hi, Rajesh;</p>\n<p>One of the reasons for providing the sample submission notebook is that it can be frustrating getting the formatting and such right when starting from scratch. We strongly recommend reusing the code from the submission notebook for writing the submission.</p>",
      "rawMarkdown": "Hi, Rajesh;\n\nOne of the reasons for providing the sample submission notebook is that it can be frustrating getting the formatting and such right when starting from scratch. We strongly recommend reusing the code from the submission notebook for writing the submission.",
      "votes": null
    },
    {
      "id": "2210749",
      "postDate": "04/05/2023 15:40:40",
      "content": "<p>I should have done that in the first place, regret and changing my code as per your suggestion. I will update if there is an error else all set. Thanks for the support. </p>",
      "rawMarkdown": "I should have done that in the first place, regret and changing my code as per your suggestion. I will update if there is an error else all set. Thanks for the support.",
      "votes": null
    },
    {
      "id": "2211639",
      "postDate": "04/06/2023 06:55:14",
      "content": "<p>Solved: you have to read in the CSV file provided by organisers and modify that. A minimal submission below:</p>\n<pre><code>df= pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\ndf[classes] = df[classes].astype(np.float32)\ndf[classes] = 0.0038\ndf.to_csv(\"submission.csv\", index=False)\n</code></pre>",
      "rawMarkdown": "Solved: you have to read in the CSV file provided by organisers and modify that. A minimal submission below:\n\n```\ndf= pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\ndf[classes] = df[classes].astype(np.float32)\ndf[classes] = 0.0038\ndf.to_csv(\"submission.csv\", index=False)\n```",
      "votes": null
    },
    {
      "id": "2216335",
      "postDate": "04/10/2023 02:46:26",
      "content": "<p>Have you solved this problem?</p>",
      "rawMarkdown": "Have you solved this problem?",
      "votes": null
    },
    {
      "id": "2216495",
      "postDate": "04/10/2023 05:59:24",
      "content": "<p>May i ask how did you solve this problem?</p>",
      "rawMarkdown": "May i ask how did you solve this problem?",
      "votes": null
    },
    {
      "id": "2217575",
      "postDate": "04/11/2023 03:15:10",
      "content": "<p>Not yet, after following some of the notebooks I revised the submission logic. Now I am getting \"Notebook Threw Exception\" as per the <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">code debug doc</a>, revisiting my pipeline and rebuild it in a way that starts with a valid submission. 🤒</p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023</a></p>",
      "rawMarkdown": "Not yet, after following some of the notebooks I revised the submission logic. Now I am getting \"Notebook Threw Exception\" as per the [code debug doc](https://www.kaggle.com/code-competition-debugging), revisiting my pipeline and rebuild it in a way that starts with a valid submission. 🤒\n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023",
      "votes": null
    },
    {
      "id": "2217839",
      "postDate": "04/11/2023 07:50:07",
      "content": "<p>I modified the code and then faced the same problem as yours🙏</p>",
      "rawMarkdown": "I modified the code and then faced the same problem as yours🙏",
      "votes": null
    },
    {
      "id": "2218089",
      "postDate": "04/11/2023 12:18:40",
      "content": "<p>I can share some perspectives from my own experience. When I set test batch size to 1, everything is fine. However, when batch size goes up, \"Notebook Threw Exception\" occurs. I am still trying to figure out the reason and will do more experiments. </p>",
      "rawMarkdown": "I can share some perspectives from my own experience. When I set test batch size to 1, everything is fine. However, when batch size goes up, \"Notebook Threw Exception\" occurs. I am still trying to figure out the reason and will do more experiments.",
      "votes": null
    },
    {
      "id": "2219087",
      "postDate": "04/12/2023 09:35:53",
      "content": "<p>HI man, sorry for the delay, I thought you solved it. Does your dataset have a 120x265 shape? What Aphysict said it's true, happens to me too. <br>\nAlso, make the train in a separate notebook, save the model, and inference in another one. <br>\nCheck out this predict approach, it's super simple, <a href=\"https://www.kaggle.com/code/morodertobias/bc23-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/morodertobias/bc23-baseline-inference</a>.</p>",
      "rawMarkdown": "HI man, sorry for the delay, I thought you solved it. Does your dataset have a 120x265 shape? What Aphysict said it's true, happens to me too. \nAlso, make the train in a separate notebook, save the model, and inference in another one. \nCheck out this predict approach, it's super simple, https://www.kaggle.com/code/morodertobias/bc23-baseline-inference.",
      "votes": null
    },
    {
      "id": "2219098",
      "postDate": "04/12/2023 09:47:42",
      "content": "<p>BINGO! I fixed it. </p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook</a></p>",
      "rawMarkdown": "BINGO! I fixed it. \n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook",
      "votes": null
    },
    {
      "id": "2222634",
      "postDate": "04/15/2023 12:45:01",
      "content": "<p>The problem is that there are many test sets, you need to design a common input and output to submit all test sets</p>",
      "rawMarkdown": "The problem is that there are many test sets, you need to design a common input and output to submit all test sets",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2209017,
      "author_name": "maxdiazbattan",
      "author_url": "",
      "post_date": "04/04/2023 12:19:35",
      "content": "<p>Hi man, double check your dataset generator, probably the issue it's there, also check your submission CSV file, it must have id's separated by 5 (for 5 seconds), eg, soundscape_29201_5, soundscape_29201_10, soundscape_29201_15, etc. And from soundscape_29201_5 to soundscape_29201_600, with this shape 120 × 265.<br>\nPerhaps the following inference notebooks can help you:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook\" target=\"_blank\">https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-pretraining-is-all-you-need-infer\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/birdclef23-pretraining-is-all-you-need-infer</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2210174,
          "author_name": "rajeshradhakrishnan",
          "author_url": "",
          "post_date": "04/05/2023 07:34:44",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2654254%2Fdba10da9c5f214b625ba0da0f61dece1%2FR1005_BirdCLEF2023.png?generation=1680679854812334&amp;alt=media\" alt=\"\"></p>\n<p>in the snapshot please find my dataframe before I dump into submission.csv and below is the code to convert my predictions (0/1) into df column values.</p>\n<p><em>One-hot encode the labels, assuming that the labels are integers</em><br>\none_hot_labels = np.eye(num_labels, dtype = int)[predictions] </p>\n<p><em>Create a DataFrame with the one-hot encoded labels, assuming that the column names are incorrect</em><br>\nlabel_columns = [f'label_{i}' for i in range(num_labels)]<br>\npred_df = pd.DataFrame(data=one_hot_labels, columns=label_columns)</p>\n<p><em>Create a mapping dictionary between the column names and the labels</em><br>\nmapping = {col_name: birds[label_idx] for label_idx, col_name in enumerate(pred_df.columns)}</p>\n<p><em>Map the column names to the correct labels</em><br>\npred_df = pred_df.rename(columns=mapping)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2209237,
      "author_name": "jedski",
      "author_url": "",
      "post_date": "04/04/2023 15:03:45",
      "content": "<p>So I'm having the same problem and I think there are some issues with the scoring code. I have the following piece of code which generates <code>submission.csv</code> and scoring of this file fails. However, if I generate exactly the same CSV file by modifying the example notebook, I get a score of 0.71. I can see that the files are identical by downloading them and diffing them locally.</p>\n<p>Any ideas?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2209534,
          "author_name": "aryankhatana",
          "author_url": "",
          "post_date": "04/04/2023 18:28:40",
          "content": "<p>Your notebook is just generating a static file for a single visible test soundscape which wouldn't work for the hidden test that is made available to the notebook during submission scoring. Hope this helps :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2210076,
              "author_name": "jedski",
              "author_url": "",
              "post_date": "04/05/2023 05:31:38",
              "content": "<p>But the example notebook (<a href=\"https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models\" target=\"_blank\">https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models</a>) does exactly the same, it generates a submission.csv with 3 rows.</p>\n<p>The competition description makes it sound like all that matters is the generated CSV file. But this is clearly not true as the same CSV file sometimes gets processed correctly and sometimes does not.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2210096,
                  "author_name": "aryankhatana",
                  "author_url": "",
                  "post_date": "04/05/2023 06:07:21",
                  "content": "<p>In the screenshot that you attached, you are doing this: ```return (f'soundscape_29201…'), so essentially, you are hardcoding for just one example in the visible test. You need to make this general so all the samples in the hidden test set have the appropriate name.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2216495,
          "author_name": "leexiang0",
          "author_url": "",
          "post_date": "04/10/2023 05:59:24",
          "content": "<p>May i ask how did you solve this problem?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2210727,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "04/05/2023 15:24:19",
      "content": "<p>Hi, Rajesh;</p>\n<p>One of the reasons for providing the sample submission notebook is that it can be frustrating getting the formatting and such right when starting from scratch. We strongly recommend reusing the code from the submission notebook for writing the submission.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2210749,
          "author_name": "rajeshradhakrishnan",
          "author_url": "",
          "post_date": "04/05/2023 15:40:40",
          "content": "<p>I should have done that in the first place, regret and changing my code as per your suggestion. I will update if there is an error else all set. Thanks for the support. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2211639,
              "author_name": "jedski",
              "author_url": "",
              "post_date": "04/06/2023 06:55:14",
              "content": "<p>Solved: you have to read in the CSV file provided by organisers and modify that. A minimal submission below:</p>\n<pre><code>df= pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\ndf[classes] = df[classes].astype(np.float32)\ndf[classes] = 0.0038\ndf.to_csv(\"submission.csv\", index=False)\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2216335,
      "author_name": "leexiang0",
      "author_url": "",
      "post_date": "04/10/2023 02:46:26",
      "content": "<p>Have you solved this problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2217575,
          "author_name": "rajeshradhakrishnan",
          "author_url": "",
          "post_date": "04/11/2023 03:15:10",
          "content": "<p>Not yet, after following some of the notebooks I revised the submission logic. Now I am getting \"Notebook Threw Exception\" as per the <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">code debug doc</a>, revisiting my pipeline and rebuild it in a way that starts with a valid submission. 🤒</p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2217839,
              "author_name": "leexiang0",
              "author_url": "",
              "post_date": "04/11/2023 07:50:07",
              "content": "<p>I modified the code and then faced the same problem as yours🙏</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2218089,
              "author_name": "aphysict",
              "author_url": "",
              "post_date": "04/11/2023 12:18:40",
              "content": "<p>I can share some perspectives from my own experience. When I set test batch size to 1, everything is fine. However, when batch size goes up, \"Notebook Threw Exception\" occurs. I am still trying to figure out the reason and will do more experiments. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2219087,
                  "author_name": "maxdiazbattan",
                  "author_url": "",
                  "post_date": "04/12/2023 09:35:53",
                  "content": "<p>HI man, sorry for the delay, I thought you solved it. Does your dataset have a 120x265 shape? What Aphysict said it's true, happens to me too. <br>\nAlso, make the train in a separate notebook, save the model, and inference in another one. <br>\nCheck out this predict approach, it's super simple, <a href=\"https://www.kaggle.com/code/morodertobias/bc23-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/morodertobias/bc23-baseline-inference</a>.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2219098,
                      "author_name": "rajeshradhakrishnan",
                      "author_url": "",
                      "post_date": "04/12/2023 09:47:42",
                      "content": "<p>BINGO! I fixed it. </p>\n<p><a href=\"https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook\" target=\"_blank\">https://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook</a></p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2222634,
      "author_name": "xuanleekaggle",
      "author_url": "",
      "post_date": "04/15/2023 12:45:01",
      "content": "<p>The problem is that there are many test sets, you need to design a common input and output to submit all test sets</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2208865": "Repeatedly, I am getting the error \n\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.\n\nI have correct the data many times still the same, any help much appreciated.\n\nMy notebook is shared, apologies if its not following any standards.\n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023\n\nAfter a couple of attempts its fixed:\n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook",
    "2209017": "Hi man, double check your dataset generator, probably the issue it's there, also check your submission CSV file, it must have id's separated by 5 (for 5 seconds), eg, soundscape_29201_5, soundscape_29201_10, soundscape_29201_15, etc. And from soundscape_29201_5 to soundscape_29201_600, with this shape 120 × 265.\nPerhaps the following inference notebooks can help you:\n\n- https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook\n- https://www.kaggle.com/code/awsaf49/birdclef23-pretraining-is-all-you-need-infer",
    "2209237": "So I'm having the same problem and I think there are some issues with the scoring code. I have the following piece of code which generates `submission.csv` and scoring of this file fails. However, if I generate exactly the same CSV file by modifying the example notebook, I get a score of 0.71. I can see that the files are identical by downloading them and diffing them locally.\n\nAny ideas?",
    "2209534": "Your notebook is just generating a static file for a single visible test soundscape which wouldn't work for the hidden test that is made available to the notebook during submission scoring. Hope this helps :)",
    "2210076": "But the example notebook (https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models) does exactly the same, it generates a submission.csv with 3 rows.\n\nThe competition description makes it sound like all that matters is the generated CSV file. But this is clearly not true as the same CSV file sometimes gets processed correctly and sometimes does not.",
    "2210096": "In the screenshot that you attached, you are doing this: ```return (f'soundscape_29201...'), so essentially, you are hardcoding for just one example in the visible test. You need to make this general so all the samples in the hidden test set have the appropriate name.",
    "2210174": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2654254%2Fdba10da9c5f214b625ba0da0f61dece1%2FR1005_BirdCLEF2023.png?generation=1680679854812334&alt=media)\n\nin the snapshot please find my dataframe before I dump into submission.csv and below is the code to convert my predictions (0/1) into df column values.\n\n*One-hot encode the labels, assuming that the labels are integers*\none_hot_labels = np.eye(num_labels, dtype = int)[predictions] \n\n*Create a DataFrame with the one-hot encoded labels, assuming that the column names are incorrect*\nlabel_columns = [f'label_{i}' for i in range(num_labels)]\npred_df = pd.DataFrame(data=one_hot_labels, columns=label_columns)\n\n*Create a mapping dictionary between the column names and the labels*\nmapping = {col_name: birds[label_idx] for label_idx, col_name in enumerate(pred_df.columns)}\n\n*Map the column names to the correct labels*\npred_df = pred_df.rename(columns=mapping)",
    "2210727": "Hi, Rajesh;\n\nOne of the reasons for providing the sample submission notebook is that it can be frustrating getting the formatting and such right when starting from scratch. We strongly recommend reusing the code from the submission notebook for writing the submission.",
    "2210749": "I should have done that in the first place, regret and changing my code as per your suggestion. I will update if there is an error else all set. Thanks for the support.",
    "2211639": "Solved: you have to read in the CSV file provided by organisers and modify that. A minimal submission below:\n\n```\ndf= pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\ndf[classes] = df[classes].astype(np.float32)\ndf[classes] = 0.0038\ndf.to_csv(\"submission.csv\", index=False)\n```",
    "2216335": "Have you solved this problem?",
    "2216495": "May i ask how did you solve this problem?",
    "2217575": "Not yet, after following some of the notebooks I revised the submission logic. Now I am getting \"Notebook Threw Exception\" as per the [code debug doc](https://www.kaggle.com/code-competition-debugging), revisiting my pipeline and rebuild it in a way that starts with a valid submission. 🤒\n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/submission-scoring-error-birdclef2023",
    "2217839": "I modified the code and then faced the same problem as yours🙏",
    "2218089": "I can share some perspectives from my own experience. When I set test batch size to 1, everything is fine. However, when batch size goes up, \"Notebook Threw Exception\" occurs. I am still trying to figure out the reason and will do more experiments.",
    "2219087": "HI man, sorry for the delay, I thought you solved it. Does your dataset have a 120x265 shape? What Aphysict said it's true, happens to me too. \nAlso, make the train in a separate notebook, save the model, and inference in another one. \nCheck out this predict approach, it's super simple, https://www.kaggle.com/code/morodertobias/bc23-baseline-inference.",
    "2219098": "BINGO! I fixed it. \n\nhttps://www.kaggle.com/code/rajeshradhakrishnan/baseline-pytorch-birdclef2023/notebook",
    "2222634": "The problem is that there are many test sets, you need to design a common input and output to submit all test sets"
  },
  "source": "meta"
}