{
  "id": 230729,
  "title": "Getting Submission CSV Not Found",
  "url": "/competitions/birdclef-2021/discussion/230729",
  "author_name": "",
  "post_date": "2021-04-05T12:11:46.445526700Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have submitted my notebook which has a submission file and in the proper format but while running for the private dataset it's showing  <strong>Submission CSV Not Found</strong>. I have no clue why this happening. <br>\nThis is how I am switching my directory from public to private</p>\n<blockquote>\n  <p>TEST = (len(list(Path(\"../input/birdclef-2021/test_soundscapes/\").glob(\"*.ogg\"))) != 0)</p>\n  <p>if TEST:<br>\n            DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")<br>\n  else:<br>\n          DATADIR = Path(\"../input/birdclef-2021/train_soundscapes/\")</p>\n</blockquote>\n<p>And this is a <strong>submission.csv</strong> file format. Please help<br>\n    row_id  birds<br>\n0    20152_SSW_5 nocall<br>\n1    20152_SSW_10    nocall<br>\n2    20152_SSW_15    nocall<br>\n3    20152_SSW_20    nocall<br>\n4    20152_SSW_25    nocall<br>\n…    … …<br>\n2395    26709_SSW_580   nocall<br>\n2396    26709_SSW_585   nocall<br>\n2397    26709_SSW_590   nocall<br>\n2398    26709_SSW_595   nocall<br>\n2399    26709_SSW_600   nocall</p>",
  "messages": [
    {
      "id": "1263432",
      "postDate": "04/05/2021 12:11:46",
      "content": "<p>I have submitted my notebook which has a submission file and in the proper format but while running for the private dataset it's showing  <strong>Submission CSV Not Found</strong>. I have no clue why this happening. <br>\nThis is how I am switching my directory from public to private</p>\n<blockquote>\n  <p>TEST = (len(list(Path(\"../input/birdclef-2021/test_soundscapes/\").glob(\"*.ogg\"))) != 0)</p>\n  <p>if TEST:<br>\n            DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")<br>\n  else:<br>\n          DATADIR = Path(\"../input/birdclef-2021/train_soundscapes/\")</p>\n</blockquote>\n<p>And this is a <strong>submission.csv</strong> file format. Please help<br>\n    row_id  birds<br>\n0    20152_SSW_5 nocall<br>\n1    20152_SSW_10    nocall<br>\n2    20152_SSW_15    nocall<br>\n3    20152_SSW_20    nocall<br>\n4    20152_SSW_25    nocall<br>\n…    … …<br>\n2395    26709_SSW_580   nocall<br>\n2396    26709_SSW_585   nocall<br>\n2397    26709_SSW_590   nocall<br>\n2398    26709_SSW_595   nocall<br>\n2399    26709_SSW_600   nocall</p>",
      "rawMarkdown": "I have submitted my notebook which has a submission file and in the proper format but while running for the private dataset it's showing  **Submission CSV Not Found**. I have no clue why this happening. \nThis is how I am switching my directory from public to private\n\n> TEST = (len(list(Path(\"../input/birdclef-2021/test_soundscapes/\").glob(\"*.ogg\"))) != 0)\n\n>if TEST:\n>           DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")\n>else:\n>         DATADIR = Path(\"../input/birdclef-2021/train_soundscapes/\")\n\nAnd this is a **submission.csv** file format. Please help\n\trow_id\tbirds\n0\t20152_SSW_5\tnocall\n1\t20152_SSW_10\tnocall\n2\t20152_SSW_15\tnocall\n3\t20152_SSW_20\tnocall\n4\t20152_SSW_25\tnocall\n...\t...\t...\n2395\t26709_SSW_580\tnocall\n2396\t26709_SSW_585\tnocall\n2397\t26709_SSW_590\tnocall\n2398\t26709_SSW_595\tnocall\n2399\t26709_SSW_600\tnocall",
      "votes": null
    },
    {
      "id": "1263961",
      "postDate": "04/05/2021 19:26:29",
      "content": "<p>Hi Nikhil,</p>\n<p>The contents of your <code>submission.csv</code> isn't in the correct format.  Specifically, you shouldn't have 1, 2, 3… at the start of the rows.  Assuming you're using pandas to generate this, you need to add <code>index=False</code> when you call <code>.to_csv</code>.  For example, if your dataframe is called <code>df_test</code> then you should use…</p>\n<p><code>df_test[[\"row_id\", \"birds\"]].to_csv(\"submission.csv\", index=False)</code></p>\n<p>Let me know if you're still having problems after that.</p>\n<p>Andrew</p>",
      "rawMarkdown": "Hi Nikhil,\n\nThe contents of your `submission.csv` isn't in the correct format.  Specifically, you shouldn't have 1, 2, 3... at the start of the rows.  Assuming you're using pandas to generate this, you need to add `index=False` when you call `.to_csv`.  For example, if your dataframe is called `df_test` then you should use...\n\n`df_test[[\"row_id\", \"birds\"]].to_csv(\"submission.csv\", index=False)`\n\nLet me know if you're still having problems after that.\n\nAndrew",
      "votes": null
    },
    {
      "id": "1264038",
      "postDate": "04/05/2021 20:25:30",
      "content": "<p><a href=\"https://www.kaggle.com/andrewrrose\" target=\"_blank\">@andrewrrose</a> Hi thanks for the reply.<br>\nThis is my code for df to csv  </p>\n<blockquote>\n  <p>prediction_df.to_csv(\"submission.csv\", index=False)</p>\n</blockquote>\n<p>I already set index = False before posting this issue. I guess there is some other issue. Let me know you want anything which can help to resolve this issue.</p>",
      "rawMarkdown": "andrewrrose Hi thanks for the reply.\nThis is my code for df to csv  \n> prediction_df.to_csv(\"submission.csv\", index=False)\n\nI already set index = False before posting this issue. I guess there is some other issue. Let me know you want anything which can help to resolve this issue.",
      "votes": null
    },
    {
      "id": "1264369",
      "postDate": "04/06/2021 05:35:08",
      "content": "<p>Three thoughts then…</p>\n<ul>\n<li>Can you post you actual submission.csv, rather than the dataframe.</li>\n<li>Try hard-coding <code>TEST = True</code> in case there's something wrong with your auto-detection of which environment you're running in.</li>\n<li>If it really can't find <code>submission.csv</code>, your code is probably crashing / running out of memory when being run for real.  See <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">Kaggle's debugging guide</a> although mostly it just says that you need to do careful debugging.</li>\n</ul>",
      "rawMarkdown": "Three thoughts then...\n\n- Can you post you actual submission.csv, rather than the dataframe.\n- Try hard-coding `TEST = True` in case there's something wrong with your auto-detection of which environment you're running in.\n- If it really can't find `submission.csv`, your code is probably crashing / running out of memory when being run for real.  See [Kaggle's debugging guide](https://www.kaggle.com/code-competition-debugging) although mostly it just says that you need to do careful debugging.",
      "votes": null
    },
    {
      "id": "1264532",
      "postDate": "04/06/2021 08:22:44",
      "content": "<p>Solved it. I am getting issue because of this   </p>\n<blockquote>\n  <p>DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")</p>\n</blockquote>\n<p>I was assuming this test dir only contains  '.ogg'  file but it also has '.txt' and '.csv' file which when passing through the particular function which is meant for only '.ogg' file throws an error.<br>\nI fixed it by setting a filter for '.ogg' file and submitted it successfully.<br>\nThanks for the help, sir. </p>",
      "rawMarkdown": "Solved it. I am getting issue because of this   \n> DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")\n\nI was assuming this test dir only contains  '.ogg'  file but it also has '.txt' and '.csv' file which when passing through the particular function which is meant for only '.ogg' file throws an error.\nI fixed it by setting a filter for '.ogg' file and submitted it successfully.\nThanks for the help, sir.",
      "votes": null
    },
    {
      "id": "1268878",
      "postDate": "04/09/2021 22:21:40",
      "content": "<p>Hi, I have a similar problem, but I already filter for '.ogg' files with </p>\n<p><code>all_audios = list(DATADIR.glob(\"*.ogg\"))</code></p>\n<p>How did you do that?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi, I have a similar problem, but I already filter for '.ogg' files with \n\n`all_audios = list(DATADIR.glob(\"*.ogg\"))`\n\nHow did you do that?\n\nThanks in advance!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1263961,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "04/05/2021 19:26:29",
      "content": "<p>Hi Nikhil,</p>\n<p>The contents of your <code>submission.csv</code> isn't in the correct format.  Specifically, you shouldn't have 1, 2, 3… at the start of the rows.  Assuming you're using pandas to generate this, you need to add <code>index=False</code> when you call <code>.to_csv</code>.  For example, if your dataframe is called <code>df_test</code> then you should use…</p>\n<p><code>df_test[[\"row_id\", \"birds\"]].to_csv(\"submission.csv\", index=False)</code></p>\n<p>Let me know if you're still having problems after that.</p>\n<p>Andrew</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1264038,
      "author_name": "coder247",
      "author_url": "",
      "post_date": "04/05/2021 20:25:30",
      "content": "<p><a href=\"https://www.kaggle.com/andrewrrose\" target=\"_blank\">@andrewrrose</a> Hi thanks for the reply.<br>\nThis is my code for df to csv  </p>\n<blockquote>\n  <p>prediction_df.to_csv(\"submission.csv\", index=False)</p>\n</blockquote>\n<p>I already set index = False before posting this issue. I guess there is some other issue. Let me know you want anything which can help to resolve this issue.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264369,
          "author_name": "andrewrrose",
          "author_url": "",
          "post_date": "04/06/2021 05:35:08",
          "content": "<p>Three thoughts then…</p>\n<ul>\n<li>Can you post you actual submission.csv, rather than the dataframe.</li>\n<li>Try hard-coding <code>TEST = True</code> in case there's something wrong with your auto-detection of which environment you're running in.</li>\n<li>If it really can't find <code>submission.csv</code>, your code is probably crashing / running out of memory when being run for real.  See <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">Kaggle's debugging guide</a> although mostly it just says that you need to do careful debugging.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264532,
          "author_name": "coder247",
          "author_url": "",
          "post_date": "04/06/2021 08:22:44",
          "content": "<p>Solved it. I am getting issue because of this   </p>\n<blockquote>\n  <p>DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")</p>\n</blockquote>\n<p>I was assuming this test dir only contains  '.ogg'  file but it also has '.txt' and '.csv' file which when passing through the particular function which is meant for only '.ogg' file throws an error.<br>\nI fixed it by setting a filter for '.ogg' file and submitted it successfully.<br>\nThanks for the help, sir. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1268878,
          "author_name": "andrescastroitm",
          "author_url": "",
          "post_date": "04/09/2021 22:21:40",
          "content": "<p>Hi, I have a similar problem, but I already filter for '.ogg' files with </p>\n<p><code>all_audios = list(DATADIR.glob(\"*.ogg\"))</code></p>\n<p>How did you do that?</p>\n<p>Thanks in advance!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1263432": "I have submitted my notebook which has a submission file and in the proper format but while running for the private dataset it's showing  **Submission CSV Not Found**. I have no clue why this happening. \nThis is how I am switching my directory from public to private\n\n> TEST = (len(list(Path(\"../input/birdclef-2021/test_soundscapes/\").glob(\"*.ogg\"))) != 0)\n\n>if TEST:\n>           DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")\n>else:\n>         DATADIR = Path(\"../input/birdclef-2021/train_soundscapes/\")\n\nAnd this is a **submission.csv** file format. Please help\n\trow_id\tbirds\n0\t20152_SSW_5\tnocall\n1\t20152_SSW_10\tnocall\n2\t20152_SSW_15\tnocall\n3\t20152_SSW_20\tnocall\n4\t20152_SSW_25\tnocall\n...\t...\t...\n2395\t26709_SSW_580\tnocall\n2396\t26709_SSW_585\tnocall\n2397\t26709_SSW_590\tnocall\n2398\t26709_SSW_595\tnocall\n2399\t26709_SSW_600\tnocall",
    "1263961": "Hi Nikhil,\n\nThe contents of your `submission.csv` isn't in the correct format.  Specifically, you shouldn't have 1, 2, 3... at the start of the rows.  Assuming you're using pandas to generate this, you need to add `index=False` when you call `.to_csv`.  For example, if your dataframe is called `df_test` then you should use...\n\n`df_test[[\"row_id\", \"birds\"]].to_csv(\"submission.csv\", index=False)`\n\nLet me know if you're still having problems after that.\n\nAndrew",
    "1264038": "andrewrrose Hi thanks for the reply.\nThis is my code for df to csv  \n> prediction_df.to_csv(\"submission.csv\", index=False)\n\nI already set index = False before posting this issue. I guess there is some other issue. Let me know you want anything which can help to resolve this issue.",
    "1264369": "Three thoughts then...\n\n- Can you post you actual submission.csv, rather than the dataframe.\n- Try hard-coding `TEST = True` in case there's something wrong with your auto-detection of which environment you're running in.\n- If it really can't find `submission.csv`, your code is probably crashing / running out of memory when being run for real.  See [Kaggle's debugging guide](https://www.kaggle.com/code-competition-debugging) although mostly it just says that you need to do careful debugging.",
    "1264532": "Solved it. I am getting issue because of this   \n> DATADIR = Path(\"../input/birdclef-2021/test_soundscapes/\")\n\nI was assuming this test dir only contains  '.ogg'  file but it also has '.txt' and '.csv' file which when passing through the particular function which is meant for only '.ogg' file throws an error.\nI fixed it by setting a filter for '.ogg' file and submitted it successfully.\nThanks for the help, sir.",
    "1268878": "Hi, I have a similar problem, but I already filter for '.ogg' files with \n\n`all_audios = list(DATADIR.glob(\"*.ogg\"))`\n\nHow did you do that?\n\nThanks in advance!"
  },
  "source": "meta"
}