{
  "id": 227225,
  "title": "For those Submission Scoring Error, Endless running of submission, Submission CSV Not Found",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/227225",
  "author_name": "",
  "post_date": "2021-03-19T11:24:27.429273500Z",
  "votes": 16,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I have submitted a notebook running for more than 5 days despite the error message like \"Submission CSV Not Found,\" \"Submission Scoring Error,\" etc.</p>\n<p>After my investigation, finally, I make a successful kernel submission.<br>\n<strong>1.</strong> The reason for the \"Submission Scoring Error\" is because the submission CSV file did not follow the scoring order.  To solve it, you need to read the sample CSV from ../input/hubmap-kidney-segmentation/test/sample_submission.csv as your template.</p>\n<p>For example, my submission with \"Submission Scoring Error\"  showed my predicted order is <br>\naa05346ff<br>\n2ec3f1bb9<br>\n3589adb90<br>\nd488c759a<br>\n57512b7f1</p>\n<p>But the right predicted order should be <br>\n2ec3f1bb9<br>\n3589adb90<br>\nd488c759a<br>\naa05346ff<br>\n57512b7f1</p>\n<p><strong>2.</strong> Submission CSV Not Found. This is because the test names of the images from the public dataset and the private are different. Previously,  I used </p>\n<p>\"p = pathlib.Path(DATA_PATH)<br>\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff'))):\"</p>\n<p>to get the test tiff image. However, it seems that the images in the private dataset did not follow the rules used in the public dataset.</p>\n<p>Thus, please read sample_submission.csv from the test as the reference.</p>\n<p>If you did not follow the rules, then your kernel submission will fall into endless running.</p>\n<p>Check the working notebook <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub</a> for example that is written by Iafoss, <a href=\"https://www.kaggle.com/tikutiku\" target=\"_blank\">https://www.kaggle.com/tikutiku</a> </p>\n<p>Thanks very much, Iafoss. <br>\nYour notebook leads to my successful submission.</p>",
  "messages": [
    {
      "id": "1244981",
      "postDate": "03/19/2021 11:24:27",
      "content": "<p>I have submitted a notebook running for more than 5 days despite the error message like \"Submission CSV Not Found,\" \"Submission Scoring Error,\" etc.</p>\n<p>After my investigation, finally, I make a successful kernel submission.<br>\n<strong>1.</strong> The reason for the \"Submission Scoring Error\" is because the submission CSV file did not follow the scoring order.  To solve it, you need to read the sample CSV from ../input/hubmap-kidney-segmentation/test/sample_submission.csv as your template.</p>\n<p>For example, my submission with \"Submission Scoring Error\"  showed my predicted order is <br>\naa05346ff<br>\n2ec3f1bb9<br>\n3589adb90<br>\nd488c759a<br>\n57512b7f1</p>\n<p>But the right predicted order should be <br>\n2ec3f1bb9<br>\n3589adb90<br>\nd488c759a<br>\naa05346ff<br>\n57512b7f1</p>\n<p><strong>2.</strong> Submission CSV Not Found. This is because the test names of the images from the public dataset and the private are different. Previously,  I used </p>\n<p>\"p = pathlib.Path(DATA_PATH)<br>\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff'))):\"</p>\n<p>to get the test tiff image. However, it seems that the images in the private dataset did not follow the rules used in the public dataset.</p>\n<p>Thus, please read sample_submission.csv from the test as the reference.</p>\n<p>If you did not follow the rules, then your kernel submission will fall into endless running.</p>\n<p>Check the working notebook <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub</a> for example that is written by Iafoss, <a href=\"https://www.kaggle.com/tikutiku\" target=\"_blank\">https://www.kaggle.com/tikutiku</a> </p>\n<p>Thanks very much, Iafoss. <br>\nYour notebook leads to my successful submission.</p>",
      "rawMarkdown": "I have submitted a notebook running for more than 5 days despite the error message like \"Submission CSV Not Found,\" \"Submission Scoring Error,\" etc.\n\nAfter my investigation, finally, I make a successful kernel submission.\n**1.** The reason for the \"Submission Scoring Error\" is because the submission CSV file did not follow the scoring order.  To solve it, you need to read the sample CSV from ../input/hubmap-kidney-segmentation/test/sample_submission.csv as your template.\n\nFor example, my submission with \"Submission Scoring Error\"  showed my predicted order is \naa05346ff\n2ec3f1bb9\n3589adb90\nd488c759a\n57512b7f1\n\nBut the right predicted order should be \n2ec3f1bb9\n3589adb90\nd488c759a\naa05346ff\n57512b7f1\n\n**2.** Submission CSV Not Found. This is because the test names of the images from the public dataset and the private are different. Previously,  I used \n\n\"p = pathlib.Path(DATA_PATH)\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff'))):\"\n\nto get the test tiff image. However, it seems that the images in the private dataset did not follow the rules used in the public dataset.\n\nThus, please read sample_submission.csv from the test as the reference.\n\nIf you did not follow the rules, then your kernel submission will fall into endless running.\n\nCheck the working notebook https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub for example that is written by Iafoss, https://www.kaggle.com/tikutiku \n\nThanks very much, Iafoss. \nYour notebook leads to my successful submission.",
      "votes": null
    },
    {
      "id": "1244998",
      "postDate": "03/19/2021 11:41:49",
      "content": "<p><a href=\"https://www.kaggle.com/goldgarudagoldgaruda\" target=\"_blank\">@goldgarudagoldgaruda</a> Congrats to solve the problem, but the notebook is written by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> :)</p>",
      "rawMarkdown": "goldgarudagoldgaruda Congrats to solve the problem, but the notebook is written by @iafoss :)",
      "votes": null
    },
    {
      "id": "1244999",
      "postDate": "03/19/2021 11:43:43",
      "content": "<p>Thanks. Updated. :)</p>",
      "rawMarkdown": "Thanks. Updated. :)",
      "votes": null
    },
    {
      "id": "1247405",
      "postDate": "03/21/2021 17:24:01",
      "content": "<p>Thanks all! This was driving me crazy. If you use glob to get the public test tiffs and ignore the private ones it works. So couldn't figure out what was going on.</p>\n<p>Also there are 2 sample_submission.csv files both seemed to work for me. Now I can start trying to use the new data instead of trying to get my submissions to work :)</p>",
      "rawMarkdown": "Thanks all! This was driving me crazy. If you use glob to get the public test tiffs and ignore the private ones it works. So couldn't figure out what was going on.\n\nAlso there are 2 sample_submission.csv files both seemed to work for me. Now I can start trying to use the new data instead of trying to get my submissions to work :)",
      "votes": null
    },
    {
      "id": "1250267",
      "postDate": "03/23/2021 22:26:31",
      "content": "<p>For me its still not working</p>",
      "rawMarkdown": "For me its still not working",
      "votes": null
    },
    {
      "id": "1250831",
      "postDate": "03/24/2021 10:01:52",
      "content": "<p>We just solved this issue now. But still we don't know why this code is working suddenly</p>\n<ol>\n<li>memory optimization</li>\n<li>save sample_submission per sample idx(seems like weird but just share it for whom need to try)</li>\n</ol>\n<pre><code>for step,idx enumerate(test_files):\n\n    sample_submission.reset_index().to_csv('/kaggle/working/submission.csv',index=False)\n    sample_submission.reset_index().to_csv('submission.csv', index = False)\n\nsample_submission = sample_submission.reset_index()\nsample_submission.to_csv('/kaggle/working/submission.csv',index=False)\nsample_submission.to_csv('submission.csv', index = False)\n</code></pre>",
      "rawMarkdown": "We just solved this issue now. But still we don't know why this code is working suddenly\n\n1. memory optimization\n2. save sample_submission per sample idx(seems like weird but just share it for whom need to try)\n\n```\nfor step,idx enumerate(test_files):\n\n    sample_submission.reset_index().to_csv('/kaggle/working/submission.csv',index=False)\n    sample_submission.reset_index().to_csv('submission.csv', index = False)\n\nsample_submission = sample_submission.reset_index()\nsample_submission.to_csv('/kaggle/working/submission.csv',index=False)\nsample_submission.to_csv('submission.csv', index = False)\n```",
      "votes": null
    },
    {
      "id": "1250917",
      "postDate": "03/24/2021 11:00:25",
      "content": "<p>Perhaps, your notebook collapsed when dealing with the private tiff. So optimize your memory at first. I suspect that there may be larger tiff images in the private dataset.</p>",
      "rawMarkdown": "Perhaps, your notebook collapsed when dealing with the private tiff. So optimize your memory at first. I suspect that there may be larger tiff images in the private dataset.",
      "votes": null
    },
    {
      "id": "1251426",
      "postDate": "03/24/2021 18:56:54",
      "content": "<p>Its probably not that, because the error is \"Notebook Threw Exception\". I had this memory problem before in the previous dataset, and i solved that mapping the image from disk. So actually im not dealing with the whole image in memory. Using try catch i discovered that the problem is in the function tifffile.imread(), so probably is not reading the file, but im using this sample_submission.csv as  template.</p>",
      "rawMarkdown": "Its probably not that, because the error is \"Notebook Threw Exception\". I had this memory problem before in the previous dataset, and i solved that mapping the image from disk. So actually im not dealing with the whole image in memory. Using try catch i discovered that the problem is in the function tifffile.imread(), so probably is not reading the file, but im using this sample_submission.csv as  template.",
      "votes": null
    },
    {
      "id": "1252785",
      "postDate": "03/26/2021 03:44:05",
      "content": "<p>I trained a different model, then followed the submission code by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, but still get the 'Submission Scoring Error'. The scoring order is right, so I don't know what is the problem. May be I try to use  both train and test code by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>. </p>",
      "rawMarkdown": "I trained a different model, then followed the submission code by @iafoss, but still get the 'Submission Scoring Error'. The scoring order is right, so I don't know what is the problem. May be I try to use  both train and test code by @iafoss.",
      "votes": null
    },
    {
      "id": "1254574",
      "postDate": "03/27/2021 20:27:27",
      "content": "<p>Have you been able to fix the issue?</p>",
      "rawMarkdown": "Have you been able to fix the issue?",
      "votes": null
    },
    {
      "id": "1257074",
      "postDate": "03/30/2021 13:56:23",
      "content": "<p>I am busy at last weekend, so I can't try other ways to solve it. May be I will try again in next days.</p>",
      "rawMarkdown": "I am busy at last weekend, so I can't try other ways to solve it. May be I will try again in next days.",
      "votes": null
    },
    {
      "id": "1260877",
      "postDate": "04/02/2021 13:38:00",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/goldgarudagoldgaruda\" target=\"_blank\">@goldgarudagoldgaruda</a> for pointing to solution and <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for the notebook!</p>\n<p>Eventually I was able to solve the \"endless submission\" problem using following:</p>\n<ul>\n<li>proper reading images with <em>rasterio</em> package.</li>\n<li>prediction order based on <em>sample_submission.csv</em>.</li>\n</ul>",
      "rawMarkdown": "Thanks a lot @goldgarudagoldgaruda for pointing to solution and @iafoss for the notebook!\n\nEventually I was able to solve the \"endless submission\" problem using following:\n- proper reading images with *rasterio* package.\n- prediction order based on *sample_submission.csv*.",
      "votes": null
    },
    {
      "id": "1261339",
      "postDate": "04/03/2021 01:09:36",
      "content": "<p>You are very welcome, I'm happy that it is solved.</p>",
      "rawMarkdown": "You are very welcome, I'm happy that it is solved.",
      "votes": null
    },
    {
      "id": "1265821",
      "postDate": "04/07/2021 08:28:02",
      "content": "<p><a href=\"https://www.kaggle.com/goldgarudagoldgaruda\" target=\"_blank\">@goldgarudagoldgaruda</a>  i follow the order same as there in sample submission. <br>\nI get sub scoring error after an hour time. <br>\nCan there be GPU out of memory error ?</p>",
      "rawMarkdown": "goldgarudagoldgaruda  i follow the order same as there in sample submission. \nI get sub scoring error after an hour time. \nCan there be GPU out of memory error ?",
      "votes": null
    },
    {
      "id": "1266024",
      "postDate": "04/07/2021 12:24:40",
      "content": "<p>That might be one possible reason. Be careful, there are two sample_submission.csv files under 2 different folders.  sub scoring error often means your csv file is not satisfying the scoring rule.</p>",
      "rawMarkdown": "That might be one possible reason. Be careful, there are two sample_submission.csv files under 2 different folders.  sub scoring error often means your csv file is not satisfying the scoring rule.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1244998,
      "author_name": "tikutiku",
      "author_url": "",
      "post_date": "03/19/2021 11:41:49",
      "content": "<p><a href=\"https://www.kaggle.com/goldgarudagoldgaruda\" target=\"_blank\">@goldgarudagoldgaruda</a> Congrats to solve the problem, but the notebook is written by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1244999,
      "author_name": "goldgarudagoldgaruda",
      "author_url": "",
      "post_date": "03/19/2021 11:43:43",
      "content": "<p>Thanks. Updated. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1247405,
      "author_name": "governor",
      "author_url": "",
      "post_date": "03/21/2021 17:24:01",
      "content": "<p>Thanks all! This was driving me crazy. If you use glob to get the public test tiffs and ignore the private ones it works. So couldn't figure out what was going on.</p>\n<p>Also there are 2 sample_submission.csv files both seemed to work for me. Now I can start trying to use the new data instead of trying to get my submissions to work :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1250267,
      "author_name": "joaorodriguez",
      "author_url": "",
      "post_date": "03/23/2021 22:26:31",
      "content": "<p>For me its still not working</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250917,
          "author_name": "goldgarudagoldgaruda",
          "author_url": "",
          "post_date": "03/24/2021 11:00:25",
          "content": "<p>Perhaps, your notebook collapsed when dealing with the private tiff. So optimize your memory at first. I suspect that there may be larger tiff images in the private dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251426,
          "author_name": "joaorodriguez",
          "author_url": "",
          "post_date": "03/24/2021 18:56:54",
          "content": "<p>Its probably not that, because the error is \"Notebook Threw Exception\". I had this memory problem before in the previous dataset, and i solved that mapping the image from disk. So actually im not dealing with the whole image in memory. Using try catch i discovered that the problem is in the function tifffile.imread(), so probably is not reading the file, but im using this sample_submission.csv as  template.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1250831,
      "author_name": "jinssaa",
      "author_url": "",
      "post_date": "03/24/2021 10:01:52",
      "content": "<p>We just solved this issue now. But still we don't know why this code is working suddenly</p>\n<ol>\n<li>memory optimization</li>\n<li>save sample_submission per sample idx(seems like weird but just share it for whom need to try)</li>\n</ol>\n<pre><code>for step,idx enumerate(test_files):\n\n    sample_submission.reset_index().to_csv('/kaggle/working/submission.csv',index=False)\n    sample_submission.reset_index().to_csv('submission.csv', index = False)\n\nsample_submission = sample_submission.reset_index()\nsample_submission.to_csv('/kaggle/working/submission.csv',index=False)\nsample_submission.to_csv('submission.csv', index = False)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1252785,
      "author_name": "ntthientn",
      "author_url": "",
      "post_date": "03/26/2021 03:44:05",
      "content": "<p>I trained a different model, then followed the submission code by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, but still get the 'Submission Scoring Error'. The scoring order is right, so I don't know what is the problem. May be I try to use  both train and test code by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1254574,
          "author_name": "nauyan",
          "author_url": "",
          "post_date": "03/27/2021 20:27:27",
          "content": "<p>Have you been able to fix the issue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1257074,
          "author_name": "ntthientn",
          "author_url": "",
          "post_date": "03/30/2021 13:56:23",
          "content": "<p>I am busy at last weekend, so I can't try other ways to solve it. May be I will try again in next days.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1260877,
      "author_name": "eytankats",
      "author_url": "",
      "post_date": "04/02/2021 13:38:00",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/goldgarudagoldgaruda\" target=\"_blank\">@goldgarudagoldgaruda</a> for pointing to solution and <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for the notebook!</p>\n<p>Eventually I was able to solve the \"endless submission\" problem using following:</p>\n<ul>\n<li>proper reading images with <em>rasterio</em> package.</li>\n<li>prediction order based on <em>sample_submission.csv</em>.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1261339,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "04/03/2021 01:09:36",
          "content": "<p>You are very welcome, I'm happy that it is solved.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1265821,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "04/07/2021 08:28:02",
      "content": "<p><a href=\"https://www.kaggle.com/goldgarudagoldgaruda\" target=\"_blank\">@goldgarudagoldgaruda</a>  i follow the order same as there in sample submission. <br>\nI get sub scoring error after an hour time. <br>\nCan there be GPU out of memory error ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1266024,
          "author_name": "goldgarudagoldgaruda",
          "author_url": "",
          "post_date": "04/07/2021 12:24:40",
          "content": "<p>That might be one possible reason. Be careful, there are two sample_submission.csv files under 2 different folders.  sub scoring error often means your csv file is not satisfying the scoring rule.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1244981": "I have submitted a notebook running for more than 5 days despite the error message like \"Submission CSV Not Found,\" \"Submission Scoring Error,\" etc.\n\nAfter my investigation, finally, I make a successful kernel submission.\n**1.** The reason for the \"Submission Scoring Error\" is because the submission CSV file did not follow the scoring order.  To solve it, you need to read the sample CSV from ../input/hubmap-kidney-segmentation/test/sample_submission.csv as your template.\n\nFor example, my submission with \"Submission Scoring Error\"  showed my predicted order is \naa05346ff\n2ec3f1bb9\n3589adb90\nd488c759a\n57512b7f1\n\nBut the right predicted order should be \n2ec3f1bb9\n3589adb90\nd488c759a\naa05346ff\n57512b7f1\n\n**2.** Submission CSV Not Found. This is because the test names of the images from the public dataset and the private are different. Previously,  I used \n\n\"p = pathlib.Path(DATA_PATH)\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff'))):\"\n\nto get the test tiff image. However, it seems that the images in the private dataset did not follow the rules used in the public dataset.\n\nThus, please read sample_submission.csv from the test as the reference.\n\nIf you did not follow the rules, then your kernel submission will fall into endless running.\n\nCheck the working notebook https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub for example that is written by Iafoss, https://www.kaggle.com/tikutiku \n\nThanks very much, Iafoss. \nYour notebook leads to my successful submission.",
    "1244998": "goldgarudagoldgaruda Congrats to solve the problem, but the notebook is written by @iafoss :)",
    "1244999": "Thanks. Updated. :)",
    "1247405": "Thanks all! This was driving me crazy. If you use glob to get the public test tiffs and ignore the private ones it works. So couldn't figure out what was going on.\n\nAlso there are 2 sample_submission.csv files both seemed to work for me. Now I can start trying to use the new data instead of trying to get my submissions to work :)",
    "1250267": "For me its still not working",
    "1250831": "We just solved this issue now. But still we don't know why this code is working suddenly\n\n1. memory optimization\n2. save sample_submission per sample idx(seems like weird but just share it for whom need to try)\n\n```\nfor step,idx enumerate(test_files):\n\n    sample_submission.reset_index().to_csv('/kaggle/working/submission.csv',index=False)\n    sample_submission.reset_index().to_csv('submission.csv', index = False)\n\nsample_submission = sample_submission.reset_index()\nsample_submission.to_csv('/kaggle/working/submission.csv',index=False)\nsample_submission.to_csv('submission.csv', index = False)\n```",
    "1250917": "Perhaps, your notebook collapsed when dealing with the private tiff. So optimize your memory at first. I suspect that there may be larger tiff images in the private dataset.",
    "1251426": "Its probably not that, because the error is \"Notebook Threw Exception\". I had this memory problem before in the previous dataset, and i solved that mapping the image from disk. So actually im not dealing with the whole image in memory. Using try catch i discovered that the problem is in the function tifffile.imread(), so probably is not reading the file, but im using this sample_submission.csv as  template.",
    "1252785": "I trained a different model, then followed the submission code by @iafoss, but still get the 'Submission Scoring Error'. The scoring order is right, so I don't know what is the problem. May be I try to use  both train and test code by @iafoss.",
    "1254574": "Have you been able to fix the issue?",
    "1257074": "I am busy at last weekend, so I can't try other ways to solve it. May be I will try again in next days.",
    "1260877": "Thanks a lot @goldgarudagoldgaruda for pointing to solution and @iafoss for the notebook!\n\nEventually I was able to solve the \"endless submission\" problem using following:\n- proper reading images with *rasterio* package.\n- prediction order based on *sample_submission.csv*.",
    "1261339": "You are very welcome, I'm happy that it is solved.",
    "1265821": "goldgarudagoldgaruda  i follow the order same as there in sample submission. \nI get sub scoring error after an hour time. \nCan there be GPU out of memory error ?",
    "1266024": "That might be one possible reason. Be careful, there are two sample_submission.csv files under 2 different folders.  sub scoring error often means your csv file is not satisfying the scoring rule."
  },
  "source": "meta"
}