{
  "id": 166317,
  "title": "Is there any way to get information on why a submission failed? (Subission Scoring Error)",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/166317",
  "author_name": "",
  "post_date": "2020-07-12T13:29:45.970008400Z",
  "votes": 6,
  "comment_count": 14,
  "views": 0,
  "content": "<p>My submissions are repeatedly failing, and I have no idea why. There is no error information reported. </p>\n\n<p>I've checked everything I can think to check, the memory usage is smallish (though I do load all of a single patient's DCM files at a time, never more than one patient is loaded at any given time), the execution time on the test set is minimal (about a minute, though submissions take longer to run - about half an hour), the submitted CSV is identical in structure to the example submission.</p>\n\n<p>I'm using a keras framework model which I'm training on my local machine and uploading to kaggle as a dataset. A variety of safe-guards are in place for encountering a corrupted DCM file.</p>\n\n<p>The only thing I can think of is that my pre-trained model is maybe not loaded in the submission environmnet, but I have no way to verify that.</p>\n\n<p>Everything works exactly as expected when run as a notebook, it only fails with no error information when I submit with \"Submission Scoring Error\".</p>\n\n<p>EDIT: Solution found, the range of results has to be <em>exactly</em> -12 to 133 weeks. Not one more or less.</p>",
  "messages": [
    {
      "id": "926075",
      "postDate": "07/12/2020 13:29:45",
      "content": "<p>My submissions are repeatedly failing, and I have no idea why. There is no error information reported. </p>\n\n<p>I've checked everything I can think to check, the memory usage is smallish (though I do load all of a single patient's DCM files at a time, never more than one patient is loaded at any given time), the execution time on the test set is minimal (about a minute, though submissions take longer to run - about half an hour), the submitted CSV is identical in structure to the example submission.</p>\n\n<p>I'm using a keras framework model which I'm training on my local machine and uploading to kaggle as a dataset. A variety of safe-guards are in place for encountering a corrupted DCM file.</p>\n\n<p>The only thing I can think of is that my pre-trained model is maybe not loaded in the submission environmnet, but I have no way to verify that.</p>\n\n<p>Everything works exactly as expected when run as a notebook, it only fails with no error information when I submit with \"Submission Scoring Error\".</p>\n\n<p>EDIT: Solution found, the range of results has to be <em>exactly</em> -12 to 133 weeks. Not one more or less.</p>",
      "rawMarkdown": "My submissions are repeatedly failing, and I have no idea why. There is no error information reported. \n\nI've checked everything I can think to check, the memory usage is smallish (though I do load all of a single patient's DCM files at a time, never more than one patient is loaded at any given time), the execution time on the test set is minimal (about a minute, though submissions take longer to run - about half an hour), the submitted CSV is identical in structure to the example submission.\n\nI'm using a keras framework model which I'm training on my local machine and uploading to kaggle as a dataset. A variety of safe-guards are in place for encountering a corrupted DCM file.\n\nThe only thing I can think of is that my pre-trained model is maybe not loaded in the submission environmnet, but I have no way to verify that.\n\nEverything works exactly as expected when run as a notebook, it only fails with no error information when I submit with \"Submission Scoring Error\".\n\n\nEDIT: Solution found, the range of results has to be *exactly* -12 to 133 weeks. Not one more or less.",
      "votes": null
    },
    {
      "id": "926492",
      "postDate": "07/12/2020 18:56:24",
      "content": "<p>Have you tried to run your code in the interactive session?</p>",
      "rawMarkdown": "Have you tried to run your code in the interactive session?",
      "votes": null
    },
    {
      "id": "926545",
      "postDate": "07/12/2020 19:17:39",
      "content": "<p>Hi, i get the \"Submission Scoring Error\" too. In my case even submitting the \"sample_submission.csv\" gives me this error.</p>",
      "rawMarkdown": "Hi, i get the \"Submission Scoring Error\" too. In my case even submitting the \"sample_submission.csv\" gives me this error.",
      "votes": null
    },
    {
      "id": "926664",
      "postDate": "07/12/2020 21:12:45",
      "content": "<p>Yes, it runs totally normally and produces the expected results.</p>",
      "rawMarkdown": "Yes, it runs totally normally and produces the expected results.",
      "votes": null
    },
    {
      "id": "927588",
      "postDate": "07/13/2020 13:53:33",
      "content": "<p>From Code Requirements: \"Submission file must be named submission.csv\".</p>",
      "rawMarkdown": "From Code Requirements: \"Submission file must be named submission.csv\".",
      "votes": null
    },
    {
      "id": "927590",
      "postDate": "07/13/2020 13:53:58",
      "content": "<p>Is your file called \"submission.csv\"?</p>",
      "rawMarkdown": "Is your file called \"submission.csv\"?",
      "votes": null
    },
    {
      "id": "927643",
      "postDate": "07/13/2020 14:16:48",
      "content": "<p>It is. I have noticed there are a couple of negative FCV values at higher week numbers, which is obviously not right. I'm testing to see if they were the cause of the error.</p>\n\n<p>Edit: They were not.</p>",
      "rawMarkdown": "It is. I have noticed there are a couple of negative FCV values at higher week numbers, which is obviously not right. I'm testing to see if they were the cause of the error.\n\nEdit: They were not.",
      "votes": null
    },
    {
      "id": "927674",
      "postDate": "07/13/2020 14:31:51",
      "content": "<p>We unfortunately can't provide highly specific debugging messages. This is indeed frustrating, but it's intentionally designed this way to prevent people from leaking information about the private test set. However, if you get really stuck, you can still treat this like a debugging problem. I've seen these tips help others in the past:</p>\n\n<ol>\n<li>Step away from the code, sleep or go for a walk, get your mind off it, then come back and examine with fresh eyes.</li>\n<li>Use a submission to verify basic fundamentals are working. E.g. you can test if your pretrained model is present by trying to load it and submit the sample submission if it succeeds, and erroring if it fails. You can also devise ways to communicate with yourself via the score or timing of your submission (e.g. intentionally submitting a <em>very</em> bad submission or <code>sleep()</code>ing a set amount of time).</li>\n<li>You can use function decorators (see <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/120375#688496\">here</a> for an example) to make your code extra robust. Work towards something that fails gracefully and still produces a score, then keep analyzing scores to determine just how often you think the code might be failing gracefully.</li>\n<li>Build sanity checks into your pipeline - is the submission you produce the same number of lines as the sample submission? Is everything that should be positive actually positive? etc.</li>\n<li>Sometimes it is easiest to start with the sample submission, as opposed to trying to recreate it by working from the test set. It can be easy to drop a row amidst all the loops, joins, grouping, and image processing.</li>\n<li>If all else fails, tear your pipeline down and rebuild it in a way that starts with a valid submission (such as submitting the sample submission), adds components one at a time, and verifies the output is still scoring. It's very common that an error is really trivial or early in the code, and once you correct it the entire pipeline is back to being functional.</li>\n</ol>\n\n<p>Writing code that works perfectly on unseen data is hard, so don't get discouraged or feel that you're the only one stuck on this. Keep at it and good luck!</p>",
      "rawMarkdown": "We unfortunately can't provide highly specific debugging messages. This is indeed frustrating, but it's intentionally designed this way to prevent people from leaking information about the private test set. However, if you get really stuck, you can still treat this like a debugging problem. I've seen these tips help others in the past:\n\n1. Step away from the code, sleep or go for a walk, get your mind off it, then come back and examine with fresh eyes.\n2. Use a submission to verify basic fundamentals are working. E.g. you can test if your pretrained model is present by trying to load it and submit the sample submission if it succeeds, and erroring if it fails. You can also devise ways to communicate with yourself via the score or timing of your submission (e.g. intentionally submitting a _very_ bad submission or `sleep()`ing a set amount of time).\n3. You can use function decorators (see [here](https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/120375#688496) for an example) to make your code extra robust. Work towards something that fails gracefully and still produces a score, then keep analyzing scores to determine just how often you think the code might be failing gracefully.\n4. Build sanity checks into your pipeline - is the submission you produce the same number of lines as the sample submission? Is everything that should be positive actually positive? etc.\n5. Sometimes it is easiest to start with the sample submission, as opposed to trying to recreate it by working from the test set. It can be easy to drop a row amidst all the loops, joins, grouping, and image processing.\n6. If all else fails, tear your pipeline down and rebuild it in a way that starts with a valid submission (such as submitting the sample submission), adds components one at a time, and verifies the output is still scoring. It's very common that an error is really trivial or early in the code, and once you correct it the entire pipeline is back to being functional.\n\nWriting code that works perfectly on unseen data is hard, so don't get discouraged or feel that you're the only one stuck on this. Keep at it and good luck!",
      "votes": null
    },
    {
      "id": "928316",
      "postDate": "07/13/2020 21:37:09",
      "content": "<p>Hey,</p>\n\n<p>I am sorry to hear that you have encouraged an issue with submitting, that is always frustrating, however i believe sooner or later you will overcome this obstacle. </p>\n\n<p>Try to check that your submission doesn`t strongly depend on provided test set. I believe issue comes from leaderboard score.  Take into account test set is only used as an example to show how to prepare submission file, after submission your notebook is re-committed with unseen data and lb score is estimated. Try to be sure that output is going be the same independently from input.</p>",
      "rawMarkdown": "Hey,\n\nI am sorry to hear that you have encouraged an issue with submitting, that is always frustrating, however i believe sooner or later you will overcome this obstacle. \n\nTry to check that your submission doesn`t strongly depend on provided test set. I believe issue comes from leaderboard score.  Take into account test set is only used as an example to show how to prepare submission file, after submission your notebook is re-committed with unseen data and lb score is estimated. Try to be sure that output is going be the same independently from input.",
      "votes": null
    },
    {
      "id": "932269",
      "postDate": "07/16/2020 22:21:38",
      "content": "<p>Oh my god, I figured it out! The range has to be <em>exactly</em> -12 to 133 weeks. I had an off-by-one error giving out 134 weeks :D</p>",
      "rawMarkdown": "Oh my god, I figured it out! The range has to be *exactly* -12 to 133 weeks. I had an off-by-one error giving out 134 weeks :D",
      "votes": null
    },
    {
      "id": "947917",
      "postDate": "07/27/2020 14:55:38",
      "content": "<p>I fixed my similar error making csv without index\n.to_csv(\"submission.csv\", index=False)</p>",
      "rawMarkdown": "I fixed my similar error making csv without index\n.to_csv(\"submission.csv\", index=False)",
      "votes": null
    },
    {
      "id": "983789",
      "postDate": "08/24/2020 15:43:38",
      "content": "<p>Does the order of data in \"Patient_Week\" column matters ?<br>\ni.e. should the Patient_Week data be in same order as of sample_submission.csv ?</p>",
      "rawMarkdown": "Does the order of data in \"Patient_Week\" column matters ?\ni.e. should the Patient_Week data be in same order as of sample_submission.csv ?",
      "votes": null
    },
    {
      "id": "983813",
      "postDate": "08/24/2020 15:56:44",
      "content": "<p>Order does not matter (in 99.99% of competitions/metrics, we have an <code>Id</code> column that is used to join to the solution)</p>",
      "rawMarkdown": "Order does not matter (in 99.99% of competitions/metrics, we have an `Id` column that is used to join to the solution)",
      "votes": null
    },
    {
      "id": "983837",
      "postDate": "08/24/2020 16:16:07",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> . I tried all the options mentioned here, still getting submission error. <br>\nBut when submitting the \"sample_submission.csv\", it is successful.</p>",
      "rawMarkdown": "Thank you @wcukierski . I tried all the options mentioned here, still getting submission error. \nBut when submitting the \"sample_submission.csv\", it is successful.",
      "votes": null
    },
    {
      "id": "986260",
      "postDate": "08/26/2020 10:22:38",
      "content": "<p>Is it possible to get \"Submission Scoring Error\" by using external dataset or information ?</p>",
      "rawMarkdown": "Is it possible to get \"Submission Scoring Error\" by using external dataset or information ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 983789,
      "author_name": "ashutosh3060",
      "author_url": "",
      "post_date": "08/24/2020 15:43:38",
      "content": "<p>Does the order of data in \"Patient_Week\" column matters ?<br>\ni.e. should the Patient_Week data be in same order as of sample_submission.csv ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 983813,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "08/24/2020 15:56:44",
          "content": "<p>Order does not matter (in 99.99% of competitions/metrics, we have an <code>Id</code> column that is used to join to the solution)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983837,
          "author_name": "ashutosh3060",
          "author_url": "",
          "post_date": "08/24/2020 16:16:07",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> . I tried all the options mentioned here, still getting submission error. <br>\nBut when submitting the \"sample_submission.csv\", it is successful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 986260,
      "author_name": "ashutosh3060",
      "author_url": "",
      "post_date": "08/26/2020 10:22:38",
      "content": "<p>Is it possible to get \"Submission Scoring Error\" by using external dataset or information ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 926492,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "07/12/2020 18:56:24",
      "content": "<p>Have you tried to run your code in the interactive session?</p>",
      "votes": null,
      "replies": [
        {
          "id": 926664,
          "author_name": "amylizzle",
          "author_url": "",
          "post_date": "07/12/2020 21:12:45",
          "content": "<p>Yes, it runs totally normally and produces the expected results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 927590,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "07/13/2020 13:53:58",
          "content": "<p>Is your file called \"submission.csv\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 927643,
          "author_name": "amylizzle",
          "author_url": "",
          "post_date": "07/13/2020 14:16:48",
          "content": "<p>It is. I have noticed there are a couple of negative FCV values at higher week numbers, which is obviously not right. I'm testing to see if they were the cause of the error.</p>\n\n<p>Edit: They were not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926545,
      "author_name": "gragasml",
      "author_url": "",
      "post_date": "07/12/2020 19:17:39",
      "content": "<p>Hi, i get the \"Submission Scoring Error\" too. In my case even submitting the \"sample_submission.csv\" gives me this error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 927588,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "07/13/2020 13:53:33",
          "content": "<p>From Code Requirements: \"Submission file must be named submission.csv\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 927674,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "07/13/2020 14:31:51",
      "content": "<p>We unfortunately can't provide highly specific debugging messages. This is indeed frustrating, but it's intentionally designed this way to prevent people from leaking information about the private test set. However, if you get really stuck, you can still treat this like a debugging problem. I've seen these tips help others in the past:</p>\n\n<ol>\n<li>Step away from the code, sleep or go for a walk, get your mind off it, then come back and examine with fresh eyes.</li>\n<li>Use a submission to verify basic fundamentals are working. E.g. you can test if your pretrained model is present by trying to load it and submit the sample submission if it succeeds, and erroring if it fails. You can also devise ways to communicate with yourself via the score or timing of your submission (e.g. intentionally submitting a <em>very</em> bad submission or <code>sleep()</code>ing a set amount of time).</li>\n<li>You can use function decorators (see <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/120375#688496\">here</a> for an example) to make your code extra robust. Work towards something that fails gracefully and still produces a score, then keep analyzing scores to determine just how often you think the code might be failing gracefully.</li>\n<li>Build sanity checks into your pipeline - is the submission you produce the same number of lines as the sample submission? Is everything that should be positive actually positive? etc.</li>\n<li>Sometimes it is easiest to start with the sample submission, as opposed to trying to recreate it by working from the test set. It can be easy to drop a row amidst all the loops, joins, grouping, and image processing.</li>\n<li>If all else fails, tear your pipeline down and rebuild it in a way that starts with a valid submission (such as submitting the sample submission), adds components one at a time, and verifies the output is still scoring. It's very common that an error is really trivial or early in the code, and once you correct it the entire pipeline is back to being functional.</li>\n</ol>\n\n<p>Writing code that works perfectly on unseen data is hard, so don't get discouraged or feel that you're the only one stuck on this. Keep at it and good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 928316,
      "author_name": "maksimbahdanchyk",
      "author_url": "",
      "post_date": "07/13/2020 21:37:09",
      "content": "<p>Hey,</p>\n\n<p>I am sorry to hear that you have encouraged an issue with submitting, that is always frustrating, however i believe sooner or later you will overcome this obstacle. </p>\n\n<p>Try to check that your submission doesn`t strongly depend on provided test set. I believe issue comes from leaderboard score.  Take into account test set is only used as an example to show how to prepare submission file, after submission your notebook is re-committed with unseen data and lb score is estimated. Try to be sure that output is going be the same independently from input.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 932269,
      "author_name": "amylizzle",
      "author_url": "",
      "post_date": "07/16/2020 22:21:38",
      "content": "<p>Oh my god, I figured it out! The range has to be <em>exactly</em> -12 to 133 weeks. I had an off-by-one error giving out 134 weeks :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 947917,
      "author_name": "kens41",
      "author_url": "",
      "post_date": "07/27/2020 14:55:38",
      "content": "<p>I fixed my similar error making csv without index\n.to_csv(\"submission.csv\", index=False)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "926075": "My submissions are repeatedly failing, and I have no idea why. There is no error information reported. \n\nI've checked everything I can think to check, the memory usage is smallish (though I do load all of a single patient's DCM files at a time, never more than one patient is loaded at any given time), the execution time on the test set is minimal (about a minute, though submissions take longer to run - about half an hour), the submitted CSV is identical in structure to the example submission.\n\nI'm using a keras framework model which I'm training on my local machine and uploading to kaggle as a dataset. A variety of safe-guards are in place for encountering a corrupted DCM file.\n\nThe only thing I can think of is that my pre-trained model is maybe not loaded in the submission environmnet, but I have no way to verify that.\n\nEverything works exactly as expected when run as a notebook, it only fails with no error information when I submit with \"Submission Scoring Error\".\n\n\nEDIT: Solution found, the range of results has to be *exactly* -12 to 133 weeks. Not one more or less.",
    "926492": "Have you tried to run your code in the interactive session?",
    "926545": "Hi, i get the \"Submission Scoring Error\" too. In my case even submitting the \"sample_submission.csv\" gives me this error.",
    "926664": "Yes, it runs totally normally and produces the expected results.",
    "927588": "From Code Requirements: \"Submission file must be named submission.csv\".",
    "927590": "Is your file called \"submission.csv\"?",
    "927643": "It is. I have noticed there are a couple of negative FCV values at higher week numbers, which is obviously not right. I'm testing to see if they were the cause of the error.\n\nEdit: They were not.",
    "927674": "We unfortunately can't provide highly specific debugging messages. This is indeed frustrating, but it's intentionally designed this way to prevent people from leaking information about the private test set. However, if you get really stuck, you can still treat this like a debugging problem. I've seen these tips help others in the past:\n\n1. Step away from the code, sleep or go for a walk, get your mind off it, then come back and examine with fresh eyes.\n2. Use a submission to verify basic fundamentals are working. E.g. you can test if your pretrained model is present by trying to load it and submit the sample submission if it succeeds, and erroring if it fails. You can also devise ways to communicate with yourself via the score or timing of your submission (e.g. intentionally submitting a _very_ bad submission or `sleep()`ing a set amount of time).\n3. You can use function decorators (see [here](https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/120375#688496) for an example) to make your code extra robust. Work towards something that fails gracefully and still produces a score, then keep analyzing scores to determine just how often you think the code might be failing gracefully.\n4. Build sanity checks into your pipeline - is the submission you produce the same number of lines as the sample submission? Is everything that should be positive actually positive? etc.\n5. Sometimes it is easiest to start with the sample submission, as opposed to trying to recreate it by working from the test set. It can be easy to drop a row amidst all the loops, joins, grouping, and image processing.\n6. If all else fails, tear your pipeline down and rebuild it in a way that starts with a valid submission (such as submitting the sample submission), adds components one at a time, and verifies the output is still scoring. It's very common that an error is really trivial or early in the code, and once you correct it the entire pipeline is back to being functional.\n\nWriting code that works perfectly on unseen data is hard, so don't get discouraged or feel that you're the only one stuck on this. Keep at it and good luck!",
    "928316": "Hey,\n\nI am sorry to hear that you have encouraged an issue with submitting, that is always frustrating, however i believe sooner or later you will overcome this obstacle. \n\nTry to check that your submission doesn`t strongly depend on provided test set. I believe issue comes from leaderboard score.  Take into account test set is only used as an example to show how to prepare submission file, after submission your notebook is re-committed with unseen data and lb score is estimated. Try to be sure that output is going be the same independently from input.",
    "932269": "Oh my god, I figured it out! The range has to be *exactly* -12 to 133 weeks. I had an off-by-one error giving out 134 weeks :D",
    "947917": "I fixed my similar error making csv without index\n.to_csv(\"submission.csv\", index=False)",
    "983789": "Does the order of data in \"Patient_Week\" column matters ?\ni.e. should the Patient_Week data be in same order as of sample_submission.csv ?",
    "983813": "Order does not matter (in 99.99% of competitions/metrics, we have an `Id` column that is used to join to the solution)",
    "983837": "Thank you @wcukierski . I tried all the options mentioned here, still getting submission error. \nBut when submitting the \"sample_submission.csv\", it is successful.",
    "986260": "Is it possible to get \"Submission Scoring Error\" by using external dataset or information ?"
  },
  "source": "meta"
}