{
  "id": 203463,
  "title": "Submission Scoring Error",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/203463",
  "author_name": "",
  "post_date": "2020-12-15T11:49:51.883974400Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>unfortunately my submission failed and i dont know why. I get the error massage <strong>Submission Scoring Error</strong>.<br>\nI analysed the values (there are no missing values) and the shape of the output file. i found no problems. <br>\nMy notebook is <a href=\"https://www.kaggle.com/drcapa/catheter-line-position-eda-resnet50\" target=\"_blank\">this</a>.<br>\nAre there any advices to handle the error?</p>\n<p>Best regards</p>",
  "messages": [
    {
      "id": "1113359",
      "postDate": "12/15/2020 11:49:51",
      "content": "<p>Hi,</p>\n<p>unfortunately my submission failed and i dont know why. I get the error massage <strong>Submission Scoring Error</strong>.<br>\nI analysed the values (there are no missing values) and the shape of the output file. i found no problems. <br>\nMy notebook is <a href=\"https://www.kaggle.com/drcapa/catheter-line-position-eda-resnet50\" target=\"_blank\">this</a>.<br>\nAre there any advices to handle the error?</p>\n<p>Best regards</p>",
      "rawMarkdown": "Hi,\n\nunfortunately my submission failed and i dont know why. I get the error massage **Submission Scoring Error**.\nI analysed the values (there are no missing values) and the shape of the output file. i found no problems. \nMy notebook is [this](https://www.kaggle.com/drcapa/catheter-line-position-eda-resnet50).\nAre there any advices to handle the error?\n\nBest regards",
      "votes": null
    },
    {
      "id": "1113409",
      "postDate": "12/15/2020 12:24:32",
      "content": "<p>May be it is because of \"there is a hidden test set (approximately 4x larger, with ~14k images) as well\"…</p>",
      "rawMarkdown": "May be it is because of \"there is a hidden test set (approximately 4x larger, with ~14k images) as well\"...",
      "votes": null
    },
    {
      "id": "1117643",
      "postDate": "12/18/2020 09:41:55",
      "content": "<p>Please make sure that the count of submission is same as the test data count. Its more good to see the sample submission with same index and same columns. </p>",
      "rawMarkdown": "Please make sure that the count of submission is same as the test data count. Its more good to see the sample submission with same index and same columns.",
      "votes": null
    },
    {
      "id": "1124514",
      "postDate": "12/24/2020 02:29:05",
      "content": "<p>How was this issue finally resolved ? I am also getting the same issue and am unable to debug the exact issue in inference</p>",
      "rawMarkdown": "How was this issue finally resolved ? I am also getting the same issue and am unable to debug the exact issue in inference",
      "votes": null
    },
    {
      "id": "1124939",
      "postDate": "12/24/2020 09:31:56",
      "content": "<p>In my case the error based on a wrong number of samples in the submission file. The known test dataset and the hidden test dataset (used for scoring) have not the same number of samples. You have to make sure that your final prediction not depends on the number of sample of the known test dataset. </p>\n<p>You can use my notebook for example. My strategy:</p>\n<ol>\n<li>Use a DataGenerator to load the data on demand (to avoid RAM issues). Here the length of test data is considered automatically. Because of the batch size there could be samples with zero values.</li>\n<li>The predictions for the zero values are nan-values. These samples i have to drop:<br>\n<img src=\"https://i.ibb.co/WccwVqD/output.png\" alt=\"\"></li>\n</ol>\n<p>I hope it helps.</p>",
      "rawMarkdown": "In my case the error based on a wrong number of samples in the submission file. The known test dataset and the hidden test dataset (used for scoring) have not the same number of samples. You have to make sure that your final prediction not depends on the number of sample of the known test dataset. \n\nYou can use my notebook for example. My strategy:\n1. Use a DataGenerator to load the data on demand (to avoid RAM issues). Here the length of test data is considered automatically. Because of the batch size there could be samples with zero values.\n2. The predictions for the zero values are nan-values. These samples i have to drop:\n![](https://i.ibb.co/WccwVqD/output.png)\n\nI hope it helps.",
      "votes": null
    },
    {
      "id": "1125142",
      "postDate": "12/24/2020 12:14:11",
      "content": "<p>ok. Though I later found the issue in my case and resolved it. I had to remove \".jpg\" from the file names before adding them into the 'StudyInstanceUID' Column</p>",
      "rawMarkdown": "ok. Though I later found the issue in my case and resolved it. I had to remove \".jpg\" from the file names before adding them into the 'StudyInstanceUID' Column",
      "votes": null
    },
    {
      "id": "1125352",
      "postDate": "12/24/2020 16:38:09",
      "content": "<p>I agree with you. This is right method to crosscheck. <a href=\"https://www.kaggle.com/imranzaman5202\" target=\"_blank\">@imranzaman5202</a> </p>",
      "rawMarkdown": "I agree with you. This is right method to crosscheck. @imranzaman5202",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1113409,
      "author_name": "drcapa",
      "author_url": "",
      "post_date": "12/15/2020 12:24:32",
      "content": "<p>May be it is because of \"there is a hidden test set (approximately 4x larger, with ~14k images) as well\"…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1117643,
      "author_name": "",
      "author_url": "",
      "post_date": "12/18/2020 09:41:55",
      "content": "<p>Please make sure that the count of submission is same as the test data count. Its more good to see the sample submission with same index and same columns. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1125352,
          "author_name": "saurabhshahane",
          "author_url": "",
          "post_date": "12/24/2020 16:38:09",
          "content": "<p>I agree with you. This is right method to crosscheck. <a href=\"https://www.kaggle.com/imranzaman5202\" target=\"_blank\">@imranzaman5202</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1124514,
      "author_name": "kabhinay",
      "author_url": "",
      "post_date": "12/24/2020 02:29:05",
      "content": "<p>How was this issue finally resolved ? I am also getting the same issue and am unable to debug the exact issue in inference</p>",
      "votes": null,
      "replies": [
        {
          "id": 1124939,
          "author_name": "drcapa",
          "author_url": "",
          "post_date": "12/24/2020 09:31:56",
          "content": "<p>In my case the error based on a wrong number of samples in the submission file. The known test dataset and the hidden test dataset (used for scoring) have not the same number of samples. You have to make sure that your final prediction not depends on the number of sample of the known test dataset. </p>\n<p>You can use my notebook for example. My strategy:</p>\n<ol>\n<li>Use a DataGenerator to load the data on demand (to avoid RAM issues). Here the length of test data is considered automatically. Because of the batch size there could be samples with zero values.</li>\n<li>The predictions for the zero values are nan-values. These samples i have to drop:<br>\n<img src=\"https://i.ibb.co/WccwVqD/output.png\" alt=\"\"></li>\n</ol>\n<p>I hope it helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125142,
          "author_name": "kabhinay",
          "author_url": "",
          "post_date": "12/24/2020 12:14:11",
          "content": "<p>ok. Though I later found the issue in my case and resolved it. I had to remove \".jpg\" from the file names before adding them into the 'StudyInstanceUID' Column</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1113359": "Hi,\n\nunfortunately my submission failed and i dont know why. I get the error massage **Submission Scoring Error**.\nI analysed the values (there are no missing values) and the shape of the output file. i found no problems. \nMy notebook is [this](https://www.kaggle.com/drcapa/catheter-line-position-eda-resnet50).\nAre there any advices to handle the error?\n\nBest regards",
    "1113409": "May be it is because of \"there is a hidden test set (approximately 4x larger, with ~14k images) as well\"...",
    "1117643": "Please make sure that the count of submission is same as the test data count. Its more good to see the sample submission with same index and same columns.",
    "1124514": "How was this issue finally resolved ? I am also getting the same issue and am unable to debug the exact issue in inference",
    "1124939": "In my case the error based on a wrong number of samples in the submission file. The known test dataset and the hidden test dataset (used for scoring) have not the same number of samples. You have to make sure that your final prediction not depends on the number of sample of the known test dataset. \n\nYou can use my notebook for example. My strategy:\n1. Use a DataGenerator to load the data on demand (to avoid RAM issues). Here the length of test data is considered automatically. Because of the batch size there could be samples with zero values.\n2. The predictions for the zero values are nan-values. These samples i have to drop:\n![](https://i.ibb.co/WccwVqD/output.png)\n\nI hope it helps.",
    "1125142": "ok. Though I later found the issue in my case and resolved it. I had to remove \".jpg\" from the file names before adding them into the 'StudyInstanceUID' Column",
    "1125352": "I agree with you. This is right method to crosscheck. @imranzaman5202"
  },
  "source": "meta"
}