{
  "id": 231332,
  "title": "Help needed to debug \"Submission Scoring Error\"",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/231332",
  "author_name": "Strideradu",
  "post_date": "2021-04-08T03:20:14.587000",
  "votes": 3,
  "comment_count": 18,
  "views": 0,
  "content": "<p>My personal kernel keep got the \"Submission Scoring Error\" when I submit to the leaderboard (although there is no error from the public part). I am trying to figure out what is the root cause and think the forum may have some people who may have insights to help me.</p>\n<p>Potential cause:</p>\n<ol>\n<li>Out of time. Unlikely, the kernel is only using 4900s for the public eval dataset, so it is unlikely the public + private will take more than 9 hours (also wondering if there will be other error showing if out of time?)</li>\n<li>Memory issue? Personally, I have no idea whether the memory will become an issue. Not sure if someone has figured out a way to debug this?</li>\n<li>Format issue. Possible but I do have a successful submission (0.2x) when I filter the data using <code>sub_df.loc[sub_df['ImageWidth'] == 2048]</code>. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\" </li>\n</ol>\n<p>PS Wondering Kaggle and organizer can provide more support on the error for the private part. This is a kind of black box and hard to debug the issue. </p>",
  "messages": [
    {
      "id": 1266733,
      "postDate": "2021-04-08T03:20:14.587Z",
      "content": "<p>My personal kernel keep got the \"Submission Scoring Error\" when I submit to the leaderboard (although there is no error from the public part). I am trying to figure out what is the root cause and think the forum may have some people who may have insights to help me.</p>\n<p>Potential cause:</p>\n<ol>\n<li>Out of time. Unlikely, the kernel is only using 4900s for the public eval dataset, so it is unlikely the public + private will take more than 9 hours (also wondering if there will be other error showing if out of time?)</li>\n<li>Memory issue? Personally, I have no idea whether the memory will become an issue. Not sure if someone has figured out a way to debug this?</li>\n<li>Format issue. Possible but I do have a successful submission (0.2x) when I filter the data using <code>sub_df.loc[sub_df['ImageWidth'] == 2048]</code>. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\" </li>\n</ol>\n<p>PS Wondering Kaggle and organizer can provide more support on the error for the private part. This is a kind of black box and hard to debug the issue. </p>",
      "rawMarkdown": "My personal kernel keep got the \"Submission Scoring Error\" when I submit to the leaderboard (although there is no error from the public part). I am trying to figure out what is the root cause and think the forum may have some people who may have insights to help me.\n\nPotential cause:\n1. Out of time. Unlikely, the kernel is only using 4900s for the public eval dataset, so it is unlikely the public + private will take more than 9 hours (also wondering if there will be other error showing if out of time?)\n2. Memory issue? Personally, I have no idea whether the memory will become an issue. Not sure if someone has figured out a way to debug this?\n3. Format issue. Possible but I do have a successful submission (0.2x) when I filter the data using `sub_df.loc[sub_df['ImageWidth'] == 2048]`. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\" \n\nPS Wondering Kaggle and organizer can provide more support on the error for the private part. This is a kind of black box and hard to debug the issue. ",
      "votes": 3
    },
    {
      "id": 1274228,
      "postDate": "2021-04-15T05:55:33.960Z",
      "content": "<p>I got \"Submission Scoring Error\" when I submitted huge submission file. Approximately 8 GiB or so including public and private.</p>",
      "rawMarkdown": "I got \"Submission Scoring Error\" when I submitted huge submission file. Approximately 8 GiB or so including public and private.",
      "votes": 1
    },
    {
      "id": 1267949,
      "postDate": "2021-04-09T01:50:21.423Z",
      "content": "<p>Found 2 issues in my code</p>\n<ol>\n<li>I hardcoded mask size to be 2048 in the dataset although I have arg for this in init() … I guess this caused the scoring error</li>\n<li>Another issue I set padding = False and that will cause an issue when segment 1728 images<br>\nNow I am save and update new submission hope this can work</li>\n</ol>",
      "rawMarkdown": "Found 2 issues in my code\n1. I hardcoded mask size to be 2048 in the dataset although I have arg for this in init() ... I guess this caused the scoring error\n2. Another issue I set padding = False and that will cause an issue when segment 1728 images\nNow I am save and update new submission hope this can work",
      "votes": 1
    },
    {
      "id": 1266881,
      "postDate": "2021-04-08T06:24:35.767Z",
      "content": "<p>In my previous submission, I got the same error message if there are more than 1 mask for the same 1 cell.<br>\nI guess “1 mask for 1 cell” is required to be scored properly.<br>\nAnother possibility is the format of mask encoding string.</p>",
      "rawMarkdown": "In my previous submission, I got the same error message if there are more than 1 mask for the same 1 cell.\nI guess “1 mask for 1 cell” is required to be scored properly.\nAnother possibility is the format of mask encoding string.",
      "votes": 1,
      "replies": [
        {
          "id": 1266909,
          "postDate": "2021-04-08T07:01:26.637Z",
          "content": "<p>Sorry what do you mean by 1 mask for 1 cell? Can you give me an example of your error submission?</p>",
          "rawMarkdown": "Sorry what do you mean by 1 mask for 1 cell? Can you give me an example of your error submission?"
        },
        {
          "id": 1266945,
          "postDate": "2021-04-08T07:34:44.040Z",
          "content": "<p>For example, suppose that \"eNoLCAgIsAQABJ4Beg==\" is the mask of the most left upper corner cell of the image id X, we have to write \"0 0.001 eNoLCAgIsAQABJ4Beg== 1 0.1 eNoLCAgIsAQABJ4Beg== 2 0.005 eNoLCAgIsAQABJ4Beg==…and so on\"<br>\nI guess we shouldn't have another mask string for that cell, like \"0 0.001 eNoLCAgIsAQABJ4Beg== 1 0.1 HIAaagsCAgIssagJIANDIiOKFA== 2 0.005 eNoLCAgIsAQABJ4Beg==\".</p>",
          "rawMarkdown": "For example, suppose that \"eNoLCAgIsAQABJ4Beg==\" is the mask of the most left upper corner cell of the image id X, we have to write \"0 0.001 eNoLCAgIsAQABJ4Beg== 1 0.1 eNoLCAgIsAQABJ4Beg== 2 0.005 eNoLCAgIsAQABJ4Beg==...and so on\"\nI guess we shouldn't have another mask string for that cell, like \"0 0.001 eNoLCAgIsAQABJ4Beg== 1 0.1 HIAaagsCAgIssagJIANDIiOKFA== 2 0.005 eNoLCAgIsAQABJ4Beg==\".",
          "votes": 2
        },
        {
          "id": 1266962,
          "postDate": "2021-04-08T07:56:12.060Z",
          "content": "<p>I see your point. Let me check the code. But my code do work if i only select 2048 images, and after that what i did is just combine [2048] + [3072] + … and output to the df</p>",
          "rawMarkdown": "I see your point. Let me check the code. But my code do work if i only select 2048 images, and after that what i did is just combine [2048] + [3072] + ... and output to the df",
          "votes": 1
        },
        {
          "id": 1266965,
          "postDate": "2021-04-08T07:58:42.697Z",
          "content": "<p>I see. How about selecting only 1728 or 3072 images?</p>",
          "rawMarkdown": "I see. How about selecting only 1728 or 3072 images?"
        },
        {
          "id": 1267883,
          "postDate": "2021-04-08T22:50:44.453Z",
          "content": "<p>Seems the issue is from those images not equal to 2048. </p>",
          "rawMarkdown": "Seems the issue is from those images not equal to 2048. "
        }
      ]
    },
    {
      "id": 1274494,
      "postDate": "2021-04-15T11:05:20.030Z",
      "content": "<p>Hiii, I'm having the same issues.</p>\n<p>I am processing batches of images of the same size, but my notebook can't handle all the data, it takes too long.<br>\nThe only sumbission I achieved is running only the most frequent size images. Now I'm trying processing the most 2 frequent sizes, but let's see if I don't get a <code>Notebook Timeout</code>.</p>",
      "rawMarkdown": "Hiii, I'm having the same issues.\n\nI am processing batches of images of the same size, but my notebook can't handle all the data, it takes too long.\nThe only sumbission I achieved is running only the most frequent size images. Now I'm trying processing the most 2 frequent sizes, but let's see if I don't get a `Notebook Timeout`.",
      "replies": [
        {
          "id": 1278408,
          "postDate": "2021-04-19T21:07:35.403Z",
          "content": "<p>I think for segmentation you shouldn't get time out. I guess you may want to optimize your code (but I think you can also try fast submit which only process the public test dataset for now)</p>",
          "rawMarkdown": "I think for segmentation you shouldn't get time out. I guess you may want to optimize your code (but I think you can also try fast submit which only process the public test dataset for now)"
        }
      ]
    },
    {
      "id": 1266971,
      "postDate": "2021-04-08T08:03:25.540Z",
      "content": "<p>Did you find the reason for the error?</p>",
      "rawMarkdown": "Did you find the reason for the error?"
    },
    {
      "id": 1266913,
      "postDate": "2021-04-08T07:04:21.337Z",
      "content": "<p><a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> Wondering if we can somehow get access to the error info for the private part. Right now we can only guess what happend to debug</p>",
      "rawMarkdown": "@lnhtrang Wondering if we can somehow get access to the error info for the private part. Right now we can only guess what happend to debug",
      "replies": [
        {
          "id": 1267129,
          "postDate": "2021-04-08T10:37:08.787Z",
          "content": "<p>Hi! I can only see your submission error here:</p>\n<pre><code>&lt;Error&gt;\n&lt;Code&gt;NoSuchKey&lt;/Code&gt;\n&lt;Message&gt;The specified key does not exist.&lt;/Message&gt;\n&lt;Details&gt;No such object: kaggle-competitions-submissions/23823/20381194.raw&lt;/Details&gt;\n&lt;/Error&gt;\n</code></pre>",
          "rawMarkdown": "Hi! I can only see your submission error here:\n```\n<Error>\n<Code>NoSuchKey</Code>\n<Message>The specified key does not exist.</Message>\n<Details>No such object: kaggle-competitions-submissions/23823/20381194.raw</Details>\n</Error>\n```",
          "votes": 1
        },
        {
          "id": 1267653,
          "postDate": "2021-04-08T17:13:48.040Z",
          "content": "<p>Thanks for the information! Seems there is some error in my code caused it cannot work for the image with 3072 size</p>",
          "rawMarkdown": "Thanks for the information! Seems there is some error in my code caused it cannot work for the image with 3072 size"
        },
        {
          "id": 1277894,
          "postDate": "2021-04-19T10:52:14.103Z",
          "content": "<p>Hi, after I come access some posts having same issue, I understand that submission notebook has to read file sample_submission.csv and produce encode string for all images and make prediction with each segmentation cell. <br>\nPlease correct my thought if it is wrong. Thank you very much.</p>",
          "rawMarkdown": "Hi, after I come access some posts having same issue, I understand that submission notebook has to read file sample_submission.csv and produce encode string for all images and make prediction with each segmentation cell. \nPlease correct my thought if it is wrong. Thank you very much."
        }
      ]
    },
    {
      "id": 1266764,
      "postDate": "2021-04-08T04:14:06.813Z",
      "content": "<pre><code>Format issue. Possible but I do have a successful submission (0.2x) when I filter the data using sub_df.loc[sub_df['ImageWidth'] == 2048]. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\"\n</code></pre>\n<p>Could it be that the mask exceeds the size of the image?</p>",
      "rawMarkdown": "```\nFormat issue. Possible but I do have a successful submission (0.2x) when I filter the data using sub_df.loc[sub_df['ImageWidth'] == 2048]. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\"\n```\nCould it be that the mask exceeds the size of the image?",
      "replies": [
        {
          "id": 1266774,
          "postDate": "2021-04-08T04:28:55.527Z",
          "content": "<p>Not sure, have you met such error before?</p>",
          "rawMarkdown": "Not sure, have you met such error before?",
          "votes": 1
        },
        {
          "id": 1266789,
          "postDate": "2021-04-08T04:44:58.823Z",
          "content": "<p>No. I am just guessing. You can find the difference between success and failure.</p>",
          "rawMarkdown": "No. I am just guessing. You can find the difference between success and failure."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1274228,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2021-04-15T05:55:33.960000",
      "content": "<p>I got \"Submission Scoring Error\" when I submitted huge submission file. Approximately 8 GiB or so including public and private.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1267949,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2021-04-09T01:50:21.423000",
      "content": "<p>Found 2 issues in my code</p>\n<ol>\n<li>I hardcoded mask size to be 2048 in the dataset although I have arg for this in init() … I guess this caused the scoring error</li>\n<li>Another issue I set padding = False and that will cause an issue when segment 1728 images<br>\nNow I am save and update new submission hope this can work</li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1266881,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-04-08T06:24:35.767000",
      "content": "<p>In my previous submission, I got the same error message if there are more than 1 mask for the same 1 cell.<br>\nI guess “1 mask for 1 cell” is required to be scored properly.<br>\nAnother possibility is the format of mask encoding string.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1266909,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-04-08T07:01:26.637000",
          "content": "<p>Sorry what do you mean by 1 mask for 1 cell? Can you give me an example of your error submission?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1266945,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-04-08T07:34:44.040000",
          "content": "<p>For example, suppose that \"eNoLCAgIsAQABJ4Beg==\" is the mask of the most left upper corner cell of the image id X, we have to write \"0 0.001 eNoLCAgIsAQABJ4Beg== 1 0.1 eNoLCAgIsAQABJ4Beg== 2 0.005 eNoLCAgIsAQABJ4Beg==…and so on\"<br>\nI guess we shouldn't have another mask string for that cell, like \"0 0.001 eNoLCAgIsAQABJ4Beg== 1 0.1 HIAaagsCAgIssagJIANDIiOKFA== 2 0.005 eNoLCAgIsAQABJ4Beg==\".</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1266962,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-04-08T07:56:12.060000",
          "content": "<p>I see your point. Let me check the code. But my code do work if i only select 2048 images, and after that what i did is just combine [2048] + [3072] + … and output to the df</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1266965,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-04-08T07:58:42.697000",
          "content": "<p>I see. How about selecting only 1728 or 3072 images?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267883,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-04-08T22:50:44.453000",
          "content": "<p>Seems the issue is from those images not equal to 2048. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1274494,
      "author_name": "glopezzz",
      "author_url": "",
      "post_date": "2021-04-15T11:05:20.030000",
      "content": "<p>Hiii, I'm having the same issues.</p>\n<p>I am processing batches of images of the same size, but my notebook can't handle all the data, it takes too long.<br>\nThe only sumbission I achieved is running only the most frequent size images. Now I'm trying processing the most 2 frequent sizes, but let's see if I don't get a <code>Notebook Timeout</code>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1278408,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-04-19T21:07:35.403000",
          "content": "<p>I think for segmentation you shouldn't get time out. I guess you may want to optimize your code (but I think you can also try fast submit which only process the public test dataset for now)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1266971,
      "author_name": "Zaakcii Ru",
      "author_url": "",
      "post_date": "2021-04-08T08:03:25.540000",
      "content": "<p>Did you find the reason for the error?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1266913,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2021-04-08T07:04:21.337000",
      "content": "<p><a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> Wondering if we can somehow get access to the error info for the private part. Right now we can only guess what happend to debug</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1267129,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-08T10:37:08.787000",
          "content": "<p>Hi! I can only see your submission error here:</p>\n<pre><code>&lt;Error&gt;\n&lt;Code&gt;NoSuchKey&lt;/Code&gt;\n&lt;Message&gt;The specified key does not exist.&lt;/Message&gt;\n&lt;Details&gt;No such object: kaggle-competitions-submissions/23823/20381194.raw&lt;/Details&gt;\n&lt;/Error&gt;\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1267653,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-04-08T17:13:48.040000",
          "content": "<p>Thanks for the information! Seems there is some error in my code caused it cannot work for the image with 3072 size</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277894,
          "author_name": "Anh-Vu Mai-Nguyen",
          "author_url": "",
          "post_date": "2021-04-19T10:52:14.103000",
          "content": "<p>Hi, after I come access some posts having same issue, I understand that submission notebook has to read file sample_submission.csv and produce encode string for all images and make prediction with each segmentation cell. <br>\nPlease correct my thought if it is wrong. Thank you very much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1266764,
      "author_name": "Alien",
      "author_url": "",
      "post_date": "2021-04-08T04:14:06.813000",
      "content": "<pre><code>Format issue. Possible but I do have a successful submission (0.2x) when I filter the data using sub_df.loc[sub_df['ImageWidth'] == 2048]. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\"\n</code></pre>\n<p>Could it be that the mask exceeds the size of the image?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1266774,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-04-08T04:28:55.527000",
          "content": "<p>Not sure, have you met such error before?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1266789,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-04-08T04:44:58.823000",
          "content": "<p>No. I am just guessing. You can find the difference between success and failure.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1266733": "My personal kernel keep got the \"Submission Scoring Error\" when I submit to the leaderboard (although there is no error from the public part). I am trying to figure out what is the root cause and think the forum may have some people who may have insights to help me.\n\nPotential cause:\n1. Out of time. Unlikely, the kernel is only using 4900s for the public eval dataset, so it is unlikely the public + private will take more than 9 hours (also wondering if there will be other error showing if out of time?)\n2. Memory issue? Personally, I have no idea whether the memory will become an issue. Not sure if someone has figured out a way to debug this?\n3. Format issue. Possible but I do have a successful submission (0.2x) when I filter the data using `sub_df.loc[sub_df['ImageWidth'] == 2048]`. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\" \n\nPS Wondering Kaggle and organizer can provide more support on the error for the private part. This is a kind of black box and hard to debug the issue. ",
    "1274228": "I got \"Submission Scoring Error\" when I submitted huge submission file. Approximately 8 GiB or so including public and private.",
    "1267949": "Found 2 issues in my code\n1. I hardcoded mask size to be 2048 in the dataset although I have arg for this in init() ... I guess this caused the scoring error\n2. Another issue I set padding = False and that will cause an issue when segment 1728 images\nNow I am save and update new submission hope this can work",
    "1266881": "In my previous submission, I got the same error message if there are more than 1 mask for the same 1 cell.\nI guess “1 mask for 1 cell” is required to be scored properly.\nAnother possibility is the format of mask encoding string.",
    "1274494": "Hiii, I'm having the same issues.\n\nI am processing batches of images of the same size, but my notebook can't handle all the data, it takes too long.\nThe only sumbission I achieved is running only the most frequent size images. Now I'm trying processing the most 2 frequent sizes, but let's see if I don't get a `Notebook Timeout`.",
    "1266971": "Did you find the reason for the error?",
    "1266913": "@lnhtrang Wondering if we can somehow get access to the error info for the private part. Right now we can only guess what happend to debug",
    "1266764": "```\nFormat issue. Possible but I do have a successful submission (0.2x) when I filter the data using sub_df.loc[sub_df['ImageWidth'] == 2048]. But later on when I changed 2048 to other image sizes and concatenate all results together then I got the \"Submission Scoring Error\"\n```\nCould it be that the mask exceeds the size of the image?"
  }
}