{
  "id": 99227,
  "title": "Common submission errors (and how to fix them)",
  "url": "/competitions/aptos2019-blindness-detection/discussion/99227",
  "author_name": "",
  "post_date": "2019-07-09T17:20:17.910013300Z",
  "votes": 31,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I lose 15 submissions until I was able to fixed these 3 issues, so it may help some people here. If you find any other common mistakes and how to solve it, please share your problem&amp;solution.</p>\n\n<p>1) <strong>Dataset update</strong>. This is probably a kernel bug (and it is very confusing). If you upload your own dataset, and update it, it might be possible that the kernel is still reading the old version (and may be incompatible with your new version code.) To fix, simply remove the dataset, and then add it again in the kernel.</p>\n\n<p>This was mentioned already in some previous discussion but I couldn’t find it now so I am sorry not to give a proper credit.</p>\n\n<p>If you use ‘output from other kernel’ directly as a dataset, this may also cause a bug, especially if that kernel is updated. It is much safer to create a new dataset. </p>\n\n<p>2) <strong>Memory error</strong>. Common cases are people try to use big image size, and store many of them in the memory. The safe way is to read image and predict one by one. In other experience, I sometimes found that Keras model can use a lot of memory. In this case, peridiocially <code>del model</code>, <code>gc.collect()</code>, <code>K.clear_session()</code> before create and load model again may help.</p>\n\n<p>3) <strong>Crop function</strong>. If your crop function simply try to crop all dark pixels out. In the private test data, it seems that there are all-dark image, and your crop function will crop everything out. Simply check the dimension of the result, if 0, return original image.</p>\n\n<p>All these bugs can result in ‘no details’, ‘submission CSV not found’, ‘score 0.000’</p>\n\n<p>Lastly, if you are not in these cases, this method may also help\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409</a></p>\n\n<p>If you was able to fix other bugs, please let us know!</p>",
  "messages": [
    {
      "id": "571506",
      "postDate": "07/09/2019 17:20:17",
      "content": "<p>I lose 15 submissions until I was able to fixed these 3 issues, so it may help some people here. If you find any other common mistakes and how to solve it, please share your problem&amp;solution.</p>\n\n<p>1) <strong>Dataset update</strong>. This is probably a kernel bug (and it is very confusing). If you upload your own dataset, and update it, it might be possible that the kernel is still reading the old version (and may be incompatible with your new version code.) To fix, simply remove the dataset, and then add it again in the kernel.</p>\n\n<p>This was mentioned already in some previous discussion but I couldn’t find it now so I am sorry not to give a proper credit.</p>\n\n<p>If you use ‘output from other kernel’ directly as a dataset, this may also cause a bug, especially if that kernel is updated. It is much safer to create a new dataset. </p>\n\n<p>2) <strong>Memory error</strong>. Common cases are people try to use big image size, and store many of them in the memory. The safe way is to read image and predict one by one. In other experience, I sometimes found that Keras model can use a lot of memory. In this case, peridiocially <code>del model</code>, <code>gc.collect()</code>, <code>K.clear_session()</code> before create and load model again may help.</p>\n\n<p>3) <strong>Crop function</strong>. If your crop function simply try to crop all dark pixels out. In the private test data, it seems that there are all-dark image, and your crop function will crop everything out. Simply check the dimension of the result, if 0, return original image.</p>\n\n<p>All these bugs can result in ‘no details’, ‘submission CSV not found’, ‘score 0.000’</p>\n\n<p>Lastly, if you are not in these cases, this method may also help\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409</a></p>\n\n<p>If you was able to fix other bugs, please let us know!</p>",
      "rawMarkdown": "I lose 15 submissions until I was able to fixed these 3 issues, so it may help some people here. If you find any other common mistakes and how to solve it, please share your problem&amp;solution.\n\n1) **Dataset update**. This is probably a kernel bug (and it is very confusing). If you upload your own dataset, and update it, it might be possible that the kernel is still reading the old version (and may be incompatible with your new version code.) To fix, simply remove the dataset, and then add it again in the kernel.\n\nThis was mentioned already in some previous discussion but I couldn’t find it now so I am sorry not to give a proper credit.\n\nIf you use ‘output from other kernel’ directly as a dataset, this may also cause a bug, especially if that kernel is updated. It is much safer to create a new dataset. \n\n2) **Memory error**. Common cases are people try to use big image size, and store many of them in the memory. The safe way is to read image and predict one by one. In other experience, I sometimes found that Keras model can use a lot of memory. In this case, peridiocially `del model`, `gc.collect()`, `K.clear_session()` before create and load model again may help.\n\n3) **Crop function**. If your crop function simply try to crop all dark pixels out. In the private test data, it seems that there are all-dark image, and your crop function will crop everything out. Simply check the dimension of the result, if 0, return original image.\n\nAll these bugs can result in ‘no details’, ‘submission CSV not found’, ‘score 0.000’\n\nLastly, if you are not in these cases, this method may also help\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409\n\nIf you was able to fix other bugs, please let us know!",
      "votes": null
    },
    {
      "id": "571778",
      "postDate": "07/10/2019 04:56:07",
      "content": "<p>And of course, another obvious bug is that the kernel is taking more than 9 hours on the private test set, because 6x larger dataset was not taken into account.</p>",
      "rawMarkdown": "And of course, another obvious bug is that the kernel is taking more than 9 hours on the private test set, because 6x larger dataset was not taken into account.",
      "votes": null
    },
    {
      "id": "571995",
      "postDate": "07/10/2019 10:13:58",
      "content": "<p>Thanks for sharing this!\nI'm the author of <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409</a>.</p>\n\n<p>I was seeing \"Kernel Threw Exception\", but simply creating new kernel and copy and paste code worked. This is really weird but worth trying. Note that I tried remove/re-add private dataset but didn't work some times.</p>",
      "rawMarkdown": "Thanks for sharing this!\nI'm the author of https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409.\n\nI was seeing \"Kernel Threw Exception\", but simply creating new kernel and copy and paste code worked. This is really weird but worth trying. Note that I tried remove/re-add private dataset but didn't work some times.",
      "votes": null
    },
    {
      "id": "572154",
      "postDate": "07/10/2019 14:42:03",
      "content": "<p>From yesterday morning I am not able to commit almost identical new versions of the same kernel because of some kernel error. Last successfull commit took less than three hours. New commits timeout after nine hours waiting. I beleive there is a problem on Kaggle. Will try to copy and paste this kernel's code to new kernel though. Thanks, higepon.</p>\n\n<p>Has anyone faced similar problems?</p>",
      "rawMarkdown": "From yesterday morning I am not able to commit almost identical new versions of the same kernel because of some kernel error. Last successfull commit took less than three hours. New commits timeout after nine hours waiting. I beleive there is a problem on Kaggle. Will try to copy and paste this kernel's code to new kernel though. Thanks, higepon.\n\nHas anyone faced similar problems?",
      "votes": null
    },
    {
      "id": "572452",
      "postDate": "07/11/2019 00:42:48",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> , <a href=\"/tanlikesmath\">@tanlikesmath</a>: Do you mean it runs on private test set as well for the public leaderboard? What is your usual submission hour? </p>\n\n<p><strong>Here is my actual question:</strong></p>\n\n<p>Isn't it enough to make predictions(or inference) on the submission file (i.e. id_code - public test set) for the public leaderboard? Obviously, it makes prediction very well on the test set (1928 samples and time 4279.8seconds ~ 1.18 hour) before making an actual submission after the kernel has been already successfully run. </p>\n\n<p><strong>My submission has been running for two to three hours now.</strong></p>",
      "rawMarkdown": "ratthachat , @tanlikesmath: Do you mean it runs on private test set as well for the public leaderboard? What is your usual submission hour? \n\n**Here is my actual question:**\n\nIsn't it enough to make predictions(or inference) on the submission file (i.e. id_code - public test set) for the public leaderboard? Obviously, it makes prediction very well on the test set (1928 samples and time 4279.8seconds ~ 1.18 hour) before making an actual submission after the kernel has been already successfully run. \n\n**My submission has been running for two to three hours now.**",
      "votes": null
    },
    {
      "id": "572461",
      "postDate": "07/11/2019 00:59:02",
      "content": "<p>This competition is a synchronous kernels-only competition. What this means is that it also runs on the private test set synchronously, but will only display public LB. If there is an error it will say there is an error.</p>\n\n<p>There are two benefits compared to previous kernels-only competitions that would do what you suggested:\n1. People will know if their dataset failed on private test set so people will not be surprised when they have no submission on the private LB and cannot fix it.\n2. The results of the competition would be out immediately after the end of the competition, as opposed to waiting for the kernels to be rerun on the private set.</p>\n\n<p>Because the kernel is rerun on the private set, which is ~6x larger, it will take ~6x longer for the submission to finish. </p>",
      "rawMarkdown": "This competition is a synchronous kernels-only competition. What this means is that it also runs on the private test set synchronously, but will only display public LB. If there is an error it will say there is an error.\n\nThere are two benefits compared to previous kernels-only competitions that would do what you suggested:\n1. People will know if their dataset failed on private test set so people will not be surprised when they have no submission on the private LB and cannot fix it.\n2. The results of the competition would be out immediately after the end of the competition, as opposed to waiting for the kernels to be rerun on the private set.\n\nBecause the kernel is rerun on the private set, which is ~6x larger, it will take ~6x longer for the submission to finish.",
      "votes": null
    },
    {
      "id": "572488",
      "postDate": "07/11/2019 01:51:21",
      "content": "<p>Thanks for the explanation brother.  <a href=\"/tanlikesmath\">@tanlikesmath</a> </p>\n\n<p>Not sure if I missed this note, I feel when kaggle describing synchronous KO competitions should have an additional line of text stating that every submission is going to be running on both public(i.e. 15% samples) as well as a private(i.e. 85% samples) test set and the results could take little longer(Approx. ~6 times).  </p>\n\n<p><strong>Now For e.g., if the public test set inference takes 1 hour would be approx. 6 hour to see an error or your public score.</strong></p>\n\n<p>Thanks,\nHimanshu</p>",
      "rawMarkdown": "Thanks for the explanation brother.  @tanlikesmath \n\nNot sure if I missed this note, I feel when kaggle describing synchronous KO competitions should have an additional line of text stating that every submission is going to be running on both public(i.e. 15% samples) as well as a private(i.e. 85% samples) test set and the results could take little longer(Approx. ~6 times).  \n\n**Now For e.g., if the public test set inference takes 1 hour would be approx. 6 hour to see an error or your public score.**\n\nThanks,\nHimanshu",
      "votes": null
    },
    {
      "id": "573559",
      "postDate": "07/12/2019 12:15:54",
      "content": "<p>I have found the reason of unsuccessful submissions. Experimenting on Kaggle platform I have turned off GPU and kernel executions became too slow. Turned on GPU back and the problem disappeared.</p>",
      "rawMarkdown": "I have found the reason of unsuccessful submissions. Experimenting on Kaggle platform I have turned off GPU and kernel executions became too slow. Turned on GPU back and the problem disappeared.",
      "votes": null
    },
    {
      "id": "574044",
      "postDate": "07/13/2019 06:41:47",
      "content": "<p>I think there are some unknown problems with add training result to the submit kernel by Kaggle's \"add dataset\". </p>\n\n<p>I used a training kernel to generate \"model.bin\" and synchronized it with \"add dataset\" tool in my submit kernel. </p>\n\n<p>At first time, I  re-\"add dataset\" when my new committed training kernel finished, and after 2-3 commits, the submit kernel will have an unknown kernel error during submit.</p>\n\n<p>After that, I found the synchronized training result in my submit kernel would be changed every time my new committed training kernel finished. It means I don't need to re-\"add dataset\" to my submit kernel and the unknown error didn't happen again. On the other hand, I think I need to take care of the version of 'model.bin'...</p>\n\n<p>I hope this information could help.</p>",
      "rawMarkdown": "I think there are some unknown problems with add training result to the submit kernel by Kaggle's \"add dataset\". \n\nI used a training kernel to generate \"model.bin\" and synchronized it with \"add dataset\" tool in my submit kernel. \n\nAt first time, I  re-\"add dataset\" when my new committed training kernel finished, and after 2-3 commits, the submit kernel will have an unknown kernel error during submit.\n\nAfter that, I found the synchronized training result in my submit kernel would be changed every time my new committed training kernel finished. It means I don't need to re-\"add dataset\" to my submit kernel and the unknown error didn't happen again. On the other hand, I think I need to take care of the version of 'model.bin'...\n\nI hope this information could help.",
      "votes": null
    },
    {
      "id": "574446",
      "postDate": "07/13/2019 20:38:11",
      "content": "<p>Neither of this are my problems. Rather I am getting fairly very very low score. Close to 0 but not 0. I dont understand what the problem might be. Like to help or give directions?</p>",
      "rawMarkdown": "Neither of this are my problems. Rather I am getting fairly very very low score. Close to 0 but not 0. I dont understand what the problem might be. Like to help or give directions?",
      "votes": null
    },
    {
      "id": "574514",
      "postDate": "07/14/2019 02:19:12",
      "content": "<p>it is useful.Thanks</p>",
      "rawMarkdown": "it is useful.Thanks",
      "votes": null
    },
    {
      "id": "577724",
      "postDate": "07/17/2019 02:26:15",
      "content": "<p><strong>UPDATE on possible failure cause #4, #5</strong></p>\n\n<p>4) <strong>Internet=On, using git clone</strong> it happen many times that if you are cloning external libraries from github with internet=ON, you can test run in the editing kernel with no problem. However, when commiting things will fail. It is tempting to think that internet=ON causes your commit to fail (since competition’s rule says that internet is not allowed)</p>\n\n<p>However, internet=ON is not the real culprit, as internet is not allowed only in the submission kernel. Many times, there is a problem in <code>.git</code> (hidden directory) in your <code>working directory</code> containing too many sub-directories which is not support by the committing kernel.</p>\n\n<p>The solution is simply <code>rm -rf [cloned repository]/.git</code></p>\n\n<p>5) <strong>Too many output files</strong> . Like 4) , even though you didn't use internet, but try to save too many output files e.g. save all individual train/test data into separate files, this can exceed maximum output file allowed for the kernel.</p>",
      "rawMarkdown": "**UPDATE on possible failure cause #4, #5**\n\n4) **Internet=On, using git clone** it happen many times that if you are cloning external libraries from github with internet=ON, you can test run in the editing kernel with no problem. However, when commiting things will fail. It is tempting to think that internet=ON causes your commit to fail (since competition’s rule says that internet is not allowed)\n\nHowever, internet=ON is not the real culprit, as internet is not allowed only in the submission kernel. Many times, there is a problem in `.git` (hidden directory) in your `working directory` containing too many sub-directories which is not support by the committing kernel.\n\nThe solution is simply `rm -rf [cloned repository]/.git`\n\n5) **Too many output files** . Like 4) , even though you didn't use internet, but try to save too many output files e.g. save all individual train/test data into separate files, this can exceed maximum output file allowed for the kernel.",
      "votes": null
    },
    {
      "id": "607504",
      "postDate": "08/25/2019 12:13:44",
      "content": "<p>I am facing a similar problem. Were you able to figure out a solution?</p>",
      "rawMarkdown": "I am facing a similar problem. Were you able to figure out a solution?",
      "votes": null
    },
    {
      "id": "608542",
      "postDate": "08/26/2019 22:50:44",
      "content": "<p>I am facing the following error during the initial kernel commit:</p>\n\n<p><code>\nError in atexit._run_exitfuncs: Traceback (most recent call last): File \"/opt/conda/lib/python3.6/logging/__init__.py\", line 1944, in shutdown h.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 882, in flush self._current_handler.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 776, in flush self.stream.flush() File \"/opt/conda/lib/python3.6/site-packages/ipykernel/iostream.py\", line 341, in flush if self.pub_thread.thread.is_alive(): AttributeError: 'NoneType' object has no attribute 'thread'\n</code>\nWhat does that mean? How can I fix it?</p>",
      "rawMarkdown": "I am facing the following error during the initial kernel commit:\n\n`\nError in atexit._run_exitfuncs: Traceback (most recent call last): File \"/opt/conda/lib/python3.6/logging/__init__.py\", line 1944, in shutdown h.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 882, in flush self._current_handler.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 776, in flush self.stream.flush() File \"/opt/conda/lib/python3.6/site-packages/ipykernel/iostream.py\", line 341, in flush if self.pub_thread.thread.is_alive(): AttributeError: 'NoneType' object has no attribute 'thread'\n`\nWhat does that mean? How can I fix it?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 571778,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "07/10/2019 04:56:07",
      "content": "<p>And of course, another obvious bug is that the kernel is taking more than 9 hours on the private test set, because 6x larger dataset was not taken into account.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 571995,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "07/10/2019 10:13:58",
      "content": "<p>Thanks for sharing this!\nI'm the author of <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409</a>.</p>\n\n<p>I was seeing \"Kernel Threw Exception\", but simply creating new kernel and copy and paste code worked. This is really weird but worth trying. Note that I tried remove/re-add private dataset but didn't work some times.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 572154,
      "author_name": "prokofyev2",
      "author_url": "",
      "post_date": "07/10/2019 14:42:03",
      "content": "<p>From yesterday morning I am not able to commit almost identical new versions of the same kernel because of some kernel error. Last successfull commit took less than three hours. New commits timeout after nine hours waiting. I beleive there is a problem on Kaggle. Will try to copy and paste this kernel's code to new kernel though. Thanks, higepon.</p>\n\n<p>Has anyone faced similar problems?</p>",
      "votes": null,
      "replies": [
        {
          "id": 607504,
          "author_name": "adeperio",
          "author_url": "",
          "post_date": "08/25/2019 12:13:44",
          "content": "<p>I am facing a similar problem. Were you able to figure out a solution?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 572452,
      "author_name": "hmnshu",
      "author_url": "",
      "post_date": "07/11/2019 00:42:48",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> , <a href=\"/tanlikesmath\">@tanlikesmath</a>: Do you mean it runs on private test set as well for the public leaderboard? What is your usual submission hour? </p>\n\n<p><strong>Here is my actual question:</strong></p>\n\n<p>Isn't it enough to make predictions(or inference) on the submission file (i.e. id_code - public test set) for the public leaderboard? Obviously, it makes prediction very well on the test set (1928 samples and time 4279.8seconds ~ 1.18 hour) before making an actual submission after the kernel has been already successfully run. </p>\n\n<p><strong>My submission has been running for two to three hours now.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 572461,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "07/11/2019 00:59:02",
          "content": "<p>This competition is a synchronous kernels-only competition. What this means is that it also runs on the private test set synchronously, but will only display public LB. If there is an error it will say there is an error.</p>\n\n<p>There are two benefits compared to previous kernels-only competitions that would do what you suggested:\n1. People will know if their dataset failed on private test set so people will not be surprised when they have no submission on the private LB and cannot fix it.\n2. The results of the competition would be out immediately after the end of the competition, as opposed to waiting for the kernels to be rerun on the private set.</p>\n\n<p>Because the kernel is rerun on the private set, which is ~6x larger, it will take ~6x longer for the submission to finish. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 572488,
          "author_name": "hmnshu",
          "author_url": "",
          "post_date": "07/11/2019 01:51:21",
          "content": "<p>Thanks for the explanation brother.  <a href=\"/tanlikesmath\">@tanlikesmath</a> </p>\n\n<p>Not sure if I missed this note, I feel when kaggle describing synchronous KO competitions should have an additional line of text stating that every submission is going to be running on both public(i.e. 15% samples) as well as a private(i.e. 85% samples) test set and the results could take little longer(Approx. ~6 times).  </p>\n\n<p><strong>Now For e.g., if the public test set inference takes 1 hour would be approx. 6 hour to see an error or your public score.</strong></p>\n\n<p>Thanks,\nHimanshu</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 573559,
      "author_name": "prokofyev2",
      "author_url": "",
      "post_date": "07/12/2019 12:15:54",
      "content": "<p>I have found the reason of unsuccessful submissions. Experimenting on Kaggle platform I have turned off GPU and kernel executions became too slow. Turned on GPU back and the problem disappeared.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 574044,
      "author_name": "wzyplus",
      "author_url": "",
      "post_date": "07/13/2019 06:41:47",
      "content": "<p>I think there are some unknown problems with add training result to the submit kernel by Kaggle's \"add dataset\". </p>\n\n<p>I used a training kernel to generate \"model.bin\" and synchronized it with \"add dataset\" tool in my submit kernel. </p>\n\n<p>At first time, I  re-\"add dataset\" when my new committed training kernel finished, and after 2-3 commits, the submit kernel will have an unknown kernel error during submit.</p>\n\n<p>After that, I found the synchronized training result in my submit kernel would be changed every time my new committed training kernel finished. It means I don't need to re-\"add dataset\" to my submit kernel and the unknown error didn't happen again. On the other hand, I think I need to take care of the version of 'model.bin'...</p>\n\n<p>I hope this information could help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 574446,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "07/13/2019 20:38:11",
      "content": "<p>Neither of this are my problems. Rather I am getting fairly very very low score. Close to 0 but not 0. I dont understand what the problem might be. Like to help or give directions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 574514,
      "author_name": "chanhu",
      "author_url": "",
      "post_date": "07/14/2019 02:19:12",
      "content": "<p>it is useful.Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 577724,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "07/17/2019 02:26:15",
      "content": "<p><strong>UPDATE on possible failure cause #4, #5</strong></p>\n\n<p>4) <strong>Internet=On, using git clone</strong> it happen many times that if you are cloning external libraries from github with internet=ON, you can test run in the editing kernel with no problem. However, when commiting things will fail. It is tempting to think that internet=ON causes your commit to fail (since competition’s rule says that internet is not allowed)</p>\n\n<p>However, internet=ON is not the real culprit, as internet is not allowed only in the submission kernel. Many times, there is a problem in <code>.git</code> (hidden directory) in your <code>working directory</code> containing too many sub-directories which is not support by the committing kernel.</p>\n\n<p>The solution is simply <code>rm -rf [cloned repository]/.git</code></p>\n\n<p>5) <strong>Too many output files</strong> . Like 4) , even though you didn't use internet, but try to save too many output files e.g. save all individual train/test data into separate files, this can exceed maximum output file allowed for the kernel.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 608542,
      "author_name": "tahsin",
      "author_url": "",
      "post_date": "08/26/2019 22:50:44",
      "content": "<p>I am facing the following error during the initial kernel commit:</p>\n\n<p><code>\nError in atexit._run_exitfuncs: Traceback (most recent call last): File \"/opt/conda/lib/python3.6/logging/__init__.py\", line 1944, in shutdown h.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 882, in flush self._current_handler.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 776, in flush self.stream.flush() File \"/opt/conda/lib/python3.6/site-packages/ipykernel/iostream.py\", line 341, in flush if self.pub_thread.thread.is_alive(): AttributeError: 'NoneType' object has no attribute 'thread'\n</code>\nWhat does that mean? How can I fix it?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "571506": "I lose 15 submissions until I was able to fixed these 3 issues, so it may help some people here. If you find any other common mistakes and how to solve it, please share your problem&amp;solution.\n\n1) **Dataset update**. This is probably a kernel bug (and it is very confusing). If you upload your own dataset, and update it, it might be possible that the kernel is still reading the old version (and may be incompatible with your new version code.) To fix, simply remove the dataset, and then add it again in the kernel.\n\nThis was mentioned already in some previous discussion but I couldn’t find it now so I am sorry not to give a proper credit.\n\nIf you use ‘output from other kernel’ directly as a dataset, this may also cause a bug, especially if that kernel is updated. It is much safer to create a new dataset. \n\n2) **Memory error**. Common cases are people try to use big image size, and store many of them in the memory. The safe way is to read image and predict one by one. In other experience, I sometimes found that Keras model can use a lot of memory. In this case, peridiocially `del model`, `gc.collect()`, `K.clear_session()` before create and load model again may help.\n\n3) **Crop function**. If your crop function simply try to crop all dark pixels out. In the private test data, it seems that there are all-dark image, and your crop function will crop everything out. Simply check the dimension of the result, if 0, return original image.\n\nAll these bugs can result in ‘no details’, ‘submission CSV not found’, ‘score 0.000’\n\nLastly, if you are not in these cases, this method may also help\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409\n\nIf you was able to fix other bugs, please let us know!",
    "571778": "And of course, another obvious bug is that the kernel is taking more than 9 hours on the private test set, because 6x larger dataset was not taken into account.",
    "571995": "Thanks for sharing this!\nI'm the author of https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98849#latest-571409.\n\nI was seeing \"Kernel Threw Exception\", but simply creating new kernel and copy and paste code worked. This is really weird but worth trying. Note that I tried remove/re-add private dataset but didn't work some times.",
    "572154": "From yesterday morning I am not able to commit almost identical new versions of the same kernel because of some kernel error. Last successfull commit took less than three hours. New commits timeout after nine hours waiting. I beleive there is a problem on Kaggle. Will try to copy and paste this kernel's code to new kernel though. Thanks, higepon.\n\nHas anyone faced similar problems?",
    "572452": "ratthachat , @tanlikesmath: Do you mean it runs on private test set as well for the public leaderboard? What is your usual submission hour? \n\n**Here is my actual question:**\n\nIsn't it enough to make predictions(or inference) on the submission file (i.e. id_code - public test set) for the public leaderboard? Obviously, it makes prediction very well on the test set (1928 samples and time 4279.8seconds ~ 1.18 hour) before making an actual submission after the kernel has been already successfully run. \n\n**My submission has been running for two to three hours now.**",
    "572461": "This competition is a synchronous kernels-only competition. What this means is that it also runs on the private test set synchronously, but will only display public LB. If there is an error it will say there is an error.\n\nThere are two benefits compared to previous kernels-only competitions that would do what you suggested:\n1. People will know if their dataset failed on private test set so people will not be surprised when they have no submission on the private LB and cannot fix it.\n2. The results of the competition would be out immediately after the end of the competition, as opposed to waiting for the kernels to be rerun on the private set.\n\nBecause the kernel is rerun on the private set, which is ~6x larger, it will take ~6x longer for the submission to finish.",
    "572488": "Thanks for the explanation brother.  @tanlikesmath \n\nNot sure if I missed this note, I feel when kaggle describing synchronous KO competitions should have an additional line of text stating that every submission is going to be running on both public(i.e. 15% samples) as well as a private(i.e. 85% samples) test set and the results could take little longer(Approx. ~6 times).  \n\n**Now For e.g., if the public test set inference takes 1 hour would be approx. 6 hour to see an error or your public score.**\n\nThanks,\nHimanshu",
    "573559": "I have found the reason of unsuccessful submissions. Experimenting on Kaggle platform I have turned off GPU and kernel executions became too slow. Turned on GPU back and the problem disappeared.",
    "574044": "I think there are some unknown problems with add training result to the submit kernel by Kaggle's \"add dataset\". \n\nI used a training kernel to generate \"model.bin\" and synchronized it with \"add dataset\" tool in my submit kernel. \n\nAt first time, I  re-\"add dataset\" when my new committed training kernel finished, and after 2-3 commits, the submit kernel will have an unknown kernel error during submit.\n\nAfter that, I found the synchronized training result in my submit kernel would be changed every time my new committed training kernel finished. It means I don't need to re-\"add dataset\" to my submit kernel and the unknown error didn't happen again. On the other hand, I think I need to take care of the version of 'model.bin'...\n\nI hope this information could help.",
    "574446": "Neither of this are my problems. Rather I am getting fairly very very low score. Close to 0 but not 0. I dont understand what the problem might be. Like to help or give directions?",
    "574514": "it is useful.Thanks",
    "577724": "**UPDATE on possible failure cause #4, #5**\n\n4) **Internet=On, using git clone** it happen many times that if you are cloning external libraries from github with internet=ON, you can test run in the editing kernel with no problem. However, when commiting things will fail. It is tempting to think that internet=ON causes your commit to fail (since competition’s rule says that internet is not allowed)\n\nHowever, internet=ON is not the real culprit, as internet is not allowed only in the submission kernel. Many times, there is a problem in `.git` (hidden directory) in your `working directory` containing too many sub-directories which is not support by the committing kernel.\n\nThe solution is simply `rm -rf [cloned repository]/.git`\n\n5) **Too many output files** . Like 4) , even though you didn't use internet, but try to save too many output files e.g. save all individual train/test data into separate files, this can exceed maximum output file allowed for the kernel.",
    "607504": "I am facing a similar problem. Were you able to figure out a solution?",
    "608542": "I am facing the following error during the initial kernel commit:\n\n`\nError in atexit._run_exitfuncs: Traceback (most recent call last): File \"/opt/conda/lib/python3.6/logging/__init__.py\", line 1944, in shutdown h.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 882, in flush self._current_handler.flush() File \"/opt/conda/lib/python3.6/site-packages/absl/logging/__init__.py\", line 776, in flush self.stream.flush() File \"/opt/conda/lib/python3.6/site-packages/ipykernel/iostream.py\", line 341, in flush if self.pub_thread.thread.is_alive(): AttributeError: 'NoneType' object has no attribute 'thread'\n`\nWhat does that mean? How can I fix it?"
  },
  "source": "meta"
}