{
  "id": 216439,
  "title": "Submission CSV Not Found today, Kaggle issue or just me?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/216439",
  "author_name": "MPWARE",
  "post_date": "2021-02-02T17:31:03.435000",
  "votes": 12,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I've a working inference kernel, I've updated/uploaded the weights of my model and I got \"<strong>Submission CSV Not Found</strong>\" error. When I look at diff between yesterday working version and today's broken version I can confirm that only path to weights has been updated. Am I alone to get this issue today? I've tried to re-submit and I've got the same error. It does not make sense, weights update cannot break the pipeline and the submission.csv and I can see it's generated on commit. </p>",
  "messages": [
    {
      "id": 1183015,
      "postDate": "2021-02-02T17:31:03.437Z",
      "content": "<p>I've a working inference kernel, I've updated/uploaded the weights of my model and I got \"<strong>Submission CSV Not Found</strong>\" error. When I look at diff between yesterday working version and today's broken version I can confirm that only path to weights has been updated. Am I alone to get this issue today? I've tried to re-submit and I've got the same error. It does not make sense, weights update cannot break the pipeline and the submission.csv and I can see it's generated on commit. </p>",
      "rawMarkdown": "I've a working inference kernel, I've updated/uploaded the weights of my model and I got \"**Submission CSV Not Found**\" error. When I look at diff between yesterday working version and today's broken version I can confirm that only path to weights has been updated. Am I alone to get this issue today? I've tried to re-submit and I've got the same error. It does not make sense, weights update cannot break the pipeline and the submission.csv and I can see it's generated on commit. ",
      "votes": 11
    },
    {
      "id": 1215663,
      "postDate": "2021-02-23T20:59:18.750Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I just experienced the exact same issue that you described…only change in my kernel is that it is a new weights file that is loaded…a follow up version with yet another weights file is working.</p>\n<p>Does not seem to happen often though…first time for me. </p>",
      "rawMarkdown": "Hi @mpware I just experienced the exact same issue that you described...only change in my kernel is that it is a new weights file that is loaded...a follow up version with yet another weights file is working.\n\nDoes not seem to happen often though...first time for me. \n",
      "votes": 1,
      "replies": [
        {
          "id": 1215731,
          "postDate": "2021-02-23T22:35:49.193Z",
          "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Are you using multiple workers? Or multiprocessing? It was the root cause of my problem. It looks <code>RasterIO</code> does not work well with multiple workers, it fails silently (not always) and the <code>submission.csv</code> file is not generated. I've been able to reproduce on commit. I'm using <code>workers=0</code> now and it works.</p>",
          "rawMarkdown": "@rsmits Are you using multiple workers? Or multiprocessing? It was the root cause of my problem. It looks `RasterIO` does not work well with multiple workers, it fails silently (not always) and the `submission.csv` file is not generated. I've been able to reproduce on commit. I'm using `workers=0` now and it works."
        },
        {
          "id": 1215741,
          "postDate": "2021-02-23T22:46:38.157Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I'am using the default workers setting which I believe is 0. So that's likely not the cause of my issue…<br>\nI'am retrying it with the model file again uploaded. So far so good..another model file from the same training run (different fold) did work so I thought I could try that.<br>\nSeems to keep working although it will be a few hours before I can see the final end result.</p>",
          "rawMarkdown": "Hi @mpware I'am using the default workers setting which I believe is 0. So that's likely not the cause of my issue...\nI'am retrying it with the model file again uploaded. So far so good..another model file from the same training run (different fold) did work so I thought I could try that.\nSeems to keep working although it will be a few hours before I can see the final end result.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1185435,
      "postDate": "2021-02-04T06:54:46.567Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Hope, so you will resolve this soon.</p>",
      "rawMarkdown": "@mpware Hope, so you will resolve this soon.\n\n",
      "votes": 2
    },
    {
      "id": 1183959,
      "postDate": "2021-02-03T10:15:03.067Z",
      "content": "<p>Hope you will resolve soon. </p>",
      "rawMarkdown": "Hope you will resolve soon. ",
      "votes": 2
    },
    {
      "id": 1183525,
      "postDate": "2021-02-03T04:14:54.847Z",
      "content": "<p>Today I've succeeded in submitting my kernel to get a score, so there seems to be little Kaggle scoring issue.</p>",
      "rawMarkdown": "Today I've succeeded in submitting my kernel to get a score, so there seems to be little Kaggle scoring issue.",
      "votes": 2,
      "replies": [
        {
          "id": 1183910,
          "postDate": "2021-02-03T09:45:03.157Z",
          "content": "<p>Thanks, I was also able to have successfull submissions in the mean time. However, it looks I'm able to reproduce the problem now on commit. Kernel succeeds (but dies indeed) and logs report:</p>\n<pre><code>File \"/opt/conda/lib/python3.7/site-packages/jupyter_client/asynchronous/channels.py\", line 46, in get_msg\nreturn await self._recv()\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/asynchronous/channels.py\", line 37, in _recv\nreturn self.session.deserialize(smsg)\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/session.py\", line 921, in deserialize\nraise ValueError(\"Duplicate Signature: %r\" % signature\n</code></pre>\n<p>I suspect a concurrent threading issue with RasterIO. I'm trying to use the 2 CPUs to speed up processing but it fails on some race conditions. I'm using thread lock but it does not seem to work as expected. If I commit again then it works. Same issue might happen on submit.</p>",
          "rawMarkdown": "Thanks, I was also able to have successfull submissions in the mean time. However, it looks I'm able to reproduce the problem now on commit. Kernel succeeds (but dies indeed) and logs report:\n\n```\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/asynchronous/channels.py\", line 46, in get_msg\nreturn await self._recv()\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/asynchronous/channels.py\", line 37, in _recv\nreturn self.session.deserialize(smsg)\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/session.py\", line 921, in deserialize\nraise ValueError(\"Duplicate Signature: %r\" % signature\n```\nI suspect a concurrent threading issue with RasterIO. I'm trying to use the 2 CPUs to speed up processing but it fails on some race conditions. I'm using thread lock but it does not seem to work as expected. If I commit again then it works. Same issue might happen on submit.",
          "votes": 4
        },
        {
          "id": 1187188,
          "postDate": "2021-02-05T09:17:19.277Z",
          "content": "<p>Confirmed, so it was me. No problem with Kaggle environment.</p>",
          "rawMarkdown": "Confirmed, so it was me. No problem with Kaggle environment.",
          "votes": 3
        },
        {
          "id": 1191868,
          "postDate": "2021-02-08T18:27:03.333Z",
          "content": "<p>RasterIO documentation claims it would work:<br>\n<a href=\"https://rasterio.readthedocs.io/en/latest/topics/concurrency.html\" target=\"_blank\">https://rasterio.readthedocs.io/en/latest/topics/concurrency.html</a></p>\n<pre><code>        with rasterio.open(outfile, \"w\", **src.profile) as dst:\n            windows = [window for ij, window in dst.block_windows()]\n\n            # We cannot write to the same file from multiple threads\n            # without causing race conditions. To safely read/write\n            # from multiple threads, we use a lock to protect the\n            # DatasetReader/Writer\n            read_lock = threading.Lock()\n            write_lock = threading.Lock()\n\n            def process(window):\n                with read_lock:\n                    src_array = src.read(window=window)\n</code></pre>\n<p>But for some unknown reason it fails randomly. Anyone experimenting similar issue?</p>\n<p>This guy had similar problem but not with rasterio:<br>\n<a href=\"https://www.kaggle.com/product-feedback/164002\" target=\"_blank\">https://www.kaggle.com/product-feedback/164002</a></p>",
          "rawMarkdown": "RasterIO documentation claims it would work:\nhttps://rasterio.readthedocs.io/en/latest/topics/concurrency.html\n\n```\n        with rasterio.open(outfile, \"w\", **src.profile) as dst:\n            windows = [window for ij, window in dst.block_windows()]\n\n            # We cannot write to the same file from multiple threads\n            # without causing race conditions. To safely read/write\n            # from multiple threads, we use a lock to protect the\n            # DatasetReader/Writer\n            read_lock = threading.Lock()\n            write_lock = threading.Lock()\n\n            def process(window):\n                with read_lock:\n                    src_array = src.read(window=window)\n```\n\nBut for some unknown reason it fails randomly. Anyone experimenting similar issue?\n\nThis guy had similar problem but not with rasterio:\nhttps://www.kaggle.com/product-feedback/164002\n"
        },
        {
          "id": 1265564,
          "postDate": "2021-04-07T03:06:02.517Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>    Will we get same error in case of oom ? I get sub csv error after 45 to 50 min .<br>\n. If there is an oom we get with in 30 min or after this much time only?</p>",
          "rawMarkdown": "@mpware    Will we get same error in case of oom ? I get sub csv error after 45 to 50 min .\n. If there is an oom we get with in 30 min or after this much time only?"
        }
      ]
    },
    {
      "id": 1277347,
      "postDate": "2021-04-18T16:30:33.427Z",
      "content": "<p>I'm having the same problem. Have you solved it? how?</p>",
      "rawMarkdown": "I'm having the same problem. Have you solved it? how?",
      "replies": [
        {
          "id": 1277349,
          "postDate": "2021-04-18T16:30:55.277Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1211404,
      "postDate": "2021-02-20T07:39:40.203Z",
      "content": "<p>I think the problem is with your end. Check again once.</p>",
      "rawMarkdown": "I think the problem is with your end. Check again once."
    },
    {
      "id": 1253543,
      "postDate": "2021-03-26T21:05:35.537Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1215663,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2021-02-23T20:59:18.750000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I just experienced the exact same issue that you described…only change in my kernel is that it is a new weights file that is loaded…a follow up version with yet another weights file is working.</p>\n<p>Does not seem to happen often though…first time for me. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1215731,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-02-23T22:35:49.193000",
          "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Are you using multiple workers? Or multiprocessing? It was the root cause of my problem. It looks <code>RasterIO</code> does not work well with multiple workers, it fails silently (not always) and the <code>submission.csv</code> file is not generated. I've been able to reproduce on commit. I'm using <code>workers=0</code> now and it works.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1215741,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2021-02-23T22:46:38.157000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I'am using the default workers setting which I believe is 0. So that's likely not the cause of my issue…<br>\nI'am retrying it with the model file again uploaded. So far so good..another model file from the same training run (different fold) did work so I thought I could try that.<br>\nSeems to keep working although it will be a few hours before I can see the final end result.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1185435,
      "author_name": "Tasneem Abdul Rahim",
      "author_url": "",
      "post_date": "2021-02-04T06:54:46.567000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Hope, so you will resolve this soon.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1183959,
      "author_name": "Adnan Zaidi",
      "author_url": "",
      "post_date": "2021-02-03T10:15:03.067000",
      "content": "<p>Hope you will resolve soon. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1183525,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-02-03T04:14:54.847000",
      "content": "<p>Today I've succeeded in submitting my kernel to get a score, so there seems to be little Kaggle scoring issue.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1183910,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-02-03T09:45:03.157000",
          "content": "<p>Thanks, I was also able to have successfull submissions in the mean time. However, it looks I'm able to reproduce the problem now on commit. Kernel succeeds (but dies indeed) and logs report:</p>\n<pre><code>File \"/opt/conda/lib/python3.7/site-packages/jupyter_client/asynchronous/channels.py\", line 46, in get_msg\nreturn await self._recv()\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/asynchronous/channels.py\", line 37, in _recv\nreturn self.session.deserialize(smsg)\nFile \"/opt/conda/lib/python3.7/site-packages/jupyter_client/session.py\", line 921, in deserialize\nraise ValueError(\"Duplicate Signature: %r\" % signature\n</code></pre>\n<p>I suspect a concurrent threading issue with RasterIO. I'm trying to use the 2 CPUs to speed up processing but it fails on some race conditions. I'm using thread lock but it does not seem to work as expected. If I commit again then it works. Same issue might happen on submit.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1187188,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-02-05T09:17:19.277000",
          "content": "<p>Confirmed, so it was me. No problem with Kaggle environment.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1191868,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-02-08T18:27:03.333000",
          "content": "<p>RasterIO documentation claims it would work:<br>\n<a href=\"https://rasterio.readthedocs.io/en/latest/topics/concurrency.html\" target=\"_blank\">https://rasterio.readthedocs.io/en/latest/topics/concurrency.html</a></p>\n<pre><code>        with rasterio.open(outfile, \"w\", **src.profile) as dst:\n            windows = [window for ij, window in dst.block_windows()]\n\n            # We cannot write to the same file from multiple threads\n            # without causing race conditions. To safely read/write\n            # from multiple threads, we use a lock to protect the\n            # DatasetReader/Writer\n            read_lock = threading.Lock()\n            write_lock = threading.Lock()\n\n            def process(window):\n                with read_lock:\n                    src_array = src.read(window=window)\n</code></pre>\n<p>But for some unknown reason it fails randomly. Anyone experimenting similar issue?</p>\n<p>This guy had similar problem but not with rasterio:<br>\n<a href=\"https://www.kaggle.com/product-feedback/164002\" target=\"_blank\">https://www.kaggle.com/product-feedback/164002</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1265564,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-04-07T03:06:02.517000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>    Will we get same error in case of oom ? I get sub csv error after 45 to 50 min .<br>\n. If there is an oom we get with in 30 min or after this much time only?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1277347,
      "author_name": "Luiz Felipe de Barros Jordão Costa",
      "author_url": "",
      "post_date": "2021-04-18T16:30:33.427000",
      "content": "<p>I'm having the same problem. Have you solved it? how?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1277349,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-18T16:30:55.277000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1211404,
      "author_name": "Himanshu Mehndiratta",
      "author_url": "",
      "post_date": "2021-02-20T07:39:40.203000",
      "content": "<p>I think the problem is with your end. Check again once.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1253543,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-26T21:05:35.537000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1183015": "I've a working inference kernel, I've updated/uploaded the weights of my model and I got \"**Submission CSV Not Found**\" error. When I look at diff between yesterday working version and today's broken version I can confirm that only path to weights has been updated. Am I alone to get this issue today? I've tried to re-submit and I've got the same error. It does not make sense, weights update cannot break the pipeline and the submission.csv and I can see it's generated on commit. ",
    "1215663": "Hi @mpware I just experienced the exact same issue that you described...only change in my kernel is that it is a new weights file that is loaded...a follow up version with yet another weights file is working.\n\nDoes not seem to happen often though...first time for me. \n",
    "1185435": "@mpware Hope, so you will resolve this soon.\n\n",
    "1183959": "Hope you will resolve soon. ",
    "1183525": "Today I've succeeded in submitting my kernel to get a score, so there seems to be little Kaggle scoring issue.",
    "1277347": "I'm having the same problem. Have you solved it? how?",
    "1211404": "I think the problem is with your end. Check again once.",
    "1253543": ""
  }
}