{
  "id": 223266,
  "title": "Submission Scoring Error Output",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/223266",
  "author_name": "",
  "post_date": "2021-03-03T06:08:30.908557700Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I haven't received this error before - I am trying to add more single models. Has any one dealt with this before. </p>",
  "messages": [
    {
      "id": "1224879",
      "postDate": "03/03/2021 06:08:30",
      "content": "<p>I haven't received this error before - I am trying to add more single models. Has any one dealt with this before. </p>",
      "rawMarkdown": "I haven't received this error before - I am trying to add more single models. Has any one dealt with this before.",
      "votes": null
    },
    {
      "id": "1224973",
      "postDate": "03/03/2021 08:00:23",
      "content": "<p>Here's some ideas that I've had in the past:</p>\n<ul>\n<li>I'll assume you've successfully submitted each script individually and you've tested that the whole script runs successfully on the visible public test set. If those things work, then that would exclude many of the most common issues.</li>\n<li>Are the different models interfering by running in the single script (some of these issues can probably be largely excluded, if this runs on the public test set, but some maybe not)? E.g. on the private test set more memory might be needed. The safest approach would be to call each script from a notebook e.g. like this <code>!ipython '../input/my-scripts/submission_model1.py'</code> (scripts stored in a dataset), write out a submission file with names like <code>submission_model1.csv</code> from each script, and then combine at the very end. That might be worth trying.</li>\n<li>Another candidate is issues with CUDA memory, where stuff accumulates until you restart the notebook (not an option here) or clear it up. For that issue, lowering batch sizes may also help, but I've had good experiences - for PyTorch - with <code>torch.cuda.empty_cache()</code> followed by <code>del ....</code> for anything that can go, followed by <code>gc.collect()</code> and then  <br>\n<code>torch.cuda.empty_cache()</code>, again.</li>\n<li>Another idea is that the way you combine models can somehow break on the test set (e.g. if you combine predictions pre-sigmoid, but for one model you get a 0 to 1 output that happens to end up going to exactly 0 or exactly 1 - due to available precision - on something in the private test set. If you then take the logit and get NaN, that could break things.</li>\n<li>Another rather unlikely idea: If you create lots of files in <code>./</code> that can be an issue (not if we talk a about a few dozen files, more if you did something like save a <code>.csv</code> for the predictions of each test set image), because the process at the end of the run might take too long with all of those (e.g. you cannot have a notebook creating a resized version of each image in this competition, unless you zip them up due to this, I suspect this could also theoretically be an issue during a submission).</li>\n</ul>\n<p>I'd be happy to learn about any other common pitfalls from others…</p>",
      "rawMarkdown": "Here's some ideas that I've had in the past:\n* I'll assume you've successfully submitted each script individually and you've tested that the whole script runs successfully on the visible public test set. If those things work, then that would exclude many of the most common issues.\n* Are the different models interfering by running in the single script (some of these issues can probably be largely excluded, if this runs on the public test set, but some maybe not)? E.g. on the private test set more memory might be needed. The safest approach would be to call each script from a notebook e.g. like this `!ipython '../input/my-scripts/submission_model1.py'` (scripts stored in a dataset), write out a submission file with names like `submission_model1.csv` from each script, and then combine at the very end. That might be worth trying.\n* Another candidate is issues with CUDA memory, where stuff accumulates until you restart the notebook (not an option here) or clear it up. For that issue, lowering batch sizes may also help, but I've had good experiences - for PyTorch - with `torch.cuda.empty_cache()` followed by `del ....` for anything that can go, followed by `gc.collect()` and then  \n `torch.cuda.empty_cache()`, again.\n* Another idea is that the way you combine models can somehow break on the test set (e.g. if you combine predictions pre-sigmoid, but for one model you get a 0 to 1 output that happens to end up going to exactly 0 or exactly 1 - due to available precision - on something in the private test set. If you then take the logit and get NaN, that could break things.\n* Another rather unlikely idea: If you create lots of files in `./` that can be an issue (not if we talk a about a few dozen files, more if you did something like save a `.csv` for the predictions of each test set image), because the process at the end of the run might take too long with all of those (e.g. you cannot have a notebook creating a resized version of each image in this competition, unless you zip them up due to this, I suspect this could also theoretically be an issue during a submission).\n\nI'd be happy to learn about any other common pitfalls from others...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1224973,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/03/2021 08:00:23",
      "content": "<p>Here's some ideas that I've had in the past:</p>\n<ul>\n<li>I'll assume you've successfully submitted each script individually and you've tested that the whole script runs successfully on the visible public test set. If those things work, then that would exclude many of the most common issues.</li>\n<li>Are the different models interfering by running in the single script (some of these issues can probably be largely excluded, if this runs on the public test set, but some maybe not)? E.g. on the private test set more memory might be needed. The safest approach would be to call each script from a notebook e.g. like this <code>!ipython '../input/my-scripts/submission_model1.py'</code> (scripts stored in a dataset), write out a submission file with names like <code>submission_model1.csv</code> from each script, and then combine at the very end. That might be worth trying.</li>\n<li>Another candidate is issues with CUDA memory, where stuff accumulates until you restart the notebook (not an option here) or clear it up. For that issue, lowering batch sizes may also help, but I've had good experiences - for PyTorch - with <code>torch.cuda.empty_cache()</code> followed by <code>del ....</code> for anything that can go, followed by <code>gc.collect()</code> and then  <br>\n<code>torch.cuda.empty_cache()</code>, again.</li>\n<li>Another idea is that the way you combine models can somehow break on the test set (e.g. if you combine predictions pre-sigmoid, but for one model you get a 0 to 1 output that happens to end up going to exactly 0 or exactly 1 - due to available precision - on something in the private test set. If you then take the logit and get NaN, that could break things.</li>\n<li>Another rather unlikely idea: If you create lots of files in <code>./</code> that can be an issue (not if we talk a about a few dozen files, more if you did something like save a <code>.csv</code> for the predictions of each test set image), because the process at the end of the run might take too long with all of those (e.g. you cannot have a notebook creating a resized version of each image in this competition, unless you zip them up due to this, I suspect this could also theoretically be an issue during a submission).</li>\n</ul>\n<p>I'd be happy to learn about any other common pitfalls from others…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1224879": "I haven't received this error before - I am trying to add more single models. Has any one dealt with this before.",
    "1224973": "Here's some ideas that I've had in the past:\n* I'll assume you've successfully submitted each script individually and you've tested that the whole script runs successfully on the visible public test set. If those things work, then that would exclude many of the most common issues.\n* Are the different models interfering by running in the single script (some of these issues can probably be largely excluded, if this runs on the public test set, but some maybe not)? E.g. on the private test set more memory might be needed. The safest approach would be to call each script from a notebook e.g. like this `!ipython '../input/my-scripts/submission_model1.py'` (scripts stored in a dataset), write out a submission file with names like `submission_model1.csv` from each script, and then combine at the very end. That might be worth trying.\n* Another candidate is issues with CUDA memory, where stuff accumulates until you restart the notebook (not an option here) or clear it up. For that issue, lowering batch sizes may also help, but I've had good experiences - for PyTorch - with `torch.cuda.empty_cache()` followed by `del ....` for anything that can go, followed by `gc.collect()` and then  \n `torch.cuda.empty_cache()`, again.\n* Another idea is that the way you combine models can somehow break on the test set (e.g. if you combine predictions pre-sigmoid, but for one model you get a 0 to 1 output that happens to end up going to exactly 0 or exactly 1 - due to available precision - on something in the private test set. If you then take the logit and get NaN, that could break things.\n* Another rather unlikely idea: If you create lots of files in `./` that can be an issue (not if we talk a about a few dozen files, more if you did something like save a `.csv` for the predictions of each test set image), because the process at the end of the run might take too long with all of those (e.g. you cannot have a notebook creating a resized version of each image in this competition, unless you zip them up due to this, I suspect this could also theoretically be an issue during a submission).\n\nI'd be happy to learn about any other common pitfalls from others..."
  },
  "source": "meta"
}