{
  "id": 384878,
  "title": "Notebook Threw Exception",
  "url": "/competitions/nfl-player-contact-detection/discussion/384878",
  "author_name": "",
  "post_date": "2023-02-10T01:26:06.704545700Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have a problem with submissions.<br>\nMy notebook is well done in public data and the log has no problem.<br>\nBut, an error \"Notebook Threw Exception\" is raised when the notebook is re-running with private data. <br>\nI save preprocessed frames in hdd for fatser data loading, it occupied 500Mb in public data. Is it to be a problem? </p>\n<p>I have no idea why the error is raised. <br>\nPlease share your experience with it.</p>",
  "messages": [
    {
      "id": "2137419",
      "postDate": "02/10/2023 01:26:06",
      "content": "<p>I have a problem with submissions.<br>\nMy notebook is well done in public data and the log has no problem.<br>\nBut, an error \"Notebook Threw Exception\" is raised when the notebook is re-running with private data. <br>\nI save preprocessed frames in hdd for fatser data loading, it occupied 500Mb in public data. Is it to be a problem? </p>\n<p>I have no idea why the error is raised. <br>\nPlease share your experience with it.</p>",
      "rawMarkdown": "I have a problem with submissions.\nMy notebook is well done in public data and the log has no problem.\nBut, an error \"Notebook Threw Exception\" is raised when the notebook is re-running with private data. \nI save preprocessed frames in hdd for fatser data loading, it occupied 500Mb in public data. Is it to be a problem? \n\nI have no idea why the error is raised. \nPlease share your experience with it.",
      "votes": null
    },
    {
      "id": "2139704",
      "postDate": "02/11/2023 03:49:21",
      "content": "<p>It can be really hard to debug because of the (intentional) lack of feedback, but here are some tips:</p>\n<ul>\n<li><p>If your notebook runs out of hdd space, then it will give a different error: \"Notebook out of disk\", so you're ok there</p></li>\n<li><p>I believe you'll get a similar error if you run out of gpu ram, but I don't 100% remember.</p></li>\n<li><p>one tip is to run just parts of your code at a time, and submit that (you'll still have to write a submission.csv file to submit, but it could be all 0s for example). That can help you narrow down what the issue is.</p></li>\n</ul>\n<p>The most common error I've personally run into has been in my dataloaders - trying to load a video frame that doesn't exist, etc, which causes it to throw an exception. </p>\n<p>I tend to add try/excepts all over the place until it works, and then slowly remove them until it breaks, and then at least I know where the error is :)</p>\n<p>Hopefully that helps!</p>",
      "rawMarkdown": "It can be really hard to debug because of the (intentional) lack of feedback, but here are some tips:\n\n- If your notebook runs out of hdd space, then it will give a different error: \"Notebook out of disk\", so you're ok there\n\n- I believe you'll get a similar error if you run out of gpu ram, but I don't 100% remember.\n\n- one tip is to run just parts of your code at a time, and submit that (you'll still have to write a submission.csv file to submit, but it could be all 0s for example). That can help you narrow down what the issue is.\n\nThe most common error I've personally run into has been in my dataloaders - trying to load a video frame that doesn't exist, etc, which causes it to throw an exception. \n\nI tend to add try/excepts all over the place until it works, and then slowly remove them until it breaks, and then at least I know where the error is :)\n\nHopefully that helps!",
      "votes": null
    },
    {
      "id": "2140216",
      "postDate": "02/11/2023 14:58:28",
      "content": "<p>Thank you for sharing your experience and tips.<br>\nI am using batch size with 1, so i believe gpu ram is not exhausted,<br>\nSo, I should debug the error by running parts of my code. <br>\nI'll share the problem if i fix it. Thanks :)</p>",
      "rawMarkdown": "Thank you for sharing your experience and tips.\nI am using batch size with 1, so i believe gpu ram is not exhausted,\nSo, I should debug the error by running parts of my code. \nI'll share the problem if i fix it. Thanks :)",
      "votes": null
    },
    {
      "id": "2140294",
      "postDate": "02/11/2023 16:11:13",
      "content": "<p>Currently, I am using my local machine for training model.<br>\nThe trained model and source code (including model definition) is uploaded in datasets as privately.<br>\nI imported the trained model and source code in inference notebook.<br>\nCode requirements of this competition includes \"Freely &amp; publicly available external data is allowed, including pre-trained models\", Is it means using privately uploaded data is not allowed? </p>",
      "rawMarkdown": "Currently, I am using my local machine for training model.\nThe trained model and source code (including model definition) is uploaded in datasets as privately.\nI imported the trained model and source code in inference notebook.\nCode requirements of this competition includes \"Freely & publicly available external data is allowed, including pre-trained models\", Is it means using privately uploaded data is not allowed?",
      "votes": null
    },
    {
      "id": "2140306",
      "postDate": "02/11/2023 16:21:49",
      "content": "<p>What you're doing is just fine - yes it's allowed.</p>\n<p>That line is referring to pre-trained models like resnet, which is trained on image net. I agree it's confusing wording :)</p>\n<p>But yes, it's fine to train your own models and upload them as private datasets.</p>",
      "rawMarkdown": "What you're doing is just fine - yes it's allowed.\n\nThat line is referring to pre-trained models like resnet, which is trained on image net. I agree it's confusing wording :)\n\nBut yes, it's fine to train your own models and upload them as private datasets.",
      "votes": null
    },
    {
      "id": "2141622",
      "postDate": "02/13/2023 00:28:10",
      "content": "<p>I fixed the error by submitting parts of my code.<br>\nThe problem was in preprocessing function and it was attribute error.<br>\nI wrote code expecting pd.Series datatype, but the variable is sometimes numpy in private data.<br>\nThank you for your adivce :)</p>",
      "rawMarkdown": "I fixed the error by submitting parts of my code.\nThe problem was in preprocessing function and it was attribute error.\nI wrote code expecting pd.Series datatype, but the variable is sometimes numpy in private data.\nThank you for your adivce :)",
      "votes": null
    },
    {
      "id": "2164864",
      "postDate": "03/01/2023 20:53:25",
      "content": "<p>Hi, are you sure about this? I'm new to the platform as well, can you point me to the rules where they mention this is allowed?</p>\n<p>If uploading locally models are allowed, then this changes everything (as well as sounding quite unfair to those without access to computing power).</p>",
      "rawMarkdown": "Hi, are you sure about this? I'm new to the platform as well, can you point me to the rules where they mention this is allowed?\n\nIf uploading locally models are allowed, then this changes everything (as well as sounding quite unfair to those without access to computing power).",
      "votes": null
    },
    {
      "id": "2164894",
      "postDate": "03/01/2023 21:15:37",
      "content": "<p>Yes, I'm very sure about this (and yes, I agree it's confusing, and Kaggle may want to think about the wording of that page since I see multiple people in each comp who are confused about it :) )</p>\n<p>Basically: for <em>most</em> competitions, you are allowed to do almost anything you want locally, and then upload it as a private dataset and use that - including pretraining models. The big exception is that you can't hand-label any <em>test</em> data. That doesn't apply so much for code competitions (where the test data is hidden), but can be a big (and gray) issue on non-code competitions, where the test data is available.</p>\n<p>For <em>some</em> competitions, there is an additional rule that you must do both your training and inference in the 9 hour (or whatever) limit - but that is always explicitly called out in the rules.</p>\n<p>Yes, that does mean that for many competitions, it is unfair to people without computing power. Kaggle is very generous with their gpus available, but it can't compete with a dedicated gpu or cluster - especially for very large data competitions like this one (probably) will be.</p>",
      "rawMarkdown": "Yes, I'm very sure about this (and yes, I agree it's confusing, and Kaggle may want to think about the wording of that page since I see multiple people in each comp who are confused about it :) )\n\nBasically: for _most_ competitions, you are allowed to do almost anything you want locally, and then upload it as a private dataset and use that - including pretraining models. The big exception is that you can't hand-label any _test_ data. That doesn't apply so much for code competitions (where the test data is hidden), but can be a big (and gray) issue on non-code competitions, where the test data is available.\n\nFor _some_ competitions, there is an additional rule that you must do both your training and inference in the 9 hour (or whatever) limit - but that is always explicitly called out in the rules.\n\nYes, that does mean that for many competitions, it is unfair to people without computing power. Kaggle is very generous with their gpus available, but it can't compete with a dedicated gpu or cluster - especially for very large data competitions like this one (probably) will be.",
      "votes": null
    },
    {
      "id": "2164948",
      "postDate": "03/01/2023 21:57:44",
      "content": "<p>That does definitely explain how every one manages to pull off 1-2 hr notebooks while I barely managed to fit in 9 hrs, as well as how I'll approach my next competitions. Thanks!</p>",
      "rawMarkdown": "That does definitely explain how every one manages to pull off 1-2 hr notebooks while I barely managed to fit in 9 hrs, as well as how I'll approach my next competitions. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2139704,
      "author_name": "chris62",
      "author_url": "",
      "post_date": "02/11/2023 03:49:21",
      "content": "<p>It can be really hard to debug because of the (intentional) lack of feedback, but here are some tips:</p>\n<ul>\n<li><p>If your notebook runs out of hdd space, then it will give a different error: \"Notebook out of disk\", so you're ok there</p></li>\n<li><p>I believe you'll get a similar error if you run out of gpu ram, but I don't 100% remember.</p></li>\n<li><p>one tip is to run just parts of your code at a time, and submit that (you'll still have to write a submission.csv file to submit, but it could be all 0s for example). That can help you narrow down what the issue is.</p></li>\n</ul>\n<p>The most common error I've personally run into has been in my dataloaders - trying to load a video frame that doesn't exist, etc, which causes it to throw an exception. </p>\n<p>I tend to add try/excepts all over the place until it works, and then slowly remove them until it breaks, and then at least I know where the error is :)</p>\n<p>Hopefully that helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2140216,
          "author_name": "wonjunjg",
          "author_url": "",
          "post_date": "02/11/2023 14:58:28",
          "content": "<p>Thank you for sharing your experience and tips.<br>\nI am using batch size with 1, so i believe gpu ram is not exhausted,<br>\nSo, I should debug the error by running parts of my code. <br>\nI'll share the problem if i fix it. Thanks :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2141622,
          "author_name": "wonjunjg",
          "author_url": "",
          "post_date": "02/13/2023 00:28:10",
          "content": "<p>I fixed the error by submitting parts of my code.<br>\nThe problem was in preprocessing function and it was attribute error.<br>\nI wrote code expecting pd.Series datatype, but the variable is sometimes numpy in private data.<br>\nThank you for your adivce :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2140294,
      "author_name": "wonjunjg",
      "author_url": "",
      "post_date": "02/11/2023 16:11:13",
      "content": "<p>Currently, I am using my local machine for training model.<br>\nThe trained model and source code (including model definition) is uploaded in datasets as privately.<br>\nI imported the trained model and source code in inference notebook.<br>\nCode requirements of this competition includes \"Freely &amp; publicly available external data is allowed, including pre-trained models\", Is it means using privately uploaded data is not allowed? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2140306,
          "author_name": "chris62",
          "author_url": "",
          "post_date": "02/11/2023 16:21:49",
          "content": "<p>What you're doing is just fine - yes it's allowed.</p>\n<p>That line is referring to pre-trained models like resnet, which is trained on image net. I agree it's confusing wording :)</p>\n<p>But yes, it's fine to train your own models and upload them as private datasets.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2164864,
              "author_name": "sirapoabchaikunsaeng",
              "author_url": "",
              "post_date": "03/01/2023 20:53:25",
              "content": "<p>Hi, are you sure about this? I'm new to the platform as well, can you point me to the rules where they mention this is allowed?</p>\n<p>If uploading locally models are allowed, then this changes everything (as well as sounding quite unfair to those without access to computing power).</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2164894,
                  "author_name": "chris62",
                  "author_url": "",
                  "post_date": "03/01/2023 21:15:37",
                  "content": "<p>Yes, I'm very sure about this (and yes, I agree it's confusing, and Kaggle may want to think about the wording of that page since I see multiple people in each comp who are confused about it :) )</p>\n<p>Basically: for <em>most</em> competitions, you are allowed to do almost anything you want locally, and then upload it as a private dataset and use that - including pretraining models. The big exception is that you can't hand-label any <em>test</em> data. That doesn't apply so much for code competitions (where the test data is hidden), but can be a big (and gray) issue on non-code competitions, where the test data is available.</p>\n<p>For <em>some</em> competitions, there is an additional rule that you must do both your training and inference in the 9 hour (or whatever) limit - but that is always explicitly called out in the rules.</p>\n<p>Yes, that does mean that for many competitions, it is unfair to people without computing power. Kaggle is very generous with their gpus available, but it can't compete with a dedicated gpu or cluster - especially for very large data competitions like this one (probably) will be.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2164948,
                      "author_name": "sirapoabchaikunsaeng",
                      "author_url": "",
                      "post_date": "03/01/2023 21:57:44",
                      "content": "<p>That does definitely explain how every one manages to pull off 1-2 hr notebooks while I barely managed to fit in 9 hrs, as well as how I'll approach my next competitions. Thanks!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2137419": "I have a problem with submissions.\nMy notebook is well done in public data and the log has no problem.\nBut, an error \"Notebook Threw Exception\" is raised when the notebook is re-running with private data. \nI save preprocessed frames in hdd for fatser data loading, it occupied 500Mb in public data. Is it to be a problem? \n\nI have no idea why the error is raised. \nPlease share your experience with it.",
    "2139704": "It can be really hard to debug because of the (intentional) lack of feedback, but here are some tips:\n\n- If your notebook runs out of hdd space, then it will give a different error: \"Notebook out of disk\", so you're ok there\n\n- I believe you'll get a similar error if you run out of gpu ram, but I don't 100% remember.\n\n- one tip is to run just parts of your code at a time, and submit that (you'll still have to write a submission.csv file to submit, but it could be all 0s for example). That can help you narrow down what the issue is.\n\nThe most common error I've personally run into has been in my dataloaders - trying to load a video frame that doesn't exist, etc, which causes it to throw an exception. \n\nI tend to add try/excepts all over the place until it works, and then slowly remove them until it breaks, and then at least I know where the error is :)\n\nHopefully that helps!",
    "2140216": "Thank you for sharing your experience and tips.\nI am using batch size with 1, so i believe gpu ram is not exhausted,\nSo, I should debug the error by running parts of my code. \nI'll share the problem if i fix it. Thanks :)",
    "2140294": "Currently, I am using my local machine for training model.\nThe trained model and source code (including model definition) is uploaded in datasets as privately.\nI imported the trained model and source code in inference notebook.\nCode requirements of this competition includes \"Freely & publicly available external data is allowed, including pre-trained models\", Is it means using privately uploaded data is not allowed?",
    "2140306": "What you're doing is just fine - yes it's allowed.\n\nThat line is referring to pre-trained models like resnet, which is trained on image net. I agree it's confusing wording :)\n\nBut yes, it's fine to train your own models and upload them as private datasets.",
    "2141622": "I fixed the error by submitting parts of my code.\nThe problem was in preprocessing function and it was attribute error.\nI wrote code expecting pd.Series datatype, but the variable is sometimes numpy in private data.\nThank you for your adivce :)",
    "2164864": "Hi, are you sure about this? I'm new to the platform as well, can you point me to the rules where they mention this is allowed?\n\nIf uploading locally models are allowed, then this changes everything (as well as sounding quite unfair to those without access to computing power).",
    "2164894": "Yes, I'm very sure about this (and yes, I agree it's confusing, and Kaggle may want to think about the wording of that page since I see multiple people in each comp who are confused about it :) )\n\nBasically: for _most_ competitions, you are allowed to do almost anything you want locally, and then upload it as a private dataset and use that - including pretraining models. The big exception is that you can't hand-label any _test_ data. That doesn't apply so much for code competitions (where the test data is hidden), but can be a big (and gray) issue on non-code competitions, where the test data is available.\n\nFor _some_ competitions, there is an additional rule that you must do both your training and inference in the 9 hour (or whatever) limit - but that is always explicitly called out in the rules.\n\nYes, that does mean that for many competitions, it is unfair to people without computing power. Kaggle is very generous with their gpus available, but it can't compete with a dedicated gpu or cluster - especially for very large data competitions like this one (probably) will be.",
    "2164948": "That does definitely explain how every one manages to pull off 1-2 hr notebooks while I barely managed to fit in 9 hrs, as well as how I'll approach my next competitions. Thanks!"
  },
  "source": "meta"
}