{
  "id": 223281,
  "title": "Does notebook running time limit include the scoring time?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/223281",
  "author_name": "Correlation",
  "post_date": "2021-03-03T07:26:55.960000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I wonder if notebook running time includes scoring time? I 'm sure my notebook running time is about 7 hours, but it‘s time out.</p>",
  "messages": [
    {
      "id": 1224956,
      "postDate": "2021-03-03T07:26:55.960Z",
      "content": "<p>I wonder if notebook running time includes scoring time? I 'm sure my notebook running time is about 7 hours, but it‘s time out.</p>",
      "rawMarkdown": "I wonder if notebook running time includes scoring time? I 'm sure my notebook running time is about 7 hours, but it‘s time out.",
      "votes": 4
    },
    {
      "id": 1225248,
      "postDate": "2021-03-03T13:19:20.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/daishu\" target=\"_blank\">@daishu</a> - Is your notebook running time 7 hours on the public test data?</p>\n<p>I have submitted many times now and have changed the way I submit many times, and I have never noticed the scoring time to be included. That being said I'm not an expert, perhaps a Kaggle admin might better answer. I did hope to share some observations with you if that's alright.</p>\n<hr>\n<p><strong>My observations are the following:</strong></p>\n<ul>\n<li>The private test set is 2-3 times larger than the public test set</li>\n<li>If your notebook takes 2 hours to run on the public test set, it will take 4-6 hours to run when you submit your <strong><code>submission.csv</code></strong></li>\n<li>To see your LB score (and get 0 on the private LB), you only need to predict on the public test images and leave the default prediction string for the private test set (this will save you 2/3rds of the submission time). Although, obviously you will need to be submitting predictions on both public and private data when it comes time for your final submissions.</li>\n<li>You only need to generate a submission.csv file on a few predictions when submitting (this will save you &gt;90% of the time it takes to run your notebook on the public test set). i.e. <strong><code>IS_DEMO=len(ss_df)==559</code></strong> and then use the flag <strong><code>IS_DEMO</code></strong> to control whether you are predicting on the entire public test set or only a few images. This works because during the Save&amp;Commit part, your <strong><code>ss_df</code></strong> will have a length of 559… however, when the notebook is run on the backend, the length of <strong><code>ss_df</code></strong> will probably be closer to 1250 or something.</li>\n</ul>\n<hr>\n<p><em>My public notebooks and those by other talented Kagglers show these observations in action… <strong>I hope this helps!</strong></em></p>",
      "rawMarkdown": "@daishu - Is your notebook running time 7 hours on the public test data?\n\nI have submitted many times now and have changed the way I submit many times, and I have never noticed the scoring time to be included. That being said I'm not an expert, perhaps a Kaggle admin might better answer. I did hope to share some observations with you if that's alright.\n\n---\n\n**My observations are the following:**\n\n- The private test set is 2-3 times larger than the public test set\n- If your notebook takes 2 hours to run on the public test set, it will take 4-6 hours to run when you submit your **`submission.csv`**\n- To see your LB score (and get 0 on the private LB), you only need to predict on the public test images and leave the default prediction string for the private test set (this will save you 2/3rds of the submission time). Although, obviously you will need to be submitting predictions on both public and private data when it comes time for your final submissions.\n- You only need to generate a submission.csv file on a few predictions when submitting (this will save you >90% of the time it takes to run your notebook on the public test set). i.e. **`IS_DEMO=len(ss_df)==559`** and then use the flag **`IS_DEMO`** to control whether you are predicting on the entire public test set or only a few images. This works because during the Save&Commit part, your **`ss_df`** will have a length of 559... however, when the notebook is run on the backend, the length of **`ss_df`** will probably be closer to 1250 or something.\n\n---\n\n*My public notebooks and those by other talented Kagglers show these observations in action... **I hope this helps!***",
      "votes": 1,
      "replies": [
        {
          "id": 1225309,
          "postDate": "2021-03-03T14:25:03.010Z",
          "content": "<p>Thanks, my notebook takes 2.5 hours on public dataset. So I guess it's about 7 hours on private dataset.</p>",
          "rawMarkdown": "Thanks, my notebook takes 2.5 hours on public dataset. So I guess it's about 7 hours on private dataset."
        },
        {
          "id": 1225704,
          "postDate": "2021-03-03T20:52:52.367Z",
          "content": "<p>Hmm, it seems to take 150*100/31=&gt;8 hrs for all test data by simple calculation.<br>\nAlthough this may depend on your model, it will take more time if private includes more 3072 images or…</p>",
          "rawMarkdown": "Hmm, it seems to take 150*100/31=>8 hrs for all test data by simple calculation.\nAlthough this may depend on your model, it will take more time if private includes more 3072 images or...",
          "votes": 1
        },
        {
          "id": 1250816,
          "postDate": "2021-03-24T09:42:55.363Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> </p>\n<p>What do you mean by this? </p>\n<p><code>You only need to generate a submission.csv file on a few predictions when submitting</code></p>\n<p>I don't get it, you are supposed to generate a submission file with all cell predictions</p>",
          "rawMarkdown": "Hi @dschettler8845 \n\nWhat do you mean by this? \n\n`You only need to generate a submission.csv file on a few predictions when submitting`\n\nI don't get it, you are supposed to generate a submission file with all cell predictions"
        },
        {
          "id": 1251055,
          "postDate": "2021-03-24T13:17:22.530Z",
          "content": "<p><a href=\"https://www.kaggle.com/glopezzz\" target=\"_blank\">@glopezzz</a> (nice avatar) - There are two submission events that occur when you want to submit. </p>\n<ol>\n<li>When you <strong>'save and run all'</strong> the notebook to generate the <strong><code>submission.csv</code></strong> file.</li>\n<li>When you submit the <strong><code>submission.csv</code></strong> file and Kaggle reruns the notebook on the backend with the Private data swapped in for the Public data (3-4 times more data).</li>\n</ol>\n<hr>\n<p>My comment …</p>\n<blockquote>\n  <p>\"You only need to generate a submission.csv file on a few predictions when submitting\"</p>\n</blockquote>\n<p>… refers to the first type of submission event (generating the <strong><code>submission.csv</code></strong> on public data). If you have a catch in your notebook that identifies whether or not you are using the public data, …</p>\n<pre><code>IS_PUBLIC == len(ss_df)==559 # I think this is what it is... I can't remember exactly\n</code></pre>\n<p>… which essentially checks that the length of the <strong><code>sample_submission.csv</code></strong> is the length of the known public <strong><code>sample_submission.csv</code></strong>, you only need to generate predictions on a couple of the rows. In fact, you don't really need to generate any predictions, I simply do that to make sure that everything looks good and so that I can demo what the predictions look like as my kernel is public. </p>\n<p>All Kaggle is looking for is a <strong><code>submission.csv</code></strong> file in the output directory. If it finds that, it will then rerun the notebook. When it reruns the notebook the previously defined catch will be <strong><code>False</code></strong> (as the length of the <strong><code>sample_submission.csv</code></strong> will be greater than 559) and then you can use that flag to make sure you evaluate ALL of the data prior to submission (or just the public part if you want to save time and probe the LB).</p>\n<hr>\n<p>I hope this makes sense! If not feel free to reply and I can clarify. My submission flow looks like this…</p>\n<ol>\n<li>Submit kernel so that <strong><code>IS_PUBLIC</code></strong> evaluates to <strong><code>True</code></strong> and I only infer on a few rows from the public data. <strong><em>(3-5 minutes to complete)</em></strong></li>\n<li>Submit the <strong><code>submission.csv</code></strong> file generated in step 1 knowing that <strong><code>IS_PUBLIC</code></strong> evaluates to <strong><code>False</code></strong> due to the presence of the private data. As such I infer on all of the data. <strong><em>(3-6 hours depending)</em></strong><ul>\n<li><em>Note that in step 2 I actually only evaluate on the Public Test data and as such my LB score is still accurate (however I score 0 on the Private LB). This speeds it up so that my submission completes in under 1 hour.</em></li></ul></li>\n</ol>",
          "rawMarkdown": "@glopezzz (nice avatar) - There are two submission events that occur when you want to submit. \n\n1. When you **'save and run all'** the notebook to generate the **`submission.csv`** file.\n2. When you submit the **`submission.csv`** file and Kaggle reruns the notebook on the backend with the Private data swapped in for the Public data (3-4 times more data).\n\n---\n\nMy comment ...\n\n> \"You only need to generate a submission.csv file on a few predictions when submitting\"\n\n... refers to the first type of submission event (generating the **`submission.csv`** on public data). If you have a catch in your notebook that identifies whether or not you are using the public data, ...\n\n```\nIS_PUBLIC == len(ss_df)==559 # I think this is what it is... I can't remember exactly\n```\n\n... which essentially checks that the length of the **`sample_submission.csv`** is the length of the known public **`sample_submission.csv`**, you only need to generate predictions on a couple of the rows. In fact, you don't really need to generate any predictions, I simply do that to make sure that everything looks good and so that I can demo what the predictions look like as my kernel is public. \n\nAll Kaggle is looking for is a **`submission.csv`** file in the output directory. If it finds that, it will then rerun the notebook. When it reruns the notebook the previously defined catch will be **`False`** (as the length of the **`sample_submission.csv`** will be greater than 559) and then you can use that flag to make sure you evaluate ALL of the data prior to submission (or just the public part if you want to save time and probe the LB).\n\n---\n\nI hope this makes sense! If not feel free to reply and I can clarify. My submission flow looks like this...\n\n1. Submit kernel so that **`IS_PUBLIC`** evaluates to **`True`** and I only infer on a few rows from the public data. ***(3-5 minutes to complete)***\n2. Submit the **`submission.csv`** file generated in step 1 knowing that **`IS_PUBLIC`** evaluates to **`False`** due to the presence of the private data. As such I infer on all of the data. ***(3-6 hours depending)***\n  * *Note that in step 2 I actually only evaluate on the Public Test data and as such my LB score is still accurate (however I score 0 on the Private LB). This speeds it up so that my submission completes in under 1 hour.*",
          "votes": 4
        },
        {
          "id": 1251086,
          "postDate": "2021-03-24T13:33:56.430Z",
          "content": "<p>Ok, I totally get it now, thanks a lot for the detailed explanation, I really appreciate it :)</p>\n<p>I've having trouble with the computing time, my notebook takes around 5h with the public test data :c<br>\nAt least with your great idea, I can save a lot of time in the 'save and commit' process.</p>\n<p>Thanks again</p>",
          "rawMarkdown": "Ok, I totally get it now, thanks a lot for the detailed explanation, I really appreciate it :)\n\nI've having trouble with the computing time, my notebook takes around 5h with the public test data :c\nAt least with your great idea, I can save a lot of time in the 'save and commit' process.\n\nThanks again",
          "votes": 2
        },
        {
          "id": 1251177,
          "postDate": "2021-03-24T14:46:27.727Z",
          "content": "<p>Glad it helped!</p>\n<p>If you aren't already, I would recommend using an approach similar to the notebooks/methodology created by <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\"><strong>linshokaku</strong></a> and refined by <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\"><strong>Raman</strong></a> for the actual Cell Segmentation… as that is often the bottleneck in speed (it was for me).</p>\n<hr>\n<p><strong><em>Please see these notebooks for more details.</em></strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\" target=\"_blank\">https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation</a></li>\n<li><a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation/notebook\" target=\"_blank\">https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation/notebook</a></li>\n<li><a href=\"https://github.com/SamusRam/HPA-Cell-Segmentation\" target=\"_blank\">https://github.com/SamusRam/HPA-Cell-Segmentation</a></li>\n</ul>",
          "rawMarkdown": "Glad it helped!\n\nIf you aren't already, I would recommend using an approach similar to the notebooks/methodology created by [**linshokaku**](https://www.kaggle.com/linshokaku) and refined by [**Raman**](https://www.kaggle.com/samusram) for the actual Cell Segmentation... as that is often the bottleneck in speed (it was for me).\n\n---\n\n***Please see these notebooks for more details.***\n\n- https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\n- https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation/notebook\n- https://github.com/SamusRam/HPA-Cell-Segmentation",
          "votes": 1
        },
        {
          "id": 1251219,
          "postDate": "2021-03-24T15:32:35.390Z",
          "content": "<p>That's awesome!! Thanks for the information, I'll be working with that from now on :)</p>",
          "rawMarkdown": "That's awesome!! Thanks for the information, I'll be working with that from now on :)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1225248,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-03-03T13:19:20.037000",
      "content": "<p><a href=\"https://www.kaggle.com/daishu\" target=\"_blank\">@daishu</a> - Is your notebook running time 7 hours on the public test data?</p>\n<p>I have submitted many times now and have changed the way I submit many times, and I have never noticed the scoring time to be included. That being said I'm not an expert, perhaps a Kaggle admin might better answer. I did hope to share some observations with you if that's alright.</p>\n<hr>\n<p><strong>My observations are the following:</strong></p>\n<ul>\n<li>The private test set is 2-3 times larger than the public test set</li>\n<li>If your notebook takes 2 hours to run on the public test set, it will take 4-6 hours to run when you submit your <strong><code>submission.csv</code></strong></li>\n<li>To see your LB score (and get 0 on the private LB), you only need to predict on the public test images and leave the default prediction string for the private test set (this will save you 2/3rds of the submission time). Although, obviously you will need to be submitting predictions on both public and private data when it comes time for your final submissions.</li>\n<li>You only need to generate a submission.csv file on a few predictions when submitting (this will save you &gt;90% of the time it takes to run your notebook on the public test set). i.e. <strong><code>IS_DEMO=len(ss_df)==559</code></strong> and then use the flag <strong><code>IS_DEMO</code></strong> to control whether you are predicting on the entire public test set or only a few images. This works because during the Save&amp;Commit part, your <strong><code>ss_df</code></strong> will have a length of 559… however, when the notebook is run on the backend, the length of <strong><code>ss_df</code></strong> will probably be closer to 1250 or something.</li>\n</ul>\n<hr>\n<p><em>My public notebooks and those by other talented Kagglers show these observations in action… <strong>I hope this helps!</strong></em></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1225309,
          "author_name": "Correlation",
          "author_url": "",
          "post_date": "2021-03-03T14:25:03.010000",
          "content": "<p>Thanks, my notebook takes 2.5 hours on public dataset. So I guess it's about 7 hours on private dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1225704,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-03-03T20:52:52.367000",
          "content": "<p>Hmm, it seems to take 150*100/31=&gt;8 hrs for all test data by simple calculation.<br>\nAlthough this may depend on your model, it will take more time if private includes more 3072 images or…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1250816,
          "author_name": "glopezzz",
          "author_url": "",
          "post_date": "2021-03-24T09:42:55.363000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> </p>\n<p>What do you mean by this? </p>\n<p><code>You only need to generate a submission.csv file on a few predictions when submitting</code></p>\n<p>I don't get it, you are supposed to generate a submission file with all cell predictions</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1251055,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-24T13:17:22.530000",
          "content": "<p><a href=\"https://www.kaggle.com/glopezzz\" target=\"_blank\">@glopezzz</a> (nice avatar) - There are two submission events that occur when you want to submit. </p>\n<ol>\n<li>When you <strong>'save and run all'</strong> the notebook to generate the <strong><code>submission.csv</code></strong> file.</li>\n<li>When you submit the <strong><code>submission.csv</code></strong> file and Kaggle reruns the notebook on the backend with the Private data swapped in for the Public data (3-4 times more data).</li>\n</ol>\n<hr>\n<p>My comment …</p>\n<blockquote>\n  <p>\"You only need to generate a submission.csv file on a few predictions when submitting\"</p>\n</blockquote>\n<p>… refers to the first type of submission event (generating the <strong><code>submission.csv</code></strong> on public data). If you have a catch in your notebook that identifies whether or not you are using the public data, …</p>\n<pre><code>IS_PUBLIC == len(ss_df)==559 # I think this is what it is... I can't remember exactly\n</code></pre>\n<p>… which essentially checks that the length of the <strong><code>sample_submission.csv</code></strong> is the length of the known public <strong><code>sample_submission.csv</code></strong>, you only need to generate predictions on a couple of the rows. In fact, you don't really need to generate any predictions, I simply do that to make sure that everything looks good and so that I can demo what the predictions look like as my kernel is public. </p>\n<p>All Kaggle is looking for is a <strong><code>submission.csv</code></strong> file in the output directory. If it finds that, it will then rerun the notebook. When it reruns the notebook the previously defined catch will be <strong><code>False</code></strong> (as the length of the <strong><code>sample_submission.csv</code></strong> will be greater than 559) and then you can use that flag to make sure you evaluate ALL of the data prior to submission (or just the public part if you want to save time and probe the LB).</p>\n<hr>\n<p>I hope this makes sense! If not feel free to reply and I can clarify. My submission flow looks like this…</p>\n<ol>\n<li>Submit kernel so that <strong><code>IS_PUBLIC</code></strong> evaluates to <strong><code>True</code></strong> and I only infer on a few rows from the public data. <strong><em>(3-5 minutes to complete)</em></strong></li>\n<li>Submit the <strong><code>submission.csv</code></strong> file generated in step 1 knowing that <strong><code>IS_PUBLIC</code></strong> evaluates to <strong><code>False</code></strong> due to the presence of the private data. As such I infer on all of the data. <strong><em>(3-6 hours depending)</em></strong><ul>\n<li><em>Note that in step 2 I actually only evaluate on the Public Test data and as such my LB score is still accurate (however I score 0 on the Private LB). This speeds it up so that my submission completes in under 1 hour.</em></li></ul></li>\n</ol>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1251086,
          "author_name": "glopezzz",
          "author_url": "",
          "post_date": "2021-03-24T13:33:56.430000",
          "content": "<p>Ok, I totally get it now, thanks a lot for the detailed explanation, I really appreciate it :)</p>\n<p>I've having trouble with the computing time, my notebook takes around 5h with the public test data :c<br>\nAt least with your great idea, I can save a lot of time in the 'save and commit' process.</p>\n<p>Thanks again</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1251177,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-24T14:46:27.727000",
          "content": "<p>Glad it helped!</p>\n<p>If you aren't already, I would recommend using an approach similar to the notebooks/methodology created by <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\"><strong>linshokaku</strong></a> and refined by <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\"><strong>Raman</strong></a> for the actual Cell Segmentation… as that is often the bottleneck in speed (it was for me).</p>\n<hr>\n<p><strong><em>Please see these notebooks for more details.</em></strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\" target=\"_blank\">https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation</a></li>\n<li><a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation/notebook\" target=\"_blank\">https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation/notebook</a></li>\n<li><a href=\"https://github.com/SamusRam/HPA-Cell-Segmentation\" target=\"_blank\">https://github.com/SamusRam/HPA-Cell-Segmentation</a></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1251219,
          "author_name": "glopezzz",
          "author_url": "",
          "post_date": "2021-03-24T15:32:35.390000",
          "content": "<p>That's awesome!! Thanks for the information, I'll be working with that from now on :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1224956": "I wonder if notebook running time includes scoring time? I 'm sure my notebook running time is about 7 hours, but it‘s time out.",
    "1225248": "@daishu - Is your notebook running time 7 hours on the public test data?\n\nI have submitted many times now and have changed the way I submit many times, and I have never noticed the scoring time to be included. That being said I'm not an expert, perhaps a Kaggle admin might better answer. I did hope to share some observations with you if that's alright.\n\n---\n\n**My observations are the following:**\n\n- The private test set is 2-3 times larger than the public test set\n- If your notebook takes 2 hours to run on the public test set, it will take 4-6 hours to run when you submit your **`submission.csv`**\n- To see your LB score (and get 0 on the private LB), you only need to predict on the public test images and leave the default prediction string for the private test set (this will save you 2/3rds of the submission time). Although, obviously you will need to be submitting predictions on both public and private data when it comes time for your final submissions.\n- You only need to generate a submission.csv file on a few predictions when submitting (this will save you >90% of the time it takes to run your notebook on the public test set). i.e. **`IS_DEMO=len(ss_df)==559`** and then use the flag **`IS_DEMO`** to control whether you are predicting on the entire public test set or only a few images. This works because during the Save&Commit part, your **`ss_df`** will have a length of 559... however, when the notebook is run on the backend, the length of **`ss_df`** will probably be closer to 1250 or something.\n\n---\n\n*My public notebooks and those by other talented Kagglers show these observations in action... **I hope this helps!***"
  }
}