{
  "id": 226145,
  "title": "A reminder to not use fast sub submissions",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226145",
  "author_name": "",
  "post_date": "2021-03-15T11:21:18.867521800Z",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<p>A kind reminder to not use fast sub submissions as your final score. </p>\n<p>For ensemble submissions, what I would do: if there are 3 models, one can run all 3 models in the same inference notebook, then I would generate each of their predictions and then apply my ensemble Techniques. Then submit. Hope I’m right!</p>",
  "messages": [
    {
      "id": "1238876",
      "postDate": "03/15/2021 11:21:18",
      "content": "<p>A kind reminder to not use fast sub submissions as your final score. </p>\n<p>For ensemble submissions, what I would do: if there are 3 models, one can run all 3 models in the same inference notebook, then I would generate each of their predictions and then apply my ensemble Techniques. Then submit. Hope I’m right!</p>",
      "rawMarkdown": "A kind reminder to not use fast sub submissions as your final score. \n\nFor ensemble submissions, what I would do: if there are 3 models, one can run all 3 models in the same inference notebook, then I would generate each of their predictions and then apply my ensemble Techniques. Then submit. Hope I’m right!",
      "votes": null
    },
    {
      "id": "1238961",
      "postDate": "03/15/2021 12:05:22",
      "content": "<p>Hello! Why do u say that ?</p>",
      "rawMarkdown": "Hello! Why do u say that ?",
      "votes": null
    },
    {
      "id": "1238964",
      "postDate": "03/15/2021 12:06:05",
      "content": "<p>Yes  you are right to do all models in one inference notebook then submit. </p>\n<p>This was from another competition but a really helpful idea for doing inference and ensembles just for reference:</p>\n<p><a href=\"https://www.kaggle.com/underwearfitting/make-final-submission-the-efficient-way\" target=\"_blank\">https://www.kaggle.com/underwearfitting/make-final-submission-the-efficient-way</a></p>",
      "rawMarkdown": "Yes  you are right to do all models in one inference notebook then submit. \n\nThis was from another competition but a really helpful idea for doing inference and ensembles just for reference:\n\nhttps://www.kaggle.com/underwearfitting/make-final-submission-the-efficient-way",
      "votes": null
    },
    {
      "id": "1239213",
      "postDate": "03/15/2021 14:40:02",
      "content": "<p>Yes.</p>\n<pre><code>print('model1.predict&gt;&gt;&gt;&gt;')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nprint('model2.predict&gt;&gt;&gt;&gt;')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nprint('model3.predict&gt;&gt;&gt;&gt;')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nsubmission[label_cols] = (p1 + p2 + p3 ) / 3.0\n</code></pre>",
      "rawMarkdown": "Yes.\n\n```\nprint('model1.predict>>>>')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nprint('model2.predict>>>>')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nprint('model3.predict>>>>')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nsubmission[label_cols] = (p1 + p2 + p3 ) / 3.0\n```",
      "votes": null
    },
    {
      "id": "1239216",
      "postDate": "03/15/2021 14:41:20",
      "content": "<p>Hello,what is going wrong if i fast submit?</p>",
      "rawMarkdown": "Hello,what is going wrong if i fast submit?",
      "votes": null
    },
    {
      "id": "1239226",
      "postDate": "03/15/2021 14:48:14",
      "content": "<p>your fast submit only knows the public lb, so it will score 0.5 on private lb</p>",
      "rawMarkdown": "your fast submit only knows the public lb, so it will score 0.5 on private lb",
      "votes": null
    },
    {
      "id": "1239321",
      "postDate": "03/15/2021 16:08:56",
      "content": "<p>Yes, that works. The main limitation is time. You can of course avoid re-predicting the public LB dataset and just load those predictions, which should be enough to squeeze at least one extra model into your ensemble (or perhaps even 2 depending on how fast your inference runs). One of the bottlenecks in inference for me is just reading the images, because there's just no option to do pre-processing for the hidden private LB data…</p>",
      "rawMarkdown": "Yes, that works. The main limitation is time. You can of course avoid re-predicting the public LB dataset and just load those predictions, which should be enough to squeeze at least one extra model into your ensemble (or perhaps even 2 depending on how fast your inference runs). One of the bottlenecks in inference for me is just reading the images, because there's just no option to do pre-processing for the hidden private LB data...",
      "votes": null
    },
    {
      "id": "1239325",
      "postDate": "03/15/2021 16:14:27",
      "content": "<p>Saves hours of GPU time and let's you submit within a few min of modifying your submission program. </p>\n<p>In this competition, I guess the check is <code>if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) &lt;= 3582):</code> (in that case, don't re-run everything and just save the sample submission as <code>submission.csv</code>), while in the <code>else</code> case you run everything. <strong>WARNING:</strong> This is at the cost that mistakes in your program will only happen during the submission run (tough to debug), so I added an extra option: <code>if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) &lt;= 3582) &amp; (FORCE_RUN==False):</code> that way you can set <code>FORCE_RUN=True</code> for testing purposes without havign to change anything else.</p>",
      "rawMarkdown": "Saves hours of GPU time and let's you submit within a few min of modifying your submission program. \n\nIn this competition, I guess the check is `if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) <= 3582):` (in that case, don't re-run everything and just save the sample submission as `submission.csv`), while in the `else` case you run everything. **WARNING:** This is at the cost that mistakes in your program will only happen during the submission run (tough to debug), so I added an extra option: `if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) <= 3582) & (FORCE_RUN==False):` that way you can set `FORCE_RUN=True` for testing purposes without havign to change anything else.",
      "votes": null
    },
    {
      "id": "1239799",
      "postDate": "03/16/2021 03:00:44",
      "content": "<p>You can just check the len of sample_submission it will be different when the submit reruns the notebook and just keep a small number to be sure the code runs and creates submission.csv OK.<br>\n<code>test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')</code><br>\n<code>if len(test) == 3582:</code><br>\n<code>test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv', nrows=10)</code></p>",
      "rawMarkdown": "You can just check the len of sample_submission it will be different when the submit reruns the notebook and just keep a small number to be sure the code runs and creates submission.csv OK.\n`test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')`\n`if len(test) == 3582:  `\n`     test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv', nrows=10)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1238961,
      "author_name": "zekunn",
      "author_url": "",
      "post_date": "03/15/2021 12:05:22",
      "content": "<p>Hello! Why do u say that ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1238964,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/15/2021 12:06:05",
      "content": "<p>Yes  you are right to do all models in one inference notebook then submit. </p>\n<p>This was from another competition but a really helpful idea for doing inference and ensembles just for reference:</p>\n<p><a href=\"https://www.kaggle.com/underwearfitting/make-final-submission-the-efficient-way\" target=\"_blank\">https://www.kaggle.com/underwearfitting/make-final-submission-the-efficient-way</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1239325,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/15/2021 16:14:27",
          "content": "<p>Saves hours of GPU time and let's you submit within a few min of modifying your submission program. </p>\n<p>In this competition, I guess the check is <code>if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) &lt;= 3582):</code> (in that case, don't re-run everything and just save the sample submission as <code>submission.csv</code>), while in the <code>else</code> case you run everything. <strong>WARNING:</strong> This is at the cost that mistakes in your program will only happen during the submission run (tough to debug), so I added an extra option: <code>if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) &lt;= 3582) &amp; (FORCE_RUN==False):</code> that way you can set <code>FORCE_RUN=True</code> for testing purposes without havign to change anything else.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239799,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "03/16/2021 03:00:44",
          "content": "<p>You can just check the len of sample_submission it will be different when the submit reruns the notebook and just keep a small number to be sure the code runs and creates submission.csv OK.<br>\n<code>test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')</code><br>\n<code>if len(test) == 3582:</code><br>\n<code>test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv', nrows=10)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239213,
      "author_name": "faisalalsrheed",
      "author_url": "",
      "post_date": "03/15/2021 14:40:02",
      "content": "<p>Yes.</p>\n<pre><code>print('model1.predict&gt;&gt;&gt;&gt;')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nprint('model2.predict&gt;&gt;&gt;&gt;')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nprint('model3.predict&gt;&gt;&gt;&gt;')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nsubmission[label_cols] = (p1 + p2 + p3 ) / 3.0\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1239216,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "03/15/2021 14:41:20",
          "content": "<p>Hello,what is going wrong if i fast submit?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239226,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/15/2021 14:48:14",
          "content": "<p>your fast submit only knows the public lb, so it will score 0.5 on private lb</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239321,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/15/2021 16:08:56",
      "content": "<p>Yes, that works. The main limitation is time. You can of course avoid re-predicting the public LB dataset and just load those predictions, which should be enough to squeeze at least one extra model into your ensemble (or perhaps even 2 depending on how fast your inference runs). One of the bottlenecks in inference for me is just reading the images, because there's just no option to do pre-processing for the hidden private LB data…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1238876": "A kind reminder to not use fast sub submissions as your final score. \n\nFor ensemble submissions, what I would do: if there are 3 models, one can run all 3 models in the same inference notebook, then I would generate each of their predictions and then apply my ensemble Techniques. Then submit. Hope I’m right!",
    "1238961": "Hello! Why do u say that ?",
    "1238964": "Yes  you are right to do all models in one inference notebook then submit. \n\nThis was from another competition but a really helpful idea for doing inference and ensembles just for reference:\n\nhttps://www.kaggle.com/underwearfitting/make-final-submission-the-efficient-way",
    "1239213": "Yes.\n\n```\nprint('model1.predict>>>>')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nprint('model2.predict>>>>')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nprint('model3.predict>>>>')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nsubmission[label_cols] = (p1 + p2 + p3 ) / 3.0\n```",
    "1239216": "Hello,what is going wrong if i fast submit?",
    "1239226": "your fast submit only knows the public lb, so it will score 0.5 on private lb",
    "1239321": "Yes, that works. The main limitation is time. You can of course avoid re-predicting the public LB dataset and just load those predictions, which should be enough to squeeze at least one extra model into your ensemble (or perhaps even 2 depending on how fast your inference runs). One of the bottlenecks in inference for me is just reading the images, because there's just no option to do pre-processing for the hidden private LB data...",
    "1239325": "Saves hours of GPU time and let's you submit within a few min of modifying your submission program. \n\nIn this competition, I guess the check is `if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) <= 3582):` (in that case, don't re-run everything and just save the sample submission as `submission.csv`), while in the `else` case you run everything. **WARNING:** This is at the cost that mistakes in your program will only happen during the submission run (tough to debug), so I added an extra option: `if (len(os.listdir('../input/ranzcr-clip-catheter-line-classification/test')) <= 3582) & (FORCE_RUN==False):` that way you can set `FORCE_RUN=True` for testing purposes without havign to change anything else.",
    "1239799": "You can just check the len of sample_submission it will be different when the submit reruns the notebook and just keep a small number to be sure the code runs and creates submission.csv OK.\n`test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')`\n`if len(test) == 3582:  `\n`     test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv', nrows=10)`"
  },
  "source": "meta"
}