{
  "id": 236369,
  "title": "[QUESTION] Submission time constraints",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/236369",
  "author_name": "",
  "post_date": "2021-05-04T01:15:46.662459600Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>How do I know for sure that my code will be able to run fast enough when the competition ends? When we submit it to the public leaderboard does it compute all the 5k test images but only grade based on 18%?</p>",
  "messages": [
    {
      "id": "1292457",
      "postDate": "05/04/2021 01:15:46",
      "content": "<p>How do I know for sure that my code will be able to run fast enough when the competition ends? When we submit it to the public leaderboard does it compute all the 5k test images but only grade based on 18%?</p>",
      "rawMarkdown": "How do I know for sure that my code will be able to run fast enough when the competition ends? When we submit it to the public leaderboard does it compute all the 5k test images but only grade based on 18%?",
      "votes": null
    },
    {
      "id": "1292957",
      "postDate": "05/04/2021 12:52:49",
      "content": "<p>Your submission is run on the full private test dataset right after you submit. There's no additional runs after end of a competition. If your submission was scored successfully on public LB, it will be scored successfully on private LB.</p>",
      "rawMarkdown": "Your submission is run on the full private test dataset right after you submit. There's no additional runs after end of a competition. If your submission was scored successfully on public LB, it will be scored successfully on private LB.",
      "votes": null
    },
    {
      "id": "1293003",
      "postDate": "05/04/2021 13:24:21",
      "content": "<p>Much appreciated <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a>  this is my first competition so I am super green! =)</p>",
      "rawMarkdown": "Much appreciated @atamazian  this is my first competition so I am super green! =)",
      "votes": null
    },
    {
      "id": "1293037",
      "postDate": "05/04/2021 13:50:19",
      "content": "<p>Is that true? I'm also really concerned about the time limit, since I plan to ensemble multiple trained models.</p>",
      "rawMarkdown": "Is that true? I'm also really concerned about the time limit, since I plan to ensemble multiple trained models.",
      "votes": null
    },
    {
      "id": "1293059",
      "postDate": "05/04/2021 14:08:22",
      "content": "<p>I think it is. First your kernel is run on all private test dataset, and then resulting csv is splitted according to public/private ratio. Both parts are ranked - first for the public LB, and second for the private LB. Private LB is hidden until a competition end.</p>",
      "rawMarkdown": "I think it is. First your kernel is run on all private test dataset, and then resulting csv is splitted according to public/private ratio. Both parts are ranked - first for the public LB, and second for the private LB. Private LB is hidden until a competition end.",
      "votes": null
    },
    {
      "id": "1293060",
      "postDate": "05/04/2021 14:09:26",
      "content": "<p>As for multiple models, I also use many models. I don't think it's possible to get good result here with only only one model, even if it's 5-folded.</p>",
      "rawMarkdown": "As for multiple models, I also use many models. I don't think it's possible to get good result here with only only one model, even if it's 5-folded.",
      "votes": null
    },
    {
      "id": "1293088",
      "postDate": "05/04/2021 14:45:06",
      "content": "<p>Thanks for your patient reply, I now understand the scoring mechanism.</p>",
      "rawMarkdown": "Thanks for your patient reply, I now understand the scoring mechanism.",
      "votes": null
    },
    {
      "id": "1293098",
      "postDate": "05/04/2021 14:52:47",
      "content": "<p>Here's the official page regarding that subject <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a></p>",
      "rawMarkdown": "Here's the official page regarding that subject https://www.kaggle.com/code-competition-debugging",
      "votes": null
    },
    {
      "id": "1295523",
      "postDate": "05/06/2021 14:16:46",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> I had the exact same question!</p>",
      "rawMarkdown": "Thank you @atamazian I had the exact same question!",
      "votes": null
    },
    {
      "id": "1298134",
      "postDate": "05/08/2021 15:19:51",
      "content": "<p>I can second that. </p>\n<p>In the Cassava Lead Disease I had a model that took between 8-9 hours to submit. The competition specified that the notebook time had to be &lt;9 hours.</p>\n<p>I was worried as the public leaderboard was being tested on 69% of the data, and there was no way that my notebook would stay &lt;9 hours after being ran on the hidden test set.</p>\n<p>Nevertheless, I selected this model to be used for final scoring to test the limits and it all worked fine.</p>",
      "rawMarkdown": "I can second that. \n\nIn the Cassava Lead Disease I had a model that took between 8-9 hours to submit. The competition specified that the notebook time had to be <9 hours.\n\n I was worried as the public leaderboard was being tested on 69% of the data, and there was no way that my notebook would stay <9 hours after being ran on the hidden test set.\n\nNevertheless, I selected this model to be used for final scoring to test the limits and it all worked fine.",
      "votes": null
    },
    {
      "id": "1306806",
      "postDate": "05/14/2021 05:24:02",
      "content": "<p>Now I have some doubts about the scoring. Actually, the scoring issue just happended in another competition: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/238075</a>, where submissions with correct public LB scores failed in the private LB, and got zero. </p>",
      "rawMarkdown": "Now I have some doubts about the scoring. Actually, the scoring issue just happended in another competition: [https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/238075](url), where submissions with correct public LB scores failed in the private LB, and got zero.",
      "votes": null
    },
    {
      "id": "1306817",
      "postDate": "05/14/2021 05:30:29",
      "content": "<p>Moreover, I don't see any evidence of running the whole test set while submitting in this meterial you provided: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a>. So is there any official statement on how the scoring system works? Running on the whole test set or part of them? </p>",
      "rawMarkdown": "Moreover, I don't see any evidence of running the whole test set while submitting in this meterial you provided: [https://www.kaggle.com/code-competition-debugging](url). So is there any official statement on how the scoring system works? Running on the whole test set or part of them?",
      "votes": null
    },
    {
      "id": "1307740",
      "postDate": "05/14/2021 16:29:34",
      "content": "<p><a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> the default for code competitions is for every submission to be scored against both the public and private test sets. The only exceptions I can think of off hand would be competitions where the data has to be updated at some point after launch. For example, the <a href=\"https://www.kaggle.com/c/jane-street-market-prediction/overview/timeline\" target=\"_blank\">Jane Street Market Prediction</a> was set up with a private test set to be gathered after all notebooks were submitted. However, this is an unusual and clearly labeled arrangement.</p>",
      "rawMarkdown": "ytepzhi the default for code competitions is for every submission to be scored against both the public and private test sets. The only exceptions I can think of off hand would be competitions where the data has to be updated at some point after launch. For example, the [Jane Street Market Prediction](https://www.kaggle.com/c/jane-street-market-prediction/overview/timeline) was set up with a private test set to be gathered after all notebooks were submitted. However, this is an unusual and clearly labeled arrangement.",
      "votes": null
    },
    {
      "id": "1307751",
      "postDate": "05/14/2021 16:37:04",
      "content": "<p>Thanks for your generous reply.</p>",
      "rawMarkdown": "Thanks for your generous reply.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1292957,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "05/04/2021 12:52:49",
      "content": "<p>Your submission is run on the full private test dataset right after you submit. There's no additional runs after end of a competition. If your submission was scored successfully on public LB, it will be scored successfully on private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1293037,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/04/2021 13:50:19",
          "content": "<p>Is that true? I'm also really concerned about the time limit, since I plan to ensemble multiple trained models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293059,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/04/2021 14:08:22",
          "content": "<p>I think it is. First your kernel is run on all private test dataset, and then resulting csv is splitted according to public/private ratio. Both parts are ranked - first for the public LB, and second for the private LB. Private LB is hidden until a competition end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293060,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/04/2021 14:09:26",
          "content": "<p>As for multiple models, I also use many models. I don't think it's possible to get good result here with only only one model, even if it's 5-folded.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293088,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/04/2021 14:45:06",
          "content": "<p>Thanks for your patient reply, I now understand the scoring mechanism.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293098,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/04/2021 14:52:47",
          "content": "<p>Here's the official page regarding that subject <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1298134,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "05/08/2021 15:19:51",
          "content": "<p>I can second that. </p>\n<p>In the Cassava Lead Disease I had a model that took between 8-9 hours to submit. The competition specified that the notebook time had to be &lt;9 hours.</p>\n<p>I was worried as the public leaderboard was being tested on 69% of the data, and there was no way that my notebook would stay &lt;9 hours after being ran on the hidden test set.</p>\n<p>Nevertheless, I selected this model to be used for final scoring to test the limits and it all worked fine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306806,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/14/2021 05:24:02",
          "content": "<p>Now I have some doubts about the scoring. Actually, the scoring issue just happended in another competition: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/238075</a>, where submissions with correct public LB scores failed in the private LB, and got zero. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306817,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/14/2021 05:30:29",
          "content": "<p>Moreover, I don't see any evidence of running the whole test set while submitting in this meterial you provided: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a>. So is there any official statement on how the scoring system works? Running on the whole test set or part of them? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1307740,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "05/14/2021 16:29:34",
          "content": "<p><a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> the default for code competitions is for every submission to be scored against both the public and private test sets. The only exceptions I can think of off hand would be competitions where the data has to be updated at some point after launch. For example, the <a href=\"https://www.kaggle.com/c/jane-street-market-prediction/overview/timeline\" target=\"_blank\">Jane Street Market Prediction</a> was set up with a private test set to be gathered after all notebooks were submitted. However, this is an unusual and clearly labeled arrangement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1307751,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/14/2021 16:37:04",
          "content": "<p>Thanks for your generous reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1293003,
      "author_name": "coldfir3",
      "author_url": "",
      "post_date": "05/04/2021 13:24:21",
      "content": "<p>Much appreciated <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a>  this is my first competition so I am super green! =)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1295523,
      "author_name": "mylonsong",
      "author_url": "",
      "post_date": "05/06/2021 14:16:46",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> I had the exact same question!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1292457": "How do I know for sure that my code will be able to run fast enough when the competition ends? When we submit it to the public leaderboard does it compute all the 5k test images but only grade based on 18%?",
    "1292957": "Your submission is run on the full private test dataset right after you submit. There's no additional runs after end of a competition. If your submission was scored successfully on public LB, it will be scored successfully on private LB.",
    "1293003": "Much appreciated @atamazian  this is my first competition so I am super green! =)",
    "1293037": "Is that true? I'm also really concerned about the time limit, since I plan to ensemble multiple trained models.",
    "1293059": "I think it is. First your kernel is run on all private test dataset, and then resulting csv is splitted according to public/private ratio. Both parts are ranked - first for the public LB, and second for the private LB. Private LB is hidden until a competition end.",
    "1293060": "As for multiple models, I also use many models. I don't think it's possible to get good result here with only only one model, even if it's 5-folded.",
    "1293088": "Thanks for your patient reply, I now understand the scoring mechanism.",
    "1293098": "Here's the official page regarding that subject https://www.kaggle.com/code-competition-debugging",
    "1295523": "Thank you @atamazian I had the exact same question!",
    "1298134": "I can second that. \n\nIn the Cassava Lead Disease I had a model that took between 8-9 hours to submit. The competition specified that the notebook time had to be <9 hours.\n\n I was worried as the public leaderboard was being tested on 69% of the data, and there was no way that my notebook would stay <9 hours after being ran on the hidden test set.\n\nNevertheless, I selected this model to be used for final scoring to test the limits and it all worked fine.",
    "1306806": "Now I have some doubts about the scoring. Actually, the scoring issue just happended in another competition: [https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/238075](url), where submissions with correct public LB scores failed in the private LB, and got zero.",
    "1306817": "Moreover, I don't see any evidence of running the whole test set while submitting in this meterial you provided: [https://www.kaggle.com/code-competition-debugging](url). So is there any official statement on how the scoring system works? Running on the whole test set or part of them?",
    "1307740": "ytepzhi the default for code competitions is for every submission to be scored against both the public and private test sets. The only exceptions I can think of off hand would be competitions where the data has to be updated at some point after launch. For example, the [Jane Street Market Prediction](https://www.kaggle.com/c/jane-street-market-prediction/overview/timeline) was set up with a private test set to be gathered after all notebooks were submitted. However, this is an unusual and clearly labeled arrangement.",
    "1307751": "Thanks for your generous reply."
  },
  "source": "meta"
}