{
  "id": 87544,
  "title": "Is it allowed to train locally at the early stage",
  "url": "/competitions/imet-2019-fgvc6/discussion/87544",
  "author_name": "",
  "post_date": "2019-04-01T13:17:43.171587900Z",
  "votes": 11,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Training in Kaggle kernels can quite time-consuming because it would take several hours to run the notebook and do some modifications and then another 9 hours to commit the kernel. I'm wondering if it's allowed to train locally and then just submit the results. Of course, I'll switch to Kaggle kernel after I get a satisfying version.</p>",
  "messages": [
    {
      "id": "505026",
      "postDate": "04/01/2019 13:17:43",
      "content": "<p>Training in Kaggle kernels can quite time-consuming because it would take several hours to run the notebook and do some modifications and then another 9 hours to commit the kernel. I'm wondering if it's allowed to train locally and then just submit the results. Of course, I'll switch to Kaggle kernel after I get a satisfying version.</p>",
      "rawMarkdown": "Training in Kaggle kernels can quite time-consuming because it would take several hours to run the notebook and do some modifications and then another 9 hours to commit the kernel. I'm wondering if it's allowed to train locally and then just submit the results. Of course, I'll switch to Kaggle kernel after I get a satisfying version.",
      "votes": null
    },
    {
      "id": "505123",
      "postDate": "04/01/2019 15:34:31",
      "content": "<p>You can't submit an uploaded file.\nThere is no upload button in the submission tab</p>",
      "rawMarkdown": "You can't submit an uploaded file.\nThere is no upload button in the submission tab",
      "votes": null
    },
    {
      "id": "505145",
      "postDate": "04/01/2019 16:02:14",
      "content": "<p>Really? I haven't tried that yet. So it doesn't seem to work.</p>",
      "rawMarkdown": "Really? I haven't tried that yet. So it doesn't seem to work.",
      "votes": null
    },
    {
      "id": "505155",
      "postDate": "04/01/2019 16:14:52",
      "content": "<p>But I think you can upload the submission file as a dataset and then you import it to the notebook and export it with to_csv method. Finally submit it.</p>",
      "rawMarkdown": "But I think you can upload the submission file as a dataset and then you import it to the notebook and export it with to_csv method. Finally submit it.",
      "votes": null
    },
    {
      "id": "505177",
      "postDate": "04/01/2019 16:46:26",
      "content": "<p>I would give it a try and maybe I will share my approach here if I succeed.</p>",
      "rawMarkdown": "I would give it a try and maybe I will share my approach here if I succeed.",
      "votes": null
    },
    {
      "id": "505200",
      "postDate": "04/01/2019 17:31:33",
      "content": "<p>I had the same question than <a href=\"/syoya1997\">@syoya1997</a> </p>\n\n<p>With the same intention of switching to a Kaggle kernel once I find the right postprocessing, etc.</p>\n\n<p><a href=\"/rinnqd\">@rinnqd</a> </p>\n\n<p>It's true that you can't submit an uploaded file.</p>\n\n<p>But you can \"upload it\" inside a python script. For example, you can use something like the following code:</p>\n\n<p>```\nmy_submission='''id,attribute_ids\n1023b2cd, 0 1 2\n1023bbaa, 0 1 2\n                    ... here would be the complete submission.csv\n7163763b, 100 101\n'''</p>\n\n<p>with open('submission.csv', 'w') as f:\n    f.write(my_submission)</p>\n\n<p>```</p>\n\n<p><a href=\"/codingcliff\">@codingcliff</a> \nSo, is it allowed to \"upload\" CSVs during the model preparation? </p>",
      "rawMarkdown": "I had the same question than @syoya1997 \n\nWith the same intention of switching to a Kaggle kernel once I find the right postprocessing, etc.\n\n@rinnqd \n\nIt's true that you can't submit an uploaded file.\n\nBut you can \"upload it\" inside a python script. For example, you can use something like the following code:\n\n\n```\nmy_submission='''id,attribute_ids\n1023b2cd, 0 1 2\n1023bbaa, 0 1 2\n\t\t\t\t\t... here would be the complete submission.csv\n7163763b, 100 101\n'''\n\nwith open('submission.csv', 'w') as f:\n    f.write(my_submission)\n    \n```\n\n@codingcliff \nSo, is it allowed to \"upload\" CSVs during the model preparation?",
      "votes": null
    },
    {
      "id": "505240",
      "postDate": "04/01/2019 19:00:54",
      "content": "<p>Or you can rely on your validation set to see the improvements while training locally and then run the code on a kernel only once in a while to check if nothing weird is going on.</p>",
      "rawMarkdown": "Or you can rely on your validation set to see the improvements while training locally and then run the code on a kernel only once in a while to check if nothing weird is going on.",
      "votes": null
    },
    {
      "id": "505306",
      "postDate": "04/01/2019 21:01:56",
      "content": "<p>Hi syoya,</p>\n\n<p>Does training a model locally and submit a checkpoint as dataset and init from it works for you?</p>",
      "rawMarkdown": "Hi syoya,\n\nDoes training a model locally and submit a checkpoint as dataset and init from it works for you?",
      "votes": null
    },
    {
      "id": "505724",
      "postDate": "04/02/2019 13:16:04",
      "content": "<p>personal opinion: this seems like a grey area with potential for exploitation. To develop and debug locally I would just create my own train-val split, and once it is working run in kernel. But putting that submission as dataset and submitting seems to be against the intent of the rule. There would need to be a way to automatically check for this type of submission as it seems very much against the rules of for a final submission. </p>",
      "rawMarkdown": "personal opinion: this seems like a grey area with potential for exploitation. To develop and debug locally I would just create my own train-val split, and once it is working run in kernel. But putting that submission as dataset and submitting seems to be against the intent of the rule. There would need to be a way to automatically check for this type of submission as it seems very much against the rules of for a final submission.",
      "votes": null
    },
    {
      "id": "505731",
      "postDate": "04/02/2019 13:26:51",
      "content": "<p>Uploading a static submission to Kernels (as opposed to source code or a model capable of predicting) will cause your code to fail when the Kernel is rerun on the held-out test set.</p>",
      "rawMarkdown": "Uploading a static submission to Kernels (as opposed to source code or a model capable of predicting) will cause your code to fail when the Kernel is rerun on the held-out test set.",
      "votes": null
    },
    {
      "id": "505894",
      "postDate": "04/02/2019 17:23:06",
      "content": "<p>Is this allowed? And do I have to make the dataset public?</p>",
      "rawMarkdown": "Is this allowed? And do I have to make the dataset public?",
      "votes": null
    },
    {
      "id": "505976",
      "postDate": "04/02/2019 20:12:54",
      "content": "<p>My guess is yes to having to make the dataset public to use it in kernels. <a href=\"/wcukierski\">@wcukierski</a>  Correct me if I'm wrong. Thanks!</p>",
      "rawMarkdown": "My guess is yes to having to make the dataset public to use it in kernels. @wcukierski  Correct me if I'm wrong. Thanks!",
      "votes": null
    },
    {
      "id": "505981",
      "postDate": "04/02/2019 20:35:34",
      "content": "<p>I forget that there is an unaccessable test set for phase 2.  Only thing I see is whether this is a behavior that should increase in frequency in the future or not.  On this point I do not have an opinion. </p>",
      "rawMarkdown": "I forget that there is an unaccessable test set for phase 2.  Only thing I see is whether this is a behavior that should increase in frequency in the future or not.  On this point I do not have an opinion.",
      "votes": null
    },
    {
      "id": "505990",
      "postDate": "04/02/2019 20:58:24",
      "content": "<blockquote>\n  <p>model capable of predicting</p>\n</blockquote>\n\n<p><a href=\"/wcukierski\">@wcukierski</a> could you please clarify, must the model be trained in the kernel, or we can upload model trained on competition train data? That's a huge difference.</p>",
      "rawMarkdown": "&gt; model capable of predicting\n\n@wcukierski could you please clarify, must the model be trained in the kernel, or we can upload model trained on competition train data? That's a huge difference.",
      "votes": null
    },
    {
      "id": "506122",
      "postDate": "04/03/2019 02:22:09",
      "content": "<p>You are permitted to train outside of Kernels and upload your model.</p>",
      "rawMarkdown": "You are permitted to train outside of Kernels and upload your model.",
      "votes": null
    },
    {
      "id": "506125",
      "postDate": "04/03/2019 02:37:59",
      "content": "<p>Is it necessary  to make model public？ Thanks </p>",
      "rawMarkdown": "Is it necessary  to make model public？ Thanks",
      "votes": null
    },
    {
      "id": "506163",
      "postDate": "04/03/2019 04:17:12",
      "content": "<p><a href=\"/wcukierski\">@wcukierski</a> thanks for the comment. This busted the myth I am having that training needs to be done from scratch in kernel environment only. <br>Because I was thinking that uploaded model could have used images/data from not mentioned sources. <br>For example in previous petfinder.my competition it is mentioned that people shouldn't use the images from petfinder.my website. if inclusion of offline trained model is allowed directly then we are opening a window for a possible misuse. <br> I think may be we should allow to include model from other kernel output at least we can track the training data.<br> I am sure kaggle will verify and will ask winners to reproduce the output. but on large scale even it will be difficult for kaggle admins to verify all medal winners code (bronze). hence for at least kernel competitions instead of offline model upload make a provision that their solution can span across kernels would be a better idea I believe. Please share your opinion. Thank you.</p>",
      "rawMarkdown": "wcukierski thanks for the comment. This busted the myth I am having that training needs to be done from scratch in kernel environment only. <br>Because I was thinking that uploaded model could have used images/data from not mentioned sources. <br>For example in previous petfinder.my competition it is mentioned that people shouldn't use the images from petfinder.my website. if inclusion of offline trained model is allowed directly then we are opening a window for a possible misuse. <br> I think may be we should allow to include model from other kernel output at least we can track the training data.<br> I am sure kaggle will verify and will ask winners to reproduce the output. but on large scale even it will be difficult for kaggle admins to verify all medal winners code (bronze). hence for at least kernel competitions instead of offline model upload make a provision that their solution can span across kernels would be a better idea I believe. Please share your opinion. Thank you.",
      "votes": null
    },
    {
      "id": "506224",
      "postDate": "04/03/2019 06:36:45",
      "content": "<blockquote>\n  <p>You are permitted to train outside of Kernels and upload your model.</p>\n</blockquote>\n\n<p>Oh thanks for clarification. Looks like this is the first kernel competition on kaggle where you are allowed to upload the trained model?\nSo we'll need to train a gazillion of models and stack, sad :(</p>",
      "rawMarkdown": "&gt; You are permitted to train outside of Kernels and upload your model.\n\nOh thanks for clarification. Looks like this is the first kernel competition on kaggle where you are allowed to upload the trained model?\nSo we'll need to train a gazillion of models and stack, sad :(",
      "votes": null
    },
    {
      "id": "506235",
      "postDate": "04/03/2019 06:56:53",
      "content": "<p>Please, clarify the following: is it allowed to train at home, upload the checkpoints privately, and use these checkpoints for stage 2?</p>\n\n<p>In this way, probably the winner will be an ensemble of several checkpoints, and the winner kernel won't perform any training, just inference.</p>",
      "rawMarkdown": "Please, clarify the following: is it allowed to train at home, upload the checkpoints privately, and use these checkpoints for stage 2?\n\nIn this way, probably the winner will be an ensemble of several checkpoints, and the winner kernel won't perform any training, just inference.",
      "votes": null
    },
    {
      "id": "506237",
      "postDate": "04/03/2019 07:01:15",
      "content": "<p>IMO, if this is a kernel competition as stated, it should not be allowed to train on competition data locally or outside of whitelisted data. Otherwise it should not be named kernel competition. <a href=\"https://www.kaggle.com/lopuhin\">@Konstantin</a> Indeed, this does make you sad :(</p>",
      "rawMarkdown": "IMO, if this is a kernel competition as stated, it should not be allowed to train on competition data locally or outside of whitelisted data. Otherwise it should not be named kernel competition. [@Konstantin](https://www.kaggle.com/lopuhin) Indeed, this does make you sad :(",
      "votes": null
    },
    {
      "id": "506273",
      "postDate": "04/03/2019 08:36:06",
      "content": "<p>I guess it would be so boring if somebody trains a lot of models locally and uploads them to stack. It loses the meaning of kernel competition :(</p>",
      "rawMarkdown": "I guess it would be so boring if somebody trains a lot of models locally and uploads them to stack. It loses the meaning of kernel competition :(",
      "votes": null
    },
    {
      "id": "506353",
      "postDate": "04/03/2019 10:46:30",
      "content": "<p>Sorry for a late reply. I'm quite busy yesterday and downloading the dataset takes me quite a while. And yes, submit a checkpoint as the dataset and init works for me indeed. However, I'm afraid that even making the dataset public cannot be a good solution. We can name the checkpoint as \"densenet120.h5\" or other seemingly public pre-trained model names and this cannot be well verified.</p>",
      "rawMarkdown": "Sorry for a late reply. I'm quite busy yesterday and downloading the dataset takes me quite a while. And yes, submit a checkpoint as the dataset and init works for me indeed. However, I'm afraid that even making the dataset public cannot be a good solution. We can name the checkpoint as \"densenet120.h5\" or other seemingly public pre-trained model names and this cannot be well verified.",
      "votes": null
    },
    {
      "id": "506356",
      "postDate": "04/03/2019 10:50:20",
      "content": "<p>Does that mean actually kernel competitions have no difference with other competitions? Then I think \"kernel\" competitions are meaningless.</p>",
      "rawMarkdown": "Does that mean actually kernel competitions have no difference with other competitions? Then I think \"kernel\" competitions are meaningless.",
      "votes": null
    },
    {
      "id": "506415",
      "postDate": "04/03/2019 12:47:40",
      "content": "<p>Our own <a href=\"/inversion\">@inversion</a> wrote a nice explanation of why we'd run Kernels-only competitions in this type of format: <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/87719\">https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/87719</a> (that applies to a different competition, but the spirit is the same here).</p>",
      "rawMarkdown": "Our own @inversion wrote a nice explanation of why we'd run Kernels-only competitions in this type of format: https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/87719 (that applies to a different competition, but the spirit is the same here).",
      "votes": null
    },
    {
      "id": "506424",
      "postDate": "04/03/2019 13:15:57",
      "content": "<p>Wow. So kernel competitions are just meant to limit inference complexity? This surprises me.</p>",
      "rawMarkdown": "Wow. So kernel competitions are just meant to limit inference complexity? This surprises me.",
      "votes": null
    },
    {
      "id": "506430",
      "postDate": "04/03/2019 13:26:02",
      "content": "<p>If that is the case here than you should make it clear in <strong>Overview</strong> under <strong>Kernel Requirements</strong> that only  the inference step must be performed in kernels (9 hours). The current statement is misleading. I wonder if anyone figured it out correctly excluding people working for kaggle</p>",
      "rawMarkdown": "If that is the case here than you should make it clear in **Overview** under **Kernel Requirements** that only  the inference step must be performed in kernels (9 hours). The current statement is misleading. I wonder if anyone figured it out correctly excluding people working for kaggle",
      "votes": null
    },
    {
      "id": "506434",
      "postDate": "04/03/2019 13:29:25",
      "content": "<blockquote>\n  <p>So kernel competitions are just meant to limit inference complexity</p>\n</blockquote>\n\n<p>It's not just about complexity. It's also:</p>\n\n<ul>\n<li>Provides true unseen test set (more representative of real-world ML)</li>\n<li>Deters hand labeling and attempts at cheating</li>\n<li>Enforces that solutions run in a nice containerized environment that the host can replicate</li>\n<li>Enforces that you can actually reproduce the submission you submit (a surprisingly common failure, even amongst top Kagglers who use version control)</li>\n<li>Avoids running a more complicated \"traditional\" two-stage competition format with a model upload step</li>\n<li>Will allow us to run more interesting/varied problem types by means of controlling the execution environment</li>\n</ul>",
      "rawMarkdown": "&gt; So kernel competitions are just meant to limit inference complexity\n\nIt's not just about complexity. It's also:\n\n- Provides true unseen test set (more representative of real-world ML)\n- Deters hand labeling and attempts at cheating\n- Enforces that solutions run in a nice containerized environment that the host can replicate\n- Enforces that you can actually reproduce the submission you submit (a surprisingly common failure, even amongst top Kagglers who use version control)\n- Avoids running a more complicated \"traditional\" two-stage competition format with a model upload step\n- Will allow us to run more interesting/varied problem types by means of controlling the execution environment",
      "votes": null
    },
    {
      "id": "506437",
      "postDate": "04/03/2019 13:32:33",
      "content": "<p><a href=\"/valanm\">@valanm</a> Thanks for the suggestion! I added \"You are permitted to train a model outside of Kernels and perform just the inference step from within Kernels. \" to that page to help clarify.</p>",
      "rawMarkdown": "valanm Thanks for the suggestion! I added \"You are permitted to train a model outside of Kernels and perform just the inference step from within Kernels. \" to that page to help clarify.",
      "votes": null
    },
    {
      "id": "506461",
      "postDate": "04/03/2019 14:05:24",
      "content": "<p>Thanks! </p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "506548",
      "postDate": "04/03/2019 16:09:34",
      "content": "<p>Thanks for the clarification. </p>",
      "rawMarkdown": "Thanks for the clarification.",
      "votes": null
    },
    {
      "id": "506684",
      "postDate": "04/03/2019 19:00:52",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!",
      "votes": null
    },
    {
      "id": "524273",
      "postDate": "04/28/2019 11:32:37",
      "content": "<p><a href=\"/wcukierski\">@wcukierski</a> </p>\n\n<p>Reading the rules I found that \"No internet access enabled\"</p>\n\n<p>I'd like to try things like parallelizing the training using internet access.</p>\n\n<p>Having read this post and the <a href=\"/inversion\">@inversion</a> one, I found that the rules are related to the inference time during stage 2.</p>\n\n<p>So, the same way I could train using my own GPU at home (but I had to do the inference using a Kernel GPU) ... following the same way of thinking: I could train using internet access in my kernel, as long as I didn't use the internet access during the inference kernel.</p>\n\n<p>Could I use internet access just during the training?</p>",
      "rawMarkdown": "wcukierski \n\nReading the rules I found that \"No internet access enabled\"\n\nI'd like to try things like parallelizing the training using internet access.\n\nHaving read this post and the @inversion one, I found that the rules are related to the inference time during stage 2.\n\nSo, the same way I could train using my own GPU at home (but I had to do the inference using a Kernel GPU) ... following the same way of thinking: I could train using internet access in my kernel, as long as I didn't use the internet access during the inference kernel.\n\nCould I use internet access just during the training?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 505123,
      "author_name": "rinnqd",
      "author_url": "",
      "post_date": "04/01/2019 15:34:31",
      "content": "<p>You can't submit an uploaded file.\nThere is no upload button in the submission tab</p>",
      "votes": null,
      "replies": [
        {
          "id": 505145,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "04/01/2019 16:02:14",
          "content": "<p>Really? I haven't tried that yet. So it doesn't seem to work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505155,
          "author_name": "rinnqd",
          "author_url": "",
          "post_date": "04/01/2019 16:14:52",
          "content": "<p>But I think you can upload the submission file as a dataset and then you import it to the notebook and export it with to_csv method. Finally submit it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505177,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "04/01/2019 16:46:26",
          "content": "<p>I would give it a try and maybe I will share my approach here if I succeed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505200,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "04/01/2019 17:31:33",
          "content": "<p>I had the same question than <a href=\"/syoya1997\">@syoya1997</a> </p>\n\n<p>With the same intention of switching to a Kaggle kernel once I find the right postprocessing, etc.</p>\n\n<p><a href=\"/rinnqd\">@rinnqd</a> </p>\n\n<p>It's true that you can't submit an uploaded file.</p>\n\n<p>But you can \"upload it\" inside a python script. For example, you can use something like the following code:</p>\n\n<p>```\nmy_submission='''id,attribute_ids\n1023b2cd, 0 1 2\n1023bbaa, 0 1 2\n                    ... here would be the complete submission.csv\n7163763b, 100 101\n'''</p>\n\n<p>with open('submission.csv', 'w') as f:\n    f.write(my_submission)</p>\n\n<p>```</p>\n\n<p><a href=\"/codingcliff\">@codingcliff</a> \nSo, is it allowed to \"upload\" CSVs during the model preparation? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 505240,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "04/01/2019 19:00:54",
      "content": "<p>Or you can rely on your validation set to see the improvements while training locally and then run the code on a kernel only once in a while to check if nothing weird is going on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 505306,
      "author_name": "codingcliff",
      "author_url": "",
      "post_date": "04/01/2019 21:01:56",
      "content": "<p>Hi syoya,</p>\n\n<p>Does training a model locally and submit a checkpoint as dataset and init from it works for you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 505894,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "04/02/2019 17:23:06",
          "content": "<p>Is this allowed? And do I have to make the dataset public?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505976,
          "author_name": "codingcliff",
          "author_url": "",
          "post_date": "04/02/2019 20:12:54",
          "content": "<p>My guess is yes to having to make the dataset public to use it in kernels. <a href=\"/wcukierski\">@wcukierski</a>  Correct me if I'm wrong. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506353,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "04/03/2019 10:46:30",
          "content": "<p>Sorry for a late reply. I'm quite busy yesterday and downloading the dataset takes me quite a while. And yes, submit a checkpoint as the dataset and init works for me indeed. However, I'm afraid that even making the dataset public cannot be a good solution. We can name the checkpoint as \"densenet120.h5\" or other seemingly public pre-trained model names and this cannot be well verified.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 505724,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "04/02/2019 13:16:04",
      "content": "<p>personal opinion: this seems like a grey area with potential for exploitation. To develop and debug locally I would just create my own train-val split, and once it is working run in kernel. But putting that submission as dataset and submitting seems to be against the intent of the rule. There would need to be a way to automatically check for this type of submission as it seems very much against the rules of for a final submission. </p>",
      "votes": null,
      "replies": [
        {
          "id": 505731,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "04/02/2019 13:26:51",
          "content": "<p>Uploading a static submission to Kernels (as opposed to source code or a model capable of predicting) will cause your code to fail when the Kernel is rerun on the held-out test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505981,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "04/02/2019 20:35:34",
          "content": "<p>I forget that there is an unaccessable test set for phase 2.  Only thing I see is whether this is a behavior that should increase in frequency in the future or not.  On this point I do not have an opinion. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505990,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "04/02/2019 20:58:24",
          "content": "<blockquote>\n  <p>model capable of predicting</p>\n</blockquote>\n\n<p><a href=\"/wcukierski\">@wcukierski</a> could you please clarify, must the model be trained in the kernel, or we can upload model trained on competition train data? That's a huge difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506122,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "04/03/2019 02:22:09",
          "content": "<p>You are permitted to train outside of Kernels and upload your model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506125,
          "author_name": "qiaojian",
          "author_url": "",
          "post_date": "04/03/2019 02:37:59",
          "content": "<p>Is it necessary  to make model public？ Thanks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506163,
          "author_name": "reachkishore",
          "author_url": "",
          "post_date": "04/03/2019 04:17:12",
          "content": "<p><a href=\"/wcukierski\">@wcukierski</a> thanks for the comment. This busted the myth I am having that training needs to be done from scratch in kernel environment only. <br>Because I was thinking that uploaded model could have used images/data from not mentioned sources. <br>For example in previous petfinder.my competition it is mentioned that people shouldn't use the images from petfinder.my website. if inclusion of offline trained model is allowed directly then we are opening a window for a possible misuse. <br> I think may be we should allow to include model from other kernel output at least we can track the training data.<br> I am sure kaggle will verify and will ask winners to reproduce the output. but on large scale even it will be difficult for kaggle admins to verify all medal winners code (bronze). hence for at least kernel competitions instead of offline model upload make a provision that their solution can span across kernels would be a better idea I believe. Please share your opinion. Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506224,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "04/03/2019 06:36:45",
          "content": "<blockquote>\n  <p>You are permitted to train outside of Kernels and upload your model.</p>\n</blockquote>\n\n<p>Oh thanks for clarification. Looks like this is the first kernel competition on kaggle where you are allowed to upload the trained model?\nSo we'll need to train a gazillion of models and stack, sad :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506235,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "04/03/2019 06:56:53",
          "content": "<p>Please, clarify the following: is it allowed to train at home, upload the checkpoints privately, and use these checkpoints for stage 2?</p>\n\n<p>In this way, probably the winner will be an ensemble of several checkpoints, and the winner kernel won't perform any training, just inference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506237,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/03/2019 07:01:15",
          "content": "<p>IMO, if this is a kernel competition as stated, it should not be allowed to train on competition data locally or outside of whitelisted data. Otherwise it should not be named kernel competition. <a href=\"https://www.kaggle.com/lopuhin\">@Konstantin</a> Indeed, this does make you sad :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506273,
          "author_name": "closears",
          "author_url": "",
          "post_date": "04/03/2019 08:36:06",
          "content": "<p>I guess it would be so boring if somebody trains a lot of models locally and uploads them to stack. It loses the meaning of kernel competition :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506356,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "04/03/2019 10:50:20",
          "content": "<p>Does that mean actually kernel competitions have no difference with other competitions? Then I think \"kernel\" competitions are meaningless.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506415,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "04/03/2019 12:47:40",
          "content": "<p>Our own <a href=\"/inversion\">@inversion</a> wrote a nice explanation of why we'd run Kernels-only competitions in this type of format: <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/87719\">https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/87719</a> (that applies to a different competition, but the spirit is the same here).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506424,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "04/03/2019 13:15:57",
          "content": "<p>Wow. So kernel competitions are just meant to limit inference complexity? This surprises me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506430,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/03/2019 13:26:02",
          "content": "<p>If that is the case here than you should make it clear in <strong>Overview</strong> under <strong>Kernel Requirements</strong> that only  the inference step must be performed in kernels (9 hours). The current statement is misleading. I wonder if anyone figured it out correctly excluding people working for kaggle</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506434,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "04/03/2019 13:29:25",
          "content": "<blockquote>\n  <p>So kernel competitions are just meant to limit inference complexity</p>\n</blockquote>\n\n<p>It's not just about complexity. It's also:</p>\n\n<ul>\n<li>Provides true unseen test set (more representative of real-world ML)</li>\n<li>Deters hand labeling and attempts at cheating</li>\n<li>Enforces that solutions run in a nice containerized environment that the host can replicate</li>\n<li>Enforces that you can actually reproduce the submission you submit (a surprisingly common failure, even amongst top Kagglers who use version control)</li>\n<li>Avoids running a more complicated \"traditional\" two-stage competition format with a model upload step</li>\n<li>Will allow us to run more interesting/varied problem types by means of controlling the execution environment</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506437,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "04/03/2019 13:32:33",
          "content": "<p><a href=\"/valanm\">@valanm</a> Thanks for the suggestion! I added \"You are permitted to train a model outside of Kernels and perform just the inference step from within Kernels. \" to that page to help clarify.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506461,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "04/03/2019 14:05:24",
          "content": "<p>Thanks! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506548,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "04/03/2019 16:09:34",
          "content": "<p>Thanks for the clarification. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 506684,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "04/03/2019 19:00:52",
      "content": "<p>thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 524273,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "04/28/2019 11:32:37",
      "content": "<p><a href=\"/wcukierski\">@wcukierski</a> </p>\n\n<p>Reading the rules I found that \"No internet access enabled\"</p>\n\n<p>I'd like to try things like parallelizing the training using internet access.</p>\n\n<p>Having read this post and the <a href=\"/inversion\">@inversion</a> one, I found that the rules are related to the inference time during stage 2.</p>\n\n<p>So, the same way I could train using my own GPU at home (but I had to do the inference using a Kernel GPU) ... following the same way of thinking: I could train using internet access in my kernel, as long as I didn't use the internet access during the inference kernel.</p>\n\n<p>Could I use internet access just during the training?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "505026": "Training in Kaggle kernels can quite time-consuming because it would take several hours to run the notebook and do some modifications and then another 9 hours to commit the kernel. I'm wondering if it's allowed to train locally and then just submit the results. Of course, I'll switch to Kaggle kernel after I get a satisfying version.",
    "505123": "You can't submit an uploaded file.\nThere is no upload button in the submission tab",
    "505145": "Really? I haven't tried that yet. So it doesn't seem to work.",
    "505155": "But I think you can upload the submission file as a dataset and then you import it to the notebook and export it with to_csv method. Finally submit it.",
    "505177": "I would give it a try and maybe I will share my approach here if I succeed.",
    "505200": "I had the same question than @syoya1997 \n\nWith the same intention of switching to a Kaggle kernel once I find the right postprocessing, etc.\n\n@rinnqd \n\nIt's true that you can't submit an uploaded file.\n\nBut you can \"upload it\" inside a python script. For example, you can use something like the following code:\n\n\n```\nmy_submission='''id,attribute_ids\n1023b2cd, 0 1 2\n1023bbaa, 0 1 2\n\t\t\t\t\t... here would be the complete submission.csv\n7163763b, 100 101\n'''\n\nwith open('submission.csv', 'w') as f:\n    f.write(my_submission)\n    \n```\n\n@codingcliff \nSo, is it allowed to \"upload\" CSVs during the model preparation?",
    "505240": "Or you can rely on your validation set to see the improvements while training locally and then run the code on a kernel only once in a while to check if nothing weird is going on.",
    "505306": "Hi syoya,\n\nDoes training a model locally and submit a checkpoint as dataset and init from it works for you?",
    "505724": "personal opinion: this seems like a grey area with potential for exploitation. To develop and debug locally I would just create my own train-val split, and once it is working run in kernel. But putting that submission as dataset and submitting seems to be against the intent of the rule. There would need to be a way to automatically check for this type of submission as it seems very much against the rules of for a final submission.",
    "505731": "Uploading a static submission to Kernels (as opposed to source code or a model capable of predicting) will cause your code to fail when the Kernel is rerun on the held-out test set.",
    "505894": "Is this allowed? And do I have to make the dataset public?",
    "505976": "My guess is yes to having to make the dataset public to use it in kernels. @wcukierski  Correct me if I'm wrong. Thanks!",
    "505981": "I forget that there is an unaccessable test set for phase 2.  Only thing I see is whether this is a behavior that should increase in frequency in the future or not.  On this point I do not have an opinion.",
    "505990": "&gt; model capable of predicting\n\n@wcukierski could you please clarify, must the model be trained in the kernel, or we can upload model trained on competition train data? That's a huge difference.",
    "506122": "You are permitted to train outside of Kernels and upload your model.",
    "506125": "Is it necessary  to make model public？ Thanks",
    "506163": "wcukierski thanks for the comment. This busted the myth I am having that training needs to be done from scratch in kernel environment only. <br>Because I was thinking that uploaded model could have used images/data from not mentioned sources. <br>For example in previous petfinder.my competition it is mentioned that people shouldn't use the images from petfinder.my website. if inclusion of offline trained model is allowed directly then we are opening a window for a possible misuse. <br> I think may be we should allow to include model from other kernel output at least we can track the training data.<br> I am sure kaggle will verify and will ask winners to reproduce the output. but on large scale even it will be difficult for kaggle admins to verify all medal winners code (bronze). hence for at least kernel competitions instead of offline model upload make a provision that their solution can span across kernels would be a better idea I believe. Please share your opinion. Thank you.",
    "506224": "&gt; You are permitted to train outside of Kernels and upload your model.\n\nOh thanks for clarification. Looks like this is the first kernel competition on kaggle where you are allowed to upload the trained model?\nSo we'll need to train a gazillion of models and stack, sad :(",
    "506235": "Please, clarify the following: is it allowed to train at home, upload the checkpoints privately, and use these checkpoints for stage 2?\n\nIn this way, probably the winner will be an ensemble of several checkpoints, and the winner kernel won't perform any training, just inference.",
    "506237": "IMO, if this is a kernel competition as stated, it should not be allowed to train on competition data locally or outside of whitelisted data. Otherwise it should not be named kernel competition. [@Konstantin](https://www.kaggle.com/lopuhin) Indeed, this does make you sad :(",
    "506273": "I guess it would be so boring if somebody trains a lot of models locally and uploads them to stack. It loses the meaning of kernel competition :(",
    "506353": "Sorry for a late reply. I'm quite busy yesterday and downloading the dataset takes me quite a while. And yes, submit a checkpoint as the dataset and init works for me indeed. However, I'm afraid that even making the dataset public cannot be a good solution. We can name the checkpoint as \"densenet120.h5\" or other seemingly public pre-trained model names and this cannot be well verified.",
    "506356": "Does that mean actually kernel competitions have no difference with other competitions? Then I think \"kernel\" competitions are meaningless.",
    "506415": "Our own @inversion wrote a nice explanation of why we'd run Kernels-only competitions in this type of format: https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/87719 (that applies to a different competition, but the spirit is the same here).",
    "506424": "Wow. So kernel competitions are just meant to limit inference complexity? This surprises me.",
    "506430": "If that is the case here than you should make it clear in **Overview** under **Kernel Requirements** that only  the inference step must be performed in kernels (9 hours). The current statement is misleading. I wonder if anyone figured it out correctly excluding people working for kaggle",
    "506434": "&gt; So kernel competitions are just meant to limit inference complexity\n\nIt's not just about complexity. It's also:\n\n- Provides true unseen test set (more representative of real-world ML)\n- Deters hand labeling and attempts at cheating\n- Enforces that solutions run in a nice containerized environment that the host can replicate\n- Enforces that you can actually reproduce the submission you submit (a surprisingly common failure, even amongst top Kagglers who use version control)\n- Avoids running a more complicated \"traditional\" two-stage competition format with a model upload step\n- Will allow us to run more interesting/varied problem types by means of controlling the execution environment",
    "506437": "valanm Thanks for the suggestion! I added \"You are permitted to train a model outside of Kernels and perform just the inference step from within Kernels. \" to that page to help clarify.",
    "506461": "Thanks!",
    "506548": "Thanks for the clarification.",
    "506684": "thanks!",
    "524273": "wcukierski \n\nReading the rules I found that \"No internet access enabled\"\n\nI'd like to try things like parallelizing the training using internet access.\n\nHaving read this post and the @inversion one, I found that the rules are related to the inference time during stage 2.\n\nSo, the same way I could train using my own GPU at home (but I had to do the inference using a Kernel GPU) ... following the same way of thinking: I could train using internet access in my kernel, as long as I didn't use the internet access during the inference kernel.\n\nCould I use internet access just during the training?"
  },
  "source": "meta"
}