{
  "id": 208895,
  "title": "How to make submissions?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/208895",
  "author_name": "",
  "post_date": "2021-01-05T14:00:07.417372500Z",
  "votes": -2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So It seems that submissions are to be done via a <code>.csv</code> file. However, If I had a private dataset (which has the preprocessed files for the model) and the model derive training data from that private dataset, would that be allowed for submissions? It seems that they have said that <code>Internet access would be disabled</code> but I think Kaggle Private Datasets wouldn't use internet, right? They would presumably be shifted and transferred onsite…</p>\n<p>Or is the expectation there that within 9 hours of GPU time, Notebook has to make, preprocess, train the model and make a submission files (because that seems a bit harsh - my preprocessing is a bit lengthy, on top of that TTA and training would suck almost all the time availabe :( …..?</p>",
  "messages": [
    {
      "id": "1139585",
      "postDate": "01/05/2021 14:00:07",
      "content": "<p>So It seems that submissions are to be done via a <code>.csv</code> file. However, If I had a private dataset (which has the preprocessed files for the model) and the model derive training data from that private dataset, would that be allowed for submissions? It seems that they have said that <code>Internet access would be disabled</code> but I think Kaggle Private Datasets wouldn't use internet, right? They would presumably be shifted and transferred onsite…</p>\n<p>Or is the expectation there that within 9 hours of GPU time, Notebook has to make, preprocess, train the model and make a submission files (because that seems a bit harsh - my preprocessing is a bit lengthy, on top of that TTA and training would suck almost all the time availabe :( …..?</p>",
      "rawMarkdown": "So It seems that submissions are to be done via a `.csv` file. However, If I had a private dataset (which has the preprocessed files for the model) and the model derive training data from that private dataset, would that be allowed for submissions? It seems that they have said that `Internet access would be disabled` but I think Kaggle Private Datasets wouldn't use internet, right? They would presumably be shifted and transferred onsite...\n\nOr is the expectation there that within 9 hours of GPU time, Notebook has to make, preprocess, train the model and make a submission files (because that seems a bit harsh - my preprocessing is a bit lengthy, on top of that TTA and training would suck almost all the time availabe :( .....?",
      "votes": null
    },
    {
      "id": "1139680",
      "postDate": "01/05/2021 15:00:06",
      "content": "<p>You could have a separate notebook for training. Save the model you train and upload it to a Kaggle dataset. Load the model and do the inference.  </p>",
      "rawMarkdown": "You could have a separate notebook for training. Save the model you train and upload it to a Kaggle dataset. Load the model and do the inference.",
      "votes": null
    },
    {
      "id": "1139796",
      "postDate": "01/05/2021 16:19:15",
      "content": "<p>Datasets - both private or public are NOT loaded using the internet.  So this is the place for any data, libraries, etc that you need that are not part of the Kaggle docker image.  You can have multiple datasets but the total size limit is 20GB if they are private.  If you run up against this limit - make some of the sets public.    There may be a limit to the size you can use on a single kernel….</p>\n<p>The CPU and/or GPU time limit applies to any kernel you run - at the end of the time limit they die.  But you can have a whole bunch of individual kernels that makeup your pipeline, training, inference and submission.  Learning to put all those small pieces together is a skill set needed if your only compute resources are Kaggle kernels.</p>",
      "rawMarkdown": "Datasets - both private or public are NOT loaded using the internet.  So this is the place for any data, libraries, etc that you need that are not part of the Kaggle docker image.  You can have multiple datasets but the total size limit is 20GB if they are private.  If you run up against this limit - make some of the sets public.    There may be a limit to the size you can use on a single kernel....\n\nThe CPU and/or GPU time limit applies to any kernel you run - at the end of the time limit they die.  But you can have a whole bunch of individual kernels that makeup your pipeline, training, inference and submission.  Learning to put all those small pieces together is a skill set needed if your only compute resources are Kaggle kernels.",
      "votes": null
    },
    {
      "id": "1139833",
      "postDate": "01/05/2021 16:43:58",
      "content": "<blockquote>\n  <p>…But you can have a whole bunch of individual kernels that makeup your pipeline, training…      </p>\n</blockquote>\n<p>So you mean that we can have a submission for more than 1 kernel, each of which would have its own 9h quota?</p>",
      "rawMarkdown": "> ...But you can have a whole bunch of individual kernels that makeup your pipeline, training...      \n\nSo you mean that we can have a submission for more than 1 kernel, each of which would have its own 9h quota?",
      "votes": null
    },
    {
      "id": "1140313",
      "postDate": "01/05/2021 23:26:06",
      "content": "<p>On my local PC I might build a very long kernel to train a model and make a submission file.  Most of the models I am training right now on my local machine are taking longer than the hours limit on Kaggle.  So to run the same on Kaggle I would need to break the single script into several pieces.  The LAST piece would do the prediction and generation of the submission file.</p>\n<p>For example - if I wanted to create 66,000 images to make a balanced data set - I would do that in one kernel.  </p>\n<p>The output (60,000 images) would be added to a private data set and a second kernel might be used for training.  But that many images might need more than 8 hours so you might save the partial trained model and start a 3td kernel to finish the training.</p>\n<p>A final kernel might be used to do the prediction and file submission.</p>",
      "rawMarkdown": "On my local PC I might build a very long kernel to train a model and make a submission file.  Most of the models I am training right now on my local machine are taking longer than the hours limit on Kaggle.  So to run the same on Kaggle I would need to break the single script into several pieces.  The LAST piece would do the prediction and generation of the submission file.\n\nFor example - if I wanted to create 66,000 images to make a balanced data set - I would do that in one kernel.  \n\nThe output (60,000 images) would be added to a private data set and a second kernel might be used for training.  But that many images might need more than 8 hours so you might save the partial trained model and start a 3td kernel to finish the training.\n\nA final kernel might be used to do the prediction and file submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1139680,
      "author_name": "mithilsalunkhe",
      "author_url": "",
      "post_date": "01/05/2021 15:00:06",
      "content": "<p>You could have a separate notebook for training. Save the model you train and upload it to a Kaggle dataset. Load the model and do the inference.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1139796,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/05/2021 16:19:15",
      "content": "<p>Datasets - both private or public are NOT loaded using the internet.  So this is the place for any data, libraries, etc that you need that are not part of the Kaggle docker image.  You can have multiple datasets but the total size limit is 20GB if they are private.  If you run up against this limit - make some of the sets public.    There may be a limit to the size you can use on a single kernel….</p>\n<p>The CPU and/or GPU time limit applies to any kernel you run - at the end of the time limit they die.  But you can have a whole bunch of individual kernels that makeup your pipeline, training, inference and submission.  Learning to put all those small pieces together is a skill set needed if your only compute resources are Kaggle kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139833,
          "author_name": "neelg007",
          "author_url": "",
          "post_date": "01/05/2021 16:43:58",
          "content": "<blockquote>\n  <p>…But you can have a whole bunch of individual kernels that makeup your pipeline, training…      </p>\n</blockquote>\n<p>So you mean that we can have a submission for more than 1 kernel, each of which would have its own 9h quota?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140313,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2021 23:26:06",
          "content": "<p>On my local PC I might build a very long kernel to train a model and make a submission file.  Most of the models I am training right now on my local machine are taking longer than the hours limit on Kaggle.  So to run the same on Kaggle I would need to break the single script into several pieces.  The LAST piece would do the prediction and generation of the submission file.</p>\n<p>For example - if I wanted to create 66,000 images to make a balanced data set - I would do that in one kernel.  </p>\n<p>The output (60,000 images) would be added to a private data set and a second kernel might be used for training.  But that many images might need more than 8 hours so you might save the partial trained model and start a 3td kernel to finish the training.</p>\n<p>A final kernel might be used to do the prediction and file submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1139585": "So It seems that submissions are to be done via a `.csv` file. However, If I had a private dataset (which has the preprocessed files for the model) and the model derive training data from that private dataset, would that be allowed for submissions? It seems that they have said that `Internet access would be disabled` but I think Kaggle Private Datasets wouldn't use internet, right? They would presumably be shifted and transferred onsite...\n\nOr is the expectation there that within 9 hours of GPU time, Notebook has to make, preprocess, train the model and make a submission files (because that seems a bit harsh - my preprocessing is a bit lengthy, on top of that TTA and training would suck almost all the time availabe :( .....?",
    "1139680": "You could have a separate notebook for training. Save the model you train and upload it to a Kaggle dataset. Load the model and do the inference.",
    "1139796": "Datasets - both private or public are NOT loaded using the internet.  So this is the place for any data, libraries, etc that you need that are not part of the Kaggle docker image.  You can have multiple datasets but the total size limit is 20GB if they are private.  If you run up against this limit - make some of the sets public.    There may be a limit to the size you can use on a single kernel....\n\nThe CPU and/or GPU time limit applies to any kernel you run - at the end of the time limit they die.  But you can have a whole bunch of individual kernels that makeup your pipeline, training, inference and submission.  Learning to put all those small pieces together is a skill set needed if your only compute resources are Kaggle kernels.",
    "1139833": "> ...But you can have a whole bunch of individual kernels that makeup your pipeline, training...      \n\nSo you mean that we can have a submission for more than 1 kernel, each of which would have its own 9h quota?",
    "1140313": "On my local PC I might build a very long kernel to train a model and make a submission file.  Most of the models I am training right now on my local machine are taking longer than the hours limit on Kaggle.  So to run the same on Kaggle I would need to break the single script into several pieces.  The LAST piece would do the prediction and generation of the submission file.\n\nFor example - if I wanted to create 66,000 images to make a balanced data set - I would do that in one kernel.  \n\nThe output (60,000 images) would be added to a private data set and a second kernel might be used for training.  But that many images might need more than 8 hours so you might save the partial trained model and start a 3td kernel to finish the training.\n\nA final kernel might be used to do the prediction and file submission."
  },
  "source": "meta"
}