{
  "id": 94064,
  "title": "Best workflow with K-Folds? ",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/94064",
  "author_name": "",
  "post_date": "2019-06-01T15:15:44.896088700Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi! </p>\n\n<p>I'm a bit new to kaggle, and I'm wondering what's the best workflow with K-folds-CV? \nAs we only have 1h to run our kernel for sub, a 5KCV means we train our model for 5 times. For me, that's so far way over 1h to reach good results. Do you guys train the localy? Or do you train the model, perform predictions in every fold, average the predictions, export the submission-file as a dataset, and just load the submission file in a new kernel and submit that? </p>\n\n<p>Thanks for the help :) </p>",
  "messages": [
    {
      "id": "541027",
      "postDate": "06/01/2019 15:15:44",
      "content": "<p>Hi! </p>\n\n<p>I'm a bit new to kaggle, and I'm wondering what's the best workflow with K-folds-CV? \nAs we only have 1h to run our kernel for sub, a 5KCV means we train our model for 5 times. For me, that's so far way over 1h to reach good results. Do you guys train the localy? Or do you train the model, perform predictions in every fold, average the predictions, export the submission-file as a dataset, and just load the submission file in a new kernel and submit that? </p>\n\n<p>Thanks for the help :) </p>",
      "rawMarkdown": "Hi! \n\nI'm a bit new to kaggle, and I'm wondering what's the best workflow with K-folds-CV? \nAs we only have 1h to run our kernel for sub, a 5KCV means we train our model for 5 times. For me, that's so far way over 1h to reach good results. Do you guys train the localy? Or do you train the model, perform predictions in every fold, average the predictions, export the submission-file as a dataset, and just load the submission file in a new kernel and submit that? \n\nThanks for the help :)",
      "votes": null
    },
    {
      "id": "541061",
      "postDate": "06/01/2019 16:25:26",
      "content": "<p>Note that the 1 hour limit is for inference only. You are free to train your model locally, or train the model in a non-submission kernel (I believe the default limit for Kaggle kernels is 8-9 hours), and then use the weights from the trained model as a dataset in your inference kernel.</p>\n\n<p>You cannot just upload the submissions file because we are going to run your inference kernels on a private test dataset so your kernel needs to run your trained model on the files that we provide to generate a submission.</p>",
      "rawMarkdown": "Note that the 1 hour limit is for inference only. You are free to train your model locally, or train the model in a non-submission kernel (I believe the default limit for Kaggle kernels is 8-9 hours), and then use the weights from the trained model as a dataset in your inference kernel.\n\nYou cannot just upload the submissions file because we are going to run your inference kernels on a private test dataset so your kernel needs to run your trained model on the files that we provide to generate a submission.",
      "votes": null
    },
    {
      "id": "541338",
      "postDate": "06/02/2019 09:01:03",
      "content": "<p>Hello <a href=\"/plakal\">@plakal</a>  when i upload trained model as dataset ,  inference kernel run finished. I want to submit, and the submit button says, can't use external data. BUT i only upload my trained model and the dataset</p>",
      "rawMarkdown": "Hello @plakal  when i upload trained model as dataset ,  inference kernel run finished. I want to submit, and the submit button says, can't use external data. BUT i only upload my trained model and the dataset",
      "votes": null
    },
    {
      "id": "541339",
      "postDate": "06/02/2019 09:02:02",
      "content": "<p><a href=\"/plakal\">@plakal</a> can you help me?</p>",
      "rawMarkdown": "plakal can you help me?",
      "votes": null
    },
    {
      "id": "541343",
      "postDate": "06/02/2019 09:15:18",
      "content": "<p><a href=\"/plakal\">@plakal</a>  I solved it, I directly load the kernel output. I should download the trained model weight, and the upload it.</p>",
      "rawMarkdown": "plakal  I solved it, I directly load the kernel output. I should download the trained model weight, and the upload it.",
      "votes": null
    },
    {
      "id": "541500",
      "postDate": "06/02/2019 15:00:30",
      "content": "<p>Hi, How do you manage kernel version for each fold? I prepare K kernel for KFold. It is messy. If I add some change, I need to change all K kernels. Do you guys have any tips?</p>",
      "rawMarkdown": "Hi, How do you manage kernel version for each fold? I prepare K kernel for KFold. It is messy. If I add some change, I need to change all K kernels. Do you guys have any tips?",
      "votes": null
    },
    {
      "id": "541502",
      "postDate": "06/02/2019 15:03:28",
      "content": "<p>You can publish your kernel output as \"dataset\". and then you can load your own dataset from your another kernel. You donnot need to download data :)</p>",
      "rawMarkdown": "You can publish your kernel output as \"dataset\". and then you can load your own dataset from your another kernel. You donnot need to download data :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 541061,
      "author_name": "plakal",
      "author_url": "",
      "post_date": "06/01/2019 16:25:26",
      "content": "<p>Note that the 1 hour limit is for inference only. You are free to train your model locally, or train the model in a non-submission kernel (I believe the default limit for Kaggle kernels is 8-9 hours), and then use the weights from the trained model as a dataset in your inference kernel.</p>\n\n<p>You cannot just upload the submissions file because we are going to run your inference kernels on a private test dataset so your kernel needs to run your trained model on the files that we provide to generate a submission.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 541338,
      "author_name": "zxyu1995",
      "author_url": "",
      "post_date": "06/02/2019 09:01:03",
      "content": "<p>Hello <a href=\"/plakal\">@plakal</a>  when i upload trained model as dataset ,  inference kernel run finished. I want to submit, and the submit button says, can't use external data. BUT i only upload my trained model and the dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 541339,
      "author_name": "zxyu1995",
      "author_url": "",
      "post_date": "06/02/2019 09:02:02",
      "content": "<p><a href=\"/plakal\">@plakal</a> can you help me?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 541343,
      "author_name": "zxyu1995",
      "author_url": "",
      "post_date": "06/02/2019 09:15:18",
      "content": "<p><a href=\"/plakal\">@plakal</a>  I solved it, I directly load the kernel output. I should download the trained model weight, and the upload it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 541502,
          "author_name": "sega1031",
          "author_url": "",
          "post_date": "06/02/2019 15:03:28",
          "content": "<p>You can publish your kernel output as \"dataset\". and then you can load your own dataset from your another kernel. You donnot need to download data :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 541500,
      "author_name": "sega1031",
      "author_url": "",
      "post_date": "06/02/2019 15:00:30",
      "content": "<p>Hi, How do you manage kernel version for each fold? I prepare K kernel for KFold. It is messy. If I add some change, I need to change all K kernels. Do you guys have any tips?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "541027": "Hi! \n\nI'm a bit new to kaggle, and I'm wondering what's the best workflow with K-folds-CV? \nAs we only have 1h to run our kernel for sub, a 5KCV means we train our model for 5 times. For me, that's so far way over 1h to reach good results. Do you guys train the localy? Or do you train the model, perform predictions in every fold, average the predictions, export the submission-file as a dataset, and just load the submission file in a new kernel and submit that? \n\nThanks for the help :)",
    "541061": "Note that the 1 hour limit is for inference only. You are free to train your model locally, or train the model in a non-submission kernel (I believe the default limit for Kaggle kernels is 8-9 hours), and then use the weights from the trained model as a dataset in your inference kernel.\n\nYou cannot just upload the submissions file because we are going to run your inference kernels on a private test dataset so your kernel needs to run your trained model on the files that we provide to generate a submission.",
    "541338": "Hello @plakal  when i upload trained model as dataset ,  inference kernel run finished. I want to submit, and the submit button says, can't use external data. BUT i only upload my trained model and the dataset",
    "541339": "plakal can you help me?",
    "541343": "plakal  I solved it, I directly load the kernel output. I should download the trained model weight, and the upload it.",
    "541500": "Hi, How do you manage kernel version for each fold? I prepare K kernel for KFold. It is messy. If I add some change, I need to change all K kernels. Do you guys have any tips?",
    "541502": "You can publish your kernel output as \"dataset\". and then you can load your own dataset from your another kernel. You donnot need to download data :)"
  },
  "source": "meta"
}