{
  "id": 382051,
  "title": "using all data to train model without validating.   vs    KFold&Ensemble",
  "url": "/competitions/otto-recommender-system/discussion/382051",
  "author_name": "",
  "post_date": "2023-01-29T11:45:24.371824900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Which one would you choose ? </p>",
  "messages": [
    {
      "id": "2120146",
      "postDate": "01/29/2023 11:45:24",
      "content": "<p>Which one would you choose ? </p>",
      "rawMarkdown": "Which one would you choose ?",
      "votes": null
    },
    {
      "id": "2120222",
      "postDate": "01/29/2023 12:41:34",
      "content": "<p>Actually I prefer both but for different steps.</p>\n<ul>\n<li>I would go with KFold-Ensemble with splitting my local set to two by session IDs (50-50, 70-30, up to you) and simulate the exact scenario for my validation as I'm preparing a submission for Kaggle. I would train my voting models using CV on the first split, then calculate the Kaggle score on the second one.</li>\n<li>Once I'm sure with the whole pipeline and robustness of my approach, I would go with the full data blindly and make my final submissions.</li>\n</ul>",
      "rawMarkdown": "Actually I prefer both but for different steps.\n\n- I would go with KFold-Ensemble with splitting my local set to two by session IDs (50-50, 70-30, up to you) and simulate the exact scenario for my validation as I'm preparing a submission for Kaggle. I would train my voting models using CV on the first split, then calculate the Kaggle score on the second one.\n- Once I'm sure with the whole pipeline and robustness of my approach, I would go with the full data blindly and make my final submissions.",
      "votes": null
    },
    {
      "id": "2120343",
      "postDate": "01/29/2023 14:35:09",
      "content": "<p>Thx. And do you mean you will choose the full data to train model to make the final submission. For my experience, ensemble is more likely better than the full data single model.</p>",
      "rawMarkdown": "Thx. And do you mean you will choose the full data to train model to make the final submission. For my experience, ensemble is more likely better than the full data single model.",
      "votes": null
    },
    {
      "id": "2120347",
      "postDate": "01/29/2023 14:40:40",
      "content": "<p>Nope, I mean:</p>\n<p>For model development:</p>\n<ul>\n<li>50-50 split the train data by sessions</li>\n<li>CV train on first half</li>\n<li>Calculate score on second half with ensembled predictions</li>\n</ul>\n<p>For final submission:</p>\n<ul>\n<li>Use all train data</li>\n<li>CV train</li>\n<li>Make ensembled predictions on test data</li>\n</ul>\n<p>Sorry if I caused misunderstanding above.</p>",
      "rawMarkdown": "Nope, I mean:\n\nFor model development:\n- 50-50 split the train data by sessions\n- CV train on first half\n- Calculate score on second half with ensembled predictions\n\nFor final submission:\n- Use all train data\n- CV train\n- Make ensembled predictions on test data\n\nSorry if I caused misunderstanding above.",
      "votes": null
    },
    {
      "id": "2120374",
      "postDate": "01/29/2023 14:57:04",
      "content": "<p>thank u. I got it clearly right now, but still don't have a model trained by the full data. I worry if want to use full data to train a single modle, I have no way of knowing the optimal hyper-parameters of the model for full the data. Maybe I will test the follow:<br>\nUsing full data:</p>\n<ul>\n<li>5 Fold to get 5 models</li>\n<li>all data to get 1 model</li>\n<li>ensemble the 6 model above</li>\n</ul>",
      "rawMarkdown": "thank u. I got it clearly right now, but still don't have a model trained by the full data. I worry if want to use full data to train a single modle, I have no way of knowing the optimal hyper-parameters of the model for full the data. Maybe I will test the follow:\nUsing full data:\n- 5 Fold to get 5 models\n- all data to get 1 model\n- ensemble the 6 model above",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2120222,
      "author_name": "nlztrk",
      "author_url": "",
      "post_date": "01/29/2023 12:41:34",
      "content": "<p>Actually I prefer both but for different steps.</p>\n<ul>\n<li>I would go with KFold-Ensemble with splitting my local set to two by session IDs (50-50, 70-30, up to you) and simulate the exact scenario for my validation as I'm preparing a submission for Kaggle. I would train my voting models using CV on the first split, then calculate the Kaggle score on the second one.</li>\n<li>Once I'm sure with the whole pipeline and robustness of my approach, I would go with the full data blindly and make my final submissions.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2120343,
          "author_name": "earlee0412",
          "author_url": "",
          "post_date": "01/29/2023 14:35:09",
          "content": "<p>Thx. And do you mean you will choose the full data to train model to make the final submission. For my experience, ensemble is more likely better than the full data single model.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2120347,
              "author_name": "nlztrk",
              "author_url": "",
              "post_date": "01/29/2023 14:40:40",
              "content": "<p>Nope, I mean:</p>\n<p>For model development:</p>\n<ul>\n<li>50-50 split the train data by sessions</li>\n<li>CV train on first half</li>\n<li>Calculate score on second half with ensembled predictions</li>\n</ul>\n<p>For final submission:</p>\n<ul>\n<li>Use all train data</li>\n<li>CV train</li>\n<li>Make ensembled predictions on test data</li>\n</ul>\n<p>Sorry if I caused misunderstanding above.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2120374,
                  "author_name": "earlee0412",
                  "author_url": "",
                  "post_date": "01/29/2023 14:57:04",
                  "content": "<p>thank u. I got it clearly right now, but still don't have a model trained by the full data. I worry if want to use full data to train a single modle, I have no way of knowing the optimal hyper-parameters of the model for full the data. Maybe I will test the follow:<br>\nUsing full data:</p>\n<ul>\n<li>5 Fold to get 5 models</li>\n<li>all data to get 1 model</li>\n<li>ensemble the 6 model above</li>\n</ul>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2120146": "Which one would you choose ?",
    "2120222": "Actually I prefer both but for different steps.\n\n- I would go with KFold-Ensemble with splitting my local set to two by session IDs (50-50, 70-30, up to you) and simulate the exact scenario for my validation as I'm preparing a submission for Kaggle. I would train my voting models using CV on the first split, then calculate the Kaggle score on the second one.\n- Once I'm sure with the whole pipeline and robustness of my approach, I would go with the full data blindly and make my final submissions.",
    "2120343": "Thx. And do you mean you will choose the full data to train model to make the final submission. For my experience, ensemble is more likely better than the full data single model.",
    "2120347": "Nope, I mean:\n\nFor model development:\n- 50-50 split the train data by sessions\n- CV train on first half\n- Calculate score on second half with ensembled predictions\n\nFor final submission:\n- Use all train data\n- CV train\n- Make ensembled predictions on test data\n\nSorry if I caused misunderstanding above.",
    "2120374": "thank u. I got it clearly right now, but still don't have a model trained by the full data. I worry if want to use full data to train a single modle, I have no way of knowing the optimal hyper-parameters of the model for full the data. Maybe I will test the follow:\nUsing full data:\n- 5 Fold to get 5 models\n- all data to get 1 model\n- ensemble the 6 model above"
  },
  "source": "meta"
}