{
  "id": 56212,
  "title": "Question about full data set training and validation ",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56212",
  "author_name": "",
  "post_date": "2018-05-07T21:16:55.055777300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Fellow competition experts,</p>\n\n<p>I have a question regarding the right way to use lightgbm. </p>\n\n<p>Let me describe the procedure I am using.</p>\n\n<p>1) Use combined day 7 and 8 data for training with day 9 data as the validation set. </p>\n\n<p>2) Save the best_iteration from step 1.</p>\n\n<p>3) Use combined day 7, 8 and 9 data for training without validation.</p>\n\n<p>4) Predict test set with the model from step 3 and the best_iteration from step 2. </p>\n\n<p>I just wonder if the approach I am using is theoretically correct.</p>",
  "messages": [
    {
      "id": "324613",
      "postDate": "05/07/2018 21:16:55",
      "content": "<p>Fellow competition experts,</p>\n\n<p>I have a question regarding the right way to use lightgbm. </p>\n\n<p>Let me describe the procedure I am using.</p>\n\n<p>1) Use combined day 7 and 8 data for training with day 9 data as the validation set. </p>\n\n<p>2) Save the best_iteration from step 1.</p>\n\n<p>3) Use combined day 7, 8 and 9 data for training without validation.</p>\n\n<p>4) Predict test set with the model from step 3 and the best_iteration from step 2. </p>\n\n<p>I just wonder if the approach I am using is theoretically correct.</p>",
      "rawMarkdown": "Fellow competition experts,\n\nI have a question regarding the right way to use lightgbm. \n\nLet me describe the procedure I am using.\n\n1) Use combined day 7 and 8 data for training with day 9 data as the validation set. \n\n2) Save the best_iteration from step 1.\n\n3) Use combined day 7, 8 and 9 data for training without validation.\n\n4) Predict test set with the model from step 3 and the best_iteration from step 2. \n\nI just wonder if the approach I am using is theoretically correct.",
      "votes": null
    },
    {
      "id": "324615",
      "postDate": "05/07/2018 21:18:21",
      "content": "<p>Yes :)</p>",
      "rawMarkdown": "Yes :)",
      "votes": null
    },
    {
      "id": "324628",
      "postDate": "05/07/2018 21:37:39",
      "content": "<p>Joe,</p>\n\n<p>Thank you very much for confirming this!</p>\n\n<p>I choose to use full data set because I am expecting extreme shakeup when the private board is revealed. \nHowever, it indeed took very long time to complete the procedure even though I can access very powerful computing resources. </p>",
      "rawMarkdown": "Joe,\n\nThank you very much for confirming this!\n\nI choose to use full data set because I am expecting extreme shakeup when the private board is revealed. \nHowever, it indeed took very long time to complete the procedure even though I can access very powerful computing resources.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324615,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/07/2018 21:18:21",
      "content": "<p>Yes :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 324628,
          "author_name": "richinmind",
          "author_url": "",
          "post_date": "05/07/2018 21:37:39",
          "content": "<p>Joe,</p>\n\n<p>Thank you very much for confirming this!</p>\n\n<p>I choose to use full data set because I am expecting extreme shakeup when the private board is revealed. \nHowever, it indeed took very long time to complete the procedure even though I can access very powerful computing resources. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "324613": "Fellow competition experts,\n\nI have a question regarding the right way to use lightgbm. \n\nLet me describe the procedure I am using.\n\n1) Use combined day 7 and 8 data for training with day 9 data as the validation set. \n\n2) Save the best_iteration from step 1.\n\n3) Use combined day 7, 8 and 9 data for training without validation.\n\n4) Predict test set with the model from step 3 and the best_iteration from step 2. \n\nI just wonder if the approach I am using is theoretically correct.",
    "324615": "Yes :)",
    "324628": "Joe,\n\nThank you very much for confirming this!\n\nI choose to use full data set because I am expecting extreme shakeup when the private board is revealed. \nHowever, it indeed took very long time to complete the procedure even though I can access very powerful computing resources."
  },
  "source": "meta"
}