{
  "id": 55473,
  "title": "Multiple Validation sets LightGBM",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55473",
  "author_name": "",
  "post_date": "2018-04-26T22:47:25.203659Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>Do you know precisely how works lightgbm when passing multiple validation set?</p>\n\n<p>If I pass 3 validation sets with the same size, would the model select the best average of all 3 validations sets? Or the best single one? Will it stop learning when all 3 models stop making improvement or only when one stops?</p>\n\n<p>I checked the documentation but I didn't find an answer.</p>\n\n<p>Thanks!\nAntoine.</p>",
  "messages": [
    {
      "id": "319824",
      "postDate": "04/26/2018 22:47:25",
      "content": "<p>Hello,</p>\n\n<p>Do you know precisely how works lightgbm when passing multiple validation set?</p>\n\n<p>If I pass 3 validation sets with the same size, would the model select the best average of all 3 validations sets? Or the best single one? Will it stop learning when all 3 models stop making improvement or only when one stops?</p>\n\n<p>I checked the documentation but I didn't find an answer.</p>\n\n<p>Thanks!\nAntoine.</p>",
      "rawMarkdown": "Hello,\n\nDo you know precisely how works lightgbm when passing multiple validation set?\n\nIf I pass 3 validation sets with the same size, would the model select the best average of all 3 validations sets? Or the best single one? Will it stop learning when all 3 models stop making improvement or only when one stops?\n\nI checked the documentation but I didn't find an answer.\n\nThanks!\nAntoine.",
      "votes": null
    },
    {
      "id": "319833",
      "postDate": "04/26/2018 23:25:19",
      "content": "<p>I assume you're asking specifically about early stopping. Lightgbm will stop training as soon as the score stops improving (for early_stopping_rounds) on any of the validation sets you pass.</p>",
      "rawMarkdown": "I assume you're asking specifically about early stopping. Lightgbm will stop training as soon as the score stops improving (for early_stopping_rounds) on any of the validation sets you pass.",
      "votes": null
    },
    {
      "id": "319845",
      "postDate": "04/27/2018 00:04:29",
      "content": "<p>Thanks Eddy, so I guess in this case, we use the best iteration on this particular validation set right? </p>\n\n<p>How do you do in this case to optimize your model when some validation sets stop after 100 rounds and other with 200 rounds? You would select the model after 100 rounds to be safe? And if you want to have good results for both validation set, you fit 2 models?</p>",
      "rawMarkdown": "Thanks Eddy, so I guess in this case, we use the best iteration on this particular validation set right? \n\nHow do you do in this case to optimize your model when some validation sets stop after 100 rounds and other with 200 rounds? You would select the model after 100 rounds to be safe? And if you want to have good results for both validation set, you fit 2 models?",
      "votes": null
    },
    {
      "id": "319848",
      "postDate": "04/27/2018 00:17:59",
      "content": "<p>Normally if you're using multiple validation sets (like in CV), you'd run the model once for each training fold and pass the out of fold set (validation) as the validation set to lgbm for early stopping. Let's say you do 5-fold CV, then you end up with 5 early stopping round numbers. </p>\n\n<p>If you want to then train a model on the entire dataset, there are different heuristics you can use to choose the number of trees based on the early stopped values you have. Personally I think the safest thing to do is take the average of the 5 numbers. It's a nice way to mitigate the risk of overfitting to any particular validation fold. I've heard of people taking the maximum, but that seems too aggressive to me.</p>",
      "rawMarkdown": "Normally if you're using multiple validation sets (like in CV), you'd run the model once for each training fold and pass the out of fold set (validation) as the validation set to lgbm for early stopping. Let's say you do 5-fold CV, then you end up with 5 early stopping round numbers. \n\nIf you want to then train a model on the entire dataset, there are different heuristics you can use to choose the number of trees based on the early stopped values you have. Personally I think the safest thing to do is take the average of the 5 numbers. It's a nice way to mitigate the risk of overfitting to any particular validation fold. I've heard of people taking the maximum, but that seems too aggressive to me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 319833,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/26/2018 23:25:19",
      "content": "<p>I assume you're asking specifically about early stopping. Lightgbm will stop training as soon as the score stops improving (for early_stopping_rounds) on any of the validation sets you pass.</p>",
      "votes": null,
      "replies": [
        {
          "id": 319845,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "04/27/2018 00:04:29",
          "content": "<p>Thanks Eddy, so I guess in this case, we use the best iteration on this particular validation set right? </p>\n\n<p>How do you do in this case to optimize your model when some validation sets stop after 100 rounds and other with 200 rounds? You would select the model after 100 rounds to be safe? And if you want to have good results for both validation set, you fit 2 models?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319848,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/27/2018 00:17:59",
          "content": "<p>Normally if you're using multiple validation sets (like in CV), you'd run the model once for each training fold and pass the out of fold set (validation) as the validation set to lgbm for early stopping. Let's say you do 5-fold CV, then you end up with 5 early stopping round numbers. </p>\n\n<p>If you want to then train a model on the entire dataset, there are different heuristics you can use to choose the number of trees based on the early stopped values you have. Personally I think the safest thing to do is take the average of the 5 numbers. It's a nice way to mitigate the risk of overfitting to any particular validation fold. I've heard of people taking the maximum, but that seems too aggressive to me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "319824": "Hello,\n\nDo you know precisely how works lightgbm when passing multiple validation set?\n\nIf I pass 3 validation sets with the same size, would the model select the best average of all 3 validations sets? Or the best single one? Will it stop learning when all 3 models stop making improvement or only when one stops?\n\nI checked the documentation but I didn't find an answer.\n\nThanks!\nAntoine.",
    "319833": "I assume you're asking specifically about early stopping. Lightgbm will stop training as soon as the score stops improving (for early_stopping_rounds) on any of the validation sets you pass.",
    "319845": "Thanks Eddy, so I guess in this case, we use the best iteration on this particular validation set right? \n\nHow do you do in this case to optimize your model when some validation sets stop after 100 rounds and other with 200 rounds? You would select the model after 100 rounds to be safe? And if you want to have good results for both validation set, you fit 2 models?",
    "319848": "Normally if you're using multiple validation sets (like in CV), you'd run the model once for each training fold and pass the out of fold set (validation) as the validation set to lgbm for early stopping. Let's say you do 5-fold CV, then you end up with 5 early stopping round numbers. \n\nIf you want to then train a model on the entire dataset, there are different heuristics you can use to choose the number of trees based on the early stopped values you have. Personally I think the safest thing to do is take the average of the 5 numbers. It's a nice way to mitigate the risk of overfitting to any particular validation fold. I've heard of people taking the maximum, but that seems too aggressive to me."
  },
  "source": "meta"
}