{
  "id": 339451,
  "title": "See how LightGBM learns as it grows trees",
  "url": "/competitions/amex-default-prediction/discussion/339451",
  "author_name": "",
  "post_date": "2022-07-24T21:12:26.283156800Z",
  "votes": 31,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I made <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>a post</strong></a> recently showing how we can visualize model predictions. Also attempted to rationalize the plot in the light of our results. The data for that plot was taken from  a neural network, and I got a question whether something similar can be done for gradient boosting machine models. The answer is yes, and here goes.</p>\n<p>In that previous example I was using activations from the penultimate neural network layer, which is a point where the network has learned everything it can from the data and it is up to the final layer to make a prediction. There is no such thing in LightGBM, but we can use leaf indices from each LightGBM tree as a fairly decent proxy. By the way, XGboost leaf indices can be used the same way.</p>\n<p>I will show two plots below. One is based on leaf indices only from the first 100 trees LightGBM has grown. At that point LightGBM has learned a great deal about the data, but not everything. The second plot will be all leaf indices up to a tree at early stopping (in this case tree #1341).</p>\n<p>As with that previous post, it could help if you right-hand click while hovering over the image and do <code>Open image in new tab</code>. In that new tab you will get a magnifying glass and should be able to zoom in and see the details.</p>\n<p><img src=\"https://i.ibb.co/C7JTs91/t-SNE-LGB-training-03.png\" alt=\"t-SNE plot\"></p>\n<p>It is impressive that after growing only 100 trees the gradient booster can already make a global separation between the two classes. Yet the boundary between data classes (which I call the zone of confusion in the other post) is fairly wide and diffuse.</p>\n<p>When the booster goes all the way to the early stopping, it has learned more about the data. In fact, at that point LightGBM has learned everything it can without overfitting. To my eyes the confusion zone is defined better. I think most of the points made in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>the original post</strong></a> apply here as well.</p>\n<p><img src=\"https://i.ibb.co/9ynmDZ3/t-SNE-LGB-training-05.png\" alt=\"t-SNE plot\"></p>",
  "messages": [
    {
      "id": "1869571",
      "postDate": "07/24/2022 21:12:26",
      "content": "<p>I made <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>a post</strong></a> recently showing how we can visualize model predictions. Also attempted to rationalize the plot in the light of our results. The data for that plot was taken from  a neural network, and I got a question whether something similar can be done for gradient boosting machine models. The answer is yes, and here goes.</p>\n<p>In that previous example I was using activations from the penultimate neural network layer, which is a point where the network has learned everything it can from the data and it is up to the final layer to make a prediction. There is no such thing in LightGBM, but we can use leaf indices from each LightGBM tree as a fairly decent proxy. By the way, XGboost leaf indices can be used the same way.</p>\n<p>I will show two plots below. One is based on leaf indices only from the first 100 trees LightGBM has grown. At that point LightGBM has learned a great deal about the data, but not everything. The second plot will be all leaf indices up to a tree at early stopping (in this case tree #1341).</p>\n<p>As with that previous post, it could help if you right-hand click while hovering over the image and do <code>Open image in new tab</code>. In that new tab you will get a magnifying glass and should be able to zoom in and see the details.</p>\n<p><img src=\"https://i.ibb.co/C7JTs91/t-SNE-LGB-training-03.png\" alt=\"t-SNE plot\"></p>\n<p>It is impressive that after growing only 100 trees the gradient booster can already make a global separation between the two classes. Yet the boundary between data classes (which I call the zone of confusion in the other post) is fairly wide and diffuse.</p>\n<p>When the booster goes all the way to the early stopping, it has learned more about the data. In fact, at that point LightGBM has learned everything it can without overfitting. To my eyes the confusion zone is defined better. I think most of the points made in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>the original post</strong></a> apply here as well.</p>\n<p><img src=\"https://i.ibb.co/9ynmDZ3/t-SNE-LGB-training-05.png\" alt=\"t-SNE plot\"></p>",
      "rawMarkdown": "I made [**a post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278) recently showing how we can visualize model predictions. Also attempted to rationalize the plot in the light of our results. The data for that plot was taken from  a neural network, and I got a question whether something similar can be done for gradient boosting machine models. The answer is yes, and here goes.\n\nIn that previous example I was using activations from the penultimate neural network layer, which is a point where the network has learned everything it can from the data and it is up to the final layer to make a prediction. There is no such thing in LightGBM, but we can use leaf indices from each LightGBM tree as a fairly decent proxy. By the way, XGboost leaf indices can be used the same way.\n\nI will show two plots below. One is based on leaf indices only from the first 100 trees LightGBM has grown. At that point LightGBM has learned a great deal about the data, but not everything. The second plot will be all leaf indices up to a tree at early stopping (in this case tree #1341).\n\nAs with that previous post, it could help if you right-hand click while hovering over the image and do `Open image in new tab`. In that new tab you will get a magnifying glass and should be able to zoom in and see the details.\n\n![t-SNE plot](https://i.ibb.co/C7JTs91/t-SNE-LGB-training-03.png)\n\nIt is impressive that after growing only 100 trees the gradient booster can already make a global separation between the two classes. Yet the boundary between data classes (which I call the zone of confusion in the other post) is fairly wide and diffuse.\n\nWhen the booster goes all the way to the early stopping, it has learned more about the data. In fact, at that point LightGBM has learned everything it can without overfitting. To my eyes the confusion zone is defined better. I think most of the points made in [**the original post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278) apply here as well.\n\n![t-SNE plot](https://i.ibb.co/9ynmDZ3/t-SNE-LGB-training-05.png)",
      "votes": null
    },
    {
      "id": "1869784",
      "postDate": "07/25/2022 03:45:44",
      "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Thanks for sharing the detailed information!!</p>",
      "rawMarkdown": "tilii7 Thanks for sharing the detailed information!!",
      "votes": null
    },
    {
      "id": "1869832",
      "postDate": "07/25/2022 04:49:22",
      "content": "<p>Sure thing! Hope you are enjoying \"broccoli\" images.</p>",
      "rawMarkdown": "Sure thing! Hope you are enjoying \"broccoli\" images.",
      "votes": null
    },
    {
      "id": "1869852",
      "postDate": "07/25/2022 05:13:28",
      "content": "<p>This is great visual and explanation <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, thanks for sharing with the community here!</p>",
      "rawMarkdown": "This is great visual and explanation @tilii7, thanks for sharing with the community here!",
      "votes": null
    },
    {
      "id": "1869857",
      "postDate": "07/25/2022 05:18:26",
      "content": "<p>Great information ! Thanks for sharing. </p>",
      "rawMarkdown": "Great information ! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1869941",
      "postDate": "07/25/2022 06:39:56",
      "content": "<p>very nice, thanks for sharing.</p>",
      "rawMarkdown": "very nice, thanks for sharing.",
      "votes": null
    },
    {
      "id": "1870249",
      "postDate": "07/25/2022 12:08:52",
      "content": "<p>Informative , thanks </p>",
      "rawMarkdown": "Informative , thanks",
      "votes": null
    },
    {
      "id": "1870380",
      "postDate": "07/25/2022 14:07:19",
      "content": "<p>Very nice, thanks for sharing.</p>",
      "rawMarkdown": "Very nice, thanks for sharing.",
      "votes": null
    },
    {
      "id": "1871645",
      "postDate": "07/26/2022 11:50:40",
      "content": "<p>Different viewpoint.</p>",
      "rawMarkdown": "Different viewpoint.",
      "votes": null
    },
    {
      "id": "1874158",
      "postDate": "07/28/2022 05:26:17",
      "content": "<p>Amazing, Thanks for sharing. <br>\nsorry for the basic question, but what do the indices represent ? <br>\nCould you please also publish the code that you used to generate this LGB/XGB Visualization (because i definitely want to incorporate this visualization practice in my notebooks too).</p>",
      "rawMarkdown": "Amazing, Thanks for sharing. \nsorry for the basic question, but what do the indices represent ? \nCould you please also publish the code that you used to generate this LGB/XGB Visualization (because i definitely want to incorporate this visualization practice in my notebooks too).",
      "votes": null
    },
    {
      "id": "1874253",
      "postDate": "07/28/2022 06:54:53",
      "content": "<p>For each gradient booster tree, the indices represent positions of the final leaf for a particular data point. They are integers, so for 100 data points and a booster that has grown 50 trees we get a 100x50 matrix of integers as an output. It is easy to get this matrix by adding <code>pred_leaf=True</code> to the <code>.predict</code> command that is normally used for predictions. Beware that you won't get actual predictions if you use this switch. This goes for both XGboost and LightGBM, though I will give you <a href=\"https://xgboost.readthedocs.io/en/stable/python/python_api.html#xgboost.Booster.predict\" target=\"_blank\"><strong>a link</strong></a> to the former. This <a href=\"https://datascience.stackexchange.com/questions/81970/in-xgboost-how-is-a-leaf-index-corresponding-to-the-particular-leaf-node-in-act\" target=\"_blank\"><strong>link</strong></a> may be helpful as well.</p>\n<p>I don't have code that is ready to share as there are various bits and pieces scattered across many different scripts.</p>",
      "rawMarkdown": "For each gradient booster tree, the indices represent positions of the final leaf for a particular data point. They are integers, so for 100 data points and a booster that has grown 50 trees we get a 100x50 matrix of integers as an output. It is easy to get this matrix by adding `pred_leaf=True` to the `.predict` command that is normally used for predictions. Beware that you won't get actual predictions if you use this switch. This goes for both XGboost and LightGBM, though I will give you [**a link**](https://xgboost.readthedocs.io/en/stable/python/python_api.html#xgboost.Booster.predict) to the former. This [**link**](https://datascience.stackexchange.com/questions/81970/in-xgboost-how-is-a-leaf-index-corresponding-to-the-particular-leaf-node-in-act) may be helpful as well.\n\nI don't have code that is ready to share as there are various bits and pieces scattered across many different scripts.",
      "votes": null
    },
    {
      "id": "1875628",
      "postDate": "07/29/2022 06:54:27",
      "content": "<p>Thanks a lot for the explanation!</p>",
      "rawMarkdown": "Thanks a lot for the explanation!",
      "votes": null
    },
    {
      "id": "1880843",
      "postDate": "08/02/2022 04:55:47",
      "content": "<p>thanks for u sharing</p>",
      "rawMarkdown": "thanks for u sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1869784,
      "author_name": "arunpurakkatt",
      "author_url": "",
      "post_date": "07/25/2022 03:45:44",
      "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Thanks for sharing the detailed information!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1869832,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/25/2022 04:49:22",
          "content": "<p>Sure thing! Hope you are enjoying \"broccoli\" images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869852,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "07/25/2022 05:13:28",
      "content": "<p>This is great visual and explanation <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, thanks for sharing with the community here!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869857,
      "author_name": "cid007",
      "author_url": "",
      "post_date": "07/25/2022 05:18:26",
      "content": "<p>Great information ! Thanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869941,
      "author_name": "rizqyad",
      "author_url": "",
      "post_date": "07/25/2022 06:39:56",
      "content": "<p>very nice, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1870249,
      "author_name": "gazu468",
      "author_url": "",
      "post_date": "07/25/2022 12:08:52",
      "content": "<p>Informative , thanks </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1870380,
      "author_name": "pablocampillo",
      "author_url": "",
      "post_date": "07/25/2022 14:07:19",
      "content": "<p>Very nice, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1871645,
      "author_name": "burakkaya35",
      "author_url": "",
      "post_date": "07/26/2022 11:50:40",
      "content": "<p>Different viewpoint.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1874158,
      "author_name": "virajkadam",
      "author_url": "",
      "post_date": "07/28/2022 05:26:17",
      "content": "<p>Amazing, Thanks for sharing. <br>\nsorry for the basic question, but what do the indices represent ? <br>\nCould you please also publish the code that you used to generate this LGB/XGB Visualization (because i definitely want to incorporate this visualization practice in my notebooks too).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1874253,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/28/2022 06:54:53",
          "content": "<p>For each gradient booster tree, the indices represent positions of the final leaf for a particular data point. They are integers, so for 100 data points and a booster that has grown 50 trees we get a 100x50 matrix of integers as an output. It is easy to get this matrix by adding <code>pred_leaf=True</code> to the <code>.predict</code> command that is normally used for predictions. Beware that you won't get actual predictions if you use this switch. This goes for both XGboost and LightGBM, though I will give you <a href=\"https://xgboost.readthedocs.io/en/stable/python/python_api.html#xgboost.Booster.predict\" target=\"_blank\"><strong>a link</strong></a> to the former. This <a href=\"https://datascience.stackexchange.com/questions/81970/in-xgboost-how-is-a-leaf-index-corresponding-to-the-particular-leaf-node-in-act\" target=\"_blank\"><strong>link</strong></a> may be helpful as well.</p>\n<p>I don't have code that is ready to share as there are various bits and pieces scattered across many different scripts.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1875628,
          "author_name": "virajkadam",
          "author_url": "",
          "post_date": "07/29/2022 06:54:27",
          "content": "<p>Thanks a lot for the explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1880843,
      "author_name": "",
      "author_url": "",
      "post_date": "08/02/2022 04:55:47",
      "content": "<p>thanks for u sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1869571": "I made [**a post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278) recently showing how we can visualize model predictions. Also attempted to rationalize the plot in the light of our results. The data for that plot was taken from  a neural network, and I got a question whether something similar can be done for gradient boosting machine models. The answer is yes, and here goes.\n\nIn that previous example I was using activations from the penultimate neural network layer, which is a point where the network has learned everything it can from the data and it is up to the final layer to make a prediction. There is no such thing in LightGBM, but we can use leaf indices from each LightGBM tree as a fairly decent proxy. By the way, XGboost leaf indices can be used the same way.\n\nI will show two plots below. One is based on leaf indices only from the first 100 trees LightGBM has grown. At that point LightGBM has learned a great deal about the data, but not everything. The second plot will be all leaf indices up to a tree at early stopping (in this case tree #1341).\n\nAs with that previous post, it could help if you right-hand click while hovering over the image and do `Open image in new tab`. In that new tab you will get a magnifying glass and should be able to zoom in and see the details.\n\n![t-SNE plot](https://i.ibb.co/C7JTs91/t-SNE-LGB-training-03.png)\n\nIt is impressive that after growing only 100 trees the gradient booster can already make a global separation between the two classes. Yet the boundary between data classes (which I call the zone of confusion in the other post) is fairly wide and diffuse.\n\nWhen the booster goes all the way to the early stopping, it has learned more about the data. In fact, at that point LightGBM has learned everything it can without overfitting. To my eyes the confusion zone is defined better. I think most of the points made in [**the original post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278) apply here as well.\n\n![t-SNE plot](https://i.ibb.co/9ynmDZ3/t-SNE-LGB-training-05.png)",
    "1869784": "tilii7 Thanks for sharing the detailed information!!",
    "1869832": "Sure thing! Hope you are enjoying \"broccoli\" images.",
    "1869852": "This is great visual and explanation @tilii7, thanks for sharing with the community here!",
    "1869857": "Great information ! Thanks for sharing.",
    "1869941": "very nice, thanks for sharing.",
    "1870249": "Informative , thanks",
    "1870380": "Very nice, thanks for sharing.",
    "1871645": "Different viewpoint.",
    "1874158": "Amazing, Thanks for sharing. \nsorry for the basic question, but what do the indices represent ? \nCould you please also publish the code that you used to generate this LGB/XGB Visualization (because i definitely want to incorporate this visualization practice in my notebooks too).",
    "1874253": "For each gradient booster tree, the indices represent positions of the final leaf for a particular data point. They are integers, so for 100 data points and a booster that has grown 50 trees we get a 100x50 matrix of integers as an output. It is easy to get this matrix by adding `pred_leaf=True` to the `.predict` command that is normally used for predictions. Beware that you won't get actual predictions if you use this switch. This goes for both XGboost and LightGBM, though I will give you [**a link**](https://xgboost.readthedocs.io/en/stable/python/python_api.html#xgboost.Booster.predict) to the former. This [**link**](https://datascience.stackexchange.com/questions/81970/in-xgboost-how-is-a-leaf-index-corresponding-to-the-particular-leaf-node-in-act) may be helpful as well.\n\nI don't have code that is ready to share as there are various bits and pieces scattered across many different scripts.",
    "1875628": "Thanks a lot for the explanation!",
    "1880843": "thanks for u sharing"
  },
  "source": "meta"
}