{
  "id": 335524,
  "title": "Lag Features Are All You Need",
  "url": "/competitions/amex-default-prediction/discussion/335524",
  "author_name": "The Devastator",
  "post_date": "2022-07-06T15:16:01.411000",
  "votes": 55,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>Lag Features  You Need</h2>\n<p>OK. Maybe not <strong>all</strong> you need. But they do <strong>improve LightGBM</strong>.</p>\n<p>I tried a simple experiment that worked pretty well:</p>\n<h3>Lag Features</h3>\n<p>In this competition, we get information about clients of AMEX over time. <br>\nMost high-scoring notebooks in this competition focused on aggregating the information per client and creating a single row of extracted features: One for each client.</p>\n<p><strong>One of such agg functions is <code>last</code></strong>.</p>\n<p>A quick examination revealed that the <code>last</code> feature is extremely powerful for predicting if a client defaults or not (well.. makes sense..). </p>\n<p>So I took this concept two steps further: </p>\n<ul>\n<li><strong>\"First\" feature:</strong> Just like the <code>last</code> feature: I added a <code>first</code> feature. </li>\n<li><strong>\"Lag\" features</strong> to capture the change over time about each client I calculated two features for every <code>first</code>, <code>last</code> pair:<ul>\n<li><strong>\"Last - First\":</strong> The change from we first see the client to the last time we see the client.</li>\n<li><strong>\"Last / First\":</strong> The fractional difference from we first see the client to the last time we see the client.</li></ul></li>\n</ul>\n<p>This improved my <code>LightGBM</code> model to the point that it overtook my whole <code>LightGBM</code> + <code>Catboost</code> + <code>XGB</code> ensemble.</p>\n<ul>\n<li><strong>The notebook can be found <a href=\"https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need/\" target=\"_blank\">here</a></strong></li>\n</ul>\n<p>I uploaded a <a href=\"https://www.kaggle.com/datasets/thedevastator/amex-fe/settings\" target=\"_blank\">dataset</a> containing the extracted lag features and updated the final <a href=\"https://www.kaggle.com/datasets/thedevastator/amex-predictions\" target=\"_blank\">model predictions</a> (only <code>LightGBM</code> this time) for everyone to play with. </p>\n<p><br></p>\n<hr>\n<p><strong>Next Experiment (currently running):</strong> <br>\nMore \"lag features\" variations - Taking into consideration other parts of the time series. <br>\nwill keep you updated.</p>\n<hr>\n<p></p>\n<blockquote>\n  <p><strong>Credits:</strong>  The model is based on <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963\" target=\"_blank\">this</a> amazing notebook.</p>\n</blockquote>",
  "messages": [
    {
      "id": 1845782,
      "postDate": "2022-07-06T15:16:01.410Z",
      "content": "<h2>Lag Features  You Need</h2>\n<p>OK. Maybe not <strong>all</strong> you need. But they do <strong>improve LightGBM</strong>.</p>\n<p>I tried a simple experiment that worked pretty well:</p>\n<h3>Lag Features</h3>\n<p>In this competition, we get information about clients of AMEX over time. <br>\nMost high-scoring notebooks in this competition focused on aggregating the information per client and creating a single row of extracted features: One for each client.</p>\n<p><strong>One of such agg functions is <code>last</code></strong>.</p>\n<p>A quick examination revealed that the <code>last</code> feature is extremely powerful for predicting if a client defaults or not (well.. makes sense..). </p>\n<p>So I took this concept two steps further: </p>\n<ul>\n<li><strong>\"First\" feature:</strong> Just like the <code>last</code> feature: I added a <code>first</code> feature. </li>\n<li><strong>\"Lag\" features</strong> to capture the change over time about each client I calculated two features for every <code>first</code>, <code>last</code> pair:<ul>\n<li><strong>\"Last - First\":</strong> The change from we first see the client to the last time we see the client.</li>\n<li><strong>\"Last / First\":</strong> The fractional difference from we first see the client to the last time we see the client.</li></ul></li>\n</ul>\n<p>This improved my <code>LightGBM</code> model to the point that it overtook my whole <code>LightGBM</code> + <code>Catboost</code> + <code>XGB</code> ensemble.</p>\n<ul>\n<li><strong>The notebook can be found <a href=\"https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need/\" target=\"_blank\">here</a></strong></li>\n</ul>\n<p>I uploaded a <a href=\"https://www.kaggle.com/datasets/thedevastator/amex-fe/settings\" target=\"_blank\">dataset</a> containing the extracted lag features and updated the final <a href=\"https://www.kaggle.com/datasets/thedevastator/amex-predictions\" target=\"_blank\">model predictions</a> (only <code>LightGBM</code> this time) for everyone to play with. </p>\n<p><br></p>\n<hr>\n<p><strong>Next Experiment (currently running):</strong> <br>\nMore \"lag features\" variations - Taking into consideration other parts of the time series. <br>\nwill keep you updated.</p>\n<hr>\n<p></p>\n<blockquote>\n  <p><strong>Credits:</strong>  The model is based on <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963\" target=\"_blank\">this</a> amazing notebook.</p>\n</blockquote>",
      "rawMarkdown": "## Lag Features ~~Are All~~ You Need\n\n\nOK. Maybe not **all** you need. But they do **improve LightGBM**.\n\nI tried a simple experiment that worked pretty well:\n\n### Lag Features\n\nIn this competition, we get information about clients of AMEX over time. \nMost high-scoring notebooks in this competition focused on aggregating the information per client and creating a single row of extracted features: One for each client.\n\n**One of such agg functions is `last`**.\n\nA quick examination revealed that the `last` feature is extremely powerful for predicting if a client defaults or not (well.. makes sense..). \n\nSo I took this concept two steps further: \n\n- **\"First\" feature:** Just like the `last` feature: I added a `first` feature. \n- **\"Lag\" features** to capture the change over time about each client I calculated two features for every `first`, `last` pair:\n     - **\"Last - First\":** The change from we first see the client to the last time we see the client.\n     - **\"Last / First\":** The fractional difference from we first see the client to the last time we see the client.\n\nThis improved my `LightGBM` model to the point that it overtook my whole `LightGBM` + `Catboost` + `XGB` ensemble.\n\n- **The notebook can be found [here](https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need/)**\n\nI uploaded a [dataset](https://www.kaggle.com/datasets/thedevastator/amex-fe/settings) containing the extracted lag features and updated the final [model predictions](https://www.kaggle.com/datasets/thedevastator/amex-predictions) (only `LightGBM` this time) for everyone to play with. \n\n<br>\n\n_____\n\n**Next Experiment (currently running):** \nMore \"lag features\" variations - Taking into consideration other parts of the time series. \nwill keep you updated.\n_____\n\n<be>\n\n\n\n> **Credits:**  The model is based on [this](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963) amazing notebook.\n\n\n",
      "votes": 54
    },
    {
      "id": 1845893,
      "postDate": "2022-07-06T16:54:53.460Z",
      "content": "<p>Missed this gorilla lying in plain sight :D. Thanks for opening my eyes.</p>",
      "rawMarkdown": "Missed this gorilla lying in plain sight :D. Thanks for opening my eyes.",
      "votes": 1
    },
    {
      "id": 1849357,
      "postDate": "2022-07-09T12:51:00.120Z",
      "content": "<p>Hi, how are you able to account for the customers with less than 13 rows (specifically customers with 1 row) for features like (last/first)?</p>",
      "rawMarkdown": "Hi, how are you able to account for the customers with less than 13 rows (specifically customers with 1 row) for features like (last/first)?"
    },
    {
      "id": 1847813,
      "postDate": "2022-07-08T06:59:43.510Z",
      "content": "<p>Thanks man. The subtraction is similar to the first difference in a ARIMA (x,1,x) model. </p>",
      "rawMarkdown": "Thanks man. The subtraction is similar to the first difference in a ARIMA (x,1,x) model. "
    },
    {
      "id": 1846317,
      "postDate": "2022-07-07T02:46:40Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1846346,
          "postDate": "2022-07-07T03:14:20.320Z",
          "content": "<p>GBT at least would get there pretty indirectly -- a tree has no direct way of understanding feature1 - feature2, only lots of iterations of something like if feature1 &gt; large_value and feature2 &lt; small_value then extract something about that spread. In general, an intrinsic limitation (but also benefit!) of trees is that they do not reason about different features as if they were on the same scale -- features scales are completely independent. To be fair, a neural network would be able to directly extract max-min as a hidden layer feature but of course suffers other drawbacks on a dataset like this. </p>\n<p>A good way to think about this is that even if your model can approximate a feature extremely well, the model should be simpler and better if that feature were just fed in directly. Another good way to think about this is that the most impactful feature engineering is often guided by the information structure that a specific model is most likely to miss / struggle to recreate from scratch. Those small adjustments are the sorts of things that matter at the 3rd or 4th decimal place, even when they may be less important in real-world modeling.</p>",
          "rawMarkdown": "GBT at least would get there pretty indirectly -- a tree has no direct way of understanding feature1 - feature2, only lots of iterations of something like if feature1 > large_value and feature2 < small_value then extract something about that spread. In general, an intrinsic limitation (but also benefit!) of trees is that they do not reason about different features as if they were on the same scale -- features scales are completely independent. To be fair, a neural network would be able to directly extract max-min as a hidden layer feature but of course suffers other drawbacks on a dataset like this. \n\nA good way to think about this is that even if your model can approximate a feature extremely well, the model should be simpler and better if that feature were just fed in directly. Another good way to think about this is that the most impactful feature engineering is often guided by the information structure that a specific model is most likely to miss / struggle to recreate from scratch. Those small adjustments are the sorts of things that matter at the 3rd or 4th decimal place, even when they may be less important in real-world modeling.",
          "votes": 17
        },
        {
          "id": 1846388,
          "postDate": "2022-07-07T03:55:53.947Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1845893,
      "author_name": "Karan Dora",
      "author_url": "",
      "post_date": "2022-07-06T16:54:53.460000",
      "content": "<p>Missed this gorilla lying in plain sight :D. Thanks for opening my eyes.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1849357,
      "author_name": "Tarrasque9",
      "author_url": "",
      "post_date": "2022-07-09T12:51:00.120000",
      "content": "<p>Hi, how are you able to account for the customers with less than 13 rows (specifically customers with 1 row) for features like (last/first)?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1847813,
      "author_name": "Ben Fung",
      "author_url": "",
      "post_date": "2022-07-08T06:59:43.510000",
      "content": "<p>Thanks man. The subtraction is similar to the first difference in a ARIMA (x,1,x) model. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1846317,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-07T02:46:40",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1846346,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2022-07-07T03:14:20.320000",
          "content": "<p>GBT at least would get there pretty indirectly -- a tree has no direct way of understanding feature1 - feature2, only lots of iterations of something like if feature1 &gt; large_value and feature2 &lt; small_value then extract something about that spread. In general, an intrinsic limitation (but also benefit!) of trees is that they do not reason about different features as if they were on the same scale -- features scales are completely independent. To be fair, a neural network would be able to directly extract max-min as a hidden layer feature but of course suffers other drawbacks on a dataset like this. </p>\n<p>A good way to think about this is that even if your model can approximate a feature extremely well, the model should be simpler and better if that feature were just fed in directly. Another good way to think about this is that the most impactful feature engineering is often guided by the information structure that a specific model is most likely to miss / struggle to recreate from scratch. Those small adjustments are the sorts of things that matter at the 3rd or 4th decimal place, even when they may be less important in real-world modeling.</p>",
          "votes": 17,
          "replies": []
        },
        {
          "id": 1846388,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-07T03:55:53.947000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1845782": "## Lag Features ~~Are All~~ You Need\n\n\nOK. Maybe not **all** you need. But they do **improve LightGBM**.\n\nI tried a simple experiment that worked pretty well:\n\n### Lag Features\n\nIn this competition, we get information about clients of AMEX over time. \nMost high-scoring notebooks in this competition focused on aggregating the information per client and creating a single row of extracted features: One for each client.\n\n**One of such agg functions is `last`**.\n\nA quick examination revealed that the `last` feature is extremely powerful for predicting if a client defaults or not (well.. makes sense..). \n\nSo I took this concept two steps further: \n\n- **\"First\" feature:** Just like the `last` feature: I added a `first` feature. \n- **\"Lag\" features** to capture the change over time about each client I calculated two features for every `first`, `last` pair:\n     - **\"Last - First\":** The change from we first see the client to the last time we see the client.\n     - **\"Last / First\":** The fractional difference from we first see the client to the last time we see the client.\n\nThis improved my `LightGBM` model to the point that it overtook my whole `LightGBM` + `Catboost` + `XGB` ensemble.\n\n- **The notebook can be found [here](https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need/)**\n\nI uploaded a [dataset](https://www.kaggle.com/datasets/thedevastator/amex-fe/settings) containing the extracted lag features and updated the final [model predictions](https://www.kaggle.com/datasets/thedevastator/amex-predictions) (only `LightGBM` this time) for everyone to play with. \n\n<br>\n\n_____\n\n**Next Experiment (currently running):** \nMore \"lag features\" variations - Taking into consideration other parts of the time series. \nwill keep you updated.\n_____\n\n<be>\n\n\n\n> **Credits:**  The model is based on [this](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963) amazing notebook.\n\n\n",
    "1845893": "Missed this gorilla lying in plain sight :D. Thanks for opening my eyes.",
    "1849357": "Hi, how are you able to account for the customers with less than 13 rows (specifically customers with 1 row) for features like (last/first)?",
    "1847813": "Thanks man. The subtraction is similar to the first difference in a ARIMA (x,1,x) model. ",
    "1846317": ""
  }
}