{
  "id": 338809,
  "title": "Use CSI and KS may avoid overfitting",
  "url": "/competitions/amex-default-prediction/discussion/338809",
  "author_name": "",
  "post_date": "2022-07-22T04:33:08.423032500Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When a model deteriorates in performance, checking distributional changes in the model variables can help with identifying possible causes. This is a step that is taken generally after one has checked Population Stability Index (PSI) and it’s not in the green zone (&lt;0.1 in general) to check that the overall population distribution can be attributed majorly to which variables.</p>\n<p>With <strong>Categorical variable</strong> can be used Characteristic stability index (CSI)</p>\n<ol>\n<li>CSI &lt; 0.1 = The variable hasn’t changed, and we can use to train the model</li>\n<li>0.1 ≤ CS1 &lt; 0.2 = The variable has slightly changed, and it is advisable to evaluate the impacts of these changes</li>\n<li>CSI ≥ 0.2 = The changes in variable are significant, and the model should not be used the characteristic in model.</li>\n</ol>\n<p>With <strong>Numeric variable</strong>, we can use CSI indicator, but it requires binned variable. To avoid, loss variable information, I intend use Kolmogorov–Smirnov test.</p>\n<p>Detail notebook at <a href=\"https://www.kaggle.com/nanabi/amex-ks-csi-nnb\" target=\"_blank\">here </a></p>\n<p>I hope get some advice! </p>",
  "messages": [
    {
      "id": "1865731",
      "postDate": "07/22/2022 04:33:08",
      "content": "<p>When a model deteriorates in performance, checking distributional changes in the model variables can help with identifying possible causes. This is a step that is taken generally after one has checked Population Stability Index (PSI) and it’s not in the green zone (&lt;0.1 in general) to check that the overall population distribution can be attributed majorly to which variables.</p>\n<p>With <strong>Categorical variable</strong> can be used Characteristic stability index (CSI)</p>\n<ol>\n<li>CSI &lt; 0.1 = The variable hasn’t changed, and we can use to train the model</li>\n<li>0.1 ≤ CS1 &lt; 0.2 = The variable has slightly changed, and it is advisable to evaluate the impacts of these changes</li>\n<li>CSI ≥ 0.2 = The changes in variable are significant, and the model should not be used the characteristic in model.</li>\n</ol>\n<p>With <strong>Numeric variable</strong>, we can use CSI indicator, but it requires binned variable. To avoid, loss variable information, I intend use Kolmogorov–Smirnov test.</p>\n<p>Detail notebook at <a href=\"https://www.kaggle.com/nanabi/amex-ks-csi-nnb\" target=\"_blank\">here </a></p>\n<p>I hope get some advice! </p>",
      "rawMarkdown": "When a model deteriorates in performance, checking distributional changes in the model variables can help with identifying possible causes. This is a step that is taken generally after one has checked Population Stability Index (PSI) and it’s not in the green zone (<0.1 in general) to check that the overall population distribution can be attributed majorly to which variables.\n\nWith **Categorical variable** can be used Characteristic stability index (CSI)\n\n1. CSI < 0.1 = The variable hasn’t changed, and we can use to train the model\n1. 0.1 ≤ CS1 < 0.2 = The variable has slightly changed, and it is advisable to evaluate the impacts of these changes\n1. CSI ≥ 0.2 = The changes in variable are significant, and the model should not be used the characteristic in model.\n\nWith **Numeric variable**, we can use CSI indicator, but it requires binned variable. To avoid, loss variable information, I intend use Kolmogorov–Smirnov test.\n\nDetail notebook at [here ](https://www.kaggle.com/nanabi/amex-ks-csi-nnb)\n\nI hope get some advice!",
      "votes": null
    },
    {
      "id": "1865969",
      "postDate": "07/22/2022 07:31:14",
      "content": "<p>Thanks for sharing !</p>\n<p>Looks like <code>S_9</code> ,<code>B_29</code>, <code>D_59</code>, <code>S_11</code> ,<code>R_1</code> ,<code>S_2</code> have significant change based on this. Interesting!</p>",
      "rawMarkdown": "Thanks for sharing !\n\nLooks like `S_9` ,`B_29`, `D_59`, `S_11` ,`R_1` ,`S_2` have significant change based on this. Interesting!",
      "votes": null
    },
    {
      "id": "1866055",
      "postDate": "07/22/2022 08:44:14",
      "content": "<p>Very nice, maybe we can use <a href=\"https://github.com/amphibian-dev/toad\" target=\"_blank\">toad</a> to do more effectual feature engineering.👀</p>",
      "rawMarkdown": "Very nice, maybe we can use [toad](https://github.com/amphibian-dev/toad) to do more effectual feature engineering.👀",
      "votes": null
    },
    {
      "id": "1874513",
      "postDate": "07/28/2022 10:24:53",
      "content": "<p>Yet. toad is used to do feature engineering with training set (may only).</p>\n<p>In this post, I checked stability of variables between training set and test set.</p>",
      "rawMarkdown": "Yet. toad is used to do feature engineering with training set (may only).\n\nIn this post, I checked stability of variables between training set and test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1865969,
      "author_name": "nitishraj",
      "author_url": "",
      "post_date": "07/22/2022 07:31:14",
      "content": "<p>Thanks for sharing !</p>\n<p>Looks like <code>S_9</code> ,<code>B_29</code>, <code>D_59</code>, <code>S_11</code> ,<code>R_1</code> ,<code>S_2</code> have significant change based on this. Interesting!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1866055,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "07/22/2022 08:44:14",
      "content": "<p>Very nice, maybe we can use <a href=\"https://github.com/amphibian-dev/toad\" target=\"_blank\">toad</a> to do more effectual feature engineering.👀</p>",
      "votes": null,
      "replies": [
        {
          "id": 1874513,
          "author_name": "nanabi",
          "author_url": "",
          "post_date": "07/28/2022 10:24:53",
          "content": "<p>Yet. toad is used to do feature engineering with training set (may only).</p>\n<p>In this post, I checked stability of variables between training set and test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1865731": "When a model deteriorates in performance, checking distributional changes in the model variables can help with identifying possible causes. This is a step that is taken generally after one has checked Population Stability Index (PSI) and it’s not in the green zone (<0.1 in general) to check that the overall population distribution can be attributed majorly to which variables.\n\nWith **Categorical variable** can be used Characteristic stability index (CSI)\n\n1. CSI < 0.1 = The variable hasn’t changed, and we can use to train the model\n1. 0.1 ≤ CS1 < 0.2 = The variable has slightly changed, and it is advisable to evaluate the impacts of these changes\n1. CSI ≥ 0.2 = The changes in variable are significant, and the model should not be used the characteristic in model.\n\nWith **Numeric variable**, we can use CSI indicator, but it requires binned variable. To avoid, loss variable information, I intend use Kolmogorov–Smirnov test.\n\nDetail notebook at [here ](https://www.kaggle.com/nanabi/amex-ks-csi-nnb)\n\nI hope get some advice!",
    "1865969": "Thanks for sharing !\n\nLooks like `S_9` ,`B_29`, `D_59`, `S_11` ,`R_1` ,`S_2` have significant change based on this. Interesting!",
    "1866055": "Very nice, maybe we can use [toad](https://github.com/amphibian-dev/toad) to do more effectual feature engineering.👀",
    "1874513": "Yet. toad is used to do feature engineering with training set (may only).\n\nIn this post, I checked stability of variables between training set and test set."
  },
  "source": "meta"
}