{
  "id": 492989,
  "title": "Don't Drop Columns With Many Categories, Instead Collapse",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/492989",
  "author_name": "",
  "post_date": "2024-04-11T17:48:25.959361700Z",
  "votes": 21,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I noticed that some popular notebooks are dropping columns with some 3000 categories, but I think that we should just collapse them</p>\n<p>Here is a quick function:</p>\n<pre><code>def collapseData(X, , originalDf, threshold=, replaceVal=\"NOTPRESENTVALUE\"):\n    dataValCounts = originalDf[].value_counts()/len(originalDf)\n    replaceDict = {i:replaceVal  i  dataValCounts[dataValCounts&lt;=threshold].}\n\n    XUnique = X[].()\n    replaceDict.({i:replaceVal  i  XUnique[~np.isin(XUnique, dataValCounts.)]})\n\n     X[].map(lambda x: replaceDict.(x, x))\n</code></pre>\n<p>I used this for the person dataset thing and significantly reduced categories from like 7,500 to just 500 (improving cv by about 0.004)</p>",
  "messages": [
    {
      "id": "2747159",
      "postDate": "04/11/2024 17:48:25",
      "content": "<p>I noticed that some popular notebooks are dropping columns with some 3000 categories, but I think that we should just collapse them</p>\n<p>Here is a quick function:</p>\n<pre><code>def collapseData(X, , originalDf, threshold=, replaceVal=\"NOTPRESENTVALUE\"):\n    dataValCounts = originalDf[].value_counts()/len(originalDf)\n    replaceDict = {i:replaceVal  i  dataValCounts[dataValCounts&lt;=threshold].}\n\n    XUnique = X[].()\n    replaceDict.({i:replaceVal  i  XUnique[~np.isin(XUnique, dataValCounts.)]})\n\n     X[].map(lambda x: replaceDict.(x, x))\n</code></pre>\n<p>I used this for the person dataset thing and significantly reduced categories from like 7,500 to just 500 (improving cv by about 0.004)</p>",
      "rawMarkdown": "I noticed that some popular notebooks are dropping columns with some 3000 categories, but I think that we should just collapse them\n\nHere is a quick function:\n```\ndef collapseData(X, column, originalDf, threshold=0.001, replaceVal=\"NOTPRESENTVALUE\"):\n    dataValCounts = originalDf[column].value_counts()/len(originalDf)\n    replaceDict = {i:replaceVal for i in dataValCounts[dataValCounts<=threshold].index}\n    \n    XUnique = X[column].unique()\n    replaceDict.update({i:replaceVal for i in XUnique[~np.isin(XUnique, dataValCounts.index)]})\n    \n    return X[column].map(lambda x: replaceDict.get(x, x))\n```\n\nI used this for the person dataset thing and significantly reduced categories from like 7,500 to just 500 (improving cv by about 0.004)",
      "votes": null
    },
    {
      "id": "2748071",
      "postDate": "04/12/2024 08:28:26",
      "content": "<p>Thanks for your suggestion. I will try it</p>",
      "rawMarkdown": "Thanks for your suggestion. I will try it",
      "votes": null
    },
    {
      "id": "2748406",
      "postDate": "04/12/2024 12:28:41",
      "content": "<p>I'm back again. It's unfortunate that I didn't see any improvement in scores before exhausting today's DAILY SUBMISSIONS. Hopefully, adjusting the threshold tomorrow will change this situation. Good luck to you!</p>",
      "rawMarkdown": "I'm back again. It's unfortunate that I didn't see any improvement in scores before exhausting today's DAILY SUBMISSIONS. Hopefully, adjusting the threshold tomorrow will change this situation. Good luck to you!",
      "votes": null
    },
    {
      "id": "2748701",
      "postDate": "04/12/2024 14:41:00",
      "content": "<p>very nice and interesting, but I have a doubt, doesn't the internal binning of lgbm take care of this large number of categories in a column? I would be very much interested in comparing this approach with the internal binning strategy of lgbm.</p>",
      "rawMarkdown": "very nice and interesting, but I have a doubt, doesn't the internal binning of lgbm take care of this large number of categories in a column? I would be very much interested in comparing this approach with the internal binning strategy of lgbm.",
      "votes": null
    },
    {
      "id": "2748760",
      "postDate": "04/12/2024 15:37:01",
      "content": "<p>Maybe, it may be the case but I actually also use this to make count encoding columns which I believe some people are talking about. Also, it's just nice to reduce the column count. 7,500 columns is just a little much.</p>",
      "rawMarkdown": "Maybe, it may be the case but I actually also use this to make count encoding columns which I believe some people are talking about. Also, it's just nice to reduce the column count. 7,500 columns is just a little much.",
      "votes": null
    },
    {
      "id": "2748956",
      "postDate": "04/12/2024 18:31:06",
      "content": "<p>This approach can also be used instead of removing columns with a number of categories greater than a certain threshold. Therefore, some columns can be kept.</p>",
      "rawMarkdown": "This approach can also be used instead of removing columns with a number of categories greater than a certain threshold. Therefore, some columns can be kept.",
      "votes": null
    },
    {
      "id": "2750179",
      "postDate": "04/13/2024 13:51:48",
      "content": "<p>Thanks for sharing will try this &amp; update if CV improves !!</p>",
      "rawMarkdown": "Thanks for sharing will try this & update if CV improves !!",
      "votes": null
    },
    {
      "id": "2809721",
      "postDate": "05/13/2024 00:28:50",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2811860",
      "postDate": "05/14/2024 00:57:50",
      "content": "<p>Which threshold did you choose?</p>",
      "rawMarkdown": "Which threshold did you choose?",
      "votes": null
    },
    {
      "id": "2811984",
      "postDate": "05/14/2024 03:48:44",
      "content": "<p>Thanks, I'll give it a try.</p>",
      "rawMarkdown": "Thanks, I'll give it a try.",
      "votes": null
    },
    {
      "id": "2812603",
      "postDate": "05/14/2024 10:23:13",
      "content": "<p>I have tried but catched this bug.It said that str div str. Can anyone help me solve it? Thanks.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20132679%2F39af4a59c6c1e4282d46aa159ce63cae%2FScreenshot%202024-05-14%20171739.png?generation=1715682156871611&amp;alt=media\"></p>",
      "rawMarkdown": "I have tried but catched this bug.It said that str div str. Can anyone help me solve it? Thanks.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20132679%2F39af4a59c6c1e4282d46aa159ce63cae%2FScreenshot%202024-05-14%20171739.png?generation=1715682156871611&alt=media)",
      "votes": null
    },
    {
      "id": "2813540",
      "postDate": "05/14/2024 20:17:18",
      "content": "<p>Here you have an encoding depending on the cardinality that covers all the categorical columns <a href=\"https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric\" target=\"_blank\">https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric</a></p>",
      "rawMarkdown": "Here you have an encoding depending on the cardinality that covers all the categorical columns https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2748071,
      "author_name": "tianya55",
      "author_url": "",
      "post_date": "04/12/2024 08:28:26",
      "content": "<p>Thanks for your suggestion. I will try it</p>",
      "votes": null,
      "replies": [
        {
          "id": 2748406,
          "author_name": "tianya55",
          "author_url": "",
          "post_date": "04/12/2024 12:28:41",
          "content": "<p>I'm back again. It's unfortunate that I didn't see any improvement in scores before exhausting today's DAILY SUBMISSIONS. Hopefully, adjusting the threshold tomorrow will change this situation. Good luck to you!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2811860,
              "author_name": "phamhoanglenguyen",
              "author_url": "",
              "post_date": "05/14/2024 00:57:50",
              "content": "<p>Which threshold did you choose?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2748701,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/12/2024 14:41:00",
      "content": "<p>very nice and interesting, but I have a doubt, doesn't the internal binning of lgbm take care of this large number of categories in a column? I would be very much interested in comparing this approach with the internal binning strategy of lgbm.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2748760,
          "author_name": "kawaiicoderuwu",
          "author_url": "",
          "post_date": "04/12/2024 15:37:01",
          "content": "<p>Maybe, it may be the case but I actually also use this to make count encoding columns which I believe some people are talking about. Also, it's just nice to reduce the column count. 7,500 columns is just a little much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2748956,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "04/12/2024 18:31:06",
          "content": "<p>This approach can also be used instead of removing columns with a number of categories greater than a certain threshold. Therefore, some columns can be kept.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2750179,
      "author_name": "kunduruanil",
      "author_url": "",
      "post_date": "04/13/2024 13:51:48",
      "content": "<p>Thanks for sharing will try this &amp; update if CV improves !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2809721,
      "author_name": "dannykan",
      "author_url": "",
      "post_date": "05/13/2024 00:28:50",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2811984,
      "author_name": "kamonabeyy",
      "author_url": "",
      "post_date": "05/14/2024 03:48:44",
      "content": "<p>Thanks, I'll give it a try.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2812603,
      "author_name": "nguynphmhongl",
      "author_url": "",
      "post_date": "05/14/2024 10:23:13",
      "content": "<p>I have tried but catched this bug.It said that str div str. Can anyone help me solve it? Thanks.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20132679%2F39af4a59c6c1e4282d46aa159ce63cae%2FScreenshot%202024-05-14%20171739.png?generation=1715682156871611&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2813540,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "05/14/2024 20:17:18",
      "content": "<p>Here you have an encoding depending on the cardinality that covers all the categorical columns <a href=\"https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric\" target=\"_blank\">https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2747159": "I noticed that some popular notebooks are dropping columns with some 3000 categories, but I think that we should just collapse them\n\nHere is a quick function:\n```\ndef collapseData(X, column, originalDf, threshold=0.001, replaceVal=\"NOTPRESENTVALUE\"):\n    dataValCounts = originalDf[column].value_counts()/len(originalDf)\n    replaceDict = {i:replaceVal for i in dataValCounts[dataValCounts<=threshold].index}\n    \n    XUnique = X[column].unique()\n    replaceDict.update({i:replaceVal for i in XUnique[~np.isin(XUnique, dataValCounts.index)]})\n    \n    return X[column].map(lambda x: replaceDict.get(x, x))\n```\n\nI used this for the person dataset thing and significantly reduced categories from like 7,500 to just 500 (improving cv by about 0.004)",
    "2748071": "Thanks for your suggestion. I will try it",
    "2748406": "I'm back again. It's unfortunate that I didn't see any improvement in scores before exhausting today's DAILY SUBMISSIONS. Hopefully, adjusting the threshold tomorrow will change this situation. Good luck to you!",
    "2748701": "very nice and interesting, but I have a doubt, doesn't the internal binning of lgbm take care of this large number of categories in a column? I would be very much interested in comparing this approach with the internal binning strategy of lgbm.",
    "2748760": "Maybe, it may be the case but I actually also use this to make count encoding columns which I believe some people are talking about. Also, it's just nice to reduce the column count. 7,500 columns is just a little much.",
    "2748956": "This approach can also be used instead of removing columns with a number of categories greater than a certain threshold. Therefore, some columns can be kept.",
    "2750179": "Thanks for sharing will try this & update if CV improves !!",
    "2809721": "Thanks for sharing.",
    "2811860": "Which threshold did you choose?",
    "2811984": "Thanks, I'll give it a try.",
    "2812603": "I have tried but catched this bug.It said that str div str. Can anyone help me solve it? Thanks.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20132679%2F39af4a59c6c1e4282d46aa159ce63cae%2FScreenshot%202024-05-14%20171739.png?generation=1715682156871611&alt=media)",
    "2813540": "Here you have an encoding depending on the cardinality that covers all the categorical columns https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric"
  },
  "source": "meta"
}