{
  "id": 308931,
  "title": "effective way for postal_code encoding method",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308931",
  "author_name": "",
  "post_date": "2022-02-21T03:24:09.925717300Z",
  "votes": 21,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In metadata of customer(<code>customers.csv</code>), postal code/address is an effective metadata, but as i have seen in most notebooks, every one is skipping this features. In practice, address is very effective for predicting sale amount: customers from big city tend to purchase more than other customers from countryside, and customers in same city may share same trending when purchasing. So i note my ideal about encoding method here for anyone who want to use postal_code/address as a feature in their model.</p>\n<h1>encoding method</h1>\n<ul>\n<li>We have so many distinct values (352899 different values) in postal_code features, so we can not use one-hot encoding. So we need to group it first.</li>\n<li>By counting rows for each postal code, we see that so many customer share same postal code. As an example:  120303 customers (8% of all customers) have same postal code. So some postal_codes have so many customers (customers came from same city), but some postal_codes have only a customer. </li>\n<li>So we have ideal here: Let group postal_code by customer size. All of postal_code had many customers will be in a group (big city group), and other postal code will be in other group.</li>\n<li>In my experience, i use below groups:<ul>\n<li>Postal_code with customer_size &gt; 1000: big_city</li>\n<li>Postal_code with customer_size &gt; 100 &amp; &lt;1000: medium city</li>\n<li>Postal_code with customer_size &gt; 10 &amp; &lt;100: small city</li>\n<li>Postal_code with customer_size &lt;10: very small city.</li></ul></li>\n</ul>\n<p>This is my code:</p>\n<pre><code>df = pd.read_csv(r\"customers.csv\")\ntest = df.postal_code.value_counts().reset_index()\ntest['bins'] = pd.cut(test['postal_code'], bins=[-1, 10, 100, 1e4], include_lowest=True)\n</code></pre>\n<p>Please upvote if it help you.<br>\nIf you have any other ideals, please share with me and others by comments </p>",
  "messages": [
    {
      "id": "1699231",
      "postDate": "02/21/2022 03:24:09",
      "content": "<p>In metadata of customer(<code>customers.csv</code>), postal code/address is an effective metadata, but as i have seen in most notebooks, every one is skipping this features. In practice, address is very effective for predicting sale amount: customers from big city tend to purchase more than other customers from countryside, and customers in same city may share same trending when purchasing. So i note my ideal about encoding method here for anyone who want to use postal_code/address as a feature in their model.</p>\n<h1>encoding method</h1>\n<ul>\n<li>We have so many distinct values (352899 different values) in postal_code features, so we can not use one-hot encoding. So we need to group it first.</li>\n<li>By counting rows for each postal code, we see that so many customer share same postal code. As an example:  120303 customers (8% of all customers) have same postal code. So some postal_codes have so many customers (customers came from same city), but some postal_codes have only a customer. </li>\n<li>So we have ideal here: Let group postal_code by customer size. All of postal_code had many customers will be in a group (big city group), and other postal code will be in other group.</li>\n<li>In my experience, i use below groups:<ul>\n<li>Postal_code with customer_size &gt; 1000: big_city</li>\n<li>Postal_code with customer_size &gt; 100 &amp; &lt;1000: medium city</li>\n<li>Postal_code with customer_size &gt; 10 &amp; &lt;100: small city</li>\n<li>Postal_code with customer_size &lt;10: very small city.</li></ul></li>\n</ul>\n<p>This is my code:</p>\n<pre><code>df = pd.read_csv(r\"customers.csv\")\ntest = df.postal_code.value_counts().reset_index()\ntest['bins'] = pd.cut(test['postal_code'], bins=[-1, 10, 100, 1e4], include_lowest=True)\n</code></pre>\n<p>Please upvote if it help you.<br>\nIf you have any other ideals, please share with me and others by comments </p>",
      "rawMarkdown": "In metadata of customer(`customers.csv`), postal code/address is an effective metadata, but as i have seen in most notebooks, every one is skipping this features. In practice, address is very effective for predicting sale amount: customers from big city tend to purchase more than other customers from countryside, and customers in same city may share same trending when purchasing. So i note my ideal about encoding method here for anyone who want to use postal_code/address as a feature in their model.\n\n# encoding method\n* We have so many distinct values (352899 different values) in postal_code features, so we can not use one-hot encoding. So we need to group it first.\n* By counting rows for each postal code, we see that so many customer share same postal code. As an example:  120303 customers (8% of all customers) have same postal code. So some postal_codes have so many customers (customers came from same city), but some postal_codes have only a customer. \n* So we have ideal here: Let group postal_code by customer size. All of postal_code had many customers will be in a group (big city group), and other postal code will be in other group.\n* In my experience, i use below groups:\n    * Postal_code with customer_size > 1000: big_city\n    * Postal_code with customer_size > 100 & <1000: medium city\n    * Postal_code with customer_size > 10 & <100: small city\n    * Postal_code with customer_size <10: very small city.\n\nThis is my code:\n\n    df = pd.read_csv(r\"customers.csv\")\n    test = df.postal_code.value_counts().reset_index()\n    test['bins'] = pd.cut(test['postal_code'], bins=[-1, 10, 100, 1e4], include_lowest=True)\n\nPlease upvote if it help you.\nIf you have any other ideals, please share with me and others by comments",
      "votes": null
    },
    {
      "id": "1700198",
      "postDate": "02/21/2022 18:22:59",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306488\" target=\"_blank\">this discussion</a> where he notes that the second biggest postal code group has only 261 customers.</p>\n<p>Seems that:</p>\n<ol>\n<li>Postal code groups are pretty small (probably not cities)</li>\n<li>Something unusual about that largest postal code</li>\n</ol>",
      "rawMarkdown": "See [this discussion](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306488) where he notes that the second biggest postal code group has only 261 customers.\n\nSeems that:\n1. Postal code groups are pretty small (probably not cities)\n2. Something unusual about that largest postal code",
      "votes": null
    },
    {
      "id": "1700509",
      "postDate": "02/22/2022 03:57:58",
      "content": "<p>i agree that the biggest group is unusual, but from viewing of data uploader, i don't think they need to group special small groups into a big group (the biggest one) after hashing. It may has different meaning from other groups (the main store/hall, or some addresses is inaccessible), but we can treat it as an outlier with a specific encoding. In any cases, i think it is still an useful feature.</p>",
      "rawMarkdown": "i agree that the biggest group is unusual, but from viewing of data uploader, i don't think they need to group special small groups into a big group (the biggest one) after hashing. It may has different meaning from other groups (the main store/hall, or some addresses is inaccessible), but we can treat it as an outlier with a specific encoding. In any cases, i think it is still an useful feature.",
      "votes": null
    },
    {
      "id": "1701600",
      "postDate": "02/22/2022 23:25:23",
      "content": "<p>Customers inside a postal code do not need to be a large set. Postal codes represent a small area and a small percentage of people inside that area may shop at H&amp;M. There could be various correlations with Postal Code.</p>",
      "rawMarkdown": "Customers inside a postal code do not need to be a large set. Postal codes represent a small area and a small percentage of people inside that area may shop at H&M. There could be various correlations with Postal Code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1700198,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "02/21/2022 18:22:59",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306488\" target=\"_blank\">this discussion</a> where he notes that the second biggest postal code group has only 261 customers.</p>\n<p>Seems that:</p>\n<ol>\n<li>Postal code groups are pretty small (probably not cities)</li>\n<li>Something unusual about that largest postal code</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1700509,
          "author_name": "astrung",
          "author_url": "",
          "post_date": "02/22/2022 03:57:58",
          "content": "<p>i agree that the biggest group is unusual, but from viewing of data uploader, i don't think they need to group special small groups into a big group (the biggest one) after hashing. It may has different meaning from other groups (the main store/hall, or some addresses is inaccessible), but we can treat it as an outlier with a specific encoding. In any cases, i think it is still an useful feature.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1701600,
      "author_name": "atulverma",
      "author_url": "",
      "post_date": "02/22/2022 23:25:23",
      "content": "<p>Customers inside a postal code do not need to be a large set. Postal codes represent a small area and a small percentage of people inside that area may shop at H&amp;M. There could be various correlations with Postal Code.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1699231": "In metadata of customer(`customers.csv`), postal code/address is an effective metadata, but as i have seen in most notebooks, every one is skipping this features. In practice, address is very effective for predicting sale amount: customers from big city tend to purchase more than other customers from countryside, and customers in same city may share same trending when purchasing. So i note my ideal about encoding method here for anyone who want to use postal_code/address as a feature in their model.\n\n# encoding method\n* We have so many distinct values (352899 different values) in postal_code features, so we can not use one-hot encoding. So we need to group it first.\n* By counting rows for each postal code, we see that so many customer share same postal code. As an example:  120303 customers (8% of all customers) have same postal code. So some postal_codes have so many customers (customers came from same city), but some postal_codes have only a customer. \n* So we have ideal here: Let group postal_code by customer size. All of postal_code had many customers will be in a group (big city group), and other postal code will be in other group.\n* In my experience, i use below groups:\n    * Postal_code with customer_size > 1000: big_city\n    * Postal_code with customer_size > 100 & <1000: medium city\n    * Postal_code with customer_size > 10 & <100: small city\n    * Postal_code with customer_size <10: very small city.\n\nThis is my code:\n\n    df = pd.read_csv(r\"customers.csv\")\n    test = df.postal_code.value_counts().reset_index()\n    test['bins'] = pd.cut(test['postal_code'], bins=[-1, 10, 100, 1e4], include_lowest=True)\n\nPlease upvote if it help you.\nIf you have any other ideals, please share with me and others by comments",
    "1700198": "See [this discussion](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306488) where he notes that the second biggest postal code group has only 261 customers.\n\nSeems that:\n1. Postal code groups are pretty small (probably not cities)\n2. Something unusual about that largest postal code",
    "1700509": "i agree that the biggest group is unusual, but from viewing of data uploader, i don't think they need to group special small groups into a big group (the biggest one) after hashing. It may has different meaning from other groups (the main store/hall, or some addresses is inaccessible), but we can treat it as an outlier with a specific encoding. In any cases, i think it is still an useful feature.",
    "1701600": "Customers inside a postal code do not need to be a large set. Postal codes represent a small area and a small percentage of people inside that area may shop at H&M. There could be various correlations with Postal Code."
  },
  "source": "meta"
}