{
  "id": 308555,
  "title": "How to get customer's \"sex\" feature",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308555",
  "author_name": "",
  "post_date": "2022-02-19T07:05:56.636406600Z",
  "votes": 45,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This is my first post to the Kaggle discussion.</p>\n<p>I think \"sex\" (male or female) is an important customer feature because<br>\nmale customer do not wear or buy the ladies fashion, and vice versa.<br>\nHowever, it seems there are no feature about sex in customer dataset.</p>\n<p>On the other hand, it seems article dataset include the information about man/ladies fashion.</p>\n<p>I'm wondering there are ways to know each customer's gender or not.</p>",
  "messages": [
    {
      "id": "1696886",
      "postDate": "02/19/2022 07:05:56",
      "content": "<p>This is my first post to the Kaggle discussion.</p>\n<p>I think \"sex\" (male or female) is an important customer feature because<br>\nmale customer do not wear or buy the ladies fashion, and vice versa.<br>\nHowever, it seems there are no feature about sex in customer dataset.</p>\n<p>On the other hand, it seems article dataset include the information about man/ladies fashion.</p>\n<p>I'm wondering there are ways to know each customer's gender or not.</p>",
      "rawMarkdown": "This is my first post to the Kaggle discussion.\n\nI think \"sex\" (male or female) is an important customer feature because\nmale customer do not wear or buy the ladies fashion, and vice versa.\nHowever, it seems there are no feature about sex in customer dataset.\n\nOn the other hand, it seems article dataset include the information about man/ladies fashion.\n\nI'm wondering there are ways to know each customer's gender or not.",
      "votes": null
    },
    {
      "id": "1697221",
      "postDate": "02/19/2022 13:08:34",
      "content": "<p>Of course, group all purchases into sections and in proportion you will see which section is in priority, respectively, it will be possible to assume gender.</p>",
      "rawMarkdown": "Of course, group all purchases into sections and in proportion you will see which section is in priority, respectively, it will be possible to assume gender.",
      "votes": null
    },
    {
      "id": "1697817",
      "postDate": "02/19/2022 21:26:57",
      "content": "<p>I explored this using the item column <code>index_group_name</code>. I mapped this column to a new column named <code>gender</code> with <code>mp = {'Ladieswear':1, 'Baby/Children':0.5, 'Menswear':0, 'Sport':0.5, 'Divided':0.5}</code> and computed <code>g = transactions.groupby('customer_id').gender.mean().reset_index()</code>. Then identified </p>\n<pre><code>female = g.loc[g.gender&gt;0.5].customer_id.values\nmale = g.loc[g.gender&lt;0.5].customer_id.values\n</code></pre>\n<p>I found 84.5% female, 4.7% male, and 10.9% unknown among the <code>1,371,980</code> unique customers</p>",
      "rawMarkdown": "I explored this using the item column `index_group_name`. I mapped this column to a new column named `gender` with `mp = {'Ladieswear':1, 'Baby/Children':0.5, 'Menswear':0, 'Sport':0.5, 'Divided':0.5}` and computed `g = transactions.groupby('customer_id').gender.mean().reset_index()`. Then identified \n\n    female = g.loc[g.gender>0.5].customer_id.values\n    male = g.loc[g.gender<0.5].customer_id.values\n\nI found 84.5% female, 4.7% male, and 10.9% unknown among the `1,371,980` unique customers",
      "votes": null
    },
    {
      "id": "1698843",
      "postDate": "02/20/2022 17:56:08",
      "content": "<p>Thanks for sharing the idea! Trying this, I found that the item column <code>section_name</code> also has suitable values: <code>Ladies H&amp;M Sport</code> [section no: 5] and <code>Men H&amp;M Sport</code> [22]. These are the subgroups of <code>index_group_name == 'Sport'</code>. It seems likely that we can use them.</p>",
      "rawMarkdown": "Thanks for sharing the idea! Trying this, I found that the item column `section_name` also has suitable values: `Ladies H&M Sport` [section no: 5] and `Men H&M Sport` [22]. These are the subgroups of `index_group_name == 'Sport'`. It seems likely that we can use them.",
      "votes": null
    },
    {
      "id": "1698848",
      "postDate": "02/20/2022 18:00:06",
      "content": "<p>Yes, that can help us. Good discovery.</p>",
      "rawMarkdown": "Yes, that can help us. Good discovery.",
      "votes": null
    },
    {
      "id": "1699949",
      "postDate": "02/21/2022 14:45:45",
      "content": "<p>Thank you for your advice!<br>\nIt's very helpful for me.</p>",
      "rawMarkdown": "Thank you for your advice!\nIt's very helpful for me.",
      "votes": null
    },
    {
      "id": "1705915",
      "postDate": "02/27/2022 01:08:25",
      "content": "<p>Doesn't that result seem skewed?</p>",
      "rawMarkdown": "Doesn't that result seem skewed?",
      "votes": null
    },
    {
      "id": "1735957",
      "postDate": "03/26/2022 18:46:22",
      "content": "<p>Not sure this feature will help a lot, because we can consider that a customer make orders for all the family, so he can buy for the kids, man or woman … Eventually a continuous variable instead of a boolean should reflect better the reality ?</p>",
      "rawMarkdown": "Not sure this feature will help a lot, because we can consider that a customer make orders for all the family, so he can buy for the kids, man or woman ... Eventually a continuous variable instead of a boolean should reflect better the reality ?",
      "votes": null
    },
    {
      "id": "1736481",
      "postDate": "03/27/2022 11:26:54",
      "content": "<p>HM target mainly women. So it’s skewed</p>",
      "rawMarkdown": "HM target mainly women. So it’s skewed",
      "votes": null
    },
    {
      "id": "1736482",
      "postDate": "03/27/2022 11:27:27",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nIt’s a nice work.thanks </p>",
      "rawMarkdown": "cdeotte \nIt’s a nice work.thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1697221,
      "author_name": "dmitryuarov",
      "author_url": "",
      "post_date": "02/19/2022 13:08:34",
      "content": "<p>Of course, group all purchases into sections and in proportion you will see which section is in priority, respectively, it will be possible to assume gender.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1697817,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2022 21:26:57",
      "content": "<p>I explored this using the item column <code>index_group_name</code>. I mapped this column to a new column named <code>gender</code> with <code>mp = {'Ladieswear':1, 'Baby/Children':0.5, 'Menswear':0, 'Sport':0.5, 'Divided':0.5}</code> and computed <code>g = transactions.groupby('customer_id').gender.mean().reset_index()</code>. Then identified </p>\n<pre><code>female = g.loc[g.gender&gt;0.5].customer_id.values\nmale = g.loc[g.gender&lt;0.5].customer_id.values\n</code></pre>\n<p>I found 84.5% female, 4.7% male, and 10.9% unknown among the <code>1,371,980</code> unique customers</p>",
      "votes": null,
      "replies": [
        {
          "id": 1698843,
          "author_name": "negoto",
          "author_url": "",
          "post_date": "02/20/2022 17:56:08",
          "content": "<p>Thanks for sharing the idea! Trying this, I found that the item column <code>section_name</code> also has suitable values: <code>Ladies H&amp;M Sport</code> [section no: 5] and <code>Men H&amp;M Sport</code> [22]. These are the subgroups of <code>index_group_name == 'Sport'</code>. It seems likely that we can use them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698848,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/20/2022 18:00:06",
          "content": "<p>Yes, that can help us. Good discovery.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1705915,
          "author_name": "atulverma",
          "author_url": "",
          "post_date": "02/27/2022 01:08:25",
          "content": "<p>Doesn't that result seem skewed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1736481,
          "author_name": "deepkun1995",
          "author_url": "",
          "post_date": "03/27/2022 11:26:54",
          "content": "<p>HM target mainly women. So it’s skewed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1736482,
          "author_name": "deepkun1995",
          "author_url": "",
          "post_date": "03/27/2022 11:27:27",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nIt’s a nice work.thanks </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1699949,
      "author_name": "tetsuro731",
      "author_url": "",
      "post_date": "02/21/2022 14:45:45",
      "content": "<p>Thank you for your advice!<br>\nIt's very helpful for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1735957,
      "author_name": "dsemmau",
      "author_url": "",
      "post_date": "03/26/2022 18:46:22",
      "content": "<p>Not sure this feature will help a lot, because we can consider that a customer make orders for all the family, so he can buy for the kids, man or woman … Eventually a continuous variable instead of a boolean should reflect better the reality ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1696886": "This is my first post to the Kaggle discussion.\n\nI think \"sex\" (male or female) is an important customer feature because\nmale customer do not wear or buy the ladies fashion, and vice versa.\nHowever, it seems there are no feature about sex in customer dataset.\n\nOn the other hand, it seems article dataset include the information about man/ladies fashion.\n\nI'm wondering there are ways to know each customer's gender or not.",
    "1697221": "Of course, group all purchases into sections and in proportion you will see which section is in priority, respectively, it will be possible to assume gender.",
    "1697817": "I explored this using the item column `index_group_name`. I mapped this column to a new column named `gender` with `mp = {'Ladieswear':1, 'Baby/Children':0.5, 'Menswear':0, 'Sport':0.5, 'Divided':0.5}` and computed `g = transactions.groupby('customer_id').gender.mean().reset_index()`. Then identified \n\n    female = g.loc[g.gender>0.5].customer_id.values\n    male = g.loc[g.gender<0.5].customer_id.values\n\nI found 84.5% female, 4.7% male, and 10.9% unknown among the `1,371,980` unique customers",
    "1698843": "Thanks for sharing the idea! Trying this, I found that the item column `section_name` also has suitable values: `Ladies H&M Sport` [section no: 5] and `Men H&M Sport` [22]. These are the subgroups of `index_group_name == 'Sport'`. It seems likely that we can use them.",
    "1698848": "Yes, that can help us. Good discovery.",
    "1699949": "Thank you for your advice!\nIt's very helpful for me.",
    "1705915": "Doesn't that result seem skewed?",
    "1735957": "Not sure this feature will help a lot, because we can consider that a customer make orders for all the family, so he can buy for the kids, man or woman ... Eventually a continuous variable instead of a boolean should reflect better the reality ?",
    "1736481": "HM target mainly women. So it’s skewed",
    "1736482": "cdeotte \nIt’s a nice work.thanks"
  },
  "source": "meta"
}