{
  "id": 57119,
  "title": "Top10 feature sharing",
  "url": "/competitions/avito-demand-prediction/discussion/57119",
  "author_name": "",
  "post_date": "2018-05-19T13:53:57.103072600Z",
  "votes": 31,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi, mates:</p>\n\n<p>I made a big decision for sharing some top10 features in my lightgbm model.\nI have tried target cluster for categorical feature, it improve my model , it plays real important feature (top 10) at lightgbm model. If you want to try, pseudo-code as following:</p>\n\n<pre><code>def add_target_cluster(daset, c,  target_c):\n  # prepare w2v sentences with target histories\n  gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n  hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n  gp_index = hist.index\n  sentences = [x.split(' ') for x in hist.values]\n  n_features = 500\n  w2v = Word2Vec(sentences=sentences, min_count=1, size=n_features)\n  w2v_feature = transeform_to_matrix(w2v, sentences)\n  # clustering\n  cluster_labels = any cluster method e.g K-Means\n  cluster_labels = pd.Series(cluster_labels, name=c + '_cluster', index=gp_index)\n  daset[c + '_cluster'] = daset[c].map(cluster_labels).fillna(-1).astype(int)\n</code></pre>\n\n<p>You can try embedding any categorical features in this competition.</p>\n\n<p>There should be a lot more to improve: e.g: change Word2Vec to Doc2Vec, try different target (price...).</p>\n\n<p>Finally, welcome to the discussion if you have any new ideas or ideas about above approach.</p>\n\n<p>Thank you for your time reading, and hope some help for you.</p>\n\n<p>Enjoy competition!</p>",
  "messages": [
    {
      "id": "330708",
      "postDate": "05/19/2018 13:53:57",
      "content": "<p>Hi, mates:</p>\n\n<p>I made a big decision for sharing some top10 features in my lightgbm model.\nI have tried target cluster for categorical feature, it improve my model , it plays real important feature (top 10) at lightgbm model. If you want to try, pseudo-code as following:</p>\n\n<pre><code>def add_target_cluster(daset, c,  target_c):\n  # prepare w2v sentences with target histories\n  gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n  hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n  gp_index = hist.index\n  sentences = [x.split(' ') for x in hist.values]\n  n_features = 500\n  w2v = Word2Vec(sentences=sentences, min_count=1, size=n_features)\n  w2v_feature = transeform_to_matrix(w2v, sentences)\n  # clustering\n  cluster_labels = any cluster method e.g K-Means\n  cluster_labels = pd.Series(cluster_labels, name=c + '_cluster', index=gp_index)\n  daset[c + '_cluster'] = daset[c].map(cluster_labels).fillna(-1).astype(int)\n</code></pre>\n\n<p>You can try embedding any categorical features in this competition.</p>\n\n<p>There should be a lot more to improve: e.g: change Word2Vec to Doc2Vec, try different target (price...).</p>\n\n<p>Finally, welcome to the discussion if you have any new ideas or ideas about above approach.</p>\n\n<p>Thank you for your time reading, and hope some help for you.</p>\n\n<p>Enjoy competition!</p>",
      "rawMarkdown": "Hi, mates:\n\nI made a big decision for sharing some top10 features in my lightgbm model.\nI have tried target cluster for categorical feature, it improve my model , it plays real important feature (top 10) at lightgbm model. If you want to try, pseudo-code as following:\n\n    def add_target_cluster(daset, c,  target_c):\n      # prepare w2v sentences with target histories\n      gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n      hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n      gp_index = hist.index\n      sentences = [x.split(' ') for x in hist.values]\n      n_features = 500\n      w2v = Word2Vec(sentences=sentences, min_count=1, size=n_features)\n      w2v_feature = transeform_to_matrix(w2v, sentences)\n      # clustering\n      cluster_labels = any cluster method e.g K-Means\n      cluster_labels = pd.Series(cluster_labels, name=c + '_cluster', index=gp_index)\n      daset[c + '_cluster'] = daset[c].map(cluster_labels).fillna(-1).astype(int)\n\nYou can try embedding any categorical features in this competition.\n\nThere should be a lot more to improve: e.g: change Word2Vec to Doc2Vec, try different target (price...).\n\nFinally, welcome to the discussion if you have any new ideas or ideas about above approach.\n\nThank you for your time reading, and hope some help for you.\n\nEnjoy competition!",
      "votes": null
    },
    {
      "id": "330840",
      "postDate": "05/19/2018 20:20:46",
      "content": "<p>Hi thanks for your sharing ! \nCan we have your fonction <code>transform_to_matrix</code> ?\nAnd last question what is <code>gkey</code> in your code ?</p>",
      "rawMarkdown": "Hi thanks for your sharing ! \nCan we have your fonction `transform_to_matrix` ?\nAnd last question what is `gkey` in your code ?",
      "votes": null
    },
    {
      "id": "331289",
      "postDate": "05/20/2018 21:23:56",
      "content": "<p>Does this perform better than mean encoding of target variable?</p>",
      "rawMarkdown": "Does this perform better than mean encoding of target variable?",
      "votes": null
    },
    {
      "id": "331460",
      "postDate": "05/21/2018 10:04:18",
      "content": "<p>It brings semantic information on your categorical variables, you can add it to your target encoding</p>",
      "rawMarkdown": "It brings semantic information on your categorical variables, you can add it to your target encoding",
      "votes": null
    },
    {
      "id": "332375",
      "postDate": "05/23/2018 04:26:07",
      "content": "<pre><code>`\ndef apply_w2v(sentences, model, num_features):\n  def _average_word_vectors(words, model, vocabulary, num_features):\n      feature_vector = np.zeros((num_features,), dtype=\"float64\")\n      n_words = 0.\n      for word in words:\n          if word in vocabulary:\n              n_words = n_words + 1.\n              feature_vector = np.add(feature_vector, model[word])\n\n      if n_words:\n          feature_vector = np.divide(feature_vector, n_words)\n      return feature_vector\n\n  vocab = set(model.wv.index2word)\n  feats = [_average_word_vectors(s, model, vocab, num_features) for s in sentences]\n return csr_matrix(np.array(feats))\n</code></pre>\n\n<p>`</p>",
      "rawMarkdown": "`\n    def apply_w2v(sentences, model, num_features):\n      def _average_word_vectors(words, model, vocabulary, num_features):\n          feature_vector = np.zeros((num_features,), dtype=\"float64\")\n          n_words = 0.\n          for word in words:\n              if word in vocabulary:\n                  n_words = n_words + 1.\n                  feature_vector = np.add(feature_vector, model[word])\n\n          if n_words:\n              feature_vector = np.divide(feature_vector, n_words)\n          return feature_vector\n\n      vocab = set(model.wv.index2word)\n      feats = [_average_word_vectors(s, model, vocab, num_features) for s in sentences]\n     return csr_matrix(np.array(feats))\n`",
      "votes": null
    },
    {
      "id": "332447",
      "postDate": "05/23/2018 07:08:32",
      "content": "<p>Thanks you ! </p>",
      "rawMarkdown": "Thanks you !",
      "votes": null
    },
    {
      "id": "332756",
      "postDate": "05/23/2018 17:03:02",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": null
    },
    {
      "id": "333813",
      "postDate": "05/25/2018 21:59:18",
      "content": "<p>Hello,</p>\n\n<p>Thanks a lot for sharing... So to make sure I understand your arguments correctly I should call it with something like ?\nadd_target_cluster(df, 'param_1', 'deal_probability')</p>\n\n<p>I have an error with this line :</p>\n\n<p>gp = df[[\"param_1\",\"deal_probability\"]].groupby(\"param_1\")[\"deal_probability\"]\nhist = gp.agg(lambda x: ' '.join(x).astype(str))</p>\n\n<blockquote>\n  <blockquote>\n    <p>TypeError: sequence item 0: expected str instance, float found</p>\n  </blockquote>\n</blockquote>\n\n<p>Am I doing something wrong ?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hello,\n\nThanks a lot for sharing... So to make sure I understand your arguments correctly I should call it with something like ?\nadd_target_cluster(df, 'param_1', 'deal_probability')\n\nI have an error with this line :\n\ngp = df[[\"param_1\",\"deal_probability\"]].groupby(\"param_1\")[\"deal_probability\"]\nhist = gp.agg(lambda x: ' '.join(x).astype(str))\n\n&gt;&gt; TypeError: sequence item 0: expected str instance, float found\n\nAm I doing something wrong ?\n\nThanks",
      "votes": null
    },
    {
      "id": "333838",
      "postDate": "05/25/2018 22:31:40",
      "content": "<pre><code>group.agg(lambda x: ' '.join(str(x)))\n</code></pre>",
      "rawMarkdown": "group.agg(lambda x: ' '.join(str(x)))",
      "votes": null
    },
    {
      "id": "334054",
      "postDate": "05/26/2018 12:18:03",
      "content": "<p>Thanks a lot Adil !</p>",
      "rawMarkdown": "Thanks a lot Adil !",
      "votes": null
    },
    {
      "id": "334257",
      "postDate": "05/26/2018 19:35:13",
      "content": "<p>What does this part do?</p>\n\n<pre><code>\n\n      gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n      hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n      gp_index = hist.index\n      sentences = [x.split(' ') for x in hist.values]\n</code></pre>",
      "rawMarkdown": "What does this part do?\n<pre><code>\n\n      gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n      hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n      gp_index = hist.index\n      sentences = [x.split(' ') for x in hist.values]\n</code></pre>",
      "votes": null
    },
    {
      "id": "334283",
      "postDate": "05/26/2018 21:50:01",
      "content": "<p>It’s an very interesting idea and I’d like to explore it bit more. Do you have a paper for this method? </p>\n\n<p>Thanks </p>",
      "rawMarkdown": "It’s an very interesting idea and I’d like to explore it bit more. Do you have a paper for this method? \n\nThanks",
      "votes": null
    },
    {
      "id": "334400",
      "postDate": "05/27/2018 08:01:59",
      "content": "<p>Hi, this part code is for building target history on on one category in time series model. You can treat it as a sentence or document which just build on on type of categorical feature's target values.</p>",
      "rawMarkdown": "Hi, this part code is for building target history on on one category in time series model. You can treat it as a sentence or document which just build on on type of categorical feature's target values.",
      "votes": null
    },
    {
      "id": "334401",
      "postDate": "05/27/2018 08:04:41",
      "content": "<p>I didn't have found any related paper using this method. However I always use it in my real work, and works well.</p>",
      "rawMarkdown": "I didn't have found any related paper using this method. However I always use it in my real work, and works well.",
      "votes": null
    },
    {
      "id": "334474",
      "postDate": "05/27/2018 14:00:56",
      "content": "<p>So for example, I choose to embed the 'city' category as a target value, I can choose 'region', 'parent_category_name', and 'category_name' as categories to embed?</p>",
      "rawMarkdown": "So for example, I choose to embed the 'city' category as a target value, I can choose 'region', 'parent_category_name', and 'category_name' as categories to embed?",
      "votes": null
    },
    {
      "id": "336799",
      "postDate": "06/01/2018 09:59:08",
      "content": "<p>You'd better chose price or deal_probability as a valuable target.\nCategory feature, you can choose city  image_top_1 or other many unique values categorical feature columns</p>",
      "rawMarkdown": "You'd better chose price or deal_probability as a valuable target.\nCategory feature, you can choose city  image_top_1 or other many unique values categorical feature columns",
      "votes": null
    },
    {
      "id": "336923",
      "postDate": "06/01/2018 14:42:25",
      "content": "<p>Thanks for this great idea. I'm wondering if you have any comments on clustering algorithms to use or alternatively if there's intelligent way to estimate optimal value for number of clusters (which is often required by K-means method). </p>",
      "rawMarkdown": "Thanks for this great idea. I'm wondering if you have any comments on clustering algorithms to use or alternatively if there's intelligent way to estimate optimal value for number of clusters (which is often required by K-means method).",
      "votes": null
    },
    {
      "id": "337652",
      "postDate": "06/03/2018 12:22:46",
      "content": "<p>Score silhouette can be interesting</p>",
      "rawMarkdown": "Score silhouette can be interesting",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 330840,
      "author_name": "adilztn",
      "author_url": "",
      "post_date": "05/19/2018 20:20:46",
      "content": "<p>Hi thanks for your sharing ! \nCan we have your fonction <code>transform_to_matrix</code> ?\nAnd last question what is <code>gkey</code> in your code ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 332375,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "05/23/2018 04:26:07",
          "content": "<pre><code>`\ndef apply_w2v(sentences, model, num_features):\n  def _average_word_vectors(words, model, vocabulary, num_features):\n      feature_vector = np.zeros((num_features,), dtype=\"float64\")\n      n_words = 0.\n      for word in words:\n          if word in vocabulary:\n              n_words = n_words + 1.\n              feature_vector = np.add(feature_vector, model[word])\n\n      if n_words:\n          feature_vector = np.divide(feature_vector, n_words)\n      return feature_vector\n\n  vocab = set(model.wv.index2word)\n  feats = [_average_word_vectors(s, model, vocab, num_features) for s in sentences]\n return csr_matrix(np.array(feats))\n</code></pre>\n\n<p>`</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332447,
          "author_name": "adilztn",
          "author_url": "",
          "post_date": "05/23/2018 07:08:32",
          "content": "<p>Thanks you ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331289,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "05/20/2018 21:23:56",
      "content": "<p>Does this perform better than mean encoding of target variable?</p>",
      "votes": null,
      "replies": [
        {
          "id": 331460,
          "author_name": "adilztn",
          "author_url": "",
          "post_date": "05/21/2018 10:04:18",
          "content": "<p>It brings semantic information on your categorical variables, you can add it to your target encoding</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332756,
      "author_name": "jmohitj",
      "author_url": "",
      "post_date": "05/23/2018 17:03:02",
      "content": "<p>thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 333813,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "05/25/2018 21:59:18",
      "content": "<p>Hello,</p>\n\n<p>Thanks a lot for sharing... So to make sure I understand your arguments correctly I should call it with something like ?\nadd_target_cluster(df, 'param_1', 'deal_probability')</p>\n\n<p>I have an error with this line :</p>\n\n<p>gp = df[[\"param_1\",\"deal_probability\"]].groupby(\"param_1\")[\"deal_probability\"]\nhist = gp.agg(lambda x: ' '.join(x).astype(str))</p>\n\n<blockquote>\n  <blockquote>\n    <p>TypeError: sequence item 0: expected str instance, float found</p>\n  </blockquote>\n</blockquote>\n\n<p>Am I doing something wrong ?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 333838,
          "author_name": "adilztn",
          "author_url": "",
          "post_date": "05/25/2018 22:31:40",
          "content": "<pre><code>group.agg(lambda x: ' '.join(str(x)))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334054,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "05/26/2018 12:18:03",
          "content": "<p>Thanks a lot Adil !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 334257,
      "author_name": "lintang",
      "author_url": "",
      "post_date": "05/26/2018 19:35:13",
      "content": "<p>What does this part do?</p>\n\n<pre><code>\n\n      gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n      hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n      gp_index = hist.index\n      sentences = [x.split(' ') for x in hist.values]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 334400,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "05/27/2018 08:01:59",
          "content": "<p>Hi, this part code is for building target history on on one category in time series model. You can treat it as a sentence or document which just build on on type of categorical feature's target values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334474,
          "author_name": "lintang",
          "author_url": "",
          "post_date": "05/27/2018 14:00:56",
          "content": "<p>So for example, I choose to embed the 'city' category as a target value, I can choose 'region', 'parent_category_name', and 'category_name' as categories to embed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 336799,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "06/01/2018 09:59:08",
          "content": "<p>You'd better chose price or deal_probability as a valuable target.\nCategory feature, you can choose city  image_top_1 or other many unique values categorical feature columns</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 334283,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "05/26/2018 21:50:01",
      "content": "<p>It’s an very interesting idea and I’d like to explore it bit more. Do you have a paper for this method? </p>\n\n<p>Thanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 334401,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "05/27/2018 08:04:41",
          "content": "<p>I didn't have found any related paper using this method. However I always use it in my real work, and works well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 336923,
      "author_name": "jyuan1986",
      "author_url": "",
      "post_date": "06/01/2018 14:42:25",
      "content": "<p>Thanks for this great idea. I'm wondering if you have any comments on clustering algorithms to use or alternatively if there's intelligent way to estimate optimal value for number of clusters (which is often required by K-means method). </p>",
      "votes": null,
      "replies": [
        {
          "id": 337652,
          "author_name": "adilztn",
          "author_url": "",
          "post_date": "06/03/2018 12:22:46",
          "content": "<p>Score silhouette can be interesting</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "330708": "Hi, mates:\n\nI made a big decision for sharing some top10 features in my lightgbm model.\nI have tried target cluster for categorical feature, it improve my model , it plays real important feature (top 10) at lightgbm model. If you want to try, pseudo-code as following:\n\n    def add_target_cluster(daset, c,  target_c):\n      # prepare w2v sentences with target histories\n      gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n      hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n      gp_index = hist.index\n      sentences = [x.split(' ') for x in hist.values]\n      n_features = 500\n      w2v = Word2Vec(sentences=sentences, min_count=1, size=n_features)\n      w2v_feature = transeform_to_matrix(w2v, sentences)\n      # clustering\n      cluster_labels = any cluster method e.g K-Means\n      cluster_labels = pd.Series(cluster_labels, name=c + '_cluster', index=gp_index)\n      daset[c + '_cluster'] = daset[c].map(cluster_labels).fillna(-1).astype(int)\n\nYou can try embedding any categorical features in this competition.\n\nThere should be a lot more to improve: e.g: change Word2Vec to Doc2Vec, try different target (price...).\n\nFinally, welcome to the discussion if you have any new ideas or ideas about above approach.\n\nThank you for your time reading, and hope some help for you.\n\nEnjoy competition!",
    "330840": "Hi thanks for your sharing ! \nCan we have your fonction `transform_to_matrix` ?\nAnd last question what is `gkey` in your code ?",
    "331289": "Does this perform better than mean encoding of target variable?",
    "331460": "It brings semantic information on your categorical variables, you can add it to your target encoding",
    "332375": "`\n    def apply_w2v(sentences, model, num_features):\n      def _average_word_vectors(words, model, vocabulary, num_features):\n          feature_vector = np.zeros((num_features,), dtype=\"float64\")\n          n_words = 0.\n          for word in words:\n              if word in vocabulary:\n                  n_words = n_words + 1.\n                  feature_vector = np.add(feature_vector, model[word])\n\n          if n_words:\n              feature_vector = np.divide(feature_vector, n_words)\n          return feature_vector\n\n      vocab = set(model.wv.index2word)\n      feats = [_average_word_vectors(s, model, vocab, num_features) for s in sentences]\n     return csr_matrix(np.array(feats))\n`",
    "332447": "Thanks you !",
    "332756": "thanks for sharing",
    "333813": "Hello,\n\nThanks a lot for sharing... So to make sure I understand your arguments correctly I should call it with something like ?\nadd_target_cluster(df, 'param_1', 'deal_probability')\n\nI have an error with this line :\n\ngp = df[[\"param_1\",\"deal_probability\"]].groupby(\"param_1\")[\"deal_probability\"]\nhist = gp.agg(lambda x: ' '.join(x).astype(str))\n\n&gt;&gt; TypeError: sequence item 0: expected str instance, float found\n\nAm I doing something wrong ?\n\nThanks",
    "333838": "group.agg(lambda x: ' '.join(str(x)))",
    "334054": "Thanks a lot Adil !",
    "334257": "What does this part do?\n<pre><code>\n\n      gp = daset.loc[tr_index, [c,target_c]].groupby(c)[target_c]\n      hist = gp.agg(lambda x: ' '.join(x).astype(str)))\n      gp_index = hist.index\n      sentences = [x.split(' ') for x in hist.values]\n</code></pre>",
    "334283": "It’s an very interesting idea and I’d like to explore it bit more. Do you have a paper for this method? \n\nThanks",
    "334400": "Hi, this part code is for building target history on on one category in time series model. You can treat it as a sentence or document which just build on on type of categorical feature's target values.",
    "334401": "I didn't have found any related paper using this method. However I always use it in my real work, and works well.",
    "334474": "So for example, I choose to embed the 'city' category as a target value, I can choose 'region', 'parent_category_name', and 'category_name' as categories to embed?",
    "336799": "You'd better chose price or deal_probability as a valuable target.\nCategory feature, you can choose city  image_top_1 or other many unique values categorical feature columns",
    "336923": "Thanks for this great idea. I'm wondering if you have any comments on clustering algorithms to use or alternatively if there's intelligent way to estimate optimal value for number of clusters (which is often required by K-means method).",
    "337652": "Score silhouette can be interesting"
  },
  "source": "meta"
}