{
  "id": 59876,
  "title": "Features around price-competitiveness of an ad.",
  "url": "/competitions/avito-demand-prediction/discussion/59876",
  "author_name": "",
  "post_date": "2018-06-28T00:57:35.379575300Z",
  "votes": 21,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Congratulations to our winners, and thanks to everyone for making this such a fun competition. I've spent a lot of time on these boards so I figured I'll <em>humbly</em> contribute something back. Here's an idea that worked for me --</p>\n\n<p>I wanted to featurize the intuition that people mostly click on ads that are price competitive. So I built thousands of clusters of \"similar\" ads and computed their mean/median prices. For example,  \"BMW 3 series 2017\" and \"good condition 2017 BMW 3\" are similar so tended to be assigned to the same cluster. Then, I computed price ratios to determine whether any given ad was over/underpriced in relation to its peer group (i.e. within city/category/etc.). The clusters themselves were created using document vectors (Doc2Vec through corruption) derived from the various text columns, and running kmeans on a GPU. Ultimately, I used a single 7-fold LGBM (0.2188 public LB) with 44 features plus the usual tfidfs. My next step: Learn how to stack ... lol.</p>",
  "messages": [
    {
      "id": "349252",
      "postDate": "06/28/2018 00:57:35",
      "content": "<p>Congratulations to our winners, and thanks to everyone for making this such a fun competition. I've spent a lot of time on these boards so I figured I'll <em>humbly</em> contribute something back. Here's an idea that worked for me --</p>\n\n<p>I wanted to featurize the intuition that people mostly click on ads that are price competitive. So I built thousands of clusters of \"similar\" ads and computed their mean/median prices. For example,  \"BMW 3 series 2017\" and \"good condition 2017 BMW 3\" are similar so tended to be assigned to the same cluster. Then, I computed price ratios to determine whether any given ad was over/underpriced in relation to its peer group (i.e. within city/category/etc.). The clusters themselves were created using document vectors (Doc2Vec through corruption) derived from the various text columns, and running kmeans on a GPU. Ultimately, I used a single 7-fold LGBM (0.2188 public LB) with 44 features plus the usual tfidfs. My next step: Learn how to stack ... lol.</p>",
      "rawMarkdown": "Congratulations to our winners, and thanks to everyone for making this such a fun competition. I've spent a lot of time on these boards so I figured I'll _humbly_ contribute something back. Here's an idea that worked for me --\n\nI wanted to featurize the intuition that people mostly click on ads that are price competitive. So I built thousands of clusters of \"similar\" ads and computed their mean/median prices. For example,  \"BMW 3 series 2017\" and \"good condition 2017 BMW 3\" are similar so tended to be assigned to the same cluster. Then, I computed price ratios to determine whether any given ad was over/underpriced in relation to its peer group (i.e. within city/category/etc.). The clusters themselves were created using document vectors (Doc2Vec through corruption) derived from the various text columns, and running kmeans on a GPU. Ultimately, I used a single 7-fold LGBM (0.2188 public LB) with 44 features plus the usual tfidfs. My next step: Learn how to stack ... lol.",
      "votes": null
    },
    {
      "id": "349263",
      "postDate": "06/28/2018 01:11:18",
      "content": "<p>Congrats on your silver medal!  I didn't have any single models &lt; 0.2190.  You are going to be unstoppable once you start stacking.</p>",
      "rawMarkdown": "Congrats on your silver medal!  I didn't have any single models &lt; 0.2190.  You are going to be unstoppable once you start stacking.",
      "votes": null
    },
    {
      "id": "349271",
      "postDate": "06/28/2018 01:19:19",
      "content": "<p>Thanks, @Harlan :-)</p>",
      "rawMarkdown": "Thanks, @Harlan :-)",
      "votes": null
    },
    {
      "id": "349357",
      "postDate": "06/28/2018 03:01:42",
      "content": "<p>Congrats eigenvector.... You would have been in a much better position if you had used an XGB and stack the results....</p>",
      "rawMarkdown": "Congrats eigenvector.... You would have been in a much better position if you had used an XGB and stack the results....",
      "votes": null
    },
    {
      "id": "349410",
      "postDate": "06/28/2018 04:46:03",
      "content": "<p>Thanks, Samrat. Congratulations to you too!</p>",
      "rawMarkdown": "Thanks, Samrat. Congratulations to you too!",
      "votes": null
    },
    {
      "id": "349425",
      "postDate": "06/28/2018 05:05:25",
      "content": "<p>Congratulations @Eigenvector and thanks for sharing. With a performance like this you chose your name rightly :-)</p>",
      "rawMarkdown": "Congratulations @Eigenvector and thanks for sharing. With a performance like this you chose your name rightly :-)",
      "votes": null
    },
    {
      "id": "350357",
      "postDate": "06/29/2018 15:57:26",
      "content": "<p>Cool, in the mid of the competition I also have similar idea to find similar ads but cannot go deeper finally. Really nice to see a workable implementation here! Will you share the code or more details later? Thanks in advance and congrats for the silver!</p>",
      "rawMarkdown": "Cool, in the mid of the competition I also have similar idea to find similar ads but cannot go deeper finally. Really nice to see a workable implementation here! Will you share the code or more details later? Thanks in advance and congrats for the silver!",
      "votes": null
    },
    {
      "id": "350391",
      "postDate": "06/29/2018 17:05:13",
      "content": "<p>Thanks for sharing your intuition, greatly pushes need for good feature engineering also Congrats on your well-deserved silver .  Can you elaborate more on Doc2Vec through \"corruption\", and where can I read more about this technique.</p>",
      "rawMarkdown": "Thanks for sharing your intuition, greatly pushes need for good feature engineering also Congrats on your well-deserved silver .  Can you elaborate more on Doc2Vec through \"corruption\", and where can I read more about this technique.",
      "votes": null
    },
    {
      "id": "350511",
      "postDate": "06/29/2018 21:42:54",
      "content": "<p>Thanks <a href=\"/bangda\">@bangda</a>. Congratulations to you for the silver too!  In terms of details, the strategy shares some elements with  the #4 solution by Joe Eddy. Of course, his solution is on steriods while I brought vitamins to the table :-) For Doc2VecC, you can find the paper + code here: <a href=\"https://github.com/mchen24/iclr2017\">https://github.com/mchen24/iclr2017</a>. This is a bit of a twist on simple averaging of word vectors. For K-means clustering, I used libKMCUDA to run on a GPU- <a href=\"https://github.com/src-d/kmcuda\">https://github.com/src-d/kmcuda</a>. For forming peer groups, I used various combinations of cluster-id + title + category + city + region etc., and filtered out what didn't stick. The features were constructed as follows: price-of-ad / median(price-of-all-ads-in-peer-group), count(*), etc. For determining cluster-ids, I tediously experimented with different settings of K, examining k-means elbows and using other heuristics to determine the right cutoffs. To the extent possible, I tried to form the right-sized clusters that only included similar ads in each cluster, no more or less. I ended up with 75000+ clusters across all categories.</p>",
      "rawMarkdown": "Thanks @bangda. Congratulations to you for the silver too!  In terms of details, the strategy shares some elements with  the #4 solution by Joe Eddy. Of course, his solution is on steriods while I brought vitamins to the table :-) For Doc2VecC, you can find the paper + code here: https://github.com/mchen24/iclr2017. This is a bit of a twist on simple averaging of word vectors. For K-means clustering, I used libKMCUDA to run on a GPU- https://github.com/src-d/kmcuda. For forming peer groups, I used various combinations of cluster-id + title + category + city + region etc., and filtered out what didn't stick. The features were constructed as follows: price-of-ad / median(price-of-all-ads-in-peer-group), count(*), etc. For determining cluster-ids, I tediously experimented with different settings of K, examining k-means elbows and using other heuristics to determine the right cutoffs. To the extent possible, I tried to form the right-sized clusters that only included similar ads in each cluster, no more or less. I ended up with 75000+ clusters across all categories.",
      "votes": null
    },
    {
      "id": "350512",
      "postDate": "06/29/2018 21:44:47",
      "content": "<p>Thanks <a href=\"/shivraj\">@shivraj</a>. Congratulations to you on your medal too! For Doc2VecC, you can find the original author's paper + code here: <a href=\"https://github.com/mchen24/iclr2017\">https://github.com/mchen24/iclr2017</a>. This is a bit of a twist on averaging of word vectors.</p>",
      "rawMarkdown": "Thanks @shivraj. Congratulations to you on your medal too! For Doc2VecC, you can find the original author's paper + code here: https://github.com/mchen24/iclr2017. This is a bit of a twist on averaging of word vectors.",
      "votes": null
    },
    {
      "id": "350513",
      "postDate": "06/29/2018 21:46:36",
      "content": "<p>Thanks YaGana! :-) Congratulations to you too on your medal, and on your impressive ranking within such a short amount of time.</p>",
      "rawMarkdown": "Thanks YaGana! :-) Congratulations to you too on your medal, and on your impressive ranking within such a short amount of time.",
      "votes": null
    },
    {
      "id": "350526",
      "postDate": "06/29/2018 22:18:25",
      "content": "<p>Thanks @Eigenvector. When I joined Kaggle in March 2017, I put in a lot of time to reach the top-500 by December. Since then finding enough time to devote to these competitions has been difficult. I really do enjoy Kaggling hence slogging along.</p>",
      "rawMarkdown": "Thanks @Eigenvector. When I joined Kaggle in March 2017, I put in a lot of time to reach the top-500 by December. Since then finding enough time to devote to these competitions has been difficult. I really do enjoy Kaggling hence slogging along.",
      "votes": null
    },
    {
      "id": "350528",
      "postDate": "06/29/2018 22:23:18",
      "content": "<p>@Eigenvector thanks for those links, I have not seen that paper before.</p>",
      "rawMarkdown": "Eigenvector thanks for those links, I have not seen that paper before.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 349263,
      "author_name": "shadowwarrior",
      "author_url": "",
      "post_date": "06/28/2018 01:11:18",
      "content": "<p>Congrats on your silver medal!  I didn't have any single models &lt; 0.2190.  You are going to be unstoppable once you start stacking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 349271,
          "author_name": "eigenvector",
          "author_url": "",
          "post_date": "06/28/2018 01:19:19",
          "content": "<p>Thanks, @Harlan :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349357,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "06/28/2018 03:01:42",
      "content": "<p>Congrats eigenvector.... You would have been in a much better position if you had used an XGB and stack the results....</p>",
      "votes": null,
      "replies": [
        {
          "id": 349410,
          "author_name": "eigenvector",
          "author_url": "",
          "post_date": "06/28/2018 04:46:03",
          "content": "<p>Thanks, Samrat. Congratulations to you too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349425,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/28/2018 05:05:25",
      "content": "<p>Congratulations @Eigenvector and thanks for sharing. With a performance like this you chose your name rightly :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 350513,
          "author_name": "eigenvector",
          "author_url": "",
          "post_date": "06/29/2018 21:46:36",
          "content": "<p>Thanks YaGana! :-) Congratulations to you too on your medal, and on your impressive ranking within such a short amount of time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350526,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "06/29/2018 22:18:25",
          "content": "<p>Thanks @Eigenvector. When I joined Kaggle in March 2017, I put in a lot of time to reach the top-500 by December. Since then finding enough time to devote to these competitions has been difficult. I really do enjoy Kaggling hence slogging along.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350357,
      "author_name": "bangdasun",
      "author_url": "",
      "post_date": "06/29/2018 15:57:26",
      "content": "<p>Cool, in the mid of the competition I also have similar idea to find similar ads but cannot go deeper finally. Really nice to see a workable implementation here! Will you share the code or more details later? Thanks in advance and congrats for the silver!</p>",
      "votes": null,
      "replies": [
        {
          "id": 350511,
          "author_name": "eigenvector",
          "author_url": "",
          "post_date": "06/29/2018 21:42:54",
          "content": "<p>Thanks <a href=\"/bangda\">@bangda</a>. Congratulations to you for the silver too!  In terms of details, the strategy shares some elements with  the #4 solution by Joe Eddy. Of course, his solution is on steriods while I brought vitamins to the table :-) For Doc2VecC, you can find the paper + code here: <a href=\"https://github.com/mchen24/iclr2017\">https://github.com/mchen24/iclr2017</a>. This is a bit of a twist on simple averaging of word vectors. For K-means clustering, I used libKMCUDA to run on a GPU- <a href=\"https://github.com/src-d/kmcuda\">https://github.com/src-d/kmcuda</a>. For forming peer groups, I used various combinations of cluster-id + title + category + city + region etc., and filtered out what didn't stick. The features were constructed as follows: price-of-ad / median(price-of-all-ads-in-peer-group), count(*), etc. For determining cluster-ids, I tediously experimented with different settings of K, examining k-means elbows and using other heuristics to determine the right cutoffs. To the extent possible, I tried to form the right-sized clusters that only included similar ads in each cluster, no more or less. I ended up with 75000+ clusters across all categories.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350528,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "06/29/2018 22:23:18",
          "content": "<p>@Eigenvector thanks for those links, I have not seen that paper before.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350391,
      "author_name": "shivrajp",
      "author_url": "",
      "post_date": "06/29/2018 17:05:13",
      "content": "<p>Thanks for sharing your intuition, greatly pushes need for good feature engineering also Congrats on your well-deserved silver .  Can you elaborate more on Doc2Vec through \"corruption\", and where can I read more about this technique.</p>",
      "votes": null,
      "replies": [
        {
          "id": 350512,
          "author_name": "eigenvector",
          "author_url": "",
          "post_date": "06/29/2018 21:44:47",
          "content": "<p>Thanks <a href=\"/shivraj\">@shivraj</a>. Congratulations to you on your medal too! For Doc2VecC, you can find the original author's paper + code here: <a href=\"https://github.com/mchen24/iclr2017\">https://github.com/mchen24/iclr2017</a>. This is a bit of a twist on averaging of word vectors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "349252": "Congratulations to our winners, and thanks to everyone for making this such a fun competition. I've spent a lot of time on these boards so I figured I'll _humbly_ contribute something back. Here's an idea that worked for me --\n\nI wanted to featurize the intuition that people mostly click on ads that are price competitive. So I built thousands of clusters of \"similar\" ads and computed their mean/median prices. For example,  \"BMW 3 series 2017\" and \"good condition 2017 BMW 3\" are similar so tended to be assigned to the same cluster. Then, I computed price ratios to determine whether any given ad was over/underpriced in relation to its peer group (i.e. within city/category/etc.). The clusters themselves were created using document vectors (Doc2Vec through corruption) derived from the various text columns, and running kmeans on a GPU. Ultimately, I used a single 7-fold LGBM (0.2188 public LB) with 44 features plus the usual tfidfs. My next step: Learn how to stack ... lol.",
    "349263": "Congrats on your silver medal!  I didn't have any single models &lt; 0.2190.  You are going to be unstoppable once you start stacking.",
    "349271": "Thanks, @Harlan :-)",
    "349357": "Congrats eigenvector.... You would have been in a much better position if you had used an XGB and stack the results....",
    "349410": "Thanks, Samrat. Congratulations to you too!",
    "349425": "Congratulations @Eigenvector and thanks for sharing. With a performance like this you chose your name rightly :-)",
    "350357": "Cool, in the mid of the competition I also have similar idea to find similar ads but cannot go deeper finally. Really nice to see a workable implementation here! Will you share the code or more details later? Thanks in advance and congrats for the silver!",
    "350391": "Thanks for sharing your intuition, greatly pushes need for good feature engineering also Congrats on your well-deserved silver .  Can you elaborate more on Doc2Vec through \"corruption\", and where can I read more about this technique.",
    "350511": "Thanks @bangda. Congratulations to you for the silver too!  In terms of details, the strategy shares some elements with  the #4 solution by Joe Eddy. Of course, his solution is on steriods while I brought vitamins to the table :-) For Doc2VecC, you can find the paper + code here: https://github.com/mchen24/iclr2017. This is a bit of a twist on simple averaging of word vectors. For K-means clustering, I used libKMCUDA to run on a GPU- https://github.com/src-d/kmcuda. For forming peer groups, I used various combinations of cluster-id + title + category + city + region etc., and filtered out what didn't stick. The features were constructed as follows: price-of-ad / median(price-of-all-ads-in-peer-group), count(*), etc. For determining cluster-ids, I tediously experimented with different settings of K, examining k-means elbows and using other heuristics to determine the right cutoffs. To the extent possible, I tried to form the right-sized clusters that only included similar ads in each cluster, no more or less. I ended up with 75000+ clusters across all categories.",
    "350512": "Thanks @shivraj. Congratulations to you on your medal too! For Doc2VecC, you can find the original author's paper + code here: https://github.com/mchen24/iclr2017. This is a bit of a twist on averaging of word vectors.",
    "350513": "Thanks YaGana! :-) Congratulations to you too on your medal, and on your impressive ranking within such a short amount of time.",
    "350526": "Thanks @Eigenvector. When I joined Kaggle in March 2017, I put in a lot of time to reach the top-500 by December. Since then finding enough time to devote to these competitions has been difficult. I really do enjoy Kaggling hence slogging along.",
    "350528": "Eigenvector thanks for those links, I have not seen that paper before."
  },
  "source": "meta"
}