{
  "id": 59886,
  "title": "The last gold solution",
  "url": "/competitions/avito-demand-prediction/writeups/win-a-gold-for-my-grandma-the-last-gold-solution",
  "author_name": "",
  "post_date": "2018-07-05T22:28:22.223Z",
  "votes": 66,
  "comment_count": 19,
  "views": 0,
  "content": "<p>First of all, thanks to Avito and Kaggle for providing this great competition. No leakage, very representative testing set, countless solutions. I believe every one enjoyed the game.</p>\n\n<p>Besides, thanks to all community contributors, like Dieter and Peter Hurford, who published helpful kernels and discussions.</p>\n\n<p>Features:\nMost of the ideas came from public kernels and discussions: tf-idf, CountVectorizor (char level), image meta features, self-trained W2V, ads time from Benjamin Minixhofer's aggregated features.</p>\n\n<p>What I did additionally is clustering ads titles based on sub-categories. For example:</p>\n\n<pre><code>sub_category = df[(df['category_name'] == 'something') &amp; (df['param_1'] == 'something') &amp; (df['param_2'] == 'something') &amp; (df['param_3'] == 'something')]\ntfidf = TfidfVectorizor()\ntfidf_vec = tfidf.fit_transform(sub_category['title'])\nkmean_cluster = KMmean()\nsub_category['title_cluster'] = kmean_cluster.fit_predict(tfidf_vec)\n</code></pre>\n\n<p>I did such calculation for all sub-catgories. After doing this, in the subcategory of iphone, all iphone 5s 32gb were in one group and iphone 7 128 GB were in another group. Then it was a good time to compare the prices and other features. So I groupby this 'title_cluster' and calculated aggregated features like price rank (which gave a big improvement), mean price, count and so on.</p>\n\n<p>I also used regex to manually extract numbers out of the title for properties (area, room numbers, floor) and automobile (vehicle years) ads. Then calculated features such as price per room, price per sqm.</p>\n\n<p>Other than that, since traditional ML models such as LightGBM are not good at handling unstructured data like text and images, I used NN (biLSTM for texts and a simple 4 layers CNN for image) to generate vector representations. Image vectors directly came from a NN that were used to predict the deal_probability. But when I tried the same strategy for texts, the generated text vector dramatically made my LightGBM overfitted. Then I tried another method: only using title and description W2V in a biLSTM NN model with MSE loss and predicting everything else: price, item_seq_number, city, region, user_id, parent_category_name, category_name, param_1, param_2, param_3. All categorical features are onehot encoded after removing low frequent entities. Therefore, X = title W2V + description W2V. y = a few hundred columns table.</p>\n\n<p>Models:\nThree layer of stacking. first layer: 11 Lightgbms, 6NNs. second layer: 3 LightGBM, 2 Ridge. third layer is just a Ridge and a LighGBM with linear average. I am a newbie of NN (my oof NN scores were very unstable) and hope to learn more NN strategies from other teams.</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "349311",
      "postDate": "06/28/2018 02:17:21",
      "content": "<p>First of all, thanks to Avito and Kaggle for providing this great competition. No leakage, very representative testing set, countless solutions. I believe every one enjoyed the game.</p>\n\n<p>Besides, thanks to all community contributors, like Dieter and Peter Hurford, who published helpful kernels and discussions.</p>\n\n<p>Features:\nMost of the ideas came from public kernels and discussions: tf-idf, CountVectorizor (char level), image meta features, self-trained W2V, ads time from Benjamin Minixhofer's aggregated features.</p>\n\n<p>What I did additionally is clustering ads titles based on sub-categories. For example:</p>\n\n<pre><code>sub_category = df[(df['category_name'] == 'something') &amp; (df['param_1'] == 'something') &amp; (df['param_2'] == 'something') &amp; (df['param_3'] == 'something')]\ntfidf = TfidfVectorizor()\ntfidf_vec = tfidf.fit_transform(sub_category['title'])\nkmean_cluster = KMmean()\nsub_category['title_cluster'] = kmean_cluster.fit_predict(tfidf_vec)\n</code></pre>\n\n<p>I did such calculation for all sub-catgories. After doing this, in the subcategory of iphone, all iphone 5s 32gb were in one group and iphone 7 128 GB were in another group. Then it was a good time to compare the prices and other features. So I groupby this 'title_cluster' and calculated aggregated features like price rank (which gave a big improvement), mean price, count and so on.</p>\n\n<p>I also used regex to manually extract numbers out of the title for properties (area, room numbers, floor) and automobile (vehicle years) ads. Then calculated features such as price per room, price per sqm.</p>\n\n<p>Other than that, since traditional ML models such as LightGBM are not good at handling unstructured data like text and images, I used NN (biLSTM for texts and a simple 4 layers CNN for image) to generate vector representations. Image vectors directly came from a NN that were used to predict the deal_probability. But when I tried the same strategy for texts, the generated text vector dramatically made my LightGBM overfitted. Then I tried another method: only using title and description W2V in a biLSTM NN model with MSE loss and predicting everything else: price, item_seq_number, city, region, user_id, parent_category_name, category_name, param_1, param_2, param_3. All categorical features are onehot encoded after removing low frequent entities. Therefore, X = title W2V + description W2V. y = a few hundred columns table.</p>\n\n<p>Models:\nThree layer of stacking. first layer: 11 Lightgbms, 6NNs. second layer: 3 LightGBM, 2 Ridge. third layer is just a Ridge and a LighGBM with linear average. I am a newbie of NN (my oof NN scores were very unstable) and hope to learn more NN strategies from other teams.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "First of all, thanks to Avito and Kaggle for providing this great competition. No leakage, very representative testing set, countless solutions. I believe every one enjoyed the game.\n\nBesides, thanks to all community contributors, like Dieter and Peter Hurford, who published helpful kernels and discussions.\n\nFeatures:\nMost of the ideas came from public kernels and discussions: tf-idf, CountVectorizor (char level), image meta features, self-trained W2V, ads time from Benjamin Minixhofer's aggregated features.\n\nWhat I did additionally is clustering ads titles based on sub-categories. For example:\n\n    sub_category = df[(df['category_name'] == 'something') &amp; (df['param_1'] == 'something') &amp; (df['param_2'] == 'something') &amp; (df['param_3'] == 'something')]\n    tfidf = TfidfVectorizor()\n    tfidf_vec = tfidf.fit_transform(sub_category['title'])\n    kmean_cluster = KMmean()\n    sub_category['title_cluster'] = kmean_cluster.fit_predict(tfidf_vec)\n\nI did such calculation for all sub-catgories. After doing this, in the subcategory of iphone, all iphone 5s 32gb were in one group and iphone 7 128 GB were in another group. Then it was a good time to compare the prices and other features. So I groupby this 'title_cluster' and calculated aggregated features like price rank (which gave a big improvement), mean price, count and so on.\n\nI also used regex to manually extract numbers out of the title for properties (area, room numbers, floor) and automobile (vehicle years) ads. Then calculated features such as price per room, price per sqm.\n\nOther than that, since traditional ML models such as LightGBM are not good at handling unstructured data like text and images, I used NN (biLSTM for texts and a simple 4 layers CNN for image) to generate vector representations. Image vectors directly came from a NN that were used to predict the deal_probability. But when I tried the same strategy for texts, the generated text vector dramatically made my LightGBM overfitted. Then I tried another method: only using title and description W2V in a biLSTM NN model with MSE loss and predicting everything else: price, item_seq_number, city, region, user_id, parent_category_name, category_name, param_1, param_2, param_3. All categorical features are onehot encoded after removing low frequent entities. Therefore, X = title W2V + description W2V. y = a few hundred columns table.\n\nModels:\nThree layer of stacking. first layer: 11 Lightgbms, 6NNs. second layer: 3 LightGBM, 2 Ridge. third layer is just a Ridge and a LighGBM with linear average. I am a newbie of NN (my oof NN scores were very unstable) and hope to learn more NN strategies from other teams.\n\nThanks",
      "votes": null
    },
    {
      "id": "349315",
      "postDate": "06/28/2018 02:23:54",
      "content": "<p>Congratulations for winning a solo gold medal in this competition!!!</p>",
      "rawMarkdown": "Congratulations for winning a solo gold medal in this competition!!!",
      "votes": null
    },
    {
      "id": "349320",
      "postDate": "06/28/2018 02:27:23",
      "content": "<p>Congratulations for solo gold medal ! Your team name is interesting at the last week of competition : ) </p>",
      "rawMarkdown": "Congratulations for solo gold medal ! Your team name is interesting at the last week of competition : )",
      "votes": null
    },
    {
      "id": "349324",
      "postDate": "06/28/2018 02:29:12",
      "content": "<p>Bummer for us to lose out on gold, but I'm really glad you were able to solo gold. That impressive accomplishment makes me feel better about our loss. I was secretly rooting for you throughout your quest up the LB.</p>",
      "rawMarkdown": "Bummer for us to lose out on gold, but I'm really glad you were able to solo gold. That impressive accomplishment makes me feel better about our loss. I was secretly rooting for you throughout your quest up the LB.",
      "votes": null
    },
    {
      "id": "349325",
      "postDate": "06/28/2018 02:29:17",
      "content": "<p>Congrats for your Gold... So, finally you showed everyone that solo gold is not very difficult :-)</p>",
      "rawMarkdown": "Congrats for your Gold... So, finally you showed everyone that solo gold is not very difficult :-)",
      "votes": null
    },
    {
      "id": "349333",
      "postDate": "06/28/2018 02:36:12",
      "content": "<p>Congrats for your impressive solo gold ! Your Grandma will be proud :)</p>",
      "rawMarkdown": "Congrats for your impressive solo gold ! Your Grandma will be proud :)",
      "votes": null
    },
    {
      "id": "349337",
      "postDate": "06/28/2018 02:38:52",
      "content": "<p>Congratulations for the solo gold medal! </p>",
      "rawMarkdown": "Congratulations for the solo gold medal!",
      "votes": null
    },
    {
      "id": "349365",
      "postDate": "06/28/2018 03:18:34",
      "content": "<p>Ayyy! Congrats on the solo gold. Truly an impressive feat 🍻</p>",
      "rawMarkdown": "Ayyy! Congrats on the solo gold. Truly an impressive feat 🍻",
      "votes": null
    },
    {
      "id": "349422",
      "postDate": "06/28/2018 04:59:33",
      "content": "<p>Congratulations @Weber on an amazing solo gold. Your solution is inspiring as well.</p>",
      "rawMarkdown": "Congratulations @Weber on an amazing solo gold. Your solution is inspiring as well.",
      "votes": null
    },
    {
      "id": "349438",
      "postDate": "06/28/2018 05:46:01",
      "content": "<p>Congratulations for your solo gold medal.  Great job!</p>",
      "rawMarkdown": "Congratulations for your solo gold medal.  Great job!",
      "votes": null
    },
    {
      "id": "349471",
      "postDate": "06/28/2018 06:56:31",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "349475",
      "postDate": "06/28/2018 07:03:23",
      "content": "<p>\"Win a gold for my grandma\", congratz. Hands down, it's the best name among all teams.\nYour nanny must be so happy.</p>",
      "rawMarkdown": "\"Win a gold for my grandma\", congratz. Hands down, it's the best name among all teams.\nYour nanny must be so happy.",
      "votes": null
    },
    {
      "id": "349842",
      "postDate": "06/28/2018 18:07:57",
      "content": "<p>Thanks! It's effort and luck making it easy.</p>",
      "rawMarkdown": "Thanks! It's effort and luck making it easy.",
      "votes": null
    },
    {
      "id": "349843",
      "postDate": "06/28/2018 18:08:34",
      "content": "<p>Haha, thanks. Just had something fun in the game</p>",
      "rawMarkdown": "Haha, thanks. Just had something fun in the game",
      "votes": null
    },
    {
      "id": "349844",
      "postDate": "06/28/2018 18:10:11",
      "content": "<p>Your team also did a great job. I was just a little bit more lucky at the end. You will be the lucky guy next time I believe.</p>",
      "rawMarkdown": "Your team also did a great job. I was just a little bit more lucky at the end. You will be the lucky guy next time I believe.",
      "votes": null
    },
    {
      "id": "349937",
      "postDate": "06/28/2018 22:14:22",
      "content": "<p>@Peter, what were the image feature you used - I think you mentioned earlier you got a .001 or .002 uplift from them ? Well done on your position, would love to see a write up</p>",
      "rawMarkdown": "Peter, what were the image feature you used - I think you mentioned earlier you got a .001 or .002 uplift from them ? Well done on your position, would love to see a write up",
      "votes": null
    },
    {
      "id": "349938",
      "postDate": "06/28/2018 22:31:54",
      "content": "<p>@Darragh, yeah the final boost from images was ~0.0015 on LightGBM. They also seemed to help our NNs significantly too. We're working on a write-up -- we're just not as fast as everyone else, haha.</p>",
      "rawMarkdown": "Darragh, yeah the final boost from images was ~0.0015 on LightGBM. They also seemed to help our NNs significantly too. We're working on a write-up -- we're just not as fast as everyone else, haha.",
      "votes": null
    },
    {
      "id": "349940",
      "postDate": "06/28/2018 22:36:36",
      "content": "<p>Take your time and make it good :) looking forward to it, good job!</p>",
      "rawMarkdown": "Take your time and make it good :) looking forward to it, good job!",
      "votes": null
    },
    {
      "id": "349985",
      "postDate": "06/29/2018 01:30:08",
      "content": "<p>I'm so glad you got a solo gold that was \"so hard to get\"! :)</p>",
      "rawMarkdown": "I'm so glad you got a solo gold that was \"so hard to get\"! :)",
      "votes": null
    },
    {
      "id": "421546",
      "postDate": "11/15/2018 05:37:57",
      "content": "<p>学长好强！！学习学习！</p>",
      "rawMarkdown": "学长好强！！学习学习！",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 349315,
      "author_name": "cczaixian",
      "author_url": "",
      "post_date": "06/28/2018 02:23:54",
      "content": "<p>Congratulations for winning a solo gold medal in this competition!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349320,
      "author_name": "mzr2017",
      "author_url": "",
      "post_date": "06/28/2018 02:27:23",
      "content": "<p>Congratulations for solo gold medal ! Your team name is interesting at the last week of competition : ) </p>",
      "votes": null,
      "replies": [
        {
          "id": 349843,
          "author_name": "wenbozhao",
          "author_url": "",
          "post_date": "06/28/2018 18:08:34",
          "content": "<p>Haha, thanks. Just had something fun in the game</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349324,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "06/28/2018 02:29:12",
      "content": "<p>Bummer for us to lose out on gold, but I'm really glad you were able to solo gold. That impressive accomplishment makes me feel better about our loss. I was secretly rooting for you throughout your quest up the LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 349844,
          "author_name": "wenbozhao",
          "author_url": "",
          "post_date": "06/28/2018 18:10:11",
          "content": "<p>Your team also did a great job. I was just a little bit more lucky at the end. You will be the lucky guy next time I believe.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349937,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "06/28/2018 22:14:22",
          "content": "<p>@Peter, what were the image feature you used - I think you mentioned earlier you got a .001 or .002 uplift from them ? Well done on your position, would love to see a write up</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349938,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "06/28/2018 22:31:54",
          "content": "<p>@Darragh, yeah the final boost from images was ~0.0015 on LightGBM. They also seemed to help our NNs significantly too. We're working on a write-up -- we're just not as fast as everyone else, haha.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349940,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "06/28/2018 22:36:36",
          "content": "<p>Take your time and make it good :) looking forward to it, good job!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349325,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "06/28/2018 02:29:17",
      "content": "<p>Congrats for your Gold... So, finally you showed everyone that solo gold is not very difficult :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 349842,
          "author_name": "wenbozhao",
          "author_url": "",
          "post_date": "06/28/2018 18:07:57",
          "content": "<p>Thanks! It's effort and luck making it easy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349333,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/28/2018 02:36:12",
      "content": "<p>Congrats for your impressive solo gold ! Your Grandma will be proud :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349337,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "06/28/2018 02:38:52",
      "content": "<p>Congratulations for the solo gold medal! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349365,
      "author_name": "yk1598",
      "author_url": "",
      "post_date": "06/28/2018 03:18:34",
      "content": "<p>Ayyy! Congrats on the solo gold. Truly an impressive feat 🍻</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349422,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/28/2018 04:59:33",
      "content": "<p>Congratulations @Weber on an amazing solo gold. Your solution is inspiring as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349438,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "06/28/2018 05:46:01",
      "content": "<p>Congratulations for your solo gold medal.  Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349471,
      "author_name": "greener",
      "author_url": "",
      "post_date": "06/28/2018 06:56:31",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349475,
      "author_name": "tmhung",
      "author_url": "",
      "post_date": "06/28/2018 07:03:23",
      "content": "<p>\"Win a gold for my grandma\", congratz. Hands down, it's the best name among all teams.\nYour nanny must be so happy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349985,
      "author_name": "u39kun",
      "author_url": "",
      "post_date": "06/29/2018 01:30:08",
      "content": "<p>I'm so glad you got a solo gold that was \"so hard to get\"! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 421546,
      "author_name": "louishgy",
      "author_url": "",
      "post_date": "11/15/2018 05:37:57",
      "content": "<p>学长好强！！学习学习！</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349311": "First of all, thanks to Avito and Kaggle for providing this great competition. No leakage, very representative testing set, countless solutions. I believe every one enjoyed the game.\n\nBesides, thanks to all community contributors, like Dieter and Peter Hurford, who published helpful kernels and discussions.\n\nFeatures:\nMost of the ideas came from public kernels and discussions: tf-idf, CountVectorizor (char level), image meta features, self-trained W2V, ads time from Benjamin Minixhofer's aggregated features.\n\nWhat I did additionally is clustering ads titles based on sub-categories. For example:\n\n    sub_category = df[(df['category_name'] == 'something') &amp; (df['param_1'] == 'something') &amp; (df['param_2'] == 'something') &amp; (df['param_3'] == 'something')]\n    tfidf = TfidfVectorizor()\n    tfidf_vec = tfidf.fit_transform(sub_category['title'])\n    kmean_cluster = KMmean()\n    sub_category['title_cluster'] = kmean_cluster.fit_predict(tfidf_vec)\n\nI did such calculation for all sub-catgories. After doing this, in the subcategory of iphone, all iphone 5s 32gb were in one group and iphone 7 128 GB were in another group. Then it was a good time to compare the prices and other features. So I groupby this 'title_cluster' and calculated aggregated features like price rank (which gave a big improvement), mean price, count and so on.\n\nI also used regex to manually extract numbers out of the title for properties (area, room numbers, floor) and automobile (vehicle years) ads. Then calculated features such as price per room, price per sqm.\n\nOther than that, since traditional ML models such as LightGBM are not good at handling unstructured data like text and images, I used NN (biLSTM for texts and a simple 4 layers CNN for image) to generate vector representations. Image vectors directly came from a NN that were used to predict the deal_probability. But when I tried the same strategy for texts, the generated text vector dramatically made my LightGBM overfitted. Then I tried another method: only using title and description W2V in a biLSTM NN model with MSE loss and predicting everything else: price, item_seq_number, city, region, user_id, parent_category_name, category_name, param_1, param_2, param_3. All categorical features are onehot encoded after removing low frequent entities. Therefore, X = title W2V + description W2V. y = a few hundred columns table.\n\nModels:\nThree layer of stacking. first layer: 11 Lightgbms, 6NNs. second layer: 3 LightGBM, 2 Ridge. third layer is just a Ridge and a LighGBM with linear average. I am a newbie of NN (my oof NN scores were very unstable) and hope to learn more NN strategies from other teams.\n\nThanks",
    "349315": "Congratulations for winning a solo gold medal in this competition!!!",
    "349320": "Congratulations for solo gold medal ! Your team name is interesting at the last week of competition : )",
    "349324": "Bummer for us to lose out on gold, but I'm really glad you were able to solo gold. That impressive accomplishment makes me feel better about our loss. I was secretly rooting for you throughout your quest up the LB.",
    "349325": "Congrats for your Gold... So, finally you showed everyone that solo gold is not very difficult :-)",
    "349333": "Congrats for your impressive solo gold ! Your Grandma will be proud :)",
    "349337": "Congratulations for the solo gold medal!",
    "349365": "Ayyy! Congrats on the solo gold. Truly an impressive feat 🍻",
    "349422": "Congratulations @Weber on an amazing solo gold. Your solution is inspiring as well.",
    "349438": "Congratulations for your solo gold medal.  Great job!",
    "349471": "Congratulations!",
    "349475": "\"Win a gold for my grandma\", congratz. Hands down, it's the best name among all teams.\nYour nanny must be so happy.",
    "349842": "Thanks! It's effort and luck making it easy.",
    "349843": "Haha, thanks. Just had something fun in the game",
    "349844": "Your team also did a great job. I was just a little bit more lucky at the end. You will be the lucky guy next time I believe.",
    "349937": "Peter, what were the image feature you used - I think you mentioned earlier you got a .001 or .002 uplift from them ? Well done on your position, would love to see a write up",
    "349938": "Darragh, yeah the final boost from images was ~0.0015 on LightGBM. They also seemed to help our NNs significantly too. We're working on a write-up -- we're just not as fast as everyone else, haha.",
    "349940": "Take your time and make it good :) looking forward to it, good job!",
    "349985": "I'm so glad you got a solo gold that was \"so hard to get\"! :)",
    "421546": "学长好强！！学习学习！"
  },
  "source": "meta"
}