{
  "id": 324070,
  "title": "1st place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070",
  "author_name": "senkin13",
  "post_date": "2022-05-10T02:08:03.127000",
  "votes": 430,
  "comment_count": 222,
  "views": 0,
  "content": "<p>Thanks to H&amp;M and Kaggle Team giving us such a wonderful competition.Thanks my team mate <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> finally we got 1st place.This competition has no leak,stable cv-lb correlation,and infinite possibilities to improve, we really enjoyed exploring the accuracy boundaries of  infinite possibilities. </p>\n<h1>Overview</h1>\n<p>The most interesting part of this competition is we need to generate train and test by ourself, the candinate generation strategy is the key to beyond limitation of accuracy ,good feature engineering or modeling could be close to the limitation .Our solution is using various retrieval strategies + feature engineering + GBDT, seems simple but powerful.</p>\n<p>We mainly generate recent popular items because fashion changing fast and has seasonality,tried to add cold start items but they will never be ranked to top12 due to lacking of interaction information.<br>\nuser and item interaction information are always the most important of recommendation problem,the features we created are almost interaction features,image and text features didn't help but should be useful for cold start problem.</p>\n<p>Almost 50% users have no transactions in recent 3months,so we created many cumulative features for them,and last week,last month,last season features for active users.</p>\n<p>We use 6 weeks data as train,last week as valid,retrieve 100 caninates for each user,it has stable cv-lb correlation.We focus on improving single lightgbm model to the last week, cv is 0.0430 and lb is 0.0362.At last week I have to rent gcp's big memory server and vast.ai's gpu server to run bigger models to get higher accuracy.</p>\n<p><a href=\"https://postimg.cc/xXLGHd2r\" target=\"_blank\"><img src=\"https://i.postimg.cc/bwBm7GVv/overview.png\" alt=\"overview.png\"></a></p>\n<h1>Retrieval(aka: candidate generation | recall)</h1>\n<p>At early stage,we focus on increasing the hit number of retrieved 100 candidates,try various strategies to cover more positive sample. </p>\n<table>\n<thead>\n<tr>\n<th>WeekNo</th>\n<th>HitNum@100</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2020-09-16</td>\n<td>39142</td>\n</tr>\n<tr>\n<td>2020-09-09</td>\n<td>38427</td>\n</tr>\n<tr>\n<td>2020-09-02</td>\n<td>41019</td>\n</tr>\n</tbody>\n</table>\n<h1>Ranking</h1>\n<ul>\n<li><p><strong>Feature Engineering</strong><br>\nbasiclly,features are created base on retrieval strategies,create user-item interaction for repurchase,create collaborative filtering score for itemcf,create similarity for embedding retrieval,create item count for popularity……</p>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Count</td>\n<td>user-item, user-category of last week/month/season/same week of last year/all, time weighted count…</td>\n</tr>\n<tr>\n<td>Time</td>\n<td>first,last days of tranactions…</td>\n</tr>\n<tr>\n<td>Mean/Max/Min</td>\n<td>aggregation of age,price,sales_channel_id…</td>\n</tr>\n<tr>\n<td>Difference/Ratio</td>\n<td>difference between age and mean age of who purchased item, ratio of one user's purchased item count and the item's count</td>\n</tr>\n<tr>\n<td>Similarity</td>\n<td>collaborative filtering score of  item2item, cosine similarity of item2item(word2vec), cosine similarity of user2item(ProNE)</td>\n</tr>\n</tbody>\n</table></li>\n<li><p><strong>Downsampling</strong><br>\nIf we retrieve 100~500 candidates for each user,the negative samples number is very big,so negative downsampling is must,we found 1 million ~ 2 million negative samples for each week has better performace .<br>\n<code>neg_samples = 1000000;seed = 42\ntrain[train['label']&gt;0].append(train[train['label']==0].sample(neg_samples, random_state=seed))\n</code></p></li>\n<li><p><strong>Model</strong><br>\nOur best single model is a lightgbm(cv:0.0441,lb:0.0367),finally we trained 5 lightgbm classifier and 7 catboost classifier for ensemble(lb:0.0371).catboost's lb score is much worse than lightgbm,lightgbm has very stable cv-lb correlation.<br>\n<a href=\"https://postimg.cc/ZBmJRVLt\" target=\"_blank\"><img src=\"https://i.postimg.cc/V6rvKhr6/cv.png\" alt=\"cv.png\"></a></p></li>\n</ul>\n<h1>Optimization</h1>\n<ol>\n<li>we use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.</li>\n<li>transform all the categorical features(including two way) to label encoding,use reduce_mem_usage,</li>\n<li>create a <strong>feature store</strong>,save intermediate features files to dictionary , final features to feather,exsiting features will not be create again</li>\n<li>split all the users to 28 group,inference simultaneously with multiple servers.</li>\n</ol>\n<h1>Machine</h1>\n<p>I have a desktop with 128G RAM,64 vcore CPU,TITAN RTX GPU.It's enough to run most of our models to win.My teammate has a desktop with 64G RAM.At last week ,we use 300G RAM gcp instances.</p>",
  "messages": [
    {
      "id": 1782928,
      "postDate": "2022-05-10T02:08:03.127Z",
      "content": "<p>Thanks to H&amp;M and Kaggle Team giving us such a wonderful competition.Thanks my team mate <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> finally we got 1st place.This competition has no leak,stable cv-lb correlation,and infinite possibilities to improve, we really enjoyed exploring the accuracy boundaries of  infinite possibilities. </p>\n<h1>Overview</h1>\n<p>The most interesting part of this competition is we need to generate train and test by ourself, the candinate generation strategy is the key to beyond limitation of accuracy ,good feature engineering or modeling could be close to the limitation .Our solution is using various retrieval strategies + feature engineering + GBDT, seems simple but powerful.</p>\n<p>We mainly generate recent popular items because fashion changing fast and has seasonality,tried to add cold start items but they will never be ranked to top12 due to lacking of interaction information.<br>\nuser and item interaction information are always the most important of recommendation problem,the features we created are almost interaction features,image and text features didn't help but should be useful for cold start problem.</p>\n<p>Almost 50% users have no transactions in recent 3months,so we created many cumulative features for them,and last week,last month,last season features for active users.</p>\n<p>We use 6 weeks data as train,last week as valid,retrieve 100 caninates for each user,it has stable cv-lb correlation.We focus on improving single lightgbm model to the last week, cv is 0.0430 and lb is 0.0362.At last week I have to rent gcp's big memory server and vast.ai's gpu server to run bigger models to get higher accuracy.</p>\n<p><a href=\"https://postimg.cc/xXLGHd2r\" target=\"_blank\"><img src=\"https://i.postimg.cc/bwBm7GVv/overview.png\" alt=\"overview.png\"></a></p>\n<h1>Retrieval(aka: candidate generation | recall)</h1>\n<p>At early stage,we focus on increasing the hit number of retrieved 100 candidates,try various strategies to cover more positive sample. </p>\n<table>\n<thead>\n<tr>\n<th>WeekNo</th>\n<th>HitNum@100</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2020-09-16</td>\n<td>39142</td>\n</tr>\n<tr>\n<td>2020-09-09</td>\n<td>38427</td>\n</tr>\n<tr>\n<td>2020-09-02</td>\n<td>41019</td>\n</tr>\n</tbody>\n</table>\n<h1>Ranking</h1>\n<ul>\n<li><p><strong>Feature Engineering</strong><br>\nbasiclly,features are created base on retrieval strategies,create user-item interaction for repurchase,create collaborative filtering score for itemcf,create similarity for embedding retrieval,create item count for popularity……</p>\n<table>\n<thead>\n<tr>\n<th>Type</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Count</td>\n<td>user-item, user-category of last week/month/season/same week of last year/all, time weighted count…</td>\n</tr>\n<tr>\n<td>Time</td>\n<td>first,last days of tranactions…</td>\n</tr>\n<tr>\n<td>Mean/Max/Min</td>\n<td>aggregation of age,price,sales_channel_id…</td>\n</tr>\n<tr>\n<td>Difference/Ratio</td>\n<td>difference between age and mean age of who purchased item, ratio of one user's purchased item count and the item's count</td>\n</tr>\n<tr>\n<td>Similarity</td>\n<td>collaborative filtering score of  item2item, cosine similarity of item2item(word2vec), cosine similarity of user2item(ProNE)</td>\n</tr>\n</tbody>\n</table></li>\n<li><p><strong>Downsampling</strong><br>\nIf we retrieve 100~500 candidates for each user,the negative samples number is very big,so negative downsampling is must,we found 1 million ~ 2 million negative samples for each week has better performace .<br>\n<code>neg_samples = 1000000;seed = 42\ntrain[train['label']&gt;0].append(train[train['label']==0].sample(neg_samples, random_state=seed))\n</code></p></li>\n<li><p><strong>Model</strong><br>\nOur best single model is a lightgbm(cv:0.0441,lb:0.0367),finally we trained 5 lightgbm classifier and 7 catboost classifier for ensemble(lb:0.0371).catboost's lb score is much worse than lightgbm,lightgbm has very stable cv-lb correlation.<br>\n<a href=\"https://postimg.cc/ZBmJRVLt\" target=\"_blank\"><img src=\"https://i.postimg.cc/V6rvKhr6/cv.png\" alt=\"cv.png\"></a></p></li>\n</ul>\n<h1>Optimization</h1>\n<ol>\n<li>we use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.</li>\n<li>transform all the categorical features(including two way) to label encoding,use reduce_mem_usage,</li>\n<li>create a <strong>feature store</strong>,save intermediate features files to dictionary , final features to feather,exsiting features will not be create again</li>\n<li>split all the users to 28 group,inference simultaneously with multiple servers.</li>\n</ol>\n<h1>Machine</h1>\n<p>I have a desktop with 128G RAM,64 vcore CPU,TITAN RTX GPU.It's enough to run most of our models to win.My teammate has a desktop with 64G RAM.At last week ,we use 300G RAM gcp instances.</p>",
      "rawMarkdown": "Thanks to H&M and Kaggle Team giving us such a wonderful competition.Thanks my team mate @h4211819 finally we got 1st place.This competition has no leak,stable cv-lb correlation,and infinite possibilities to improve, we really enjoyed exploring the accuracy boundaries of  infinite possibilities. \n\n# Overview\nThe most interesting part of this competition is we need to generate train and test by ourself, the candinate generation strategy is the key to beyond limitation of accuracy ,good feature engineering or modeling could be close to the limitation .Our solution is using various retrieval strategies + feature engineering + GBDT, seems simple but powerful.\n\nWe mainly generate recent popular items because fashion changing fast and has seasonality,tried to add cold start items but they will never be ranked to top12 due to lacking of interaction information.\nuser and item interaction information are always the most important of recommendation problem,the features we created are almost interaction features,image and text features didn't help but should be useful for cold start problem.\n\nAlmost 50% users have no transactions in recent 3months,so we created many cumulative features for them,and last week,last month,last season features for active users.\n\nWe use 6 weeks data as train,last week as valid,retrieve 100 caninates for each user,it has stable cv-lb correlation.We focus on improving single lightgbm model to the last week, cv is 0.0430 and lb is 0.0362.At last week I have to rent gcp's big memory server and vast.ai's gpu server to run bigger models to get higher accuracy.\n\n[![overview.png](https://i.postimg.cc/bwBm7GVv/overview.png)](https://postimg.cc/xXLGHd2r)\n\n# Retrieval(aka: candidate generation | recall)\nAt early stage,we focus on increasing the hit number of retrieved 100 candidates,try various strategies to cover more positive sample. \n| WeekNo | HitNum@100 |\n| --- | --- |\n|  2020-09-16|  39142|\n|  2020-09-09|  38427|\n|  2020-09-02|  41019|\n\n# Ranking\n- **Feature Engineering**\nbasiclly,features are created base on retrieval strategies,create user-item interaction for repurchase,create collaborative filtering score for itemcf,create similarity for embedding retrieval,create item count for popularity......\n| Type | Description|\n| --- | --- |\n| Count |  user-item, user-category of last week/month/season/same week of last year/all, time weighted count... |\n| Time|  first,last days of tranactions...|\n| Mean/Max/Min| aggregation of age,price,sales_channel_id... |\n| Difference/Ratio|  difference between age and mean age of who purchased item, ratio of one user's purchased item count and the item's count|\n| Similarity| collaborative filtering score of  item2item, cosine similarity of item2item(word2vec), cosine similarity of user2item(ProNE) |\n\n- **Downsampling**\nIf we retrieve 100~500 candidates for each user,the negative samples number is very big,so negative downsampling is must,we found 1 million ~ 2 million negative samples for each week has better performace .\n`neg_samples = 1000000;seed = 42\ntrain[train['label']>0].append(train[train['label']==0].sample(neg_samples, random_state=seed))\n`\n\n- **Model**\nOur best single model is a lightgbm(cv:0.0441,lb:0.0367),finally we trained 5 lightgbm classifier and 7 catboost classifier for ensemble(lb:0.0371).catboost's lb score is much worse than lightgbm,lightgbm has very stable cv-lb correlation.\n[![cv.png](https://i.postimg.cc/V6rvKhr6/cv.png)](https://postimg.cc/ZBmJRVLt)\n\n# Optimization\n1. we use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.\n2. transform all the categorical features(including two way) to label encoding,use reduce_mem_usage,\n3. create a **feature store**,save intermediate features files to dictionary , final features to feather,exsiting features will not be create again\n4. split all the users to 28 group,inference simultaneously with multiple servers.\n\n# Machine\nI have a desktop with 128G RAM,64 vcore CPU,TITAN RTX GPU.It's enough to run most of our models to win.My teammate has a desktop with 64G RAM.At last week ,we use 300G RAM gcp instances.\n\n",
      "votes": 430
    },
    {
      "id": 1783076,
      "postDate": "2022-05-10T05:25:35.497Z",
      "content": "<p>It is an amazing win and huge congrats to you!</p>\n<p>Gradient Boosted Trees &gt;= NN in recommender systems.</p>\n<p>Gradient boosted trees do not receive nearly as much coverage in papers &amp; research compared to their capabilities and importance for real-life production use-cases vs Neural Networks! And why? Because they do not resemble our brain and hence no promise of the \"general AI\"?</p>\n<p>I admit, a little bit of self confirmation bias speaks through me, because the LGBM + Manual Feature Engineering is also the technology choice that we made at my company for the recommender system. How many discussions we had that maybe there is some hidden \"gold\" in the user-item interactions… </p>\n<p>Turns out there is none. It is a super strong result to see this win! Thank you guys! You made my day!</p>",
      "rawMarkdown": "It is an amazing win and huge congrats to you!\n\nGradient Boosted Trees >= NN in recommender systems.\n\nGradient boosted trees do not receive nearly as much coverage in papers & research compared to their capabilities and importance for real-life production use-cases vs Neural Networks! And why? Because they do not resemble our brain and hence no promise of the \"general AI\"?\n\nI admit, a little bit of self confirmation bias speaks through me, because the LGBM + Manual Feature Engineering is also the technology choice that we made at my company for the recommender system. How many discussions we had that maybe there is some hidden \"gold\" in the user-item interactions... \n\nTurns out there is none. It is a super strong result to see this win! Thank you guys! You made my day!",
      "votes": 24,
      "replies": [
        {
          "id": 1783375,
          "postDate": "2022-05-10T10:52:59.917Z",
          "content": "<p>Totally agree with you, technology should be appropriate for business not just seems modern or fashionable</p>",
          "rawMarkdown": "Totally agree with you, technology should be appropriate for business not just seems modern or fashionable",
          "votes": 13
        },
        {
          "id": 1784501,
          "postDate": "2022-05-11T08:01:50.383Z",
          "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> I think one reason for success of gradient boosting models(over nns) for this competition is that we don't have as much compute power and as much time to try&amp;tune as engineering teams in big companys. We also have less data(e.g. exposed but not bought samples). Big companys do use nns for recommendation(though gbms are also used in industry).</p>",
          "rawMarkdown": "@narsil I think one reason for success of gradient boosting models(over nns) for this competition is that we don't have as much compute power and as much time to try&tune as engineering teams in big companys. We also have less data(e.g. exposed but not bought samples). Big companys do use nns for recommendation(though gbms are also used in industry).",
          "votes": 1,
          "replies": [
            {
              "id": 3167004,
              "postDate": "2025-04-01T07:58:40.530Z",
              "content": "<p>so you mean that gradient boosting models is better than nns ? I wonder if there are any stronger nn models can do better </p>",
              "rawMarkdown": "so you mean that gradient boosting models is better than nns ? I wonder if there are any stronger nn models can do better "
            }
          ]
        },
        {
          "id": 1786665,
          "postDate": "2022-05-13T05:58:54.443Z",
          "content": "<p><a href=\"https://www.kaggle.com/homoalways\" target=\"_blank\">@homoalways</a> Even in big companies, Trees + Manual FE have some big advantages over NNs. Cost, where you point out, is one of them. Also, manually engineered features can be easily debugged. They provide an additional layer of visibility into the model. By contrast, imagine debugging equivalent embeddings (pre final layer) in NNs.  </p>",
          "rawMarkdown": "@homoalways Even in big companies, Trees + Manual FE have some big advantages over NNs. Cost, where you point out, is one of them. Also, manually engineered features can be easily debugged. They provide an additional layer of visibility into the model. By contrast, imagine debugging equivalent embeddings (pre final layer) in NNs.  "
        }
      ]
    },
    {
      "id": 1783627,
      "postDate": "2022-05-10T14:32:42.967Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> &amp; <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> !  Great solution. </p>\n<p>Have you tried FIL library for LightGBM inference? In my tests it can be 10x faster than cpu inference.</p>",
      "rawMarkdown": "Congrats @senkin13 & @h4211819 !  Great solution. \n\nHave you tried FIL library for LightGBM inference? In my tests it can be 10x faster than cpu inference.",
      "votes": 10,
      "replies": [
        {
          "id": 1784163,
          "postDate": "2022-05-11T01:33:02.393Z",
          "content": "<p>never heard of FIL library,thanks for sharing</p>",
          "rawMarkdown": "never heard of FIL library,thanks for sharing"
        },
        {
          "id": 1784551,
          "postDate": "2022-05-11T08:44:30.517Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1783033,
      "postDate": "2022-05-10T04:25:05.173Z",
      "content": "<p>thanks for your great solution sharing…looking forward your code sharing..</p>",
      "rawMarkdown": "thanks for your great solution sharing...looking forward your code sharing..",
      "votes": 7
    },
    {
      "id": 1783012,
      "postDate": "2022-05-10T04:07:08.727Z",
      "content": "<p>Congratulations! Very solid work<br>\nOne question is what's the improvement of graph embedding in your recalling and ranking? That's a part which I wanna try but didn't have enough time.<br>\nThanks for your sharing.</p>",
      "rawMarkdown": "Congratulations! Very solid work\nOne question is what's the improvement of graph embedding in your recalling and ranking? That's a part which I wanna try but didn't have enough time.\nThanks for your sharing.",
      "votes": 5,
      "replies": [
        {
          "id": 1783378,
          "postDate": "2022-05-10T10:55:16.380Z",
          "content": "<p>about 0.002 up</p>",
          "rawMarkdown": "about 0.002 up",
          "votes": 1
        },
        {
          "id": 1783507,
          "postDate": "2022-05-10T12:45:59.947Z",
          "content": "<p>Wow, that's really powerful!</p>",
          "rawMarkdown": "Wow, that's really powerful!"
        }
      ]
    },
    {
      "id": 1784180,
      "postDate": "2022-05-11T02:15:10.040Z",
      "content": "<p>Very organized pipline and strong win. For some ideas, I also came up with, but because of the ineffeciency of the pipeline and I didn't organize everything as well as you did, so I came up so many bugs all along the way. The idea iteration speed is also not that satisfying. Learn from the best! Thanks for the sharing again!</p>",
      "rawMarkdown": "Very organized pipline and strong win. For some ideas, I also came up with, but because of the ineffeciency of the pipeline and I didn't organize everything as well as you did, so I came up so many bugs all along the way. The idea iteration speed is also not that satisfying. Learn from the best! Thanks for the sharing again!",
      "votes": 6
    },
    {
      "id": 1783023,
      "postDate": "2022-05-10T04:17:28.887Z",
      "content": "<p>Congratulations on winning the competition, it's really amazing.<br>\nWe were having difficulty dealing with 6 weeks of data, so I can imagine how challenging it'd have been to manage so much of data.</p>",
      "rawMarkdown": "Congratulations on winning the competition, it's really amazing.\nWe were having difficulty dealing with 6 weeks of data, so I can imagine how challenging it'd have been to manage so much of data.",
      "votes": 3
    },
    {
      "id": 1791221,
      "postDate": "2022-05-15T18:39:57.153Z",
      "content": "<p>In your chart, where you write <code>Top N</code>, do you use the same <code>N</code> for each retrieval method?</p>",
      "rawMarkdown": "In your chart, where you write `Top N`, do you use the same `N` for each retrieval method?",
      "votes": 4,
      "replies": [
        {
          "id": 1791819,
          "postDate": "2022-05-16T12:02:20.813Z",
          "content": "<p>different N by adjusting manually to optimize the hit number</p>",
          "rawMarkdown": "different N by adjusting manually to optimize the hit number",
          "votes": 2
        },
        {
          "id": 1792492,
          "postDate": "2022-05-17T02:38:29.267Z",
          "content": "<p>Thank you!</p>\n<p>What do you do if some customers don't have 'N' candidates for a given strategy?<br>\nFor example, if you first take top 10 repurchase candidates, what do you do for customers who have less than 20 previously purchased items?</p>",
          "rawMarkdown": "Thank you!\n\nWhat do you do if some customers don't have 'N' candidates for a given strategy?\nFor example, if you first take top 10 repurchase candidates, what do you do for customers who have less than 20 previously purchased items?"
        },
        {
          "id": 1793842,
          "postDate": "2022-05-18T09:20:43.393Z",
          "content": "<p>we don't need to take same N for each strategy,make sure total candidates is 100.At least we have top 100 popular items for all users.</p>",
          "rawMarkdown": "we don't need to take same N for each strategy,make sure total candidates is 100.At least we have top 100 popular items for all users."
        }
      ]
    },
    {
      "id": 1785247,
      "postDate": "2022-05-12T00:37:25.930Z",
      "content": "<p>This was super helpful for a begginer like me. Thanks My Besto Friendo.</p>",
      "rawMarkdown": "This was super helpful for a begginer like me. Thanks My Besto Friendo.",
      "votes": 4
    },
    {
      "id": 1783255,
      "postDate": "2022-05-10T08:24:33.553Z",
      "content": "<p>Congrats! I like the mention of a feature store. It seems very useful to not recalculate the features every time.</p>",
      "rawMarkdown": "Congrats! I like the mention of a feature store. It seems very useful to not recalculate the features every time.",
      "votes": 4,
      "replies": [
        {
          "id": 1783364,
          "postDate": "2022-05-10T10:38:47.730Z",
          "content": "<p>Yes,although I spent much time on creating pipeline and feature store,it helped us do more experiments without bugs</p>",
          "rawMarkdown": "Yes,although I spent much time on creating pipeline and feature store,it helped us do more experiments without bugs",
          "votes": 2
        },
        {
          "id": 1783477,
          "postDate": "2022-05-10T12:26:12.797Z",
          "content": "<p>would you please sharing the \" feature store\" demo code ?</p>",
          "rawMarkdown": "would you please sharing the \" feature store\" demo code ?",
          "votes": 1
        },
        {
          "id": 1784164,
          "postDate": "2022-05-11T01:34:12.703Z",
          "content": "<p>it's just a concept.<br>\nWhat is a feature store?<br>\nThe short version - a data management layer for machine learning that allows to share &amp; discover features and create more effective machine learning pipelines.</p>",
          "rawMarkdown": "it's just a concept.\nWhat is a feature store?\nThe short version - a data management layer for machine learning that allows to share & discover features and create more effective machine learning pipelines.",
          "votes": 5
        },
        {
          "id": 1784735,
          "postDate": "2022-05-11T12:32:17.583Z",
          "content": "<p>well, thanks</p>",
          "rawMarkdown": "well, thanks"
        }
      ]
    },
    {
      "id": 1782955,
      "postDate": "2022-05-10T02:52:55.223Z",
      "content": "<p>Congratulations! I didn't see anything special about it; just a great solution where every single process is deeply validated and refined, including optimizing the development! Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations! I didn't see anything special about it; just a great solution where every single process is deeply validated and refined, including optimizing the development! Thanks for sharing!",
      "votes": 4
    },
    {
      "id": 1794237,
      "postDate": "2022-05-18T15:58:11.693Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 1792871,
      "postDate": "2022-05-17T11:33:26.203Z",
      "content": "<p>Congratulations!<br>\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model …… Until the recall model is trained with the training set of the first six weeks and the final validation set is ranked to get the result.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations!\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model ...... Until the recall model is trained with the training set of the first six weeks and the final validation set is ranked to get the result.\n\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1793837,
          "postDate": "2022-05-18T09:14:51.193Z",
          "content": "<p>the simplest model is using last week data as samples of ranking model,recall top N candidates from previous data by rules or model,the overlap of recall data and last week data is positive samples,other recall data is negative samples.<br>\nI sugguest you read the code of kernel someone have publiced their code.</p>",
          "rawMarkdown": "the simplest model is using last week data as samples of ranking model,recall top N candidates from previous data by rules or model,the overlap of recall data and last week data is positive samples,other recall data is negative samples.\nI sugguest you read the code of kernel someone have publiced their code.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1792502,
      "postDate": "2022-05-17T02:52:21.093Z",
      "content": "<p>Congratulations on your first place finish!</p>\n<p>Very happy for you - you've put in years of kaggle work, and hard work finally paid off!</p>",
      "rawMarkdown": "Congratulations on your first place finish!\n\nVery happy for you - you've put in years of kaggle work, and hard work finally paid off!",
      "votes": 1
    },
    {
      "id": 1792498,
      "postDate": "2022-05-17T02:48:10.570Z",
      "content": "<p>What retrieval did you use for customers without history?</p>\n<p>Just 12 most popular items, or something more?<br>\nOn your CV, do you know what part of the score was attributable to the new customers?</p>",
      "rawMarkdown": "What retrieval did you use for customers without history?\n\nJust 12 most popular items, or something more?\nOn your CV, do you know what part of the score was attributable to the new customers?",
      "votes": 1,
      "replies": [
        {
          "id": 1793840,
          "postDate": "2022-05-18T09:17:34.847Z",
          "content": "<p>just most popular items.I have not checked but I think everyone have no big difference of new customers' score because it's hard to predict.</p>",
          "rawMarkdown": "just most popular items.I have not checked but I think everyone have no big difference of new customers' score because it's hard to predict.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1792313,
      "postDate": "2022-05-16T20:13:11.610Z",
      "content": "<p>Interesting :) Thx for sharing</p>",
      "rawMarkdown": "Interesting :) Thx for sharing",
      "votes": 1
    },
    {
      "id": 1791811,
      "postDate": "2022-05-16T11:53:33.487Z",
      "content": "<p>Great work! Thank you for this knowledge</p>",
      "rawMarkdown": "Great work! Thank you for this knowledge",
      "votes": 1
    },
    {
      "id": 1791010,
      "postDate": "2022-05-15T15:09:33.527Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": 1
    },
    {
      "id": 1791007,
      "postDate": "2022-05-15T15:07:40.357Z",
      "content": "<p>Congratulations! Thank you for your sharing! This solution is really nice!</p>",
      "rawMarkdown": "Congratulations! Thank you for your sharing! This solution is really nice!",
      "votes": 1
    },
    {
      "id": 1790882,
      "postDate": "2022-05-15T12:48:01.797Z",
      "content": "<p>Amazing job! Please share more soon. </p>",
      "rawMarkdown": "Amazing job! Please share more soon. ",
      "votes": 1
    },
    {
      "id": 1790549,
      "postDate": "2022-05-15T04:44:17.390Z",
      "content": "<p>Looks interesting. Thanks for sharing. 👊</p>",
      "rawMarkdown": "Looks interesting. Thanks for sharing. 👊",
      "votes": 1
    },
    {
      "id": 1790535,
      "postDate": "2022-05-15T04:32:38.947Z",
      "content": "<p>Congratulation and Thanks for the sharing 👍</p>",
      "rawMarkdown": "Congratulation and Thanks for the sharing 👍",
      "votes": 1
    },
    {
      "id": 1790328,
      "postDate": "2022-05-14T19:15:04.820Z",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work\n",
      "votes": 1
    },
    {
      "id": 1789946,
      "postDate": "2022-05-14T11:24:11.417Z",
      "content": "<p>THX for sharing!!!</p>",
      "rawMarkdown": "THX for sharing!!!\n",
      "votes": 1
    },
    {
      "id": 1789569,
      "postDate": "2022-05-14T03:00:26.503Z",
      "content": "<p>Great work.</p>",
      "rawMarkdown": "Great work.",
      "votes": 1
    },
    {
      "id": 1787343,
      "postDate": "2022-05-13T19:47:01.180Z",
      "content": "<p>great work!</p>",
      "rawMarkdown": "great work!",
      "votes": 1
    },
    {
      "id": 1787271,
      "postDate": "2022-05-13T18:10:59.220Z",
      "content": "<p>Very well explained. Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Very well explained. Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1787158,
      "postDate": "2022-05-13T16:05:51.927Z",
      "content": "<p>Hi, amazing work!! Thank so much for sharing!! </p>\n<p>Could I ask you a question?<br>\nIn the case of cosine similarity, are you using the closest or another distance to the items bought in the previous weeks?</p>\n<p>Thank you again!</p>",
      "rawMarkdown": "Hi, amazing work!! Thank so much for sharing!! \n\nCould I ask you a question?\nIn the case of cosine similarity, are you using the closest or another distance to the items bought in the previous weeks?\n\nThank you again!",
      "votes": 1,
      "replies": [
        {
          "id": 1791827,
          "postDate": "2022-05-16T12:07:44.490Z",
          "content": "<p>cosine similarity between user's latest purchased ,2nd latest purchased items and candidate item</p>",
          "rawMarkdown": " cosine similarity between user's latest purchased ,2nd latest purchased items and candidate item"
        }
      ]
    },
    {
      "id": 1787153,
      "postDate": "2022-05-13T16:03:00.137Z",
      "content": "<p>Hello, congratulations for the work.</p>\n<p>But I have a question. How does item2item integrate with lightgbm? Is it a feature engineering itself or an intermediate step for another process?</p>\n<p>Thanks for sharing knowledge!</p>",
      "rawMarkdown": "Hello, congratulations for the work.\n\nBut I have a question. How does item2item integrate with lightgbm? Is it a feature engineering itself or an intermediate step for another process?\n\nThanks for sharing knowledge!",
      "votes": 1,
      "replies": [
        {
          "id": 1791832,
          "postDate": "2022-05-16T12:09:26.473Z",
          "content": "<p>it's feature engineering, cosine similarity between user's latest purchased ,2nd latest purchased items and candidate item</p>",
          "rawMarkdown": "it's feature engineering, cosine similarity between user's latest purchased ,2nd latest purchased items and candidate item"
        }
      ]
    },
    {
      "id": 1786735,
      "postDate": "2022-05-13T08:05:04.417Z",
      "content": "<p>great work!</p>",
      "rawMarkdown": "great work!",
      "votes": 1
    },
    {
      "id": 1786667,
      "postDate": "2022-05-13T06:00:02.337Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": 1
    },
    {
      "id": 1786497,
      "postDate": "2022-05-13T00:47:35.707Z",
      "content": "<p>Nice, good work.</p>",
      "rawMarkdown": "Nice, good work.",
      "votes": 1
    },
    {
      "id": 1786438,
      "postDate": "2022-05-12T22:34:18.533Z",
      "content": "<p>outstanding performance and a wonderful write up! huge congrats and thank you for sharing your insights!!! 🙏</p>\n<p>Could I please ask you how did you handle duplicate purchases? Do you ever predict duplicate items?</p>\n<p>I created \"baskets\" of purchases by each customer per week and removed duplicates, but not sure if that is a good way to handle this?</p>\n<p>Also, if I am reading this right, you used some number of weeks for train, used the last week for validation, and used the model trained like this to predict on test. Is my understanding correct? You didn't retrain the models on the full dataset (up unto including the validation week, the last week of the train set) once you found good hyperparams to train with?</p>\n<p>Thank you very much for all your help on this! 🙂</p>",
      "rawMarkdown": "outstanding performance and a wonderful write up! huge congrats and thank you for sharing your insights!!! 🙏\n\nCould I please ask you how did you handle duplicate purchases? Do you ever predict duplicate items?\n\nI created \"baskets\" of purchases by each customer per week and removed duplicates, but not sure if that is a good way to handle this?\n\nAlso, if I am reading this right, you used some number of weeks for train, used the last week for validation, and used the model trained like this to predict on test. Is my understanding correct? You didn't retrain the models on the full dataset (up unto including the validation week, the last week of the train set) once you found good hyperparams to train with?\n\nThank you very much for all your help on this! 🙂",
      "votes": 1,
      "replies": [
        {
          "id": 1786499,
          "postDate": "2022-05-13T00:52:19.573Z",
          "content": "<p>I removed duplicate purchases by each customer per week,but keep them if they existed in different weeks.I retrained full dataset with more rounds,about 1.1X .</p>",
          "rawMarkdown": "I removed duplicate purchases by each customer per week,but keep them if they existed in different weeks.I retrained full dataset with more rounds,about 1.1X .",
          "votes": 1
        },
        {
          "id": 1786582,
          "postDate": "2022-05-13T03:31:17.913Z",
          "content": "<p>Thank you very much for your reply!!!!! 🙂🙏</p>",
          "rawMarkdown": "Thank you very much for your reply!!!!! 🙂🙏"
        }
      ]
    },
    {
      "id": 1786412,
      "postDate": "2022-05-12T21:32:29.380Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 1786218,
      "postDate": "2022-05-12T17:33:59.083Z",
      "content": "<p>Amazing work!!</p>",
      "rawMarkdown": "Amazing work!!",
      "votes": 1
    },
    {
      "id": 1785998,
      "postDate": "2022-05-12T14:43:53.250Z",
      "content": "<p>Great work! Congratulations!</p>",
      "rawMarkdown": "Great work! Congratulations!",
      "votes": 1
    },
    {
      "id": 1785408,
      "postDate": "2022-05-12T04:37:07.027Z",
      "content": "<p>congratulation</p>",
      "rawMarkdown": "congratulation",
      "votes": 1
    },
    {
      "id": 1785292,
      "postDate": "2022-05-12T02:05:57.930Z",
      "content": "<p>Congratulations! Thanks for sharing the solution. </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing the solution. ",
      "votes": 1
    },
    {
      "id": 1784919,
      "postDate": "2022-05-11T15:34:50.580Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": 1
    },
    {
      "id": 1784487,
      "postDate": "2022-05-11T07:53:33.230Z",
      "content": "<p>Congrats. Did you process your features&amp;recall candidates with gpu using cudf? Or you use cpu to do these? I got cuda out of memory error if I use gpu(p100, 16GB gpu memory) to recall and merge features for all data.</p>",
      "rawMarkdown": "Congrats. Did you process your features&recall candidates with gpu using cudf? Or you use cpu to do these? I got cuda out of memory error if I use gpu(p100, 16GB gpu memory) to recall and merge features for all data.",
      "votes": 1,
      "replies": [
        {
          "id": 1784620,
          "postDate": "2022-05-11T10:19:07.110Z",
          "content": "<p>I installed cudf in my windows but failed in the past.Just use cpu.I used joblib to parallel computing and multiple notebooks to process data.</p>",
          "rawMarkdown": "I installed cudf in my windows but failed in the past.Just use cpu.I used joblib to parallel computing and multiple notebooks to process data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1784428,
      "postDate": "2022-05-11T07:13:59.130Z",
      "content": "<p>공유해 주셔서 감사합니다!</p>",
      "rawMarkdown": "공유해 주셔서 감사합니다!",
      "votes": 1
    },
    {
      "id": 1784361,
      "postDate": "2022-05-11T06:09:49.813Z",
      "content": "<p>Congrats on winner <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and team!</p>",
      "rawMarkdown": "Congrats on winner @senkin13 and team!",
      "votes": 1
    },
    {
      "id": 1784232,
      "postDate": "2022-05-11T03:41:51.740Z",
      "content": "<p>Congratulation! and thank you for sharing！</p>",
      "rawMarkdown": "Congratulation! and thank you for sharing！",
      "votes": 1
    },
    {
      "id": 1784155,
      "postDate": "2022-05-11T01:18:08.877Z",
      "content": "<p>Congratulations and thank you for sharing, this is very insightful.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing, this is very insightful.",
      "votes": 1
    },
    {
      "id": 1784053,
      "postDate": "2022-05-10T22:19:57.340Z",
      "content": "<p>Congratulations!  Thanks for sharing!<br>\nGreat idia about to reduce work with data is to take the latest popular product and according to seasonal characteristics!</p>",
      "rawMarkdown": "Congratulations!  Thanks for sharing!\nGreat idia about to reduce work with data is to take the latest popular product and according to seasonal characteristics!",
      "votes": 1
    },
    {
      "id": 1783928,
      "postDate": "2022-05-10T19:27:35.077Z",
      "content": "<p>This is a great write-up! You guys definitely deserve the first place with all the hard work you put in. I love how you experimented with different retrieval strategies and features to get the most out of your model. Keep up the good work!</p>",
      "rawMarkdown": "This is a great write-up! You guys definitely deserve the first place with all the hard work you put in. I love how you experimented with different retrieval strategies and features to get the most out of your model. Keep up the good work!",
      "votes": 1
    },
    {
      "id": 1783867,
      "postDate": "2022-05-10T18:38:55.087Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 1783818,
      "postDate": "2022-05-10T17:53:01.190Z",
      "content": "<p>Hats off to you !</p>",
      "rawMarkdown": "Hats off to you !",
      "votes": 1
    },
    {
      "id": 1783797,
      "postDate": "2022-05-10T17:18:03.823Z",
      "content": "<p>Congratulations on winning the 1st place! It's mind blowing to see such diverse concepts being used to get the best result possible. For someone like me who is just getting started in this field, this is highly inspiring to put more effort and learn in this bottomless ocean of a field. Thank you.</p>",
      "rawMarkdown": "Congratulations on winning the 1st place! It's mind blowing to see such diverse concepts being used to get the best result possible. For someone like me who is just getting started in this field, this is highly inspiring to put more effort and learn in this bottomless ocean of a field. Thank you.",
      "votes": 1
    },
    {
      "id": 1783751,
      "postDate": "2022-05-10T16:37:40.437Z",
      "content": "<p>Congratulation! and thank you for sharing yours!</p>",
      "rawMarkdown": "Congratulation! and thank you for sharing yours!",
      "votes": 1
    },
    {
      "id": 1783713,
      "postDate": "2022-05-10T16:06:21.983Z",
      "content": "<p>Congratulations on the first place and thanks for sharing your approach!</p>",
      "rawMarkdown": "Congratulations on the first place and thanks for sharing your approach!",
      "votes": 1
    },
    {
      "id": 1783712,
      "postDate": "2022-05-10T16:05:34.580Z",
      "content": "<p>Congratulations and thank you for sharing.<br>\nI have a question about GBDT.<br>\nYou said you used LightGBM and CatBoost.<br>\nDidn't you try to use XGBoost?<br>\nIn my environment, XGBoost is as good as LightGBM and can be used for the ensemble.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing.\nI have a question about GBDT.\nYou said you used LightGBM and CatBoost.\nDidn't you try to use XGBoost?\nIn my environment, XGBoost is as good as LightGBM and can be used for the ensemble.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1784166,
          "postDate": "2022-05-11T01:42:10.263Z",
          "content": "<p>I tried XGBoost-gpu once ,worse than lightgbm,and slower than catboost,so I didn't try more.</p>",
          "rawMarkdown": "I tried XGBoost-gpu once ,worse than lightgbm,and slower than catboost,so I didn't try more.",
          "votes": 2
        },
        {
          "id": 1784554,
          "postDate": "2022-05-11T08:45:20.610Z",
          "content": "<p>I understand.<br>\nThank you for your reply.</p>",
          "rawMarkdown": "I understand.\nThank you for your reply."
        }
      ]
    },
    {
      "id": 1783658,
      "postDate": "2022-05-10T15:00:47.080Z",
      "content": "<p>Congratulations! I tried deep learning models for recall and rank. But none of the models get a good score. </p>",
      "rawMarkdown": "Congratulations! I tried deep learning models for recall and rank. But none of the models get a good score. ",
      "votes": 1,
      "replies": [
        {
          "id": 1784165,
          "postDate": "2022-05-11T01:38:23.207Z",
          "content": "<p>I tried FM for recall and MLP,transformer for rank,didn't work.Luckily I gave up soon to focus on GBDT.</p>",
          "rawMarkdown": "I tried FM for recall and MLP,transformer for rank,didn't work.Luckily I gave up soon to focus on GBDT.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1783593,
      "postDate": "2022-05-10T14:02:15.153Z",
      "content": "<p>congratulation! what seems to be weird always be the best of bests </p>",
      "rawMarkdown": "congratulation! what seems to be weird always be the best of bests ",
      "votes": 1
    },
    {
      "id": 1783577,
      "postDate": "2022-05-10T13:48:31.497Z",
      "content": "<p>Congratulations! thanks for taking the time to write this up.</p>",
      "rawMarkdown": "Congratulations! thanks for taking the time to write this up.",
      "votes": 1
    },
    {
      "id": 1783500,
      "postDate": "2022-05-10T12:41:04.543Z",
      "content": "<p>Congratulation! and thank for sharing yours!</p>",
      "rawMarkdown": "Congratulation! and thank for sharing yours!",
      "votes": 1
    },
    {
      "id": 1783447,
      "postDate": "2022-05-10T12:03:10.640Z",
      "content": "<p>Congratulation! Thanks for sharing! 💪💪💪</p>",
      "rawMarkdown": "Congratulation! Thanks for sharing! 💪💪💪",
      "votes": 1
    },
    {
      "id": 1783390,
      "postDate": "2022-05-10T11:06:09.973Z",
      "content": "<p>Really well done - thanks for the explanation on approach. Congratulations.</p>",
      "rawMarkdown": "Really well done - thanks for the explanation on approach. Congratulations.",
      "votes": 1
    },
    {
      "id": 1783386,
      "postDate": "2022-05-10T11:02:44.767Z",
      "content": "<p>learnt a lot from this!</p>",
      "rawMarkdown": "learnt a lot from this!\n",
      "votes": 1
    },
    {
      "id": 1783359,
      "postDate": "2022-05-10T10:34:53.040Z",
      "content": "<p>Congratulations! Thanks for sharing and explanations! and looking forward to a code sharing ;)</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing and explanations! and looking forward to a code sharing ;)",
      "votes": 1
    },
    {
      "id": 1783301,
      "postDate": "2022-05-10T09:24:48.520Z",
      "content": "<p>Congratulation! and thank for sharing yours!</p>",
      "rawMarkdown": "Congratulation! and thank for sharing yours!",
      "votes": 1
    },
    {
      "id": 1783267,
      "postDate": "2022-05-10T08:45:10.657Z",
      "content": "<p>Really well done - thanks for the explanation on approach. Congratulations.</p>",
      "rawMarkdown": "Really well done - thanks for the explanation on approach. Congratulations.",
      "votes": 1
    },
    {
      "id": 1783253,
      "postDate": "2022-05-10T08:22:01.833Z",
      "content": "<p>Congrats on first place!<br>\nI used to use item2vec's similarity feature, but my score doesn't even come close your score!<br>\nIf you don't mind, I'd like to know your thoughts on which particular features gave you the best results.</p>",
      "rawMarkdown": "Congrats on first place!\nI used to use item2vec's similarity feature, but my score doesn't even come close your score!\nIf you don't mind, I'd like to know your thoughts on which particular features gave you the best results.",
      "votes": 1,
      "replies": [
        {
          "id": 1783365,
          "postDate": "2022-05-10T10:42:12.550Z",
          "content": "<p>no magic features,just add more and more useful features,finally we have 340+ features</p>",
          "rawMarkdown": "no magic features,just add more and more useful features,finally we have 340+ features",
          "votes": 1
        }
      ]
    },
    {
      "id": 1783217,
      "postDate": "2022-05-10T07:50:45.507Z",
      "content": "<p>Congratulation! and thank for sharing yours!</p>",
      "rawMarkdown": "Congratulation! and thank for sharing yours!",
      "votes": 1
    },
    {
      "id": 1783213,
      "postDate": "2022-05-10T07:49:40.873Z",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> <br>\nCongratulations! One question - considering stable cv-lb correlation, and using last week as valid ( also assuming that you used last week data as it appeared in the transactions csv, basically in order of purchase date), would it fair to assume that LB data is also based on purchase date? Probably no distinction for same day purchase but otherwise on purchase date.</p>",
      "rawMarkdown": "@senkin13 \nCongratulations! One question - considering stable cv-lb correlation, and using last week as valid ( also assuming that you used last week data as it appeared in the transactions csv, basically in order of purchase date), would it fair to assume that LB data is also based on purchase date? Probably no distinction for same day purchase but otherwise on purchase date.",
      "votes": 1,
      "replies": [
        {
          "id": 1783368,
          "postDate": "2022-05-10T10:45:48.653Z",
          "content": "<p>public and private leaderboard data should be splitted by user groups,the data size is big enough to have a approximate distribution</p>",
          "rawMarkdown": "public and private leaderboard data should be splitted by user groups,the data size is big enough to have a approximate distribution",
          "votes": 1
        }
      ]
    },
    {
      "id": 1783210,
      "postDate": "2022-05-10T07:49:02.593Z",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> wow fantastic work! congratulations</p>",
      "rawMarkdown": "@senkin13 wow fantastic work! congratulations",
      "votes": 1
    },
    {
      "id": 1783185,
      "postDate": "2022-05-10T07:25:27.993Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 1783139,
      "postDate": "2022-05-10T06:43:10.087Z",
      "content": "<p>Congratulation for winning!!!<br>\nI also split test_data into 28 groups.it was hard work.😂</p>",
      "rawMarkdown": "Congratulation for winning!!!\nI also split test_data into 28 groups.it was hard work.😂",
      "votes": 1
    },
    {
      "id": 1783137,
      "postDate": "2022-05-10T06:42:47.877Z",
      "content": "<p>Congrats on the amazing win <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> </p>\n<p>Thanks a lot for the detailed solution sharing. </p>",
      "rawMarkdown": "Congrats on the amazing win @senkin13 and @h4211819 \n\nThanks a lot for the detailed solution sharing. ",
      "votes": 1
    },
    {
      "id": 1783075,
      "postDate": "2022-05-10T05:21:00.677Z",
      "content": "<p>Congrats for the WIN!</p>",
      "rawMarkdown": "Congrats for the WIN!",
      "votes": 1
    },
    {
      "id": 1783070,
      "postDate": "2022-05-10T05:16:21.503Z",
      "content": "<p>Congratulation to winning.<br>\nIf you have paper of ProNE, can you share them?</p>",
      "rawMarkdown": "Congratulation to winning.\nIf you have paper of ProNE, can you share them?",
      "votes": 1,
      "replies": [
        {
          "id": 1783129,
          "postDate": "2022-05-10T06:27:49.520Z",
          "content": "<p><a href=\"https://www.ijcai.org/proceedings/2019/0594.pdf\" target=\"_blank\">https://www.ijcai.org/proceedings/2019/0594.pdf</a></p>",
          "rawMarkdown": "https://www.ijcai.org/proceedings/2019/0594.pdf",
          "votes": 1
        },
        {
          "id": 1808756,
          "postDate": "2022-06-02T06:13:16.267Z",
          "content": "<p>thanks for your reply</p>",
          "rawMarkdown": "thanks for your reply"
        }
      ]
    },
    {
      "id": 1783069,
      "postDate": "2022-05-10T05:16:02.860Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 1783028,
      "postDate": "2022-05-10T04:21:06.443Z",
      "content": "<p>Congratulations! Nice work! Could you elaborate the way to utilize graph embedding to generate candidates? Did you use it for some models as inputs? Thanks for sharing?</p>",
      "rawMarkdown": "Congratulations! Nice work! Could you elaborate the way to utilize graph embedding to generate candidates? Did you use it for some models as inputs? Thanks for sharing?",
      "votes": 1,
      "replies": [
        {
          "id": 1783054,
          "postDate": "2022-05-10T04:59:34.687Z",
          "content": "<p>you can use deepwalk, node2vec, LINE, ProNE…to train a user-item bipartite graph to output embeddings</p>",
          "rawMarkdown": "you can use deepwalk, node2vec, LINE, ProNE...to train a user-item bipartite graph to output embeddings",
          "votes": 1
        },
        {
          "id": 1783125,
          "postDate": "2022-05-10T06:24:51.923Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1785479,
          "postDate": "2022-05-12T06:12:30.360Z",
          "content": "<p>so what‘s the network embedding method in your case?</p>",
          "rawMarkdown": "so what‘s the network embedding method in your case?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1782970,
      "postDate": "2022-05-10T03:18:39.157Z",
      "content": "<p>Congrats for the champinship!</p>",
      "rawMarkdown": "Congrats for the champinship!",
      "votes": 1
    },
    {
      "id": 1782966,
      "postDate": "2022-05-10T03:11:54.747Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! ",
      "votes": 1
    },
    {
      "id": 1782939,
      "postDate": "2022-05-10T02:17:46.007Z",
      "content": "<p>Congrats for your No 1, Teather Zhan and Teather Kan. Great solution and nice write up.</p>",
      "rawMarkdown": "Congrats for your No 1, Teather Zhan and Teather Kan. Great solution and nice write up.",
      "votes": 1
    },
    {
      "id": 1791444,
      "postDate": "2022-05-16T03:11:00.140Z",
      "content": "<p>Congrats🎉🎉  and thanks for sharing your great work. It's quite impressive😄. <br>\nIf you don't mind, I want to ask some question about your work.</p>\n<ol>\n<li><p>How did you trained ProNE graph embedding model using bipartite graph(using two kinds of nodes : user and item)? I've checked the github for ProNE, it's basically using homogeneous information. Can you explain more detail about the preprocessing for bipartite graph using hm dataset or any reference about this?</p></li>\n<li><p>Could you explain more detail about the optimization of the ranking model ? If there are various kinds of candidates from different strategies, how did you set the target variables for each candidates?</p></li>\n</ol>\n<p>Again, thanks for your sharing!</p>",
      "rawMarkdown": "Congrats🎉🎉  and thanks for sharing your great work. It's quite impressive😄. \nIf you don't mind, I want to ask some question about your work.\n\n1. How did you trained ProNE graph embedding model using bipartite graph(using two kinds of nodes : user and item)? I've checked the github for ProNE, it's basically using homogeneous information. Can you explain more detail about the preprocessing for bipartite graph using hm dataset or any reference about this?\n\n2. Could you explain more detail about the optimization of the ranking model ? If there are various kinds of candidates from different strategies, how did you set the target variables for each candidates?\n\nAgain, thanks for your sharing!",
      "votes": 2,
      "replies": [
        {
          "id": 1791818,
          "postDate": "2022-05-16T12:01:05.593Z",
          "content": "<p>you can refer to this github, nodes  are user and item. <a href=\"https://github.com/shenweichen/GraphEmbedding\" target=\"_blank\">https://github.com/shenweichen/GraphEmbedding</a>.</p>\n<p>we don't need to set the target variables for different strategies,just merge them to top 100 deduplicated candidates for each user,if user-item pair existed in next week,the target is 1 else 0.</p>",
          "rawMarkdown": "you can refer to this github, nodes  are user and item. https://github.com/shenweichen/GraphEmbedding.\n\nwe don't need to set the target variables for different strategies,just merge them to top 100 deduplicated candidates for each user,if user-item pair existed in next week,the target is 1 else 0.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1786816,
      "postDate": "2022-05-13T10:22:49.080Z",
      "content": "<p>Congratulations and thanks for sharing!<br>\nI have one question:<br>\nhow to calculate <code>HitNum@100</code> in the Retrieval module?Is it equal to the number of positive samples in the top100 recall items?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!\nI have one question:\nhow to calculate `HitNum@100` in the Retrieval module?Is it equal to the number of positive samples in the top100 recall items?\n",
      "votes": 2,
      "replies": [
        {
          "id": 1787066,
          "postDate": "2022-05-13T14:49:41.460Z",
          "content": "<p>yes,equal to the number of positive samples in the top100 recall items</p>",
          "rawMarkdown": "yes,equal to the number of positive samples in the top100 recall items",
          "votes": 2
        },
        {
          "id": 1787102,
          "postDate": "2022-05-13T15:19:07.377Z",
          "content": "<p>thanks for your reply!</p>",
          "rawMarkdown": "thanks for your reply!"
        }
      ]
    },
    {
      "id": 1785765,
      "postDate": "2022-05-12T11:50:57.600Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 2
    },
    {
      "id": 1785559,
      "postDate": "2022-05-12T07:59:22.077Z",
      "content": "<p>Epic work!!</p>",
      "rawMarkdown": "Epic work!!\n",
      "votes": 2
    },
    {
      "id": 1785496,
      "postDate": "2022-05-12T06:23:32.553Z",
      "content": "<p>Great work!, what's \"the logistic regression with categorical information\"?</p>",
      "rawMarkdown": "Great work!, what's \"the logistic regression with categorical information\"?",
      "votes": 2,
      "replies": [
        {
          "id": 1785782,
          "postDate": "2022-05-12T12:09:18.393Z",
          "content": "<p>train a logistic regression model to retrieve from top 1000 popular candidates to 50~200 candidates</p>",
          "rawMarkdown": "train a logistic regression model to retrieve from top 1000 popular candidates to 50~200 candidates",
          "votes": 1
        },
        {
          "id": 1785828,
          "postDate": "2022-05-12T12:46:36.090Z",
          "content": "<p>ohhhhhhh… great, just like  a pre-ranker?</p>",
          "rawMarkdown": "ohhhhhhh... great, just like  a pre-ranker?"
        },
        {
          "id": 1785896,
          "postDate": "2022-05-12T13:29:07.960Z",
          "content": "<p>yes,ike a pre-ranker</p>",
          "rawMarkdown": "yes,ike a pre-ranker",
          "votes": 1
        },
        {
          "id": 1792496,
          "postDate": "2022-05-17T02:46:41.680Z",
          "content": "<p>Thank you for sharing this!</p>\n<p>#1<br>\nBy \"categorical information\" do you mean department/category/product features for items, and <br>\ncounts/percentages of previous purchases in those categories for users? Or something else?</p>\n<p>#2<br>\nAre 1000 and 50~200 the actual numbers?</p>\n<p>So you'd have 60 million candidates to train for each week (60k purchasing customers X 1,000 popular candidates)?<br>\nAnd 50 of your hundred candidates for each customer were from this retrieval method?</p>",
          "rawMarkdown": "Thank you for sharing this!\n\n\\#1\nBy \"categorical information\" do you mean department/category/product features for items, and \ncounts/percentages of previous purchases in those categories for users? Or something else?\n\n\\#2\nAre 1000 and 50~200 the actual numbers?\n\nSo you'd have 60 million candidates to train for each week (60k purchasing customers X 1,000 popular candidates)?\nAnd 50 of your hundred candidates for each customer were from this retrieval method?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1785319,
      "postDate": "2022-05-12T02:51:07.350Z",
      "content": "<p>Congratulation! and thank you for sharing!<br>\nIf you don't mind, I'd like to know how to calc collaborative filtering score of item2item.<br>\nI would be glad to know if you have any reference materials for the collaborative filtering score of item2item.</p>",
      "rawMarkdown": "Congratulation! and thank you for sharing!\nIf you don't mind, I'd like to know how to calc collaborative filtering score of item2item.\nI would be glad to know if you have any reference materials for the collaborative filtering score of item2item.",
      "votes": 2,
      "replies": [
        {
          "id": 1785497,
          "postDate": "2022-05-12T06:23:32.603Z",
          "content": "<p>you can see this article<br>\n<a href=\"https://songxia-sophia.medium.com/collaborative-filtering-recommendation-with-co-occurrence-algorithm-dea583e12e2a\" target=\"_blank\">https://songxia-sophia.medium.com/collaborative-filtering-recommendation-with-co-occurrence-algorithm-dea583e12e2a</a></p>",
          "rawMarkdown": "you can see this article\nhttps://songxia-sophia.medium.com/collaborative-filtering-recommendation-with-co-occurrence-algorithm-dea583e12e2a",
          "votes": 2
        },
        {
          "id": 1791927,
          "postDate": "2022-05-16T13:35:04.983Z",
          "content": "<p>Thank you! Thank you! I learned a lot!!</p>",
          "rawMarkdown": "Thank you! Thank you! I learned a lot!!"
        }
      ]
    },
    {
      "id": 1782994,
      "postDate": "2022-05-10T03:43:50.520Z",
      "content": "<p>Congrats on 1st and sharing an impressive solution 🙌 Looking forward to seeing it be expanded on to learn more from it.</p>",
      "rawMarkdown": "Congrats on 1st and sharing an impressive solution 🙌 Looking forward to seeing it be expanded on to learn more from it.",
      "votes": 2
    },
    {
      "id": 1782974,
      "postDate": "2022-05-10T03:22:55.877Z",
      "content": "<p>Nice job both of you two grandmasters.</p>",
      "rawMarkdown": "Nice job both of you two grandmasters.",
      "votes": 2
    },
    {
      "id": 1782973,
      "postDate": "2022-05-10T03:21:41.447Z",
      "content": "<p>Congratulations! If it isn't my mistake you guys led the competition from the start. That's impressive! Could you please describe the used hardware for your solution?</p>",
      "rawMarkdown": "Congratulations! If it isn't my mistake you guys led the competition from the start. That's impressive! Could you please describe the used hardware for your solution?",
      "votes": 2,
      "replies": [
        {
          "id": 1783049,
          "postDate": "2022-05-10T04:53:03.730Z",
          "content": "<p>I have a desktop with 128G RAM,64 vcore CPU,TITAN RTX GPU.It's enough to run most of our models to win.But bigger model need 300G RAM so I have to use gcp.</p>",
          "rawMarkdown": "I have a desktop with 128G RAM,64 vcore CPU,TITAN RTX GPU.It's enough to run most of our models to win.But bigger model need 300G RAM so I have to use gcp.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1784787,
      "postDate": "2022-05-11T13:29:01.560Z",
      "content": "<p>Congratulation! and thank you for sharing！<br>\nAnd one question - How do you merge various retrieval strategies into 100 candidates for each user?<br>\nIn my solution, I just recall 200 candidates from each recall strategy for each user, and then just merge these items, and feed into the rank model. I find when I add more recall strategy, my score does not improve, so I am wondering if there's something wrong with the way I merge each recall candidates(just merge not limit or rough rank).</p>",
      "rawMarkdown": "Congratulation! and thank you for sharing！\nAnd one question - How do you merge various retrieval strategies into 100 candidates for each user?\nIn my solution, I just recall 200 candidates from each recall strategy for each user, and then just merge these items, and feed into the rank model. I find when I add more recall strategy, my score does not improve, so I am wondering if there's something wrong with the way I merge each recall candidates(just merge not limit or rough rank).",
      "replies": [
        {
          "id": 1784881,
          "postDate": "2022-05-11T14:57:47.697Z",
          "content": "<p>adjusting top N of each retrieval strategie and strategies priority manually ,then merge them to select top 100,optimize the hit number(positive samples)</p>",
          "rawMarkdown": "adjusting top N of each retrieval strategie and strategies priority manually ,then merge them to select top 100,optimize the hit number(positive samples)"
        },
        {
          "id": 1784912,
          "postDate": "2022-05-11T15:28:22.700Z",
          "content": "<p>Thank you for your reply. </p>",
          "rawMarkdown": "Thank you for your reply. "
        }
      ]
    },
    {
      "id": 2480501,
      "postDate": "2023-10-13T10:59:43.283Z",
      "content": "<p>can you please share your full notebook <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a>. i need it for educational purpose.</p>",
      "rawMarkdown": "can you please share your full notebook @h4211819. i need it for educational purpose."
    },
    {
      "id": 2072382,
      "postDate": "2022-12-22T04:01:27.960Z",
      "content": "<p>great solution! looking forward your code sharing..</p>",
      "rawMarkdown": "great solution! looking forward your code sharing.."
    },
    {
      "id": 2006153,
      "postDate": "2022-10-27T13:01:46.703Z",
      "content": "<p>Thank you very much for sharing your work!! respect! But ı clouldnt understand this part <br>\n\"  gender_sale.loc[gender_sale[1]&gt;=0.8, 'gender'] = 1<br>\ngender_sale.loc[gender_sale[2]&gt;=0.8, 'gender'] = 2   \" </p>\n<p>Could you explain that part please?</p>",
      "rawMarkdown": "Thank you very much for sharing your work!! respect! But ı clouldnt understand this part \n\"  gender_sale.loc[gender_sale[1]>=0.8, 'gender'] = 1\ngender_sale.loc[gender_sale[2]>=0.8, 'gender'] = 2   \" \n\nCould you explain that part please?"
    },
    {
      "id": 1853511,
      "postDate": "2022-07-12T23:42:26.090Z",
      "content": "<p>I Love it, you deserve it  ❤️❤️</p>",
      "rawMarkdown": "I Love it, you deserve it  ❤️❤️"
    },
    {
      "id": 1837021,
      "postDate": "2022-06-29T09:22:37.897Z",
      "content": "<p>After reading your explanation carefully, I found that I benefited a lot, especially from the design of LGBM + manual feature engineering</p>",
      "rawMarkdown": "After reading your explanation carefully, I found that I benefited a lot, especially from the design of LGBM + manual feature engineering"
    },
    {
      "id": 1819921,
      "postDate": "2022-06-14T08:31:14.303Z",
      "content": "<p>on your ranking stage - is it one model that serves for every user ? or rather each user_id has its own model ? <br>\nI guess if you make one model for all users - the user-specific features will not be pronounced. </p>",
      "rawMarkdown": "on your ranking stage - is it one model that serves for every user ? or rather each user_id has its own model ? \nI guess if you make one model for all users - the user-specific features will not be pronounced. ",
      "replies": [
        {
          "id": 1821336,
          "postDate": "2022-06-15T13:02:06.233Z",
          "content": "<p>one model for every user</p>",
          "rawMarkdown": " one model for every user",
          "votes": 1
        }
      ]
    },
    {
      "id": 1818641,
      "postDate": "2022-06-13T02:51:52.237Z",
      "content": "<p>We still have a lot to learn from Gradient Boosted Trees.</p>\n<p>Thanks Senkin13.</p>",
      "rawMarkdown": "We still have a lot to learn from Gradient Boosted Trees.\n\nThanks Senkin13."
    },
    {
      "id": 1818371,
      "postDate": "2022-06-12T15:38:13.343Z",
      "content": "<p>Very helpful and congratulations!</p>",
      "rawMarkdown": "Very helpful and congratulations!"
    },
    {
      "id": 1812540,
      "postDate": "2022-06-06T01:46:42.633Z",
      "content": "<p>How much time it took you to train the model?<br>\nHom much time it took you to resolve all the challenge?</p>",
      "rawMarkdown": "How much time it took you to train the model?\nHom much time it took you to resolve all the challenge?",
      "replies": [
        {
          "id": 1821352,
          "postDate": "2022-06-15T13:14:43.173Z",
          "content": "<p>30 minutes ~ 1 day due to the data size. total 250hours  to resolve all the challenge.</p>",
          "rawMarkdown": "30 minutes ~ 1 day due to the data size. total 250hours  to resolve all the challenge."
        }
      ]
    },
    {
      "id": 1798389,
      "postDate": "2022-05-23T01:39:51.827Z",
      "content": "<p>Thank you for sharing this is really gr8.</p>",
      "rawMarkdown": "Thank you for sharing this is really gr8."
    },
    {
      "id": 1798195,
      "postDate": "2022-05-22T19:00:47.940Z",
      "content": "<p>Just epic!</p>",
      "rawMarkdown": "Just epic!"
    },
    {
      "id": 1797662,
      "postDate": "2022-05-22T08:20:01.243Z",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!"
    },
    {
      "id": 1797129,
      "postDate": "2022-05-21T15:14:17.043Z",
      "content": "<p>dude, great stuff!!</p>",
      "rawMarkdown": "dude, great stuff!!"
    },
    {
      "id": 1796296,
      "postDate": "2022-05-20T17:14:13.357Z",
      "content": "<p>Good job, thx for sharing!</p>",
      "rawMarkdown": "Good job, thx for sharing!"
    },
    {
      "id": 1795937,
      "postDate": "2022-05-20T09:23:27.103Z",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!"
    },
    {
      "id": 1795533,
      "postDate": "2022-05-19T21:43:24.257Z",
      "content": "<p>Hello, I would like to ask you a question with total ignorance of the subject.</p>\n<p>when you mention that you validate in 28 groups, what do you do is divide the clients into 28 groups and get the HitNum@100 of each client, and then average the results of each group?</p>\n<p>If so, should all records from the same client, positive and negative, be in the same group?</p>\n<p>thanks for your answer</p>",
      "rawMarkdown": "Hello, I would like to ask you a question with total ignorance of the subject.\n\nwhen you mention that you validate in 28 groups, what do you do is divide the clients into 28 groups and get the HitNum@100 of each client, and then average the results of each group?\n\nIf so, should all records from the same client, positive and negative, be in the same group?\n\nthanks for your answer",
      "replies": [
        {
          "id": 1809165,
          "postDate": "2022-06-02T13:12:44.207Z",
          "content": "<p>28 groups only for test data, calculate HitNum@100 of all valid data </p>",
          "rawMarkdown": "28 groups only for test data, calculate HitNum@100 of all valid data "
        }
      ]
    },
    {
      "id": 1795003,
      "postDate": "2022-05-19T10:32:22.193Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!"
    },
    {
      "id": 1794745,
      "postDate": "2022-05-19T06:02:10.687Z",
      "content": "<p>Congratulations ! and Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations ! and Thanks for sharing."
    },
    {
      "id": 1794624,
      "postDate": "2022-05-19T02:41:28.217Z",
      "content": "<p>thx for sharing :) </p>",
      "rawMarkdown": "thx for sharing :) "
    },
    {
      "id": 1784469,
      "postDate": "2022-05-11T07:41:22.707Z",
      "content": "<p>Thanks for sharing !<br>\nwhat do cold start items mean ?</p>",
      "rawMarkdown": "Thanks for sharing !\nwhat do cold start items mean ?",
      "replies": [
        {
          "id": 1784621,
          "postDate": "2022-05-11T10:21:07.613Z",
          "content": "<p>cold start means new articles that have not purchased before</p>",
          "rawMarkdown": "cold start means new articles that have not purchased before"
        }
      ]
    },
    {
      "id": 2872308,
      "postDate": "2024-06-14T17:46:28.353Z",
      "content": "<p>Is it possible to share your final AUC/accuracy ?</p>",
      "rawMarkdown": "Is it possible to share your final AUC/accuracy ?",
      "isDeleted": true
    },
    {
      "id": 1809773,
      "postDate": "2022-06-03T03:36:33.067Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1821354,
          "postDate": "2022-06-15T13:18:26.810Z",
          "content": "<ol>\n<li>manually adjust the different top N of different strategies to 100~500 to maximize the hit number.</li>\n<li>we use the scores used in retrieval strategies as features for ranking model.</li>\n</ol>",
          "rawMarkdown": "1. manually adjust the different top N of different strategies to 100~500 to maximize the hit number.\n2. we use the scores used in retrieval strategies as features for ranking model."
        }
      ]
    },
    {
      "id": 1797817,
      "postDate": "2022-05-22T11:55:38.407Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1795135,
      "postDate": "2022-05-19T13:07:18.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1789521,
      "postDate": "2022-05-14T01:36:41.713Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1787325,
      "postDate": "2022-05-13T19:26:27.290Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1784520,
      "postDate": "2022-05-11T08:15:03.570Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1787061,
      "postDate": "2022-05-13T14:47:12.070Z",
      "content": "<p>Great work.</p>\n<p>Thank you for sharing!</p>",
      "rawMarkdown": "Great work.\n\nThank you for sharing!\n\n",
      "votes": 3
    },
    {
      "id": 1793416,
      "postDate": "2022-05-17T21:44:43.107Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!\n",
      "votes": 1
    },
    {
      "id": 1793069,
      "postDate": "2022-05-17T15:14:33.400Z",
      "content": "<p>Great job! Thanks</p>",
      "rawMarkdown": "Great job! Thanks",
      "votes": 1
    },
    {
      "id": 1792964,
      "postDate": "2022-05-17T13:13:01.813Z",
      "content": "<p>Great! Thank you for sharing.</p>",
      "rawMarkdown": "Great! Thank you for sharing.\n",
      "votes": 1
    },
    {
      "id": 1792351,
      "postDate": "2022-05-16T21:24:12.073Z",
      "content": "<p>Nice Work ! Thanks for sharing !</p>",
      "rawMarkdown": "Nice Work ! Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1791833,
      "postDate": "2022-05-16T12:09:45.757Z",
      "content": "<p>Very interesting project; Thank You!</p>",
      "rawMarkdown": "Very interesting project; Thank You!",
      "votes": 1
    },
    {
      "id": 1791508,
      "postDate": "2022-05-16T05:04:12.707Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1791325,
      "postDate": "2022-05-15T22:08:06.473Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": 1
    },
    {
      "id": 1791098,
      "postDate": "2022-05-15T16:34:08.380Z",
      "content": "<p>Great! Thank you for sharing. </p>",
      "rawMarkdown": "Great! Thank you for sharing. ",
      "votes": 1
    },
    {
      "id": 1791027,
      "postDate": "2022-05-15T15:24:29.420Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": 1
    },
    {
      "id": 1790628,
      "postDate": "2022-05-15T06:35:45.070Z",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "rawMarkdown": "Congratulations and thanks for sharing",
      "votes": 1
    },
    {
      "id": 1790595,
      "postDate": "2022-05-15T05:37:07.670Z",
      "content": "<p>Great Work : ) <br>\nThank you for sharing</p>",
      "rawMarkdown": "Great Work : ) \nThank you for sharing",
      "votes": 1
    },
    {
      "id": 1789966,
      "postDate": "2022-05-14T11:48:31.233Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": 1
    },
    {
      "id": 1789799,
      "postDate": "2022-05-14T08:38:33.440Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1787095,
      "postDate": "2022-05-13T15:14:59.287Z",
      "content": "<p>Great work! Thanks for sharing!</p>",
      "rawMarkdown": "Great work! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1787070,
      "postDate": "2022-05-13T14:53:04.400Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1786801,
      "postDate": "2022-05-13T09:50:07.967Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1786726,
      "postDate": "2022-05-13T07:39:37.583Z",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "rawMarkdown": "Congratulations and thanks for sharing",
      "votes": 1
    },
    {
      "id": 1786400,
      "postDate": "2022-05-12T21:07:51.630Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1786354,
      "postDate": "2022-05-12T19:34:15.123Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1785945,
      "postDate": "2022-05-12T14:05:18.793Z",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing! ",
      "votes": 1
    },
    {
      "id": 1785404,
      "postDate": "2022-05-12T04:29:20.957Z",
      "content": "<p>Great Work! Thanks for sharing.</p>",
      "rawMarkdown": "Great Work! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1785321,
      "postDate": "2022-05-12T02:55:06.087Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1785230,
      "postDate": "2022-05-12T00:13:32.433Z",
      "content": "<p>Awesome! Thanks for sharing.</p>",
      "rawMarkdown": "Awesome! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1784896,
      "postDate": "2022-05-11T15:09:57.753Z",
      "content": "<p>Thanks for sharing ! </p>",
      "rawMarkdown": "Thanks for sharing ! ",
      "votes": 1
    },
    {
      "id": 1784884,
      "postDate": "2022-05-11T15:02:16.900Z",
      "content": "<p>Thank you very much for sharing!</p>",
      "rawMarkdown": "Thank you very much for sharing!",
      "votes": 1
    },
    {
      "id": 1784880,
      "postDate": "2022-05-11T14:55:19.300Z",
      "content": "<p>Great Work! Thanks for sharing.</p>",
      "rawMarkdown": "Great Work! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1784668,
      "postDate": "2022-05-11T11:14:08.687Z",
      "content": "<p>Great! Thanks for your sharing!</p>",
      "rawMarkdown": "Great! Thanks for your sharing!",
      "votes": 1
    },
    {
      "id": 1784519,
      "postDate": "2022-05-11T08:14:56.973Z",
      "content": "<p>Great! Thanks for your sharing!</p>",
      "rawMarkdown": "Great! Thanks for your sharing!",
      "votes": 1
    },
    {
      "id": 1784328,
      "postDate": "2022-05-11T05:49:29.063Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1784256,
      "postDate": "2022-05-11T04:12:58.667Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1783388,
      "postDate": "2022-05-10T11:05:10.007Z",
      "content": "<p>Awesome stuff, thanks for sharing 👍</p>",
      "rawMarkdown": "Awesome stuff, thanks for sharing 👍",
      "votes": 1
    },
    {
      "id": 1783292,
      "postDate": "2022-05-10T09:07:08.440Z",
      "content": "<p>Congratulations! thanks for sharing</p>",
      "rawMarkdown": "Congratulations! thanks for sharing",
      "votes": 1
    },
    {
      "id": 1782946,
      "postDate": "2022-05-10T02:33:04.197Z",
      "content": "<p>congrats, thanks for sharing!</p>",
      "rawMarkdown": "congrats, thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1785884,
      "postDate": "2022-05-12T13:21:59.097Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": 2
    },
    {
      "id": 3128124,
      "postDate": "2025-02-19T09:27:26.320Z",
      "content": "<p>great job！</p>",
      "rawMarkdown": "great job！"
    },
    {
      "id": 2073769,
      "postDate": "2022-12-23T11:52:54.073Z",
      "content": "<p>Impressive work! Thanks for sharing.</p>",
      "rawMarkdown": "Impressive work! Thanks for sharing."
    },
    {
      "id": 1812400,
      "postDate": "2022-06-05T18:46:59.183Z",
      "content": "<p>Nice work! Thanks for sharing. </p>",
      "rawMarkdown": "Nice work! Thanks for sharing. "
    },
    {
      "id": 1800000,
      "postDate": "2022-05-24T14:02:40.420Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing"
    },
    {
      "id": 1798905,
      "postDate": "2022-05-23T13:49:01.917Z",
      "content": "<p>Thanks!Great!</p>",
      "rawMarkdown": "Thanks!Great!"
    },
    {
      "id": 1798572,
      "postDate": "2022-05-23T06:43:33.357Z",
      "content": "<p>Great!!!! Thank you for sharing</p>",
      "rawMarkdown": "Great!!!! Thank you for sharing"
    },
    {
      "id": 1798373,
      "postDate": "2022-05-23T01:21:15.830Z",
      "content": "<p>Thanks for sharing:)</p>",
      "rawMarkdown": "Thanks for sharing:)"
    },
    {
      "id": 1797837,
      "postDate": "2022-05-22T12:18:31.197Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1797717,
      "postDate": "2022-05-22T09:24:33.283Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1797515,
      "postDate": "2022-05-22T02:56:18.583Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1796597,
      "postDate": "2022-05-21T01:21:04.537Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing."
    },
    {
      "id": 1796054,
      "postDate": "2022-05-20T12:31:17.847Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1795976,
      "postDate": "2022-05-20T10:23:41.220Z",
      "content": "<p>Great! Thank you for sharing.</p>",
      "rawMarkdown": "Great! Thank you for sharing."
    },
    {
      "id": 1795639,
      "postDate": "2022-05-20T02:07:44.863Z",
      "content": "<p>Congratulations!<br>\nThanks for sharing .</p>",
      "rawMarkdown": "Congratulations!\nThanks for sharing ."
    },
    {
      "id": 1795282,
      "postDate": "2022-05-19T15:42:47.023Z",
      "content": "<p>Thanks for Sharing</p>",
      "rawMarkdown": "Thanks for Sharing\n"
    },
    {
      "id": 1795269,
      "postDate": "2022-05-19T15:33:59.543Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!"
    },
    {
      "id": 1795137,
      "postDate": "2022-05-19T13:08:05.720Z",
      "content": "<p>Great Work! Thanks for sharing</p>",
      "rawMarkdown": "Great Work! Thanks for sharing"
    },
    {
      "id": 1795014,
      "postDate": "2022-05-19T10:46:10.380Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    },
    {
      "id": 1794900,
      "postDate": "2022-05-19T08:50:58.320Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1794853,
      "postDate": "2022-05-19T08:04:30.847Z",
      "content": "<p>Nice work thanks!</p>",
      "rawMarkdown": "Nice work thanks!"
    },
    {
      "id": 1794837,
      "postDate": "2022-05-19T07:53:39.327Z",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!"
    },
    {
      "id": 1794707,
      "postDate": "2022-05-19T04:46:06.110Z",
      "content": "<p>Thank you so much for sharing</p>",
      "rawMarkdown": "Thank you so much for sharing"
    },
    {
      "id": 1794122,
      "postDate": "2022-05-18T14:07:33.073Z",
      "content": "<p>Great work. Thank you.</p>",
      "rawMarkdown": "Great work. Thank you."
    },
    {
      "id": 1793985,
      "postDate": "2022-05-18T12:08:09.197Z",
      "content": "<p>great job! thanks</p>",
      "rawMarkdown": "great job! thanks"
    },
    {
      "id": 1793924,
      "postDate": "2022-05-18T11:05:52.813Z",
      "content": "<p>Great job! Thanks for sharing!</p>",
      "rawMarkdown": "Great job! Thanks for sharing!"
    },
    {
      "id": 1793872,
      "postDate": "2022-05-18T09:48:37.720Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": " Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 1783076,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2022-05-10T05:25:35.497000",
      "content": "<p>It is an amazing win and huge congrats to you!</p>\n<p>Gradient Boosted Trees &gt;= NN in recommender systems.</p>\n<p>Gradient boosted trees do not receive nearly as much coverage in papers &amp; research compared to their capabilities and importance for real-life production use-cases vs Neural Networks! And why? Because they do not resemble our brain and hence no promise of the \"general AI\"?</p>\n<p>I admit, a little bit of self confirmation bias speaks through me, because the LGBM + Manual Feature Engineering is also the technology choice that we made at my company for the recommender system. How many discussions we had that maybe there is some hidden \"gold\" in the user-item interactions… </p>\n<p>Turns out there is none. It is a super strong result to see this win! Thank you guys! You made my day!</p>",
      "votes": 24,
      "replies": [
        {
          "id": 1783375,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-10T10:52:59.917000",
          "content": "<p>Totally agree with you, technology should be appropriate for business not just seems modern or fashionable</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 1784501,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2022-05-11T08:01:50.383000",
          "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> I think one reason for success of gradient boosting models(over nns) for this competition is that we don't have as much compute power and as much time to try&amp;tune as engineering teams in big companys. We also have less data(e.g. exposed but not bought samples). Big companys do use nns for recommendation(though gbms are also used in industry).</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3167004,
              "author_name": "canfaneat",
              "author_url": "",
              "post_date": "2025-04-01T07:58:40.530000",
              "content": "<p>so you mean that gradient boosting models is better than nns ? I wonder if there are any stronger nn models can do better </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1786665,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2022-05-13T05:58:54.443000",
          "content": "<p><a href=\"https://www.kaggle.com/homoalways\" target=\"_blank\">@homoalways</a> Even in big companies, Trees + Manual FE have some big advantages over NNs. Cost, where you point out, is one of them. Also, manually engineered features can be easily debugged. They provide an additional layer of visibility into the model. By contrast, imagine debugging equivalent embeddings (pre final layer) in NNs.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1783627,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2022-05-10T14:32:42.967000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> &amp; <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> !  Great solution. </p>\n<p>Have you tried FIL library for LightGBM inference? In my tests it can be 10x faster than cpu inference.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1784163,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-11T01:33:02.393000",
          "content": "<p>never heard of FIL library,thanks for sharing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1784551,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T08:44:30.517000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1783033,
      "author_name": "biubiuG",
      "author_url": "",
      "post_date": "2022-05-10T04:25:05.173000",
      "content": "<p>thanks for your great solution sharing…looking forward your code sharing..</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1783012,
      "author_name": "sirius",
      "author_url": "",
      "post_date": "2022-05-10T04:07:08.727000",
      "content": "<p>Congratulations! Very solid work<br>\nOne question is what's the improvement of graph embedding in your recalling and ranking? That's a part which I wanna try but didn't have enough time.<br>\nThanks for your sharing.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1783378,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-10T10:55:16.380000",
          "content": "<p>about 0.002 up</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1783507,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-10T12:45:59.947000",
          "content": "<p>Wow, that's really powerful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1784180,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2022-05-11T02:15:10.040000",
      "content": "<p>Very organized pipline and strong win. For some ideas, I also came up with, but because of the ineffeciency of the pipeline and I didn't organize everything as well as you did, so I came up so many bugs all along the way. The idea iteration speed is also not that satisfying. Learn from the best! Thanks for the sharing again!</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1783023,
      "author_name": "tarick.morty",
      "author_url": "",
      "post_date": "2022-05-10T04:17:28.887000",
      "content": "<p>Congratulations on winning the competition, it's really amazing.<br>\nWe were having difficulty dealing with 6 weeks of data, so I can imagine how challenging it'd have been to manage so much of data.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1791221,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-05-15T18:39:57.153000",
      "content": "<p>In your chart, where you write <code>Top N</code>, do you use the same <code>N</code> for each retrieval method?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1791819,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-16T12:02:20.813000",
          "content": "<p>different N by adjusting manually to optimize the hit number</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1792492,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-05-17T02:38:29.267000",
          "content": "<p>Thank you!</p>\n<p>What do you do if some customers don't have 'N' candidates for a given strategy?<br>\nFor example, if you first take top 10 repurchase candidates, what do you do for customers who have less than 20 previously purchased items?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1793842,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-18T09:20:43.393000",
          "content": "<p>we don't need to take same N for each strategy,make sure total candidates is 100.At least we have top 100 popular items for all users.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1785247,
      "author_name": "Rohit Kumar",
      "author_url": "",
      "post_date": "2022-05-12T00:37:25.930000",
      "content": "<p>This was super helpful for a begginer like me. Thanks My Besto Friendo.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1783255,
      "author_name": "Paweł Jankiewicz",
      "author_url": "",
      "post_date": "2022-05-10T08:24:33.553000",
      "content": "<p>Congrats! I like the mention of a feature store. It seems very useful to not recalculate the features every time.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1783364,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-10T10:38:47.730000",
          "content": "<p>Yes,although I spent much time on creating pipeline and feature store,it helped us do more experiments without bugs</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1783477,
          "author_name": "biubiuG",
          "author_url": "",
          "post_date": "2022-05-10T12:26:12.797000",
          "content": "<p>would you please sharing the \" feature store\" demo code ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1784164,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-11T01:34:12.703000",
          "content": "<p>it's just a concept.<br>\nWhat is a feature store?<br>\nThe short version - a data management layer for machine learning that allows to share &amp; discover features and create more effective machine learning pipelines.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1784735,
          "author_name": "biubiuG",
          "author_url": "",
          "post_date": "2022-05-11T12:32:17.583000",
          "content": "<p>well, thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1782955,
      "author_name": "myaun",
      "author_url": "",
      "post_date": "2022-05-10T02:52:55.223000",
      "content": "<p>Congratulations! I didn't see anything special about it; just a great solution where every single process is deeply validated and refined, including optimizing the development! Thanks for sharing!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1794237,
      "author_name": "evancallaghan",
      "author_url": "",
      "post_date": "2022-05-18T15:58:11.693000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1792871,
      "author_name": "Ruiqi Gao",
      "author_url": "",
      "post_date": "2022-05-17T11:33:26.203000",
      "content": "<p>Congratulations!<br>\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model …… Until the recall model is trained with the training set of the first six weeks and the final validation set is ranked to get the result.</p>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1793837,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-18T09:14:51.193000",
          "content": "<p>the simplest model is using last week data as samples of ranking model,recall top N candidates from previous data by rules or model,the overlap of recall data and last week data is positive samples,other recall data is negative samples.<br>\nI sugguest you read the code of kernel someone have publiced their code.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1792502,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-05-17T02:52:21.093000",
      "content": "<p>Congratulations on your first place finish!</p>\n<p>Very happy for you - you've put in years of kaggle work, and hard work finally paid off!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1792498,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-05-17T02:48:10.570000",
      "content": "<p>What retrieval did you use for customers without history?</p>\n<p>Just 12 most popular items, or something more?<br>\nOn your CV, do you know what part of the score was attributable to the new customers?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1793840,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-18T09:17:34.847000",
          "content": "<p>just most popular items.I have not checked but I think everyone have no big difference of new customers' score because it's hard to predict.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1792313,
      "author_name": "Grzesiek Zyśk",
      "author_url": "",
      "post_date": "2022-05-16T20:13:11.610000",
      "content": "<p>Interesting :) Thx for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791811,
      "author_name": "Wahyu Dwi Prasetio",
      "author_url": "",
      "post_date": "2022-05-16T11:53:33.487000",
      "content": "<p>Great work! Thank you for this knowledge</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791010,
      "author_name": "I110143105",
      "author_url": "",
      "post_date": "2022-05-15T15:09:33.527000",
      "content": "<p>Great Work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791007,
      "author_name": "KIYAKIM",
      "author_url": "",
      "post_date": "2022-05-15T15:07:40.357000",
      "content": "<p>Congratulations! Thank you for your sharing! This solution is really nice!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1790882,
      "author_name": "Rashmita Karak",
      "author_url": "",
      "post_date": "2022-05-15T12:48:01.797000",
      "content": "<p>Amazing job! Please share more soon. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1790549,
      "author_name": "Gargeya Sharma",
      "author_url": "",
      "post_date": "2022-05-15T04:44:17.390000",
      "content": "<p>Looks interesting. Thanks for sharing. 👊</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1790535,
      "author_name": "Prakarn Kidngun",
      "author_url": "",
      "post_date": "2022-05-15T04:32:38.947000",
      "content": "<p>Congratulation and Thanks for the sharing 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1790328,
      "author_name": "Violence Detection",
      "author_url": "",
      "post_date": "2022-05-14T19:15:04.820000",
      "content": "<p>great work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1789946,
      "author_name": "Jingping Pei",
      "author_url": "",
      "post_date": "2022-05-14T11:24:11.417000",
      "content": "<p>THX for sharing!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1789569,
      "author_name": "ahmed samy",
      "author_url": "",
      "post_date": "2022-05-14T03:00:26.503000",
      "content": "<p>Great work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787343,
      "author_name": "Naomi Ugwumba",
      "author_url": "",
      "post_date": "2022-05-13T19:47:01.180000",
      "content": "<p>great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787271,
      "author_name": "dxpNinja",
      "author_url": "",
      "post_date": "2022-05-13T18:10:59.220000",
      "content": "<p>Very well explained. Congratulations and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787158,
      "author_name": "Sofía Cavallo",
      "author_url": "",
      "post_date": "2022-05-13T16:05:51.927000",
      "content": "<p>Hi, amazing work!! Thank so much for sharing!! </p>\n<p>Could I ask you a question?<br>\nIn the case of cosine similarity, are you using the closest or another distance to the items bought in the previous weeks?</p>\n<p>Thank you again!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1791827,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-16T12:07:44.490000",
          "content": "<p>cosine similarity between user's latest purchased ,2nd latest purchased items and candidate item</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1787153,
      "author_name": "Edgar Sanchez",
      "author_url": "",
      "post_date": "2022-05-13T16:03:00.137000",
      "content": "<p>Hello, congratulations for the work.</p>\n<p>But I have a question. How does item2item integrate with lightgbm? Is it a feature engineering itself or an intermediate step for another process?</p>\n<p>Thanks for sharing knowledge!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1791832,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2022-05-16T12:09:26.473000",
          "content": "<p>it's feature engineering, cosine similarity between user's latest purchased ,2nd latest purchased items and candidate item</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1786735,
      "author_name": "Karol Szczukowski",
      "author_url": "",
      "post_date": "2022-05-13T08:05:04.417000",
      "content": "<p>great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786667,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T06:00:02.337000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786497,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T00:47:35.707000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786438,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T22:34:18.533000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1786499,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-13T00:52:19.573000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1786582,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-13T03:31:17.913000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1786412,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T21:32:29.380000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786218,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T17:33:59.083000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785998,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T14:43:53.250000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785408,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T04:37:07.027000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785292,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T02:05:57.930000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784919,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T15:34:50.580000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784487,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T07:53:33.230000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1784620,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T10:19:07.110000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1784428,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T07:13:59.130000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784361,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T06:09:49.813000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784232,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T03:41:51.740000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784155,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T01:18:08.877000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784053,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T22:19:57.340000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783928,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T19:27:35.077000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783867,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T18:38:55.087000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783818,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T17:53:01.190000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783797,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T17:18:03.823000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783751,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T16:37:40.437000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783713,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T16:06:21.983000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783712,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T16:05:34.580000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1784166,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T01:42:10.263000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1784554,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T08:45:20.610000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1783658,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T15:00:47.080000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1784165,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T01:38:23.207000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1783593,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T14:02:15.153000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783577,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T13:48:31.497000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783500,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T12:41:04.543000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783447,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T12:03:10.640000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783390,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T11:06:09.973000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783386,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T11:02:44.767000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783359,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T10:34:53.040000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783301,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T09:24:48.520000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783267,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T08:45:10.657000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783253,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T08:22:01.833000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1783365,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-10T10:42:12.550000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1783217,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T07:50:45.507000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783213,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T07:49:40.873000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1783368,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-10T10:45:48.653000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1783210,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T07:49:02.593000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783185,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T07:25:27.993000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783139,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T06:43:10.087000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783137,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T06:42:47.877000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783075,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T05:21:00.677000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783070,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T05:16:21.503000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1783129,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-10T06:27:49.520000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1808756,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-02T06:13:16.267000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1783069,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T05:16:02.860000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783028,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T04:21:06.443000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1783054,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-10T04:59:34.687000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1783125,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-10T06:24:51.923000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1785479,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-12T06:12:30.360000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1782970,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T03:18:39.157000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1782966,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T03:11:54.747000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1782939,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T02:17:46.007000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791444,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-16T03:11:00.140000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1791818,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-16T12:01:05.593000",
          "content": "",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1786816,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T10:22:49.080000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1787066,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-13T14:49:41.460000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1787102,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-13T15:19:07.377000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1785765,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T11:50:57.600000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1785559,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T07:59:22.077000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1785496,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T06:23:32.553000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1785782,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-12T12:09:18.393000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1785828,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-12T12:46:36.090000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1785896,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-12T13:29:07.960000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1792496,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-17T02:46:41.680000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1785319,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T02:51:07.350000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1785497,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-12T06:23:32.603000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1791927,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-16T13:35:04.983000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1782994,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T03:43:50.520000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1782974,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T03:22:55.877000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1782973,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T03:21:41.447000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1783049,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-10T04:53:03.730000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1784787,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T13:29:01.560000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1784881,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T14:57:47.697000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1784912,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T15:28:22.700000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2480501,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-13T10:59:43.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2072382,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-22T04:01:27.960000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2006153,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-27T13:01:46.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1853511,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-12T23:42:26.090000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1837021,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-29T09:22:37.897000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1819921,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-14T08:31:14.303000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1821336,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-15T13:02:06.233000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1818641,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-13T02:51:52.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1818371,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-12T15:38:13.343000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1812540,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-06T01:46:42.633000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1821352,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-15T13:14:43.173000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1798389,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-23T01:39:51.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1798195,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T19:00:47.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1797662,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T08:20:01.243000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1797129,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-21T15:14:17.043000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1796296,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-20T17:14:13.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795937,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-20T09:23:27.103000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795533,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T21:43:24.257000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1809165,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-02T13:12:44.207000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1795003,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T10:32:22.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794745,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T06:02:10.687000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794624,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T02:41:28.217000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1784469,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T07:41:22.707000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1784621,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-11T10:21:07.613000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2872308,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-14T17:46:28.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809773,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T03:36:33.067000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1821354,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-15T13:18:26.810000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1797817,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T11:55:38.407000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795135,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T13:07:18.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1789521,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-14T01:36:41.713000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787325,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T19:26:27.290000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784520,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T08:15:03.570000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787061,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T14:47:12.070000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1793416,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-17T21:44:43.107000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1793069,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-17T15:14:33.400000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1792964,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-17T13:13:01.813000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1792351,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-16T21:24:12.073000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791833,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-16T12:09:45.757000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791508,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-16T05:04:12.707000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791325,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-15T22:08:06.473000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791098,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-15T16:34:08.380000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1791027,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-15T15:24:29.420000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1790628,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-15T06:35:45.070000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1790595,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-15T05:37:07.670000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1789966,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-14T11:48:31.233000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1789799,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-14T08:38:33.440000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787095,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T15:14:59.287000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1787070,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T14:53:04.400000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786801,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T09:50:07.967000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786726,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T07:39:37.583000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786400,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T21:07:51.630000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1786354,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T19:34:15.123000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785945,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T14:05:18.793000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785404,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T04:29:20.957000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785321,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T02:55:06.087000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785230,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T00:13:32.433000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784896,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T15:09:57.753000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784884,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T15:02:16.900000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784880,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T14:55:19.300000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784668,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T11:14:08.687000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784519,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T08:14:56.973000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784328,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T05:49:29.063000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1784256,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T04:12:58.667000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783388,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T11:05:10.007000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1783292,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T09:07:08.440000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1782946,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T02:33:04.197000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1785884,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-12T13:21:59.097000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3128124,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-02-19T09:27:26.320000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2073769,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-23T11:52:54.073000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1812400,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T18:46:59.183000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1800000,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-24T14:02:40.420000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1798905,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-23T13:49:01.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1798572,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-23T06:43:33.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1798373,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-23T01:21:15.830000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1797837,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T12:18:31.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1797717,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T09:24:33.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1797515,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T02:56:18.583000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1796597,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-21T01:21:04.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1796054,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-20T12:31:17.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795976,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-20T10:23:41.220000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795639,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-20T02:07:44.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795282,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T15:42:47.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795269,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T15:33:59.543000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795137,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T13:08:05.720000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1795014,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T10:46:10.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794900,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T08:50:58.320000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794853,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T08:04:30.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794837,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T07:53:39.327000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794707,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-19T04:46:06.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1794122,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-18T14:07:33.073000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1793985,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-18T12:08:09.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1793924,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-18T11:05:52.813000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1793872,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-18T09:48:37.720000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1782928": "Thanks to H&M and Kaggle Team giving us such a wonderful competition.Thanks my team mate @h4211819 finally we got 1st place.This competition has no leak,stable cv-lb correlation,and infinite possibilities to improve, we really enjoyed exploring the accuracy boundaries of  infinite possibilities. \n\n# Overview\nThe most interesting part of this competition is we need to generate train and test by ourself, the candinate generation strategy is the key to beyond limitation of accuracy ,good feature engineering or modeling could be close to the limitation .Our solution is using various retrieval strategies + feature engineering + GBDT, seems simple but powerful.\n\nWe mainly generate recent popular items because fashion changing fast and has seasonality,tried to add cold start items but they will never be ranked to top12 due to lacking of interaction information.\nuser and item interaction information are always the most important of recommendation problem,the features we created are almost interaction features,image and text features didn't help but should be useful for cold start problem.\n\nAlmost 50% users have no transactions in recent 3months,so we created many cumulative features for them,and last week,last month,last season features for active users.\n\nWe use 6 weeks data as train,last week as valid,retrieve 100 caninates for each user,it has stable cv-lb correlation.We focus on improving single lightgbm model to the last week, cv is 0.0430 and lb is 0.0362.At last week I have to rent gcp's big memory server and vast.ai's gpu server to run bigger models to get higher accuracy.\n\n[![overview.png](https://i.postimg.cc/bwBm7GVv/overview.png)](https://postimg.cc/xXLGHd2r)\n\n# Retrieval(aka: candidate generation | recall)\nAt early stage,we focus on increasing the hit number of retrieved 100 candidates,try various strategies to cover more positive sample. \n| WeekNo | HitNum@100 |\n| --- | --- |\n|  2020-09-16|  39142|\n|  2020-09-09|  38427|\n|  2020-09-02|  41019|\n\n# Ranking\n- **Feature Engineering**\nbasiclly,features are created base on retrieval strategies,create user-item interaction for repurchase,create collaborative filtering score for itemcf,create similarity for embedding retrieval,create item count for popularity......\n| Type | Description|\n| --- | --- |\n| Count |  user-item, user-category of last week/month/season/same week of last year/all, time weighted count... |\n| Time|  first,last days of tranactions...|\n| Mean/Max/Min| aggregation of age,price,sales_channel_id... |\n| Difference/Ratio|  difference between age and mean age of who purchased item, ratio of one user's purchased item count and the item's count|\n| Similarity| collaborative filtering score of  item2item, cosine similarity of item2item(word2vec), cosine similarity of user2item(ProNE) |\n\n- **Downsampling**\nIf we retrieve 100~500 candidates for each user,the negative samples number is very big,so negative downsampling is must,we found 1 million ~ 2 million negative samples for each week has better performace .\n`neg_samples = 1000000;seed = 42\ntrain[train['label']>0].append(train[train['label']==0].sample(neg_samples, random_state=seed))\n`\n\n- **Model**\nOur best single model is a lightgbm(cv:0.0441,lb:0.0367),finally we trained 5 lightgbm classifier and 7 catboost classifier for ensemble(lb:0.0371).catboost's lb score is much worse than lightgbm,lightgbm has very stable cv-lb correlation.\n[![cv.png](https://i.postimg.cc/V6rvKhr6/cv.png)](https://postimg.cc/ZBmJRVLt)\n\n# Optimization\n1. we use TreeLite to accelerate lightgbm inference speed (2X faster),catboost-gpu is 30X faster than lightgbm-cpu inference.\n2. transform all the categorical features(including two way) to label encoding,use reduce_mem_usage,\n3. create a **feature store**,save intermediate features files to dictionary , final features to feather,exsiting features will not be create again\n4. split all the users to 28 group,inference simultaneously with multiple servers.\n\n# Machine\nI have a desktop with 128G RAM,64 vcore CPU,TITAN RTX GPU.It's enough to run most of our models to win.My teammate has a desktop with 64G RAM.At last week ,we use 300G RAM gcp instances.\n\n",
    "1783076": "It is an amazing win and huge congrats to you!\n\nGradient Boosted Trees >= NN in recommender systems.\n\nGradient boosted trees do not receive nearly as much coverage in papers & research compared to their capabilities and importance for real-life production use-cases vs Neural Networks! And why? Because they do not resemble our brain and hence no promise of the \"general AI\"?\n\nI admit, a little bit of self confirmation bias speaks through me, because the LGBM + Manual Feature Engineering is also the technology choice that we made at my company for the recommender system. How many discussions we had that maybe there is some hidden \"gold\" in the user-item interactions... \n\nTurns out there is none. It is a super strong result to see this win! Thank you guys! You made my day!",
    "1783627": "Congrats @senkin13 & @h4211819 !  Great solution. \n\nHave you tried FIL library for LightGBM inference? In my tests it can be 10x faster than cpu inference.",
    "1783033": "thanks for your great solution sharing...looking forward your code sharing..",
    "1783012": "Congratulations! Very solid work\nOne question is what's the improvement of graph embedding in your recalling and ranking? That's a part which I wanna try but didn't have enough time.\nThanks for your sharing.",
    "1784180": "Very organized pipline and strong win. For some ideas, I also came up with, but because of the ineffeciency of the pipeline and I didn't organize everything as well as you did, so I came up so many bugs all along the way. The idea iteration speed is also not that satisfying. Learn from the best! Thanks for the sharing again!",
    "1783023": "Congratulations on winning the competition, it's really amazing.\nWe were having difficulty dealing with 6 weeks of data, so I can imagine how challenging it'd have been to manage so much of data.",
    "1791221": "In your chart, where you write `Top N`, do you use the same `N` for each retrieval method?",
    "1785247": "This was super helpful for a begginer like me. Thanks My Besto Friendo.",
    "1783255": "Congrats! I like the mention of a feature store. It seems very useful to not recalculate the features every time.",
    "1782955": "Congratulations! I didn't see anything special about it; just a great solution where every single process is deeply validated and refined, including optimizing the development! Thanks for sharing!",
    "1794237": "Great work!",
    "1792871": "Congratulations!\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model ...... Until the recall model is trained with the training set of the first six weeks and the final validation set is ranked to get the result.\n\nThanks!",
    "1792502": "Congratulations on your first place finish!\n\nVery happy for you - you've put in years of kaggle work, and hard work finally paid off!",
    "1792498": "What retrieval did you use for customers without history?\n\nJust 12 most popular items, or something more?\nOn your CV, do you know what part of the score was attributable to the new customers?",
    "1792313": "Interesting :) Thx for sharing",
    "1791811": "Great work! Thank you for this knowledge",
    "1791010": "Great Work!",
    "1791007": "Congratulations! Thank you for your sharing! This solution is really nice!",
    "1790882": "Amazing job! Please share more soon. ",
    "1790549": "Looks interesting. Thanks for sharing. 👊",
    "1790535": "Congratulation and Thanks for the sharing 👍",
    "1790328": "great work\n",
    "1789946": "THX for sharing!!!\n",
    "1789569": "Great work.",
    "1787343": "great work!",
    "1787271": "Very well explained. Congratulations and thanks for sharing!",
    "1787158": "Hi, amazing work!! Thank so much for sharing!! \n\nCould I ask you a question?\nIn the case of cosine similarity, are you using the closest or another distance to the items bought in the previous weeks?\n\nThank you again!",
    "1787153": "Hello, congratulations for the work.\n\nBut I have a question. How does item2item integrate with lightgbm? Is it a feature engineering itself or an intermediate step for another process?\n\nThanks for sharing knowledge!",
    "1786735": "great work!",
    "1786667": "Great Work!",
    "1786497": "Nice, good work.",
    "1786438": "outstanding performance and a wonderful write up! huge congrats and thank you for sharing your insights!!! 🙏\n\nCould I please ask you how did you handle duplicate purchases? Do you ever predict duplicate items?\n\nI created \"baskets\" of purchases by each customer per week and removed duplicates, but not sure if that is a good way to handle this?\n\nAlso, if I am reading this right, you used some number of weeks for train, used the last week for validation, and used the model trained like this to predict on test. Is my understanding correct? You didn't retrain the models on the full dataset (up unto including the validation week, the last week of the train set) once you found good hyperparams to train with?\n\nThank you very much for all your help on this! 🙂",
    "1786412": "congratulations!",
    "1786218": "Amazing work!!",
    "1785998": "Great work! Congratulations!",
    "1785408": "congratulation",
    "1785292": "Congratulations! Thanks for sharing the solution. ",
    "1784919": "Great Work!",
    "1784487": "Congrats. Did you process your features&recall candidates with gpu using cudf? Or you use cpu to do these? I got cuda out of memory error if I use gpu(p100, 16GB gpu memory) to recall and merge features for all data.",
    "1784428": "공유해 주셔서 감사합니다!",
    "1784361": "Congrats on winner @senkin13 and team!",
    "1784232": "Congratulation! and thank you for sharing！",
    "1784155": "Congratulations and thank you for sharing, this is very insightful.",
    "1784053": "Congratulations!  Thanks for sharing!\nGreat idia about to reduce work with data is to take the latest popular product and according to seasonal characteristics!",
    "1783928": "This is a great write-up! You guys definitely deserve the first place with all the hard work you put in. I love how you experimented with different retrieval strategies and features to get the most out of your model. Keep up the good work!",
    "1783867": "Congratulations!",
    "1783818": "Hats off to you !",
    "1783797": "Congratulations on winning the 1st place! It's mind blowing to see such diverse concepts being used to get the best result possible. For someone like me who is just getting started in this field, this is highly inspiring to put more effort and learn in this bottomless ocean of a field. Thank you.",
    "1783751": "Congratulation! and thank you for sharing yours!",
    "1783713": "Congratulations on the first place and thanks for sharing your approach!",
    "1783712": "Congratulations and thank you for sharing.\nI have a question about GBDT.\nYou said you used LightGBM and CatBoost.\nDidn't you try to use XGBoost?\nIn my environment, XGBoost is as good as LightGBM and can be used for the ensemble.\n",
    "1783658": "Congratulations! I tried deep learning models for recall and rank. But none of the models get a good score. ",
    "1783593": "congratulation! what seems to be weird always be the best of bests ",
    "1783577": "Congratulations! thanks for taking the time to write this up.",
    "1783500": "Congratulation! and thank for sharing yours!",
    "1783447": "Congratulation! Thanks for sharing! 💪💪💪",
    "1783390": "Really well done - thanks for the explanation on approach. Congratulations.",
    "1783386": "learnt a lot from this!\n",
    "1783359": "Congratulations! Thanks for sharing and explanations! and looking forward to a code sharing ;)",
    "1783301": "Congratulation! and thank for sharing yours!",
    "1783267": "Really well done - thanks for the explanation on approach. Congratulations.",
    "1783253": "Congrats on first place!\nI used to use item2vec's similarity feature, but my score doesn't even come close your score!\nIf you don't mind, I'd like to know your thoughts on which particular features gave you the best results.",
    "1783217": "Congratulation! and thank for sharing yours!",
    "1783213": "@senkin13 \nCongratulations! One question - considering stable cv-lb correlation, and using last week as valid ( also assuming that you used last week data as it appeared in the transactions csv, basically in order of purchase date), would it fair to assume that LB data is also based on purchase date? Probably no distinction for same day purchase but otherwise on purchase date.",
    "1783210": "@senkin13 wow fantastic work! congratulations",
    "1783185": "Congratulations!",
    "1783139": "Congratulation for winning!!!\nI also split test_data into 28 groups.it was hard work.😂",
    "1783137": "Congrats on the amazing win @senkin13 and @h4211819 \n\nThanks a lot for the detailed solution sharing. ",
    "1783075": "Congrats for the WIN!",
    "1783070": "Congratulation to winning.\nIf you have paper of ProNE, can you share them?",
    "1783069": "Congratulations!",
    "1783028": "Congratulations! Nice work! Could you elaborate the way to utilize graph embedding to generate candidates? Did you use it for some models as inputs? Thanks for sharing?",
    "1782970": "Congrats for the champinship!",
    "1782966": "Congratulations! ",
    "1782939": "Congrats for your No 1, Teather Zhan and Teather Kan. Great solution and nice write up.",
    "1791444": "Congrats🎉🎉  and thanks for sharing your great work. It's quite impressive😄. \nIf you don't mind, I want to ask some question about your work.\n\n1. How did you trained ProNE graph embedding model using bipartite graph(using two kinds of nodes : user and item)? I've checked the github for ProNE, it's basically using homogeneous information. Can you explain more detail about the preprocessing for bipartite graph using hm dataset or any reference about this?\n\n2. Could you explain more detail about the optimization of the ranking model ? If there are various kinds of candidates from different strategies, how did you set the target variables for each candidates?\n\nAgain, thanks for your sharing!",
    "1786816": "Congratulations and thanks for sharing!\nI have one question:\nhow to calculate `HitNum@100` in the Retrieval module?Is it equal to the number of positive samples in the top100 recall items?\n",
    "1785765": "Congratulations!",
    "1785559": "Epic work!!\n",
    "1785496": "Great work!, what's \"the logistic regression with categorical information\"?",
    "1785319": "Congratulation! and thank you for sharing!\nIf you don't mind, I'd like to know how to calc collaborative filtering score of item2item.\nI would be glad to know if you have any reference materials for the collaborative filtering score of item2item.",
    "1782994": "Congrats on 1st and sharing an impressive solution 🙌 Looking forward to seeing it be expanded on to learn more from it.",
    "1782974": "Nice job both of you two grandmasters.",
    "1782973": "Congratulations! If it isn't my mistake you guys led the competition from the start. That's impressive! Could you please describe the used hardware for your solution?",
    "1784787": "Congratulation! and thank you for sharing！\nAnd one question - How do you merge various retrieval strategies into 100 candidates for each user?\nIn my solution, I just recall 200 candidates from each recall strategy for each user, and then just merge these items, and feed into the rank model. I find when I add more recall strategy, my score does not improve, so I am wondering if there's something wrong with the way I merge each recall candidates(just merge not limit or rough rank).",
    "2480501": "can you please share your full notebook @h4211819. i need it for educational purpose.",
    "2072382": "great solution! looking forward your code sharing..",
    "2006153": "Thank you very much for sharing your work!! respect! But ı clouldnt understand this part \n\"  gender_sale.loc[gender_sale[1]>=0.8, 'gender'] = 1\ngender_sale.loc[gender_sale[2]>=0.8, 'gender'] = 2   \" \n\nCould you explain that part please?",
    "1853511": "I Love it, you deserve it  ❤️❤️",
    "1837021": "After reading your explanation carefully, I found that I benefited a lot, especially from the design of LGBM + manual feature engineering",
    "1819921": "on your ranking stage - is it one model that serves for every user ? or rather each user_id has its own model ? \nI guess if you make one model for all users - the user-specific features will not be pronounced. ",
    "1818641": "We still have a lot to learn from Gradient Boosted Trees.\n\nThanks Senkin13.",
    "1818371": "Very helpful and congratulations!",
    "1812540": "How much time it took you to train the model?\nHom much time it took you to resolve all the challenge?",
    "1798389": "Thank you for sharing this is really gr8.",
    "1798195": "Just epic!",
    "1797662": "Great job!",
    "1797129": "dude, great stuff!!",
    "1796296": "Good job, thx for sharing!",
    "1795937": "Great job!",
    "1795533": "Hello, I would like to ask you a question with total ignorance of the subject.\n\nwhen you mention that you validate in 28 groups, what do you do is divide the clients into 28 groups and get the HitNum@100 of each client, and then average the results of each group?\n\nIf so, should all records from the same client, positive and negative, be in the same group?\n\nthanks for your answer",
    "1795003": "Great work!",
    "1794745": "Congratulations ! and Thanks for sharing.",
    "1794624": "thx for sharing :) ",
    "1784469": "Thanks for sharing !\nwhat do cold start items mean ?",
    "2872308": "Is it possible to share your final AUC/accuracy ?",
    "1809773": "",
    "1797817": "",
    "1795135": "",
    "1789521": "",
    "1787325": "",
    "1784520": "",
    "1787061": "Great work.\n\nThank you for sharing!\n\n",
    "1793416": "Thanks for sharing!\n",
    "1793069": "Great job! Thanks",
    "1792964": "Great! Thank you for sharing.\n",
    "1792351": "Nice Work ! Thanks for sharing !",
    "1791833": "Very interesting project; Thank You!",
    "1791508": "Thanks for sharing.",
    "1791325": "Thank you for sharing.",
    "1791098": "Great! Thank you for sharing. ",
    "1791027": "Thank you for sharing!",
    "1790628": "Congratulations and thanks for sharing",
    "1790595": "Great Work : ) \nThank you for sharing",
    "1789966": "Thank you for sharing!",
    "1789799": "thanks for sharing!",
    "1787095": "Great work! Thanks for sharing!",
    "1787070": "thanks for sharing!",
    "1786801": "Thanks for sharing!",
    "1786726": "Congratulations and thanks for sharing",
    "1786400": "Congratulations and thanks for sharing!",
    "1786354": "Thanks for sharing!",
    "1785945": "Thanks for sharing! ",
    "1785404": "Great Work! Thanks for sharing.",
    "1785321": "Thanks for sharing!",
    "1785230": "Awesome! Thanks for sharing.",
    "1784896": "Thanks for sharing ! ",
    "1784884": "Thank you very much for sharing!",
    "1784880": "Great Work! Thanks for sharing.",
    "1784668": "Great! Thanks for your sharing!",
    "1784519": "Great! Thanks for your sharing!",
    "1784328": "Thanks for sharing !",
    "1784256": "Thanks for sharing !",
    "1783388": "Awesome stuff, thanks for sharing 👍",
    "1783292": "Congratulations! thanks for sharing",
    "1782946": "congrats, thanks for sharing!",
    "1785884": "Thanks for sharing!!",
    "3128124": "great job！",
    "2073769": "Impressive work! Thanks for sharing.",
    "1812400": "Nice work! Thanks for sharing. ",
    "1800000": "thanks for sharing",
    "1798905": "Thanks!Great!",
    "1798572": "Great!!!! Thank you for sharing",
    "1798373": "Thanks for sharing:)",
    "1797837": "Thanks for sharing!",
    "1797717": "Thanks for sharing!",
    "1797515": "Thanks for sharing!",
    "1796597": "Thank you for sharing.",
    "1796054": "Thanks for sharing!",
    "1795976": "Great! Thank you for sharing.",
    "1795639": "Congratulations!\nThanks for sharing .",
    "1795282": "Thanks for Sharing\n",
    "1795269": "Thanks for sharing!!",
    "1795137": "Great Work! Thanks for sharing",
    "1795014": "Thank you for sharing!",
    "1794900": "Thanks for sharing!",
    "1794853": "Nice work thanks!",
    "1794837": "thank you!",
    "1794707": "Thank you so much for sharing",
    "1794122": "Great work. Thank you.",
    "1793985": "great job! thanks",
    "1793924": "Great job! Thanks for sharing!",
    "1793872": " Thank you for sharing."
  }
}