{
  "id": 384022,
  "title": "1st place solution",
  "url": "/competitions/otto-recommender-system/discussion/384022",
  "author_name": "mrkmakr",
  "post_date": "2023-02-06T08:14:44.865000",
  "votes": 207,
  "comment_count": 41,
  "views": 0,
  "content": "<p>Thank you very much for organizing this fun competition.<br>\nThe problem set up was relatively close to my actual work, and I was glad to learn a lot.</p>\n<h2>Candidates</h2>\n<p>The average number of candidates is around 1200.</p>\n<ul>\n<li>visited aids in session</li>\n<li>covisitation matrix<ul>\n<li>use multiple versions with different weighing by type and aggregation period</li>\n<li>apply covisitation matrix at multiple times like beam search</li></ul></li>\n<li>NN that predicts subsequent aids<ul>\n<li>use multiple versions to create candidates and rerank features</li>\n<li>NN structure is MLP or transformer (there was no big difference)</li>\n<li>I tried to focus on samples that are not predicted well</li>\n<li>I used the same embedding for x_aid and y_aid.</li>\n<li>I used multiple aids in future as positive targets.</li>\n<li>I used prediction target aid type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.</li>\n<li>some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.</li></ul></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Fb39461ef4b594a17799b4b06f1ae6fb2%2F2023-02-06%2017.12.38.png?generation=1675671185056988&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Faff8177d9e576a6e40a08ed9f80f2140%2F2023-02-06%2017.04.33.png?generation=1675670729353616&amp;alt=media\" alt=\"\"></p>\n<h2>Reranker</h2>\n<h3>model</h3>\n<p>single LGBMRanker : LB 0.604  <br>\nensemble of 9 LGBMRankers with different hyperparameters : LB 0.605  <br>\nI performed ensemble by averaging the predicted scores of the rankers.  </p>\n<ul>\n<li>I haven't tested if this is a better method than voting etc.</li>\n</ul>\n<h3>features</h3>\n<ul>\n<li>session * aid<ul>\n<li>rank by covisitation matrix at candidate generation</li>\n<li>cosine similarity by NN at candidate generation</li>\n<li>aid info in the session (when it appeared, what type it is, etc)</li></ul></li>\n<li>aid<ul>\n<li>popularity of aids<ul>\n<li>It worked well when ranked</li>\n<li>calculated by multiple time windows</li></ul></li>\n<li>ratio of types</li></ul></li>\n<li>session<ul>\n<li>length</li>\n<li>aid dupplication rate</li>\n<li>ts between the last aid and the second last aid</li></ul></li>\n</ul>\n<p>about 200 features were created  <br>\nselect about 100 features for each target by lgbm gain importance to reduce memory usage  </p>\n<h4>negative sampling rate</h4>\n<p>clicks : 5% <br>\ncarts  : 25%  <br>\norders : 40%  <br>\nI set these values so that the training data can be handled by my machine (the data size is around 35GB for each).</p>\n<h2>Cv strategy</h2>\n<p>I followed radek's set up. <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991</a>  <br>\nI can get almost perfect correlation between local validation and LB.  <br>\nFor quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.</p>\n<h2>ablation study</h2>\n<p>ablation study by local validation.  <br>\nInformation that is involved in both candidate generation and reranker features is removed from both.</p>\n<table>\n<thead>\n<tr>\n<th>condition</th>\n<th>clicks_recall@20</th>\n<th>carts_recall@20</th>\n<th>orders_recall@20</th>\n<th>weighted_recall@20</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>my solution (LB604)</td>\n<td>0.556607</td>\n<td>0.436375</td>\n<td>0.669644</td>\n<td>0.588359</td>\n</tr>\n<tr>\n<td>without visited aid</td>\n<td>0.555677</td>\n<td>0.435616</td>\n<td>0.666456</td>\n<td>0.586126</td>\n</tr>\n<tr>\n<td>without covisitation</td>\n<td>0.547493</td>\n<td>0.430180</td>\n<td>0.665553</td>\n<td>0.583136</td>\n</tr>\n<tr>\n<td>without nn</td>\n<td>0.544811</td>\n<td>0.429904</td>\n<td>0.666004</td>\n<td>0.583055</td>\n</tr>\n<tr>\n<td>without aid feats</td>\n<td>0.550472</td>\n<td>0.433442</td>\n<td>0.666275</td>\n<td>0.584845</td>\n</tr>\n<tr>\n<td>without session feats</td>\n<td>0.555922</td>\n<td>0.435805</td>\n<td>0.669734</td>\n<td>0.588174</td>\n</tr>\n<tr>\n<td>only single nn</td>\n<td>0.532279</td>\n<td>0.410148</td>\n<td>0.564768</td>\n<td>0.515133</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2131588,
      "postDate": "2023-02-06T08:14:44.867Z",
      "content": "<p>Thank you very much for organizing this fun competition.<br>\nThe problem set up was relatively close to my actual work, and I was glad to learn a lot.</p>\n<h2>Candidates</h2>\n<p>The average number of candidates is around 1200.</p>\n<ul>\n<li>visited aids in session</li>\n<li>covisitation matrix<ul>\n<li>use multiple versions with different weighing by type and aggregation period</li>\n<li>apply covisitation matrix at multiple times like beam search</li></ul></li>\n<li>NN that predicts subsequent aids<ul>\n<li>use multiple versions to create candidates and rerank features</li>\n<li>NN structure is MLP or transformer (there was no big difference)</li>\n<li>I tried to focus on samples that are not predicted well</li>\n<li>I used the same embedding for x_aid and y_aid.</li>\n<li>I used multiple aids in future as positive targets.</li>\n<li>I used prediction target aid type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.</li>\n<li>some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.</li></ul></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Fb39461ef4b594a17799b4b06f1ae6fb2%2F2023-02-06%2017.12.38.png?generation=1675671185056988&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Faff8177d9e576a6e40a08ed9f80f2140%2F2023-02-06%2017.04.33.png?generation=1675670729353616&amp;alt=media\" alt=\"\"></p>\n<h2>Reranker</h2>\n<h3>model</h3>\n<p>single LGBMRanker : LB 0.604  <br>\nensemble of 9 LGBMRankers with different hyperparameters : LB 0.605  <br>\nI performed ensemble by averaging the predicted scores of the rankers.  </p>\n<ul>\n<li>I haven't tested if this is a better method than voting etc.</li>\n</ul>\n<h3>features</h3>\n<ul>\n<li>session * aid<ul>\n<li>rank by covisitation matrix at candidate generation</li>\n<li>cosine similarity by NN at candidate generation</li>\n<li>aid info in the session (when it appeared, what type it is, etc)</li></ul></li>\n<li>aid<ul>\n<li>popularity of aids<ul>\n<li>It worked well when ranked</li>\n<li>calculated by multiple time windows</li></ul></li>\n<li>ratio of types</li></ul></li>\n<li>session<ul>\n<li>length</li>\n<li>aid dupplication rate</li>\n<li>ts between the last aid and the second last aid</li></ul></li>\n</ul>\n<p>about 200 features were created  <br>\nselect about 100 features for each target by lgbm gain importance to reduce memory usage  </p>\n<h4>negative sampling rate</h4>\n<p>clicks : 5% <br>\ncarts  : 25%  <br>\norders : 40%  <br>\nI set these values so that the training data can be handled by my machine (the data size is around 35GB for each).</p>\n<h2>Cv strategy</h2>\n<p>I followed radek's set up. <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991</a>  <br>\nI can get almost perfect correlation between local validation and LB.  <br>\nFor quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.</p>\n<h2>ablation study</h2>\n<p>ablation study by local validation.  <br>\nInformation that is involved in both candidate generation and reranker features is removed from both.</p>\n<table>\n<thead>\n<tr>\n<th>condition</th>\n<th>clicks_recall@20</th>\n<th>carts_recall@20</th>\n<th>orders_recall@20</th>\n<th>weighted_recall@20</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>my solution (LB604)</td>\n<td>0.556607</td>\n<td>0.436375</td>\n<td>0.669644</td>\n<td>0.588359</td>\n</tr>\n<tr>\n<td>without visited aid</td>\n<td>0.555677</td>\n<td>0.435616</td>\n<td>0.666456</td>\n<td>0.586126</td>\n</tr>\n<tr>\n<td>without covisitation</td>\n<td>0.547493</td>\n<td>0.430180</td>\n<td>0.665553</td>\n<td>0.583136</td>\n</tr>\n<tr>\n<td>without nn</td>\n<td>0.544811</td>\n<td>0.429904</td>\n<td>0.666004</td>\n<td>0.583055</td>\n</tr>\n<tr>\n<td>without aid feats</td>\n<td>0.550472</td>\n<td>0.433442</td>\n<td>0.666275</td>\n<td>0.584845</td>\n</tr>\n<tr>\n<td>without session feats</td>\n<td>0.555922</td>\n<td>0.435805</td>\n<td>0.669734</td>\n<td>0.588174</td>\n</tr>\n<tr>\n<td>only single nn</td>\n<td>0.532279</td>\n<td>0.410148</td>\n<td>0.564768</td>\n<td>0.515133</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thank you very much for organizing this fun competition.\nThe problem set up was relatively close to my actual work, and I was glad to learn a lot.\n\n## Candidates\nThe average number of candidates is around 1200.\n- visited aids in session\n- covisitation matrix\n  - use multiple versions with different weighing by type and aggregation period\n  - apply covisitation matrix at multiple times like beam search\n- NN that predicts subsequent aids\n   - use multiple versions to create candidates and rerank features\n   - NN structure is MLP or transformer (there was no big difference)\n   - I tried to focus on samples that are not predicted well\n   - I used the same embedding for x_aid and y_aid.\n   - I used multiple aids in future as positive targets.\n   - I used prediction target aid type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.\n   - some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Fb39461ef4b594a17799b4b06f1ae6fb2%2F2023-02-06%2017.12.38.png?generation=1675671185056988&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Faff8177d9e576a6e40a08ed9f80f2140%2F2023-02-06%2017.04.33.png?generation=1675670729353616&alt=media)\n\n\n## Reranker\n### model\n\nsingle LGBMRanker : LB 0.604  \nensemble of 9 LGBMRankers with different hyperparameters : LB 0.605  \nI performed ensemble by averaging the predicted scores of the rankers.  \n- I haven't tested if this is a better method than voting etc.\n\n### features\n- session * aid\n    - rank by covisitation matrix at candidate generation\n    - cosine similarity by NN at candidate generation\n    - aid info in the session (when it appeared, what type it is, etc)\n- aid\n    - popularity of aids\n        - It worked well when ranked\n        - calculated by multiple time windows\n    - ratio of types\n- session\n    - length\n    - aid dupplication rate\n    - ts between the last aid and the second last aid\n\nabout 200 features were created  \nselect about 100 features for each target by lgbm gain importance to reduce memory usage  \n\n#### negative sampling rate\nclicks : 5% \ncarts  : 25%  \norders : 40%  \nI set these values so that the training data can be handled by my machine (the data size is around 35GB for each).\n\n## Cv strategy\nI followed radek's set up. https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991  \nI can get almost perfect correlation between local validation and LB.  \nFor quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.\n\n## ablation study\nablation study by local validation.  \nInformation that is involved in both candidate generation and reranker features is removed from both.\n\n|  condition  |  clicks_recall@20  | carts_recall@20 | orders_recall@20 | weighted_recall@20 |\n| ----------- | -----------------  | --------------- | ---------------- | ------------------ |\n|  my solution (LB604)             | 0.556607 | 0.436375 | 0.669644 | 0.588359 |\n|  without visited aid             | 0.555677 | 0.435616 | 0.666456 | 0.586126 |\n|  without covisitation            | 0.547493 | 0.430180 | 0.665553 | 0.583136 |\n|  without nn                      | 0.544811 | 0.429904 | 0.666004 | 0.583055 |\n|  without aid feats               | 0.550472 | 0.433442 | 0.666275 | 0.584845 |\n|  without session feats           | 0.555922 | 0.435805 | 0.669734 | 0.588174 |\n|  only single nn                     | 0.532279 | 0.410148 | 0.564768 | 0.515133 |",
      "votes": 207
    },
    {
      "id": 2131649,
      "postDate": "2023-02-06T08:57:37.410Z",
      "content": "<p>Congrats on the result.  I worked on an extremely similar method.  It is quite incredible how close it is to yours.</p>\n<p>One things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.</p>\n<p>The one thing I didn't do was the topk for sampling negatives. I sampled a large number of negatives (up to 6000). Maybe using topk is what makes your solution so good.</p>\n<p>What surprised me is that NN prediciting next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.</p>",
      "rawMarkdown": "Congrats on the result.  I worked on an extremely similar method.  It is quite incredible how close it is to yours.\n\nOne things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.\n\nThe one thing I didn't do was the topk for sampling negatives. I sampled a large number of negatives (up to 6000). Maybe using topk is what makes your solution so good.\n\nWhat surprised me is that NN prediciting next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.\n",
      "votes": 5,
      "replies": [
        {
          "id": 2131673,
          "postDate": "2023-02-06T09:16:31.897Z",
          "content": "<p>Thank you for your comments!</p>\n<blockquote>\n  <p>One things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.</p>\n</blockquote>\n<p>I actually did same thing.<br>\nI also used the same embedding for x_aid and y_aid.</p>\n<blockquote>\n  <p>What surprised me is that NN predicitin next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.</p>\n</blockquote>\n<p>I used multiple aids in future as positive targets, not just the next click.<br>\nAnd I uses prediction target type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.<br>\n(I have added more explanations to the main text in response to your question)</p>",
          "rawMarkdown": "Thank you for your comments!\n\n> One things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.\n\nI actually did same thing.\nI also used the same embedding for x_aid and y_aid.\n\n\n> What surprised me is that NN predicitin next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.\n\nI used multiple aids in future as positive targets, not just the next click.\nAnd I uses prediction target type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.\n(I have added more explanations to the main text in response to your question)",
          "votes": 3,
          "replies": [
            {
              "id": 2131686,
              "postDate": "2023-02-06T09:29:09.130Z",
              "content": "<p>Me too, in some variant I predict n next events.  I also have some covisit matrix model, where all events in a window of size n must have high cosine and all the rest is negative.</p>\n<p>I have mixed feelings now that I see that I was on the right track!  </p>",
              "rawMarkdown": "Me too, in some variant I predict n next events.  I also have some covisit matrix model, where all events in a window of size n must have high cosine and all the rest is negative.\n\nI have mixed feelings now that I see that I was on the right track!  ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2132943,
      "postDate": "2023-02-07T06:25:40.633Z",
      "content": "<p>Thanks for sharing and congratuations</p>\n<blockquote>\n  <p>For quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.</p>\n</blockquote>\n<p>The time needed for an iteration really depressed me: I thought it's too long to finished an iteration, and I was thinking maybe I should use pyspark with clusters.  Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration? </p>\n<p>Thanks</p>",
      "rawMarkdown": "Thanks for sharing and congratuations\n\n\n> For quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.\n\n\nThe time needed for an iteration really depressed me: I thought it's too long to finished an iteration, and I was thinking maybe I should use pyspark with clusters.  Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration? \n\nThanks",
      "votes": 3,
      "replies": [
        {
          "id": 2132975,
          "postDate": "2023-02-07T06:55:21.933Z",
          "content": "<p>Thank you for your comments!</p>\n<blockquote>\n  <p>Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration?  </p>\n</blockquote>\n<p>The time taken for my sampled local validation is.  </p>\n<ul>\n<li>About 1.5 hours to combine multiple candidates and features from each logic ,  </li>\n<li>About 1.5 hours to train and evaluate a reranking model.  </li>\n</ul>\n<p>Candidates and features by individual logic are calculated once and reused, which is not included in time above. (like training and inference of nn or covisitation matrix, etc.)<br>\nI didn't use any special tools.</p>",
          "rawMarkdown": "\nThank you for your comments!\n\n> Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration?  \n  \nThe time taken for my sampled local validation is.  \n- About 1.5 hours to combine multiple candidates and features from each logic ,  \n- About 1.5 hours to train and evaluate a reranking model.  \n \nCandidates and features by individual logic are calculated once and reused, which is not included in time above. (like training and inference of nn or covisitation matrix, etc.)\n\nI didn't use any special tools.",
          "votes": 6
        },
        {
          "id": 2133594,
          "postDate": "2023-02-07T14:05:55.587Z",
          "content": "<p>Hi xiaoye,<br>\nYou mention special tools in your post. I am curious: do you use special tools in your work, and are there any that you might recommend. Are they usually necessary or are they only for convenience?</p>",
          "rawMarkdown": "Hi xiaoye,\nYou mention special tools in your post. I am curious: do you use special tools in your work, and are there any that you might recommend. Are they usually necessary or are they only for convenience?",
          "replies": [
            {
              "id": 2217709,
              "postDate": "2023-04-11T05:44:08.550Z",
              "content": "<p>Hi. Jack. sorry for the late reply.</p>\n<p>Actually I don't know some special tools. I just found it was too slow to iterate.  </p>\n<p>Polars is one of the tools that I learned from this competition, and it really help. I also found some people used bigquery from GCP</p>\n<p>Hope it helps</p>",
              "rawMarkdown": "Hi. Jack. sorry for the late reply.\n\nActually I don't know some special tools. I just found it was too slow to iterate.  \n\nPolars is one of the tools that I learned from this competition, and it really help. I also found some people used bigquery from GCP\n\nHope it helps"
            }
          ]
        }
      ]
    },
    {
      "id": 2134170,
      "postDate": "2023-02-07T19:59:22.430Z",
      "content": "<p>Congrats!! Nice writeup thanks for the insights and the ablation study! </p>",
      "rawMarkdown": "Congrats!! Nice writeup thanks for the insights and the ablation study! ",
      "votes": 1
    },
    {
      "id": 2131854,
      "postDate": "2023-02-06T12:48:11.620Z",
      "content": "<p>Great work on going solo! Congratulations on earning a gold medal on your own!</p>",
      "rawMarkdown": "Great work on going solo! Congratulations on earning a gold medal on your own!",
      "votes": 1
    },
    {
      "id": 2131670,
      "postDate": "2023-02-06T09:12:35.397Z",
      "content": "<p>Congrats on the win! It is very helpful writeup!<br>\nBTW, I have a few questions.</p>\n<ol>\n<li>There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?</li>\n<li>What was the Recall when forecasting with NN only?</li>\n</ol>",
      "rawMarkdown": "Congrats on the win! It is very helpful writeup!\nBTW, I have a few questions.\n1. There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?\n2. What was the Recall when forecasting with NN only?",
      "votes": 1,
      "replies": [
        {
          "id": 2131763,
          "postDate": "2023-02-06T10:29:46.223Z",
          "content": "<p>Thank you for your comments!</p>\n<blockquote>\n  <ol>\n  <li>There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?</li>\n  </ol>\n</blockquote>\n<p>I didn't do anything special.<br>\nI did only common things like perform negative sampling, have all features in uint32 type, and reduce features by importance.<br>\nI used 256GB RAM machine for training with full data.</p>\n<blockquote>\n  <ol>\n  <li>What was the Recall when forecasting with NN only?</li>\n  </ol>\n</blockquote>\n<p>I have added NN only local score in the ablation study section.</p>",
          "rawMarkdown": "Thank you for your comments!\n\n> 1. There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?\n\nI didn't do anything special.\nI did only common things like perform negative sampling, have all features in uint32 type, and reduce features by importance.\nI used 256GB RAM machine for training with full data.\n\n> 2. What was the Recall when forecasting with NN only?\n\nI have added NN only local score in the ablation study section.",
          "votes": 2,
          "replies": [
            {
              "id": 3081060,
              "postDate": "2024-12-26T06:41:03.940Z",
              "content": "<p>Congrats on the win! It is very helpful writeup! Excuse me, is ablation research a controlled variable?</p>",
              "rawMarkdown": "Congrats on the win! It is very helpful writeup! Excuse me, is ablation research a controlled variable?"
            }
          ]
        }
      ]
    },
    {
      "id": 2131639,
      "postDate": "2023-02-06T08:52:52.673Z",
      "content": "<p>Thank you for the write-up! The ablation study is a fantastic source of knowledge!</p>",
      "rawMarkdown": "Thank you for the write-up! The ablation study is a fantastic source of knowledge!",
      "votes": 1
    },
    {
      "id": 2132193,
      "postDate": "2023-02-06T16:48:40.183Z",
      "content": "<p>Thanks for sharing. Awesome work. Solo rank 1 is a rare phenomenon. 🎉</p>",
      "rawMarkdown": "Thanks for sharing. Awesome work. Solo rank 1 is a rare phenomenon. 🎉",
      "votes": 2
    },
    {
      "id": 2131781,
      "postDate": "2023-02-06T11:10:35.087Z",
      "content": "<p>Congratulations!<br>\nThe number of candidates, ~1200, is very impressive! Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.<br>\nAnd in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?</p>",
      "rawMarkdown": "Congratulations!\nThe number of candidates, ~1200, is very impressive! Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.\nAnd in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?",
      "votes": 2,
      "replies": [
        {
          "id": 2131813,
          "postDate": "2023-02-06T12:01:10.810Z",
          "content": "<blockquote>\n  <p>The number of candidates, ~1200, is very impressive! </p>\n</blockquote>\n<p>the 1200 candidates include candidates for clicks, carts and orders. <br>\nI didn’t explicitly separate which candidate is for which type, and this may be one reason of the very big candidates number.</p>\n<blockquote>\n  <p>Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.</p>\n</blockquote>\n<p>I didnt test it.<br>\nThis is sligtly difficult to conduct a good experiment now.<br>\nI dont know a good way to reduce candidates from multiple strategies by not using ranker ML model.</p>\n<blockquote>\n  <p>And in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?</p>\n</blockquote>\n<p>assuming hh = ts // 3600 // 1000,</p>\n<ul>\n<li>hh (categorical feature)</li>\n<li>hh // 24 (categorical feature)</li>\n<li>sin(hh % 24 / 24 * np.pi * 2) (numerical feature)</li>\n<li>cos(hh % 24 / 24 * np.pi * 2) (numerical feature)</li>\n</ul>\n<p>I have these values for each aid in each session, and I concat these time features timeseries with aid embedding timeseries.</p>",
          "rawMarkdown": "> The number of candidates, ~1200, is very impressive! \n\nthe 1200 candidates include candidates for clicks, carts and orders. \nI didn’t explicitly separate which candidate is for which type, and this may be one reason of the very big candidates number.\n\n> Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.\n\nI didnt test it.\nThis is sligtly difficult to conduct a good experiment now.\nI dont know a good way to reduce candidates from multiple strategies by not using ranker ML model.\n\n> And in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?\n\nassuming hh = ts // 3600 // 1000,\n- hh (categorical feature)\n- hh // 24 (categorical feature)\n- sin(hh % 24 / 24 * np.pi * 2) (numerical feature)\n- cos(hh % 24 / 24 * np.pi * 2) (numerical feature)\n\nI have these values for each aid in each session, and I concat these time features timeseries with aid embedding timeseries.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2131696,
      "postDate": "2023-02-06T09:40:54.843Z",
      "content": "<p>very impressive ! I trained a GNN relatively similar to your method, the difference is that I considered an heterogenous graph with only USERs and their visited AIDs, I didn't exploit the item CF part.<br>\n thanks for sharing.</p>",
      "rawMarkdown": "very impressive ! I trained a GNN relatively similar to your method, the difference is that I considered an heterogenous graph with only USERs and their visited AIDs, I didn't exploit the item CF part.\n thanks for sharing.",
      "votes": 2
    },
    {
      "id": 3081080,
      "postDate": "2024-12-26T07:12:35.107Z",
      "content": "<p>Thank you for sharing, it was very rewarding, but there are still some areas that I don't quite understand.</p>\n<p>1.I was wondering how to generate candidates, that is, how to imerge the results of different recall methods, such as visited aids in session,covisitation matrix and NN that predicts subsequent aids.<br>\n2.you mentioned in nn models,  some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.<br>\nI don't quite understand this method, non visited aids should be negative samples. It means some models use non visited aids as negative samples, and some use both non visited aids and random samples. Is my understanding correct and Is this to avoid confusion?~</p>",
      "rawMarkdown": "Thank you for sharing, it was very rewarding, but there are still some areas that I don't quite understand.\n\n1.I was wondering how to generate candidates, that is, how to imerge the results of different recall methods, such as visited aids in session,covisitation matrix and NN that predicts subsequent aids.\n2.you mentioned in nn models,  some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.\nI don't quite understand this method, non visited aids should be negative samples. It means some models use non visited aids as negative samples, and some use both non visited aids and random samples. Is my understanding correct and Is this to avoid confusion?~"
    },
    {
      "id": 2657257,
      "postDate": "2024-02-18T12:01:25.600Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mrkmakr\" target=\"_blank\">@mrkmakr</a> ,</p>\n<p>Can you share the model latency when inferring results for a single user session ? Thanks ! </p>",
      "rawMarkdown": "Hi @mrkmakr ,\n\nCan you share the model latency when inferring results for a single user session ? Thanks ! "
    },
    {
      "id": 2341069,
      "postDate": "2023-07-11T20:48:29.647Z",
      "content": "<p><a href=\"https://www.kaggle.com/mrkmakr\" target=\"_blank\">@mrkmakr</a> hey, awesome write up! Can you explain the choice of <code>(avg + min) / 2</code> to aggregate the positive cosine similarities?</p>",
      "rawMarkdown": "@mrkmakr hey, awesome write up! Can you explain the choice of `(avg + min) / 2` to aggregate the positive cosine similarities?"
    },
    {
      "id": 2165510,
      "postDate": "2023-03-02T08:48:30.877Z",
      "content": "<p>good job, learned a lot!</p>",
      "rawMarkdown": "good job, learned a lot!"
    },
    {
      "id": 2165117,
      "postDate": "2023-03-02T02:15:27.793Z",
      "content": "<p>Thanks for sharing and congratuations. \"NN structure is MLP or transformer (there was no big difference)\"<br>\ncan you explain how to use transformer in this competition. Thank you very much</p>",
      "rawMarkdown": "Thanks for sharing and congratuations. \"NN structure is MLP or transformer (there was no big difference)\"\ncan you explain how to use transformer in this competition. Thank you very much"
    },
    {
      "id": 2142264,
      "postDate": "2023-02-13T12:16:47.023Z",
      "content": "<p>effective and solid methods, congratulations</p>",
      "rawMarkdown": "effective and solid methods, congratulations"
    },
    {
      "id": 2141769,
      "postDate": "2023-02-13T04:52:10.233Z",
      "content": "<p>WOW, congrats on first place! Your solution sounds like a well-designed and thoroughly tested approach, taking into consideration both the candidate generation and the reranker. It's impressive that you were able to create over 200 features, and it's great to see the results of your ablation study and how removing certain elements affected the performance.</p>\n<p>Your focus on negative sampling rates and finding a balance between having enough data to train your model and being able to handle the data on your machine was smart. Thanks for sharing!</p>",
      "rawMarkdown": "WOW, congrats on first place! Your solution sounds like a well-designed and thoroughly tested approach, taking into consideration both the candidate generation and the reranker. It's impressive that you were able to create over 200 features, and it's great to see the results of your ablation study and how removing certain elements affected the performance.\n\nYour focus on negative sampling rates and finding a balance between having enough data to train your model and being able to handle the data on your machine was smart. Thanks for sharing!"
    },
    {
      "id": 2140187,
      "postDate": "2023-02-11T14:26:45.573Z",
      "content": "<p>Let's not boast about the others. The flow chart is really clear, which is also great👍👍</p>",
      "rawMarkdown": "Let's not boast about the others. The flow chart is really clear, which is also great👍👍"
    },
    {
      "id": 2136405,
      "postDate": "2023-02-09T10:09:33.443Z",
      "content": "<p>Nice approach to the solution</p>",
      "rawMarkdown": "Nice approach to the solution"
    },
    {
      "id": 2135954,
      "postDate": "2023-02-09T01:30:04.207Z",
      "content": "<p>Congrats!! Thank you for the great work!</p>",
      "rawMarkdown": "Congrats!! Thank you for the great work!"
    },
    {
      "id": 2135866,
      "postDate": "2023-02-08T22:56:25.673Z",
      "content": "<p>Congrats! Really nice write up!</p>",
      "rawMarkdown": "Congrats! Really nice write up!"
    },
    {
      "id": 2135708,
      "postDate": "2023-02-08T20:24:05.693Z",
      "content": "<p>Hi! Can you tell me which site/tool you made the flowchart drawings on? Congratz, by the way :)</p>",
      "rawMarkdown": "Hi! Can you tell me which site/tool you made the flowchart drawings on? Congratz, by the way :)"
    },
    {
      "id": 2134283,
      "postDate": "2023-02-07T21:40:06.523Z",
      "content": "<p><a href=\"https://www.kaggle.com/mrkmakr\" target=\"_blank\">@mrkmakr</a> really amazing work!! I really love to give attention to every detail and try to understand 1st place solutions. Thank you for sharing🎉💪🤘</p>",
      "rawMarkdown": "@mrkmakr really amazing work!! I really love to give attention to every detail and try to understand 1st place solutions. Thank you for sharing🎉💪🤘"
    },
    {
      "id": 2132133,
      "postDate": "2023-02-06T16:05:07.650Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!"
    },
    {
      "id": 2131687,
      "postDate": "2023-02-06T09:31:35.840Z",
      "content": "<p>Great solution - clean and effective. Thanks for sharing!</p>",
      "rawMarkdown": "Great solution - clean and effective. Thanks for sharing!"
    },
    {
      "id": 2132547,
      "postDate": "2023-02-06T21:36:51.983Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2647769,
      "postDate": "2024-02-11T18:42:11.260Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 2220225,
      "postDate": "2023-04-13T07:53:26.403Z",
      "content": "<p>Congrats, thanks for sharing</p>",
      "rawMarkdown": "Congrats, thanks for sharing"
    },
    {
      "id": 2136296,
      "postDate": "2023-02-09T08:19:01.600Z",
      "content": "<p>Thanks for sharing and congratuations</p>",
      "rawMarkdown": "Thanks for sharing and congratuations"
    },
    {
      "id": 2136200,
      "postDate": "2023-02-09T06:51:20.033Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing\n"
    },
    {
      "id": 2136103,
      "postDate": "2023-02-09T05:01:18.173Z",
      "content": "<p>Thanks for sharing and congrats!</p>",
      "rawMarkdown": "Thanks for sharing and congrats!"
    },
    {
      "id": 2135147,
      "postDate": "2023-02-08T13:33:28.013Z",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)"
    },
    {
      "id": 2134700,
      "postDate": "2023-02-08T07:49:29.193Z",
      "content": "<p>Thanks for sharing. </p>",
      "rawMarkdown": "Thanks for sharing. "
    },
    {
      "id": 2133774,
      "postDate": "2023-02-07T15:35:16.553Z",
      "content": "<p>cool! Thanks for sharing</p>",
      "rawMarkdown": "cool! Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 2131649,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2023-02-06T08:57:37.410000",
      "content": "<p>Congrats on the result.  I worked on an extremely similar method.  It is quite incredible how close it is to yours.</p>\n<p>One things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.</p>\n<p>The one thing I didn't do was the topk for sampling negatives. I sampled a large number of negatives (up to 6000). Maybe using topk is what makes your solution so good.</p>\n<p>What surprised me is that NN prediciting next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2131673,
          "author_name": "mrkmakr",
          "author_url": "",
          "post_date": "2023-02-06T09:16:31.897000",
          "content": "<p>Thank you for your comments!</p>\n<blockquote>\n  <p>One things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.</p>\n</blockquote>\n<p>I actually did same thing.<br>\nI also used the same embedding for x_aid and y_aid.</p>\n<blockquote>\n  <p>What surprised me is that NN predicitin next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.</p>\n</blockquote>\n<p>I used multiple aids in future as positive targets, not just the next click.<br>\nAnd I uses prediction target type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.<br>\n(I have added more explanations to the main text in response to your question)</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2131686,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-02-06T09:29:09.130000",
              "content": "<p>Me too, in some variant I predict n next events.  I also have some covisit matrix model, where all events in a window of size n must have high cosine and all the rest is negative.</p>\n<p>I have mixed feelings now that I see that I was on the right track!  </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2132943,
      "author_name": "xiaoye",
      "author_url": "",
      "post_date": "2023-02-07T06:25:40.633000",
      "content": "<p>Thanks for sharing and congratuations</p>\n<blockquote>\n  <p>For quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.</p>\n</blockquote>\n<p>The time needed for an iteration really depressed me: I thought it's too long to finished an iteration, and I was thinking maybe I should use pyspark with clusters.  Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration? </p>\n<p>Thanks</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2132975,
          "author_name": "mrkmakr",
          "author_url": "",
          "post_date": "2023-02-07T06:55:21.933000",
          "content": "<p>Thank you for your comments!</p>\n<blockquote>\n  <p>Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration?  </p>\n</blockquote>\n<p>The time taken for my sampled local validation is.  </p>\n<ul>\n<li>About 1.5 hours to combine multiple candidates and features from each logic ,  </li>\n<li>About 1.5 hours to train and evaluate a reranking model.  </li>\n</ul>\n<p>Candidates and features by individual logic are calculated once and reused, which is not included in time above. (like training and inference of nn or covisitation matrix, etc.)<br>\nI didn't use any special tools.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2133594,
          "author_name": "Jack E Martin",
          "author_url": "",
          "post_date": "2023-02-07T14:05:55.587000",
          "content": "<p>Hi xiaoye,<br>\nYou mention special tools in your post. I am curious: do you use special tools in your work, and are there any that you might recommend. Are they usually necessary or are they only for convenience?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2217709,
              "author_name": "xiaoye",
              "author_url": "",
              "post_date": "2023-04-11T05:44:08.550000",
              "content": "<p>Hi. Jack. sorry for the late reply.</p>\n<p>Actually I don't know some special tools. I just found it was too slow to iterate.  </p>\n<p>Polars is one of the tools that I learned from this competition, and it really help. I also found some people used bigquery from GCP</p>\n<p>Hope it helps</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2134170,
      "author_name": "EnricRovira",
      "author_url": "",
      "post_date": "2023-02-07T19:59:22.430000",
      "content": "<p>Congrats!! Nice writeup thanks for the insights and the ablation study! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2131854,
      "author_name": "Taha Muhammad Haider",
      "author_url": "",
      "post_date": "2023-02-06T12:48:11.620000",
      "content": "<p>Great work on going solo! Congratulations on earning a gold medal on your own!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2131670,
      "author_name": "shimacos",
      "author_url": "",
      "post_date": "2023-02-06T09:12:35.397000",
      "content": "<p>Congrats on the win! It is very helpful writeup!<br>\nBTW, I have a few questions.</p>\n<ol>\n<li>There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?</li>\n<li>What was the Recall when forecasting with NN only?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2131763,
          "author_name": "mrkmakr",
          "author_url": "",
          "post_date": "2023-02-06T10:29:46.223000",
          "content": "<p>Thank you for your comments!</p>\n<blockquote>\n  <ol>\n  <li>There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?</li>\n  </ol>\n</blockquote>\n<p>I didn't do anything special.<br>\nI did only common things like perform negative sampling, have all features in uint32 type, and reduce features by importance.<br>\nI used 256GB RAM machine for training with full data.</p>\n<blockquote>\n  <ol>\n  <li>What was the Recall when forecasting with NN only?</li>\n  </ol>\n</blockquote>\n<p>I have added NN only local score in the ablation study section.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3081060,
              "author_name": "huang_jia_li",
              "author_url": "",
              "post_date": "2024-12-26T06:41:03.940000",
              "content": "<p>Congrats on the win! It is very helpful writeup! Excuse me, is ablation research a controlled variable?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2131639,
      "author_name": "Piotr Gabrys",
      "author_url": "",
      "post_date": "2023-02-06T08:52:52.673000",
      "content": "<p>Thank you for the write-up! The ablation study is a fantastic source of knowledge!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2132193,
      "author_name": "Parth Tiwary",
      "author_url": "",
      "post_date": "2023-02-06T16:48:40.183000",
      "content": "<p>Thanks for sharing. Awesome work. Solo rank 1 is a rare phenomenon. 🎉</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2131781,
      "author_name": "sirius",
      "author_url": "",
      "post_date": "2023-02-06T11:10:35.087000",
      "content": "<p>Congratulations!<br>\nThe number of candidates, ~1200, is very impressive! Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.<br>\nAnd in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2131813,
          "author_name": "mrkmakr",
          "author_url": "",
          "post_date": "2023-02-06T12:01:10.810000",
          "content": "<blockquote>\n  <p>The number of candidates, ~1200, is very impressive! </p>\n</blockquote>\n<p>the 1200 candidates include candidates for clicks, carts and orders. <br>\nI didn’t explicitly separate which candidate is for which type, and this may be one reason of the very big candidates number.</p>\n<blockquote>\n  <p>Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.</p>\n</blockquote>\n<p>I didnt test it.<br>\nThis is sligtly difficult to conduct a good experiment now.<br>\nI dont know a good way to reduce candidates from multiple strategies by not using ranker ML model.</p>\n<blockquote>\n  <p>And in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?</p>\n</blockquote>\n<p>assuming hh = ts // 3600 // 1000,</p>\n<ul>\n<li>hh (categorical feature)</li>\n<li>hh // 24 (categorical feature)</li>\n<li>sin(hh % 24 / 24 * np.pi * 2) (numerical feature)</li>\n<li>cos(hh % 24 / 24 * np.pi * 2) (numerical feature)</li>\n</ul>\n<p>I have these values for each aid in each session, and I concat these time features timeseries with aid embedding timeseries.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2131696,
      "author_name": "Rayan-aay",
      "author_url": "",
      "post_date": "2023-02-06T09:40:54.843000",
      "content": "<p>very impressive ! I trained a GNN relatively similar to your method, the difference is that I considered an heterogenous graph with only USERs and their visited AIDs, I didn't exploit the item CF part.<br>\n thanks for sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3081080,
      "author_name": "huang_jia_li",
      "author_url": "",
      "post_date": "2024-12-26T07:12:35.107000",
      "content": "<p>Thank you for sharing, it was very rewarding, but there are still some areas that I don't quite understand.</p>\n<p>1.I was wondering how to generate candidates, that is, how to imerge the results of different recall methods, such as visited aids in session,covisitation matrix and NN that predicts subsequent aids.<br>\n2.you mentioned in nn models,  some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.<br>\nI don't quite understand this method, non visited aids should be negative samples. It means some models use non visited aids as negative samples, and some use both non visited aids and random samples. Is my understanding correct and Is this to avoid confusion?~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2657257,
      "author_name": "Ishaan",
      "author_url": "",
      "post_date": "2024-02-18T12:01:25.600000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mrkmakr\" target=\"_blank\">@mrkmakr</a> ,</p>\n<p>Can you share the model latency when inferring results for a single user session ? Thanks ! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2341069,
      "author_name": "JJFurf",
      "author_url": "",
      "post_date": "2023-07-11T20:48:29.647000",
      "content": "<p><a href=\"https://www.kaggle.com/mrkmakr\" target=\"_blank\">@mrkmakr</a> hey, awesome write up! Can you explain the choice of <code>(avg + min) / 2</code> to aggregate the positive cosine similarities?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2165510,
      "author_name": "DDD",
      "author_url": "",
      "post_date": "2023-03-02T08:48:30.877000",
      "content": "<p>good job, learned a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2165117,
      "author_name": "caicai",
      "author_url": "",
      "post_date": "2023-03-02T02:15:27.793000",
      "content": "<p>Thanks for sharing and congratuations. \"NN structure is MLP or transformer (there was no big difference)\"<br>\ncan you explain how to use transformer in this competition. Thank you very much</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2142264,
      "author_name": "Adil Ramadhan",
      "author_url": "",
      "post_date": "2023-02-13T12:16:47.023000",
      "content": "<p>effective and solid methods, congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2141769,
      "author_name": "Ryan",
      "author_url": "",
      "post_date": "2023-02-13T04:52:10.233000",
      "content": "<p>WOW, congrats on first place! Your solution sounds like a well-designed and thoroughly tested approach, taking into consideration both the candidate generation and the reranker. It's impressive that you were able to create over 200 features, and it's great to see the results of your ablation study and how removing certain elements affected the performance.</p>\n<p>Your focus on negative sampling rates and finding a balance between having enough data to train your model and being able to handle the data on your machine was smart. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2140187,
      "author_name": "Benny",
      "author_url": "",
      "post_date": "2023-02-11T14:26:45.573000",
      "content": "<p>Let's not boast about the others. The flow chart is really clear, which is also great👍👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2136405,
      "author_name": "Shahbaz Qaiser",
      "author_url": "",
      "post_date": "2023-02-09T10:09:33.443000",
      "content": "<p>Nice approach to the solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2135954,
      "author_name": "Yuhua Lu",
      "author_url": "",
      "post_date": "2023-02-09T01:30:04.207000",
      "content": "<p>Congrats!! Thank you for the great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2135866,
      "author_name": "01001",
      "author_url": "",
      "post_date": "2023-02-08T22:56:25.673000",
      "content": "<p>Congrats! Really nice write up!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2135708,
      "author_name": "Leandro Destefani",
      "author_url": "",
      "post_date": "2023-02-08T20:24:05.693000",
      "content": "<p>Hi! Can you tell me which site/tool you made the flowchart drawings on? Congratz, by the way :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2134283,
      "author_name": "George Papachristou",
      "author_url": "",
      "post_date": "2023-02-07T21:40:06.523000",
      "content": "<p><a href=\"https://www.kaggle.com/mrkmakr\" target=\"_blank\">@mrkmakr</a> really amazing work!! I really love to give attention to every detail and try to understand 1st place solutions. Thank you for sharing🎉💪🤘</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2132133,
      "author_name": "ArmanGrewal007",
      "author_url": "",
      "post_date": "2023-02-06T16:05:07.650000",
      "content": "<p>Congratulations!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2131687,
      "author_name": "Grigol",
      "author_url": "",
      "post_date": "2023-02-06T09:31:35.840000",
      "content": "<p>Great solution - clean and effective. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2132547,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-06T21:36:51.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2647769,
      "author_name": "Roberto Viegas Dutra",
      "author_url": "",
      "post_date": "2024-02-11T18:42:11.260000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2220225,
      "author_name": "hd00000",
      "author_url": "",
      "post_date": "2023-04-13T07:53:26.403000",
      "content": "<p>Congrats, thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2136296,
      "author_name": "Sounder",
      "author_url": "",
      "post_date": "2023-02-09T08:19:01.600000",
      "content": "<p>Thanks for sharing and congratuations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2136200,
      "author_name": "JEEVANATHAN V",
      "author_url": "",
      "post_date": "2023-02-09T06:51:20.033000",
      "content": "<p>thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2136103,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-09T05:01:18.173000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2135147,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-08T13:33:28.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2134700,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-08T07:49:29.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2133774,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-07T15:35:16.553000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2131588": "Thank you very much for organizing this fun competition.\nThe problem set up was relatively close to my actual work, and I was glad to learn a lot.\n\n## Candidates\nThe average number of candidates is around 1200.\n- visited aids in session\n- covisitation matrix\n  - use multiple versions with different weighing by type and aggregation period\n  - apply covisitation matrix at multiple times like beam search\n- NN that predicts subsequent aids\n   - use multiple versions to create candidates and rerank features\n   - NN structure is MLP or transformer (there was no big difference)\n   - I tried to focus on samples that are not predicted well\n   - I used the same embedding for x_aid and y_aid.\n   - I used multiple aids in future as positive targets.\n   - I used prediction target aid type information when calculating session embedding so that session embedding is adjusted according to the prediction target aid type.\n   - some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Fb39461ef4b594a17799b4b06f1ae6fb2%2F2023-02-06%2017.12.38.png?generation=1675671185056988&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1019365%2Faff8177d9e576a6e40a08ed9f80f2140%2F2023-02-06%2017.04.33.png?generation=1675670729353616&alt=media)\n\n\n## Reranker\n### model\n\nsingle LGBMRanker : LB 0.604  \nensemble of 9 LGBMRankers with different hyperparameters : LB 0.605  \nI performed ensemble by averaging the predicted scores of the rankers.  \n- I haven't tested if this is a better method than voting etc.\n\n### features\n- session * aid\n    - rank by covisitation matrix at candidate generation\n    - cosine similarity by NN at candidate generation\n    - aid info in the session (when it appeared, what type it is, etc)\n- aid\n    - popularity of aids\n        - It worked well when ranked\n        - calculated by multiple time windows\n    - ratio of types\n- session\n    - length\n    - aid dupplication rate\n    - ts between the last aid and the second last aid\n\nabout 200 features were created  \nselect about 100 features for each target by lgbm gain importance to reduce memory usage  \n\n#### negative sampling rate\nclicks : 5% \ncarts  : 25%  \norders : 40%  \nI set these values so that the training data can be handled by my machine (the data size is around 35GB for each).\n\n## Cv strategy\nI followed radek's set up. https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991  \nI can get almost perfect correlation between local validation and LB.  \nFor quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.\n\n## ablation study\nablation study by local validation.  \nInformation that is involved in both candidate generation and reranker features is removed from both.\n\n|  condition  |  clicks_recall@20  | carts_recall@20 | orders_recall@20 | weighted_recall@20 |\n| ----------- | -----------------  | --------------- | ---------------- | ------------------ |\n|  my solution (LB604)             | 0.556607 | 0.436375 | 0.669644 | 0.588359 |\n|  without visited aid             | 0.555677 | 0.435616 | 0.666456 | 0.586126 |\n|  without covisitation            | 0.547493 | 0.430180 | 0.665553 | 0.583136 |\n|  without nn                      | 0.544811 | 0.429904 | 0.666004 | 0.583055 |\n|  without aid feats               | 0.550472 | 0.433442 | 0.666275 | 0.584845 |\n|  without session feats           | 0.555922 | 0.435805 | 0.669734 | 0.588174 |\n|  only single nn                     | 0.532279 | 0.410148 | 0.564768 | 0.515133 |",
    "2131649": "Congrats on the result.  I worked on an extremely similar method.  It is quite incredible how close it is to yours.\n\nOne things I did, not clear if you did it too from your post, is that I used the same embedding for x_aid and y_aid.\n\nThe one thing I didn't do was the topk for sampling negatives. I sampled a large number of negatives (up to 6000). Maybe using topk is what makes your solution so good.\n\nWhat surprised me is that NN prediciting next click were useful for candidate generation for orders and carts as well. I did train a couple of model to predict next cart and next order, but they don't improve much.\n",
    "2132943": "Thanks for sharing and congratuations\n\n\n> For quick iteration of improvements, I conducted experiments by training with 5% of the data and evaluating with other 10% of the data.\n\n\nThe time needed for an iteration really depressed me: I thought it's too long to finished an iteration, and I was thinking maybe I should use pyspark with clusters.  Just curious to know, how long does it take you for one iteration? did you use any tools to accelerate your iteration? \n\nThanks",
    "2134170": "Congrats!! Nice writeup thanks for the insights and the ablation study! ",
    "2131854": "Great work on going solo! Congratulations on earning a gold medal on your own!",
    "2131670": "Congrats on the win! It is very helpful writeup!\nBTW, I have a few questions.\n1. There seem to be a very large number of candidates for 1200. Did you try anything during the training to reduce the memory?\n2. What was the Recall when forecasting with NN only?",
    "2131639": "Thank you for the write-up! The ablation study is a fantastic source of knowledge!",
    "2132193": "Thanks for sharing. Awesome work. Solo rank 1 is a rare phenomenon. 🎉",
    "2131781": "Congratulations!\nThe number of candidates, ~1200, is very impressive! Do you have the CV/LB scores of some smaller num? e.g. 200, as many others do.\nAnd in the NN model, what are the x_timeinfos extractly? Time diff between the action and last action?",
    "2131696": "very impressive ! I trained a GNN relatively similar to your method, the difference is that I considered an heterogenous graph with only USERs and their visited AIDs, I didn't exploit the item CF part.\n thanks for sharing.",
    "3081080": "Thank you for sharing, it was very rewarding, but there are still some areas that I don't quite understand.\n\n1.I was wondering how to generate candidates, that is, how to imerge the results of different recall methods, such as visited aids in session,covisitation matrix and NN that predicts subsequent aids.\n2.you mentioned in nn models,  some models are trained by using only non visited aids as targets to avoid overlapping information with revisitation based candidates and features.\nI don't quite understand this method, non visited aids should be negative samples. It means some models use non visited aids as negative samples, and some use both non visited aids and random samples. Is my understanding correct and Is this to avoid confusion?~",
    "2657257": "Hi @mrkmakr ,\n\nCan you share the model latency when inferring results for a single user session ? Thanks ! ",
    "2341069": "@mrkmakr hey, awesome write up! Can you explain the choice of `(avg + min) / 2` to aggregate the positive cosine similarities?",
    "2165510": "good job, learned a lot!",
    "2165117": "Thanks for sharing and congratuations. \"NN structure is MLP or transformer (there was no big difference)\"\ncan you explain how to use transformer in this competition. Thank you very much",
    "2142264": "effective and solid methods, congratulations",
    "2141769": "WOW, congrats on first place! Your solution sounds like a well-designed and thoroughly tested approach, taking into consideration both the candidate generation and the reranker. It's impressive that you were able to create over 200 features, and it's great to see the results of your ablation study and how removing certain elements affected the performance.\n\nYour focus on negative sampling rates and finding a balance between having enough data to train your model and being able to handle the data on your machine was smart. Thanks for sharing!",
    "2140187": "Let's not boast about the others. The flow chart is really clear, which is also great👍👍",
    "2136405": "Nice approach to the solution",
    "2135954": "Congrats!! Thank you for the great work!",
    "2135866": "Congrats! Really nice write up!",
    "2135708": "Hi! Can you tell me which site/tool you made the flowchart drawings on? Congratz, by the way :)",
    "2134283": "@mrkmakr really amazing work!! I really love to give attention to every detail and try to understand 1st place solutions. Thank you for sharing🎉💪🤘",
    "2132133": "Congratulations!!",
    "2131687": "Great solution - clean and effective. Thanks for sharing!",
    "2132547": "",
    "2647769": "Thanks for sharing",
    "2220225": "Congrats, thanks for sharing",
    "2136296": "Thanks for sharing and congratuations",
    "2136200": "thanks for sharing\n",
    "2136103": "Thanks for sharing and congrats!",
    "2135147": "Thanks for sharing :)",
    "2134700": "Thanks for sharing. ",
    "2133774": "cool! Thanks for sharing"
  }
}