{
  "id": 382851,
  "title": "10th Place Solution",
  "url": "/competitions/otto-recommender-system/writeups/learn-recsys-x-gbdt-10th-place-solution",
  "author_name": "",
  "post_date": "2023-02-10T01:01:55.857Z",
  "votes": 31,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thanks OTTO &amp; Kaggle to host such interesting competition, my solution is straight forward, just candidate generation + reranker, a short summary as below.</p>\n<h6>Candidate Generation</h6>\n<ul>\n<li><strong>re-visit</strong> - all visited items</li>\n<li><strong>co-visit1</strong> - based on the public notebook with 2 improvement. Thanks <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>    <ol>\n<li>candidate generation by weight decay, e.g. a session has N items(idx1 by the reverse ts) &amp; each item has M pairs of co-visit item(idx2 by time weight or type weight), score is calculated as 1 / idx1 / idx2 for each item &amp; candidate pair, then sum the score as candidate score</li>\n<li>co-visit features for rerank model - counts &amp; probs pairwise item co-occurrences, index distance &amp; time distance(consider order or not), etc.</li></ol></li>\n<li><strong>co-visit2(buy2buy)</strong> - only carts or orders</li>\n<li><strong>co-visit3(next item)</strong> - only consider last item when generate candidate</li>\n<li><strong>sknn1</strong> - simple sknn with 100 most similair &amp; recent session </li>\n<li><strong>sknn2</strong> - a variant of sknn named as Sequence and Time Aware Neighborhood, reference <a href=\"https://arxiv.org/abs/1910.12781\" target=\"_blank\">https://arxiv.org/abs/1910.12781</a> for both sknn implementation</li>\n</ul>\n<h6>Re-Rank Model</h6>\n<ul>\n<li><strong>aid features</strong> - simple stat by time window(7, 14, 21)</li>\n<li><strong>session features</strong> - length, aid count, count by type, time to the end of prediction period</li>\n<li><strong>revisit aid X session</strong>, clicks/carts/orders count, absolute/relative position</li>\n<li><strong>co-visit features</strong> - as mentioned before, generate a lot feature to describe pairwise item co-occurrences</li>\n<li><strong>sknn features</strong> - only candidate score/rank</li>\n<li><strong>similarity features</strong> - similarity between candidate item and session by item embedding of word2vec &amp; implicit' s BPR</li>\n</ul>\n<p>Finally, 3 models(clicks/carts/orders) is trained by LightGBM(lambdarank)</p>",
  "messages": [
    {
      "id": "2124826",
      "postDate": "02/01/2023 08:32:34",
      "content": "<p>Thanks OTTO &amp; Kaggle to host such interesting competition, my solution is straight forward, just candidate generation + reranker, a short summary as below.</p>\n<h6>Candidate Generation</h6>\n<ul>\n<li><strong>re-visit</strong> - all visited items</li>\n<li><strong>co-visit1</strong> - based on the public notebook with 2 improvement. Thanks <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>    <ol>\n<li>candidate generation by weight decay, e.g. a session has N items(idx1 by the reverse ts) &amp; each item has M pairs of co-visit item(idx2 by time weight or type weight), score is calculated as 1 / idx1 / idx2 for each item &amp; candidate pair, then sum the score as candidate score</li>\n<li>co-visit features for rerank model - counts &amp; probs pairwise item co-occurrences, index distance &amp; time distance(consider order or not), etc.</li></ol></li>\n<li><strong>co-visit2(buy2buy)</strong> - only carts or orders</li>\n<li><strong>co-visit3(next item)</strong> - only consider last item when generate candidate</li>\n<li><strong>sknn1</strong> - simple sknn with 100 most similair &amp; recent session </li>\n<li><strong>sknn2</strong> - a variant of sknn named as Sequence and Time Aware Neighborhood, reference <a href=\"https://arxiv.org/abs/1910.12781\" target=\"_blank\">https://arxiv.org/abs/1910.12781</a> for both sknn implementation</li>\n</ul>\n<h6>Re-Rank Model</h6>\n<ul>\n<li><strong>aid features</strong> - simple stat by time window(7, 14, 21)</li>\n<li><strong>session features</strong> - length, aid count, count by type, time to the end of prediction period</li>\n<li><strong>revisit aid X session</strong>, clicks/carts/orders count, absolute/relative position</li>\n<li><strong>co-visit features</strong> - as mentioned before, generate a lot feature to describe pairwise item co-occurrences</li>\n<li><strong>sknn features</strong> - only candidate score/rank</li>\n<li><strong>similarity features</strong> - similarity between candidate item and session by item embedding of word2vec &amp; implicit' s BPR</li>\n</ul>\n<p>Finally, 3 models(clicks/carts/orders) is trained by LightGBM(lambdarank)</p>",
      "rawMarkdown": "Thanks OTTO & Kaggle to host such interesting competition, my solution is straight forward, just candidate generation + reranker, a short summary as below.\n\n###### Candidate Generation\n\n- **re-visit** - all visited items\n- **co-visit1** - based on the public notebook with 2 improvement. Thanks @radek1 and @cdeotte    \n    1. candidate generation by weight decay, e.g. a session has N items(idx1 by the reverse ts) & each item has M pairs of co-visit item(idx2 by time weight or type weight), score is calculated as 1 / idx1 / idx2 for each item & candidate pair, then sum the score as candidate score\n    2. co-visit features for rerank model - counts & probs pairwise item co-occurrences, index distance & time distance(consider order or not), etc.\n- **co-visit2(buy2buy)** - only carts or orders\n- **co-visit3(next item)** - only consider last item when generate candidate\n- **sknn1** - simple sknn with 100 most similair & recent session \n- **sknn2** - a variant of sknn named as Sequence and Time Aware Neighborhood, reference https://arxiv.org/abs/1910.12781 for both sknn implementation\n\n###### Re-Rank Model\n\n- **aid features** - simple stat by time window(7, 14, 21)\n- **session features** - length, aid count, count by type, time to the end of prediction period\n- **revisit aid X session**, clicks/carts/orders count, absolute/relative position\n- **co-visit features** - as mentioned before, generate a lot feature to describe pairwise item co-occurrences\n- **sknn features** - only candidate score/rank\n- **similarity features** - similarity between candidate item and session by item embedding of word2vec & implicit' s BPR\n\nFinally, 3 models(clicks/carts/orders) is trained by LightGBM(lambdarank)",
      "votes": null
    },
    {
      "id": "2124886",
      "postDate": "02/01/2023 09:23:15",
      "content": "<p>Congrats for solo gold!</p>",
      "rawMarkdown": "Congrats for solo gold!",
      "votes": null
    },
    {
      "id": "2125043",
      "postDate": "02/01/2023 11:37:03",
      "content": "<p>Congratulation!👍 I have a question about similarity features. Do you present the session embedding by using the average  of the historical aid sequence? or else.  </p>",
      "rawMarkdown": "Congratulation!👍 I have a question about similarity features. Do you present the session embedding by using the average  of the historical aid sequence? or else.",
      "votes": null
    },
    {
      "id": "2125051",
      "postDate": "02/01/2023 11:43:23",
      "content": "<p>Thanks. I calculate the cosine distance between the candidate and last item, and max, min, mean to each item in the session.</p>",
      "rawMarkdown": "Thanks. I calculate the cosine distance between the candidate and last item, and max, min, mean to each item in the session.",
      "votes": null
    },
    {
      "id": "2125061",
      "postDate": "02/01/2023 11:56:09",
      "content": "<p>Thanks, lucky solo, let's find a chance to team up.</p>",
      "rawMarkdown": "Thanks, lucky solo, let's find a chance to team up.",
      "votes": null
    },
    {
      "id": "2125113",
      "postDate": "02/01/2023 12:38:46",
      "content": "<p>You didn't take only 20/50/last day features, used all aids in session?</p>",
      "rawMarkdown": "You didn't take only 20/50/last day features, used all aids in session?",
      "votes": null
    },
    {
      "id": "2125124",
      "postDate": "02/01/2023 12:48:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gongbi\" target=\"_blank\">@gongbi</a> congrats for solo gold!</p>\n<p>one question, which library do you use for sknn?</p>",
      "rawMarkdown": "Hi @gongbi congrats for solo gold!\n\none question, which library do you use for sknn?",
      "votes": null
    },
    {
      "id": "2125281",
      "postDate": "02/01/2023 15:12:32",
      "content": "<p>just implement it by myself, the logic is simple.</p>",
      "rawMarkdown": "just implement it by myself, the logic is simple.",
      "votes": null
    },
    {
      "id": "2125308",
      "postDate": "02/01/2023 15:28:14",
      "content": "<p>Congras and thanks for sharing. May I ask what training data and label are you using ?? RADEK's CV data here -&gt; <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991</a> or you use the whole train and test data.</p>",
      "rawMarkdown": "Congras and thanks for sharing. May I ask what training data and label are you using ?? RADEK's CV data here -> https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991 or you use the whole train and test data.",
      "votes": null
    },
    {
      "id": "2125323",
      "postDate": "02/01/2023 15:34:59",
      "content": "<p>I built my own code to prepare the training data and label, logic should be same. </p>",
      "rawMarkdown": "I built my own code to prepare the training data and label, logic should be same.",
      "votes": null
    },
    {
      "id": "2125396",
      "postDate": "02/01/2023 16:22:20",
      "content": "<p>Yes, i only take last 40 aid, not selected carefully. And I have another group of feature to take carts or orders only.</p>",
      "rawMarkdown": "Yes, i only take last 40 aid, not selected carefully. And I have another group of feature to take carts or orders only.",
      "votes": null
    },
    {
      "id": "2126127",
      "postDate": "02/02/2023 05:33:13",
      "content": "<p>Congrats on your solo gold! Amazing work!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Congrats on your solo gold! Amazing work!\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2126802",
      "postDate": "02/02/2023 13:07:41",
      "content": "<p>What is sknn? What is his full name?</p>",
      "rawMarkdown": "What is sknn? What is his full name?",
      "votes": null
    },
    {
      "id": "2126814",
      "postDate": "02/02/2023 13:20:24",
      "content": "<p>Noted, thanks for your reply</p>",
      "rawMarkdown": "Noted, thanks for your reply",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2124886,
      "author_name": "chenxin1991",
      "author_url": "",
      "post_date": "02/01/2023 09:23:15",
      "content": "<p>Congrats for solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125061,
          "author_name": "gongbi",
          "author_url": "",
          "post_date": "02/01/2023 11:56:09",
          "content": "<p>Thanks, lucky solo, let's find a chance to team up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125043,
      "author_name": "deeeeeeeplearning",
      "author_url": "",
      "post_date": "02/01/2023 11:37:03",
      "content": "<p>Congratulation!👍 I have a question about similarity features. Do you present the session embedding by using the average  of the historical aid sequence? or else.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2125051,
          "author_name": "gongbi",
          "author_url": "",
          "post_date": "02/01/2023 11:43:23",
          "content": "<p>Thanks. I calculate the cosine distance between the candidate and last item, and max, min, mean to each item in the session.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2125113,
              "author_name": "artemfedorov",
              "author_url": "",
              "post_date": "02/01/2023 12:38:46",
              "content": "<p>You didn't take only 20/50/last day features, used all aids in session?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2125396,
                  "author_name": "gongbi",
                  "author_url": "",
                  "post_date": "02/01/2023 16:22:20",
                  "content": "<p>Yes, i only take last 40 aid, not selected carefully. And I have another group of feature to take carts or orders only.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2125124,
      "author_name": "trasibulo",
      "author_url": "",
      "post_date": "02/01/2023 12:48:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gongbi\" target=\"_blank\">@gongbi</a> congrats for solo gold!</p>\n<p>one question, which library do you use for sknn?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125281,
          "author_name": "gongbi",
          "author_url": "",
          "post_date": "02/01/2023 15:12:32",
          "content": "<p>just implement it by myself, the logic is simple.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125308,
      "author_name": "weiqiu",
      "author_url": "",
      "post_date": "02/01/2023 15:28:14",
      "content": "<p>Congras and thanks for sharing. May I ask what training data and label are you using ?? RADEK's CV data here -&gt; <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991</a> or you use the whole train and test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125323,
          "author_name": "gongbi",
          "author_url": "",
          "post_date": "02/01/2023 15:34:59",
          "content": "<p>I built my own code to prepare the training data and label, logic should be same. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2126814,
              "author_name": "weiqiu",
              "author_url": "",
              "post_date": "02/02/2023 13:20:24",
              "content": "<p>Noted, thanks for your reply</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2126127,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "02/02/2023 05:33:13",
      "content": "<p>Congrats on your solo gold! Amazing work!</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2126802,
      "author_name": "yasso1",
      "author_url": "",
      "post_date": "02/02/2023 13:07:41",
      "content": "<p>What is sknn? What is his full name?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2124826": "Thanks OTTO & Kaggle to host such interesting competition, my solution is straight forward, just candidate generation + reranker, a short summary as below.\n\n###### Candidate Generation\n\n- **re-visit** - all visited items\n- **co-visit1** - based on the public notebook with 2 improvement. Thanks @radek1 and @cdeotte    \n    1. candidate generation by weight decay, e.g. a session has N items(idx1 by the reverse ts) & each item has M pairs of co-visit item(idx2 by time weight or type weight), score is calculated as 1 / idx1 / idx2 for each item & candidate pair, then sum the score as candidate score\n    2. co-visit features for rerank model - counts & probs pairwise item co-occurrences, index distance & time distance(consider order or not), etc.\n- **co-visit2(buy2buy)** - only carts or orders\n- **co-visit3(next item)** - only consider last item when generate candidate\n- **sknn1** - simple sknn with 100 most similair & recent session \n- **sknn2** - a variant of sknn named as Sequence and Time Aware Neighborhood, reference https://arxiv.org/abs/1910.12781 for both sknn implementation\n\n###### Re-Rank Model\n\n- **aid features** - simple stat by time window(7, 14, 21)\n- **session features** - length, aid count, count by type, time to the end of prediction period\n- **revisit aid X session**, clicks/carts/orders count, absolute/relative position\n- **co-visit features** - as mentioned before, generate a lot feature to describe pairwise item co-occurrences\n- **sknn features** - only candidate score/rank\n- **similarity features** - similarity between candidate item and session by item embedding of word2vec & implicit' s BPR\n\nFinally, 3 models(clicks/carts/orders) is trained by LightGBM(lambdarank)",
    "2124886": "Congrats for solo gold!",
    "2125043": "Congratulation!👍 I have a question about similarity features. Do you present the session embedding by using the average  of the historical aid sequence? or else.",
    "2125051": "Thanks. I calculate the cosine distance between the candidate and last item, and max, min, mean to each item in the session.",
    "2125061": "Thanks, lucky solo, let's find a chance to team up.",
    "2125113": "You didn't take only 20/50/last day features, used all aids in session?",
    "2125124": "Hi @gongbi congrats for solo gold!\n\none question, which library do you use for sknn?",
    "2125281": "just implement it by myself, the logic is simple.",
    "2125308": "Congras and thanks for sharing. May I ask what training data and label are you using ?? RADEK's CV data here -> https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991 or you use the whole train and test data.",
    "2125323": "I built my own code to prepare the training data and label, logic should be same.",
    "2125396": "Yes, i only take last 40 aid, not selected carefully. And I have another group of feature to take carts or orders only.",
    "2126127": "Congrats on your solo gold! Amazing work!\n\nThe Devastator.",
    "2126802": "What is sknn? What is his full name?",
    "2126814": "Noted, thanks for your reply"
  },
  "source": "meta"
}