{
  "id": 382790,
  "title": "2nd Place Solution(ONODERA part)",
  "url": "/competitions/otto-recommender-system/discussion/382790",
  "author_name": "ONODERA",
  "post_date": "2023-02-01T02:53:54.110000",
  "votes": 90,
  "comment_count": 34,
  "views": 0,
  "content": "<p>First of all, thank you for launching and organizing this terrific competition <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>.</p>\n<p>I wanted to be first, but I don't really care so far.<br>\nI'd like to explain my part.</p>\n<h3>Candidates</h3>\n<p>When I teamed up with <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>, he already has great candidates compared to mine.<br>\nSo I decided to use his candidates.</p>\n<h3>Features</h3>\n<h4>Item2item Features</h4>\n<p>Also <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> already has splendid features, but there is room for improvement regarding CF features.<br>\nSo I focused on item2item features and that consists of</p>\n<ul>\n<li>count</li>\n<li>time difference</li>\n<li>sequence difference(invented by <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>)</li>\n<li>2 kind of weighted above features</li>\n<li>Aggregation of above<br>\nIn total, I got 93 features. After this, I could generate almost 5k features using different combination(e.g. click to order, cart to order, etc…)<br>\nI use just 400~500 features eventually.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F209f3518a4ec064f4bbd8e3c0d1677d0%2F2023-02-17%205.10.31.png?generation=1676579857970702&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h4>1st stage prediction Features</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F6041fac8d79269776d88f8350774950a%2F2023-02-17%205.10.47.png?generation=1676579712049232&amp;alt=media\" alt=\"\"></p>\n<h4>Pseudo Event Features</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Feac4b9290d28bf9725b59969745f0459%2F2023-02-17%205.11.11.png?generation=1676579770675159&amp;alt=media\" alt=\"\"></p>\n<h3>Models</h3>\n<p>I used XGBoost and CatBoost.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Ffc93d88efbabc29bb619f7a8d5cd8858%2F2023-02-17%205.10.08.png?generation=1676579918058966&amp;alt=media\" alt=\"\"></p>\n<h3>Pipeline</h3>\n<p>After that 2nd stage, we blended our result ( <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>, <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> ) by rank.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F7c88b62d96eb550b66d392c4a9d46413%2F2023-02-03%209.07.35.png?generation=1675382887632198&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382839\" target=\"_blank\">my teammate's solution</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Fb89de93b7508e672fc007b50232dea1f%2F2023-02-17%205.09.49.png?generation=1676580022020358&amp;alt=media\" alt=\"\"></p>\n<h3>Acknowledgments</h3>\n<p>If I hadn't used cuDF and cuML, I couldn't manage a lot of experiments.<br>\nThanks <a href=\"https://rapids.ai/index.html\" target=\"_blank\">RAPIDS</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F2fad8cb1ed63c9d91ed4822fbdf133e4%2FRAPIDS-logo-white.png?generation=1675217750447580&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2124468,
      "postDate": "2023-02-01T02:53:54.110Z",
      "content": "<p>First of all, thank you for launching and organizing this terrific competition <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>.</p>\n<p>I wanted to be first, but I don't really care so far.<br>\nI'd like to explain my part.</p>\n<h3>Candidates</h3>\n<p>When I teamed up with <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>, he already has great candidates compared to mine.<br>\nSo I decided to use his candidates.</p>\n<h3>Features</h3>\n<h4>Item2item Features</h4>\n<p>Also <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> already has splendid features, but there is room for improvement regarding CF features.<br>\nSo I focused on item2item features and that consists of</p>\n<ul>\n<li>count</li>\n<li>time difference</li>\n<li>sequence difference(invented by <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>)</li>\n<li>2 kind of weighted above features</li>\n<li>Aggregation of above<br>\nIn total, I got 93 features. After this, I could generate almost 5k features using different combination(e.g. click to order, cart to order, etc…)<br>\nI use just 400~500 features eventually.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F209f3518a4ec064f4bbd8e3c0d1677d0%2F2023-02-17%205.10.31.png?generation=1676579857970702&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h4>1st stage prediction Features</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F6041fac8d79269776d88f8350774950a%2F2023-02-17%205.10.47.png?generation=1676579712049232&amp;alt=media\" alt=\"\"></p>\n<h4>Pseudo Event Features</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Feac4b9290d28bf9725b59969745f0459%2F2023-02-17%205.11.11.png?generation=1676579770675159&amp;alt=media\" alt=\"\"></p>\n<h3>Models</h3>\n<p>I used XGBoost and CatBoost.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Ffc93d88efbabc29bb619f7a8d5cd8858%2F2023-02-17%205.10.08.png?generation=1676579918058966&amp;alt=media\" alt=\"\"></p>\n<h3>Pipeline</h3>\n<p>After that 2nd stage, we blended our result ( <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>, <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> ) by rank.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F7c88b62d96eb550b66d392c4a9d46413%2F2023-02-03%209.07.35.png?generation=1675382887632198&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382839\" target=\"_blank\">my teammate's solution</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Fb89de93b7508e672fc007b50232dea1f%2F2023-02-17%205.09.49.png?generation=1676580022020358&amp;alt=media\" alt=\"\"></p>\n<h3>Acknowledgments</h3>\n<p>If I hadn't used cuDF and cuML, I couldn't manage a lot of experiments.<br>\nThanks <a href=\"https://rapids.ai/index.html\" target=\"_blank\">RAPIDS</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F2fad8cb1ed63c9d91ed4822fbdf133e4%2FRAPIDS-logo-white.png?generation=1675217750447580&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, thank you for launching and organizing this terrific competition @pnormann.\n\nI wanted to be first, but I don't really care so far.\nI'd like to explain my part.\n\n### Candidates\nWhen I teamed up with @psilogram, he already has great candidates compared to mine.\nSo I decided to use his candidates.\n\n### Features\n#### Item2item Features\nAlso @psilogram already has splendid features, but there is room for improvement regarding CF features.\nSo I focused on item2item features and that consists of\n- count\n- time difference\n- sequence difference(invented by @psilogram)\n- 2 kind of weighted above features\n- Aggregation of above\nIn total, I got 93 features. After this, I could generate almost 5k features using different combination(e.g. click to order, cart to order, etc...)\nI use just 400~500 features eventually.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F209f3518a4ec064f4bbd8e3c0d1677d0%2F2023-02-17%205.10.31.png?generation=1676579857970702&alt=media)\n\n#### 1st stage prediction Features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F6041fac8d79269776d88f8350774950a%2F2023-02-17%205.10.47.png?generation=1676579712049232&alt=media)\n\n#### Pseudo Event Features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Feac4b9290d28bf9725b59969745f0459%2F2023-02-17%205.11.11.png?generation=1676579770675159&alt=media)\n\n\n### Models\nI used XGBoost and CatBoost.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Ffc93d88efbabc29bb619f7a8d5cd8858%2F2023-02-17%205.10.08.png?generation=1676579918058966&alt=media)\n\n### Pipeline\nAfter that 2nd stage, we blended our result ( @senkin13, @h4211819 ) by rank.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F7c88b62d96eb550b66d392c4a9d46413%2F2023-02-03%209.07.35.png?generation=1675382887632198&alt=media)\n[my teammate's solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382839)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Fb89de93b7508e672fc007b50232dea1f%2F2023-02-17%205.09.49.png?generation=1676580022020358&alt=media)\n\n### Acknowledgments\nIf I hadn't used cuDF and cuML, I couldn't manage a lot of experiments.\nThanks [RAPIDS](https://rapids.ai/index.html)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F2fad8cb1ed63c9d91ed4822fbdf133e4%2FRAPIDS-logo-white.png?generation=1675217750447580&alt=media)",
      "votes": 90
    },
    {
      "id": 2124812,
      "postDate": "2023-02-01T08:21:44.030Z",
      "content": "<p>A few thoughts on this competition. </p>\n<ul>\n<li><p>I must say it's really nice to have a competition where CV/public LB/private LB scores are all aligned. This was due primarily to the large amount of data provided, but also to the way train/test was split. Kudos to the organizers. </p></li>\n<li><p>This competition was different from some previous recommender competitions because it was based almost entirely on item-item and user-item features as opposed to user-user features. Towards the end of the competition, I spent quite a bit of time generating candidates and features based on user (session) similarity but couldn't find anything that helped the cv score. Stepping back and thinking what this means, I guess Otto shoppers are all unique individuals with unique requirements, not bland automatons. That's a comforting take-away.</p></li>\n<li><p>Like <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>, I used cudf and was again impressed with its speed and ease-of-use. I also used the RAPIDS cugraph package, which is also great.</p></li>\n<li><p>I started with LGB for my ranker models, but quickly switched to XGB due to its better gpu support. Towards the end of the competition, I tried Catboost and was pleasantly surprised by its speed and accuracy. In the end, all three models produced similar scores, but Catboost was slightly better than LGB and XGB, and XGB and CB with GPU were faster than LGB.</p></li>\n</ul>\n<p>Last but not least, I didn't have much time to devote to the competition in the final weeks, so many thanks to my teammates ( <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>, <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>, <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> ) for carrying me over the finish line.</p>",
      "rawMarkdown": "A few thoughts on this competition. \n\n- I must say it's really nice to have a competition where CV/public LB/private LB scores are all aligned. This was due primarily to the large amount of data provided, but also to the way train/test was split. Kudos to the organizers. \n\n- This competition was different from some previous recommender competitions because it was based almost entirely on item-item and user-item features as opposed to user-user features. Towards the end of the competition, I spent quite a bit of time generating candidates and features based on user (session) similarity but couldn't find anything that helped the cv score. Stepping back and thinking what this means, I guess Otto shoppers are all unique individuals with unique requirements, not bland automatons. That's a comforting take-away.\n\n- Like @onodera, I used cudf and was again impressed with its speed and ease-of-use. I also used the RAPIDS cugraph package, which is also great.\n\n- I started with LGB for my ranker models, but quickly switched to XGB due to its better gpu support. Towards the end of the competition, I tried Catboost and was pleasantly surprised by its speed and accuracy. In the end, all three models produced similar scores, but Catboost was slightly better than LGB and XGB, and XGB and CB with GPU were faster than LGB.\n\nLast but not least, I didn't have much time to devote to the competition in the final weeks, so many thanks to my teammates ( @onodera, @senkin13, @h4211819 ) for carrying me over the finish line.",
      "votes": 12,
      "replies": [
        {
          "id": 2124868,
          "postDate": "2023-02-01T09:07:44.460Z",
          "content": "<p>sknn did well on my side which mainly search the most similar &amp; recent K session(by aid overlapping)</p>",
          "rawMarkdown": "sknn did well on my side which mainly search the most similar & recent K session(by aid overlapping)",
          "replies": [
            {
              "id": 2125121,
              "postDate": "2023-02-01T12:46:08.763Z",
              "content": "<p>Interesting. How did you convert the session into vectors?</p>",
              "rawMarkdown": "Interesting. How did you convert the session into vectors?",
              "votes": 1
            },
            {
              "id": 2125311,
              "postDate": "2023-02-01T15:29:10.680Z",
              "content": "<p>I don't convert it to vector when calculate the similarity of 2 sessions as aid is high cardinality. Just simply convert all aid of session as set, then use the union operator to calculate the overlap. </p>\n<p><code>len(s1_aids &amp; s2_aids) / sqrt(len(s1_aids)) * sqrt(len(s2_aids))</code></p>",
              "rawMarkdown": "I don't convert it to vector when calculate the similarity of 2 sessions as aid is high cardinality. Just simply convert all aid of session as set, then use the union operator to calculate the overlap. \n\n`len(s1_aids & s2_aids) / sqrt(len(s1_aids)) * sqrt(len(s2_aids))`",
              "votes": 1
            },
            {
              "id": 2125349,
              "postDate": "2023-02-01T15:49:55.960Z",
              "content": "<p>I see, so something like jaccard similarity. Does this similarity metric have a name? </p>",
              "rawMarkdown": "I see, so something like jaccard similarity. Does this similarity metric have a name? "
            },
            {
              "id": 2125360,
              "postDate": "2023-02-01T16:00:52.833Z",
              "content": "<p>typo in previous message, it should be intersection. It's a cosine distance likely.</p>",
              "rawMarkdown": "typo in previous message, it should be intersection. It's a cosine distance likely."
            }
          ]
        },
        {
          "id": 2125116,
          "postDate": "2023-02-01T12:41:51.547Z",
          "content": "<p>Congratulations! 🎉🎉🎉🎉🎉🎉🎉🎉<br>\nConsider publishing the code for the solution? <br>\nI'd like to study your team's solution carefully.</p>",
          "rawMarkdown": "Congratulations! 🎉🎉🎉🎉🎉🎉🎉🎉\nConsider publishing the code for the solution? \nI'd like to study your team's solution carefully."
        },
        {
          "id": 2125306,
          "postDate": "2023-02-01T15:26:44.120Z",
          "content": "<p>Congrats! Could you please elaborate on the usage of cugraph? Thank you.</p>",
          "rawMarkdown": "Congrats! Could you please elaborate on the usage of cugraph? Thank you.",
          "replies": [
            {
              "id": 2125357,
              "postDate": "2023-02-01T15:57:33.487Z",
              "content": "<p>I used cugraph to build a graph of items that appeared together in the same session. Each node was an item and edges represented a pair of items appearing in the same session (with session truncation to avoid too many edges from sessions with lots of items). Then I used the the cugraph jaccard function to measure the jaccard similarity between pairs of nodes (e.g., how many common nodes were in the two subnets centered around each node in the pair) and used these jaccard scores to generate candidates and as features for the ranker. The jaccard scores turned out to be among the strongest features.</p>",
              "rawMarkdown": "I used cugraph to build a graph of items that appeared together in the same session. Each node was an item and edges represented a pair of items appearing in the same session (with session truncation to avoid too many edges from sessions with lots of items). Then I used the the cugraph jaccard function to measure the jaccard similarity between pairs of nodes (e.g., how many common nodes were in the two subnets centered around each node in the pair) and used these jaccard scores to generate candidates and as features for the ranker. The jaccard scores turned out to be among the strongest features.",
              "votes": 4
            },
            {
              "id": 2125539,
              "postDate": "2023-02-01T17:39:53.740Z",
              "content": "<blockquote>\n  <p>The jaccard scores turned out to be among the strongest features.</p>\n</blockquote>\n<p>Awesome! Nice use of cugraph features!</p>",
              "rawMarkdown": ">The jaccard scores turned out to be among the strongest features.\n\nAwesome! Nice use of cugraph features!",
              "votes": 2
            }
          ]
        },
        {
          "id": 2126094,
          "postDate": "2023-02-02T05:02:22.143Z",
          "content": "<p>Congratulation for 2nd rank. </p>\n<blockquote>\n  <p>sequence difference(invented by <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>)</p>\n</blockquote>\n<p>Can you explain more about this feature?</p>",
          "rawMarkdown": "Congratulation for 2nd rank. \n\n> sequence difference(invented by @psilogram)\n\nCan you explain more about this feature?",
          "replies": [
            {
              "id": 2126119,
              "postDate": "2023-02-02T05:27:41.967Z",
              "content": "<p>Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].<br>\nWe can generate distance features from this.<br>\ne.g. distance between aid d and aid a is 3(3-0)</p>",
              "rawMarkdown": "Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].\nWe can generate distance features from this.\ne.g. distance between aid d and aid a is 3(3-0)",
              "votes": 2
            },
            {
              "id": 2126179,
              "postDate": "2023-02-02T05:53:11.903Z",
              "content": "<p><a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> I see. However, I'm starting to be confused about what CF features mean.<br>\nBut since this features are your part, let me ask in other thread.</p>",
              "rawMarkdown": "@onodera I see. However, I'm starting to be confused about what CF features mean.\nBut since this features are your part, let me ask in other thread."
            }
          ]
        }
      ]
    },
    {
      "id": 2124647,
      "postDate": "2023-02-01T05:54:34.800Z",
      "content": "<p>Hi, will you open source your code later? I want to learn the details</p>",
      "rawMarkdown": "Hi, will you open source your code later? I want to learn the details",
      "votes": 12
    },
    {
      "id": 2124529,
      "postDate": "2023-02-01T04:13:47.847Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> 2nd place curse continues lol</p>",
      "rawMarkdown": "Congrats @onodera 2nd place curse continues lol",
      "votes": 6
    },
    {
      "id": 2125908,
      "postDate": "2023-02-02T01:43:41.903Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and team. Great job jumping upward from public to private LB!</p>",
      "rawMarkdown": "Congratulations @onodera and team. Great job jumping upward from public to private LB!",
      "votes": 1
    },
    {
      "id": 2147737,
      "postDate": "2023-02-16T20:43:48.920Z",
      "content": "<p>I added some figures. I hope this helps you understand.</p>",
      "rawMarkdown": "I added some figures. I hope this helps you understand."
    },
    {
      "id": 2126187,
      "postDate": "2023-02-02T05:56:59.580Z",
      "content": "<p>I'm not familiar with CF feature. So, what this means? Is this session feature or item feature?<br>\nConsidering from your answer, this seems to describe kind of item2item features but is this correct?</p>\n<blockquote>\n  <p>Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].<br>\n  We can generate distance features from this.<br>\n  e.g. distance between aid d and aid a is 3(3-0)</p>\n</blockquote>",
      "rawMarkdown": "I'm not familiar with CF feature. So, what this means? Is this session feature or item feature?\nConsidering from your answer, this seems to describe kind of item2item features but is this correct?\n\n> Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].\n> We can generate distance features from this.\n> e.g. distance between aid d and aid a is 3(3-0)",
      "replies": [
        {
          "id": 2126203,
          "postDate": "2023-02-02T06:02:35.273Z",
          "content": "<p>Yeah, exactly</p>",
          "rawMarkdown": "Yeah, exactly",
          "replies": [
            {
              "id": 2126232,
              "postDate": "2023-02-02T06:18:03.340Z",
              "content": "<p>Well, then, how to calculate this feature from 1st stage output?</p>\n<p>According to <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> 's comment, 1st stage output are like this:</p>\n<p>[session, aid, aid_candidate, jaccard, …]</p>\n<p>how to add e.g. sequence difference between 2 items in a session to this table?</p>",
              "rawMarkdown": "Well, then, how to calculate this feature from 1st stage output?\n\nAccording to @psilogram 's comment, 1st stage output are like this:\n\n[session, aid, aid_candidate, jaccard, ...]\n\nhow to add e.g. sequence difference between 2 items in a session to this table?"
            },
            {
              "id": 2126245,
              "postDate": "2023-02-02T06:28:28.717Z",
              "content": "<p>Oh, I think I'm  start to understand.</p>\n<blockquote>\n  <p>aid, aid_candidate</p>\n</blockquote>\n<p>Since all these pairs are AIDs in the same session, you just added session difference of these items, right?</p>",
              "rawMarkdown": "Oh, I think I'm  start to understand.\n\n> aid, aid_candidate\n\nSince all these pairs are AIDs in the same session, you just added session difference of these items, right?"
            },
            {
              "id": 2126264,
              "postDate": "2023-02-02T06:38:08.483Z",
              "content": "<p>Yes. FYI, the primary keys of my item2item features are [aid_x(aid), type_x, aid_y(aid_candidate), type_y], so I could aggregate them again. That's why I had a lot of features.</p>",
              "rawMarkdown": "Yes. FYI, the primary keys of my item2item features are [aid_x(aid), type_x, aid_y(aid_candidate), type_y], so I could aggregate them again. That's why I had a lot of features.",
              "votes": 2
            },
            {
              "id": 2126296,
              "postDate": "2023-02-02T06:50:04.570Z",
              "content": "<p>I see. Thanks.</p>\n<p>Can I also ask:</p>\n<ol>\n<li>how many events in the 2nd stage input</li>\n<li>how much time to calculate generating features</li>\n<li>Your machine spec</li>\n</ol>",
              "rawMarkdown": "I see. Thanks.\n\nCan I also ask:\n\n1. how many events in the 2nd stage input\n2. how much time to calculate generating features\n3. Your machine spec"
            },
            {
              "id": 2126343,
              "postDate": "2023-02-02T07:15:52.897Z",
              "content": "<ol>\n<li>same as 1st stage. All events.</li>\n<li>Just 1 hour or so.</li>\n<li>DGX(4X V100)</li>\n</ol>",
              "rawMarkdown": "1. same as 1st stage. All events.\n2. Just 1 hour or so.\n3. DGX(4X V100)",
              "votes": 2
            }
          ]
        },
        {
          "id": 2126359,
          "postDate": "2023-02-02T07:32:27.393Z",
          "content": "<p>Just to be clear, all of these features (jaccard, sequence distance, event type, time between events, event recency, etc.) are just different ways of weighting pairs when building the co-visitation matrices. The intuition for sequence distance is that item events that occur close together in a session should be more significant than events separated by other events.</p>",
          "rawMarkdown": "Just to be clear, all of these features (jaccard, sequence distance, event type, time between events, event recency, etc.) are just different ways of weighting pairs when building the co-visitation matrices. The intuition for sequence distance is that item events that occur close together in a session should be more significant than events separated by other events.",
          "votes": 2,
          "replies": [
            {
              "id": 2126375,
              "postDate": "2023-02-02T07:42:55.757Z",
              "content": "<p>I see, so does 1st stage model generate candidate solely from this co-visitation matrix? If so, how to optimize weight from these features?</p>",
              "rawMarkdown": "I see, so does 1st stage model generate candidate solely from this co-visitation matrix? If so, how to optimize weight from these features?"
            },
            {
              "id": 2126406,
              "postDate": "2023-02-02T07:54:28.223Z",
              "content": "<p>Yes, candidates are generated almost entirely by joining the events that have already occurred in the session to the various co-visitation matrices. Weight optimization was done by trial-and-error. Using cudf, I could generate 20-30 co-visit matrices in about 2 minutes, and the candidate collection took another couple minutes, so iteration was relatively quick.</p>",
              "rawMarkdown": "Yes, candidates are generated almost entirely by joining the events that have already occurred in the session to the various co-visitation matrices. Weight optimization was done by trial-and-error. Using cudf, I could generate 20-30 co-visit matrices in about 2 minutes, and the candidate collection took another couple minutes, so iteration was relatively quick.",
              "votes": 1
            },
            {
              "id": 2126413,
              "postDate": "2023-02-02T07:59:08.687Z",
              "content": "<p><a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> I see. So you used similar approach that Chris does. Thanks.</p>",
              "rawMarkdown": "@psilogram I see. So you used similar approach that Chris does. Thanks."
            }
          ]
        }
      ]
    },
    {
      "id": 2126130,
      "postDate": "2023-02-02T05:35:28.487Z",
      "content": "<p>Congratulations King of 2nd place on Kaggle</p>",
      "rawMarkdown": "Congratulations King of 2nd place on Kaggle"
    },
    {
      "id": 2125842,
      "postDate": "2023-02-01T23:43:33.757Z",
      "content": "<p>Congratulations! In your nice visualized flow chart I see a very high recall score in 1st stage, is this corresponding to recall@20 only with candidate generation? Is so, could you help provide more tips on how you achieve such high recall without using ranker? Thanks.</p>",
      "rawMarkdown": "Congratulations! In your nice visualized flow chart I see a very high recall score in 1st stage, is this corresponding to recall@20 only with candidate generation? Is so, could you help provide more tips on how you achieve such high recall without using ranker? Thanks."
    },
    {
      "id": 2124498,
      "postDate": "2023-02-01T03:27:58.267Z",
      "content": "<p>Good example of aggregated data.</p>",
      "rawMarkdown": "Good example of aggregated data."
    },
    {
      "id": 2124855,
      "postDate": "2023-02-01T08:57:30.300Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2124659,
      "postDate": "2023-02-01T06:03:48.907Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2219888,
      "postDate": "2023-04-13T01:14:57.383Z",
      "content": "<p>Thanks for the write-up</p>",
      "rawMarkdown": "Thanks for the write-up"
    }
  ],
  "comments": [
    {
      "id": 2124812,
      "author_name": "Silogram",
      "author_url": "",
      "post_date": "2023-02-01T08:21:44.030000",
      "content": "<p>A few thoughts on this competition. </p>\n<ul>\n<li><p>I must say it's really nice to have a competition where CV/public LB/private LB scores are all aligned. This was due primarily to the large amount of data provided, but also to the way train/test was split. Kudos to the organizers. </p></li>\n<li><p>This competition was different from some previous recommender competitions because it was based almost entirely on item-item and user-item features as opposed to user-user features. Towards the end of the competition, I spent quite a bit of time generating candidates and features based on user (session) similarity but couldn't find anything that helped the cv score. Stepping back and thinking what this means, I guess Otto shoppers are all unique individuals with unique requirements, not bland automatons. That's a comforting take-away.</p></li>\n<li><p>Like <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>, I used cudf and was again impressed with its speed and ease-of-use. I also used the RAPIDS cugraph package, which is also great.</p></li>\n<li><p>I started with LGB for my ranker models, but quickly switched to XGB due to its better gpu support. Towards the end of the competition, I tried Catboost and was pleasantly surprised by its speed and accuracy. In the end, all three models produced similar scores, but Catboost was slightly better than LGB and XGB, and XGB and CB with GPU were faster than LGB.</p></li>\n</ul>\n<p>Last but not least, I didn't have much time to devote to the competition in the final weeks, so many thanks to my teammates ( <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>, <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>, <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> ) for carrying me over the finish line.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2124868,
          "author_name": "MichaelG",
          "author_url": "",
          "post_date": "2023-02-01T09:07:44.460000",
          "content": "<p>sknn did well on my side which mainly search the most similar &amp; recent K session(by aid overlapping)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2125121,
              "author_name": "Silogram",
              "author_url": "",
              "post_date": "2023-02-01T12:46:08.763000",
              "content": "<p>Interesting. How did you convert the session into vectors?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2125311,
              "author_name": "MichaelG",
              "author_url": "",
              "post_date": "2023-02-01T15:29:10.680000",
              "content": "<p>I don't convert it to vector when calculate the similarity of 2 sessions as aid is high cardinality. Just simply convert all aid of session as set, then use the union operator to calculate the overlap. </p>\n<p><code>len(s1_aids &amp; s2_aids) / sqrt(len(s1_aids)) * sqrt(len(s2_aids))</code></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2125349,
              "author_name": "Silogram",
              "author_url": "",
              "post_date": "2023-02-01T15:49:55.960000",
              "content": "<p>I see, so something like jaccard similarity. Does this similarity metric have a name? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2125360,
              "author_name": "MichaelG",
              "author_url": "",
              "post_date": "2023-02-01T16:00:52.833000",
              "content": "<p>typo in previous message, it should be intersection. It's a cosine distance likely.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2125116,
          "author_name": "kgxiao",
          "author_url": "",
          "post_date": "2023-02-01T12:41:51.547000",
          "content": "<p>Congratulations! 🎉🎉🎉🎉🎉🎉🎉🎉<br>\nConsider publishing the code for the solution? <br>\nI'd like to study your team's solution carefully.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2125306,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2023-02-01T15:26:44.120000",
          "content": "<p>Congrats! Could you please elaborate on the usage of cugraph? Thank you.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2125357,
              "author_name": "Silogram",
              "author_url": "",
              "post_date": "2023-02-01T15:57:33.487000",
              "content": "<p>I used cugraph to build a graph of items that appeared together in the same session. Each node was an item and edges represented a pair of items appearing in the same session (with session truncation to avoid too many edges from sessions with lots of items). Then I used the the cugraph jaccard function to measure the jaccard similarity between pairs of nodes (e.g., how many common nodes were in the two subnets centered around each node in the pair) and used these jaccard scores to generate candidates and as features for the ranker. The jaccard scores turned out to be among the strongest features.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2125539,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-02-01T17:39:53.740000",
              "content": "<blockquote>\n  <p>The jaccard scores turned out to be among the strongest features.</p>\n</blockquote>\n<p>Awesome! Nice use of cugraph features!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2126094,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-02-02T05:02:22.143000",
          "content": "<p>Congratulation for 2nd rank. </p>\n<blockquote>\n  <p>sequence difference(invented by <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>)</p>\n</blockquote>\n<p>Can you explain more about this feature?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2126119,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2023-02-02T05:27:41.967000",
              "content": "<p>Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].<br>\nWe can generate distance features from this.<br>\ne.g. distance between aid d and aid a is 3(3-0)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2126179,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-02-02T05:53:11.903000",
              "content": "<p><a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> I see. However, I'm starting to be confused about what CF features mean.<br>\nBut since this features are your part, let me ask in other thread.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2124647,
      "author_name": "Apa",
      "author_url": "",
      "post_date": "2023-02-01T05:54:34.800000",
      "content": "<p>Hi, will you open source your code later? I want to learn the details</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 2124529,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-02-01T04:13:47.847000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> 2nd place curse continues lol</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2125908,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-02-02T01:43:41.903000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and team. Great job jumping upward from public to private LB!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2147737,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2023-02-16T20:43:48.920000",
      "content": "<p>I added some figures. I hope this helps you understand.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2126187,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2023-02-02T05:56:59.580000",
      "content": "<p>I'm not familiar with CF feature. So, what this means? Is this session feature or item feature?<br>\nConsidering from your answer, this seems to describe kind of item2item features but is this correct?</p>\n<blockquote>\n  <p>Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].<br>\n  We can generate distance features from this.<br>\n  e.g. distance between aid d and aid a is 3(3-0)</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 2126203,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2023-02-02T06:02:35.273000",
          "content": "<p>Yeah, exactly</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2126232,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-02-02T06:18:03.340000",
              "content": "<p>Well, then, how to calculate this feature from 1st stage output?</p>\n<p>According to <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> 's comment, 1st stage output are like this:</p>\n<p>[session, aid, aid_candidate, jaccard, …]</p>\n<p>how to add e.g. sequence difference between 2 items in a session to this table?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2126245,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-02-02T06:28:28.717000",
              "content": "<p>Oh, I think I'm  start to understand.</p>\n<blockquote>\n  <p>aid, aid_candidate</p>\n</blockquote>\n<p>Since all these pairs are AIDs in the same session, you just added session difference of these items, right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2126264,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2023-02-02T06:38:08.483000",
              "content": "<p>Yes. FYI, the primary keys of my item2item features are [aid_x(aid), type_x, aid_y(aid_candidate), type_y], so I could aggregate them again. That's why I had a lot of features.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2126296,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-02-02T06:50:04.570000",
              "content": "<p>I see. Thanks.</p>\n<p>Can I also ask:</p>\n<ol>\n<li>how many events in the 2nd stage input</li>\n<li>how much time to calculate generating features</li>\n<li>Your machine spec</li>\n</ol>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2126343,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2023-02-02T07:15:52.897000",
              "content": "<ol>\n<li>same as 1st stage. All events.</li>\n<li>Just 1 hour or so.</li>\n<li>DGX(4X V100)</li>\n</ol>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2126359,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2023-02-02T07:32:27.393000",
          "content": "<p>Just to be clear, all of these features (jaccard, sequence distance, event type, time between events, event recency, etc.) are just different ways of weighting pairs when building the co-visitation matrices. The intuition for sequence distance is that item events that occur close together in a session should be more significant than events separated by other events.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2126375,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-02-02T07:42:55.757000",
              "content": "<p>I see, so does 1st stage model generate candidate solely from this co-visitation matrix? If so, how to optimize weight from these features?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2126406,
              "author_name": "Silogram",
              "author_url": "",
              "post_date": "2023-02-02T07:54:28.223000",
              "content": "<p>Yes, candidates are generated almost entirely by joining the events that have already occurred in the session to the various co-visitation matrices. Weight optimization was done by trial-and-error. Using cudf, I could generate 20-30 co-visit matrices in about 2 minutes, and the candidate collection took another couple minutes, so iteration was relatively quick.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2126413,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-02-02T07:59:08.687000",
              "content": "<p><a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> I see. So you used similar approach that Chris does. Thanks.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2126130,
      "author_name": "Raheem Nasirudeen",
      "author_url": "",
      "post_date": "2023-02-02T05:35:28.487000",
      "content": "<p>Congratulations King of 2nd place on Kaggle</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2125842,
      "author_name": "Sherry_pomi",
      "author_url": "",
      "post_date": "2023-02-01T23:43:33.757000",
      "content": "<p>Congratulations! In your nice visualized flow chart I see a very high recall score in 1st stage, is this corresponding to recall@20 only with candidate generation? Is so, could you help provide more tips on how you achieve such high recall without using ranker? Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2124498,
      "author_name": "Oleg Zakh",
      "author_url": "",
      "post_date": "2023-02-01T03:27:58.267000",
      "content": "<p>Good example of aggregated data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2124855,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-01T08:57:30.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2124659,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-01T06:03:48.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2219888,
      "author_name": "Tigran Hakobyan",
      "author_url": "",
      "post_date": "2023-04-13T01:14:57.383000",
      "content": "<p>Thanks for the write-up</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2124468": "First of all, thank you for launching and organizing this terrific competition @pnormann.\n\nI wanted to be first, but I don't really care so far.\nI'd like to explain my part.\n\n### Candidates\nWhen I teamed up with @psilogram, he already has great candidates compared to mine.\nSo I decided to use his candidates.\n\n### Features\n#### Item2item Features\nAlso @psilogram already has splendid features, but there is room for improvement regarding CF features.\nSo I focused on item2item features and that consists of\n- count\n- time difference\n- sequence difference(invented by @psilogram)\n- 2 kind of weighted above features\n- Aggregation of above\nIn total, I got 93 features. After this, I could generate almost 5k features using different combination(e.g. click to order, cart to order, etc...)\nI use just 400~500 features eventually.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F209f3518a4ec064f4bbd8e3c0d1677d0%2F2023-02-17%205.10.31.png?generation=1676579857970702&alt=media)\n\n#### 1st stage prediction Features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F6041fac8d79269776d88f8350774950a%2F2023-02-17%205.10.47.png?generation=1676579712049232&alt=media)\n\n#### Pseudo Event Features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Feac4b9290d28bf9725b59969745f0459%2F2023-02-17%205.11.11.png?generation=1676579770675159&alt=media)\n\n\n### Models\nI used XGBoost and CatBoost.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Ffc93d88efbabc29bb619f7a8d5cd8858%2F2023-02-17%205.10.08.png?generation=1676579918058966&alt=media)\n\n### Pipeline\nAfter that 2nd stage, we blended our result ( @senkin13, @h4211819 ) by rank.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F7c88b62d96eb550b66d392c4a9d46413%2F2023-02-03%209.07.35.png?generation=1675382887632198&alt=media)\n[my teammate's solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382839)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2Fb89de93b7508e672fc007b50232dea1f%2F2023-02-17%205.09.49.png?generation=1676580022020358&alt=media)\n\n### Acknowledgments\nIf I hadn't used cuDF and cuML, I couldn't manage a lot of experiments.\nThanks [RAPIDS](https://rapids.ai/index.html)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F2fad8cb1ed63c9d91ed4822fbdf133e4%2FRAPIDS-logo-white.png?generation=1675217750447580&alt=media)",
    "2124812": "A few thoughts on this competition. \n\n- I must say it's really nice to have a competition where CV/public LB/private LB scores are all aligned. This was due primarily to the large amount of data provided, but also to the way train/test was split. Kudos to the organizers. \n\n- This competition was different from some previous recommender competitions because it was based almost entirely on item-item and user-item features as opposed to user-user features. Towards the end of the competition, I spent quite a bit of time generating candidates and features based on user (session) similarity but couldn't find anything that helped the cv score. Stepping back and thinking what this means, I guess Otto shoppers are all unique individuals with unique requirements, not bland automatons. That's a comforting take-away.\n\n- Like @onodera, I used cudf and was again impressed with its speed and ease-of-use. I also used the RAPIDS cugraph package, which is also great.\n\n- I started with LGB for my ranker models, but quickly switched to XGB due to its better gpu support. Towards the end of the competition, I tried Catboost and was pleasantly surprised by its speed and accuracy. In the end, all three models produced similar scores, but Catboost was slightly better than LGB and XGB, and XGB and CB with GPU were faster than LGB.\n\nLast but not least, I didn't have much time to devote to the competition in the final weeks, so many thanks to my teammates ( @onodera, @senkin13, @h4211819 ) for carrying me over the finish line.",
    "2124647": "Hi, will you open source your code later? I want to learn the details",
    "2124529": "Congrats @onodera 2nd place curse continues lol",
    "2125908": "Congratulations @onodera and team. Great job jumping upward from public to private LB!",
    "2147737": "I added some figures. I hope this helps you understand.",
    "2126187": "I'm not familiar with CF feature. So, what this means? Is this session feature or item feature?\nConsidering from your answer, this seems to describe kind of item2item features but is this correct?\n\n> Sure. Let's say a session has clicked aid [a, b, c, d] and sequence is [0, 1, 2, 3].\n> We can generate distance features from this.\n> e.g. distance between aid d and aid a is 3(3-0)",
    "2126130": "Congratulations King of 2nd place on Kaggle",
    "2125842": "Congratulations! In your nice visualized flow chart I see a very high recall score in 1st stage, is this corresponding to recall@20 only with candidate generation? Is so, could you help provide more tips on how you achieve such high recall without using ranker? Thanks.",
    "2124498": "Good example of aggregated data.",
    "2124855": "",
    "2124659": "",
    "2219888": "Thanks for the write-up"
  }
}