{
  "id": 383130,
  "title": "9th Place Solution",
  "url": "/competitions/otto-recommender-system/writeups/rotto-9th-place-solution",
  "author_name": "",
  "post_date": "2023-02-09T16:10:09.083Z",
  "votes": 43,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'd like to thank the Kaggle staff and the OTTO team for organizing this interesting competition!</p>\n<p>I'm relieved that I was able to make it through to the end, even though my score did not improve much in the final stages of the competition.</p>\n<p>Here is my solution</p>\n<h2>Candidate generation</h2>\n<ul>\n<li><strong>re-visit</strong> - all items from the session's history</li>\n<li><strong>co-visitation matrix</strong> <ul>\n<li>any2any, click2click, click2cart, click2order, (click or cart)2order, (cart or order)2order, e.t.c.</li>\n<li>I tried various patterns to create the co-visitation matrix. <br>\ne.g. actions that do not take time into account, actions immediately after an action, within 5 or 10 actions, within 5 or 10 minutes, e.t.c.</li></ul></li>\n<li><strong>create click2click(only consider the next item) graph and apply ProNE</strong> <ul>\n<li>The idea came from the hypothesis that items clicked on immediately after an item is clicked are similar to each other.</li>\n<li>I created a two-column DataFrame (an item and the item clicked immediately after) and ran it through ProNE. The number of dimensions of ProNE was 1000, and the more dimensions I increased, the more accurate the candidate recall became.</li>\n<li>Retrieve top-k aids by ProNE embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')</li></ul></li>\n<li><strong>word2vec</strong><ul>\n<li>Trained word2vec model with aid sequences (Used gensim, size=50 or 100)</li>\n<li>Retrieve top-k aids by w2v embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')</li></ul></li>\n</ul>\n<h2>Re-Ranking</h2>\n<h3>Model</h3>\n<ul>\n<li>LGBMRanker (lambdarank) </li>\n<li>I created one model each to predict clicks, carts, and orders.</li>\n</ul>\n<h3>CV</h3>\n<ul>\n<li>I created validation sets with the host's old version scripts</li>\n<li>I created 100 candidates per session when training</li>\n<li>candidate recall<ul>\n<li>click: 0.6622</li>\n<li>cart: 0.5113</li>\n<li>order: 0.7059</li></ul></li>\n<li>CV <ul>\n<li>click: 0.5601</li>\n<li>cart: 0.4414</li>\n<li>order: 0.6664</li></ul></li>\n<li>I created 300 candidates per session when inferencing</li>\n<li>Public LB: 0.603, Private LB: 0.603</li>\n</ul>\n<h3>Feature</h3>\n<ul>\n<li><strong>session features</strong><ul>\n<li>type count by session (type='clicks' or 'carts' or 'orders') </li>\n<li>number of unique types by session (type='clicks' or 'carts' or 'orders') </li>\n<li>type mean by session ('clicks'=1, 'carts'=2, 'orders'=3 and mean by session) </li></ul></li>\n<li><strong>aid features</strong> <ul>\n<li>type count within all sessions (type='clicks' or 'carts' or 'orders') </li>\n<li>type count within test sessions (type='clicks' or 'carts' or 'orders') </li>\n<li>click to cart rate within same sessions, cart to order rate within same sessions, click to order rate within same sessions</li></ul></li>\n<li><strong>session x aid features</strong><ul>\n<li>co-visitation count, rate, time-weighted count, rate, e.t.c</li>\n<li><strong>similarity</strong><ul>\n<li>cosine similarity between candidate item and last item of the session (w2v, ProNE)</li>\n<li>cosine similarity between candidate item and the second item from the back of the session (w2v, ProNE)</li>\n<li>cosine similarity between candidate item and all items of the session (w2v, ProNE)</li>\n<li>click2click Jaccard index score</li></ul></li></ul></li>\n</ul>\n<h2>What did not work well</h2>\n<ul>\n<li>Create candidates with node2vec</li>\n<li>GRU</li>\n<li>RecVAE</li>\n<li>SAR</li>\n<li>BPR</li>\n<li>Pseudo Labeling </li>\n</ul>",
  "messages": [
    {
      "id": "2126572",
      "postDate": "02/02/2023 10:07:47",
      "content": "<p>I'd like to thank the Kaggle staff and the OTTO team for organizing this interesting competition!</p>\n<p>I'm relieved that I was able to make it through to the end, even though my score did not improve much in the final stages of the competition.</p>\n<p>Here is my solution</p>\n<h2>Candidate generation</h2>\n<ul>\n<li><strong>re-visit</strong> - all items from the session's history</li>\n<li><strong>co-visitation matrix</strong> <ul>\n<li>any2any, click2click, click2cart, click2order, (click or cart)2order, (cart or order)2order, e.t.c.</li>\n<li>I tried various patterns to create the co-visitation matrix. <br>\ne.g. actions that do not take time into account, actions immediately after an action, within 5 or 10 actions, within 5 or 10 minutes, e.t.c.</li></ul></li>\n<li><strong>create click2click(only consider the next item) graph and apply ProNE</strong> <ul>\n<li>The idea came from the hypothesis that items clicked on immediately after an item is clicked are similar to each other.</li>\n<li>I created a two-column DataFrame (an item and the item clicked immediately after) and ran it through ProNE. The number of dimensions of ProNE was 1000, and the more dimensions I increased, the more accurate the candidate recall became.</li>\n<li>Retrieve top-k aids by ProNE embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')</li></ul></li>\n<li><strong>word2vec</strong><ul>\n<li>Trained word2vec model with aid sequences (Used gensim, size=50 or 100)</li>\n<li>Retrieve top-k aids by w2v embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')</li></ul></li>\n</ul>\n<h2>Re-Ranking</h2>\n<h3>Model</h3>\n<ul>\n<li>LGBMRanker (lambdarank) </li>\n<li>I created one model each to predict clicks, carts, and orders.</li>\n</ul>\n<h3>CV</h3>\n<ul>\n<li>I created validation sets with the host's old version scripts</li>\n<li>I created 100 candidates per session when training</li>\n<li>candidate recall<ul>\n<li>click: 0.6622</li>\n<li>cart: 0.5113</li>\n<li>order: 0.7059</li></ul></li>\n<li>CV <ul>\n<li>click: 0.5601</li>\n<li>cart: 0.4414</li>\n<li>order: 0.6664</li></ul></li>\n<li>I created 300 candidates per session when inferencing</li>\n<li>Public LB: 0.603, Private LB: 0.603</li>\n</ul>\n<h3>Feature</h3>\n<ul>\n<li><strong>session features</strong><ul>\n<li>type count by session (type='clicks' or 'carts' or 'orders') </li>\n<li>number of unique types by session (type='clicks' or 'carts' or 'orders') </li>\n<li>type mean by session ('clicks'=1, 'carts'=2, 'orders'=3 and mean by session) </li></ul></li>\n<li><strong>aid features</strong> <ul>\n<li>type count within all sessions (type='clicks' or 'carts' or 'orders') </li>\n<li>type count within test sessions (type='clicks' or 'carts' or 'orders') </li>\n<li>click to cart rate within same sessions, cart to order rate within same sessions, click to order rate within same sessions</li></ul></li>\n<li><strong>session x aid features</strong><ul>\n<li>co-visitation count, rate, time-weighted count, rate, e.t.c</li>\n<li><strong>similarity</strong><ul>\n<li>cosine similarity between candidate item and last item of the session (w2v, ProNE)</li>\n<li>cosine similarity between candidate item and the second item from the back of the session (w2v, ProNE)</li>\n<li>cosine similarity between candidate item and all items of the session (w2v, ProNE)</li>\n<li>click2click Jaccard index score</li></ul></li></ul></li>\n</ul>\n<h2>What did not work well</h2>\n<ul>\n<li>Create candidates with node2vec</li>\n<li>GRU</li>\n<li>RecVAE</li>\n<li>SAR</li>\n<li>BPR</li>\n<li>Pseudo Labeling </li>\n</ul>",
      "rawMarkdown": "I'd like to thank the Kaggle staff and the OTTO team for organizing this interesting competition!\n\nI'm relieved that I was able to make it through to the end, even though my score did not improve much in the final stages of the competition.\n\nHere is my solution\n\n## Candidate generation\n- **re-visit** - all items from the session's history\n- **co-visitation matrix** \n   - any2any, click2click, click2cart, click2order, (click or cart)2order, (cart or order)2order, e.t.c.\n   - I tried various patterns to create the co-visitation matrix. \ne.g. actions that do not take time into account, actions immediately after an action, within 5 or 10 actions, within 5 or 10 minutes, e.t.c.\n- **create click2click(only consider the next item) graph and apply ProNE** \n   - The idea came from the hypothesis that items clicked on immediately after an item is clicked are similar to each other.\n   - I created a two-column DataFrame (an item and the item clicked immediately after) and ran it through ProNE. The number of dimensions of ProNE was 1000, and the more dimensions I increased, the more accurate the candidate recall became.\n   - Retrieve top-k aids by ProNE embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')\n- **word2vec**\n   - Trained word2vec model with aid sequences (Used gensim, size=50 or 100)\n   - Retrieve top-k aids by w2v embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')\n\n## Re-Ranking\n### Model\n- LGBMRanker (lambdarank) \n- I created one model each to predict clicks, carts, and orders.\n\n### CV\n- I created validation sets with the host's old version scripts\n- I created 100 candidates per session when training\n- candidate recall\n   - click: 0.6622\n   - cart: 0.5113\n   - order: 0.7059\n- CV \n   - click: 0.5601\n   - cart: 0.4414\n   - order: 0.6664\n- I created 300 candidates per session when inferencing\n- Public LB: 0.603, Private LB: 0.603\n\n### Feature\n- **session features**\n   - type count by session (type='clicks' or 'carts' or 'orders') \n   - number of unique types by session (type='clicks' or 'carts' or 'orders') \n   - type mean by session ('clicks'=1, 'carts'=2, 'orders'=3 and mean by session) \n- **aid features** \n   - type count within all sessions (type='clicks' or 'carts' or 'orders') \n   - type count within test sessions (type='clicks' or 'carts' or 'orders') \n   - click to cart rate within same sessions, cart to order rate within same sessions, click to order rate within same sessions\n- **session x aid features**\n   - co-visitation count, rate, time-weighted count, rate, e.t.c\n   - **similarity**\n      - cosine similarity between candidate item and last item of the session (w2v, ProNE)\n      - cosine similarity between candidate item and the second item from the back of the session (w2v, ProNE)\n      - cosine similarity between candidate item and all items of the session (w2v, ProNE)\n      - click2click Jaccard index score\n\n## What did not work well\n- Create candidates with node2vec\n- GRU\n- RecVAE\n- SAR\n- BPR\n- Pseudo Labeling",
      "votes": null
    },
    {
      "id": "2126801",
      "postDate": "02/02/2023 13:07:24",
      "content": "<p>Amazing ! you lead the top for a long time with such a clear and fast forward approach. I also tried node2vec to compute similarity scores as a variables to my LGBM.</p>\n<p>Can you further explain how did you choose 100 candidates from all your co-visitation matrices generation ?</p>",
      "rawMarkdown": "Amazing ! you lead the top for a long time with such a clear and fast forward approach. I also tried node2vec to compute similarity scores as a variables to my LGBM.\n\nCan you further explain how did you choose 100 candidates from all your co-visitation matrices generation ?",
      "votes": null
    },
    {
      "id": "2127539",
      "postDate": "02/03/2023 00:42:17",
      "content": "<p>Thanks Rayan-aay.<br>\nI repeated trial and error by changing the ratio of co-visitation matrix, ProNE, and w2v to maximize candidate recall. Finally, I created candidates by ensembling the following three patterns with the following code from H&amp;M competition.</p>\n<ol>\n<li>top-20 ProNE aids, top-30 w2v aids, click2click within 5 actions (top 75 aids), any2any (top 150 aids)</li>\n<li>top-20 ProNE aids, top-30 w2v aids, any2(cart or order) within 10 minutes (top 50 aids), click2click within 5 actions (top 85 aids), any2any (top 150 aids)</li>\n<li>top-20 ProNE aids, top-30 w2v aids, (click or cart)2(cart or order) within 10 actions (top 50 aids), click2click within 1 minute (top 85 aids), any2any (top 150 aids)</li>\n</ol>\n<p>I created a model for each of clicks, carts, and orders, but did not change the candidate for each of them. I used W = [1,1,1].</p>\n<pre><code>def cust_blend(dt, W = [1,1,1]):\n    #Global ensemble weights\n    #W = [1.15,0.95,0.85]\n\n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['prediction0'].split())\n    REC.append(dt['prediction1'].split())\n    REC.append(dt['prediction2'].split())\n\n    #Create a dictionary of items recommended. \n    #Assign a weight according the order of appearance and multiply by global weights\n    res = {}\n    for M in range(len(REC)):\n        for n, v in enumerate(REC[M]):\n            if v in res:\n                res[v] += (W[M]/(n+1))\n            else:\n                res[v] = (W[M]/(n+1))\n\n    # Sort dictionary by item weights\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n\n    # Return the top 100 items only\n    return ' '.join(res[:100])\n</code></pre>\n<p>reference: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/titericz/h-m-ensembling-how-to</a></p>",
      "rawMarkdown": "Thanks Rayan-aay.\nI repeated trial and error by changing the ratio of co-visitation matrix, ProNE, and w2v to maximize candidate recall. Finally, I created candidates by ensembling the following three patterns with the following code from H&M competition.\n1. top-20 ProNE aids, top-30 w2v aids, click2click within 5 actions (top 75 aids), any2any (top 150 aids)\n2. top-20 ProNE aids, top-30 w2v aids, any2(cart or order) within 10 minutes (top 50 aids), click2click within 5 actions (top 85 aids), any2any (top 150 aids)\n3. top-20 ProNE aids, top-30 w2v aids, (click or cart)2(cart or order) within 10 actions (top 50 aids), click2click within 1 minute (top 85 aids), any2any (top 150 aids)\n\nI created a model for each of clicks, carts, and orders, but did not change the candidate for each of them. I used W = [1,1,1].\n```\ndef cust_blend(dt, W = [1,1,1]):\n    #Global ensemble weights\n    #W = [1.15,0.95,0.85]\n    \n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['prediction0'].split())\n    REC.append(dt['prediction1'].split())\n    REC.append(dt['prediction2'].split())\n    \n    #Create a dictionary of items recommended. \n    #Assign a weight according the order of appearance and multiply by global weights\n    res = {}\n    for M in range(len(REC)):\n        for n, v in enumerate(REC[M]):\n            if v in res:\n                res[v] += (W[M]/(n+1))\n            else:\n                res[v] = (W[M]/(n+1))\n    \n    # Sort dictionary by item weights\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n    \n    # Return the top 100 items only\n    return ' '.join(res[:100])\n```\n\nreference: [https://www.kaggle.com/code/titericz/h-m-ensembling-how-to](url)",
      "votes": null
    },
    {
      "id": "2127619",
      "postDate": "02/03/2023 01:56:19",
      "content": "<p>Solo gold medal! Congras. In our experiments, SAR (original SAS without type(click,cart,and order), we add type like bert through type embedding) and BPR worked and were both important features. But we used these methods only to generate features and not to recall. Maybe our scores boost a little bit if use them to generate candidates. Thank you for your sharing!</p>",
      "rawMarkdown": "Solo gold medal! Congras. In our experiments, SAR (original SAS without type(click,cart,and order), we add type like bert through type embedding) and BPR worked and were both important features. But we used these methods only to generate features and not to recall. Maybe our scores boost a little bit if use them to generate candidates. Thank you for your sharing!",
      "votes": null
    },
    {
      "id": "2128134",
      "postDate": "02/03/2023 13:44:35",
      "content": "<p>Congratulation on solo gold medal!</p>\n<p>Could you elaborate more on ProNE? Which library did you use to train ProNE?</p>",
      "rawMarkdown": "Congratulation on solo gold medal!\n\nCould you elaborate more on ProNE? Which library did you use to train ProNE?",
      "votes": null
    },
    {
      "id": "2130258",
      "postDate": "02/05/2023 10:06:02",
      "content": "<p>Thanks A.Sato.<br>\nI used the following library<br>\n<a href=\"url\" target=\"_blank\">https://github.com/THUDM/ProNE</a></p>",
      "rawMarkdown": "Thanks A.Sato.\nI used the following library\n[https://github.com/THUDM/ProNE](url)",
      "votes": null
    },
    {
      "id": "3115552",
      "postDate": "02/05/2025 03:59:52",
      "content": "<p>Waoo this is what we called maximum impact. You did well in the competition , please how did you creat a model for each of clicks, carts, and orders</p>",
      "rawMarkdown": "Waoo this is what we called maximum impact. You did well in the competition , please how did you creat a model for each of clicks, carts, and orders",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2126801,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "02/02/2023 13:07:24",
      "content": "<p>Amazing ! you lead the top for a long time with such a clear and fast forward approach. I also tried node2vec to compute similarity scores as a variables to my LGBM.</p>\n<p>Can you further explain how did you choose 100 candidates from all your co-visitation matrices generation ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2127539,
          "author_name": "pirotototo4023",
          "author_url": "",
          "post_date": "02/03/2023 00:42:17",
          "content": "<p>Thanks Rayan-aay.<br>\nI repeated trial and error by changing the ratio of co-visitation matrix, ProNE, and w2v to maximize candidate recall. Finally, I created candidates by ensembling the following three patterns with the following code from H&amp;M competition.</p>\n<ol>\n<li>top-20 ProNE aids, top-30 w2v aids, click2click within 5 actions (top 75 aids), any2any (top 150 aids)</li>\n<li>top-20 ProNE aids, top-30 w2v aids, any2(cart or order) within 10 minutes (top 50 aids), click2click within 5 actions (top 85 aids), any2any (top 150 aids)</li>\n<li>top-20 ProNE aids, top-30 w2v aids, (click or cart)2(cart or order) within 10 actions (top 50 aids), click2click within 1 minute (top 85 aids), any2any (top 150 aids)</li>\n</ol>\n<p>I created a model for each of clicks, carts, and orders, but did not change the candidate for each of them. I used W = [1,1,1].</p>\n<pre><code>def cust_blend(dt, W = [1,1,1]):\n    #Global ensemble weights\n    #W = [1.15,0.95,0.85]\n\n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['prediction0'].split())\n    REC.append(dt['prediction1'].split())\n    REC.append(dt['prediction2'].split())\n\n    #Create a dictionary of items recommended. \n    #Assign a weight according the order of appearance and multiply by global weights\n    res = {}\n    for M in range(len(REC)):\n        for n, v in enumerate(REC[M]):\n            if v in res:\n                res[v] += (W[M]/(n+1))\n            else:\n                res[v] = (W[M]/(n+1))\n\n    # Sort dictionary by item weights\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n\n    # Return the top 100 items only\n    return ' '.join(res[:100])\n</code></pre>\n<p>reference: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/titericz/h-m-ensembling-how-to</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3115552,
              "author_name": "graceegbe12",
              "author_url": "",
              "post_date": "02/05/2025 03:59:52",
              "content": "<p>Waoo this is what we called maximum impact. You did well in the competition , please how did you creat a model for each of clicks, carts, and orders</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2127619,
      "author_name": "toyou2u",
      "author_url": "",
      "post_date": "02/03/2023 01:56:19",
      "content": "<p>Solo gold medal! Congras. In our experiments, SAR (original SAS without type(click,cart,and order), we add type like bert through type embedding) and BPR worked and were both important features. But we used these methods only to generate features and not to recall. Maybe our scores boost a little bit if use them to generate candidates. Thank you for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2128134,
      "author_name": "tommy1028",
      "author_url": "",
      "post_date": "02/03/2023 13:44:35",
      "content": "<p>Congratulation on solo gold medal!</p>\n<p>Could you elaborate more on ProNE? Which library did you use to train ProNE?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2130258,
          "author_name": "pirotototo4023",
          "author_url": "",
          "post_date": "02/05/2023 10:06:02",
          "content": "<p>Thanks A.Sato.<br>\nI used the following library<br>\n<a href=\"url\" target=\"_blank\">https://github.com/THUDM/ProNE</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2126572": "I'd like to thank the Kaggle staff and the OTTO team for organizing this interesting competition!\n\nI'm relieved that I was able to make it through to the end, even though my score did not improve much in the final stages of the competition.\n\nHere is my solution\n\n## Candidate generation\n- **re-visit** - all items from the session's history\n- **co-visitation matrix** \n   - any2any, click2click, click2cart, click2order, (click or cart)2order, (cart or order)2order, e.t.c.\n   - I tried various patterns to create the co-visitation matrix. \ne.g. actions that do not take time into account, actions immediately after an action, within 5 or 10 actions, within 5 or 10 minutes, e.t.c.\n- **create click2click(only consider the next item) graph and apply ProNE** \n   - The idea came from the hypothesis that items clicked on immediately after an item is clicked are similar to each other.\n   - I created a two-column DataFrame (an item and the item clicked immediately after) and ran it through ProNE. The number of dimensions of ProNE was 1000, and the more dimensions I increased, the more accurate the candidate recall became.\n   - Retrieve top-k aids by ProNE embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')\n- **word2vec**\n   - Trained word2vec model with aid sequences (Used gensim, size=50 or 100)\n   - Retrieve top-k aids by w2v embeddings (Used cuml.neighbors.NearestNeighbors and metric='cosine')\n\n## Re-Ranking\n### Model\n- LGBMRanker (lambdarank) \n- I created one model each to predict clicks, carts, and orders.\n\n### CV\n- I created validation sets with the host's old version scripts\n- I created 100 candidates per session when training\n- candidate recall\n   - click: 0.6622\n   - cart: 0.5113\n   - order: 0.7059\n- CV \n   - click: 0.5601\n   - cart: 0.4414\n   - order: 0.6664\n- I created 300 candidates per session when inferencing\n- Public LB: 0.603, Private LB: 0.603\n\n### Feature\n- **session features**\n   - type count by session (type='clicks' or 'carts' or 'orders') \n   - number of unique types by session (type='clicks' or 'carts' or 'orders') \n   - type mean by session ('clicks'=1, 'carts'=2, 'orders'=3 and mean by session) \n- **aid features** \n   - type count within all sessions (type='clicks' or 'carts' or 'orders') \n   - type count within test sessions (type='clicks' or 'carts' or 'orders') \n   - click to cart rate within same sessions, cart to order rate within same sessions, click to order rate within same sessions\n- **session x aid features**\n   - co-visitation count, rate, time-weighted count, rate, e.t.c\n   - **similarity**\n      - cosine similarity between candidate item and last item of the session (w2v, ProNE)\n      - cosine similarity between candidate item and the second item from the back of the session (w2v, ProNE)\n      - cosine similarity between candidate item and all items of the session (w2v, ProNE)\n      - click2click Jaccard index score\n\n## What did not work well\n- Create candidates with node2vec\n- GRU\n- RecVAE\n- SAR\n- BPR\n- Pseudo Labeling",
    "2126801": "Amazing ! you lead the top for a long time with such a clear and fast forward approach. I also tried node2vec to compute similarity scores as a variables to my LGBM.\n\nCan you further explain how did you choose 100 candidates from all your co-visitation matrices generation ?",
    "2127539": "Thanks Rayan-aay.\nI repeated trial and error by changing the ratio of co-visitation matrix, ProNE, and w2v to maximize candidate recall. Finally, I created candidates by ensembling the following three patterns with the following code from H&M competition.\n1. top-20 ProNE aids, top-30 w2v aids, click2click within 5 actions (top 75 aids), any2any (top 150 aids)\n2. top-20 ProNE aids, top-30 w2v aids, any2(cart or order) within 10 minutes (top 50 aids), click2click within 5 actions (top 85 aids), any2any (top 150 aids)\n3. top-20 ProNE aids, top-30 w2v aids, (click or cart)2(cart or order) within 10 actions (top 50 aids), click2click within 1 minute (top 85 aids), any2any (top 150 aids)\n\nI created a model for each of clicks, carts, and orders, but did not change the candidate for each of them. I used W = [1,1,1].\n```\ndef cust_blend(dt, W = [1,1,1]):\n    #Global ensemble weights\n    #W = [1.15,0.95,0.85]\n    \n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['prediction0'].split())\n    REC.append(dt['prediction1'].split())\n    REC.append(dt['prediction2'].split())\n    \n    #Create a dictionary of items recommended. \n    #Assign a weight according the order of appearance and multiply by global weights\n    res = {}\n    for M in range(len(REC)):\n        for n, v in enumerate(REC[M]):\n            if v in res:\n                res[v] += (W[M]/(n+1))\n            else:\n                res[v] = (W[M]/(n+1))\n    \n    # Sort dictionary by item weights\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n    \n    # Return the top 100 items only\n    return ' '.join(res[:100])\n```\n\nreference: [https://www.kaggle.com/code/titericz/h-m-ensembling-how-to](url)",
    "2127619": "Solo gold medal! Congras. In our experiments, SAR (original SAS without type(click,cart,and order), we add type like bert through type embedding) and BPR worked and were both important features. But we used these methods only to generate features and not to recall. Maybe our scores boost a little bit if use them to generate candidates. Thank you for your sharing!",
    "2128134": "Congratulation on solo gold medal!\n\nCould you elaborate more on ProNE? Which library did you use to train ProNE?",
    "2130258": "Thanks A.Sato.\nI used the following library\n[https://github.com/THUDM/ProNE](url)",
    "3115552": "Waoo this is what we called maximum impact. You did well in the competition , please how did you creat a model for each of clicks, carts, and orders"
  },
  "source": "meta"
}