{
  "id": 381469,
  "title": "Top 200 : Empirical results/remarks about training the ranker.",
  "url": "/competitions/otto-recommender-system/discussion/381469",
  "author_name": "",
  "post_date": "2023-01-26T20:39:51.832805300Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>-No performance improvements when using  1% of negatives samples vs all negative samples. --&gt; Use subsampling.</p>\n<p>-When using a relative low learning rate ( 0.1 ) my carts recall got higher whereas my orders recall got lower, and vice-versa --&gt; maybe train a slow learner for carts and fast learner for orders  ?</p>\n<p>-Interactions features are more important than users and items features ( from the feature importance of the GBDT ).</p>\n<p>-The Ranker might be better than heuristic approaches in some cases, and vice-versa --&gt; you can use both of them.</p>\n<p>-Scale pos weight parameter ( ratio between negatives and positive samples ) does not help  improving my recall score.</p>\n<p>This competition is kind of painful when using only kaggle kernel 😤</p>",
  "messages": [
    {
      "id": "2116944",
      "postDate": "01/26/2023 20:39:51",
      "content": "<p>-No performance improvements when using  1% of negatives samples vs all negative samples. --&gt; Use subsampling.</p>\n<p>-When using a relative low learning rate ( 0.1 ) my carts recall got higher whereas my orders recall got lower, and vice-versa --&gt; maybe train a slow learner for carts and fast learner for orders  ?</p>\n<p>-Interactions features are more important than users and items features ( from the feature importance of the GBDT ).</p>\n<p>-The Ranker might be better than heuristic approaches in some cases, and vice-versa --&gt; you can use both of them.</p>\n<p>-Scale pos weight parameter ( ratio between negatives and positive samples ) does not help  improving my recall score.</p>\n<p>This competition is kind of painful when using only kaggle kernel 😤</p>",
      "rawMarkdown": "No performance improvements when using  1% of negatives samples vs all negative samples. --> Use subsampling.\n\n-When using a relative low learning rate ( 0.1 ) my carts recall got higher whereas my orders recall got lower, and vice-versa --> maybe train a slow learner for carts and fast learner for orders  ?\n\n-Interactions features are more important than users and items features ( from the feature importance of the GBDT ).\n\n-The Ranker might be better than heuristic approaches in some cases, and vice-versa --> you can use both of them.\n\n-Scale pos weight parameter ( ratio between negatives and positive samples ) does not help  improving my recall score.\n\n\nThis competition is kind of painful when using only kaggle kernel 😤",
      "votes": null
    },
    {
      "id": "2116996",
      "postDate": "01/26/2023 21:49:52",
      "content": "<p>Tuning downsample percentage also didn't work for me. I can manipulate CV with it, but cannot improve LB.</p>",
      "rawMarkdown": "Tuning downsample percentage also didn't work for me. I can manipulate CV with it, but cannot improve LB.",
      "votes": null
    },
    {
      "id": "2117006",
      "postDate": "01/26/2023 22:17:21",
      "content": "<p>Totally agree on this point, thanks for the remark !</p>",
      "rawMarkdown": "Totally agree on this point, thanks for the remark !",
      "votes": null
    },
    {
      "id": "2117117",
      "postDate": "01/27/2023 02:53:58",
      "content": "<p>But improved the training speed when you have a lot of candidates 😄</p>",
      "rawMarkdown": "But improved the training speed when you have a lot of candidates 😄",
      "votes": null
    },
    {
      "id": "2117118",
      "postDate": "01/27/2023 02:57:40",
      "content": "<p>I am surprised that you can do this competition using only kaggle kernel😂</p>",
      "rawMarkdown": "I am surprised that you can do this competition using only kaggle kernel😂",
      "votes": null
    },
    {
      "id": "2117441",
      "postDate": "01/27/2023 09:22:29",
      "content": "<p>I can assure you it's true 😂. I have become an expert in CHUNK, I'm also using a dozen notebooks for the processing of each step.</p>\n<p>If you don't mind <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>, What are you using for this competition ( i mean your setup ) ? I heard about google colab Pro + in some threads but don't know if it's enough for kaggle competitions, however I'm sure that it's better than my actual setup on kaggle kernel XD</p>",
      "rawMarkdown": "I can assure you it's true 😂. I have become an expert in CHUNK, I'm also using a dozen notebooks for the processing of each step.\n\nIf you don't mind @buumoo, What are you using for this competition ( i mean your setup ) ? I heard about google colab Pro + in some threads but don't know if it's enough for kaggle competitions, however I'm sure that it's better than my actual setup on kaggle kernel XD",
      "votes": null
    },
    {
      "id": "2117454",
      "postDate": "01/27/2023 09:30:44",
      "content": "<p><a href=\"https://www.kaggle.com/gongbi\" target=\"_blank\">@gongbi</a>  How many candidates did you extract from the co-visitation matrices ??</p>",
      "rawMarkdown": "gongbi  How many candidates did you extract from the co-visitation matrices ??",
      "votes": null
    },
    {
      "id": "2117516",
      "postDate": "01/27/2023 10:28:03",
      "content": "<p>My personal pc is poor as well, only 48gb ram with a 3070 GPU. So I have to rent cloud GPU server to do this competition. Previously I used google colab before it changed from monthly subscription to current compute units system. It becomes much more expensive in this way, for me at least, which depands on how much of your usage.</p>\n<p>For this competition, 128gb ram with a 3090 GPU is enough for all my methods.<br>\nGCP and AWS are secure but expensive.<br>\nSome small private cloud platforms are much cheaper, although you may face some security issue. Maybe your code/solution will be stolen. such as <a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/337853\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/337853</a>, so I will not advertise or mention any of them here.</p>",
      "rawMarkdown": "My personal pc is poor as well, only 48gb ram with a 3070 GPU. So I have to rent cloud GPU server to do this competition. Previously I used google colab before it changed from monthly subscription to current compute units system. It becomes much more expensive in this way, for me at least, which depands on how much of your usage.\n\nFor this competition, 128gb ram with a 3090 GPU is enough for all my methods.\nGCP and AWS are secure but expensive.\nSome small private cloud platforms are much cheaper, although you may face some security issue. Maybe your code/solution will be stolen. such as https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/337853, so I will not advertise or mention any of them here.",
      "votes": null
    },
    {
      "id": "2118502",
      "postDate": "01/28/2023 04:02:10",
      "content": "<p>May I ask how you do neg sample? by filter out those who do not have any cart/order labels?(e.g. A session whose 100 candidates all fail to hit the groundtruth)</p>",
      "rawMarkdown": "May I ask how you do neg sample? by filter out those who do not have any cart/order labels?(e.g. A session whose 100 candidates all fail to hit the groundtruth)",
      "votes": null
    },
    {
      "id": "2118567",
      "postDate": "01/28/2023 05:48:35",
      "content": "<p><a href=\"https://www.kaggle.com/yuzhang0422\" target=\"_blank\">@yuzhang0422</a> Exactly, I just randomly remove samples that have target=0. I do that for each dataframe (clicks, carts and orders):</p>\n<pre><code>print(cand_carts.shape)\n\ncand_carts = pl.concat([\n             cand_carts.filter(~(pl.col(\"cart_target\") == 0)),\n             cand_carts.filter((pl.col(\"cart_target\") == 0)).sample(frac=CFG.carts_frac, seed=CFG.seed)\n             ]).sort(['session'])\n\nprint(cand_carts.shape)\n</code></pre>",
      "rawMarkdown": "yuzhang0422 Exactly, I just randomly remove samples that have target=0. I do that for each dataframe (clicks, carts and orders):\n\n```\nprint(cand_carts.shape)\n\ncand_carts = pl.concat([\n             cand_carts.filter(~(pl.col(\"cart_target\") == 0)),\n             cand_carts.filter((pl.col(\"cart_target\") == 0)).sample(frac=CFG.carts_frac, seed=CFG.seed)\n             ]).sort(['session'])\n\nprint(cand_carts.shape)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2116996,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "01/26/2023 21:49:52",
      "content": "<p>Tuning downsample percentage also didn't work for me. I can manipulate CV with it, but cannot improve LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2117006,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/26/2023 22:17:21",
          "content": "<p>Totally agree on this point, thanks for the remark !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2117117,
          "author_name": "gongbi",
          "author_url": "",
          "post_date": "01/27/2023 02:53:58",
          "content": "<p>But improved the training speed when you have a lot of candidates 😄</p>",
          "votes": null,
          "replies": [
            {
              "id": 2117454,
              "author_name": "rayanaay",
              "author_url": "",
              "post_date": "01/27/2023 09:30:44",
              "content": "<p><a href=\"https://www.kaggle.com/gongbi\" target=\"_blank\">@gongbi</a>  How many candidates did you extract from the co-visitation matrices ??</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2118502,
          "author_name": "yuzhang0422",
          "author_url": "",
          "post_date": "01/28/2023 04:02:10",
          "content": "<p>May I ask how you do neg sample? by filter out those who do not have any cart/order labels?(e.g. A session whose 100 candidates all fail to hit the groundtruth)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2118567,
              "author_name": "hinepo",
              "author_url": "",
              "post_date": "01/28/2023 05:48:35",
              "content": "<p><a href=\"https://www.kaggle.com/yuzhang0422\" target=\"_blank\">@yuzhang0422</a> Exactly, I just randomly remove samples that have target=0. I do that for each dataframe (clicks, carts and orders):</p>\n<pre><code>print(cand_carts.shape)\n\ncand_carts = pl.concat([\n             cand_carts.filter(~(pl.col(\"cart_target\") == 0)),\n             cand_carts.filter((pl.col(\"cart_target\") == 0)).sample(frac=CFG.carts_frac, seed=CFG.seed)\n             ]).sort(['session'])\n\nprint(cand_carts.shape)\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2117118,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "01/27/2023 02:57:40",
      "content": "<p>I am surprised that you can do this competition using only kaggle kernel😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 2117441,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/27/2023 09:22:29",
          "content": "<p>I can assure you it's true 😂. I have become an expert in CHUNK, I'm also using a dozen notebooks for the processing of each step.</p>\n<p>If you don't mind <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>, What are you using for this competition ( i mean your setup ) ? I heard about google colab Pro + in some threads but don't know if it's enough for kaggle competitions, however I'm sure that it's better than my actual setup on kaggle kernel XD</p>",
          "votes": null,
          "replies": [
            {
              "id": 2117516,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "01/27/2023 10:28:03",
              "content": "<p>My personal pc is poor as well, only 48gb ram with a 3070 GPU. So I have to rent cloud GPU server to do this competition. Previously I used google colab before it changed from monthly subscription to current compute units system. It becomes much more expensive in this way, for me at least, which depands on how much of your usage.</p>\n<p>For this competition, 128gb ram with a 3090 GPU is enough for all my methods.<br>\nGCP and AWS are secure but expensive.<br>\nSome small private cloud platforms are much cheaper, although you may face some security issue. Maybe your code/solution will be stolen. such as <a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/337853\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/337853</a>, so I will not advertise or mention any of them here.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2116944": "No performance improvements when using  1% of negatives samples vs all negative samples. --> Use subsampling.\n\n-When using a relative low learning rate ( 0.1 ) my carts recall got higher whereas my orders recall got lower, and vice-versa --> maybe train a slow learner for carts and fast learner for orders  ?\n\n-Interactions features are more important than users and items features ( from the feature importance of the GBDT ).\n\n-The Ranker might be better than heuristic approaches in some cases, and vice-versa --> you can use both of them.\n\n-Scale pos weight parameter ( ratio between negatives and positive samples ) does not help  improving my recall score.\n\n\nThis competition is kind of painful when using only kaggle kernel 😤",
    "2116996": "Tuning downsample percentage also didn't work for me. I can manipulate CV with it, but cannot improve LB.",
    "2117006": "Totally agree on this point, thanks for the remark !",
    "2117117": "But improved the training speed when you have a lot of candidates 😄",
    "2117118": "I am surprised that you can do this competition using only kaggle kernel😂",
    "2117441": "I can assure you it's true 😂. I have become an expert in CHUNK, I'm also using a dozen notebooks for the processing of each step.\n\nIf you don't mind @buumoo, What are you using for this competition ( i mean your setup ) ? I heard about google colab Pro + in some threads but don't know if it's enough for kaggle competitions, however I'm sure that it's better than my actual setup on kaggle kernel XD",
    "2117454": "gongbi  How many candidates did you extract from the co-visitation matrices ??",
    "2117516": "My personal pc is poor as well, only 48gb ram with a 3070 GPU. So I have to rent cloud GPU server to do this competition. Previously I used google colab before it changed from monthly subscription to current compute units system. It becomes much more expensive in this way, for me at least, which depands on how much of your usage.\n\nFor this competition, 128gb ram with a 3090 GPU is enough for all my methods.\nGCP and AWS are secure but expensive.\nSome small private cloud platforms are much cheaper, although you may face some security issue. Maybe your code/solution will be stolen. such as https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/337853, so I will not advertise or mention any of them here.",
    "2118502": "May I ask how you do neg sample? by filter out those who do not have any cart/order labels?(e.g. A session whose 100 candidates all fail to hit the groundtruth)",
    "2118567": "yuzhang0422 Exactly, I just randomly remove samples that have target=0. I do that for each dataframe (clicks, carts and orders):\n\n```\nprint(cand_carts.shape)\n\ncand_carts = pl.concat([\n             cand_carts.filter(~(pl.col(\"cart_target\") == 0)),\n             cand_carts.filter((pl.col(\"cart_target\") == 0)).sample(frac=CFG.carts_frac, seed=CFG.seed)\n             ]).sort(['session'])\n\nprint(cand_carts.shape)\n```"
  },
  "source": "meta"
}