{
  "id": 370116,
  "title": "Surprising LB 0.587 !!!Share some experimental results ~~",
  "url": "/competitions/otto-recommender-system/discussion/370116",
  "author_name": "Try Harder",
  "post_date": "2022-12-03T06:07:06.817000",
  "votes": 112,
  "comment_count": 35,
  "views": 0,
  "content": "<p>I have recall 200 candidates for each user(user mean session in this competition). Recall rate for each type as follow:</p>\n<ul>\n<li>clicks recall = 0.6506179886006195</li>\n<li>carts recall = 0.527631391786734</li>\n<li>orders recall = 0.7216145392798664<br>\nAnd generate 300+ features for each user and item pair.In training stage, I downsample positive:negative as 1:20 for training .Because of  too large clicks  data, I just train rank model for carts and orders . I got 0.581 !<br>\nCV for top20 recall  as bellow:</li>\n<li>clicks  not compute</li>\n<li>carts recall = 0.499646129454172</li>\n<li>orders recall = 0.6954449845676549</li>\n</ul>\n<p>========update===========<br>\nI delete aid col in features and got 0.585 !<br>\nCV for top20 recall  as bellow:</p>\n<ul>\n<li>clicks  not compute</li>\n<li>carts recall = 0.501434563438234</li>\n<li>orders recall = 0.6963035783251358</li>\n</ul>\n<p>========update===========<br>\nI rank clicks and get 0.587 in LB</p>",
  "messages": [
    {
      "id": 2053339,
      "postDate": "2022-12-03T06:07:06.817Z",
      "content": "<p>I have recall 200 candidates for each user(user mean session in this competition). Recall rate for each type as follow:</p>\n<ul>\n<li>clicks recall = 0.6506179886006195</li>\n<li>carts recall = 0.527631391786734</li>\n<li>orders recall = 0.7216145392798664<br>\nAnd generate 300+ features for each user and item pair.In training stage, I downsample positive:negative as 1:20 for training .Because of  too large clicks  data, I just train rank model for carts and orders . I got 0.581 !<br>\nCV for top20 recall  as bellow:</li>\n<li>clicks  not compute</li>\n<li>carts recall = 0.499646129454172</li>\n<li>orders recall = 0.6954449845676549</li>\n</ul>\n<p>========update===========<br>\nI delete aid col in features and got 0.585 !<br>\nCV for top20 recall  as bellow:</p>\n<ul>\n<li>clicks  not compute</li>\n<li>carts recall = 0.501434563438234</li>\n<li>orders recall = 0.6963035783251358</li>\n</ul>\n<p>========update===========<br>\nI rank clicks and get 0.587 in LB</p>",
      "rawMarkdown": "I have recall 200 candidates for each user(user mean session in this competition). Recall rate for each type as follow:\n- clicks recall = 0.6506179886006195\n- carts recall = 0.527631391786734\n- orders recall = 0.7216145392798664\nAnd generate 300+ features for each user and item pair.In training stage, I downsample positive:negative as 1:20 for training .Because of  too large clicks  data, I just train rank model for carts and orders . I got 0.581 !\nCV for top20 recall  as bellow:\n-  clicks  not compute\n- carts recall = 0.499646129454172\n- orders recall = 0.6954449845676549\n\n========update===========\nI delete aid col in features and got 0.585 !\nCV for top20 recall  as bellow:\n-  clicks  not compute\n- carts recall = 0.501434563438234\n- orders recall = 0.6963035783251358\n\n========update===========\nI rank clicks and get 0.587 in LB",
      "votes": 111
    },
    {
      "id": 2054331,
      "postDate": "2022-12-04T03:22:09.113Z",
      "content": "<p>I saw many people mention about <code>user</code>, however, I couldn't find the <code>user</code> in the data…</p>",
      "rawMarkdown": "I saw many people mention about `user`, however, I couldn't find the `user` in the data...",
      "votes": 3,
      "replies": [
        {
          "id": 2054332,
          "postDate": "2022-12-04T03:24:41.727Z",
          "content": "<p>The column \"session\" actually means \"user\". See discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "The column \"session\" actually means \"user\". See discussion [here][1]\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138",
          "votes": 4
        },
        {
          "id": 2054333,
          "postDate": "2022-12-04T03:25:52.623Z",
          "content": "<p>Thanks, Chris, super useful👍</p>",
          "rawMarkdown": "Thanks, Chris, super useful👍",
          "votes": 1
        },
        {
          "id": 2059118,
          "postDate": "2022-12-08T13:49:36.370Z",
          "content": "<p>Can I assume that session id = user id?</p>",
          "rawMarkdown": "Can I assume that session id = user id?",
          "votes": 1
        },
        {
          "id": 2059121,
          "postDate": "2022-12-08T13:50:27.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/guanghan\" target=\"_blank\">@guanghan</a> yes</p>",
          "rawMarkdown": "@guanghan yes",
          "replies": [
            {
              "id": 2059745,
              "postDate": "2022-12-09T07:21:08.530Z",
              "content": "<p>thanks so much for great advice</p>",
              "rawMarkdown": "thanks so much for great advice"
            }
          ]
        }
      ]
    },
    {
      "id": 2098989,
      "postDate": "2023-01-14T02:37:31.693Z",
      "content": "<p>Really thanks for sharing your experiment. I got 0.584 with 78 features using the same negative sampling methods as yours:)</p>",
      "rawMarkdown": "Really thanks for sharing your experiment. I got 0.584 with 78 features using the same negative sampling methods as yours:)",
      "votes": 1
    },
    {
      "id": 2077450,
      "postDate": "2022-12-27T15:31:00.607Z",
      "content": "<p>Your recall scores of candidates are super high, with the strategies in the public notebooks shared by Chris. The recall scores of 200 candidates each are as following:</p>\n<pre><code>clicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\n</code></pre>\n<p>I'm trying to add the candidates generated by <code>item2vec</code></p>",
      "rawMarkdown": "Your recall scores of candidates are super high, with the strategies in the public notebooks shared by Chris. The recall scores of 200 candidates each are as following:\n\n```\nclicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\n```\n\nI'm trying to add the candidates generated by `item2vec`",
      "votes": 1
    },
    {
      "id": 2075106,
      "postDate": "2022-12-25T04:24:25.440Z",
      "content": "<p>Top 200:<br>\norders recall = 0.7216145392798664</p>\n<p>CV for top20 recall as bellow:<br>\norders recall = 0.6954449845676549</p>\n<p>So the reranker only loses &lt; 0.03 points after reranking(top200 to top20)? The reranker is so good?</p>",
      "rawMarkdown": "Top 200:\norders recall = 0.7216145392798664\n\nCV for top20 recall as bellow:\norders recall = 0.6954449845676549\n\nSo the reranker only loses < 0.03 points after reranking(top200 to top20)? The reranker is so good?",
      "votes": 1
    },
    {
      "id": 2054603,
      "postDate": "2022-12-04T09:22:37.840Z",
      "content": "<p>Thanks for sharing, that are great scores. Do you have local CV scores as well?</p>",
      "rawMarkdown": "Thanks for sharing, that are great scores. Do you have local CV scores as well?",
      "votes": 1,
      "replies": [
        {
          "id": 2057338,
          "postDate": "2022-12-07T01:40:35.267Z",
          "content": "<p>I use <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> CV set. It can be seen <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">CV</a></p>",
          "rawMarkdown": "I use @cdeotte CV set. It can be seen [CV](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2060584,
      "postDate": "2022-12-10T06:20:26.520Z",
      "content": "<p>Oh ! This is almost the winning solution for this competition.</p>",
      "rawMarkdown": "Oh ! This is almost the winning solution for this competition.",
      "votes": 2
    },
    {
      "id": 2103601,
      "postDate": "2023-01-17T09:14:21.700Z",
      "content": "<p>hello,i'm a newbie,and i have a stupid question : what is the meaning of \"downsample positive:negative:, negative means no \"clicks carts orders\" ? but the data only include positive. i don't .understand .thank for your sharing !</p>",
      "rawMarkdown": "hello,i'm a newbie,and i have a stupid question : what is the meaning of \"downsample positive:negative:, negative means no \"clicks carts orders\" ? but the data only include positive. i don't .understand .thank for your sharing !",
      "replies": [
        {
          "id": 2103643,
          "postDate": "2023-01-17T09:40:07.327Z",
          "content": "<p>Normally people will retrieve a lot of item candidates based on the user's historical behavior, and among these candidates some are the correct items and majority are negtive. In order to reduce the data size, we have to sample only part of the negtive samples.</p>",
          "rawMarkdown": "Normally people will retrieve a lot of item candidates based on the user's historical behavior, and among these candidates some are the correct items and majority are negtive. In order to reduce the data size, we have to sample only part of the negtive samples.",
          "votes": 4,
          "replies": [
            {
              "id": 2103778,
              "postDate": "2023-01-17T11:34:25.877Z",
              "content": "<p>thanks a lot !</p>",
              "rawMarkdown": "thanks a lot !"
            }
          ]
        }
      ]
    },
    {
      "id": 2080272,
      "postDate": "2022-12-30T03:18:42.217Z",
      "content": "<p>so you trained two models, one for carts and the other for orders, the clicks used rules?</p>",
      "rawMarkdown": "so you trained two models, one for carts and the other for orders, the clicks used rules?"
    },
    {
      "id": 2070469,
      "postDate": "2022-12-20T03:25:06.957Z",
      "content": "<p>Hello, I would like to ask what is the downsampling method you use. Direct downsampling will lose all the data of some sessions. Thank you very much.</p>",
      "rawMarkdown": "Hello, I would like to ask what is the downsampling method you use. Direct downsampling will lose all the data of some sessions. Thank you very much."
    },
    {
      "id": 2066003,
      "postDate": "2022-12-15T09:24:32.347Z",
      "content": "<p>Thank you for sharing and congrats on the high score!<br>\nDo you happen to measure classifier evaluation metrics such as ROC AUC, PR AUC, AP in the validation of these models?</p>\n<p>i am curious on the correlation between classifier metrics to the recall@20 for each model<br>\ni found difficulty on increasing recall@20 even though i got huge bump on the classifier metrics when adding more relevant features to the model</p>\n<p>my setup is: </p>\n<ol>\n<li>use 80 candidates (only from covisitation)</li>\n<li>160-ish features of user, user-item and similarity distance</li>\n<li>the model relies much to  covisitation weights</li>\n</ol>\n<p>click model AP bump from 12% to 30%<br>\ncart model AP bump from 33% to 40%<br>\norder model AP bump from 50% to 71%</p>\n<p>the CV recall@20 is still 0.564</p>",
      "rawMarkdown": "Thank you for sharing and congrats on the high score!\nDo you happen to measure classifier evaluation metrics such as ROC AUC, PR AUC, AP in the validation of these models?\n\ni am curious on the correlation between classifier metrics to the recall@20 for each model\ni found difficulty on increasing recall@20 even though i got huge bump on the classifier metrics when adding more relevant features to the model\n\nmy setup is: \n1. use 80 candidates (only from covisitation)\n2. 160-ish features of user, user-item and similarity distance\n3. the model relies much to  covisitation weights\n\nclick model AP bump from 12% to 30%\ncart model AP bump from 33% to 40%\norder model AP bump from 50% to 71%\n\nthe CV recall@20 is still 0.564",
      "replies": [
        {
          "id": 2070034,
          "postDate": "2022-12-19T14:50:47.560Z",
          "content": "<p>Hello, could u tell us some features? </p>",
          "rawMarkdown": "Hello, could u tell us some features? "
        }
      ]
    },
    {
      "id": 2062357,
      "postDate": "2022-12-12T03:27:23.130Z",
      "content": "<p>300+ features! that's quite a lot, could you share some strong features?</p>",
      "rawMarkdown": "300+ features! that's quite a lot, could you share some strong features?"
    },
    {
      "id": 2054564,
      "postDate": "2022-12-04T08:40:47.693Z",
      "content": "<p>Great work.  What would be your score if you had just used covisit matrices ?</p>",
      "rawMarkdown": "Great work.  What would be your score if you had just used covisit matrices ?",
      "replies": [
        {
          "id": 2057333,
          "postDate": "2022-12-07T01:35:27.673Z",
          "content": "<p>I had 0.576 in LB for only covisit matrices.</p>",
          "rawMarkdown": "I had 0.576 in LB for only covisit matrices.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2054460,
      "postDate": "2022-12-04T06:54:35.463Z",
      "content": "<p>Nice Work!  thanks for sharing the exp records.  I have two quetions :</p>\n<ol>\n<li>\"I delete aid col in features\" means delete the candidates aid id in feature?</li>\n<li>\"I rank clicks and get 0.587 in LB\" means use the same pipeline train click rank model?</li>\n</ol>\n<p>Thanks again for sharing ,  great job!</p>",
      "rawMarkdown": "Nice Work!  thanks for sharing the exp records.  I have two quetions :\n1. \"I delete aid col in features\" means delete the candidates aid id in feature?\n2. \"I rank clicks and get 0.587 in LB\" means use the same pipeline train click rank model?\n\nThanks again for sharing ,  great job!",
      "replies": [
        {
          "id": 2054472,
          "postDate": "2022-12-04T07:05:17.273Z",
          "content": "<p>Thanks.1.Yes, it can increase the generalization ability of the model and don't let the model just remember the aid id. 2. Yes, I train three model  separately. And they are the same  pipeline,just different from training data.</p>",
          "rawMarkdown": "Thanks.1.Yes, it can increase the generalization ability of the model and don't let the model just remember the aid id. 2. Yes, I train three model  separately. And they are the same  pipeline,just different from training data.",
          "votes": 3
        },
        {
          "id": 2054475,
          "postDate": "2022-12-04T07:13:27.520Z",
          "content": "<p>Thanks！</p>\n<p>Your message must have at least 10 characters.</p>",
          "rawMarkdown": "Thanks！\n\nYour message must have at least 10 characters.\n"
        }
      ]
    },
    {
      "id": 2054301,
      "postDate": "2022-12-04T02:28:29.317Z",
      "content": "<p>Thanks for your sharing, can I join your team ? I'm ML engineering based in Guangzhou.</p>",
      "rawMarkdown": "Thanks for your sharing, can I join your team ? I'm ML engineering based in Guangzhou.",
      "replies": [
        {
          "id": 2054352,
          "postDate": "2022-12-04T04:17:48.097Z",
          "content": "<p>Thanks for your offer, I have teamed up.</p>",
          "rawMarkdown": "Thanks for your offer, I have teamed up."
        }
      ]
    },
    {
      "id": 2053844,
      "postDate": "2022-12-03T16:49:01.163Z",
      "content": "<p>Hey! Congrats to ur great results!<br>\nMay I ask two questions:<br>\nWhat do you mean exactly by <br>\n\"downsample positive:negative as 1:20 for training .Because of too large clicks data, I just train rank model for carts and orders\"</p>\n<ul>\n<li>what is meant by downsampling? Im not familiar with that term.</li>\n<li>regarding the second point: do I understand correctly that you use for clicks usual Co visitation matrix and for the other ones discard all clicks from training data and just use carts and orders than train seperately a ranker for each. The one with positive examples = carts grund truth and the other with orders ground truth?</li>\n</ul>\n<p>Tnx for clarifying I am sorry but I'm still beginner and want to understand what u did better :)</p>",
      "rawMarkdown": "Hey! Congrats to ur great results!\nMay I ask two questions:\nWhat do you mean exactly by \n\"downsample positive:negative as 1:20 for training .Because of too large clicks data, I just train rank model for carts and orders\"\n \n- what is meant by downsampling? Im not familiar with that term.\n- regarding the second point: do I understand correctly that you use for clicks usual Co visitation matrix and for the other ones discard all clicks from training data and just use carts and orders than train seperately a ranker for each. The one with positive examples = carts grund truth and the other with orders ground truth?\n\nTnx for clarifying I am sorry but I'm still beginner and want to understand what u did better :)",
      "replies": [
        {
          "id": 2053847,
          "postDate": "2022-12-03T16:51:00.913Z",
          "content": "<p>And maybe another small question: 300 features seems a lot to me. Did you engineer them urself or did u use some kind of library to generate additional feats</p>",
          "rawMarkdown": "And maybe another small question: 300 features seems a lot to me. Did you engineer them urself or did u use some kind of library to generate additional feats"
        },
        {
          "id": 2054357,
          "postDate": "2022-12-04T04:24:08.693Z",
          "content": "<p>I think \"downsampling\" in my discussion means that  choose 1 sample which user clicks, carts or orders and corresponding  20 samples which don't because of limitations of memory. Train seperately a ranker for each type. Later maybe I will consider together.</p>",
          "rawMarkdown": "I think \"downsampling\" in my discussion means that  choose 1 sample which user clicks, carts or orders and corresponding  20 samples which don't because of limitations of memory. Train seperately a ranker for each type. Later maybe I will consider together.\n"
        }
      ]
    },
    {
      "id": 2053650,
      "postDate": "2022-12-03T13:13:09.283Z",
      "content": "<p>Nice work, great scores! What would your LB be if you submitted clicks too?</p>",
      "rawMarkdown": "Nice work, great scores! What would your LB be if you submitted clicks too?",
      "replies": [
        {
          "id": 2053659,
          "postDate": "2022-12-03T13:22:25.020Z",
          "content": "<p>Thanks. I learn a lot from your notebooks. Now clicks are only generated  by co-visited matrix, about 0.5 recall rate in CV. If I rank the clicks, maybe  about 0.01 up in LB.</p>",
          "rawMarkdown": "Thanks. I learn a lot from your notebooks. Now clicks are only generated  by co-visited matrix, about 0.5 recall rate in CV. If I rank the clicks, maybe  about 0.01 up in LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2056910,
      "postDate": "2022-12-06T15:06:52.297Z",
      "content": "<p>May I ask about the memory cost and time cost, and the hardware?</p>",
      "rawMarkdown": "May I ask about the memory cost and time cost, and the hardware?",
      "isDeleted": true,
      "replies": [
        {
          "id": 2057337,
          "postDate": "2022-12-07T01:38:38.767Z",
          "content": "<p>128 GB RAM. 1 day for generating features and  1 day for  training (My teammate can optimize to 10 + minutes for features ).</p>",
          "rawMarkdown": "128 GB RAM. 1 day for generating features and  1 day for  training (My teammate can optimize to 10 + minutes for features ).",
          "votes": 4,
          "replies": [
            {
              "id": 2065175,
              "postDate": "2022-12-14T11:11:50.787Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2054331,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2022-12-04T03:22:09.113000",
      "content": "<p>I saw many people mention about <code>user</code>, however, I couldn't find the <code>user</code> in the data…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2054332,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-04T03:24:41.727000",
          "content": "<p>The column \"session\" actually means \"user\". See discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138\" target=\"_blank\">here</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2054333,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2022-12-04T03:25:52.623000",
          "content": "<p>Thanks, Chris, super useful👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059118,
          "author_name": "willsionwwwwww",
          "author_url": "",
          "post_date": "2022-12-08T13:49:36.370000",
          "content": "<p>Can I assume that session id = user id?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059121,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-08T13:50:27.677000",
          "content": "<p><a href=\"https://www.kaggle.com/guanghan\" target=\"_blank\">@guanghan</a> yes</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2059745,
              "author_name": "willsionwwwwww",
              "author_url": "",
              "post_date": "2022-12-09T07:21:08.530000",
              "content": "<p>thanks so much for great advice</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2098989,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2023-01-14T02:37:31.693000",
      "content": "<p>Really thanks for sharing your experiment. I got 0.584 with 78 features using the same negative sampling methods as yours:)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2077450,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2022-12-27T15:31:00.607000",
      "content": "<p>Your recall scores of candidates are super high, with the strategies in the public notebooks shared by Chris. The recall scores of 200 candidates each are as following:</p>\n<pre><code>clicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\n</code></pre>\n<p>I'm trying to add the candidates generated by <code>item2vec</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2075106,
      "author_name": "Homoalways",
      "author_url": "",
      "post_date": "2022-12-25T04:24:25.440000",
      "content": "<p>Top 200:<br>\norders recall = 0.7216145392798664</p>\n<p>CV for top20 recall as bellow:<br>\norders recall = 0.6954449845676549</p>\n<p>So the reranker only loses &lt; 0.03 points after reranking(top200 to top20)? The reranker is so good?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2054603,
      "author_name": "BenediktSchifferer",
      "author_url": "",
      "post_date": "2022-12-04T09:22:37.840000",
      "content": "<p>Thanks for sharing, that are great scores. Do you have local CV scores as well?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2057338,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-07T01:40:35.267000",
          "content": "<p>I use <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> CV set. It can be seen <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">CV</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2060584,
      "author_name": "toshi_k",
      "author_url": "",
      "post_date": "2022-12-10T06:20:26.520000",
      "content": "<p>Oh ! This is almost the winning solution for this competition.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2103601,
      "author_name": "cghsncg",
      "author_url": "",
      "post_date": "2023-01-17T09:14:21.700000",
      "content": "<p>hello,i'm a newbie,and i have a stupid question : what is the meaning of \"downsample positive:negative:, negative means no \"clicks carts orders\" ? but the data only include positive. i don't .understand .thank for your sharing !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2103643,
          "author_name": "BUUMOO",
          "author_url": "",
          "post_date": "2023-01-17T09:40:07.327000",
          "content": "<p>Normally people will retrieve a lot of item candidates based on the user's historical behavior, and among these candidates some are the correct items and majority are negtive. In order to reduce the data size, we have to sample only part of the negtive samples.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2103778,
              "author_name": "cghsncg",
              "author_url": "",
              "post_date": "2023-01-17T11:34:25.877000",
              "content": "<p>thanks a lot !</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2080272,
      "author_name": "jxlijunhao",
      "author_url": "",
      "post_date": "2022-12-30T03:18:42.217000",
      "content": "<p>so you trained two models, one for carts and the other for orders, the clicks used rules?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2070469,
      "author_name": "caicai",
      "author_url": "",
      "post_date": "2022-12-20T03:25:06.957000",
      "content": "<p>Hello, I would like to ask what is the downsampling method you use. Direct downsampling will lose all the data of some sessions. Thank you very much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2066003,
      "author_name": "Aji Samudra",
      "author_url": "",
      "post_date": "2022-12-15T09:24:32.347000",
      "content": "<p>Thank you for sharing and congrats on the high score!<br>\nDo you happen to measure classifier evaluation metrics such as ROC AUC, PR AUC, AP in the validation of these models?</p>\n<p>i am curious on the correlation between classifier metrics to the recall@20 for each model<br>\ni found difficulty on increasing recall@20 even though i got huge bump on the classifier metrics when adding more relevant features to the model</p>\n<p>my setup is: </p>\n<ol>\n<li>use 80 candidates (only from covisitation)</li>\n<li>160-ish features of user, user-item and similarity distance</li>\n<li>the model relies much to  covisitation weights</li>\n</ol>\n<p>click model AP bump from 12% to 30%<br>\ncart model AP bump from 33% to 40%<br>\norder model AP bump from 50% to 71%</p>\n<p>the CV recall@20 is still 0.564</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2070034,
          "author_name": "deeeeeeeplearning",
          "author_url": "",
          "post_date": "2022-12-19T14:50:47.560000",
          "content": "<p>Hello, could u tell us some features? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2062357,
      "author_name": "Lukan",
      "author_url": "",
      "post_date": "2022-12-12T03:27:23.130000",
      "content": "<p>300+ features! that's quite a lot, could you share some strong features?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2054564,
      "author_name": "NikhilMishra",
      "author_url": "",
      "post_date": "2022-12-04T08:40:47.693000",
      "content": "<p>Great work.  What would be your score if you had just used covisit matrices ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2057333,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-07T01:35:27.673000",
          "content": "<p>I had 0.576 in LB for only covisit matrices.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2054460,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-12-04T06:54:35.463000",
      "content": "<p>Nice Work!  thanks for sharing the exp records.  I have two quetions :</p>\n<ol>\n<li>\"I delete aid col in features\" means delete the candidates aid id in feature?</li>\n<li>\"I rank clicks and get 0.587 in LB\" means use the same pipeline train click rank model?</li>\n</ol>\n<p>Thanks again for sharing ,  great job!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2054472,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-04T07:05:17.273000",
          "content": "<p>Thanks.1.Yes, it can increase the generalization ability of the model and don't let the model just remember the aid id. 2. Yes, I train three model  separately. And they are the same  pipeline,just different from training data.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2054475,
          "author_name": "KKY",
          "author_url": "",
          "post_date": "2022-12-04T07:13:27.520000",
          "content": "<p>Thanks！</p>\n<p>Your message must have at least 10 characters.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2054301,
      "author_name": "Qucy Wei",
      "author_url": "",
      "post_date": "2022-12-04T02:28:29.317000",
      "content": "<p>Thanks for your sharing, can I join your team ? I'm ML engineering based in Guangzhou.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2054352,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-04T04:17:48.097000",
          "content": "<p>Thanks for your offer, I have teamed up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2053844,
      "author_name": "Simon Veitner",
      "author_url": "",
      "post_date": "2022-12-03T16:49:01.163000",
      "content": "<p>Hey! Congrats to ur great results!<br>\nMay I ask two questions:<br>\nWhat do you mean exactly by <br>\n\"downsample positive:negative as 1:20 for training .Because of too large clicks data, I just train rank model for carts and orders\"</p>\n<ul>\n<li>what is meant by downsampling? Im not familiar with that term.</li>\n<li>regarding the second point: do I understand correctly that you use for clicks usual Co visitation matrix and for the other ones discard all clicks from training data and just use carts and orders than train seperately a ranker for each. The one with positive examples = carts grund truth and the other with orders ground truth?</li>\n</ul>\n<p>Tnx for clarifying I am sorry but I'm still beginner and want to understand what u did better :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2053847,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-03T16:51:00.913000",
          "content": "<p>And maybe another small question: 300 features seems a lot to me. Did you engineer them urself or did u use some kind of library to generate additional feats</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2054357,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-04T04:24:08.693000",
          "content": "<p>I think \"downsampling\" in my discussion means that  choose 1 sample which user clicks, carts or orders and corresponding  20 samples which don't because of limitations of memory. Train seperately a ranker for each type. Later maybe I will consider together.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2053650,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-12-03T13:13:09.283000",
      "content": "<p>Nice work, great scores! What would your LB be if you submitted clicks too?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2053659,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-03T13:22:25.020000",
          "content": "<p>Thanks. I learn a lot from your notebooks. Now clicks are only generated  by co-visited matrix, about 0.5 recall rate in CV. If I rank the clicks, maybe  about 0.01 up in LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2056910,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-06T15:06:52.297000",
      "content": "<p>May I ask about the memory cost and time cost, and the hardware?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2057337,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-12-07T01:38:38.767000",
          "content": "<p>128 GB RAM. 1 day for generating features and  1 day for  training (My teammate can optimize to 10 + minutes for features ).</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2065175,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-12-14T11:11:50.787000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2053339": "I have recall 200 candidates for each user(user mean session in this competition). Recall rate for each type as follow:\n- clicks recall = 0.6506179886006195\n- carts recall = 0.527631391786734\n- orders recall = 0.7216145392798664\nAnd generate 300+ features for each user and item pair.In training stage, I downsample positive:negative as 1:20 for training .Because of  too large clicks  data, I just train rank model for carts and orders . I got 0.581 !\nCV for top20 recall  as bellow:\n-  clicks  not compute\n- carts recall = 0.499646129454172\n- orders recall = 0.6954449845676549\n\n========update===========\nI delete aid col in features and got 0.585 !\nCV for top20 recall  as bellow:\n-  clicks  not compute\n- carts recall = 0.501434563438234\n- orders recall = 0.6963035783251358\n\n========update===========\nI rank clicks and get 0.587 in LB",
    "2054331": "I saw many people mention about `user`, however, I couldn't find the `user` in the data...",
    "2098989": "Really thanks for sharing your experiment. I got 0.584 with 78 features using the same negative sampling methods as yours:)",
    "2077450": "Your recall scores of candidates are super high, with the strategies in the public notebooks shared by Chris. The recall scores of 200 candidates each are as following:\n\n```\nclicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\n```\n\nI'm trying to add the candidates generated by `item2vec`",
    "2075106": "Top 200:\norders recall = 0.7216145392798664\n\nCV for top20 recall as bellow:\norders recall = 0.6954449845676549\n\nSo the reranker only loses < 0.03 points after reranking(top200 to top20)? The reranker is so good?",
    "2054603": "Thanks for sharing, that are great scores. Do you have local CV scores as well?",
    "2060584": "Oh ! This is almost the winning solution for this competition.",
    "2103601": "hello,i'm a newbie,and i have a stupid question : what is the meaning of \"downsample positive:negative:, negative means no \"clicks carts orders\" ? but the data only include positive. i don't .understand .thank for your sharing !",
    "2080272": "so you trained two models, one for carts and the other for orders, the clicks used rules?",
    "2070469": "Hello, I would like to ask what is the downsampling method you use. Direct downsampling will lose all the data of some sessions. Thank you very much.",
    "2066003": "Thank you for sharing and congrats on the high score!\nDo you happen to measure classifier evaluation metrics such as ROC AUC, PR AUC, AP in the validation of these models?\n\ni am curious on the correlation between classifier metrics to the recall@20 for each model\ni found difficulty on increasing recall@20 even though i got huge bump on the classifier metrics when adding more relevant features to the model\n\nmy setup is: \n1. use 80 candidates (only from covisitation)\n2. 160-ish features of user, user-item and similarity distance\n3. the model relies much to  covisitation weights\n\nclick model AP bump from 12% to 30%\ncart model AP bump from 33% to 40%\norder model AP bump from 50% to 71%\n\nthe CV recall@20 is still 0.564",
    "2062357": "300+ features! that's quite a lot, could you share some strong features?",
    "2054564": "Great work.  What would be your score if you had just used covisit matrices ?",
    "2054460": "Nice Work!  thanks for sharing the exp records.  I have two quetions :\n1. \"I delete aid col in features\" means delete the candidates aid id in feature?\n2. \"I rank clicks and get 0.587 in LB\" means use the same pipeline train click rank model?\n\nThanks again for sharing ,  great job!",
    "2054301": "Thanks for your sharing, can I join your team ? I'm ML engineering based in Guangzhou.",
    "2053844": "Hey! Congrats to ur great results!\nMay I ask two questions:\nWhat do you mean exactly by \n\"downsample positive:negative as 1:20 for training .Because of too large clicks data, I just train rank model for carts and orders\"\n \n- what is meant by downsampling? Im not familiar with that term.\n- regarding the second point: do I understand correctly that you use for clicks usual Co visitation matrix and for the other ones discard all clicks from training data and just use carts and orders than train seperately a ranker for each. The one with positive examples = carts grund truth and the other with orders ground truth?\n\nTnx for clarifying I am sorry but I'm still beginner and want to understand what u did better :)",
    "2053650": "Nice work, great scores! What would your LB be if you submitted clicks too?",
    "2056910": "May I ask about the memory cost and time cost, and the hardware?"
  }
}