{
  "id": 374324,
  "title": "How many candidates did you generates and how did they perform?",
  "url": "/competitions/otto-recommender-system/discussion/374324",
  "author_name": "Lukan",
  "post_date": "2022-12-26T15:46:51.970000",
  "votes": 13,
  "comment_count": 19,
  "views": 0,
  "content": "<p>In my experiment, I use past event, co-visit matrix to generate candidates, each session has about 52 candidates in average, and here is the CV result of my recall method:</p>\n<pre><code>clicks recall = \ncarts recall = \norders recall = \n=============\nOverall Recall = \n=============\n</code></pre>\n<p>then I create 60 features and bulid my ranking model, it got 0.581 in LB.<br>\nWhat about your results?</p>",
  "messages": [
    {
      "id": 2076552,
      "postDate": "2022-12-26T15:46:51.970Z",
      "content": "<p>In my experiment, I use past event, co-visit matrix to generate candidates, each session has about 52 candidates in average, and here is the CV result of my recall method:</p>\n<pre><code>clicks recall = \ncarts recall = \norders recall = \n=============\nOverall Recall = \n=============\n</code></pre>\n<p>then I create 60 features and bulid my ranking model, it got 0.581 in LB.<br>\nWhat about your results?</p>",
      "rawMarkdown": "In my experiment, I use past event, co-visit matrix to generate candidates, each session has about 52 candidates in average, and here is the CV result of my recall method:\n```python\nclicks recall = 0.6101419852876675\ncarts recall = 0.4724449332329544\norders recall = 0.6862270709185676\n=============\nOverall Recall = 0.6144839210497937\n=============\n```\nthen I create 60 features and bulid my ranking model, it got 0.581 in LB.\nWhat about your results?",
      "votes": 13
    },
    {
      "id": 2082884,
      "postDate": "2023-01-02T04:20:45.380Z",
      "content": "<p>i create 100+ feature, but the rank model doesn't perform well, the good feature i don't find, but some feature i think it's good run too slow,how many time do you generate features?</p>",
      "rawMarkdown": "i create 100+ feature, but the rank model doesn't perform well, the good feature i don't find, but some feature i think it's good run too slow,how many time do you generate features?"
    },
    {
      "id": 2077768,
      "postDate": "2022-12-27T21:29:12.297Z",
      "content": "<p>Did you build for each type ( clicks/carts/orders) one specific classifier for each ?</p>\n<p>Another question is about a classifier I trained ( lightgbm ) using only clicks from co-visitation matrix of clicks, the model converge, but when computing the recall, I got a worser one.</p>",
      "rawMarkdown": "Did you build for each type ( clicks/carts/orders) one specific classifier for each ?\n\nAnother question is about a classifier I trained ( lightgbm ) using only clicks from co-visitation matrix of clicks, the model converge, but when computing the recall, I got a worser one.\n",
      "replies": [
        {
          "id": 2078340,
          "postDate": "2022-12-28T08:58:42.343Z",
          "content": "<ol>\n<li>Yes, I build tree ranking models</li>\n<li>Maybe you should make more features? If you make the right feature, the recall after ranking model should be better than before</li>\n</ol>",
          "rawMarkdown": "1. Yes, I build tree ranking models\n2. Maybe you should make more features? If you make the right feature, the recall after ranking model should be better than before",
          "replies": [
            {
              "id": 2078371,
              "postDate": "2022-12-28T09:30:43.263Z",
              "content": "<p>Ok thanks.</p>\n<p>What about the AUC of your tree ? can you give us an estimation of the AUC at the  last iteration ?</p>",
              "rawMarkdown": "Ok thanks.\n\nWhat about the AUC of your tree ? can you give us an estimation of the AUC at the  last iteration ?"
            },
            {
              "id": 2078682,
              "postDate": "2022-12-28T14:32:37.303Z",
              "content": "<p>Sorry, I hadn't use the AUC as my eval metric, while my ndcg@20 is around 0.88, binary log loss is about 0.3</p>",
              "rawMarkdown": "Sorry, I hadn't use the AUC as my eval metric, while my ndcg@20 is around 0.88, binary log loss is about 0.3",
              "votes": 1
            },
            {
              "id": 2080789,
              "postDate": "2022-12-30T13:38:58.807Z",
              "content": "<p>Awesome, did you downsample negative candidates ??</p>",
              "rawMarkdown": "Awesome, did you downsample negative candidates ??"
            },
            {
              "id": 2082791,
              "postDate": "2023-01-02T00:12:19.503Z",
              "content": "<p>Also, my binary log loss is under 0.3, but still my validation score is not good</p>",
              "rawMarkdown": "Also, my binary log loss is under 0.3, but still my validation score is not good\n"
            },
            {
              "id": 2082966,
              "postDate": "2023-01-02T06:59:16.740Z",
              "content": "<p>Yes, I downsample negative candidates, and the binary log loss is depend on the positive/negative rate of the samples</p>",
              "rawMarkdown": "Yes, I downsample negative candidates, and the binary log loss is depend on the positive/negative rate of the samples"
            },
            {
              "id": 2086698,
              "postDate": "2023-01-05T00:17:22.320Z",
              "content": "<p>Ok ! I have one last question.<br>\nFor the candidates generation, did you use both train+test to generate the co-visitation matrix ? or did you generate one for each ? </p>",
              "rawMarkdown": "Ok ! I have one last question.\nFor the candidates generation, did you use both train+test to generate the co-visitation matrix ? or did you generate one for each ? "
            }
          ]
        }
      ]
    },
    {
      "id": 2077371,
      "postDate": "2022-12-27T14:19:33.020Z",
      "content": "<p>Good job, I have a question, did you use the co visitation matrices to make the features for the rank model?</p>",
      "rawMarkdown": "Good job, I have a question, did you use the co visitation matrices to make the features for the rank model?",
      "replies": [
        {
          "id": 2077397,
          "postDate": "2022-12-27T14:38:51.283Z",
          "content": "<p>Not yet, but I use the rank of co-visit candidates, it helps</p>",
          "rawMarkdown": "Not yet, but I use the rank of co-visit candidates, it helps",
          "replies": [
            {
              "id": 2082992,
              "postDate": "2023-01-02T07:16:30.090Z",
              "content": "<p>Thank you! Can I join your team? I have been confused by some problems for a long time, and I hope to get help from you.</p>",
              "rawMarkdown": "Thank you! Can I join your team? I have been confused by some problems for a long time, and I hope to get help from you."
            }
          ]
        }
      ]
    },
    {
      "id": 2076582,
      "postDate": "2022-12-26T16:05:55.443Z",
      "content": "<p>What the <code>objective</code> and <code>eval_metric</code> did you use for the ranking model? I found some people said the rank model's <code>eval_metric</code> doesn't correlate well with the recall20</p>",
      "rawMarkdown": "What the `objective` and `eval_metric` did you use for the ranking model? I found some people said the rank model's `eval_metric` doesn't correlate well with the recall20",
      "replies": [
        {
          "id": 2076952,
          "postDate": "2022-12-27T03:35:42.637Z",
          "content": "<p>I use binary classifier</p>",
          "rawMarkdown": "I use binary classifier",
          "replies": [
            {
              "id": 2077187,
              "postDate": "2022-12-27T10:20:34.247Z",
              "content": "<p>Thanks, do you willing to share the strategy to get high recall click candidates? I can only get recall = 0.584 for 200 candidates using the strategy in public notebooks.</p>",
              "rawMarkdown": "Thanks, do you willing to share the strategy to get high recall click candidates? I can only get recall = 0.584 for 200 candidates using the strategy in public notebooks."
            },
            {
              "id": 2077311,
              "postDate": "2022-12-27T12:55:47.333Z",
              "content": "<p>I also use the strategy in public notebooks, with some modification, the best public notebooks can get 0.576 in LB with just 20 candidates, I think 200+ candidates can easily get recall over 0.6+, maybe there are something wrong with your code?</p>",
              "rawMarkdown": "I also use the strategy in public notebooks, with some modification, the best public notebooks can get 0.576 in LB with just 20 candidates, I think 200+ candidates can easily get recall over 0.6+, maybe there are something wrong with your code?"
            },
            {
              "id": 2077459,
              "postDate": "2022-12-27T15:41:48.863Z",
              "content": "<p>I mean the recall for clicks. The recall scores of 200 candidates for each type are:</p>\n<pre><code>clicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\noverall recall = 0.62310\n</code></pre>",
              "rawMarkdown": "I mean the recall for clicks. The recall scores of 200 candidates for each type are:\n\n```\nclicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\noverall recall = 0.62310\n```"
            },
            {
              "id": 2077975,
              "postDate": "2022-12-28T02:08:14.190Z",
              "content": "<p>Thanks for your reply. After modifying some parts of the public notebook, my clicks recall@200 increased to 0.6155</p>",
              "rawMarkdown": "Thanks for your reply. After modifying some parts of the public notebook, my clicks recall@200 increased to 0.6155"
            }
          ]
        }
      ]
    },
    {
      "id": 2077292,
      "postDate": "2022-12-27T12:36:51.003Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2082884,
      "author_name": "wcq_glhf",
      "author_url": "",
      "post_date": "2023-01-02T04:20:45.380000",
      "content": "<p>i create 100+ feature, but the rank model doesn't perform well, the good feature i don't find, but some feature i think it's good run too slow,how many time do you generate features?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2077768,
      "author_name": "Rayan-aay",
      "author_url": "",
      "post_date": "2022-12-27T21:29:12.297000",
      "content": "<p>Did you build for each type ( clicks/carts/orders) one specific classifier for each ?</p>\n<p>Another question is about a classifier I trained ( lightgbm ) using only clicks from co-visitation matrix of clicks, the model converge, but when computing the recall, I got a worser one.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2078340,
          "author_name": "Lukan",
          "author_url": "",
          "post_date": "2022-12-28T08:58:42.343000",
          "content": "<ol>\n<li>Yes, I build tree ranking models</li>\n<li>Maybe you should make more features? If you make the right feature, the recall after ranking model should be better than before</li>\n</ol>",
          "votes": 0,
          "replies": [
            {
              "id": 2078371,
              "author_name": "Rayan-aay",
              "author_url": "",
              "post_date": "2022-12-28T09:30:43.263000",
              "content": "<p>Ok thanks.</p>\n<p>What about the AUC of your tree ? can you give us an estimation of the AUC at the  last iteration ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2078682,
              "author_name": "Lukan",
              "author_url": "",
              "post_date": "2022-12-28T14:32:37.303000",
              "content": "<p>Sorry, I hadn't use the AUC as my eval metric, while my ndcg@20 is around 0.88, binary log loss is about 0.3</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2080789,
              "author_name": "Rayan-aay",
              "author_url": "",
              "post_date": "2022-12-30T13:38:58.807000",
              "content": "<p>Awesome, did you downsample negative candidates ??</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082791,
              "author_name": "Rayan-aay",
              "author_url": "",
              "post_date": "2023-01-02T00:12:19.503000",
              "content": "<p>Also, my binary log loss is under 0.3, but still my validation score is not good</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082966,
              "author_name": "Lukan",
              "author_url": "",
              "post_date": "2023-01-02T06:59:16.740000",
              "content": "<p>Yes, I downsample negative candidates, and the binary log loss is depend on the positive/negative rate of the samples</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2086698,
              "author_name": "Rayan-aay",
              "author_url": "",
              "post_date": "2023-01-05T00:17:22.320000",
              "content": "<p>Ok ! I have one last question.<br>\nFor the candidates generation, did you use both train+test to generate the co-visitation matrix ? or did you generate one for each ? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2077371,
      "author_name": "Dong",
      "author_url": "",
      "post_date": "2022-12-27T14:19:33.020000",
      "content": "<p>Good job, I have a question, did you use the co visitation matrices to make the features for the rank model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2077397,
          "author_name": "Lukan",
          "author_url": "",
          "post_date": "2022-12-27T14:38:51.283000",
          "content": "<p>Not yet, but I use the rank of co-visit candidates, it helps</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2082992,
              "author_name": "Dong",
              "author_url": "",
              "post_date": "2023-01-02T07:16:30.090000",
              "content": "<p>Thank you! Can I join your team? I have been confused by some problems for a long time, and I hope to get help from you.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2076582,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2022-12-26T16:05:55.443000",
      "content": "<p>What the <code>objective</code> and <code>eval_metric</code> did you use for the ranking model? I found some people said the rank model's <code>eval_metric</code> doesn't correlate well with the recall20</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2076952,
          "author_name": "Lukan",
          "author_url": "",
          "post_date": "2022-12-27T03:35:42.637000",
          "content": "<p>I use binary classifier</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2077187,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2022-12-27T10:20:34.247000",
              "content": "<p>Thanks, do you willing to share the strategy to get high recall click candidates? I can only get recall = 0.584 for 200 candidates using the strategy in public notebooks.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2077311,
              "author_name": "Lukan",
              "author_url": "",
              "post_date": "2022-12-27T12:55:47.333000",
              "content": "<p>I also use the strategy in public notebooks, with some modification, the best public notebooks can get 0.576 in LB with just 20 candidates, I think 200+ candidates can easily get recall over 0.6+, maybe there are something wrong with your code?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2077459,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2022-12-27T15:41:48.863000",
              "content": "<p>I mean the recall for clicks. The recall scores of 200 candidates for each type are:</p>\n<pre><code>clicks recall = 0.58486\ncarts recall = 0.49270\norders recall = 0.69467\noverall recall = 0.62310\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2077975,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2022-12-28T02:08:14.190000",
              "content": "<p>Thanks for your reply. After modifying some parts of the public notebook, my clicks recall@200 increased to 0.6155</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2077292,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-27T12:36:51.003000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2076552": "In my experiment, I use past event, co-visit matrix to generate candidates, each session has about 52 candidates in average, and here is the CV result of my recall method:\n```python\nclicks recall = 0.6101419852876675\ncarts recall = 0.4724449332329544\norders recall = 0.6862270709185676\n=============\nOverall Recall = 0.6144839210497937\n=============\n```\nthen I create 60 features and bulid my ranking model, it got 0.581 in LB.\nWhat about your results?",
    "2082884": "i create 100+ feature, but the rank model doesn't perform well, the good feature i don't find, but some feature i think it's good run too slow,how many time do you generate features?",
    "2077768": "Did you build for each type ( clicks/carts/orders) one specific classifier for each ?\n\nAnother question is about a classifier I trained ( lightgbm ) using only clicks from co-visitation matrix of clicks, the model converge, but when computing the recall, I got a worser one.\n",
    "2077371": "Good job, I have a question, did you use the co visitation matrices to make the features for the rank model?",
    "2076582": "What the `objective` and `eval_metric` did you use for the ranking model? I found some people said the rank model's `eval_metric` doesn't correlate well with the recall20",
    "2077292": ""
  }
}