{
  "id": 368732,
  "title": "How do you train Ranking model?",
  "url": "/competitions/otto-recommender-system/discussion/368732",
  "author_name": "",
  "post_date": "2022-11-27T12:13:49.563888600Z",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm trying to build ranking model using following methods, but it's not working well.</p>\n<blockquote>\n  <p><strong>Dataset  &amp; CV Strategy</strong></p>\n  <ul>\n  <li>Use 2w and 3w labels as training data.<br>\n  (22-08-07 ~ 22-08-21)</li>\n  <li>Use 4w labels as validation data.<br>\n  (22-08-21 ~ 22-08-28)</li>\n  <li>Training with 10% of total session.</li>\n  <li>Create candidates and features by transaction data prior to the week of the label.</li>\n  </ul>\n  <p><strong>Candidate</strong></p>\n  <ul>\n  <li>Past event candidates</li>\n  <li>Co-visitation matrix candidates</li>\n  <li>popular item candidates</li>\n  </ul>\n  <p><strong>Features</strong></p>\n  <ul>\n  <li>aid_features, session_features, session×aid_features  <br>\n  I refer to this <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893\" target=\"_blank\">thread</a>. </li>\n  </ul>\n  <p><strong>Model</strong></p>\n  <ul>\n  <li>LGBM Classifier（training models for each event）</li>\n  </ul>\n</blockquote>\n<p>How do you build models and datasets for training?<br>\nIn particular, I would like to know about the following two points</p>\n<ol>\n<li>Which period of data is used for the training labels?</li>\n<li>Do you do any sampling for training?</li>\n</ol>\n<p>I'd be glad to know.</p>",
  "messages": [
    {
      "id": "2045520",
      "postDate": "11/27/2022 12:13:49",
      "content": "<p>I'm trying to build ranking model using following methods, but it's not working well.</p>\n<blockquote>\n  <p><strong>Dataset  &amp; CV Strategy</strong></p>\n  <ul>\n  <li>Use 2w and 3w labels as training data.<br>\n  (22-08-07 ~ 22-08-21)</li>\n  <li>Use 4w labels as validation data.<br>\n  (22-08-21 ~ 22-08-28)</li>\n  <li>Training with 10% of total session.</li>\n  <li>Create candidates and features by transaction data prior to the week of the label.</li>\n  </ul>\n  <p><strong>Candidate</strong></p>\n  <ul>\n  <li>Past event candidates</li>\n  <li>Co-visitation matrix candidates</li>\n  <li>popular item candidates</li>\n  </ul>\n  <p><strong>Features</strong></p>\n  <ul>\n  <li>aid_features, session_features, session×aid_features  <br>\n  I refer to this <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893\" target=\"_blank\">thread</a>. </li>\n  </ul>\n  <p><strong>Model</strong></p>\n  <ul>\n  <li>LGBM Classifier（training models for each event）</li>\n  </ul>\n</blockquote>\n<p>How do you build models and datasets for training?<br>\nIn particular, I would like to know about the following two points</p>\n<ol>\n<li>Which period of data is used for the training labels?</li>\n<li>Do you do any sampling for training?</li>\n</ol>\n<p>I'd be glad to know.</p>",
      "rawMarkdown": "I'm trying to build ranking model using following methods, but it's not working well.\n\n> **Dataset  & CV Strategy**\n- Use 2w and 3w labels as training data.\n(22-08-07 ~ 22-08-21)\n- Use 4w labels as validation data.\n(22-08-21 ~ 22-08-28)\n- Training with 10% of total session.\n- Create candidates and features by transaction data prior to the week of the label.\n\n> **Candidate**\n- Past event candidates\n- Co-visitation matrix candidates\n- popular item candidates\n\n> **Features**\n- aid_features, session_features, session×aid_features  \nI refer to this [thread](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893). \n\n> **Model**\n- LGBM Classifier（training models for each event）\n\nHow do you build models and datasets for training?\nIn particular, I would like to know about the following two points\n\n1. Which period of data is used for the training labels?\n2. Do you do any sampling for training?\n\nI'd be glad to know.",
      "votes": null
    },
    {
      "id": "2045622",
      "postDate": "11/27/2022 14:14:40",
      "content": "<p>How many candidates do you generate for each user? What are your hit rates for clicks,carts and orders. I think you maybe  have a problem for candidate generation. I use word2vector,bpr,itemcf for candidate generation except past event candidates which all have low hit rates.</p>",
      "rawMarkdown": "How many candidates do you generate for each user? What are your hit rates for clicks,carts and orders. I think you maybe  have a problem for candidate generation. I use word2vector,bpr,itemcf for candidate generation except past event candidates which all have low hit rates.",
      "votes": null
    },
    {
      "id": "2045625",
      "postDate": "11/27/2022 14:19:05",
      "content": "<p>i use the approach shared by radek in his notebook.</p>",
      "rawMarkdown": "i use the approach shared by radek in his notebook.",
      "votes": null
    },
    {
      "id": "2045960",
      "postDate": "11/27/2022 20:11:37",
      "content": "<p>Thanks. I'm sure there will be some useful information in comment thread. I will take a look.</p>",
      "rawMarkdown": "Thanks. I'm sure there will be some useful information in comment thread. I will take a look.",
      "votes": null
    },
    {
      "id": "2045966",
      "postDate": "11/27/2022 20:23:40",
      "content": "<p>Also the associated threads point to useful techniques and Tricks.</p>",
      "rawMarkdown": "Also the associated threads point to useful techniques and Tricks.",
      "votes": null
    },
    {
      "id": "2045968",
      "postDate": "11/27/2022 20:25:49",
      "content": "<p>Data for 4w (22-08-21 ~ 22-08-28) are as follows.</p>\n<blockquote>\n  <p>session count: 685.496（Randomly sampled 10%.）<br>\n  candidates count mean: 96.6<br>\n  click hit rate: 0.005<br>\n  cart hit rate: 0.0015<br>\n  order hit rate: 0.0012</p>\n</blockquote>\n<p>Is it lower than your hit rate?</p>",
      "rawMarkdown": "Data for 4w (22-08-21 ~ 22-08-28) are as follows.\n\n> session count: 685.496（Randomly sampled 10%.）\ncandidates count mean: 96.6\nclick hit rate: 0.005\ncart hit rate: 0.0015\norder hit rate: 0.0012\n\nIs it lower than your hit rate?",
      "votes": null
    },
    {
      "id": "2045983",
      "postDate": "11/27/2022 20:42:49",
      "content": "<p>Could you report your recall@20 per clicks/carts/orders</p>",
      "rawMarkdown": "Could you report your recall@20 per clicks/carts/orders",
      "votes": null
    },
    {
      "id": "2045992",
      "postDate": "11/27/2022 21:01:42",
      "content": "<p>CV scores of  4w(22-08-21 ~ 22-08-28)  are as follows.<br>\nHowever, LB score predicted by this model was 0.562</p>\n<p>The LB/CV correlation is not good, so I suspect there is a problem with the training data.</p>\n<hr>\n<p>clicks: 0.5091<br>\ncarts: 0.4105<br>\norders: 0.6499</p>\n<p>Score: 0.5640</p>",
      "rawMarkdown": "CV scores of  4w(22-08-21 ~ 22-08-28)  are as follows.\nHowever, LB score predicted by this model was 0.562\n\nThe LB/CV correlation is not good, so I suspect there is a problem with the training data.\n\n-------------------\nclicks: 0.5091\ncarts: 0.4105\norders: 0.6499\n\nScore: 0.5640",
      "votes": null
    },
    {
      "id": "2046019",
      "postDate": "11/27/2022 21:17:49",
      "content": "<p>For now I have worse CV but I guess this is mostly due to the fact that I can just take top10 candiates from Co visitation matrix as I get memory error otherwise. May I ask how you incoporate canditates from Co visitation. For now I generate and than concat the resulting dataframe with canditates and afterwards count how many times each item was suggested by co visitation matrix. I guess this is not smart way and I could largely improve CV if I can take more canditates.. may I ask how you do that step? Maybe I should stop the counting of suggestion as that is where I get memory error. I will try.<br>\nIf I get Fürther improvement in CV I will report here. There is many features one can think of i didnt include so far, so there is long way to go.<br>\nMy current CV:<br>\nclicks recall = 0.5020045182833257<br>\ncarts recall = 0.4034401767965001<br>\norders recall = 0.6254552302403743</p>\n<p>Overall Recall = 0.5465056430115072</p>",
      "rawMarkdown": "For now I have worse CV but I guess this is mostly due to the fact that I can just take top10 candiates from Co visitation matrix as I get memory error otherwise. May I ask how you incoporate canditates from Co visitation. For now I generate and than concat the resulting dataframe with canditates and afterwards count how many times each item was suggested by co visitation matrix. I guess this is not smart way and I could largely improve CV if I can take more canditates.. may I ask how you do that step? Maybe I should stop the counting of suggestion as that is where I get memory error. I will try.\nIf I get Fürther improvement in CV I will report here. There is many features one can think of i didnt include so far, so there is long way to go.\nMy current CV:\nclicks recall = 0.5020045182833257\ncarts recall = 0.4034401767965001\norders recall = 0.6254552302403743\n\nOverall Recall = 0.5465056430115072",
      "votes": null
    },
    {
      "id": "2046235",
      "postDate": "11/28/2022 03:23:56",
      "content": "<p><a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> <br>\nI use each of the Top50 candidates from 4 different Co_visitation_matrixes as candidates.Since some of the matrix candidates are mixed with past aids, I remove duplicates at the end of the candidate generation process.<br>\nIn addition, sessions are sampled as a precautionary measure against memory errors.（10%）</p>\n<p>btw, what is the LB score for the model with CV 0.5465?</p>",
      "rawMarkdown": "simonveitner \nI use each of the Top50 candidates from 4 different Co_visitation_matrixes as candidates.Since some of the matrix candidates are mixed with past aids, I remove duplicates at the end of the candidate generation process.\nIn addition, sessions are sampled as a precautionary measure against memory errors.（10%）\n\nbtw, what is the LB score for the model with CV 0.5465?",
      "votes": null
    },
    {
      "id": "2046396",
      "postDate": "11/28/2022 07:03:11",
      "content": "<p>I didn't submit this model so far to LB. Will do so later and report on leadeboard. Yea, the dublicates I also remove before giving to my model. For now I only use CLICKS matrix for the clicks ranker and only BUY matrix for the buy ranker. I need to include more and will try sampling session as an approach. Tnx! I will let you know abou4 New CV</p>",
      "rawMarkdown": "I didn't submit this model so far to LB. Will do so later and report on leadeboard. Yea, the dublicates I also remove before giving to my model. For now I only use CLICKS matrix for the clicks ranker and only BUY matrix for the buy ranker. I need to include more and will try sampling session as an approach. Tnx! I will let you know abou4 New CV",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2045622,
      "author_name": "jiahongxie",
      "author_url": "",
      "post_date": "11/27/2022 14:14:40",
      "content": "<p>How many candidates do you generate for each user? What are your hit rates for clicks,carts and orders. I think you maybe  have a problem for candidate generation. I use word2vector,bpr,itemcf for candidate generation except past event candidates which all have low hit rates.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2045968,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "11/27/2022 20:25:49",
          "content": "<p>Data for 4w (22-08-21 ~ 22-08-28) are as follows.</p>\n<blockquote>\n  <p>session count: 685.496（Randomly sampled 10%.）<br>\n  candidates count mean: 96.6<br>\n  click hit rate: 0.005<br>\n  cart hit rate: 0.0015<br>\n  order hit rate: 0.0012</p>\n</blockquote>\n<p>Is it lower than your hit rate?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2045983,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "11/27/2022 20:42:49",
          "content": "<p>Could you report your recall@20 per clicks/carts/orders</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2045992,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "11/27/2022 21:01:42",
          "content": "<p>CV scores of  4w(22-08-21 ~ 22-08-28)  are as follows.<br>\nHowever, LB score predicted by this model was 0.562</p>\n<p>The LB/CV correlation is not good, so I suspect there is a problem with the training data.</p>\n<hr>\n<p>clicks: 0.5091<br>\ncarts: 0.4105<br>\norders: 0.6499</p>\n<p>Score: 0.5640</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046019,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "11/27/2022 21:17:49",
          "content": "<p>For now I have worse CV but I guess this is mostly due to the fact that I can just take top10 candiates from Co visitation matrix as I get memory error otherwise. May I ask how you incoporate canditates from Co visitation. For now I generate and than concat the resulting dataframe with canditates and afterwards count how many times each item was suggested by co visitation matrix. I guess this is not smart way and I could largely improve CV if I can take more canditates.. may I ask how you do that step? Maybe I should stop the counting of suggestion as that is where I get memory error. I will try.<br>\nIf I get Fürther improvement in CV I will report here. There is many features one can think of i didnt include so far, so there is long way to go.<br>\nMy current CV:<br>\nclicks recall = 0.5020045182833257<br>\ncarts recall = 0.4034401767965001<br>\norders recall = 0.6254552302403743</p>\n<p>Overall Recall = 0.5465056430115072</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046235,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "11/28/2022 03:23:56",
          "content": "<p><a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> <br>\nI use each of the Top50 candidates from 4 different Co_visitation_matrixes as candidates.Since some of the matrix candidates are mixed with past aids, I remove duplicates at the end of the candidate generation process.<br>\nIn addition, sessions are sampled as a precautionary measure against memory errors.（10%）</p>\n<p>btw, what is the LB score for the model with CV 0.5465?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046396,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "11/28/2022 07:03:11",
          "content": "<p>I didn't submit this model so far to LB. Will do so later and report on leadeboard. Yea, the dublicates I also remove before giving to my model. For now I only use CLICKS matrix for the clicks ranker and only BUY matrix for the buy ranker. I need to include more and will try sampling session as an approach. Tnx! I will let you know abou4 New CV</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2045625,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "11/27/2022 14:19:05",
      "content": "<p>i use the approach shared by radek in his notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2045960,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "11/27/2022 20:11:37",
          "content": "<p>Thanks. I'm sure there will be some useful information in comment thread. I will take a look.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2045966,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "11/27/2022 20:23:40",
          "content": "<p>Also the associated threads point to useful techniques and Tricks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2045520": "I'm trying to build ranking model using following methods, but it's not working well.\n\n> **Dataset  & CV Strategy**\n- Use 2w and 3w labels as training data.\n(22-08-07 ~ 22-08-21)\n- Use 4w labels as validation data.\n(22-08-21 ~ 22-08-28)\n- Training with 10% of total session.\n- Create candidates and features by transaction data prior to the week of the label.\n\n> **Candidate**\n- Past event candidates\n- Co-visitation matrix candidates\n- popular item candidates\n\n> **Features**\n- aid_features, session_features, session×aid_features  \nI refer to this [thread](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2030893). \n\n> **Model**\n- LGBM Classifier（training models for each event）\n\nHow do you build models and datasets for training?\nIn particular, I would like to know about the following two points\n\n1. Which period of data is used for the training labels?\n2. Do you do any sampling for training?\n\nI'd be glad to know.",
    "2045622": "How many candidates do you generate for each user? What are your hit rates for clicks,carts and orders. I think you maybe  have a problem for candidate generation. I use word2vector,bpr,itemcf for candidate generation except past event candidates which all have low hit rates.",
    "2045625": "i use the approach shared by radek in his notebook.",
    "2045960": "Thanks. I'm sure there will be some useful information in comment thread. I will take a look.",
    "2045966": "Also the associated threads point to useful techniques and Tricks.",
    "2045968": "Data for 4w (22-08-21 ~ 22-08-28) are as follows.\n\n> session count: 685.496（Randomly sampled 10%.）\ncandidates count mean: 96.6\nclick hit rate: 0.005\ncart hit rate: 0.0015\norder hit rate: 0.0012\n\nIs it lower than your hit rate?",
    "2045983": "Could you report your recall@20 per clicks/carts/orders",
    "2045992": "CV scores of  4w(22-08-21 ~ 22-08-28)  are as follows.\nHowever, LB score predicted by this model was 0.562\n\nThe LB/CV correlation is not good, so I suspect there is a problem with the training data.\n\n-------------------\nclicks: 0.5091\ncarts: 0.4105\norders: 0.6499\n\nScore: 0.5640",
    "2046019": "For now I have worse CV but I guess this is mostly due to the fact that I can just take top10 candiates from Co visitation matrix as I get memory error otherwise. May I ask how you incoporate canditates from Co visitation. For now I generate and than concat the resulting dataframe with canditates and afterwards count how many times each item was suggested by co visitation matrix. I guess this is not smart way and I could largely improve CV if I can take more canditates.. may I ask how you do that step? Maybe I should stop the counting of suggestion as that is where I get memory error. I will try.\nIf I get Fürther improvement in CV I will report here. There is many features one can think of i didnt include so far, so there is long way to go.\nMy current CV:\nclicks recall = 0.5020045182833257\ncarts recall = 0.4034401767965001\norders recall = 0.6254552302403743\n\nOverall Recall = 0.5465056430115072",
    "2046235": "simonveitner \nI use each of the Top50 candidates from 4 different Co_visitation_matrixes as candidates.Since some of the matrix candidates are mixed with past aids, I remove duplicates at the end of the candidate generation process.\nIn addition, sessions are sampled as a precautionary measure against memory errors.（10%）\n\nbtw, what is the LB score for the model with CV 0.5465?",
    "2046396": "I didn't submit this model so far to LB. Will do so later and report on leadeboard. Yea, the dublicates I also remove before giving to my model. For now I only use CLICKS matrix for the clicks ranker and only BUY matrix for the buy ranker. I need to include more and will try sampling session as an approach. Tnx! I will let you know abou4 New CV"
  },
  "source": "meta"
}