{
  "id": 379177,
  "title": "Something goes wrong on my local cv",
  "url": "/competitions/otto-recommender-system/discussion/379177",
  "author_name": "",
  "post_date": "2023-01-18T14:29:48.338544400Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I use test.parquet as validA, test_label.parquet as validB, and use validB to set label for validA's candidate. After that I feed the data to LGBMRanker, and then use trained model to predict the validA(use the model to predict the valid part in GroupKFold, each candidate will be predict only once). </p>\n<p>Finally I compared it to test_label.parquet to get local cv. But it seems something wrong. My local cv is only 0.10 but my LB can get up to 0.50. I don't know which part is wrong, it doesn't seem to be a feature problem, nor should it be an overfitting problem.</p>",
  "messages": [
    {
      "id": "2105468",
      "postDate": "01/18/2023 14:29:48",
      "content": "<p>I use test.parquet as validA, test_label.parquet as validB, and use validB to set label for validA's candidate. After that I feed the data to LGBMRanker, and then use trained model to predict the validA(use the model to predict the valid part in GroupKFold, each candidate will be predict only once). </p>\n<p>Finally I compared it to test_label.parquet to get local cv. But it seems something wrong. My local cv is only 0.10 but my LB can get up to 0.50. I don't know which part is wrong, it doesn't seem to be a feature problem, nor should it be an overfitting problem.</p>",
      "rawMarkdown": "I use test.parquet as validA, test_label.parquet as validB, and use validB to set label for validA's candidate. After that I feed the data to LGBMRanker, and then use trained model to predict the validA(use the model to predict the valid part in GroupKFold, each candidate will be predict only once). \n\nFinally I compared it to test_label.parquet to get local cv. But it seems something wrong. My local cv is only 0.10 but my LB can get up to 0.50. I don't know which part is wrong, it doesn't seem to be a feature problem, nor should it be an overfitting problem.",
      "votes": null
    },
    {
      "id": "2107528",
      "postDate": "01/19/2023 20:52:08",
      "content": "<p>How many candidates did you generate using the co-visitation matrix ? </p>\n<p>For the features, you have to generate :</p>\n<p>Users Features from validA and test data<br>\nItems Features from (Train + ValidA) and ( all data ).<br>\nInteractions Features from validA and test data</p>",
      "rawMarkdown": "How many candidates did you generate using the co-visitation matrix ? \n\nFor the features, you have to generate :\n\nUsers Features from validA and test data\nItems Features from (Train + ValidA) and ( all data ).\nInteractions Features from validA and test data",
      "votes": null
    },
    {
      "id": "2108401",
      "postDate": "01/20/2023 14:00:30",
      "content": "<p>50 for each click/cart/order. And our features generation is the same as your mentioned. We don't know where the problem is.</p>",
      "rawMarkdown": "50 for each click/cart/order. And our features generation is the same as your mentioned. We don't know where the problem is.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2107528,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/19/2023 20:52:08",
      "content": "<p>How many candidates did you generate using the co-visitation matrix ? </p>\n<p>For the features, you have to generate :</p>\n<p>Users Features from validA and test data<br>\nItems Features from (Train + ValidA) and ( all data ).<br>\nInteractions Features from validA and test data</p>",
      "votes": null,
      "replies": [
        {
          "id": 2108401,
          "author_name": "kimoyami",
          "author_url": "",
          "post_date": "01/20/2023 14:00:30",
          "content": "<p>50 for each click/cart/order. And our features generation is the same as your mentioned. We don't know where the problem is.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2105468": "I use test.parquet as validA, test_label.parquet as validB, and use validB to set label for validA's candidate. After that I feed the data to LGBMRanker, and then use trained model to predict the validA(use the model to predict the valid part in GroupKFold, each candidate will be predict only once). \n\nFinally I compared it to test_label.parquet to get local cv. But it seems something wrong. My local cv is only 0.10 but my LB can get up to 0.50. I don't know which part is wrong, it doesn't seem to be a feature problem, nor should it be an overfitting problem.",
    "2107528": "How many candidates did you generate using the co-visitation matrix ? \n\nFor the features, you have to generate :\n\nUsers Features from validA and test data\nItems Features from (Train + ValidA) and ( all data ).\nInteractions Features from validA and test data",
    "2108401": "50 for each click/cart/order. And our features generation is the same as your mentioned. We don't know where the problem is."
  },
  "source": "meta"
}