{
  "id": 380732,
  "title": "Similarities between users and candidates",
  "url": "/competitions/otto-recommender-system/discussion/380732",
  "author_name": "",
  "post_date": "2023-01-23T23:47:14.647743500Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello community, I computed similarity scores between users and candidates, using embeddings extracted from ALL the dataset ( included my validation B ).</p>\n<p>I wanted to know if it is a leak ? Because, in my orders recall , those variables extracted from ALL the dataset, boosted my local validation from 0.644 to 0.685.</p>\n<p>I'm asking this question because my LB with this Ranker is bad, it scored 0.565.<br>\nI was wondering if it was a problem of GBDT parameters such as :</p>\n<ul>\n<li>n estimators.</li>\n<li>L1 L2 Reg.</li>\n<li>Booster.</li>\n</ul>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "2112905",
      "postDate": "01/23/2023 23:47:14",
      "content": "<p>Hello community, I computed similarity scores between users and candidates, using embeddings extracted from ALL the dataset ( included my validation B ).</p>\n<p>I wanted to know if it is a leak ? Because, in my orders recall , those variables extracted from ALL the dataset, boosted my local validation from 0.644 to 0.685.</p>\n<p>I'm asking this question because my LB with this Ranker is bad, it scored 0.565.<br>\nI was wondering if it was a problem of GBDT parameters such as :</p>\n<ul>\n<li>n estimators.</li>\n<li>L1 L2 Reg.</li>\n<li>Booster.</li>\n</ul>\n<p>Thank you</p>",
      "rawMarkdown": "Hello community, I computed similarity scores between users and candidates, using embeddings extracted from ALL the dataset ( included my validation B ).\n\nI wanted to know if it is a leak ? Because, in my orders recall , those variables extracted from ALL the dataset, boosted my local validation from 0.644 to 0.685.\n\nI'm asking this question because my LB with this Ranker is bad, it scored 0.565.\nI was wondering if it was a problem of GBDT parameters such as :\n- n estimators.\n- L1 L2 Reg.\n- Booster.\n\n\n\nThank you",
      "votes": null
    },
    {
      "id": "2113095",
      "postDate": "01/24/2023 04:40:33",
      "content": "<p>Of course it's a leak. You have to create two embeddings for both validation and submission. Validation embeddings should include validation sessions and sessions before that. Submission embeddings can be trained with entire dataset though.</p>",
      "rawMarkdown": "Of course it's a leak. You have to create two embeddings for both validation and submission. Validation embeddings should include validation sessions and sessions before that. Submission embeddings can be trained with entire dataset though.",
      "votes": null
    },
    {
      "id": "2113774",
      "postDate": "01/24/2023 13:52:49",
      "content": "<p>Ok ! can't we use leakage of submission embeddings for validation ? </p>",
      "rawMarkdown": "Ok ! can't we use leakage of submission embeddings for validation ?",
      "votes": null
    },
    {
      "id": "2113799",
      "postDate": "01/24/2023 14:12:53",
      "content": "<p>You can use it but you will underestimate your generalization error.</p>",
      "rawMarkdown": "You can use it but you will underestimate your generalization error.",
      "votes": null
    },
    {
      "id": "2113807",
      "postDate": "01/24/2023 14:16:42",
      "content": "<p>Great ! thank you <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  for all this useful information !</p>",
      "rawMarkdown": "Great ! thank you @gunesevitan  for all this useful information !",
      "votes": null
    },
    {
      "id": "2114381",
      "postDate": "01/24/2023 23:58:56",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>   if you don't mind, can you tell me if you noticed a strong recall improvement when using a high learning rate and a small number of boosting iterations ? <br>\nIn fact ,I noticed a better recall within orders when using high learning rate ( 0.5 ) while having a better recall within carts with smaller ones ( 0.01 )<br>\nthank you</p>",
      "rawMarkdown": "gunesevitan   if you don't mind, can you tell me if you noticed a strong recall improvement when using a high learning rate and a small number of boosting iterations ? \nIn fact ,I noticed a better recall within orders when using high learning rate ( 0.5 ) while having a better recall within carts with smaller ones ( 0.01 )\nthank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2113095,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/24/2023 04:40:33",
      "content": "<p>Of course it's a leak. You have to create two embeddings for both validation and submission. Validation embeddings should include validation sessions and sessions before that. Submission embeddings can be trained with entire dataset though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2113774,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/24/2023 13:52:49",
          "content": "<p>Ok ! can't we use leakage of submission embeddings for validation ? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2113799,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "01/24/2023 14:12:53",
              "content": "<p>You can use it but you will underestimate your generalization error.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2113807,
                  "author_name": "rayanaay",
                  "author_url": "",
                  "post_date": "01/24/2023 14:16:42",
                  "content": "<p>Great ! thank you <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  for all this useful information !</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2114381,
                      "author_name": "rayanaay",
                      "author_url": "",
                      "post_date": "01/24/2023 23:58:56",
                      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>   if you don't mind, can you tell me if you noticed a strong recall improvement when using a high learning rate and a small number of boosting iterations ? <br>\nIn fact ,I noticed a better recall within orders when using high learning rate ( 0.5 ) while having a better recall within carts with smaller ones ( 0.01 )<br>\nthank you</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2112905": "Hello community, I computed similarity scores between users and candidates, using embeddings extracted from ALL the dataset ( included my validation B ).\n\nI wanted to know if it is a leak ? Because, in my orders recall , those variables extracted from ALL the dataset, boosted my local validation from 0.644 to 0.685.\n\nI'm asking this question because my LB with this Ranker is bad, it scored 0.565.\nI was wondering if it was a problem of GBDT parameters such as :\n- n estimators.\n- L1 L2 Reg.\n- Booster.\n\n\n\nThank you",
    "2113095": "Of course it's a leak. You have to create two embeddings for both validation and submission. Validation embeddings should include validation sessions and sessions before that. Submission embeddings can be trained with entire dataset though.",
    "2113774": "Ok ! can't we use leakage of submission embeddings for validation ?",
    "2113799": "You can use it but you will underestimate your generalization error.",
    "2113807": "Great ! thank you @gunesevitan  for all this useful information !",
    "2114381": "gunesevitan   if you don't mind, can you tell me if you noticed a strong recall improvement when using a high learning rate and a small number of boosting iterations ? \nIn fact ,I noticed a better recall within orders when using high learning rate ( 0.5 ) while having a better recall within carts with smaller ones ( 0.01 )\nthank you"
  },
  "source": "meta"
}