{
  "id": 315238,
  "title": "Why is LB lower when the TOPK value is higher？",
  "url": "/competitions/happy-whale-and-dolphin/discussion/315238",
  "author_name": "",
  "post_date": "2022-03-27T03:37:46.462977800Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>When my ValTopK value is 0.54, my Lb is 0.746, but when my ValTopK reaches 0.59, LB is only 0.744. All configurations are B5 and fold0. </p>",
  "messages": [
    {
      "id": "1736150",
      "postDate": "03/27/2022 03:37:46",
      "content": "<p>When my ValTopK value is 0.54, my Lb is 0.746, but when my ValTopK reaches 0.59, LB is only 0.744. All configurations are B5 and fold0. </p>",
      "rawMarkdown": "When my ValTopK value is 0.54, my Lb is 0.746, but when my ValTopK reaches 0.59, LB is only 0.744. All configurations are B5 and fold0.",
      "votes": null
    },
    {
      "id": "1736151",
      "postDate": "03/27/2022 03:38:58",
      "content": "<p>B5 and fold0</p>",
      "rawMarkdown": "B5 and fold0",
      "votes": null
    },
    {
      "id": "1736992",
      "postDate": "03/28/2022 03:04:50",
      "content": "<p>ValTopK -- k=5? Seems your ValTopK is very different from LB. My top5 acc is above 60 and lb is also about 0.6.</p>",
      "rawMarkdown": "ValTopK -- k=5? Seems your ValTopK is very different from LB. My top5 acc is above 60 and lb is also about 0.6.",
      "votes": null
    },
    {
      "id": "1737108",
      "postDate": "03/28/2022 06:00:58",
      "content": "<p>My top5 reaches 0.60 and lb is about 0.76</p>",
      "rawMarkdown": "My top5 reaches 0.60 and lb is about 0.76",
      "votes": null
    },
    {
      "id": "1737133",
      "postDate": "03/28/2022 06:15:44",
      "content": "<p>I've seen this happen as well. If I understand your question correctly, your ValTopK is being computed for each epoch using arcface loss. Your validation set metric is based on cosine distance to an arcface embedding feature. However, if you are following the inference process used in most public kernels, your test set predictions are based on cosine similarity to every image in the training set. So you have multiple embeddings for each individual_id, not just the one representative embedding for each individual_id from the arcface margin product weights.</p>\n<p>So I think you're observing something that <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> (currently #2 on leaderboard) hinted at in a comment saying that managing overfitting was very important (<a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315129#1735584\" target=\"_blank\">comment here</a>). The model weights that maximize your submission approach may be more sensitive to overfitting and peak long before the ValTopK starts decreasing.</p>",
      "rawMarkdown": "I've seen this happen as well. If I understand your question correctly, your ValTopK is being computed for each epoch using arcface loss. Your validation set metric is based on cosine distance to an arcface embedding feature. However, if you are following the inference process used in most public kernels, your test set predictions are based on cosine similarity to every image in the training set. So you have multiple embeddings for each individual_id, not just the one representative embedding for each individual_id from the arcface margin product weights.\n\nSo I think you're observing something that @biglafe (currently #2 on leaderboard) hinted at in a comment saying that managing overfitting was very important ([comment here](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315129#1735584)). The model weights that maximize your submission approach may be more sensitive to overfitting and peak long before the ValTopK starts decreasing.",
      "votes": null
    },
    {
      "id": "1737155",
      "postDate": "03/28/2022 06:40:11",
      "content": "<p>Are you using PyTorch? Would you mind sharing some ideas for inference? Thanks.</p>",
      "rawMarkdown": "Are you using PyTorch? Would you mind sharing some ideas for inference? Thanks.",
      "votes": null
    },
    {
      "id": "1738052",
      "postDate": "03/29/2022 01:24:57",
      "content": "<p>Yes, I've been trying to solve the overfitting problem, but I'm a novice and have little experience.</p>",
      "rawMarkdown": "Yes, I've been trying to solve the overfitting problem, but I'm a novice and have little experience.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1736151,
      "author_name": "jianguotang",
      "author_url": "",
      "post_date": "03/27/2022 03:38:58",
      "content": "<p>B5 and fold0</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1736992,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "03/28/2022 03:04:50",
      "content": "<p>ValTopK -- k=5? Seems your ValTopK is very different from LB. My top5 acc is above 60 and lb is also about 0.6.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1737108,
          "author_name": "jianguotang",
          "author_url": "",
          "post_date": "03/28/2022 06:00:58",
          "content": "<p>My top5 reaches 0.60 and lb is about 0.76</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1737155,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "03/28/2022 06:40:11",
          "content": "<p>Are you using PyTorch? Would you mind sharing some ideas for inference? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1737133,
      "author_name": "rturley",
      "author_url": "",
      "post_date": "03/28/2022 06:15:44",
      "content": "<p>I've seen this happen as well. If I understand your question correctly, your ValTopK is being computed for each epoch using arcface loss. Your validation set metric is based on cosine distance to an arcface embedding feature. However, if you are following the inference process used in most public kernels, your test set predictions are based on cosine similarity to every image in the training set. So you have multiple embeddings for each individual_id, not just the one representative embedding for each individual_id from the arcface margin product weights.</p>\n<p>So I think you're observing something that <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> (currently #2 on leaderboard) hinted at in a comment saying that managing overfitting was very important (<a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315129#1735584\" target=\"_blank\">comment here</a>). The model weights that maximize your submission approach may be more sensitive to overfitting and peak long before the ValTopK starts decreasing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1738052,
          "author_name": "jianguotang",
          "author_url": "",
          "post_date": "03/29/2022 01:24:57",
          "content": "<p>Yes, I've been trying to solve the overfitting problem, but I'm a novice and have little experience.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1736150": "When my ValTopK value is 0.54, my Lb is 0.746, but when my ValTopK reaches 0.59, LB is only 0.744. All configurations are B5 and fold0.",
    "1736151": "B5 and fold0",
    "1736992": "ValTopK -- k=5? Seems your ValTopK is very different from LB. My top5 acc is above 60 and lb is also about 0.6.",
    "1737108": "My top5 reaches 0.60 and lb is about 0.76",
    "1737133": "I've seen this happen as well. If I understand your question correctly, your ValTopK is being computed for each epoch using arcface loss. Your validation set metric is based on cosine distance to an arcface embedding feature. However, if you are following the inference process used in most public kernels, your test set predictions are based on cosine similarity to every image in the training set. So you have multiple embeddings for each individual_id, not just the one representative embedding for each individual_id from the arcface margin product weights.\n\nSo I think you're observing something that @biglafe (currently #2 on leaderboard) hinted at in a comment saying that managing overfitting was very important ([comment here](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315129#1735584)). The model weights that maximize your submission approach may be more sensitive to overfitting and peak long before the ValTopK starts decreasing.",
    "1737155": "Are you using PyTorch? Would you mind sharing some ideas for inference? Thanks.",
    "1738052": "Yes, I've been trying to solve the overfitting problem, but I'm a novice and have little experience."
  },
  "source": "meta"
}