{
  "id": 560101,
  "title": "Significant differences in LB",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/560101",
  "author_name": "",
  "post_date": "2025-01-29T11:51:25.727772300Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We are at the end of the competition.<br>\nCV and LB do not correlate well and I am not sure about the choice of files to submit.<br>\nHow much difference in LB scores would be considered a significant difference?　Should we trust CV or LB?<br>\nI would like to hear your opinion.</p>",
  "messages": [
    {
      "id": "3109875",
      "postDate": "01/29/2025 11:51:25",
      "content": "<p>We are at the end of the competition.<br>\nCV and LB do not correlate well and I am not sure about the choice of files to submit.<br>\nHow much difference in LB scores would be considered a significant difference?　Should we trust CV or LB?<br>\nI would like to hear your opinion.</p>",
      "rawMarkdown": "We are at the end of the competition.\nCV and LB do not correlate well and I am not sure about the choice of files to submit.\nHow much difference in LB scores would be considered a significant difference?　Should we trust CV or LB?\nI would like to hear your opinion.",
      "votes": null
    },
    {
      "id": "3109886",
      "postDate": "01/29/2025 12:10:44",
      "content": "<p>For me CV and LB correlate fairly well within a series of models. What I mean by that is if I have two models with the same or similar architecture, I can typically rank (not necessarily predict) how they will do on the LB. </p>\n<p>If using a connected components scheme, plot the F-Beta vs a range of thresholds. The flatter those curves, the better the correlation has been for me. When your models are very sensitive to thresholds, there will naturally be more variation. I think over/under fitting of thresholds could be a real issue with LB correlations and I think my models probably have some of this.</p>\n<p>As for CV vs LB, there are 500 samples in the LB set and the public leaderboard has 26% of this data. If we <em>assume</em> this was split by sample, the public LB is validation on 130 samples. We only have 7 to train with. So I use CV to guide my experiments and trust LB for final models.</p>",
      "rawMarkdown": "For me CV and LB correlate fairly well within a series of models. What I mean by that is if I have two models with the same or similar architecture, I can typically rank (not necessarily predict) how they will do on the LB. \n\nIf using a connected components scheme, plot the F-Beta vs a range of thresholds. The flatter those curves, the better the correlation has been for me. When your models are very sensitive to thresholds, there will naturally be more variation. I think over/under fitting of thresholds could be a real issue with LB correlations and I think my models probably have some of this.\n\nAs for CV vs LB, there are 500 samples in the LB set and the public leaderboard has 26% of this data. If we *assume* this was split by sample, the public LB is validation on 130 samples. We only have 7 to train with. So I use CV to guide my experiments and trust LB for final models.",
      "votes": null
    },
    {
      "id": "3109893",
      "postDate": "01/29/2025 12:19:36",
      "content": "<p>Thanks for your great reply, the idea of plotting the F-Beta score by threshold is new to me. I will try it as soon as possible!</p>",
      "rawMarkdown": "Thanks for your great reply, the idea of plotting the F-Beta score by threshold is new to me. I will try it as soon as possible!",
      "votes": null
    },
    {
      "id": "3109899",
      "postDate": "01/29/2025 12:22:02",
      "content": "<p>It is very helpful :) good luck 🤞</p>",
      "rawMarkdown": "It is very helpful :) good luck 🤞",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3109886,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "01/29/2025 12:10:44",
      "content": "<p>For me CV and LB correlate fairly well within a series of models. What I mean by that is if I have two models with the same or similar architecture, I can typically rank (not necessarily predict) how they will do on the LB. </p>\n<p>If using a connected components scheme, plot the F-Beta vs a range of thresholds. The flatter those curves, the better the correlation has been for me. When your models are very sensitive to thresholds, there will naturally be more variation. I think over/under fitting of thresholds could be a real issue with LB correlations and I think my models probably have some of this.</p>\n<p>As for CV vs LB, there are 500 samples in the LB set and the public leaderboard has 26% of this data. If we <em>assume</em> this was split by sample, the public LB is validation on 130 samples. We only have 7 to train with. So I use CV to guide my experiments and trust LB for final models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3109893,
          "author_name": "yoshinarikawashima",
          "author_url": "",
          "post_date": "01/29/2025 12:19:36",
          "content": "<p>Thanks for your great reply, the idea of plotting the F-Beta score by threshold is new to me. I will try it as soon as possible!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3109899,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "01/29/2025 12:22:02",
              "content": "<p>It is very helpful :) good luck 🤞</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3109875": "We are at the end of the competition.\nCV and LB do not correlate well and I am not sure about the choice of files to submit.\nHow much difference in LB scores would be considered a significant difference?　Should we trust CV or LB?\nI would like to hear your opinion.",
    "3109886": "For me CV and LB correlate fairly well within a series of models. What I mean by that is if I have two models with the same or similar architecture, I can typically rank (not necessarily predict) how they will do on the LB. \n\nIf using a connected components scheme, plot the F-Beta vs a range of thresholds. The flatter those curves, the better the correlation has been for me. When your models are very sensitive to thresholds, there will naturally be more variation. I think over/under fitting of thresholds could be a real issue with LB correlations and I think my models probably have some of this.\n\nAs for CV vs LB, there are 500 samples in the LB set and the public leaderboard has 26% of this data. If we *assume* this was split by sample, the public LB is validation on 130 samples. We only have 7 to train with. So I use CV to guide my experiments and trust LB for final models.",
    "3109893": "Thanks for your great reply, the idea of plotting the F-Beta score by threshold is new to me. I will try it as soon as possible!",
    "3109899": "It is very helpful :) good luck 🤞"
  },
  "source": "meta"
}