{
  "id": 617903,
  "title": "CV / LB gap - massive?",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/617903",
  "author_name": "",
  "post_date": "2025-11-12T15:08:40.652056700Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>CV is GroupKFold. Averaging across 5 folds.</p>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>CV (oF1)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>All \"authentic\"</td>\n<td>~50% (approx half images are authentic)</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>CNN1</td>\n<td>0.5348</td>\n<td>0.300</td>\n</tr>\n<tr>\n<td>CNN2</td>\n<td>0.704622</td>\n<td>0.301</td>\n</tr>\n<tr>\n<td>CNN3</td>\n<td>0.4301</td>\n<td>0.234</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "3320477",
      "postDate": "11/12/2025 15:08:40",
      "content": "<p>CV is GroupKFold. Averaging across 5 folds.</p>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>CV (oF1)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>All \"authentic\"</td>\n<td>~50% (approx half images are authentic)</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>CNN1</td>\n<td>0.5348</td>\n<td>0.300</td>\n</tr>\n<tr>\n<td>CNN2</td>\n<td>0.704622</td>\n<td>0.301</td>\n</tr>\n<tr>\n<td>CNN3</td>\n<td>0.4301</td>\n<td>0.234</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "CV is GroupKFold. Averaging across 5 folds.\n\n| Approach         | CV (oF1)  | LB    |\n|------------------|------------|-------|\n| All \"authentic\"  | ~50% (approx half images are authentic) | 0.303 |\n| CNN1             | 0.5348     | 0.300 |\n| CNN2             | 0.704622   | 0.301 |\n| CNN3             | 0.4301     | 0.234 |",
      "votes": null
    },
    {
      "id": "3320503",
      "postDate": "11/12/2025 15:31:25",
      "content": "<p>I think we could add a column to the LB score that solely represents the scores from successful detections, meaning it excludes the scores for \"authentic\" instances. I believe this might reveal some new information.Meanwhile, thank you for sharing, good luck.</p>",
      "rawMarkdown": "I think we could add a column to the LB score that solely represents the scores from successful detections, meaning it excludes the scores for \"authentic\" instances. I believe this might reveal some new information.Meanwhile, thank you for sharing, good luck.",
      "votes": null
    },
    {
      "id": "3320663",
      "postDate": "11/12/2025 17:39:23",
      "content": "<p>I created a nearly perfect <code>corn-detector</code>. No false positives in trainig data, masks are nice (not super accurate but better than nothing).\nAfter submitting it, I'm getting a worse score than <code>All authentic</code>. It may mean two facts:</p>\n<ol>\n<li>No corn-like images in LB</li>\n<li>My approach produces multiple false positives in LB, lowering the baseline scores. </li>\n</ol>\n<p>For me, this is the most challenging part, the images in LB may be way different from what was provided for training, so your approach should be somehow agnostic to the image, caring only about duplicates.</p>",
      "rawMarkdown": "I created a nearly perfect `corn-detector`. No false positives in trainig data, masks are nice (not super accurate but better than nothing).\nAfter submitting it, I'm getting a worse score than `All authentic`. It may mean two facts:\n1. No corn-like images in LB\n2. My approach produces multiple false positives in LB, lowering the baseline scores. \n\nFor me, this is the most challenging part, the images in LB may be way different from what was provided for training, so your approach should be somehow agnostic to the image, caring only about duplicates.",
      "votes": null
    },
    {
      "id": "3320677",
      "postDate": "11/12/2025 17:52:08",
      "content": "<p>Yes, your observation is correct. At least for the current LB (Leaderboard) scenario, among the 1,100 images in the test set, corn cobs might only account for a very small portion, or there might be none at all. Alternatively, the test set could contain many images that didn't appear in the training set, such as 45.png, which is a statistical chart-like image. \nHowever, I believe there are two points we need to pay attention to:</p>\n<ol>\n<li><p>The final test set will have twice the number of images as the current one, and we have absolutely no idea about its composition—whether it will include corn cobs or other types of images.</p></li>\n<li><p>We might need to focus more on creating a model that can handle various types of images, rather than concentrating solely on exploring the composition of the test set.</p></li>\n</ol>\n<p>Thank you for sharing, and good luck!</p>",
      "rawMarkdown": "Yes, your observation is correct. At least for the current LB (Leaderboard) scenario, among the 1,100 images in the test set, corn cobs might only account for a very small portion, or there might be none at all. Alternatively, the test set could contain many images that didn't appear in the training set, such as 45.png, which is a statistical chart-like image. \nHowever, I believe there are two points we need to pay attention to:\n\n1. The final test set will have twice the number of images as the current one, and we have absolutely no idea about its composition—whether it will include corn cobs or other types of images.\n\n2. We might need to focus more on creating a model that can handle various types of images, rather than concentrating solely on exploring the composition of the test set.\n\nThank you for sharing, and good luck!",
      "votes": null
    },
    {
      "id": "3345200",
      "postDate": "11/23/2025 09:49:06",
      "content": "<p>It may be some domain shift between train and LB. In this case, stain nomralization, agressive color augmentation and CutMix may help.</p>",
      "rawMarkdown": "It may be some domain shift between train and LB. In this case, stain nomralization, agressive color augmentation and CutMix may help.",
      "votes": null
    },
    {
      "id": "3346923",
      "postDate": "11/24/2025 18:23:33",
      "content": "<p>Are u sure that you are using the correct metrics for CV? I.e., are you including all masks per image when there are multiple?</p>",
      "rawMarkdown": "Are u sure that you are using the correct metrics for CV? I.e., are you including all masks per image when there are multiple?",
      "votes": null
    },
    {
      "id": "3381107",
      "postDate": "12/23/2025 18:38:07",
      "content": "<p>Have you tried creating a multi-panel image detector? I believe most images in the test set are multi-panel, so when we convert them to 224 or 512 even they get absolutely horrible resolution </p>",
      "rawMarkdown": "Have you tried creating a multi-panel image detector? I believe most images in the test set are multi-panel, so when we convert them to 224 or 512 even they get absolutely horrible resolution",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3320503,
      "author_name": "qifeihhh666",
      "author_url": "",
      "post_date": "11/12/2025 15:31:25",
      "content": "<p>I think we could add a column to the LB score that solely represents the scores from successful detections, meaning it excludes the scores for \"authentic\" instances. I believe this might reveal some new information.Meanwhile, thank you for sharing, good luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3320663,
      "author_name": "melgor",
      "author_url": "",
      "post_date": "11/12/2025 17:39:23",
      "content": "<p>I created a nearly perfect <code>corn-detector</code>. No false positives in trainig data, masks are nice (not super accurate but better than nothing).\nAfter submitting it, I'm getting a worse score than <code>All authentic</code>. It may mean two facts:</p>\n<ol>\n<li>No corn-like images in LB</li>\n<li>My approach produces multiple false positives in LB, lowering the baseline scores. </li>\n</ol>\n<p>For me, this is the most challenging part, the images in LB may be way different from what was provided for training, so your approach should be somehow agnostic to the image, caring only about duplicates.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3320677,
          "author_name": "qifeihhh666",
          "author_url": "",
          "post_date": "11/12/2025 17:52:08",
          "content": "<p>Yes, your observation is correct. At least for the current LB (Leaderboard) scenario, among the 1,100 images in the test set, corn cobs might only account for a very small portion, or there might be none at all. Alternatively, the test set could contain many images that didn't appear in the training set, such as 45.png, which is a statistical chart-like image. \nHowever, I believe there are two points we need to pay attention to:</p>\n<ol>\n<li><p>The final test set will have twice the number of images as the current one, and we have absolutely no idea about its composition—whether it will include corn cobs or other types of images.</p></li>\n<li><p>We might need to focus more on creating a model that can handle various types of images, rather than concentrating solely on exploring the composition of the test set.</p></li>\n</ol>\n<p>Thank you for sharing, and good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3381107,
          "author_name": "theodorlu",
          "author_url": "",
          "post_date": "12/23/2025 18:38:07",
          "content": "<p>Have you tried creating a multi-panel image detector? I believe most images in the test set are multi-panel, so when we convert them to 224 or 512 even they get absolutely horrible resolution </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3345200,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "11/23/2025 09:49:06",
      "content": "<p>It may be some domain shift between train and LB. In this case, stain nomralization, agressive color augmentation and CutMix may help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3346923,
      "author_name": "eliplutchok",
      "author_url": "",
      "post_date": "11/24/2025 18:23:33",
      "content": "<p>Are u sure that you are using the correct metrics for CV? I.e., are you including all masks per image when there are multiple?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3320477": "CV is GroupKFold. Averaging across 5 folds.\n\n| Approach         | CV (oF1)  | LB    |\n|------------------|------------|-------|\n| All \"authentic\"  | ~50% (approx half images are authentic) | 0.303 |\n| CNN1             | 0.5348     | 0.300 |\n| CNN2             | 0.704622   | 0.301 |\n| CNN3             | 0.4301     | 0.234 |",
    "3320503": "I think we could add a column to the LB score that solely represents the scores from successful detections, meaning it excludes the scores for \"authentic\" instances. I believe this might reveal some new information.Meanwhile, thank you for sharing, good luck.",
    "3320663": "I created a nearly perfect `corn-detector`. No false positives in trainig data, masks are nice (not super accurate but better than nothing).\nAfter submitting it, I'm getting a worse score than `All authentic`. It may mean two facts:\n1. No corn-like images in LB\n2. My approach produces multiple false positives in LB, lowering the baseline scores. \n\nFor me, this is the most challenging part, the images in LB may be way different from what was provided for training, so your approach should be somehow agnostic to the image, caring only about duplicates.",
    "3320677": "Yes, your observation is correct. At least for the current LB (Leaderboard) scenario, among the 1,100 images in the test set, corn cobs might only account for a very small portion, or there might be none at all. Alternatively, the test set could contain many images that didn't appear in the training set, such as 45.png, which is a statistical chart-like image. \nHowever, I believe there are two points we need to pay attention to:\n\n1. The final test set will have twice the number of images as the current one, and we have absolutely no idea about its composition—whether it will include corn cobs or other types of images.\n\n2. We might need to focus more on creating a model that can handle various types of images, rather than concentrating solely on exploring the composition of the test set.\n\nThank you for sharing, and good luck!",
    "3345200": "It may be some domain shift between train and LB. In this case, stain nomralization, agressive color augmentation and CutMix may help.",
    "3346923": "Are u sure that you are using the correct metrics for CV? I.e., are you including all masks per image when there are multiple?",
    "3381107": "Have you tried creating a multi-panel image detector? I believe most images in the test set are multi-panel, so when we convert them to 224 or 512 even they get absolutely horrible resolution"
  },
  "source": "meta"
}