{
  "id": 658623,
  "title": "Surprising Gap between Local Validation Score (0.722) and in Public LB (0.103)",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/658623",
  "author_name": "",
  "post_date": "2025-12-11T13:14:09.310929100Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>For validation, I used the exact score function provided in the competition, so I expected the numbers to be much closer. I am seeing a big difference between <strong>my local validation score (~0.72) and the public leaderboard score (~0.10)</strong>. I have checked the common mistakes like mask size, RLE format, and thresholding, and they seem fine.</p>\n<p>So I’m wondering if anyone else has faced a similar drop. </p>\n<p>Thanks!!</p>",
  "messages": [
    {
      "id": "3371414",
      "postDate": "12/11/2025 13:14:09",
      "content": "<p>Hi everyone,</p>\n<p>For validation, I used the exact score function provided in the competition, so I expected the numbers to be much closer. I am seeing a big difference between <strong>my local validation score (~0.72) and the public leaderboard score (~0.10)</strong>. I have checked the common mistakes like mask size, RLE format, and thresholding, and they seem fine.</p>\n<p>So I’m wondering if anyone else has faced a similar drop. </p>\n<p>Thanks!!</p>",
      "rawMarkdown": "Hi everyone,\n\nFor validation, I used the exact score function provided in the competition, so I expected the numbers to be much closer. I am seeing a big difference between **my local validation score (~0.72) and the public leaderboard score (~0.10)**. I have checked the common mistakes like mask size, RLE format, and thresholding, and they seem fine.\n\nSo I’m wondering if anyone else has faced a similar drop. \n\nThanks!!",
      "votes": null
    },
    {
      "id": "3372012",
      "postDate": "12/11/2025 18:55:09",
      "content": "<p>Make sure you are groupkfold on the image. You dont want to train on forged imageA.png and then validate on authentic imageA.png. It may cause leak. 5fold model averaged should be able to break 0.3 on LB</p>",
      "rawMarkdown": "Make sure you are groupkfold on the image. You dont want to train on forged imageA.png and then validate on authentic imageA.png. It may cause leak. 5fold model averaged should be able to break 0.3 on LB",
      "votes": null
    },
    {
      "id": "3372601",
      "postDate": "12/12/2025 05:23:56",
      "content": "<p>yes this challenge has this issue, i think this is because of a domain shift in the training data that they have given us and the testing data that they are using for testing our models on the private dataset</p>",
      "rawMarkdown": "yes this challenge has this issue, i think this is because of a domain shift in the training data that they have given us and the testing data that they are using for testing our models on the private dataset",
      "votes": null
    },
    {
      "id": "3376888",
      "postDate": "12/15/2025 11:09:33",
      "content": "<p>Doesn't seem to help. I'm using case_id as group but still get scores that are too high (<a href=\"https://www.kaggle.com/code/ravaghi/scientific-image-forgery-detection-u-net-1-2\" target=\"_blank\">~0.6</a>).</p>",
      "rawMarkdown": "Doesn't seem to help. I'm using case_id as group but still get scores that are too high ([~0.6](https://www.kaggle.com/code/ravaghi/scientific-image-forgery-detection-u-net-1-2)).",
      "votes": null
    },
    {
      "id": "3376901",
      "postDate": "12/15/2025 11:36:15",
      "content": "<p>Your own notebook (Dinov2) breaks 0.3 with a single model so it's definitely possible. Some of my trained models also broke 0.3 but never scored as high as yours. My current score uses your model as a baseline then just applies postprocessing.</p>",
      "rawMarkdown": "Your own notebook (Dinov2) breaks 0.3 with a single model so it's definitely possible. Some of my trained models also broke 0.3 but never scored as high as yours. My current score uses your model as a baseline then just applies postprocessing.",
      "votes": null
    },
    {
      "id": "3376903",
      "postDate": "12/15/2025 11:44:07",
      "content": "<p>I haven't done cross-validation with DINOv2 yet, so I don't know the CV score for that model. I was actually referring to my public U-Net model, which scores 0.16 on the public LB and has a CV score of ~0.6. I plan to train DINOv2 with cross-validation as well, but first I need to fix the leakage issues I'm currently facing.</p>",
      "rawMarkdown": "I haven't done cross-validation with DINOv2 yet, so I don't know the CV score for that model. I was actually referring to my public U-Net model, which scores 0.16 on the public LB and has a CV score of ~0.6. I plan to train DINOv2 with cross-validation as well, but first I need to fix the leakage issues I'm currently facing.",
      "votes": null
    },
    {
      "id": "3376910",
      "postDate": "12/15/2025 12:00:06",
      "content": "<p>GroupKFold on case_id with a pretrained UNet, average across 5 models, and tune the mask thresholds according to local CV has yielded me 0.30X to 0.31X. Average across 5 folds brought LB from 0.26X to 0.3, so I think that part is quite important.</p>",
      "rawMarkdown": "GroupKFold on case_id with a pretrained UNet, average across 5 models, and tune the mask thresholds according to local CV has yielded me 0.30X to 0.31X. Average across 5 folds brought LB from 0.26X to 0.3, so I think that part is quite important.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3372012,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "12/11/2025 18:55:09",
      "content": "<p>Make sure you are groupkfold on the image. You dont want to train on forged imageA.png and then validate on authentic imageA.png. It may cause leak. 5fold model averaged should be able to break 0.3 on LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 3376888,
          "author_name": "ravaghi",
          "author_url": "",
          "post_date": "12/15/2025 11:09:33",
          "content": "<p>Doesn't seem to help. I'm using case_id as group but still get scores that are too high (<a href=\"https://www.kaggle.com/code/ravaghi/scientific-image-forgery-detection-u-net-1-2\" target=\"_blank\">~0.6</a>).</p>",
          "votes": null,
          "replies": [
            {
              "id": 3376901,
              "author_name": "returnofsputnik",
              "author_url": "",
              "post_date": "12/15/2025 11:36:15",
              "content": "<p>Your own notebook (Dinov2) breaks 0.3 with a single model so it's definitely possible. Some of my trained models also broke 0.3 but never scored as high as yours. My current score uses your model as a baseline then just applies postprocessing.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3376903,
                  "author_name": "ravaghi",
                  "author_url": "",
                  "post_date": "12/15/2025 11:44:07",
                  "content": "<p>I haven't done cross-validation with DINOv2 yet, so I don't know the CV score for that model. I was actually referring to my public U-Net model, which scores 0.16 on the public LB and has a CV score of ~0.6. I plan to train DINOv2 with cross-validation as well, but first I need to fix the leakage issues I'm currently facing.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3376910,
                      "author_name": "returnofsputnik",
                      "author_url": "",
                      "post_date": "12/15/2025 12:00:06",
                      "content": "<p>GroupKFold on case_id with a pretrained UNet, average across 5 models, and tune the mask thresholds according to local CV has yielded me 0.30X to 0.31X. Average across 5 folds brought LB from 0.26X to 0.3, so I think that part is quite important.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3372601,
      "author_name": "mrearthworm",
      "author_url": "",
      "post_date": "12/12/2025 05:23:56",
      "content": "<p>yes this challenge has this issue, i think this is because of a domain shift in the training data that they have given us and the testing data that they are using for testing our models on the private dataset</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3371414": "Hi everyone,\n\nFor validation, I used the exact score function provided in the competition, so I expected the numbers to be much closer. I am seeing a big difference between **my local validation score (~0.72) and the public leaderboard score (~0.10)**. I have checked the common mistakes like mask size, RLE format, and thresholding, and they seem fine.\n\nSo I’m wondering if anyone else has faced a similar drop. \n\nThanks!!",
    "3372012": "Make sure you are groupkfold on the image. You dont want to train on forged imageA.png and then validate on authentic imageA.png. It may cause leak. 5fold model averaged should be able to break 0.3 on LB",
    "3372601": "yes this challenge has this issue, i think this is because of a domain shift in the training data that they have given us and the testing data that they are using for testing our models on the private dataset",
    "3376888": "Doesn't seem to help. I'm using case_id as group but still get scores that are too high ([~0.6](https://www.kaggle.com/code/ravaghi/scientific-image-forgery-detection-u-net-1-2)).",
    "3376901": "Your own notebook (Dinov2) breaks 0.3 with a single model so it's definitely possible. Some of my trained models also broke 0.3 but never scored as high as yours. My current score uses your model as a baseline then just applies postprocessing.",
    "3376903": "I haven't done cross-validation with DINOv2 yet, so I don't know the CV score for that model. I was actually referring to my public U-Net model, which scores 0.16 on the public LB and has a CV score of ~0.6. I plan to train DINOv2 with cross-validation as well, but first I need to fix the leakage issues I'm currently facing.",
    "3376910": "GroupKFold on case_id with a pretrained UNet, average across 5 models, and tune the mask thresholds according to local CV has yielded me 0.30X to 0.31X. Average across 5 folds brought LB from 0.26X to 0.3, so I think that part is quite important."
  },
  "source": "meta"
}