{
  "id": 613765,
  "title": "Anyone observing discrepancies between local validation and LB?",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613765",
  "author_name": "",
  "post_date": "2025-10-29T15:27:52.651592100Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone, I have an inference pipeline (zero-shot, the main numerical parameter to be adjusted is a real number threshold of a score in which to flag a forgery or not, at <a href=\"https://www.kaggle.com/code/carsoncheng/sam-dice-score-thresholding-lb-0-257\" target=\"_blank\">https://www.kaggle.com/code/carsoncheng/sam-dice-score-thresholding-lb-0-257</a> ) that scores around 0.35 in local validation (the train dataset) (my local validation setup used sampling with 0.3 probability of flagging an authentic sample, just like in the leaderboard), and yet only scores 0.257 in the leaderboard. Worse still, when I zero out all the masks of predicting the forged samples, the score is 0.252, which means that only 0.005 score has gone into flagging the mask of the forged samples. My local validation is reporting around 85% recall on authentic samples (which is similar to the 84% recall on the leaderboard) and 0.10 score coming from the masks, but in the leaderboard, the scores for the masks all but disappeared.</p>\n<p>First and foremost, am I sending my masks to the submission.csv in my notebook the correct way? (the pipeline constructs masks that matches the corresponding shapes in the train_masks.csv and then used the official rle_encode to produce the mask, although with the massive mask score discrepancies I think it deserves a double check)</p>\n<p>If the mask format in my notebook is correct, I wonder whether a significant data drift between train and test is happening here; I've looked at the only test image available (45.png) being a figure with multiple subfigures and bar charts that could confuse the detection (that has hopefully been accounted for in my submissions with mask area over bounding box thresholding, in the notebook producing the aforementioned scores). But still I would not expect such a disparity between validation and LB scores. Even though it says the LB scores are not meaningful I think it's still something to be aware of. Has anyone experienced this and has anyone come up with some solution or hypothesis to this? Thanks 🙏</p>",
  "messages": [
    {
      "id": "3308512",
      "postDate": "10/29/2025 15:27:52",
      "content": "<p>Hi everyone, I have an inference pipeline (zero-shot, the main numerical parameter to be adjusted is a real number threshold of a score in which to flag a forgery or not, at <a href=\"https://www.kaggle.com/code/carsoncheng/sam-dice-score-thresholding-lb-0-257\" target=\"_blank\">https://www.kaggle.com/code/carsoncheng/sam-dice-score-thresholding-lb-0-257</a> ) that scores around 0.35 in local validation (the train dataset) (my local validation setup used sampling with 0.3 probability of flagging an authentic sample, just like in the leaderboard), and yet only scores 0.257 in the leaderboard. Worse still, when I zero out all the masks of predicting the forged samples, the score is 0.252, which means that only 0.005 score has gone into flagging the mask of the forged samples. My local validation is reporting around 85% recall on authentic samples (which is similar to the 84% recall on the leaderboard) and 0.10 score coming from the masks, but in the leaderboard, the scores for the masks all but disappeared.</p>\n<p>First and foremost, am I sending my masks to the submission.csv in my notebook the correct way? (the pipeline constructs masks that matches the corresponding shapes in the train_masks.csv and then used the official rle_encode to produce the mask, although with the massive mask score discrepancies I think it deserves a double check)</p>\n<p>If the mask format in my notebook is correct, I wonder whether a significant data drift between train and test is happening here; I've looked at the only test image available (45.png) being a figure with multiple subfigures and bar charts that could confuse the detection (that has hopefully been accounted for in my submissions with mask area over bounding box thresholding, in the notebook producing the aforementioned scores). But still I would not expect such a disparity between validation and LB scores. Even though it says the LB scores are not meaningful I think it's still something to be aware of. Has anyone experienced this and has anyone come up with some solution or hypothesis to this? Thanks 🙏</p>",
      "rawMarkdown": "Hi everyone, I have an inference pipeline (zero-shot, the main numerical parameter to be adjusted is a real number threshold of a score in which to flag a forgery or not, at https://www.kaggle.com/code/carsoncheng/sam-dice-score-thresholding-lb-0-257 ) that scores around 0.35 in local validation (the train dataset) (my local validation setup used sampling with 0.3 probability of flagging an authentic sample, just like in the leaderboard), and yet only scores 0.257 in the leaderboard. Worse still, when I zero out all the masks of predicting the forged samples, the score is 0.252, which means that only 0.005 score has gone into flagging the mask of the forged samples. My local validation is reporting around 85% recall on authentic samples (which is similar to the 84% recall on the leaderboard) and 0.10 score coming from the masks, but in the leaderboard, the scores for the masks all but disappeared.\n\nFirst and foremost, am I sending my masks to the submission.csv in my notebook the correct way? (the pipeline constructs masks that matches the corresponding shapes in the train_masks.csv and then used the official rle_encode to produce the mask, although with the massive mask score discrepancies I think it deserves a double check)\n\nIf the mask format in my notebook is correct, I wonder whether a significant data drift between train and test is happening here; I've looked at the only test image available (45.png) being a figure with multiple subfigures and bar charts that could confuse the detection (that has hopefully been accounted for in my submissions with mask area over bounding box thresholding, in the notebook producing the aforementioned scores). But still I would not expect such a disparity between validation and LB scores. Even though it says the LB scores are not meaningful I think it's still something to be aware of. Has anyone experienced this and has anyone come up with some solution or hypothesis to this? Thanks 🙏",
      "votes": null
    },
    {
      "id": "3308560",
      "postDate": "10/29/2025 16:43:29",
      "content": "<p>This is a highly challenging competition, and for image data, there is bound to be a certain discrepancy between local scores and LB (Leaderboard) scores. However, the gap between 0.005 and 0.100 is indeed enormous. I suspect that there are many new types of images in the test set, just like the first image with a case_id of 45, which resembles a data analysis chart. We haven't encountered such images in the training set. Nevertheless, I believe the core of this competition is to design a robust algorithm or model that can detect copy-move in any type of image.Good Luck！</p>",
      "rawMarkdown": "This is a highly challenging competition, and for image data, there is bound to be a certain discrepancy between local scores and LB (Leaderboard) scores. However, the gap between 0.005 and 0.100 is indeed enormous. I suspect that there are many new types of images in the test set, just like the first image with a case_id of 45, which resembles a data analysis chart. We haven't encountered such images in the training set. Nevertheless, I believe the core of this competition is to design a robust algorithm or model that can detect copy-move in any type of image.Good Luck！",
      "votes": null
    },
    {
      "id": "3340717",
      "postDate": "11/19/2025 17:05:45",
      "content": "<p>helpful,I met the big gap as well</p>",
      "rawMarkdown": "helpful,I met the big gap as well",
      "votes": null
    },
    {
      "id": "3341553",
      "postDate": "11/20/2025 09:01:53",
      "content": "<p>Sorry could you confirm that, when submitting the notebook, the folder \"…rgery-detection/test_images\" will contain the private test images from which the score will be computed? And that my notebook will generate a private submission.csv file?</p>",
      "rawMarkdown": "Sorry could you confirm that, when submitting the notebook, the folder \"...rgery-detection/test_images\" will contain the private test images from which the score will be computed? And that my notebook will generate a private submission.csv file?",
      "votes": null
    },
    {
      "id": "3350857",
      "postDate": "11/27/2025 21:53:45",
      "content": "<p>The answer is yes.   the folder \"…rgery-detection/test_images\" will contain the private test images from which the score will be computed and your notebook will generate a private submission.csv file.</p>",
      "rawMarkdown": "The answer is yes.   the folder \"…rgery-detection/test_images\" will contain the private test images from which the score will be computed and your notebook will generate a private submission.csv file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3308560,
      "author_name": "qifeihhh666",
      "author_url": "",
      "post_date": "10/29/2025 16:43:29",
      "content": "<p>This is a highly challenging competition, and for image data, there is bound to be a certain discrepancy between local scores and LB (Leaderboard) scores. However, the gap between 0.005 and 0.100 is indeed enormous. I suspect that there are many new types of images in the test set, just like the first image with a case_id of 45, which resembles a data analysis chart. We haven't encountered such images in the training set. Nevertheless, I believe the core of this competition is to design a robust algorithm or model that can detect copy-move in any type of image.Good Luck！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3340717,
      "author_name": "handudu",
      "author_url": "",
      "post_date": "11/19/2025 17:05:45",
      "content": "<p>helpful,I met the big gap as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3341553,
      "author_name": "andreshzapke",
      "author_url": "",
      "post_date": "11/20/2025 09:01:53",
      "content": "<p>Sorry could you confirm that, when submitting the notebook, the folder \"…rgery-detection/test_images\" will contain the private test images from which the score will be computed? And that my notebook will generate a private submission.csv file?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3350857,
          "author_name": "hamzahabduljalil",
          "author_url": "",
          "post_date": "11/27/2025 21:53:45",
          "content": "<p>The answer is yes.   the folder \"…rgery-detection/test_images\" will contain the private test images from which the score will be computed and your notebook will generate a private submission.csv file.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3308512": "Hi everyone, I have an inference pipeline (zero-shot, the main numerical parameter to be adjusted is a real number threshold of a score in which to flag a forgery or not, at https://www.kaggle.com/code/carsoncheng/sam-dice-score-thresholding-lb-0-257 ) that scores around 0.35 in local validation (the train dataset) (my local validation setup used sampling with 0.3 probability of flagging an authentic sample, just like in the leaderboard), and yet only scores 0.257 in the leaderboard. Worse still, when I zero out all the masks of predicting the forged samples, the score is 0.252, which means that only 0.005 score has gone into flagging the mask of the forged samples. My local validation is reporting around 85% recall on authentic samples (which is similar to the 84% recall on the leaderboard) and 0.10 score coming from the masks, but in the leaderboard, the scores for the masks all but disappeared.\n\nFirst and foremost, am I sending my masks to the submission.csv in my notebook the correct way? (the pipeline constructs masks that matches the corresponding shapes in the train_masks.csv and then used the official rle_encode to produce the mask, although with the massive mask score discrepancies I think it deserves a double check)\n\nIf the mask format in my notebook is correct, I wonder whether a significant data drift between train and test is happening here; I've looked at the only test image available (45.png) being a figure with multiple subfigures and bar charts that could confuse the detection (that has hopefully been accounted for in my submissions with mask area over bounding box thresholding, in the notebook producing the aforementioned scores). But still I would not expect such a disparity between validation and LB scores. Even though it says the LB scores are not meaningful I think it's still something to be aware of. Has anyone experienced this and has anyone come up with some solution or hypothesis to this? Thanks 🙏",
    "3308560": "This is a highly challenging competition, and for image data, there is bound to be a certain discrepancy between local scores and LB (Leaderboard) scores. However, the gap between 0.005 and 0.100 is indeed enormous. I suspect that there are many new types of images in the test set, just like the first image with a case_id of 45, which resembles a data analysis chart. We haven't encountered such images in the training set. Nevertheless, I believe the core of this competition is to design a robust algorithm or model that can detect copy-move in any type of image.Good Luck！",
    "3340717": "helpful,I met the big gap as well",
    "3341553": "Sorry could you confirm that, when submitting the notebook, the folder \"...rgery-detection/test_images\" will contain the private test images from which the score will be computed? And that my notebook will generate a private submission.csv file?",
    "3350857": "The answer is yes.   the folder \"…rgery-detection/test_images\" will contain the private test images from which the score will be computed and your notebook will generate a private submission.csv file."
  },
  "source": "meta"
}