{
  "id": 635546,
  "title": "Discussion about Segmentation score",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/635546",
  "author_name": "",
  "post_date": "2025-11-20T13:25:23.034525900Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>When I submitted my code, I assumed all the images were forged and predicted the corresponding RLE encodings, but the score was only 0.03. This indicates that my image segmentation model performed very poorly, but it seemed to perform quite well during local training. Has anyone encountered this problem? Is it because the mask data in the test data is very different from that in the training data, or is it because my inference code is wrong?</p>",
  "messages": [
    {
      "id": "3341843",
      "postDate": "11/20/2025 13:25:23",
      "content": "<p>When I submitted my code, I assumed all the images were forged and predicted the corresponding RLE encodings, but the score was only 0.03. This indicates that my image segmentation model performed very poorly, but it seemed to perform quite well during local training. Has anyone encountered this problem? Is it because the mask data in the test data is very different from that in the training data, or is it because my inference code is wrong?</p>",
      "rawMarkdown": "When I submitted my code, I assumed all the images were forged and predicted the corresponding RLE encodings, but the score was only 0.03. This indicates that my image segmentation model performed very poorly, but it seemed to perform quite well during local training. Has anyone encountered this problem? Is it because the mask data in the test data is very different from that in the training data, or is it because my inference code is wrong?",
      "votes": null
    },
    {
      "id": "3380919",
      "postDate": "12/23/2025 11:08:15",
      "content": "<p>Test dataset is completely different from train and much harder</p>",
      "rawMarkdown": "Test dataset is completely different from train and much harder",
      "votes": null
    },
    {
      "id": "3389617",
      "postDate": "01/11/2026 15:18:30",
      "content": "<p>I don't know man, in local I am getting competition metric 0.45-0.5. But when I am submitting, I am getting 0.03. During my training I didn't include supplement images, because all the images inside train are cropped versions of actual images. When I validated my model on supplement images, it was giving the same score on submission. So I am assuming the test dataset has a very high resemblance to supplement images.</p>\n<p>If you see the supplement images are high resolutions or you can say multi-panel, multiple train_images (panel) are sticked together like a matplotlib grid plot, I thought of doing a tiling based inference but no luck, the portions that I am cropping are becoming different from the train images. </p>\n<p>Feeling really bad ⚰️</p>",
      "rawMarkdown": "I don't know man, in local I am getting competition metric 0.45-0.5. But when I am submitting, I am getting 0.03. During my training I didn't include supplement images, because all the images inside train are cropped versions of actual images. When I validated my model on supplement images, it was giving the same score on submission. So I am assuming the test dataset has a very high resemblance to supplement images.\n\nIf you see the supplement images are high resolutions or you can say multi-panel, multiple train_images (panel) are sticked together like a matplotlib grid plot, I thought of doing a tiling based inference but no luck, the portions that I am cropping are becoming different from the train images. \n\nFeeling really bad ⚰️",
      "votes": null
    },
    {
      "id": "3389645",
      "postDate": "01/11/2026 16:13:43",
      "content": "<p>I developed an agnostic model. It performs very well during the test phase. It is successful on both single-panel and multi-panel images. However, it gets a score of 0.25 on the leaderboard. I think the test data has different characteristics. Many people are using DINOv-based solutions and getting high scores. But DINOv-based models, especially on the supplementary multi-panel images, color them incorrectly. Your score is very low. I think there is a mistake in the submission format.</p>",
      "rawMarkdown": "I developed an agnostic model. It performs very well during the test phase. It is successful on both single-panel and multi-panel images. However, it gets a score of 0.25 on the leaderboard. I think the test data has different characteristics. Many people are using DINOv-based solutions and getting high scores. But DINOv-based models, especially on the supplementary multi-panel images, color them incorrectly. Your score is very low. I think there is a mistake in the submission format.",
      "votes": null
    },
    {
      "id": "3389767",
      "postDate": "01/11/2026 20:52:04",
      "content": "<p>There is probably no error here, it's just that the test data is very different. </p>",
      "rawMarkdown": "There is probably no error here, it's just that the test data is very different.",
      "votes": null
    },
    {
      "id": "3390706",
      "postDate": "01/13/2026 19:47:45",
      "content": "<p>Is there a way to use supplement masks?</p>",
      "rawMarkdown": "Is there a way to use supplement masks?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3380919,
      "author_name": "theodorlu",
      "author_url": "",
      "post_date": "12/23/2025 11:08:15",
      "content": "<p>Test dataset is completely different from train and much harder</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3389617,
      "author_name": "hotsonhonet",
      "author_url": "",
      "post_date": "01/11/2026 15:18:30",
      "content": "<p>I don't know man, in local I am getting competition metric 0.45-0.5. But when I am submitting, I am getting 0.03. During my training I didn't include supplement images, because all the images inside train are cropped versions of actual images. When I validated my model on supplement images, it was giving the same score on submission. So I am assuming the test dataset has a very high resemblance to supplement images.</p>\n<p>If you see the supplement images are high resolutions or you can say multi-panel, multiple train_images (panel) are sticked together like a matplotlib grid plot, I thought of doing a tiling based inference but no luck, the portions that I am cropping are becoming different from the train images. </p>\n<p>Feeling really bad ⚰️</p>",
      "votes": null,
      "replies": [
        {
          "id": 3389645,
          "author_name": "musapeker",
          "author_url": "",
          "post_date": "01/11/2026 16:13:43",
          "content": "<p>I developed an agnostic model. It performs very well during the test phase. It is successful on both single-panel and multi-panel images. However, it gets a score of 0.25 on the leaderboard. I think the test data has different characteristics. Many people are using DINOv-based solutions and getting high scores. But DINOv-based models, especially on the supplementary multi-panel images, color them incorrectly. Your score is very low. I think there is a mistake in the submission format.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3389767,
              "author_name": "aitbekovalibek",
              "author_url": "",
              "post_date": "01/11/2026 20:52:04",
              "content": "<p>There is probably no error here, it's just that the test data is very different. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3390706,
                  "author_name": "hotsonhonet",
                  "author_url": "",
                  "post_date": "01/13/2026 19:47:45",
                  "content": "<p>Is there a way to use supplement masks?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3341843": "When I submitted my code, I assumed all the images were forged and predicted the corresponding RLE encodings, but the score was only 0.03. This indicates that my image segmentation model performed very poorly, but it seemed to perform quite well during local training. Has anyone encountered this problem? Is it because the mask data in the test data is very different from that in the training data, or is it because my inference code is wrong?",
    "3380919": "Test dataset is completely different from train and much harder",
    "3389617": "I don't know man, in local I am getting competition metric 0.45-0.5. But when I am submitting, I am getting 0.03. During my training I didn't include supplement images, because all the images inside train are cropped versions of actual images. When I validated my model on supplement images, it was giving the same score on submission. So I am assuming the test dataset has a very high resemblance to supplement images.\n\nIf you see the supplement images are high resolutions or you can say multi-panel, multiple train_images (panel) are sticked together like a matplotlib grid plot, I thought of doing a tiling based inference but no luck, the portions that I am cropping are becoming different from the train images. \n\nFeeling really bad ⚰️",
    "3389645": "I developed an agnostic model. It performs very well during the test phase. It is successful on both single-panel and multi-panel images. However, it gets a score of 0.25 on the leaderboard. I think the test data has different characteristics. Many people are using DINOv-based solutions and getting high scores. But DINOv-based models, especially on the supplementary multi-panel images, color them incorrectly. Your score is very low. I think there is a mistake in the submission format.",
    "3389767": "There is probably no error here, it's just that the test data is very different.",
    "3390706": "Is there a way to use supplement masks?"
  },
  "source": "meta"
}