{
  "id": 613694,
  "title": "Possible Labeling Error: Computed Mask vs. Ground Truth Mask Mismatch",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613694",
  "author_name": "Youssef Ouertani",
  "post_date": "2025-10-29T00:53:59.834000",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>While exploring the dataset, I noticed a potential labeling issue. I computed a mask by comparing the authentic and forged images (pixels that differ). This computed mask should match the ground truth mask in terms of the number and location of forged regions.</p>\n<p>However, in some examples I checked, the results do not align:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F2e3736737aaff737e0354dea9808c03f%2FScreenshot%20from%202025-10-29%2012-17-32.png?generation=1761736681041566&amp;alt=media\" alt=\"\"></p>\n<p>The computed mask identifies one red object, while the ground truth mask shows three. However, only two should be present — the red object in the middle should not be labeled.</p>\n<p>Since the ground truth is expected to represent all forged regions and copied objects, this inconsistency suggests that some annotations may be incorrect or incomplete.</p>\n<p>Could the organizers confirm whether this labeling issue will be corrected? Additionally, if similar errors exist in the test set, they could impact evaluation fairness — is there a plan to mitigate this risk?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 3308254,
      "postDate": "2025-10-29T00:53:59.833Z",
      "content": "<p>Hi everyone,</p>\n<p>While exploring the dataset, I noticed a potential labeling issue. I computed a mask by comparing the authentic and forged images (pixels that differ). This computed mask should match the ground truth mask in terms of the number and location of forged regions.</p>\n<p>However, in some examples I checked, the results do not align:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F2e3736737aaff737e0354dea9808c03f%2FScreenshot%20from%202025-10-29%2012-17-32.png?generation=1761736681041566&amp;alt=media\" alt=\"\"></p>\n<p>The computed mask identifies one red object, while the ground truth mask shows three. However, only two should be present — the red object in the middle should not be labeled.</p>\n<p>Since the ground truth is expected to represent all forged regions and copied objects, this inconsistency suggests that some annotations may be incorrect or incomplete.</p>\n<p>Could the organizers confirm whether this labeling issue will be corrected? Additionally, if similar errors exist in the test set, they could impact evaluation fairness — is there a plan to mitigate this risk?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi everyone,\n\nWhile exploring the dataset, I noticed a potential labeling issue. I computed a mask by comparing the authentic and forged images (pixels that differ). This computed mask should match the ground truth mask in terms of the number and location of forged regions.\n\nHowever, in some examples I checked, the results do not align:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F2e3736737aaff737e0354dea9808c03f%2FScreenshot%20from%202025-10-29%2012-17-32.png?generation=1761736681041566&alt=media)\n\nThe computed mask identifies one red object, while the ground truth mask shows three. However, only two should be present — the red object in the middle should not be labeled.\n\nSince the ground truth is expected to represent all forged regions and copied objects, this inconsistency suggests that some annotations may be incorrect or incomplete.\n\nCould the organizers confirm whether this labeling issue will be corrected? Additionally, if similar errors exist in the test set, they could impact evaluation fairness — is there a plan to mitigate this risk?\n\nThanks!\n",
      "votes": 11
    },
    {
      "id": 3308430,
      "postDate": "2025-10-29T12:08:54.883Z",
      "content": "<p>Thanks for pointing this out!</p>\n<p>I can confirm that this is an extremely rare case in the training set, and it does not occur in the test set. The test set was annotated from actual forgeries found in research papers, while the training set was created by our team.</p>\n<p>In the sample you highlighted, the red areas from the two top cells (left and middle) were copied together as a single object. After being rotated, this new piece was pasted to create the forgery on the right. The red cell on the right only contains a partial view of the combined elements because the pasted piece was placed partially outside the image boundary.</p>\n<p>Thanks again for the catch! I hope that clarification helps.</p>",
      "rawMarkdown": "Thanks for pointing this out!\n\nI can confirm that this is an extremely rare case in the training set, and it does not occur in the test set. The test set was annotated from actual forgeries found in research papers, while the training set was created by our team.\n\nIn the sample you highlighted, the red areas from the two top cells (left and middle) were copied together as a single object. After being rotated, this new piece was pasted to create the forgery on the right. The red cell on the right only contains a partial view of the combined elements because the pasted piece was placed partially outside the image boundary.\n\nThanks again for the catch! I hope that clarification helps.",
      "votes": 3,
      "replies": [
        {
          "id": 3311595,
          "postDate": "2025-11-05T11:23:59.193Z",
          "content": "<p>mdam… that's what I thought. I've already submitted 70+ submissions- I've used R-CNN, YOLO, DeepLab, and all the models don't predict even at high conf. This means that test does not compete with train at all.</p>",
          "rawMarkdown": "mdam... that's what I thought. I've already submitted 70+ submissions- I've used R-CNN, YOLO, DeepLab, and all the models don't predict even at high conf. This means that test does not compete with train at all."
        }
      ]
    },
    {
      "id": 3308345,
      "postDate": "2025-10-29T06:58:02.990Z",
      "content": "<p>Did you notice that mask.shape is (2, 712, 414) for image 10070 ?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F5d50533d330783150061e9829e0f11e2%2Fa.png?generation=1761721063184691&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Did you notice that mask.shape is (2, 712, 414) for image 10070 ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F5d50533d330783150061e9829e0f11e2%2Fa.png?generation=1761721063184691&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 3308416,
          "postDate": "2025-10-29T11:24:09.727Z",
          "content": "<p>Thanks, I updated the visualization for the 2-channel mask. The annotation issue remains though , the GT shows three red regions, but only two are valid (the middle one shouldn’t be labeled).</p>",
          "rawMarkdown": "Thanks, I updated the visualization for the 2-channel mask. The annotation issue remains though , the GT shows three red regions, but only two are valid (the middle one shouldn’t be labeled)."
        }
      ]
    },
    {
      "id": 3308333,
      "postDate": "2025-10-29T05:59:10.370Z",
      "content": "<p>This is a good catch thanks for raising this</p>",
      "rawMarkdown": "This is a good catch thanks for raising this\n"
    },
    {
      "id": 3311410,
      "postDate": "2025-11-05T00:31:30.850Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3308430,
      "author_name": "João Phillipe Cardenuto",
      "author_url": "",
      "post_date": "2025-10-29T12:08:54.883000",
      "content": "<p>Thanks for pointing this out!</p>\n<p>I can confirm that this is an extremely rare case in the training set, and it does not occur in the test set. The test set was annotated from actual forgeries found in research papers, while the training set was created by our team.</p>\n<p>In the sample you highlighted, the red areas from the two top cells (left and middle) were copied together as a single object. After being rotated, this new piece was pasted to create the forgery on the right. The red cell on the right only contains a partial view of the combined elements because the pasted piece was placed partially outside the image boundary.</p>\n<p>Thanks again for the catch! I hope that clarification helps.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3311595,
          "author_name": "Antonoof",
          "author_url": "",
          "post_date": "2025-11-05T11:23:59.193000",
          "content": "<p>mdam… that's what I thought. I've already submitted 70+ submissions- I've used R-CNN, YOLO, DeepLab, and all the models don't predict even at high conf. This means that test does not compete with train at all.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3308345,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "2025-10-29T06:58:02.990000",
      "content": "<p>Did you notice that mask.shape is (2, 712, 414) for image 10070 ?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F5d50533d330783150061e9829e0f11e2%2Fa.png?generation=1761721063184691&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 3308416,
          "author_name": "Youssef Ouertani",
          "author_url": "",
          "post_date": "2025-10-29T11:24:09.727000",
          "content": "<p>Thanks, I updated the visualization for the 2-channel mask. The annotation issue remains though , the GT shows three red regions, but only two are valid (the middle one shouldn’t be labeled).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3308333,
      "author_name": "Manel ALOUI",
      "author_url": "",
      "post_date": "2025-10-29T05:59:10.370000",
      "content": "<p>This is a good catch thanks for raising this</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3311410,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-11-05T00:31:30.850000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3308254": "Hi everyone,\n\nWhile exploring the dataset, I noticed a potential labeling issue. I computed a mask by comparing the authentic and forged images (pixels that differ). This computed mask should match the ground truth mask in terms of the number and location of forged regions.\n\nHowever, in some examples I checked, the results do not align:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6498106%2F2e3736737aaff737e0354dea9808c03f%2FScreenshot%20from%202025-10-29%2012-17-32.png?generation=1761736681041566&alt=media)\n\nThe computed mask identifies one red object, while the ground truth mask shows three. However, only two should be present — the red object in the middle should not be labeled.\n\nSince the ground truth is expected to represent all forged regions and copied objects, this inconsistency suggests that some annotations may be incorrect or incomplete.\n\nCould the organizers confirm whether this labeling issue will be corrected? Additionally, if similar errors exist in the test set, they could impact evaluation fairness — is there a plan to mitigate this risk?\n\nThanks!\n",
    "3308430": "Thanks for pointing this out!\n\nI can confirm that this is an extremely rare case in the training set, and it does not occur in the test set. The test set was annotated from actual forgeries found in research papers, while the training set was created by our team.\n\nIn the sample you highlighted, the red areas from the two top cells (left and middle) were copied together as a single object. After being rotated, this new piece was pasted to create the forgery on the right. The red cell on the right only contains a partial view of the combined elements because the pasted piece was placed partially outside the image boundary.\n\nThanks again for the catch! I hope that clarification helps.",
    "3308345": "Did you notice that mask.shape is (2, 712, 414) for image 10070 ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F5d50533d330783150061e9829e0f11e2%2Fa.png?generation=1761721063184691&alt=media)",
    "3308333": "This is a good catch thanks for raising this\n",
    "3311410": ""
  }
}