{
  "id": 617968,
  "title": "Supplemental train data added",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/617968",
  "author_name": "",
  "post_date": "2025-11-12T19:16:55.847961400Z",
  "votes": 18,
  "comment_count": 13,
  "views": 0,
  "content": "<p>We've supplemented the train set with new images and masks. These new images were acquired and labeled using the exact process that will be used for the final test. Hopefully this new data alleviates the concerns about distribution shift that have been raised in the forums.</p>\n<p>Edit: The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:</p>\n<p>1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.</p>\n<p>In some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.</p>",
  "messages": [
    {
      "id": "3320762",
      "postDate": "11/12/2025 19:16:55",
      "content": "<p>We've supplemented the train set with new images and masks. These new images were acquired and labeled using the exact process that will be used for the final test. Hopefully this new data alleviates the concerns about distribution shift that have been raised in the forums.</p>\n<p>Edit: The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:</p>\n<p>1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.</p>\n<p>In some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.</p>",
      "rawMarkdown": "We've supplemented the train set with new images and masks. These new images were acquired and labeled using the exact process that will be used for the final test. Hopefully this new data alleviates the concerns about distribution shift that have been raised in the forums.\n\nEdit: The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:\n\n1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.\n\nIn some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.",
      "votes": null
    },
    {
      "id": "3342378",
      "postDate": "11/20/2025 22:02:07",
      "content": "<p>I just want to acknowledge that one of the supplemental test images has an unlabeled dupe. So should we expect some rare cases of accidentally not labeled?    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238844%2F0b4677bad76ae200448f5777062d5fe0%2Fdupe1.png?generation=1763676102856517&amp;alt=media\" alt=\"\"></p>\n<p>Please see attached images</p>",
      "rawMarkdown": "I just want to acknowledge that one of the supplemental test images has an unlabeled dupe. So should we expect some rare cases of accidentally not labeled?    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238844%2F0b4677bad76ae200448f5777062d5fe0%2Fdupe1.png?generation=1763676102856517&alt=media)\n\nPlease see attached images",
      "votes": null
    },
    {
      "id": "3363937",
      "postDate": "12/06/2025 07:21:16",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Why are all the images in supplemental_images forged? Is it possible to provide their corresponding authentic images?</p>",
      "rawMarkdown": "sohier Why are all the images in supplemental_images forged? Is it possible to provide their corresponding authentic images?",
      "votes": null
    },
    {
      "id": "3369159",
      "postDate": "12/09/2025 19:07:31",
      "content": "<p>Thanks for the catch! We are reviewing the entire supplemental training data, checking for any other possible non-reported forgery in these images. </p>",
      "rawMarkdown": "Thanks for the catch! We are reviewing the entire supplemental training data, checking for any other possible non-reported forgery in these images.",
      "votes": null
    },
    {
      "id": "3369163",
      "postDate": "12/09/2025 19:11:33",
      "content": "<p>All images in the supplemental_images set were collected and annotated from actual reported cases of scientific image manipulation. Since these are real-world forgeries, we are not able to provide their authentic versions.</p>\n<p>On the other hand, the training data was created by our team. This is why we were able to provide both the authentic versions (prior to our manipulation) and the forged ones for that specific set.</p>",
      "rawMarkdown": "All images in the supplemental_images set were collected and annotated from actual reported cases of scientific image manipulation. Since these are real-world forgeries, we are not able to provide their authentic versions.\n\nOn the other hand, the training data was created by our team. This is why we were able to provide both the authentic versions (prior to our manipulation) and the forged ones for that specific set.",
      "votes": null
    },
    {
      "id": "3370290",
      "postDate": "12/10/2025 17:21:24",
      "content": "<p>The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:</p>\n<p>1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.</p>\n<p>In some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.</p>",
      "rawMarkdown": "The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:\n\n1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.\n\nIn some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.",
      "votes": null
    },
    {
      "id": "3377974",
      "postDate": "12/17/2025 09:00:59",
      "content": "<blockquote>\n  <p>These new images were acquired and labeled using the exact process that will be used for the final test.</p>\n</blockquote>\n<p>I just started to look into this competition and I am quite confused by this statement. As others have pointed out, the supplemental data is completely different than the training data (panels instead of images). Does the hidden test set contain panels or images (or both)?</p>\n<ul>\n<li>If it only contains panels, what is the reason behind providing the training data?</li>\n<li>If it only contains images, what is the reason behind providing the supplemental data?</li>\n<li>If it contains both panels and images, why provide training data and supplemental data separately?</li>\n</ul>\n<p>Thanks a lot!</p>",
      "rawMarkdown": "> These new images were acquired and labeled using the exact process that will be used for the final test.\n\nI just started to look into this competition and I am quite confused by this statement. As others have pointed out, the supplemental data is completely different than the training data (panels instead of images). Does the hidden test set contain panels or images (or both)?\n\n- If it only contains panels, what is the reason behind providing the training data?\n- If it only contains images, what is the reason behind providing the supplemental data?\n- If it contains both panels and images, why provide training data and supplemental data separately?\n\nThanks a lot!",
      "votes": null
    },
    {
      "id": "3378018",
      "postDate": "12/17/2025 11:34:03",
      "content": "<p>You should expect both types of images (multi-panel and single-panel) in the testset.</p>\n<p>The main difference between the datasets is their origin:</p>\n<p>For the training data, we used authentic biomedical images to create their manipulated versions. For this process, we focused on single-panel images. This is why you can find both the authentic (non-forged) and their manipulated versions in the training set.</p>\n<p>For the supplemental and test data, we collected them from actual reported cases of scientific image manipulation. They reflect real-world complexity and may include multi-panel figures and single-panel images.</p>",
      "rawMarkdown": "You should expect both types of images (multi-panel and single-panel) in the testset.\n\nThe main difference between the datasets is their origin:\n\nFor the training data, we used authentic biomedical images to create their manipulated versions. For this process, we focused on single-panel images. This is why you can find both the authentic (non-forged) and their manipulated versions in the training set.\n\nFor the supplemental and test data, we collected them from actual reported cases of scientific image manipulation. They reflect real-world complexity and may include multi-panel figures and single-panel images.",
      "votes": null
    },
    {
      "id": "3379114",
      "postDate": "12/19/2025 07:31:29",
      "content": "<p>Yeah I thought the training date was random copy-move since the forged objects were located mostly in meaningless location ocurances… </p>",
      "rawMarkdown": "Yeah I thought the training date was random copy-move since the forged objects were located mostly in meaningless location ocurances...",
      "votes": null
    },
    {
      "id": "3380507",
      "postDate": "12/22/2025 13:59:47",
      "content": "<p>Is the testing set mostly multi-panel or single-panel? </p>",
      "rawMarkdown": "Is the testing set mostly multi-panel or single-panel?",
      "votes": null
    },
    {
      "id": "3380968",
      "postDate": "12/23/2025 13:14:30",
      "content": "<p>Unfortunately, to prevent data leakage, we cannot provide specific details about the composition of the current test set.</p>\n<p>Additionally, once the first stage of the competition closes, we will start collecting new images to include in the Phase 2 test set (focusing on cases reported after the 1st stage closes). We cannot predict exactly what image types will be found during that period.</p>\n<p>But, the images in the supplemental set are good candidates of the real-world complexity you can expect in the final test set.</p>\n<p>I strongly recommend that you do not overfit your model to a specific format (like single-panel vs. multi-panel). Instead, try to solve the problem with a generic and robust approach. Scientific images, especially in the biomedical field, can be quite complex and represent completely different experiments, yet the types of forgery applied tend to be similar across them.</p>\n<p>Hope that helps!</p>",
      "rawMarkdown": "Unfortunately, to prevent data leakage, we cannot provide specific details about the composition of the current test set.\n\nAdditionally, once the first stage of the competition closes, we will start collecting new images to include in the Phase 2 test set (focusing on cases reported after the 1st stage closes). We cannot predict exactly what image types will be found during that period.\n\nBut, the images in the supplemental set are good candidates of the real-world complexity you can expect in the final test set.\n\nI strongly recommend that you do not overfit your model to a specific format (like single-panel vs. multi-panel). Instead, try to solve the problem with a generic and robust approach. Scientific images, especially in the biomedical field, can be quite complex and represent completely different experiments, yet the types of forgery applied tend to be similar across them.\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "3380971",
      "postDate": "12/23/2025 13:37:01",
      "content": "<p>Scientific Image Forgery Detection</p>",
      "rawMarkdown": "Scientific Image Forgery Detection",
      "votes": null
    },
    {
      "id": "3381129",
      "postDate": "12/23/2025 19:55:20",
      "content": "<p>The supplemental images are very different from the train images. If the supplemental images are closer to the true test images than this competition essentially just gives 48 images to train on. Which leaves 2 options it seems: </p>\n<ol>\n<li>use a non-end-to-end approach</li>\n<li>find more data.</li>\n</ol>",
      "rawMarkdown": "The supplemental images are very different from the train images. If the supplemental images are closer to the true test images than this competition essentially just gives 48 images to train on. Which leaves 2 options it seems: \n1. use a non-end-to-end approach\n2. find more data.",
      "votes": null
    },
    {
      "id": "3381133",
      "postDate": "12/23/2025 20:07:33",
      "content": "<p>Yes, but images in the train dataset have nothing to do with images in the test dataset. Why even provide it? </p>",
      "rawMarkdown": "Yes, but images in the train dataset have nothing to do with images in the test dataset. Why even provide it?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3342378,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "11/20/2025 22:02:07",
      "content": "<p>I just want to acknowledge that one of the supplemental test images has an unlabeled dupe. So should we expect some rare cases of accidentally not labeled?    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238844%2F0b4677bad76ae200448f5777062d5fe0%2Fdupe1.png?generation=1763676102856517&amp;alt=media\" alt=\"\"></p>\n<p>Please see attached images</p>",
      "votes": null,
      "replies": [
        {
          "id": 3369159,
          "author_name": "joophillipecardenuto",
          "author_url": "",
          "post_date": "12/09/2025 19:07:31",
          "content": "<p>Thanks for the catch! We are reviewing the entire supplemental training data, checking for any other possible non-reported forgery in these images. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3363937,
      "author_name": "yanzhou06",
      "author_url": "",
      "post_date": "12/06/2025 07:21:16",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Why are all the images in supplemental_images forged? Is it possible to provide their corresponding authentic images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3369163,
          "author_name": "joophillipecardenuto",
          "author_url": "",
          "post_date": "12/09/2025 19:11:33",
          "content": "<p>All images in the supplemental_images set were collected and annotated from actual reported cases of scientific image manipulation. Since these are real-world forgeries, we are not able to provide their authentic versions.</p>\n<p>On the other hand, the training data was created by our team. This is why we were able to provide both the authentic versions (prior to our manipulation) and the forged ones for that specific set.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3379114,
              "author_name": "jirkaborovec",
              "author_url": "",
              "post_date": "12/19/2025 07:31:29",
              "content": "<p>Yeah I thought the training date was random copy-move since the forged objects were located mostly in meaningless location ocurances… </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3370290,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "12/10/2025 17:21:24",
      "content": "<p>The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:</p>\n<p>1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.</p>\n<p>In some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3377974,
      "author_name": "ekarais",
      "author_url": "",
      "post_date": "12/17/2025 09:00:59",
      "content": "<blockquote>\n  <p>These new images were acquired and labeled using the exact process that will be used for the final test.</p>\n</blockquote>\n<p>I just started to look into this competition and I am quite confused by this statement. As others have pointed out, the supplemental data is completely different than the training data (panels instead of images). Does the hidden test set contain panels or images (or both)?</p>\n<ul>\n<li>If it only contains panels, what is the reason behind providing the training data?</li>\n<li>If it only contains images, what is the reason behind providing the supplemental data?</li>\n<li>If it contains both panels and images, why provide training data and supplemental data separately?</li>\n</ul>\n<p>Thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3378018,
          "author_name": "joophillipecardenuto",
          "author_url": "",
          "post_date": "12/17/2025 11:34:03",
          "content": "<p>You should expect both types of images (multi-panel and single-panel) in the testset.</p>\n<p>The main difference between the datasets is their origin:</p>\n<p>For the training data, we used authentic biomedical images to create their manipulated versions. For this process, we focused on single-panel images. This is why you can find both the authentic (non-forged) and their manipulated versions in the training set.</p>\n<p>For the supplemental and test data, we collected them from actual reported cases of scientific image manipulation. They reflect real-world complexity and may include multi-panel figures and single-panel images.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3380507,
              "author_name": "theodorlu",
              "author_url": "",
              "post_date": "12/22/2025 13:59:47",
              "content": "<p>Is the testing set mostly multi-panel or single-panel? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3380968,
                  "author_name": "joophillipecardenuto",
                  "author_url": "",
                  "post_date": "12/23/2025 13:14:30",
                  "content": "<p>Unfortunately, to prevent data leakage, we cannot provide specific details about the composition of the current test set.</p>\n<p>Additionally, once the first stage of the competition closes, we will start collecting new images to include in the Phase 2 test set (focusing on cases reported after the 1st stage closes). We cannot predict exactly what image types will be found during that period.</p>\n<p>But, the images in the supplemental set are good candidates of the real-world complexity you can expect in the final test set.</p>\n<p>I strongly recommend that you do not overfit your model to a specific format (like single-panel vs. multi-panel). Instead, try to solve the problem with a generic and robust approach. Scientific images, especially in the biomedical field, can be quite complex and represent completely different experiments, yet the types of forgery applied tend to be similar across them.</p>\n<p>Hope that helps!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3381133,
                      "author_name": "theodorlu",
                      "author_url": "",
                      "post_date": "12/23/2025 20:07:33",
                      "content": "<p>Yes, but images in the train dataset have nothing to do with images in the test dataset. Why even provide it? </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3380971,
      "author_name": "khushikyad001",
      "author_url": "",
      "post_date": "12/23/2025 13:37:01",
      "content": "<p>Scientific Image Forgery Detection</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3381129,
      "author_name": "eliplutchok",
      "author_url": "",
      "post_date": "12/23/2025 19:55:20",
      "content": "<p>The supplemental images are very different from the train images. If the supplemental images are closer to the true test images than this competition essentially just gives 48 images to train on. Which leaves 2 options it seems: </p>\n<ol>\n<li>use a non-end-to-end approach</li>\n<li>find more data.</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3320762": "We've supplemented the train set with new images and masks. These new images were acquired and labeled using the exact process that will be used for the final test. Hopefully this new data alleviates the concerns about distribution shift that have been raised in the forums.\n\nEdit: The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:\n\n1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.\n\nIn some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.",
    "3342378": "I just want to acknowledge that one of the supplemental test images has an unlabeled dupe. So should we expect some rare cases of accidentally not labeled?    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238844%2F0b4677bad76ae200448f5777062d5fe0%2Fdupe1.png?generation=1763676102856517&alt=media)\n\nPlease see attached images",
    "3363937": "sohier Why are all the images in supplemental_images forged? Is it possible to provide their corresponding authentic images?",
    "3369159": "Thanks for the catch! We are reviewing the entire supplemental training data, checking for any other possible non-reported forgery in these images.",
    "3369163": "All images in the supplemental_images set were collected and annotated from actual reported cases of scientific image manipulation. Since these are real-world forgeries, we are not able to provide their authentic versions.\n\nOn the other hand, the training data was created by our team. This is why we were able to provide both the authentic versions (prior to our manipulation) and the forged ones for that specific set.",
    "3370290": "The labels for the supplemental data have been refined with another manual review pass. Expect three images with major updates and seven with minor changes. The hosts triple-checked the supplemental images with the following process:\n\n1- Reviewing the original reporting guides.\n2- Screening every image manually for candidate issues.\n3- Running five different copy-move detectors models, including state-of-the-art models.\n\nIn some cases, these detectors (which were not applied before) identified duplications that were not documented in the original image reports. All detector results were manually verified.",
    "3377974": "> These new images were acquired and labeled using the exact process that will be used for the final test.\n\nI just started to look into this competition and I am quite confused by this statement. As others have pointed out, the supplemental data is completely different than the training data (panels instead of images). Does the hidden test set contain panels or images (or both)?\n\n- If it only contains panels, what is the reason behind providing the training data?\n- If it only contains images, what is the reason behind providing the supplemental data?\n- If it contains both panels and images, why provide training data and supplemental data separately?\n\nThanks a lot!",
    "3378018": "You should expect both types of images (multi-panel and single-panel) in the testset.\n\nThe main difference between the datasets is their origin:\n\nFor the training data, we used authentic biomedical images to create their manipulated versions. For this process, we focused on single-panel images. This is why you can find both the authentic (non-forged) and their manipulated versions in the training set.\n\nFor the supplemental and test data, we collected them from actual reported cases of scientific image manipulation. They reflect real-world complexity and may include multi-panel figures and single-panel images.",
    "3379114": "Yeah I thought the training date was random copy-move since the forged objects were located mostly in meaningless location ocurances...",
    "3380507": "Is the testing set mostly multi-panel or single-panel?",
    "3380968": "Unfortunately, to prevent data leakage, we cannot provide specific details about the composition of the current test set.\n\nAdditionally, once the first stage of the competition closes, we will start collecting new images to include in the Phase 2 test set (focusing on cases reported after the 1st stage closes). We cannot predict exactly what image types will be found during that period.\n\nBut, the images in the supplemental set are good candidates of the real-world complexity you can expect in the final test set.\n\nI strongly recommend that you do not overfit your model to a specific format (like single-panel vs. multi-panel). Instead, try to solve the problem with a generic and robust approach. Scientific images, especially in the biomedical field, can be quite complex and represent completely different experiments, yet the types of forgery applied tend to be similar across them.\n\nHope that helps!",
    "3380971": "Scientific Image Forgery Detection",
    "3381129": "The supplemental images are very different from the train images. If the supplemental images are closer to the true test images than this competition essentially just gives 48 images to train on. Which leaves 2 options it seems: \n1. use a non-end-to-end approach\n2. find more data.",
    "3381133": "Yes, but images in the train dataset have nothing to do with images in the test dataset. Why even provide it?"
  },
  "source": "meta"
}