{
  "id": 613052,
  "title": "Welcome to the Scientific Image Forgery Detection Challenge!",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613052",
  "author_name": "João Phillipe Cardenuto",
  "post_date": "2025-10-23T23:48:59.292000",
  "votes": 14,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hello everyone, and a warm welcome to the competition!</p>\n<p>This is a milestone we've been working on for a long time. Having been deeply involved in the problem of scientific image integrity, I know firsthand how critical this issue is. We are incredibly excited to host this challenge and bring together a community of brilliant minds to tackle it.</p>\n<p>This is more than just a competition; we see it as a foundational step. Our hope is to build a lasting community dedicated to developing computational solutions for research integrity. We truly believe that the models and ideas generated here can have a significant, positive impact.</p>\n<p>I am personally very excited to see the novel approaches and creative solutions you will all design.</p>\n<p>Please use this forum as your main hub. Ask questions and connect with fellow participants. We are here to support you.</p>\n<p>Good luck, and thank you for joining this important mission!</p>",
  "messages": [
    {
      "id": 3306014,
      "postDate": "2025-10-23T23:48:59.293Z",
      "content": "<p>Hello everyone, and a warm welcome to the competition!</p>\n<p>This is a milestone we've been working on for a long time. Having been deeply involved in the problem of scientific image integrity, I know firsthand how critical this issue is. We are incredibly excited to host this challenge and bring together a community of brilliant minds to tackle it.</p>\n<p>This is more than just a competition; we see it as a foundational step. Our hope is to build a lasting community dedicated to developing computational solutions for research integrity. We truly believe that the models and ideas generated here can have a significant, positive impact.</p>\n<p>I am personally very excited to see the novel approaches and creative solutions you will all design.</p>\n<p>Please use this forum as your main hub. Ask questions and connect with fellow participants. We are here to support you.</p>\n<p>Good luck, and thank you for joining this important mission!</p>",
      "rawMarkdown": "Hello everyone, and a warm welcome to the competition!\n\nThis is a milestone we've been working on for a long time. Having been deeply involved in the problem of scientific image integrity, I know firsthand how critical this issue is. We are incredibly excited to host this challenge and bring together a community of brilliant minds to tackle it.\n\nThis is more than just a competition; we see it as a foundational step. Our hope is to build a lasting community dedicated to developing computational solutions for research integrity. We truly believe that the models and ideas generated here can have a significant, positive impact.\n\nI am personally very excited to see the novel approaches and creative solutions you will all design.\n\nPlease use this forum as your main hub. Ask questions and connect with fellow participants. We are here to support you.\n\nGood luck, and thank you for joining this important mission!",
      "votes": 14
    },
    {
      "id": 3381936,
      "postDate": "2025-12-26T02:01:44.900Z",
      "content": "<p>Question:</p>\n<blockquote>\n  <p>A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n  A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.</p>\n</blockquote>\n<p>Does this mean during the forecasting phase (private leaderboard), we should expect roughly 2,200 images total? Just want to confirm. So if our submission works successfully on current LB (~1,100 images) in &lt; 4 hours, then it would likely work in private leaderboard which would have a runtime of 9 hours.</p>\n<p>thanks</p>",
      "rawMarkdown": "Question:\n\n> A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n> A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.\n\nDoes this mean during the forecasting phase (private leaderboard), we should expect roughly 2,200 images total? Just want to confirm. So if our submission works successfully on current LB (~1,100 images) in < 4 hours, then it would likely work in private leaderboard which would have a runtime of 9 hours.\n\nthanks\n",
      "votes": 1,
      "replies": [
        {
          "id": 3382069,
          "postDate": "2025-12-26T11:29:01.503Z",
          "content": "<p>Yes, expect the final forecasting test set to be roughly 2,200 images.</p>\n<p>If your model currently respects the 4-hour runtime for the current set, you should be well within the time limit for the forecasting phase.</p>",
          "rawMarkdown": "Yes, expect the final forecasting test set to be roughly 2,200 images.\n\nIf your model currently respects the 4-hour runtime for the current set, you should be well within the time limit for the forecasting phase."
        }
      ]
    },
    {
      "id": 3366114,
      "postDate": "2025-12-07T14:44:46.613Z",
      "content": "<p>Will all images be png in test phase? Or should we account for non-png. I am asking because I have hardcoded .png file extension in my code.</p>",
      "rawMarkdown": "Will all images be png in test phase? Or should we account for non-png. I am asking because I have hardcoded .png file extension in my code.",
      "votes": 1,
      "replies": [
        {
          "id": 3369143,
          "postDate": "2025-12-09T18:55:33.767Z",
          "content": "<p>Yes, you should account for non-PNG images.</p>\n<p>Although most images are expected to be in PNG format, it is strongly recommended that you do not hardcode this. It is safer to make your code flexible enough to handle JPEG and other image formats as well.</p>",
          "rawMarkdown": "Yes, you should account for non-PNG images.\n\nAlthough most images are expected to be in PNG format, it is strongly recommended that you do not hardcode this. It is safer to make your code flexible enough to handle JPEG and other image formats as well.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3391889,
      "postDate": "2026-01-15T22:01:49.133Z",
      "content": "<p>After thorough examination, the images must be uploaded immediately to a secure remote server. One way to detect manipulation is to copy and rotate the object and check for repeated patterns. However, if the object was sourced from another image that I cannot access, that approach is not possible. And if pixels are added one by one using consistent RGB values, the manipulation becomes even harder to detect.</p>\n<p>On top of that, it is impossible to determine whether what we are seeing is physically plausible or biologically plausible. Living cells can look very similar, yet they are also unique. I do not see a fully reliable solution, and if I build a detection pipeline using an existing dataset to avoid these issues, I am concerned it will overfit to dataset-specific artifacts and end up flagging legitimate images as “fake.”</p>\n<p>The only practical alternative would be to detect some known ways of producing fake data, but that could give humanity a false belief in “truth”: any methods that remain undetected would effectively go unpunished, creating a misleading sense of security.</p>",
      "rawMarkdown": "After thorough examination, the images must be uploaded immediately to a secure remote server. One way to detect manipulation is to copy and rotate the object and check for repeated patterns. However, if the object was sourced from another image that I cannot access, that approach is not possible. And if pixels are added one by one using consistent RGB values, the manipulation becomes even harder to detect.\n\nOn top of that, it is impossible to determine whether what we are seeing is physically plausible or biologically plausible. Living cells can look very similar, yet they are also unique. I do not see a fully reliable solution, and if I build a detection pipeline using an existing dataset to avoid these issues, I am concerned it will overfit to dataset-specific artifacts and end up flagging legitimate images as “fake.”\n\nThe only practical alternative would be to detect some known ways of producing fake data, but that could give humanity a false belief in “truth”: any methods that remain undetected would effectively go unpunished, creating a misleading sense of security.",
      "replies": [
        {
          "id": 3391892,
          "postDate": "2026-01-15T22:06:30.233Z",
          "content": "<p>And I would much prefer a similar identification challenge on a dataset where we can say, with certainty, that everything is true—i.e., where the image acquisition methodology is fully documented (staining/coloration protocols, camera models and settings, illumination and spectral characteristics, optics, exposure, calibration, preprocessing, and any post-processing), along with clear ground-truth labels and provenance.</p>",
          "rawMarkdown": "And I would much prefer a similar identification challenge on a dataset where we can say, with certainty, that everything is true—i.e., where the image acquisition methodology is fully documented (staining/coloration protocols, camera models and settings, illumination and spectral characteristics, optics, exposure, calibration, preprocessing, and any post-processing), along with clear ground-truth labels and provenance."
        }
      ]
    },
    {
      "id": 3391770,
      "postDate": "2026-01-15T16:35:40.093Z",
      "content": "<p>will the training images folder be available in the testing phase?\nMy submission uses them to train a model on them and I need to know if they are available or i need to upload it to kaggle. <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> </p>",
      "rawMarkdown": "will the training images folder be available in the testing phase?\nMy submission uses them to train a model on them and I need to know if they are available or i need to upload it to kaggle. @joophillipecardenuto ",
      "replies": [
        {
          "id": 3391853,
          "postDate": "2026-01-15T20:00:11.090Z",
          "content": "<p>Yes, the training images will be available during the forecasting phase.</p>\n<p>The same notebook you submitted should work during that phase.</p>",
          "rawMarkdown": "Yes, the training images will be available during the forecasting phase.\n\nThe same notebook you submitted should work during that phase."
        }
      ]
    },
    {
      "id": 3388278,
      "postDate": "2026-01-08T14:46:25.083Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a>, I have a question.\nWhen using external data or pre-trained models, do we need to care about licenses, such as commercial use? Thank you.</p>",
      "rawMarkdown": "Hi @joophillipecardenuto, I have a question.\nWhen using external data or pre-trained models, do we need to care about licenses, such as commercial use? Thank you.",
      "replies": [
        {
          "id": 3388337,
          "postDate": "2026-01-08T16:41:18.617Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/moritake04\" target=\"_blank\">@moritake04</a>,</p>\n<p>Yes, you do need to be careful with licenses.</p>\n<p>According to the competition rules (rule 5), the winning solution must be open-sourced. If your solution depends on a pre-trained model or data with a restrictive license (one that prevents you from legally open-sourcing your code), you would likely be unable to fulfill this requirement.</p>\n<p>According to Rule 6, you must ensure that any external data you use is publicly available and equally accessible to all participants at no cost.</p>\n<p>Essentially, other participants should be able to access the external data and code to reproduce your results after the competition is finished.</p>\n<p>Please check Rules 5 and 6 here: <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/rules\" target=\"_blank\">https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/rules</a></p>",
          "rawMarkdown": "Hi @moritake04,\n\nYes, you do need to be careful with licenses.\n\nAccording to the competition rules (rule 5), the winning solution must be open-sourced. If your solution depends on a pre-trained model or data with a restrictive license (one that prevents you from legally open-sourcing your code), you would likely be unable to fulfill this requirement.\n\nAccording to Rule 6, you must ensure that any external data you use is publicly available and equally accessible to all participants at no cost.\n\nEssentially, other participants should be able to access the external data and code to reproduce your results after the competition is finished.\n\nPlease check Rules 5 and 6 here: https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/rules",
          "votes": 1,
          "replies": [
            {
              "id": 3388486,
              "postDate": "2026-01-09T00:29:55.287Z",
              "content": "<p>Do you allow Dinov3 models?\n<a href=\"https://arxiv.org/pdf/2508.10104\" target=\"_blank\">https://arxiv.org/pdf/2508.10104</a>\n<a href=\"https://huggingface.co/collections/facebook/dinov3\" target=\"_blank\">https://huggingface.co/collections/facebook/dinov3</a>\n<a href=\"https://www.kaggle.com/models/keras/dinov3\" target=\"_blank\">https://www.kaggle.com/models/keras/dinov3</a></p>",
              "rawMarkdown": "Do you allow Dinov3 models?\nhttps://arxiv.org/pdf/2508.10104\nhttps://huggingface.co/collections/facebook/dinov3\nhttps://www.kaggle.com/models/keras/dinov3"
            },
            {
              "id": 3388664,
              "postDate": "2026-01-09T10:12:33.200Z",
              "content": "<p>No problem with using DinoV3. You are allowed to use it.</p>",
              "rawMarkdown": "No problem with using DinoV3. You are allowed to use it.",
              "votes": 2
            },
            {
              "id": 3388696,
              "postDate": "2026-01-09T11:53:54.520Z",
              "content": "<p>Even if you cannot build a CC BY 4.0 project on top of dinov3 license?</p>",
              "rawMarkdown": "Even if you cannot build a CC BY 4.0 project on top of dinov3 license?"
            },
            {
              "id": 3388701,
              "postDate": "2026-01-09T11:58:34.933Z",
              "content": "<p>What about YOLOv11?</p>",
              "rawMarkdown": "What about YOLOv11?"
            },
            {
              "id": 3388819,
              "postDate": "2026-01-09T16:43:58.240Z",
              "content": "<p>Yes, you are also allowed to use YOLOv11.</p>\n<p>But, unfortunately, I won't be able to confirm every model individually here.</p>\n<p>As a rule of thumb, check if other participants would be able to access your external data and code to reproduce your results after the competition is finished.</p>\n<p>Hope that helps!</p>",
              "rawMarkdown": "Yes, you are also allowed to use YOLOv11.\n\nBut, unfortunately, I won't be able to confirm every model individually here.\n\nAs a rule of thumb, check if other participants would be able to access your external data and code to reproduce your results after the competition is finished.\n\nHope that helps!"
            },
            {
              "id": 3418516,
              "postDate": "2026-03-08T11:04:58.467Z",
              "content": "<p>Hello <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> ! I was wondering if you would be interested in being a featured author/data scientist on our company's website? </p>\n<p>Couldnt find a way to reach you personally</p>",
              "rawMarkdown": "Hello @wowfattie ! I was wondering if you would be interested in being a featured author/data scientist on our company's website? \n\nCouldnt find a way to reach you personally"
            }
          ]
        }
      ]
    },
    {
      "id": 3383001,
      "postDate": "2025-12-29T05:12:31.097Z",
      "content": "<p>Question:</p>\n<blockquote>\n  <p>A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.</p>\n  <p>A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.</p>\n</blockquote>\n<p>will test set in the forecasting phase include all (or a proportion) of the images from the current LB test set?</p>",
      "rawMarkdown": "Question:\n>A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n\n>A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.\n\nwill test set in the forecasting phase include all (or a proportion) of the images from the current LB test set?",
      "replies": [
        {
          "id": 3383255,
          "postDate": "2025-12-29T17:38:34.780Z",
          "content": "<p>Unfortunately, we cannot disclose more information regarding the forecasting test set at this stage.</p>",
          "rawMarkdown": "Unfortunately, we cannot disclose more information regarding the forecasting test set at this stage."
        }
      ]
    },
    {
      "id": 3383915,
      "postDate": "2025-12-31T05:26:15.410Z",
      "content": "<p>Hello,</p>\n<p>In the test set, should we expect that every manipulated image always contains a pair of duplicated regions (i.e., two masks corresponding to a duplication / copy-move operation)?</p>\n<p>Or can the test data also include manipulations without a corresponding duplicated region, such as cleaning, retouching, or other single-region edits, where no matching counterpart exists?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hello,\n\nIn the test set, should we expect that every manipulated image always contains a pair of duplicated regions (i.e., two masks corresponding to a duplication / copy-move operation)?\n\nOr can the test data also include manipulations without a corresponding duplicated region, such as cleaning, retouching, or other single-region edits, where no matching counterpart exists?\n\nThanks!",
      "isDeleted": true,
      "replies": [
        {
          "id": 3384147,
          "postDate": "2025-12-31T15:37:25.603Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/houstonhaozhang\" target=\"_blank\">@houstonhaozhang</a>,</p>\n<p>In this competition, we focus exclusively on copy-move forgeries. Single-region edits are not considered.</p>\n<p>But, be aware that cleaning (object removal) can still be performed using copy-move (by copying background over an object). Check this <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/614069\" target=\"_blank\">post</a></p>\n<p>So, expect to find duplicated regions in the masks for both the training and test sets.</p>\n<p>Also, note that a single source region can be cloned multiple times, meaning a single object (e.g., a cell) can be copied to several different locations within the same image.</p>\n<p>Hope that help :)</p>",
          "rawMarkdown": "Hi @houstonhaozhang,\n\nIn this competition, we focus exclusively on copy-move forgeries. Single-region edits are not considered.\n\nBut, be aware that cleaning (object removal) can still be performed using copy-move (by copying background over an object). Check this [post](https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/614069)\n\nSo, expect to find duplicated regions in the masks for both the training and test sets.\n\nAlso, note that a single source region can be cloned multiple times, meaning a single object (e.g., a cell) can be copied to several different locations within the same image.\n\nHope that help :)",
          "replies": [
            {
              "id": 3384867,
              "postDate": "2026-01-02T06:39:54.533Z",
              "content": "<p>Thanks for your information!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F31346843%2F3bc602093125a93ff35d0bfe9c0fa765%2FSnipaste_2026-01-02_01-37-02.jpg?generation=1767335858312330&amp;alt=media\" alt=\"\">\nI would like to clarify whether there may be a labeling issue in this sample. ↑\nThanks.</p>",
              "rawMarkdown": "Thanks for your information!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F31346843%2F3bc602093125a93ff35d0bfe9c0fa765%2FSnipaste_2026-01-02_01-37-02.jpg?generation=1767335858312330&alt=media)\nI would like to clarify whether there may be a labeling issue in this sample. ↑\nThanks.",
              "isDeleted": true
            },
            {
              "id": 3384999,
              "postDate": "2026-01-02T12:09:22.677Z",
              "content": "<p>These cells were selected together, and in this case, one of them was placed outside the image boundaries, which is a very rare anomaly.</p>\n<p>Check the discussion on this <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613694\" target=\"_blank\">topic</a> for more details.</p>",
              "rawMarkdown": "These cells were selected together, and in this case, one of them was placed outside the image boundaries, which is a very rare anomaly.\n\nCheck the discussion on this [topic](https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613694) for more details."
            },
            {
              "id": 3385278,
              "postDate": "2026-01-02T22:32:47.130Z",
              "content": "<p>Got it. Thanks for the information!</p>",
              "rawMarkdown": "Got it. Thanks for the information!",
              "isDeleted": true
            },
            {
              "id": 3387354,
              "postDate": "2026-01-06T19:24:35.727Z",
              "content": "<p>Hi João,</p>\n<p>From the training set masks, I found that most copy–paste regions appear to involve mainly translation and slight rotation. I did not see clear evidence of other common operations, such as scaling, flipping, edge smoothing, or brightness/contrast adjustment.</p>\n<p>Could you please clarify whether the private leaderboard test set follows a similar manipulation pattern?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F31346843%2F600e46dbb812c669852091df24981917%2FSnipaste_2026-01-06_14-23-42.jpg?generation=1767727441841159&amp;alt=media\" alt=\"\">\\</p>\n<p>Thanks.</p>",
              "rawMarkdown": "Hi João,\n\nFrom the training set masks, I found that most copy–paste regions appear to involve mainly translation and slight rotation. I did not see clear evidence of other common operations, such as scaling, flipping, edge smoothing, or brightness/contrast adjustment.\n\nCould you please clarify whether the private leaderboard test set follows a similar manipulation pattern?![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F31346843%2F600e46dbb812c669852091df24981917%2FSnipaste_2026-01-06_14-23-42.jpg?generation=1767727441841159&alt=media)\\\n\nThanks.",
              "isDeleted": true
            },
            {
              "id": 3387401,
              "postDate": "2026-01-06T23:38:05.167Z",
              "content": "<p>Hi,</p>\n<p>All the transformations you mentioned are common in scientific image manipulation. So, you should expect them to be present in the test set.</p>\n<p>However, we cannot disclose the proportions of these operations in the test set.</p>",
              "rawMarkdown": "Hi,\n\nAll the transformations you mentioned are common in scientific image manipulation. So, you should expect them to be present in the test set.\n\nHowever, we cannot disclose the proportions of these operations in the test set."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3381936,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2025-12-26T02:01:44.900000",
      "content": "<p>Question:</p>\n<blockquote>\n  <p>A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n  A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.</p>\n</blockquote>\n<p>Does this mean during the forecasting phase (private leaderboard), we should expect roughly 2,200 images total? Just want to confirm. So if our submission works successfully on current LB (~1,100 images) in &lt; 4 hours, then it would likely work in private leaderboard which would have a runtime of 9 hours.</p>\n<p>thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3382069,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2025-12-26T11:29:01.503000",
          "content": "<p>Yes, expect the final forecasting test set to be roughly 2,200 images.</p>\n<p>If your model currently respects the 4-hour runtime for the current set, you should be well within the time limit for the forecasting phase.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3366114,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2025-12-07T14:44:46.613000",
      "content": "<p>Will all images be png in test phase? Or should we account for non-png. I am asking because I have hardcoded .png file extension in my code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3369143,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2025-12-09T18:55:33.767000",
          "content": "<p>Yes, you should account for non-PNG images.</p>\n<p>Although most images are expected to be in PNG format, it is strongly recommended that you do not hardcode this. It is safer to make your code flexible enough to handle JPEG and other image formats as well.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3391889,
      "author_name": "Raphael Delattre",
      "author_url": "",
      "post_date": "2026-01-15T22:01:49.133000",
      "content": "<p>After thorough examination, the images must be uploaded immediately to a secure remote server. One way to detect manipulation is to copy and rotate the object and check for repeated patterns. However, if the object was sourced from another image that I cannot access, that approach is not possible. And if pixels are added one by one using consistent RGB values, the manipulation becomes even harder to detect.</p>\n<p>On top of that, it is impossible to determine whether what we are seeing is physically plausible or biologically plausible. Living cells can look very similar, yet they are also unique. I do not see a fully reliable solution, and if I build a detection pipeline using an existing dataset to avoid these issues, I am concerned it will overfit to dataset-specific artifacts and end up flagging legitimate images as “fake.”</p>\n<p>The only practical alternative would be to detect some known ways of producing fake data, but that could give humanity a false belief in “truth”: any methods that remain undetected would effectively go unpunished, creating a misleading sense of security.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3391892,
          "author_name": "Raphael Delattre",
          "author_url": "",
          "post_date": "2026-01-15T22:06:30.233000",
          "content": "<p>And I would much prefer a similar identification challenge on a dataset where we can say, with certainty, that everything is true—i.e., where the image acquisition methodology is fully documented (staining/coloration protocols, camera models and settings, illumination and spectral characteristics, optics, exposure, calibration, preprocessing, and any post-processing), along with clear ground-truth labels and provenance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3391770,
      "author_name": "weke",
      "author_url": "",
      "post_date": "2026-01-15T16:35:40.093000",
      "content": "<p>will the training images folder be available in the testing phase?\nMy submission uses them to train a model on them and I need to know if they are available or i need to upload it to kaggle. <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3391853,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2026-01-15T20:00:11.090000",
          "content": "<p>Yes, the training images will be available during the forecasting phase.</p>\n<p>The same notebook you submitted should work during that phase.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3388278,
      "author_name": "moritake04",
      "author_url": "",
      "post_date": "2026-01-08T14:46:25.083000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a>, I have a question.\nWhen using external data or pre-trained models, do we need to care about licenses, such as commercial use? Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3388337,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2026-01-08T16:41:18.617000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/moritake04\" target=\"_blank\">@moritake04</a>,</p>\n<p>Yes, you do need to be careful with licenses.</p>\n<p>According to the competition rules (rule 5), the winning solution must be open-sourced. If your solution depends on a pre-trained model or data with a restrictive license (one that prevents you from legally open-sourcing your code), you would likely be unable to fulfill this requirement.</p>\n<p>According to Rule 6, you must ensure that any external data you use is publicly available and equally accessible to all participants at no cost.</p>\n<p>Essentially, other participants should be able to access the external data and code to reproduce your results after the competition is finished.</p>\n<p>Please check Rules 5 and 6 here: <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/rules\" target=\"_blank\">https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/rules</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3388486,
              "author_name": "Guanshuo Xu",
              "author_url": "",
              "post_date": "2026-01-09T00:29:55.287000",
              "content": "<p>Do you allow Dinov3 models?\n<a href=\"https://arxiv.org/pdf/2508.10104\" target=\"_blank\">https://arxiv.org/pdf/2508.10104</a>\n<a href=\"https://huggingface.co/collections/facebook/dinov3\" target=\"_blank\">https://huggingface.co/collections/facebook/dinov3</a>\n<a href=\"https://www.kaggle.com/models/keras/dinov3\" target=\"_blank\">https://www.kaggle.com/models/keras/dinov3</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3388664,
              "author_name": "João Phillipe Cardenuto",
              "author_url": "",
              "post_date": "2026-01-09T10:12:33.200000",
              "content": "<p>No problem with using DinoV3. You are allowed to use it.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3388696,
              "author_name": "weke",
              "author_url": "",
              "post_date": "2026-01-09T11:53:54.520000",
              "content": "<p>Even if you cannot build a CC BY 4.0 project on top of dinov3 license?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3388701,
              "author_name": "weke",
              "author_url": "",
              "post_date": "2026-01-09T11:58:34.933000",
              "content": "<p>What about YOLOv11?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3388819,
              "author_name": "João Phillipe Cardenuto",
              "author_url": "",
              "post_date": "2026-01-09T16:43:58.240000",
              "content": "<p>Yes, you are also allowed to use YOLOv11.</p>\n<p>But, unfortunately, I won't be able to confirm every model individually here.</p>\n<p>As a rule of thumb, check if other participants would be able to access your external data and code to reproduce your results after the competition is finished.</p>\n<p>Hope that helps!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3418516,
              "author_name": "Brainydaps",
              "author_url": "",
              "post_date": "2026-03-08T11:04:58.467000",
              "content": "<p>Hello <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> ! I was wondering if you would be interested in being a featured author/data scientist on our company's website? </p>\n<p>Couldnt find a way to reach you personally</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3383001,
      "author_name": "Kh0a",
      "author_url": "",
      "post_date": "2025-12-29T05:12:31.097000",
      "content": "<p>Question:</p>\n<blockquote>\n  <p>A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.</p>\n  <p>A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.</p>\n</blockquote>\n<p>will test set in the forecasting phase include all (or a proportion) of the images from the current LB test set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3383255,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2025-12-29T17:38:34.780000",
          "content": "<p>Unfortunately, we cannot disclose more information regarding the forecasting test set at this stage.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3383915,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-12-31T05:26:15.410000",
      "content": "<p>Hello,</p>\n<p>In the test set, should we expect that every manipulated image always contains a pair of duplicated regions (i.e., two masks corresponding to a duplication / copy-move operation)?</p>\n<p>Or can the test data also include manipulations without a corresponding duplicated region, such as cleaning, retouching, or other single-region edits, where no matching counterpart exists?</p>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3384147,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2025-12-31T15:37:25.603000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/houstonhaozhang\" target=\"_blank\">@houstonhaozhang</a>,</p>\n<p>In this competition, we focus exclusively on copy-move forgeries. Single-region edits are not considered.</p>\n<p>But, be aware that cleaning (object removal) can still be performed using copy-move (by copying background over an object). Check this <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/614069\" target=\"_blank\">post</a></p>\n<p>So, expect to find duplicated regions in the masks for both the training and test sets.</p>\n<p>Also, note that a single source region can be cloned multiple times, meaning a single object (e.g., a cell) can be copied to several different locations within the same image.</p>\n<p>Hope that help :)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3384867,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-01-02T06:39:54.533000",
              "content": "<p>Thanks for your information!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F31346843%2F3bc602093125a93ff35d0bfe9c0fa765%2FSnipaste_2026-01-02_01-37-02.jpg?generation=1767335858312330&amp;alt=media\" alt=\"\">\nI would like to clarify whether there may be a labeling issue in this sample. ↑\nThanks.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3384999,
              "author_name": "João Phillipe Cardenuto",
              "author_url": "",
              "post_date": "2026-01-02T12:09:22.677000",
              "content": "<p>These cells were selected together, and in this case, one of them was placed outside the image boundaries, which is a very rare anomaly.</p>\n<p>Check the discussion on this <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613694\" target=\"_blank\">topic</a> for more details.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3385278,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-01-02T22:32:47.130000",
              "content": "<p>Got it. Thanks for the information!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3387354,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-01-06T19:24:35.727000",
              "content": "<p>Hi João,</p>\n<p>From the training set masks, I found that most copy–paste regions appear to involve mainly translation and slight rotation. I did not see clear evidence of other common operations, such as scaling, flipping, edge smoothing, or brightness/contrast adjustment.</p>\n<p>Could you please clarify whether the private leaderboard test set follows a similar manipulation pattern?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F31346843%2F600e46dbb812c669852091df24981917%2FSnipaste_2026-01-06_14-23-42.jpg?generation=1767727441841159&amp;alt=media\" alt=\"\">\\</p>\n<p>Thanks.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3387401,
              "author_name": "João Phillipe Cardenuto",
              "author_url": "",
              "post_date": "2026-01-06T23:38:05.167000",
              "content": "<p>Hi,</p>\n<p>All the transformations you mentioned are common in scientific image manipulation. So, you should expect them to be present in the test set.</p>\n<p>However, we cannot disclose the proportions of these operations in the test set.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3306014": "Hello everyone, and a warm welcome to the competition!\n\nThis is a milestone we've been working on for a long time. Having been deeply involved in the problem of scientific image integrity, I know firsthand how critical this issue is. We are incredibly excited to host this challenge and bring together a community of brilliant minds to tackle it.\n\nThis is more than just a competition; we see it as a foundational step. Our hope is to build a lasting community dedicated to developing computational solutions for research integrity. We truly believe that the models and ideas generated here can have a significant, positive impact.\n\nI am personally very excited to see the novel approaches and creative solutions you will all design.\n\nPlease use this forum as your main hub. Ask questions and connect with fellow participants. We are here to support you.\n\nGood luck, and thank you for joining this important mission!",
    "3381936": "Question:\n\n> A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n> A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.\n\nDoes this mean during the forecasting phase (private leaderboard), we should expect roughly 2,200 images total? Just want to confirm. So if our submission works successfully on current LB (~1,100 images) in < 4 hours, then it would likely work in private leaderboard which would have a runtime of 9 hours.\n\nthanks\n",
    "3366114": "Will all images be png in test phase? Or should we account for non-png. I am asking because I have hardcoded .png file extension in my code.",
    "3391889": "After thorough examination, the images must be uploaded immediately to a secure remote server. One way to detect manipulation is to copy and rotate the object and check for repeated patterns. However, if the object was sourced from another image that I cannot access, that approach is not possible. And if pixels are added one by one using consistent RGB values, the manipulation becomes even harder to detect.\n\nOn top of that, it is impossible to determine whether what we are seeing is physically plausible or biologically plausible. Living cells can look very similar, yet they are also unique. I do not see a fully reliable solution, and if I build a detection pipeline using an existing dataset to avoid these issues, I am concerned it will overfit to dataset-specific artifacts and end up flagging legitimate images as “fake.”\n\nThe only practical alternative would be to detect some known ways of producing fake data, but that could give humanity a false belief in “truth”: any methods that remain undetected would effectively go unpunished, creating a misleading sense of security.",
    "3391770": "will the training images folder be available in the testing phase?\nMy submission uses them to train a model on them and I need to know if they are available or i need to upload it to kaggle. @joophillipecardenuto ",
    "3388278": "Hi @joophillipecardenuto, I have a question.\nWhen using external data or pre-trained models, do we need to care about licenses, such as commercial use? Thank you.",
    "3383001": "Question:\n>A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n\n>A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.\n\nwill test set in the forecasting phase include all (or a proportion) of the images from the current LB test set?",
    "3383915": "Hello,\n\nIn the test set, should we expect that every manipulated image always contains a pair of duplicated regions (i.e., two masks corresponding to a duplication / copy-move operation)?\n\nOr can the test data also include manipulations without a corresponding duplicated region, such as cleaning, retouching, or other single-region edits, where no matching counterpart exists?\n\nThanks!"
  }
}