{
  "id": 613586,
  "title": "The \"research papers leaderboard scores during this phase are not meaningful.\"",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613586",
  "author_name": "",
  "post_date": "2025-10-28T05:48:38.828253100Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Dear Admins,</p>\n<p>Since the leaderboard scores for are explicitly stated as non-indicative during this phase, I'm unclear on the rationale for encouraging submissions right now, especially when the scoring mechanism offers no valuable insights or feedback. It strikes me as counterproductive to invest effort in training models and submitting blindly under these conditions. </p>",
  "messages": [
    {
      "id": "3307916",
      "postDate": "10/28/2025 05:48:38",
      "content": "<p>Dear Admins,</p>\n<p>Since the leaderboard scores for are explicitly stated as non-indicative during this phase, I'm unclear on the rationale for encouraging submissions right now, especially when the scoring mechanism offers no valuable insights or feedback. It strikes me as counterproductive to invest effort in training models and submitting blindly under these conditions. </p>",
      "rawMarkdown": "Dear Admins,\n\nSince the leaderboard scores for are explicitly stated as non-indicative during this phase, I'm unclear on the rationale for encouraging submissions right now, especially when the scoring mechanism offers no valuable insights or feedback. It strikes me as counterproductive to invest effort in training models and submitting blindly under these conditions.",
      "votes": null
    },
    {
      "id": "3308453",
      "postDate": "10/29/2025 13:08:51",
      "content": "<p>Thanks for sharing your concern.</p>\n<p>This competition has two phases to ensure a fair and robust evaluation. The final winners will be decided in Phase 2.</p>\n<p>Phase 1 (Training)\nThe current leaderboard is based on a test set of images from real copy-move forgery cases collected from publicly available research papers. This phase is intended for developing and tuning your models.</p>\n<p>Phase 2 (Forecasting - Final Validation)\nTo find the best solution, Phase 2 will add new images to the test set. These new images will be built only from scientific images published after Phase 1 has closed. This ensures that there is no possible data leakage; since we are dealing with public papers (in Phase 1), this approach guarantees the new data is unseen by all participants.</p>\n<p>Important: Your model will only be eligible for Phase 2 validation if you successfully submit it as a notebook during Phase 1 😉</p>",
      "rawMarkdown": "Thanks for sharing your concern.\n\nThis competition has two phases to ensure a fair and robust evaluation. The final winners will be decided in Phase 2.\n\nPhase 1 (Training)\nThe current leaderboard is based on a test set of images from real copy-move forgery cases collected from publicly available research papers. This phase is intended for developing and tuning your models.\n\nPhase 2 (Forecasting - Final Validation)\nTo find the best solution, Phase 2 will add new images to the test set. These new images will be built only from scientific images published after Phase 1 has closed. This ensures that there is no possible data leakage; since we are dealing with public papers (in Phase 1), this approach guarantees the new data is unseen by all participants.\n\nImportant: Your model will only be eligible for Phase 2 validation if you successfully submit it as a notebook during Phase 1 😉",
      "votes": null
    },
    {
      "id": "3308521",
      "postDate": "10/29/2025 15:38:07",
      "content": "<p>\"This phase is intended for developing and tuning your models.\" Once again … no real tuning is possible. </p>",
      "rawMarkdown": "\"This phase is intended for developing and tuning your models.\" Once again ... no real tuning is possible.",
      "votes": null
    },
    {
      "id": "3336911",
      "postDate": "11/18/2025 17:31:10",
      "content": "<p>When will the second phase start?</p>",
      "rawMarkdown": "When will the second phase start?",
      "votes": null
    },
    {
      "id": "3339955",
      "postDate": "11/19/2025 10:22:08",
      "content": "<p>The second phase starts after January 15th.</p>\n<p>Please note that during the second phase, participants will not be able to submit models or solutions.</p>\n<p>During this phase, the organizers will collect new cases of scientific forgery that have not been reported before January 15th.</p>",
      "rawMarkdown": "The second phase starts after January 15th.\n\nPlease note that during the second phase, participants will not be able to submit models or solutions.\n\nDuring this phase, the organizers will collect new cases of scientific forgery that have not been reported before January 15th.",
      "votes": null
    },
    {
      "id": "3342125",
      "postDate": "11/20/2025 16:45:03",
      "content": "<p><a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> \nhi, I have a question.\nRegarding the images in the dataset to be added in Phase 2, will they be biology-related images similar to those currently provided?</p>",
      "rawMarkdown": "joophillipecardenuto \nhi, I have a question.\nRegarding the images in the dataset to be added in Phase 2, will they be biology-related images similar to those currently provided?",
      "votes": null
    },
    {
      "id": "3359132",
      "postDate": "12/01/2025 17:17:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/npinpi\" target=\"_blank\">@npinpi</a>. Phase 2 consists entirely of biomedical research images, with the same types of forgeries as in Phase 1. You can expect images similar to those provided as supplemental training data. However, please be aware that we will collect these images only after the first phase of the competition has finished.</p>",
      "rawMarkdown": "Hi @npinpi. Phase 2 consists entirely of biomedical research images, with the same types of forgeries as in Phase 1. You can expect images similar to those provided as supplemental training data. However, please be aware that we will collect these images only after the first phase of the competition has finished.",
      "votes": null
    },
    {
      "id": "3359151",
      "postDate": "12/01/2025 17:21:45",
      "content": "<p>got it.\nThank you for providing information.</p>\n<p>I'm surprised that so many forgery images were gathered in such a short time… 😂</p>",
      "rawMarkdown": "got it.\nThank you for providing information.\n\nI'm surprised that so many forgery images were gathered in such a short time… 😂",
      "votes": null
    },
    {
      "id": "3388505",
      "postDate": "01/09/2026 02:23:50",
      "content": "<p>Thanks for the explanation. I have a question regarding the statement: \"Phase 2 will add new images to the test set.\"</p>\n<p>Does this imply that the final evaluation dataset will consist of both the current Phase 1 (Public) images AND the newly added images?</p>\n<p>If the Phase 1 images are included in the final scoring, wouldn't this raise fairness concerns? Since Phase 1 data comes from publicly available papers, participants might overfit their models to the current public leaderboard data. Could you please clarify if the final ranking will rely solely on the new unseen images, or a combination of both?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Thanks for the explanation. I have a question regarding the statement: \"Phase 2 will add new images to the test set.\"\n\nDoes this imply that the final evaluation dataset will consist of both the current Phase 1 (Public) images AND the newly added images?\n\nIf the Phase 1 images are included in the final scoring, wouldn't this raise fairness concerns? Since Phase 1 data comes from publicly available papers, participants might overfit their models to the current public leaderboard data. Could you please clarify if the final ranking will rely solely on the new unseen images, or a combination of both?\n\nThanks!",
      "votes": null
    },
    {
      "id": "3388532",
      "postDate": "01/09/2026 03:32:46",
      "content": "<p>i asked a similar question like this to the host, you can find it <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613052#3383001\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "i asked a similar question like this to the host, you can find it [here](https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613052#3383001)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3308453,
      "author_name": "joophillipecardenuto",
      "author_url": "",
      "post_date": "10/29/2025 13:08:51",
      "content": "<p>Thanks for sharing your concern.</p>\n<p>This competition has two phases to ensure a fair and robust evaluation. The final winners will be decided in Phase 2.</p>\n<p>Phase 1 (Training)\nThe current leaderboard is based on a test set of images from real copy-move forgery cases collected from publicly available research papers. This phase is intended for developing and tuning your models.</p>\n<p>Phase 2 (Forecasting - Final Validation)\nTo find the best solution, Phase 2 will add new images to the test set. These new images will be built only from scientific images published after Phase 1 has closed. This ensures that there is no possible data leakage; since we are dealing with public papers (in Phase 1), this approach guarantees the new data is unseen by all participants.</p>\n<p>Important: Your model will only be eligible for Phase 2 validation if you successfully submit it as a notebook during Phase 1 😉</p>",
      "votes": null,
      "replies": [
        {
          "id": 3308521,
          "author_name": "solomonk",
          "author_url": "",
          "post_date": "10/29/2025 15:38:07",
          "content": "<p>\"This phase is intended for developing and tuning your models.\" Once again … no real tuning is possible. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3336911,
          "author_name": "gesila",
          "author_url": "",
          "post_date": "11/18/2025 17:31:10",
          "content": "<p>When will the second phase start?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3339955,
              "author_name": "joophillipecardenuto",
              "author_url": "",
              "post_date": "11/19/2025 10:22:08",
              "content": "<p>The second phase starts after January 15th.</p>\n<p>Please note that during the second phase, participants will not be able to submit models or solutions.</p>\n<p>During this phase, the organizers will collect new cases of scientific forgery that have not been reported before January 15th.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3342125,
          "author_name": "npinpi",
          "author_url": "",
          "post_date": "11/20/2025 16:45:03",
          "content": "<p><a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> \nhi, I have a question.\nRegarding the images in the dataset to be added in Phase 2, will they be biology-related images similar to those currently provided?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3359132,
              "author_name": "joophillipecardenuto",
              "author_url": "",
              "post_date": "12/01/2025 17:17:48",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/npinpi\" target=\"_blank\">@npinpi</a>. Phase 2 consists entirely of biomedical research images, with the same types of forgeries as in Phase 1. You can expect images similar to those provided as supplemental training data. However, please be aware that we will collect these images only after the first phase of the competition has finished.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3359151,
                  "author_name": "npinpi",
                  "author_url": "",
                  "post_date": "12/01/2025 17:21:45",
                  "content": "<p>got it.\nThank you for providing information.</p>\n<p>I'm surprised that so many forgery images were gathered in such a short time… 😂</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3388505,
          "author_name": "",
          "author_url": "",
          "post_date": "01/09/2026 02:23:50",
          "content": "<p>Thanks for the explanation. I have a question regarding the statement: \"Phase 2 will add new images to the test set.\"</p>\n<p>Does this imply that the final evaluation dataset will consist of both the current Phase 1 (Public) images AND the newly added images?</p>\n<p>If the Phase 1 images are included in the final scoring, wouldn't this raise fairness concerns? Since Phase 1 data comes from publicly available papers, participants might overfit their models to the current public leaderboard data. Could you please clarify if the final ranking will rely solely on the new unseen images, or a combination of both?</p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3388532,
              "author_name": "llkh0a",
              "author_url": "",
              "post_date": "01/09/2026 03:32:46",
              "content": "<p>i asked a similar question like this to the host, you can find it <a href=\"https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613052#3383001\" target=\"_blank\">here</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3307916": "Dear Admins,\n\nSince the leaderboard scores for are explicitly stated as non-indicative during this phase, I'm unclear on the rationale for encouraging submissions right now, especially when the scoring mechanism offers no valuable insights or feedback. It strikes me as counterproductive to invest effort in training models and submitting blindly under these conditions.",
    "3308453": "Thanks for sharing your concern.\n\nThis competition has two phases to ensure a fair and robust evaluation. The final winners will be decided in Phase 2.\n\nPhase 1 (Training)\nThe current leaderboard is based on a test set of images from real copy-move forgery cases collected from publicly available research papers. This phase is intended for developing and tuning your models.\n\nPhase 2 (Forecasting - Final Validation)\nTo find the best solution, Phase 2 will add new images to the test set. These new images will be built only from scientific images published after Phase 1 has closed. This ensures that there is no possible data leakage; since we are dealing with public papers (in Phase 1), this approach guarantees the new data is unseen by all participants.\n\nImportant: Your model will only be eligible for Phase 2 validation if you successfully submit it as a notebook during Phase 1 😉",
    "3308521": "\"This phase is intended for developing and tuning your models.\" Once again ... no real tuning is possible.",
    "3336911": "When will the second phase start?",
    "3339955": "The second phase starts after January 15th.\n\nPlease note that during the second phase, participants will not be able to submit models or solutions.\n\nDuring this phase, the organizers will collect new cases of scientific forgery that have not been reported before January 15th.",
    "3342125": "joophillipecardenuto \nhi, I have a question.\nRegarding the images in the dataset to be added in Phase 2, will they be biology-related images similar to those currently provided?",
    "3359132": "Hi @npinpi. Phase 2 consists entirely of biomedical research images, with the same types of forgeries as in Phase 1. You can expect images similar to those provided as supplemental training data. However, please be aware that we will collect these images only after the first phase of the competition has finished.",
    "3359151": "got it.\nThank you for providing information.\n\nI'm surprised that so many forgery images were gathered in such a short time… 😂",
    "3388505": "Thanks for the explanation. I have a question regarding the statement: \"Phase 2 will add new images to the test set.\"\n\nDoes this imply that the final evaluation dataset will consist of both the current Phase 1 (Public) images AND the newly added images?\n\nIf the Phase 1 images are included in the final scoring, wouldn't this raise fairness concerns? Since Phase 1 data comes from publicly available papers, participants might overfit their models to the current public leaderboard data. Could you please clarify if the final ranking will rely solely on the new unseen images, or a combination of both?\n\nThanks!",
    "3388532": "i asked a similar question like this to the host, you can find it [here](https://www.kaggle.com/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613052#3383001)"
  },
  "source": "meta"
}