{
  "id": 657899,
  "title": "Huge differences between train and supplemental images",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/657899",
  "author_name": "",
  "post_date": "2025-12-10T20:35:40.201279900Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Upon reviewing the data, I identified two significant problems: </p>\n<ul>\n<li>Data distribution shift. The training and supplemental datasets appear to come from different sources. Training images are standalone, biology-related images, while supplemental images appear to be screenshots from academic papers containing multiple sub-images. These include text, bar charts, arrows, and other elements that could themselves be flagged as manipulations since they are overlaid on the original images.</li>\n<li>Task shift. The two datasets imply fundamentally different tasks. Training images suggest a forgery detection task: identifying where someone has copy-pasted a non-rectangular object (e.g., a single cell or a banana :))), leaving traces such as noise patterns, compression artifacts, or rough edges. The (kind of) pixel-perfect masks tell this as a segmentation problem with an element of matching. Supplemental images, on the other hand, contain large rectangular masks with multiple \"objects\". They appear to involve blind copy-paste of entire images or large regions rather than discrete objects. This implies a duplicate matching task where noise analysis is irrelevant. Also I found no non-rectangular masks in the supplemental set.</li>\n</ul>\n<p>These observations raise two questions:</p>\n<ul>\n<li>What is the actual task? If the test data follows the supplemental distribution, the training data may be insufficient or misaligned for effective model training. Should we expect rectangular masks? </li>\n<li>Which distribution does the leaderboard test set follow? My early experiments suggest the public test set mimics the training data, but does this hold for the private test set?</li>\n</ul>",
  "messages": [
    {
      "id": "3370480",
      "postDate": "12/10/2025 20:35:40",
      "content": "<p>Upon reviewing the data, I identified two significant problems: </p>\n<ul>\n<li>Data distribution shift. The training and supplemental datasets appear to come from different sources. Training images are standalone, biology-related images, while supplemental images appear to be screenshots from academic papers containing multiple sub-images. These include text, bar charts, arrows, and other elements that could themselves be flagged as manipulations since they are overlaid on the original images.</li>\n<li>Task shift. The two datasets imply fundamentally different tasks. Training images suggest a forgery detection task: identifying where someone has copy-pasted a non-rectangular object (e.g., a single cell or a banana :))), leaving traces such as noise patterns, compression artifacts, or rough edges. The (kind of) pixel-perfect masks tell this as a segmentation problem with an element of matching. Supplemental images, on the other hand, contain large rectangular masks with multiple \"objects\". They appear to involve blind copy-paste of entire images or large regions rather than discrete objects. This implies a duplicate matching task where noise analysis is irrelevant. Also I found no non-rectangular masks in the supplemental set.</li>\n</ul>\n<p>These observations raise two questions:</p>\n<ul>\n<li>What is the actual task? If the test data follows the supplemental distribution, the training data may be insufficient or misaligned for effective model training. Should we expect rectangular masks? </li>\n<li>Which distribution does the leaderboard test set follow? My early experiments suggest the public test set mimics the training data, but does this hold for the private test set?</li>\n</ul>",
      "rawMarkdown": "Upon reviewing the data, I identified two significant problems: \n\n- Data distribution shift. The training and supplemental datasets appear to come from different sources. Training images are standalone, biology-related images, while supplemental images appear to be screenshots from academic papers containing multiple sub-images. These include text, bar charts, arrows, and other elements that could themselves be flagged as manipulations since they are overlaid on the original images.\n- Task shift. The two datasets imply fundamentally different tasks. Training images suggest a forgery detection task: identifying where someone has copy-pasted a non-rectangular object (e.g., a single cell or a banana :))), leaving traces such as noise patterns, compression artifacts, or rough edges. The (kind of) pixel-perfect masks tell this as a segmentation problem with an element of matching. Supplemental images, on the other hand, contain large rectangular masks with multiple \"objects\". They appear to involve blind copy-paste of entire images or large regions rather than discrete objects. This implies a duplicate matching task where noise analysis is irrelevant. Also I found no non-rectangular masks in the supplemental set.\n\nThese observations raise two questions:\n- What is the actual task? If the test data follows the supplemental distribution, the training data may be insufficient or misaligned for effective model training. Should we expect rectangular masks? \n- Which distribution does the leaderboard test set follow? My early experiments suggest the public test set mimics the training data, but does this hold for the private test set?",
      "votes": null
    },
    {
      "id": "3372608",
      "postDate": "12/12/2025 05:28:24",
      "content": "<p>Exactly! and this is why models trained on the training data which they gave us is not performing good on the private dataset they that have put up for scoring our models, its because of this distribution shift that we have no idea what it is in their private dataset, so we cant train a model on this without knowing about it, this is why the entire leaderboard is stuck around the 0.303 baseline, (with the top player being at 0.349, only 0.046 points ahead of the 0.303 baseline) this is because you get a score of 0.303 on their private testing data if you simply pass all of your predictions as \"AUTHENTIC\" so this is the all authentic baseline and nobody has figured out a solution to this problem yet because they didnt give us enough data to train our models on the type of data which they will be testing us on.</p>",
      "rawMarkdown": "Exactly! and this is why models trained on the training data which they gave us is not performing good on the private dataset they that have put up for scoring our models, its because of this distribution shift that we have no idea what it is in their private dataset, so we cant train a model on this without knowing about it, this is why the entire leaderboard is stuck around the 0.303 baseline, (with the top player being at 0.349, only 0.046 points ahead of the 0.303 baseline) this is because you get a score of 0.303 on their private testing data if you simply pass all of your predictions as \"AUTHENTIC\" so this is the all authentic baseline and nobody has figured out a solution to this problem yet because they didnt give us enough data to train our models on the type of data which they will be testing us on.",
      "votes": null
    },
    {
      "id": "3380545",
      "postDate": "12/22/2025 15:50:20",
      "content": "<p>I bet this competition is some sort of guessing game. All we know there are no corn or bananas in the test dataset, it's probably multi-panel images from real scientific forgeries</p>",
      "rawMarkdown": "I bet this competition is some sort of guessing game. All we know there are no corn or bananas in the test dataset, it's probably multi-panel images from real scientific forgeries",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3372608,
      "author_name": "mrearthworm",
      "author_url": "",
      "post_date": "12/12/2025 05:28:24",
      "content": "<p>Exactly! and this is why models trained on the training data which they gave us is not performing good on the private dataset they that have put up for scoring our models, its because of this distribution shift that we have no idea what it is in their private dataset, so we cant train a model on this without knowing about it, this is why the entire leaderboard is stuck around the 0.303 baseline, (with the top player being at 0.349, only 0.046 points ahead of the 0.303 baseline) this is because you get a score of 0.303 on their private testing data if you simply pass all of your predictions as \"AUTHENTIC\" so this is the all authentic baseline and nobody has figured out a solution to this problem yet because they didnt give us enough data to train our models on the type of data which they will be testing us on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3380545,
      "author_name": "theodorlu",
      "author_url": "",
      "post_date": "12/22/2025 15:50:20",
      "content": "<p>I bet this competition is some sort of guessing game. All we know there are no corn or bananas in the test dataset, it's probably multi-panel images from real scientific forgeries</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3370480": "Upon reviewing the data, I identified two significant problems: \n\n- Data distribution shift. The training and supplemental datasets appear to come from different sources. Training images are standalone, biology-related images, while supplemental images appear to be screenshots from academic papers containing multiple sub-images. These include text, bar charts, arrows, and other elements that could themselves be flagged as manipulations since they are overlaid on the original images.\n- Task shift. The two datasets imply fundamentally different tasks. Training images suggest a forgery detection task: identifying where someone has copy-pasted a non-rectangular object (e.g., a single cell or a banana :))), leaving traces such as noise patterns, compression artifacts, or rough edges. The (kind of) pixel-perfect masks tell this as a segmentation problem with an element of matching. Supplemental images, on the other hand, contain large rectangular masks with multiple \"objects\". They appear to involve blind copy-paste of entire images or large regions rather than discrete objects. This implies a duplicate matching task where noise analysis is irrelevant. Also I found no non-rectangular masks in the supplemental set.\n\nThese observations raise two questions:\n- What is the actual task? If the test data follows the supplemental distribution, the training data may be insufficient or misaligned for effective model training. Should we expect rectangular masks? \n- Which distribution does the leaderboard test set follow? My early experiments suggest the public test set mimics the training data, but does this hold for the private test set?",
    "3372608": "Exactly! and this is why models trained on the training data which they gave us is not performing good on the private dataset they that have put up for scoring our models, its because of this distribution shift that we have no idea what it is in their private dataset, so we cant train a model on this without knowing about it, this is why the entire leaderboard is stuck around the 0.303 baseline, (with the top player being at 0.349, only 0.046 points ahead of the 0.303 baseline) this is because you get a score of 0.303 on their private testing data if you simply pass all of your predictions as \"AUTHENTIC\" so this is the all authentic baseline and nobody has figured out a solution to this problem yet because they didnt give us enough data to train our models on the type of data which they will be testing us on.",
    "3380545": "I bet this competition is some sort of guessing game. All we know there are no corn or bananas in the test dataset, it's probably multi-panel images from real scientific forgeries"
  },
  "source": "meta"
}