{
  "id": 668521,
  "title": "RSIID Dataset",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/668521",
  "author_name": "",
  "post_date": "2026-01-17T07:22:49.682882200Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://zenodo.org/records/15095089\" target=\"_blank\">https://zenodo.org/records/15095089</a></p>\n<p>It was surprising to me that this dataset didn't improve my lb score that much. I had a hunch that private test set will be mostly inter/intra-panel forgeries, so I also used this dataset for training my models. My best public lb score was 0.311 which is 0.008 better than all authentic prediction. Anyone used this dataset and get a decent boost?</p>",
  "messages": [
    {
      "id": "3392594",
      "postDate": "01/17/2026 07:22:49",
      "content": "<p><a href=\"https://zenodo.org/records/15095089\" target=\"_blank\">https://zenodo.org/records/15095089</a></p>\n<p>It was surprising to me that this dataset didn't improve my lb score that much. I had a hunch that private test set will be mostly inter/intra-panel forgeries, so I also used this dataset for training my models. My best public lb score was 0.311 which is 0.008 better than all authentic prediction. Anyone used this dataset and get a decent boost?</p>",
      "rawMarkdown": "https://zenodo.org/records/15095089\n\nIt was surprising to me that this dataset didn't improve my lb score that much. I had a hunch that private test set will be mostly inter/intra-panel forgeries, so I also used this dataset for training my models. My best public lb score was 0.311 which is 0.008 better than all authentic prediction. Anyone used this dataset and get a decent boost?",
      "votes": null
    },
    {
      "id": "3392631",
      "postDate": "01/17/2026 09:01:55",
      "content": "<p>I used it. However, the data was completely forged. As a result, the models I trained tended to consistently find forged data. I tried to achieve balance using different techniques with authentic data. But I couldn't do it. I stopped using it.</p>",
      "rawMarkdown": "I used it. However, the data was completely forged. As a result, the models I trained tended to consistently find forged data. I tried to achieve balance using different techniques with authentic data. But I couldn't do it. I stopped using it.",
      "votes": null
    },
    {
      "id": "3392640",
      "postDate": "01/17/2026 09:30:09",
      "content": "<p>Yes, but you could have created authentic versions of them programmatically. For every forged sample, I created an authentic version by simply overwriting the pristine panel inside the directory.</p>",
      "rawMarkdown": "Yes, but you could have created authentic versions of them programmatically. For every forged sample, I created an authentic version by simply overwriting the pristine panel inside the directory.",
      "votes": null
    },
    {
      "id": "3392654",
      "postDate": "01/17/2026 09:57:38",
      "content": "<p>That makes sense. I also created synthetic authentic data by combining images from the single panel. One of the issues is that the forgeries in the supplementary images are different from the forgeries in the RSIID Dataset. The RSIID Dataset consists of fully synthetic data created by researchers; it is not real data. The forgeries in the supplementary images are much more professional. For this reason, training with RSIID Dataset data may not have worked well on the final data.</p>",
      "rawMarkdown": "That makes sense. I also created synthetic authentic data by combining images from the single panel. One of the issues is that the forgeries in the supplementary images are different from the forgeries in the RSIID Dataset. The RSIID Dataset consists of fully synthetic data created by researchers; it is not real data. The forgeries in the supplementary images are much more professional. For this reason, training with RSIID Dataset data may not have worked well on the final data.",
      "votes": null
    },
    {
      "id": "3392776",
      "postDate": "01/17/2026 13:25:40",
      "content": "<p>Note some of the rsiid dataset had some identical images with trainset. So if you didnt dedupe there was leakage in your groupkfold validation. Ultimately i did not focus a lot on it but judging from what you say, i should have to get some extra points boost</p>",
      "rawMarkdown": "Note some of the rsiid dataset had some identical images with trainset. So if you didnt dedupe there was leakage in your groupkfold validation. Ultimately i did not focus a lot on it but judging from what you say, i should have to get some extra points boost",
      "votes": null
    },
    {
      "id": "3393884",
      "postDate": "01/20/2026 04:22:57",
      "content": "<p>I used this dataset and got a score of 0.314 using Dino v2+semantic segmentation head. Ultimately, the problem does not seem to be the copy-move dataset, it might be how to generate copy-move inference using this data since its not classic semantic segmentation. So I tried contrastive modeling but got worse scores (0.2-0.3).</p>",
      "rawMarkdown": "I used this dataset and got a score of 0.314 using Dino v2+semantic segmentation head. Ultimately, the problem does not seem to be the copy-move dataset, it might be how to generate copy-move inference using this data since its not classic semantic segmentation. So I tried contrastive modeling but got worse scores (0.2-0.3).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3392631,
      "author_name": "musapeker",
      "author_url": "",
      "post_date": "01/17/2026 09:01:55",
      "content": "<p>I used it. However, the data was completely forged. As a result, the models I trained tended to consistently find forged data. I tried to achieve balance using different techniques with authentic data. But I couldn't do it. I stopped using it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3392640,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "01/17/2026 09:30:09",
          "content": "<p>Yes, but you could have created authentic versions of them programmatically. For every forged sample, I created an authentic version by simply overwriting the pristine panel inside the directory.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3392654,
              "author_name": "musapeker",
              "author_url": "",
              "post_date": "01/17/2026 09:57:38",
              "content": "<p>That makes sense. I also created synthetic authentic data by combining images from the single panel. One of the issues is that the forgeries in the supplementary images are different from the forgeries in the RSIID Dataset. The RSIID Dataset consists of fully synthetic data created by researchers; it is not real data. The forgeries in the supplementary images are much more professional. For this reason, training with RSIID Dataset data may not have worked well on the final data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3392776,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "01/17/2026 13:25:40",
      "content": "<p>Note some of the rsiid dataset had some identical images with trainset. So if you didnt dedupe there was leakage in your groupkfold validation. Ultimately i did not focus a lot on it but judging from what you say, i should have to get some extra points boost</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3393884,
      "author_name": "tusharsharmads",
      "author_url": "",
      "post_date": "01/20/2026 04:22:57",
      "content": "<p>I used this dataset and got a score of 0.314 using Dino v2+semantic segmentation head. Ultimately, the problem does not seem to be the copy-move dataset, it might be how to generate copy-move inference using this data since its not classic semantic segmentation. So I tried contrastive modeling but got worse scores (0.2-0.3).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3392594": "https://zenodo.org/records/15095089\n\nIt was surprising to me that this dataset didn't improve my lb score that much. I had a hunch that private test set will be mostly inter/intra-panel forgeries, so I also used this dataset for training my models. My best public lb score was 0.311 which is 0.008 better than all authentic prediction. Anyone used this dataset and get a decent boost?",
    "3392631": "I used it. However, the data was completely forged. As a result, the models I trained tended to consistently find forged data. I tried to achieve balance using different techniques with authentic data. But I couldn't do it. I stopped using it.",
    "3392640": "Yes, but you could have created authentic versions of them programmatically. For every forged sample, I created an authentic version by simply overwriting the pristine panel inside the directory.",
    "3392654": "That makes sense. I also created synthetic authentic data by combining images from the single panel. One of the issues is that the forgeries in the supplementary images are different from the forgeries in the RSIID Dataset. The RSIID Dataset consists of fully synthetic data created by researchers; it is not real data. The forgeries in the supplementary images are much more professional. For this reason, training with RSIID Dataset data may not have worked well on the final data.",
    "3392776": "Note some of the rsiid dataset had some identical images with trainset. So if you didnt dedupe there was leakage in your groupkfold validation. Ultimately i did not focus a lot on it but judging from what you say, i should have to get some extra points boost",
    "3393884": "I used this dataset and got a score of 0.314 using Dino v2+semantic segmentation head. Ultimately, the problem does not seem to be the copy-move dataset, it might be how to generate copy-move inference using this data since its not classic semantic segmentation. So I tried contrastive modeling but got worse scores (0.2-0.3)."
  },
  "source": "meta"
}