{
  "id": 621903,
  "title": "Is a cropped images dataset useful?",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/621903",
  "author_name": "Andres H. Zapke",
  "post_date": "2025-11-15T11:17:19.506000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>i am reading through the notebooks and came across a note addressing the pixel class imbalance (there is many more non-forged 0 than forged 1 pixels in the masks dataset), and how that negatively impacts the training due to the NN bias towards 0.</p>\n<p>I thought to create a balanced dataset, cropping the forged regions out of the masks/images and also adding non-forged images. However the image context would be probably lost, as i imagine that every decoder/segmenter is based on finding similarities between two objects in one image. Also, im not sure which balance we are seeking, since adding 50/50 forged vs non-forged images would also lead to pixel imbalance. </p>\n<p>Then i thought one could first find similar patches on the image, and pass cropped region-pairs to the decoder/segmenter for training, but this also loses the image context and wouldnt work for deletions/removals. Any thoughts or ideas on this strategy?</p>\n<p>Do you think uploading a dataset containing only the forged regions could be of any use ?</p>\n<p>Where i found the note: <a href=\"https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example\" target=\"_blank\">https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example</a>\nCoarse grained (pseudo-mask) from \"painting\" only the similar patches: <a href=\"https://www.kaggle.com/code/imaadmahmood/sift-elite-v3-1-micro-tuned\" target=\"_blank\">https://www.kaggle.com/code/imaadmahmood/sift-elite-v3-1-micro-tuned</a></p>",
  "messages": [
    {
      "id": 3326643,
      "postDate": "2025-11-15T11:17:19.507Z",
      "content": "<p>Hi,</p>\n<p>i am reading through the notebooks and came across a note addressing the pixel class imbalance (there is many more non-forged 0 than forged 1 pixels in the masks dataset), and how that negatively impacts the training due to the NN bias towards 0.</p>\n<p>I thought to create a balanced dataset, cropping the forged regions out of the masks/images and also adding non-forged images. However the image context would be probably lost, as i imagine that every decoder/segmenter is based on finding similarities between two objects in one image. Also, im not sure which balance we are seeking, since adding 50/50 forged vs non-forged images would also lead to pixel imbalance. </p>\n<p>Then i thought one could first find similar patches on the image, and pass cropped region-pairs to the decoder/segmenter for training, but this also loses the image context and wouldnt work for deletions/removals. Any thoughts or ideas on this strategy?</p>\n<p>Do you think uploading a dataset containing only the forged regions could be of any use ?</p>\n<p>Where i found the note: <a href=\"https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example\" target=\"_blank\">https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example</a>\nCoarse grained (pseudo-mask) from \"painting\" only the similar patches: <a href=\"https://www.kaggle.com/code/imaadmahmood/sift-elite-v3-1-micro-tuned\" target=\"_blank\">https://www.kaggle.com/code/imaadmahmood/sift-elite-v3-1-micro-tuned</a></p>",
      "rawMarkdown": "Hi,\n\ni am reading through the notebooks and came across a note addressing the pixel class imbalance (there is many more non-forged 0 than forged 1 pixels in the masks dataset), and how that negatively impacts the training due to the NN bias towards 0.\n\nI thought to create a balanced dataset, cropping the forged regions out of the masks/images and also adding non-forged images. However the image context would be probably lost, as i imagine that every decoder/segmenter is based on finding similarities between two objects in one image. Also, im not sure which balance we are seeking, since adding 50/50 forged vs non-forged images would also lead to pixel imbalance. \n\nThen i thought one could first find similar patches on the image, and pass cropped region-pairs to the decoder/segmenter for training, but this also loses the image context and wouldnt work for deletions/removals. Any thoughts or ideas on this strategy?\n\nDo you think uploading a dataset containing only the forged regions could be of any use ?\n\nWhere i found the note: https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example\nCoarse grained (pseudo-mask) from \"painting\" only the similar patches: https://www.kaggle.com/code/imaadmahmood/sift-elite-v3-1-micro-tuned",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3326643": "Hi,\n\ni am reading through the notebooks and came across a note addressing the pixel class imbalance (there is many more non-forged 0 than forged 1 pixels in the masks dataset), and how that negatively impacts the training due to the NN bias towards 0.\n\nI thought to create a balanced dataset, cropping the forged regions out of the masks/images and also adding non-forged images. However the image context would be probably lost, as i imagine that every decoder/segmenter is based on finding similarities between two objects in one image. Also, im not sure which balance we are seeking, since adding 50/50 forged vs non-forged images would also lead to pixel imbalance. \n\nThen i thought one could first find similar patches on the image, and pass cropped region-pairs to the decoder/segmenter for training, but this also loses the image context and wouldnt work for deletions/removals. Any thoughts or ideas on this strategy?\n\nDo you think uploading a dataset containing only the forged regions could be of any use ?\n\nWhere i found the note: https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example\nCoarse grained (pseudo-mask) from \"painting\" only the similar patches: https://www.kaggle.com/code/imaadmahmood/sift-elite-v3-1-micro-tuned"
  }
}