{
  "id": 613990,
  "title": "The score does not improve beyond 0.303 significantly!!! Why?",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/613990",
  "author_name": "",
  "post_date": "2025-10-31T09:29:14.978638900Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The score doesn’t improve beyond 0.303. This seems to happen because the model fails to predict the forged pixels correctly. Since the number of non-forged pixels is much higher, the loss function (like BCE) tends to focus on the majority class and just ignores the forged regions to minimize the overall loss.</p>\n<p>As a result, the model basically predicts everything as authentic, which explains why we get a constant 0.303 F1 score, both in my dummy submission and my toy CNN example.</p>\n<p>Dummy submission (all authentic), F1 = 0.303\n<a href=\"https://www.kaggle.com/code/vinothkumarsekar89/dummy-submission-all-authentic\" target=\"_blank\">Dummy Submission</a></p>\n<p>Toy CNN example, F1 = 0.303\n<a href=\"https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example\" target=\"_blank\">A Toy CNN Example</a></p>\n<p>It also looks like roughly 30% of the test samples are authentic, which matches the score. The model simply ends up predicting everything as authentic.</p>\n<p>This is also clear from the baseline model’s (Toy CNN Example) CV. It shows F1 = 0.0000 on forged pixels (and F1 = 1.000 on authentic ones) due to:</p>\n<ul>\n<li>Severe pixel-level class imbalance (most pixels are non-forged)</li>\n<li>BCE loss may not handle such imbalance well</li>\n<li>Too simple CNN architecture</li>\n</ul>\n<p>Next Steps / Ideas:</p>\n<ul>\n<li>Try better loss functions that handle imbalance</li>\n<li>Explore deeper architectures like U-Net or similar </li>\n</ul>\n<p>Please let me know your opinion.!!! </p>",
  "messages": [
    {
      "id": "3309240",
      "postDate": "10/31/2025 09:29:14",
      "content": "<p>The score doesn’t improve beyond 0.303. This seems to happen because the model fails to predict the forged pixels correctly. Since the number of non-forged pixels is much higher, the loss function (like BCE) tends to focus on the majority class and just ignores the forged regions to minimize the overall loss.</p>\n<p>As a result, the model basically predicts everything as authentic, which explains why we get a constant 0.303 F1 score, both in my dummy submission and my toy CNN example.</p>\n<p>Dummy submission (all authentic), F1 = 0.303\n<a href=\"https://www.kaggle.com/code/vinothkumarsekar89/dummy-submission-all-authentic\" target=\"_blank\">Dummy Submission</a></p>\n<p>Toy CNN example, F1 = 0.303\n<a href=\"https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example\" target=\"_blank\">A Toy CNN Example</a></p>\n<p>It also looks like roughly 30% of the test samples are authentic, which matches the score. The model simply ends up predicting everything as authentic.</p>\n<p>This is also clear from the baseline model’s (Toy CNN Example) CV. It shows F1 = 0.0000 on forged pixels (and F1 = 1.000 on authentic ones) due to:</p>\n<ul>\n<li>Severe pixel-level class imbalance (most pixels are non-forged)</li>\n<li>BCE loss may not handle such imbalance well</li>\n<li>Too simple CNN architecture</li>\n</ul>\n<p>Next Steps / Ideas:</p>\n<ul>\n<li>Try better loss functions that handle imbalance</li>\n<li>Explore deeper architectures like U-Net or similar </li>\n</ul>\n<p>Please let me know your opinion.!!! </p>",
      "rawMarkdown": "The score doesn’t improve beyond 0.303. This seems to happen because the model fails to predict the forged pixels correctly. Since the number of non-forged pixels is much higher, the loss function (like BCE) tends to focus on the majority class and just ignores the forged regions to minimize the overall loss.\n\nAs a result, the model basically predicts everything as authentic, which explains why we get a constant 0.303 F1 score, both in my dummy submission and my toy CNN example.\n\nDummy submission (all authentic), F1 = 0.303\n[Dummy Submission](https://www.kaggle.com/code/vinothkumarsekar89/dummy-submission-all-authentic)\n\nToy CNN example, F1 = 0.303\n[A Toy CNN Example](https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example)\n\nIt also looks like roughly 30% of the test samples are authentic, which matches the score. The model simply ends up predicting everything as authentic.\n\nThis is also clear from the baseline model’s (Toy CNN Example) CV. It shows F1 = 0.0000 on forged pixels (and F1 = 1.000 on authentic ones) due to:\n\n- Severe pixel-level class imbalance (most pixels are non-forged)\n- BCE loss may not handle such imbalance well\n- Too simple CNN architecture\n\nNext Steps / Ideas:\n- Try better loss functions that handle imbalance\n- Explore deeper architectures like U-Net or similar \n\nPlease let me know your opinion.!!!",
      "votes": null
    },
    {
      "id": "3309341",
      "postDate": "10/31/2025 13:53:24",
      "content": "<p>Thank you for sharing. Your viewpoint is correct. However, for images like corn cobs, it's still possible to make a relatively good distinction. But I believe the current bottleneck of the model lies in its inability to effectively differentiate between images with and without copy-move operations, which results in a relatively low score.Also, regarding your mention of designing a more complex model, I'm looking forward to your performance. Good luck!</p>",
      "rawMarkdown": "Thank you for sharing. Your viewpoint is correct. However, for images like corn cobs, it's still possible to make a relatively good distinction. But I believe the current bottleneck of the model lies in its inability to effectively differentiate between images with and without copy-move operations, which results in a relatively low score.Also, regarding your mention of designing a more complex model, I'm looking forward to your performance. Good luck!",
      "votes": null
    },
    {
      "id": "3309387",
      "postDate": "10/31/2025 15:33:35",
      "content": "<p>To deal with imbalance issue I have been using dice loss averaged with BCE although I am not sure if this is the best approach.</p>",
      "rawMarkdown": "To deal with imbalance issue I have been using dice loss averaged with BCE although I am not sure if this is the best approach.",
      "votes": null
    },
    {
      "id": "3311234",
      "postDate": "11/04/2025 14:53:05",
      "content": "<p>Try to increase weight of minority class. This approch works very well in most tasks.</p>",
      "rawMarkdown": "Try to increase weight of minority class. This approch works very well in most tasks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3309341,
      "author_name": "qifeihhh666",
      "author_url": "",
      "post_date": "10/31/2025 13:53:24",
      "content": "<p>Thank you for sharing. Your viewpoint is correct. However, for images like corn cobs, it's still possible to make a relatively good distinction. But I believe the current bottleneck of the model lies in its inability to effectively differentiate between images with and without copy-move operations, which results in a relatively low score.Also, regarding your mention of designing a more complex model, I'm looking forward to your performance. Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3309387,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "10/31/2025 15:33:35",
      "content": "<p>To deal with imbalance issue I have been using dice loss averaged with BCE although I am not sure if this is the best approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3311234,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "11/04/2025 14:53:05",
      "content": "<p>Try to increase weight of minority class. This approch works very well in most tasks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3309240": "The score doesn’t improve beyond 0.303. This seems to happen because the model fails to predict the forged pixels correctly. Since the number of non-forged pixels is much higher, the loss function (like BCE) tends to focus on the majority class and just ignores the forged regions to minimize the overall loss.\n\nAs a result, the model basically predicts everything as authentic, which explains why we get a constant 0.303 F1 score, both in my dummy submission and my toy CNN example.\n\nDummy submission (all authentic), F1 = 0.303\n[Dummy Submission](https://www.kaggle.com/code/vinothkumarsekar89/dummy-submission-all-authentic)\n\nToy CNN example, F1 = 0.303\n[A Toy CNN Example](https://www.kaggle.com/code/vinothkumarsekar89/recod-ai-sifd-a-toy-cnn-example)\n\nIt also looks like roughly 30% of the test samples are authentic, which matches the score. The model simply ends up predicting everything as authentic.\n\nThis is also clear from the baseline model’s (Toy CNN Example) CV. It shows F1 = 0.0000 on forged pixels (and F1 = 1.000 on authentic ones) due to:\n\n- Severe pixel-level class imbalance (most pixels are non-forged)\n- BCE loss may not handle such imbalance well\n- Too simple CNN architecture\n\nNext Steps / Ideas:\n- Try better loss functions that handle imbalance\n- Explore deeper architectures like U-Net or similar \n\nPlease let me know your opinion.!!!",
    "3309341": "Thank you for sharing. Your viewpoint is correct. However, for images like corn cobs, it's still possible to make a relatively good distinction. But I believe the current bottleneck of the model lies in its inability to effectively differentiate between images with and without copy-move operations, which results in a relatively low score.Also, regarding your mention of designing a more complex model, I'm looking forward to your performance. Good luck!",
    "3309387": "To deal with imbalance issue I have been using dice loss averaged with BCE although I am not sure if this is the best approach.",
    "3311234": "Try to increase weight of minority class. This approch works very well in most tasks."
  },
  "source": "meta"
}