{
  "id": 614069,
  "title": "Getting Deeper into Scientific Image Copy-Move Forgeries",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/614069",
  "author_name": "",
  "post_date": "2025-10-31T20:53:13.813843200Z",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This post gives a quick overview and a few insights that might help you think about the problem from a forensic perspective.</p>\n<h2>1. What is Copy-Move Forgery (CMF)?</h2>\n<p>A <strong>copy-move forgery</strong> happens when someone copies a region of an image and pastes it somewhere else <em>within the same image</em>.</p>\n<p>The reasons for doing this vary.</p>\n<ul>\n<li><p>Sometimes it’s to <strong>fabricate results</strong>,  duplicating objects to make an experiment look more convincing.</p></li>\n<li><p>Other times it’s to <strong>hide something</strong>, such as covering up a defect or unwanted artifact by pasting the background over it.</p>\n<p><strong>Example 1 – Object duplication</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fb6463c0fd4e3f456e216417c30a51471%2Fduplication.png?generation=1761943705342104&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<p><strong>Example 2 – Object Cleaning</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F71c66c3830fcd38a4ba9fa778bfc1279%2Fcleaning.png?generation=1761943749703657&amp;alt=media\" alt=\"\"></p>\n<h2>2. How Forgeries Are Made</h2>\n<p>Bad actors know detectors exist, so they try to make detection difficult. Real-world CMFs are rarely simple “copy and paste” edits.</p>\n<h3>a. Geometric Transformations</h3>\n<p>The copied region is almost never identical. It’s usually transformed with homographic transformations to blend naturally into the new context.\n Common transformations include:</p>\n<ul>\n<li><strong>Rotation</strong> (e.g., a few degrees)</li>\n<li><strong>Flipping</strong> (horizontal or vertical)</li>\n<li><strong>Scaling</strong> (shrinking or enlarging)</li>\n<li>Or a <strong>combination</strong> of these (e.g., rotated + scaled + flipped)</li>\n</ul>\n<h3>b. Region Size</h3>\n<p>CMF occurs on regions of any size, from small patches to large portions of the image.</p>\n<h3>c. Post-Processing &amp; Obfuscation</h3>\n<p>After pasting, the region is often adjusted to hide the manipulation:</p>\n<ul>\n<li><strong>Edge smoothing or feathering</strong> to blend the pasted area with its surroundings.</li>\n<li><strong>Brightness or contrast adjustments</strong> to match lighting conditions.</li>\n<li>Sometimes, <strong>extra noise</strong> is added over the image to mask inconsistencies.</li>\n</ul>\n<p>These tweaks attempt the pasted region to “disappear” visually and statistically (considering noise stats).</p>\n<h3>d. Resolution</h3>\n<p>Most real forgeries occur on <strong>mid-to-low resolution</strong> images.\nScaling and pixelation artifacts naturally confuse detectors and humans and make differences harder to spot.</p>\n<h2>3. Why Scientific Images Are Tricky</h2>\n<p>Scientific images bring their own set of challenges. Many contains <strong>multiple sub-images or panels</strong>, and each of these can act as a distraction.</p>\n<p>For example, a multi-panel figure might have several plots or microscopy crops arranged together. The repeated structures, similar textures, labels, and grid lines can all look like CMFs — but they’re actually benign duplicates.</p>\n<p>Some examples:</p>\n<ul>\n<li><strong>Bar graphs</strong> often contain identical-looking bars (not forgeries).</li>\n<li><strong>Microscopy images</strong> show repeating cellular patterns.</li>\n<li><strong>Grids and labels</strong> introduce sharp edges and text that can confuse detectors.</li>\n</ul>\n<p>So, context really matters. A duplicated region isn’t always a forgery.</p>\n<h2>4. What This Means for Your Models</h2>\n<p>Here are a few takeaways when designing your approach:</p>\n<ol>\n<li><strong>Be robust to transformations.</strong>\nThe copied region might be rotated, scaled, or flipped. Models need to handle these geometric changes.</li>\n<li><strong>Look for post-processing traces.</strong>\nEven if the pasted region blends visually, subtle differences in texture, noise, or local statistics may remain.</li>\n<li><strong>Use context.</strong>\nNot every similar patch is malicious. Your model should learn to distinguish between <em>real duplication</em> and <em>natural repetition</em> (e.g., bars in a graph, repeated patterns in microscopy).</li>\n</ol>\n<hr>\n<p>These are just some starting points for your ideas, don’t feel limited by them.</p>\n<p>More resources:</p>\n<ol>\n<li><p><a href=\"https://link.springer.com/article/10.1007/s11948-022-00391-4\" target=\"_blank\">Benchmarking Scientific Image Forgery Detectors</a></p></li>\n<li><p><a href=\"https://rupress.org/jcb/article/166/1/11/34064/What-s-in-a-picture-The-temptation-of-image\" target=\"_blank\">What's in a picture? The temptation of image manipulation</a></p></li>\n</ol>\n<p>Feel free to share your thoughts and experiences with this problem  :)</p>\n<p>Good luck, and have fun!</p>",
  "messages": [
    {
      "id": "3309489",
      "postDate": "10/31/2025 20:53:13",
      "content": "<p>This post gives a quick overview and a few insights that might help you think about the problem from a forensic perspective.</p>\n<h2>1. What is Copy-Move Forgery (CMF)?</h2>\n<p>A <strong>copy-move forgery</strong> happens when someone copies a region of an image and pastes it somewhere else <em>within the same image</em>.</p>\n<p>The reasons for doing this vary.</p>\n<ul>\n<li><p>Sometimes it’s to <strong>fabricate results</strong>,  duplicating objects to make an experiment look more convincing.</p></li>\n<li><p>Other times it’s to <strong>hide something</strong>, such as covering up a defect or unwanted artifact by pasting the background over it.</p>\n<p><strong>Example 1 – Object duplication</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fb6463c0fd4e3f456e216417c30a51471%2Fduplication.png?generation=1761943705342104&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<p><strong>Example 2 – Object Cleaning</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F71c66c3830fcd38a4ba9fa778bfc1279%2Fcleaning.png?generation=1761943749703657&amp;alt=media\" alt=\"\"></p>\n<h2>2. How Forgeries Are Made</h2>\n<p>Bad actors know detectors exist, so they try to make detection difficult. Real-world CMFs are rarely simple “copy and paste” edits.</p>\n<h3>a. Geometric Transformations</h3>\n<p>The copied region is almost never identical. It’s usually transformed with homographic transformations to blend naturally into the new context.\n Common transformations include:</p>\n<ul>\n<li><strong>Rotation</strong> (e.g., a few degrees)</li>\n<li><strong>Flipping</strong> (horizontal or vertical)</li>\n<li><strong>Scaling</strong> (shrinking or enlarging)</li>\n<li>Or a <strong>combination</strong> of these (e.g., rotated + scaled + flipped)</li>\n</ul>\n<h3>b. Region Size</h3>\n<p>CMF occurs on regions of any size, from small patches to large portions of the image.</p>\n<h3>c. Post-Processing &amp; Obfuscation</h3>\n<p>After pasting, the region is often adjusted to hide the manipulation:</p>\n<ul>\n<li><strong>Edge smoothing or feathering</strong> to blend the pasted area with its surroundings.</li>\n<li><strong>Brightness or contrast adjustments</strong> to match lighting conditions.</li>\n<li>Sometimes, <strong>extra noise</strong> is added over the image to mask inconsistencies.</li>\n</ul>\n<p>These tweaks attempt the pasted region to “disappear” visually and statistically (considering noise stats).</p>\n<h3>d. Resolution</h3>\n<p>Most real forgeries occur on <strong>mid-to-low resolution</strong> images.\nScaling and pixelation artifacts naturally confuse detectors and humans and make differences harder to spot.</p>\n<h2>3. Why Scientific Images Are Tricky</h2>\n<p>Scientific images bring their own set of challenges. Many contains <strong>multiple sub-images or panels</strong>, and each of these can act as a distraction.</p>\n<p>For example, a multi-panel figure might have several plots or microscopy crops arranged together. The repeated structures, similar textures, labels, and grid lines can all look like CMFs — but they’re actually benign duplicates.</p>\n<p>Some examples:</p>\n<ul>\n<li><strong>Bar graphs</strong> often contain identical-looking bars (not forgeries).</li>\n<li><strong>Microscopy images</strong> show repeating cellular patterns.</li>\n<li><strong>Grids and labels</strong> introduce sharp edges and text that can confuse detectors.</li>\n</ul>\n<p>So, context really matters. A duplicated region isn’t always a forgery.</p>\n<h2>4. What This Means for Your Models</h2>\n<p>Here are a few takeaways when designing your approach:</p>\n<ol>\n<li><strong>Be robust to transformations.</strong>\nThe copied region might be rotated, scaled, or flipped. Models need to handle these geometric changes.</li>\n<li><strong>Look for post-processing traces.</strong>\nEven if the pasted region blends visually, subtle differences in texture, noise, or local statistics may remain.</li>\n<li><strong>Use context.</strong>\nNot every similar patch is malicious. Your model should learn to distinguish between <em>real duplication</em> and <em>natural repetition</em> (e.g., bars in a graph, repeated patterns in microscopy).</li>\n</ol>\n<hr>\n<p>These are just some starting points for your ideas, don’t feel limited by them.</p>\n<p>More resources:</p>\n<ol>\n<li><p><a href=\"https://link.springer.com/article/10.1007/s11948-022-00391-4\" target=\"_blank\">Benchmarking Scientific Image Forgery Detectors</a></p></li>\n<li><p><a href=\"https://rupress.org/jcb/article/166/1/11/34064/What-s-in-a-picture-The-temptation-of-image\" target=\"_blank\">What's in a picture? The temptation of image manipulation</a></p></li>\n</ol>\n<p>Feel free to share your thoughts and experiences with this problem  :)</p>\n<p>Good luck, and have fun!</p>",
      "rawMarkdown": "This post gives a quick overview and a few insights that might help you think about the problem from a forensic perspective.\n\n## 1. What is Copy-Move Forgery (CMF)?\n\nA **copy-move forgery** happens when someone copies a region of an image and pastes it somewhere else *within the same image*.\n\nThe reasons for doing this vary.\n\n- Sometimes it’s to **fabricate results**,  duplicating objects to make an experiment look more convincing.\n- Other times it’s to **hide something**, such as covering up a defect or unwanted artifact by pasting the background over it.\n\n\n **Example 1 – Object duplication**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fb6463c0fd4e3f456e216417c30a51471%2Fduplication.png?generation=1761943705342104&alt=media)\n\n**Example 2 – Object Cleaning**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F71c66c3830fcd38a4ba9fa778bfc1279%2Fcleaning.png?generation=1761943749703657&alt=media)\n\n## 2. How Forgeries Are Made\n\nBad actors know detectors exist, so they try to make detection difficult. Real-world CMFs are rarely simple “copy and paste” edits.\n\n### a. Geometric Transformations\n\nThe copied region is almost never identical. It’s usually transformed with homographic transformations to blend naturally into the new context.\n Common transformations include:\n\n- **Rotation** (e.g., a few degrees)\n- **Flipping** (horizontal or vertical)\n- **Scaling** (shrinking or enlarging)\n- Or a **combination** of these (e.g., rotated + scaled + flipped)\n\n### b. Region Size\n\nCMF occurs on regions of any size, from small patches to large portions of the image.\n\n### c. Post-Processing & Obfuscation\n\nAfter pasting, the region is often adjusted to hide the manipulation:\n\n- **Edge smoothing or feathering** to blend the pasted area with its surroundings.\n- **Brightness or contrast adjustments** to match lighting conditions.\n- Sometimes, **extra noise** is added over the image to mask inconsistencies.\n\nThese tweaks attempt the pasted region to “disappear” visually and statistically (considering noise stats).\n\n### d. Resolution\n\nMost real forgeries occur on **mid-to-low resolution** images.\nScaling and pixelation artifacts naturally confuse detectors and humans and make differences harder to spot.\n\n\n\n## 3. Why Scientific Images Are Tricky\n\nScientific images bring their own set of challenges. Many contains **multiple sub-images or panels**, and each of these can act as a distraction.\n\nFor example, a multi-panel figure might have several plots or microscopy crops arranged together. The repeated structures, similar textures, labels, and grid lines can all look like CMFs — but they’re actually benign duplicates.\n\nSome examples:\n\n- **Bar graphs** often contain identical-looking bars (not forgeries).\n- **Microscopy images** show repeating cellular patterns.\n- **Grids and labels** introduce sharp edges and text that can confuse detectors.\n\nSo, context really matters. A duplicated region isn’t always a forgery.\n\n\n\n## 4. What This Means for Your Models\n\nHere are a few takeaways when designing your approach:\n\n1. **Be robust to transformations.**\n   The copied region might be rotated, scaled, or flipped. Models need to handle these geometric changes.\n2. **Look for post-processing traces.**\n   Even if the pasted region blends visually, subtle differences in texture, noise, or local statistics may remain.\n3. **Use context.**\n   Not every similar patch is malicious. Your model should learn to distinguish between *real duplication* and *natural repetition* (e.g., bars in a graph, repeated patterns in microscopy).\n\n\n\n------\n\nThese are just some starting points for your ideas, don’t feel limited by them.\n\nMore resources:\n\n1. [Benchmarking Scientific Image Forgery Detectors](https://link.springer.com/article/10.1007/s11948-022-00391-4)\n\n2. [What's in a picture? The temptation of image manipulation](https://rupress.org/jcb/article/166/1/11/34064/What-s-in-a-picture-The-temptation-of-image)\n\n\n\nFeel free to share your thoughts and experiences with this problem  :)\n\n\n\nGood luck, and have fun!",
      "votes": null
    },
    {
      "id": "3309894",
      "postDate": "11/01/2025 17:12:33",
      "content": "<p>My thoughts:</p>\n<p>1) Will there be scaling transformations in the hidden test set? I don't recall seeing any in the training set.</p>\n<p>For reference, I've been tinkering with postprocessing technique that identifies my predicted masks, then makes embedding for each individual mask and then tries to cluster masks, and only accepts clusters if there are at least 2 masks in the cluster (implies there is a copy-forge), otherwise it removes the mask prediction. Right now, this postprocessing is not effective, but maybe with better model predictions/more tinkering it will be?</p>\n<p>2) Will there be panel data in the hidden test set? Don't see any in the training set. If so, can we get some expectation of the percentage, that way we can prepare validation set properly balanced?</p>\n<p>3) In the literature I see basically 2 approaches: finetune a pretrained image net to predict the masks (good at identifying visually similar things, but struggles when you have a photo of lots of similar-looking cells), OR train a new CNN with small stride, small amount of pooling to try and capture rough edges around the copy-pasted region, that way it doesn't get confused between similar-looking things that aren't a direct copy &amp; paste. I expect the best solution would be a combination of both of these approaches (so far, they both perform similarly for me, but haven't tried a blend/stack). Do you agree?</p>\n<p>4) Seemingly adding \"patch embedding features\" (SIFT, Zernike, HardNet, etc.) simply as extra channels in the inputs not helping. Not sure if others experienced the same. I think these are probably more useful in postprocessing but I haven't experimented as much. Do you have any advice for these?</p>\n<p>5) I found warmup pretraining on other copy-forge datasets (CASIA, Defacto) barely helpful at all. Is it because the domain of biological photos is different, or perhaps the copy-forge techniques are different from the ones here?</p>\n<p>6) Will the hidden test set distribution of authentic vs inauthentic be similar to 50/50 like our training dataset?</p>",
      "rawMarkdown": "My thoughts:\n\n1) Will there be scaling transformations in the hidden test set? I don't recall seeing any in the training set.\n\nFor reference, I've been tinkering with postprocessing technique that identifies my predicted masks, then makes embedding for each individual mask and then tries to cluster masks, and only accepts clusters if there are at least 2 masks in the cluster (implies there is a copy-forge), otherwise it removes the mask prediction. Right now, this postprocessing is not effective, but maybe with better model predictions/more tinkering it will be?\n\n2) Will there be panel data in the hidden test set? Don't see any in the training set. If so, can we get some expectation of the percentage, that way we can prepare validation set properly balanced?\n\n3) In the literature I see basically 2 approaches: finetune a pretrained image net to predict the masks (good at identifying visually similar things, but struggles when you have a photo of lots of similar-looking cells), OR train a new CNN with small stride, small amount of pooling to try and capture rough edges around the copy-pasted region, that way it doesn't get confused between similar-looking things that aren't a direct copy & paste. I expect the best solution would be a combination of both of these approaches (so far, they both perform similarly for me, but haven't tried a blend/stack). Do you agree?\n\n4) Seemingly adding \"patch embedding features\" (SIFT, Zernike, HardNet, etc.) simply as extra channels in the inputs not helping. Not sure if others experienced the same. I think these are probably more useful in postprocessing but I haven't experimented as much. Do you have any advice for these?\n\n5) I found warmup pretraining on other copy-forge datasets (CASIA, Defacto) barely helpful at all. Is it because the domain of biological photos is different, or perhaps the copy-forge techniques are different from the ones here?\n\n6) Will the hidden test set distribution of authentic vs inauthentic be similar to 50/50 like our training dataset?",
      "votes": null
    },
    {
      "id": "3310185",
      "postDate": "11/02/2025 10:10:24",
      "content": "<p>I also used the criterion of \"at least 2 masks match, otherwise exclude them\" in my post-processing code. However, regarding the dataset, it seems we can't delve too deeply into its composition. The organizers appear to expect us to build a copy-move detection model capable of handling any scenario, which is truly an enormous challenge.</p>",
      "rawMarkdown": "I also used the criterion of \"at least 2 masks match, otherwise exclude them\" in my post-processing code. However, regarding the dataset, it seems we can't delve too deeply into its composition. The organizers appear to expect us to build a copy-move detection model capable of handling any scenario, which is truly an enormous challenge.",
      "votes": null
    },
    {
      "id": "3310498",
      "postDate": "11/03/2025 04:27:29",
      "content": "<p>Hello how are you i am maheen riaz i am ml professional and ai professional i have some issues in my note i submit 11 notebook but there is erro <a href=\"https://www.kaggle.com/code/maheenriaz1122/rluc-sifd-eda-train-inference?scriptVersionId=272997301\" target=\"_blank\">https://www.kaggle.com/code/maheenriaz1122/rluc-sifd-eda-train-inference?scriptVersionId=272997301</a></p>",
      "rawMarkdown": "Hello how are you i am maheen riaz i am ml professional and ai professional i have some issues in my note i submit 11 notebook but there is erro https://www.kaggle.com/code/maheenriaz1122/rluc-sifd-eda-train-inference?scriptVersionId=272997301",
      "votes": null
    },
    {
      "id": "3310667",
      "postDate": "11/03/2025 12:51:19",
      "content": "<p>Hello! I think it's because your submission is in the wrong format:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26368162%2Fb76e45c6e9459fb427839f0b8944c2a5%2FScreenshot%202025-11-03%20134910.png?generation=1762174160713075&amp;alt=media\" alt=\"img\"></p>\n<p><code>endoded_pixels</code> should be <code>annotation</code> a and if it is not \"authentic\", it should be encoded in RLE format </p>\n<p>Exemple:</p>\n<pre><code>id,\n,authentic\n,\n</code></pre>",
      "rawMarkdown": "Hello! I think it's because your submission is in the wrong format:\n\n![img](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26368162%2Fb76e45c6e9459fb427839f0b8944c2a5%2FScreenshot%202025-11-03%20134910.png?generation=1762174160713075&alt=media)\n\n`endoded_pixels` should be `annotation` a and if it is not \"authentic\", it should be encoded in RLE format \n\nExemple:\n\n```\ncase_id,annotation\n1,authentic\n2,\"[123 4]\"\n```",
      "votes": null
    },
    {
      "id": "3328640",
      "postDate": "11/15/2025 22:58:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> \nCan you give little more clarification on the <strong>what the mask represent</strong> in the <strong>Example 2 – Object Cleaning</strong> ?</p>",
      "rawMarkdown": "Hi @joophillipecardenuto \nCan you give little more clarification on the **what the mask represent** in the **Example 2 – Object Cleaning** ?",
      "votes": null
    },
    {
      "id": "3333867",
      "postDate": "11/17/2025 10:26:00",
      "content": "<p>Hi,</p>\n<p>In both cases, the masks represent the copied and source areas.</p>\n<p>The cleaning case differs from the other only in that the object is being hidden from the image, rather than duplicated. In this scenario, a region of the background is copied and pasted over a cell (or any other object) to cover it.</p>\n<p>If you adjust the image colors, it makes it easier to verify that the background has been copied over the cell. By looking at the differences in the neighborhood (surrounding area) of that operation, you can often find artifacts of the forgery.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fddccd5dd9bc1e504b2a8c08f4f22d657%2Finbox_11373480_34c9647262711972cd65240743dcdcd8_inbox_11373480_151df79a678026afe02902ece04bd009_background-forgery.png?generation=1763374959900741&amp;alt=media\" alt=\"\"></p>\n<p>Hope that helps!</p>",
      "rawMarkdown": "Hi,\n\nIn both cases, the masks represent the copied and source areas.\n\nThe cleaning case differs from the other only in that the object is being hidden from the image, rather than duplicated. In this scenario, a region of the background is copied and pasted over a cell (or any other object) to cover it.\n\nIf you adjust the image colors, it makes it easier to verify that the background has been copied over the cell. By looking at the differences in the neighborhood (surrounding area) of that operation, you can often find artifacts of the forgery.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fddccd5dd9bc1e504b2a8c08f4f22d657%2Finbox_11373480_34c9647262711972cd65240743dcdcd8_inbox_11373480_151df79a678026afe02902ece04bd009_background-forgery.png?generation=1763374959900741&alt=media)\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "3350810",
      "postDate": "11/27/2025 20:19:56",
      "content": "<p>Hello,</p>\n<ol>\n<li><p>In the first example, a few of the red-circled nuclei seem unique and do not match any duplicated region in the image. I’m unsure whether the red markings represent only the pasted (forged) area or both the original source and the pasted region. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14946445%2F2f9b44714a11a8c784c334f5112f8fe8%2FScreenshot%202025-11-27%20210353.png?generation=1764277086315160&amp;alt=media\" alt=\"here\"></p></li>\n<li><p>Can the pasted forged region be cropped from the source or only limited to: scaling, rotation, flipping and combination  as you mentioned?</p></li>\n</ol>",
      "rawMarkdown": "Hello,\n\n1. In the first example, a few of the red-circled nuclei seem unique and do not match any duplicated region in the image. I’m unsure whether the red markings represent only the pasted (forged) area or both the original source and the pasted region. ![here](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14946445%2F2f9b44714a11a8c784c334f5112f8fe8%2FScreenshot%202025-11-27%20210353.png?generation=1764277086315160&alt=media)\n\n2. Can the pasted forged region be cropped from the source or only limited to: scaling, rotation, flipping and combination  as you mentioned?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3309894,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "11/01/2025 17:12:33",
      "content": "<p>My thoughts:</p>\n<p>1) Will there be scaling transformations in the hidden test set? I don't recall seeing any in the training set.</p>\n<p>For reference, I've been tinkering with postprocessing technique that identifies my predicted masks, then makes embedding for each individual mask and then tries to cluster masks, and only accepts clusters if there are at least 2 masks in the cluster (implies there is a copy-forge), otherwise it removes the mask prediction. Right now, this postprocessing is not effective, but maybe with better model predictions/more tinkering it will be?</p>\n<p>2) Will there be panel data in the hidden test set? Don't see any in the training set. If so, can we get some expectation of the percentage, that way we can prepare validation set properly balanced?</p>\n<p>3) In the literature I see basically 2 approaches: finetune a pretrained image net to predict the masks (good at identifying visually similar things, but struggles when you have a photo of lots of similar-looking cells), OR train a new CNN with small stride, small amount of pooling to try and capture rough edges around the copy-pasted region, that way it doesn't get confused between similar-looking things that aren't a direct copy &amp; paste. I expect the best solution would be a combination of both of these approaches (so far, they both perform similarly for me, but haven't tried a blend/stack). Do you agree?</p>\n<p>4) Seemingly adding \"patch embedding features\" (SIFT, Zernike, HardNet, etc.) simply as extra channels in the inputs not helping. Not sure if others experienced the same. I think these are probably more useful in postprocessing but I haven't experimented as much. Do you have any advice for these?</p>\n<p>5) I found warmup pretraining on other copy-forge datasets (CASIA, Defacto) barely helpful at all. Is it because the domain of biological photos is different, or perhaps the copy-forge techniques are different from the ones here?</p>\n<p>6) Will the hidden test set distribution of authentic vs inauthentic be similar to 50/50 like our training dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3310185,
          "author_name": "qifeihhh666",
          "author_url": "",
          "post_date": "11/02/2025 10:10:24",
          "content": "<p>I also used the criterion of \"at least 2 masks match, otherwise exclude them\" in my post-processing code. However, regarding the dataset, it seems we can't delve too deeply into its composition. The organizers appear to expect us to build a copy-move detection model capable of handling any scenario, which is truly an enormous challenge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3310498,
      "author_name": "maheenriaz1122",
      "author_url": "",
      "post_date": "11/03/2025 04:27:29",
      "content": "<p>Hello how are you i am maheen riaz i am ml professional and ai professional i have some issues in my note i submit 11 notebook but there is erro <a href=\"https://www.kaggle.com/code/maheenriaz1122/rluc-sifd-eda-train-inference?scriptVersionId=272997301\" target=\"_blank\">https://www.kaggle.com/code/maheenriaz1122/rluc-sifd-eda-train-inference?scriptVersionId=272997301</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3310667,
          "author_name": "isaacmenard",
          "author_url": "",
          "post_date": "11/03/2025 12:51:19",
          "content": "<p>Hello! I think it's because your submission is in the wrong format:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26368162%2Fb76e45c6e9459fb427839f0b8944c2a5%2FScreenshot%202025-11-03%20134910.png?generation=1762174160713075&amp;alt=media\" alt=\"img\"></p>\n<p><code>endoded_pixels</code> should be <code>annotation</code> a and if it is not \"authentic\", it should be encoded in RLE format </p>\n<p>Exemple:</p>\n<pre><code>id,\n,authentic\n,\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3328640,
      "author_name": "manojkumars00",
      "author_url": "",
      "post_date": "11/15/2025 22:58:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/joophillipecardenuto\" target=\"_blank\">@joophillipecardenuto</a> \nCan you give little more clarification on the <strong>what the mask represent</strong> in the <strong>Example 2 – Object Cleaning</strong> ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3333867,
          "author_name": "joophillipecardenuto",
          "author_url": "",
          "post_date": "11/17/2025 10:26:00",
          "content": "<p>Hi,</p>\n<p>In both cases, the masks represent the copied and source areas.</p>\n<p>The cleaning case differs from the other only in that the object is being hidden from the image, rather than duplicated. In this scenario, a region of the background is copied and pasted over a cell (or any other object) to cover it.</p>\n<p>If you adjust the image colors, it makes it easier to verify that the background has been copied over the cell. By looking at the differences in the neighborhood (surrounding area) of that operation, you can often find artifacts of the forgery.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fddccd5dd9bc1e504b2a8c08f4f22d657%2Finbox_11373480_34c9647262711972cd65240743dcdcd8_inbox_11373480_151df79a678026afe02902ece04bd009_background-forgery.png?generation=1763374959900741&amp;alt=media\" alt=\"\"></p>\n<p>Hope that helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3350810,
      "author_name": "hamzahabduljalil",
      "author_url": "",
      "post_date": "11/27/2025 20:19:56",
      "content": "<p>Hello,</p>\n<ol>\n<li><p>In the first example, a few of the red-circled nuclei seem unique and do not match any duplicated region in the image. I’m unsure whether the red markings represent only the pasted (forged) area or both the original source and the pasted region. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14946445%2F2f9b44714a11a8c784c334f5112f8fe8%2FScreenshot%202025-11-27%20210353.png?generation=1764277086315160&amp;alt=media\" alt=\"here\"></p></li>\n<li><p>Can the pasted forged region be cropped from the source or only limited to: scaling, rotation, flipping and combination  as you mentioned?</p></li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3309489": "This post gives a quick overview and a few insights that might help you think about the problem from a forensic perspective.\n\n## 1. What is Copy-Move Forgery (CMF)?\n\nA **copy-move forgery** happens when someone copies a region of an image and pastes it somewhere else *within the same image*.\n\nThe reasons for doing this vary.\n\n- Sometimes it’s to **fabricate results**,  duplicating objects to make an experiment look more convincing.\n- Other times it’s to **hide something**, such as covering up a defect or unwanted artifact by pasting the background over it.\n\n\n **Example 1 – Object duplication**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fb6463c0fd4e3f456e216417c30a51471%2Fduplication.png?generation=1761943705342104&alt=media)\n\n**Example 2 – Object Cleaning**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F71c66c3830fcd38a4ba9fa778bfc1279%2Fcleaning.png?generation=1761943749703657&alt=media)\n\n## 2. How Forgeries Are Made\n\nBad actors know detectors exist, so they try to make detection difficult. Real-world CMFs are rarely simple “copy and paste” edits.\n\n### a. Geometric Transformations\n\nThe copied region is almost never identical. It’s usually transformed with homographic transformations to blend naturally into the new context.\n Common transformations include:\n\n- **Rotation** (e.g., a few degrees)\n- **Flipping** (horizontal or vertical)\n- **Scaling** (shrinking or enlarging)\n- Or a **combination** of these (e.g., rotated + scaled + flipped)\n\n### b. Region Size\n\nCMF occurs on regions of any size, from small patches to large portions of the image.\n\n### c. Post-Processing & Obfuscation\n\nAfter pasting, the region is often adjusted to hide the manipulation:\n\n- **Edge smoothing or feathering** to blend the pasted area with its surroundings.\n- **Brightness or contrast adjustments** to match lighting conditions.\n- Sometimes, **extra noise** is added over the image to mask inconsistencies.\n\nThese tweaks attempt the pasted region to “disappear” visually and statistically (considering noise stats).\n\n### d. Resolution\n\nMost real forgeries occur on **mid-to-low resolution** images.\nScaling and pixelation artifacts naturally confuse detectors and humans and make differences harder to spot.\n\n\n\n## 3. Why Scientific Images Are Tricky\n\nScientific images bring their own set of challenges. Many contains **multiple sub-images or panels**, and each of these can act as a distraction.\n\nFor example, a multi-panel figure might have several plots or microscopy crops arranged together. The repeated structures, similar textures, labels, and grid lines can all look like CMFs — but they’re actually benign duplicates.\n\nSome examples:\n\n- **Bar graphs** often contain identical-looking bars (not forgeries).\n- **Microscopy images** show repeating cellular patterns.\n- **Grids and labels** introduce sharp edges and text that can confuse detectors.\n\nSo, context really matters. A duplicated region isn’t always a forgery.\n\n\n\n## 4. What This Means for Your Models\n\nHere are a few takeaways when designing your approach:\n\n1. **Be robust to transformations.**\n   The copied region might be rotated, scaled, or flipped. Models need to handle these geometric changes.\n2. **Look for post-processing traces.**\n   Even if the pasted region blends visually, subtle differences in texture, noise, or local statistics may remain.\n3. **Use context.**\n   Not every similar patch is malicious. Your model should learn to distinguish between *real duplication* and *natural repetition* (e.g., bars in a graph, repeated patterns in microscopy).\n\n\n\n------\n\nThese are just some starting points for your ideas, don’t feel limited by them.\n\nMore resources:\n\n1. [Benchmarking Scientific Image Forgery Detectors](https://link.springer.com/article/10.1007/s11948-022-00391-4)\n\n2. [What's in a picture? The temptation of image manipulation](https://rupress.org/jcb/article/166/1/11/34064/What-s-in-a-picture-The-temptation-of-image)\n\n\n\nFeel free to share your thoughts and experiences with this problem  :)\n\n\n\nGood luck, and have fun!",
    "3309894": "My thoughts:\n\n1) Will there be scaling transformations in the hidden test set? I don't recall seeing any in the training set.\n\nFor reference, I've been tinkering with postprocessing technique that identifies my predicted masks, then makes embedding for each individual mask and then tries to cluster masks, and only accepts clusters if there are at least 2 masks in the cluster (implies there is a copy-forge), otherwise it removes the mask prediction. Right now, this postprocessing is not effective, but maybe with better model predictions/more tinkering it will be?\n\n2) Will there be panel data in the hidden test set? Don't see any in the training set. If so, can we get some expectation of the percentage, that way we can prepare validation set properly balanced?\n\n3) In the literature I see basically 2 approaches: finetune a pretrained image net to predict the masks (good at identifying visually similar things, but struggles when you have a photo of lots of similar-looking cells), OR train a new CNN with small stride, small amount of pooling to try and capture rough edges around the copy-pasted region, that way it doesn't get confused between similar-looking things that aren't a direct copy & paste. I expect the best solution would be a combination of both of these approaches (so far, they both perform similarly for me, but haven't tried a blend/stack). Do you agree?\n\n4) Seemingly adding \"patch embedding features\" (SIFT, Zernike, HardNet, etc.) simply as extra channels in the inputs not helping. Not sure if others experienced the same. I think these are probably more useful in postprocessing but I haven't experimented as much. Do you have any advice for these?\n\n5) I found warmup pretraining on other copy-forge datasets (CASIA, Defacto) barely helpful at all. Is it because the domain of biological photos is different, or perhaps the copy-forge techniques are different from the ones here?\n\n6) Will the hidden test set distribution of authentic vs inauthentic be similar to 50/50 like our training dataset?",
    "3310185": "I also used the criterion of \"at least 2 masks match, otherwise exclude them\" in my post-processing code. However, regarding the dataset, it seems we can't delve too deeply into its composition. The organizers appear to expect us to build a copy-move detection model capable of handling any scenario, which is truly an enormous challenge.",
    "3310498": "Hello how are you i am maheen riaz i am ml professional and ai professional i have some issues in my note i submit 11 notebook but there is erro https://www.kaggle.com/code/maheenriaz1122/rluc-sifd-eda-train-inference?scriptVersionId=272997301",
    "3310667": "Hello! I think it's because your submission is in the wrong format:\n\n![img](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26368162%2Fb76e45c6e9459fb427839f0b8944c2a5%2FScreenshot%202025-11-03%20134910.png?generation=1762174160713075&alt=media)\n\n`endoded_pixels` should be `annotation` a and if it is not \"authentic\", it should be encoded in RLE format \n\nExemple:\n\n```\ncase_id,annotation\n1,authentic\n2,\"[123 4]\"\n```",
    "3328640": "Hi @joophillipecardenuto \nCan you give little more clarification on the **what the mask represent** in the **Example 2 – Object Cleaning** ?",
    "3333867": "Hi,\n\nIn both cases, the masks represent the copied and source areas.\n\nThe cleaning case differs from the other only in that the object is being hidden from the image, rather than duplicated. In this scenario, a region of the background is copied and pasted over a cell (or any other object) to cover it.\n\nIf you adjust the image colors, it makes it easier to verify that the background has been copied over the cell. By looking at the differences in the neighborhood (surrounding area) of that operation, you can often find artifacts of the forgery.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2Fddccd5dd9bc1e504b2a8c08f4f22d657%2Finbox_11373480_34c9647262711972cd65240743dcdcd8_inbox_11373480_151df79a678026afe02902ece04bd009_background-forgery.png?generation=1763374959900741&alt=media)\n\nHope that helps!",
    "3350810": "Hello,\n\n1. In the first example, a few of the red-circled nuclei seem unique and do not match any duplicated region in the image. I’m unsure whether the red markings represent only the pasted (forged) area or both the original source and the pasted region. ![here](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14946445%2F2f9b44714a11a8c784c334f5112f8fe8%2FScreenshot%202025-11-27%20210353.png?generation=1764277086315160&alt=media)\n\n2. Can the pasted forged region be cropped from the source or only limited to: scaling, rotation, flipping and combination  as you mentioned?"
  },
  "source": "meta"
}