{
  "id": 641092,
  "title": "The competition metric is extremely unfair. ",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/641092",
  "author_name": "",
  "post_date": "2025-11-26T14:36:34.559176800Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Question to competition host: Do you really want it like this? Or is there an error in the competion metric?</p>\n<p>If there are two or more instances of masks, there is a huge penalty in the metric, if you predict all mask correctly, but not in different instances.</p>\n<p>Example:</p>\n<p>mask = mask with two instances (shape = '[2, 529, 696]')</p>\n<p>put all masks in one (instance) mask:          melted = mask[0,:,:] + mask[1,:,:]</p>\n<pre><code>     evaluate_single_image(rle_encode(mask), rle_encode(melted), shape) =  !!!!!\n</code></pre>\n<p>When you add a dimension it gets slightly better:</p>\n<pre><code>     evaluate_single_image(rle_encode(mask), rle_encode(np.expand_dims(melted, axis=)), shape) =  !!!!!\n</code></pre>\n<p>I think this is the main reason for the bad results in the leaderboard. It is much easier to predict 'authentic' for a score of 0.303 than to predict a perfect mask with incorrect differentiation in instances.</p>\n<p>I have uploaded a notebook for demonstration.</p>\n<p><a href=\"https://www.kaggle.com/code/jensvannahl/test-metric-recod-ai\" target=\"_blank\">https://www.kaggle.com/code/jensvannahl/test-metric-recod-ai</a></p>",
  "messages": [
    {
      "id": "3349216",
      "postDate": "11/26/2025 14:36:34",
      "content": "<p>Question to competition host: Do you really want it like this? Or is there an error in the competion metric?</p>\n<p>If there are two or more instances of masks, there is a huge penalty in the metric, if you predict all mask correctly, but not in different instances.</p>\n<p>Example:</p>\n<p>mask = mask with two instances (shape = '[2, 529, 696]')</p>\n<p>put all masks in one (instance) mask:          melted = mask[0,:,:] + mask[1,:,:]</p>\n<pre><code>     evaluate_single_image(rle_encode(mask), rle_encode(melted), shape) =  !!!!!\n</code></pre>\n<p>When you add a dimension it gets slightly better:</p>\n<pre><code>     evaluate_single_image(rle_encode(mask), rle_encode(np.expand_dims(melted, axis=)), shape) =  !!!!!\n</code></pre>\n<p>I think this is the main reason for the bad results in the leaderboard. It is much easier to predict 'authentic' for a score of 0.303 than to predict a perfect mask with incorrect differentiation in instances.</p>\n<p>I have uploaded a notebook for demonstration.</p>\n<p><a href=\"https://www.kaggle.com/code/jensvannahl/test-metric-recod-ai\" target=\"_blank\">https://www.kaggle.com/code/jensvannahl/test-metric-recod-ai</a></p>",
      "rawMarkdown": "Question to competition host: Do you really want it like this? Or is there an error in the competion metric?\n\nIf there are two or more instances of masks, there is a huge penalty in the metric, if you predict all mask correctly, but not in different instances.\n\nExample:\n\nmask = mask with two instances (shape = '[2, 529, 696]')\n\nput all masks in one (instance) mask:          melted = mask[0,:,:] + mask[1,:,:]\n\n         evaluate_single_image(rle_encode(mask), rle_encode(melted), shape) = 0 !!!!!\n\nWhen you add a dimension it gets slightly better:\n\n         evaluate_single_image(rle_encode(mask), rle_encode(np.expand_dims(melted, axis=0)), shape) = 0.34 !!!!!\n\nI think this is the main reason for the bad results in the leaderboard. It is much easier to predict 'authentic' for a score of 0.303 than to predict a perfect mask with incorrect differentiation in instances.\n\nI have uploaded a notebook for demonstration.\n\nhttps://www.kaggle.com/code/jensvannahl/test-metric-recod-ai",
      "votes": null
    },
    {
      "id": "3350992",
      "postDate": "11/28/2025 04:20:09",
      "content": "<p>This post just made me realize that I have been submitting \"melted\" version of all my masks. What the heck. This might be the most important post-processing you can do to your outputs to improve your score. Look at your output segmentation, identify duplicate groups, and transform it from a (1, H, W) submission to a (n, H, W) for each duplicate group you identify.</p>",
      "rawMarkdown": "This post just made me realize that I have been submitting \"melted\" version of all my masks. What the heck. This might be the most important post-processing you can do to your outputs to improve your score. Look at your output segmentation, identify duplicate groups, and transform it from a (1, H, W) submission to a (n, H, W) for each duplicate group you identify.",
      "votes": null
    },
    {
      "id": "3351437",
      "postDate": "11/28/2025 12:07:56",
      "content": "<p>Update: for my sub, not so big gains, only like .001 to .002. Also OP, the input to competition metric rle_encode expects something in the shape of (c, H, W). So you should always pass in with expand_dims.  Take a look at the rle_string with and without.</p>\n<p>Note that the popular public notebook redefines rle_encode to work with a (H, W) input instead</p>",
      "rawMarkdown": "Update: for my sub, not so big gains, only like .001 to .002. Also OP, the input to competition metric rle_encode expects something in the shape of (c, H, W). So you should always pass in with expand_dims.  Take a look at the rle_string with and without.\n\nNote that the popular public notebook redefines rle_encode to work with a (H, W) input instead",
      "votes": null
    },
    {
      "id": "3351951",
      "postDate": "11/28/2025 18:50:34",
      "content": "<p>There might be not a lot of duplicated groups. That is why the gain is very small. I also have a very small gain. </p>",
      "rawMarkdown": "There might be not a lot of duplicated groups. That is why the gain is very small. I also have a very small gain.",
      "votes": null
    },
    {
      "id": "3352095",
      "postDate": "11/28/2025 22:28:44",
      "content": "<p>rle_encode function expects masks is list[npt.NDArray]</p>\n<p>the prediction should be in 3 dimensions, 2 dimensions just causes score to 0, and actually the function should throw error.</p>",
      "rawMarkdown": "rle_encode function expects masks is list[npt.NDArray]\n\nthe prediction should be in 3 dimensions, 2 dimensions just causes score to 0, and actually the function should throw error.",
      "votes": null
    },
    {
      "id": "3353434",
      "postDate": "11/29/2025 18:23:30",
      "content": "<p>Actually for my sub not so big gains, too. Maybe there are not so much masks with more than 1 instance in the test set as in the training set.\nHow did you manage to transform ‚melted‘ versions of masks to different instances? I trained a Siamese Network with Triplet loss, which achieved 88%accuracy.</p>",
      "rawMarkdown": "Actually for my sub not so big gains, too. Maybe there are not so much masks with more than 1 instance in the test set as in the training set.\nHow did you manage to transform ‚melted‘ versions of masks to different instances? I trained a Siamese Network with Triplet loss, which achieved 88%accuracy.",
      "votes": null
    },
    {
      "id": "3353671",
      "postDate": "11/29/2025 21:25:11",
      "content": "<p>May I ask how you used Siamese network for binary segmentation?</p>",
      "rawMarkdown": "May I ask how you used Siamese network for binary segmentation?",
      "votes": null
    },
    {
      "id": "3354754",
      "postDate": "11/30/2025 12:35:37",
      "content": "<p>I used it for the conversion from 'flat' masks with shape (1, w, h) into instance masks with shape (number of instances, w, h).\nThe Siamese network is feeded with the masks of the training set. It extracts each single mask, takes one mask as anchor, another mask from the same instance as positive, and a mask from a different instance as negative. \nAt inference time, you take the 'flat' mask and can measure the similarity by calculating the euclidean distance between each single mask.</p>",
      "rawMarkdown": "I used it for the conversion from 'flat' masks with shape (1, w, h) into instance masks with shape (number of instances, w, h).\nThe Siamese network is feeded with the masks of the training set. It extracts each single mask, takes one mask as anchor, another mask from the same instance as positive, and a mask from a different instance as negative. \nAt inference time, you take the 'flat' mask and can measure the similarity by calculating the euclidean distance between each single mask.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3350992,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "11/28/2025 04:20:09",
      "content": "<p>This post just made me realize that I have been submitting \"melted\" version of all my masks. What the heck. This might be the most important post-processing you can do to your outputs to improve your score. Look at your output segmentation, identify duplicate groups, and transform it from a (1, H, W) submission to a (n, H, W) for each duplicate group you identify.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3351437,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "11/28/2025 12:07:56",
          "content": "<p>Update: for my sub, not so big gains, only like .001 to .002. Also OP, the input to competition metric rle_encode expects something in the shape of (c, H, W). So you should always pass in with expand_dims.  Take a look at the rle_string with and without.</p>\n<p>Note that the popular public notebook redefines rle_encode to work with a (H, W) input instead</p>",
          "votes": null,
          "replies": [
            {
              "id": 3351951,
              "author_name": "hamzahabduljalil",
              "author_url": "",
              "post_date": "11/28/2025 18:50:34",
              "content": "<p>There might be not a lot of duplicated groups. That is why the gain is very small. I also have a very small gain. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3353434,
              "author_name": "jensvannahl",
              "author_url": "",
              "post_date": "11/29/2025 18:23:30",
              "content": "<p>Actually for my sub not so big gains, too. Maybe there are not so much masks with more than 1 instance in the test set as in the training set.\nHow did you manage to transform ‚melted‘ versions of masks to different instances? I trained a Siamese Network with Triplet loss, which achieved 88%accuracy.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3353671,
                  "author_name": "hamzahabduljalil",
                  "author_url": "",
                  "post_date": "11/29/2025 21:25:11",
                  "content": "<p>May I ask how you used Siamese network for binary segmentation?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3354754,
                      "author_name": "jensvannahl",
                      "author_url": "",
                      "post_date": "11/30/2025 12:35:37",
                      "content": "<p>I used it for the conversion from 'flat' masks with shape (1, w, h) into instance masks with shape (number of instances, w, h).\nThe Siamese network is feeded with the masks of the training set. It extracts each single mask, takes one mask as anchor, another mask from the same instance as positive, and a mask from a different instance as negative. \nAt inference time, you take the 'flat' mask and can measure the similarity by calculating the euclidean distance between each single mask.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3352095,
      "author_name": "xuyuan",
      "author_url": "",
      "post_date": "11/28/2025 22:28:44",
      "content": "<p>rle_encode function expects masks is list[npt.NDArray]</p>\n<p>the prediction should be in 3 dimensions, 2 dimensions just causes score to 0, and actually the function should throw error.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3349216": "Question to competition host: Do you really want it like this? Or is there an error in the competion metric?\n\nIf there are two or more instances of masks, there is a huge penalty in the metric, if you predict all mask correctly, but not in different instances.\n\nExample:\n\nmask = mask with two instances (shape = '[2, 529, 696]')\n\nput all masks in one (instance) mask:          melted = mask[0,:,:] + mask[1,:,:]\n\n         evaluate_single_image(rle_encode(mask), rle_encode(melted), shape) = 0 !!!!!\n\nWhen you add a dimension it gets slightly better:\n\n         evaluate_single_image(rle_encode(mask), rle_encode(np.expand_dims(melted, axis=0)), shape) = 0.34 !!!!!\n\nI think this is the main reason for the bad results in the leaderboard. It is much easier to predict 'authentic' for a score of 0.303 than to predict a perfect mask with incorrect differentiation in instances.\n\nI have uploaded a notebook for demonstration.\n\nhttps://www.kaggle.com/code/jensvannahl/test-metric-recod-ai",
    "3350992": "This post just made me realize that I have been submitting \"melted\" version of all my masks. What the heck. This might be the most important post-processing you can do to your outputs to improve your score. Look at your output segmentation, identify duplicate groups, and transform it from a (1, H, W) submission to a (n, H, W) for each duplicate group you identify.",
    "3351437": "Update: for my sub, not so big gains, only like .001 to .002. Also OP, the input to competition metric rle_encode expects something in the shape of (c, H, W). So you should always pass in with expand_dims.  Take a look at the rle_string with and without.\n\nNote that the popular public notebook redefines rle_encode to work with a (H, W) input instead",
    "3351951": "There might be not a lot of duplicated groups. That is why the gain is very small. I also have a very small gain.",
    "3352095": "rle_encode function expects masks is list[npt.NDArray]\n\nthe prediction should be in 3 dimensions, 2 dimensions just causes score to 0, and actually the function should throw error.",
    "3353434": "Actually for my sub not so big gains, too. Maybe there are not so much masks with more than 1 instance in the test set as in the training set.\nHow did you manage to transform ‚melted‘ versions of masks to different instances? I trained a Siamese Network with Triplet loss, which achieved 88%accuracy.",
    "3353671": "May I ask how you used Siamese network for binary segmentation?",
    "3354754": "I used it for the conversion from 'flat' masks with shape (1, w, h) into instance masks with shape (number of instances, w, h).\nThe Siamese network is feeded with the masks of the training set. It extracts each single mask, takes one mask as anchor, another mask from the same instance as positive, and a mask from a different instance as negative. \nAt inference time, you take the 'flat' mask and can measure the similarity by calculating the euclidean distance between each single mask."
  },
  "source": "meta"
}