{
  "id": 665875,
  "title": "Help needed: RLE submission format check (multiple connected components + authentic example)",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/665875",
  "author_name": "",
  "post_date": "2026-01-04T08:17:48.128600200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,\nI’m having trouble confirming the correct RLE format for this competition’s submission.</p>\n<p>Below is a snippet from my submission CSV as an example. It includes:</p>\n<p>one sample with three connected components in the mask (so the RLE contains multiple segments), and</p>\n<p>one authentic sample (no forged region).</p>\n<p>Could someone please help me verify whether my submission format is correct? Also, is it possible that the submission format description on the competition page is ambiguous or contains a mistake?</p>\n<p>case_id,annotation\n1,authentic\n2,\"[123 4]\"</p>\n<p>Here is my CSV example (numbers are from my actual submission):</p>\n<p>case_id,annotation\n13977,\"[1500356, 110, 1501738, 110, 1503120, 110, 1504502, 110, 1505884, 110, 1507266, 110, 1508648, 110, 1510030, 110, 1511412, 110, 1512794, 110, 1514176, 110, 1515558, 110, 1516940, 110, 1518322, 110, 1519704, 110, 1521086, 110, 1522468, 110, 1523850, 110];[1428492, 101, 1429874, 101, 1431256, 101, 1432638, 101, 1434020, 101, 1435402, 101, 1436784, 101, 1438166, 101, 1439548, 101, 1440930, 101, 1442312, 101, 1443694, 101, 1445076, 101, 1446458, 101, 1447840, 101, 1449222, 101, 1450604, 101, 1451986, 101, 1453368, 101, 1454750, 101, 1456132, 101, 1457514, 101, 1458896, 101, 1460278, 101, 1461660, 101, 1463042, 101, 1464424, 101, 1465806, 101, 1467188, 101, 1468570, 101, 1469952, 101, 1471334, 101, 1472716, 101, 1474098, 101, 1475480, 101];[627074, 99, 628456, 99, 629838, 99, 631220, 99, 632602, 99, 633984, 99, 635366, 99, 636748, 99, 638130, 99, 639512, 99, 640894, 99, 642276, 99, 643658, 99, 645040, 99, 646422, 99, 647804, 99, 649186, 99, 650568, 99, 651950, 99, 653332, 99, 654714, 99, 656096, 99, 657478, 99, 658860, 99, 660242, 99, 661624, 99, 663006, 99, 664388, 99, 665770, 99, 667152, 99, 668534, 99, 669916, 99, 671298, 99, 672680, 99, 674062, 99, 675444, 99, 676826, 99, 678208, 99]\"\n15257,authentic</p>\n<p>Thanks a lot!</p>",
  "messages": [
    {
      "id": "3385909",
      "postDate": "01/04/2026 08:17:48",
      "content": "<p>Hi everyone,\nI’m having trouble confirming the correct RLE format for this competition’s submission.</p>\n<p>Below is a snippet from my submission CSV as an example. It includes:</p>\n<p>one sample with three connected components in the mask (so the RLE contains multiple segments), and</p>\n<p>one authentic sample (no forged region).</p>\n<p>Could someone please help me verify whether my submission format is correct? Also, is it possible that the submission format description on the competition page is ambiguous or contains a mistake?</p>\n<p>case_id,annotation\n1,authentic\n2,\"[123 4]\"</p>\n<p>Here is my CSV example (numbers are from my actual submission):</p>\n<p>case_id,annotation\n13977,\"[1500356, 110, 1501738, 110, 1503120, 110, 1504502, 110, 1505884, 110, 1507266, 110, 1508648, 110, 1510030, 110, 1511412, 110, 1512794, 110, 1514176, 110, 1515558, 110, 1516940, 110, 1518322, 110, 1519704, 110, 1521086, 110, 1522468, 110, 1523850, 110];[1428492, 101, 1429874, 101, 1431256, 101, 1432638, 101, 1434020, 101, 1435402, 101, 1436784, 101, 1438166, 101, 1439548, 101, 1440930, 101, 1442312, 101, 1443694, 101, 1445076, 101, 1446458, 101, 1447840, 101, 1449222, 101, 1450604, 101, 1451986, 101, 1453368, 101, 1454750, 101, 1456132, 101, 1457514, 101, 1458896, 101, 1460278, 101, 1461660, 101, 1463042, 101, 1464424, 101, 1465806, 101, 1467188, 101, 1468570, 101, 1469952, 101, 1471334, 101, 1472716, 101, 1474098, 101, 1475480, 101];[627074, 99, 628456, 99, 629838, 99, 631220, 99, 632602, 99, 633984, 99, 635366, 99, 636748, 99, 638130, 99, 639512, 99, 640894, 99, 642276, 99, 643658, 99, 645040, 99, 646422, 99, 647804, 99, 649186, 99, 650568, 99, 651950, 99, 653332, 99, 654714, 99, 656096, 99, 657478, 99, 658860, 99, 660242, 99, 661624, 99, 663006, 99, 664388, 99, 665770, 99, 667152, 99, 668534, 99, 669916, 99, 671298, 99, 672680, 99, 674062, 99, 675444, 99, 676826, 99, 678208, 99]\"\n15257,authentic</p>\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Hi everyone,\nI’m having trouble confirming the correct RLE format for this competition’s submission.\n\nBelow is a snippet from my submission CSV as an example. It includes:\n\none sample with three connected components in the mask (so the RLE contains multiple segments), and\n\none authentic sample (no forged region).\n\nCould someone please help me verify whether my submission format is correct? Also, is it possible that the submission format description on the competition page is ambiguous or contains a mistake?\n\ncase_id,annotation\n1,authentic\n2,\"[123 4]\"\n\nHere is my CSV example (numbers are from my actual submission):\n\ncase_id,annotation\n13977,\"[1500356, 110, 1501738, 110, 1503120, 110, 1504502, 110, 1505884, 110, 1507266, 110, 1508648, 110, 1510030, 110, 1511412, 110, 1512794, 110, 1514176, 110, 1515558, 110, 1516940, 110, 1518322, 110, 1519704, 110, 1521086, 110, 1522468, 110, 1523850, 110];[1428492, 101, 1429874, 101, 1431256, 101, 1432638, 101, 1434020, 101, 1435402, 101, 1436784, 101, 1438166, 101, 1439548, 101, 1440930, 101, 1442312, 101, 1443694, 101, 1445076, 101, 1446458, 101, 1447840, 101, 1449222, 101, 1450604, 101, 1451986, 101, 1453368, 101, 1454750, 101, 1456132, 101, 1457514, 101, 1458896, 101, 1460278, 101, 1461660, 101, 1463042, 101, 1464424, 101, 1465806, 101, 1467188, 101, 1468570, 101, 1469952, 101, 1471334, 101, 1472716, 101, 1474098, 101, 1475480, 101];[627074, 99, 628456, 99, 629838, 99, 631220, 99, 632602, 99, 633984, 99, 635366, 99, 636748, 99, 638130, 99, 639512, 99, 640894, 99, 642276, 99, 643658, 99, 645040, 99, 646422, 99, 647804, 99, 649186, 99, 650568, 99, 651950, 99, 653332, 99, 654714, 99, 656096, 99, 657478, 99, 658860, 99, 660242, 99, 661624, 99, 663006, 99, 664388, 99, 665770, 99, 667152, 99, 668534, 99, 669916, 99, 671298, 99, 672680, 99, 674062, 99, 675444, 99, 676826, 99, 678208, 99]\"\n15257,authentic\n\nThanks a lot!",
      "votes": null
    },
    {
      "id": "3386049",
      "postDate": "01/04/2026 13:48:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sch01ar\" target=\"_blank\">@sch01ar</a> ,</p>\n<p>1- Submission Format:</p>\n<p>When properly parsed with newlines, your submission format is correct </p>\n<ul>\n<li>authentic: This is the correct label for images predicted with no problems.</li>\n<li>RLE Format: Your example for case 13977 correctly uses semicolons (;) to separate the distinct masks (channels).</li>\n</ul>\n<p>⚠️ If you are ever in doubt about the RLE submission file format, check the official implementation here: <a href=\"https://www.kaggle.com/code/metric/recodai-f1/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/recodai-f1/notebook</a></p>\n<p>2- Metric calculation:</p>\n<p>As this could also be a doubt from other participants, I will include some more info here:</p>\n<p>This is the side-by-side visualization of your prediction and the annotation for case 13977:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F5b7e7398ca257b925f92e3b5d35657d8%2Fvis.png?generation=1767533547596680&amp;alt=media\" alt=\"\"></p>\n<p>Note that your CSV/Prediction has 3 channels (visualized as Lime, Magenta, and Orange), whereas the Ground Truth (GT) has only 2 channels (visualized as Red and Cyan).</p>\n<p>The metric optimizes the maximum F1 score between the GT and Prediction channels. It matches each GT channel to the single best Prediction channel.</p>\n<p>In your example, the Magenta (Ch1) and Orange (Ch2) predictions appear to correspond to the same forgery instance (the Red GT channel).</p>\n<p>Because they are split into separate channels in your prediction, the metric likely matches only one of them (e.g., Orange) to the RED GT and ignores the other.</p>\n<p>In this case, your score would double if you merge the source/copies channels that belong to the same forgery instance.</p>\n<p>3- Overview Page Sample:</p>\n<p>You are correct. The sample on the overview page is missing a comma. It should be:</p>\n<pre><code>case_id,annotation\n1,authentic\n2,\"[123, 4]\"\n</code></pre>\n<p>Thanks for pointing that out!</p>\n<p>Hope that helps :)</p>",
      "rawMarkdown": "Hi @sch01ar ,\n\n1- Submission Format:\n\nWhen properly parsed with newlines, your submission format is correct \n- authentic: This is the correct label for images predicted with no problems.\n- RLE Format: Your example for case 13977 correctly uses semicolons (;) to separate the distinct masks (channels).\n\n⚠️ If you are ever in doubt about the RLE submission file format, check the official implementation here: https://www.kaggle.com/code/metric/recodai-f1/notebook\n\n2- Metric calculation:\n\nAs this could also be a doubt from other participants, I will include some more info here:\n\nThis is the side-by-side visualization of your prediction and the annotation for case 13977:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F5b7e7398ca257b925f92e3b5d35657d8%2Fvis.png?generation=1767533547596680&alt=media)\n\nNote that your CSV/Prediction has 3 channels (visualized as Lime, Magenta, and Orange), whereas the Ground Truth (GT) has only 2 channels (visualized as Red and Cyan).\n\nThe metric optimizes the maximum F1 score between the GT and Prediction channels. It matches each GT channel to the single best Prediction channel.\n\nIn your example, the Magenta (Ch1) and Orange (Ch2) predictions appear to correspond to the same forgery instance (the Red GT channel).\n\nBecause they are split into separate channels in your prediction, the metric likely matches only one of them (e.g., Orange) to the RED GT and ignores the other.\n\nIn this case, your score would double if you merge the source/copies channels that belong to the same forgery instance.\n\n3- Overview Page Sample:\n\n You are correct. The sample on the overview page is missing a comma. It should be:\n\n```\ncase_id,annotation\n1,authentic\n2,\"[123, 4]\"\n```\n\nThanks for pointing that out!\n\nHope that helps :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3386049,
      "author_name": "joophillipecardenuto",
      "author_url": "",
      "post_date": "01/04/2026 13:48:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sch01ar\" target=\"_blank\">@sch01ar</a> ,</p>\n<p>1- Submission Format:</p>\n<p>When properly parsed with newlines, your submission format is correct </p>\n<ul>\n<li>authentic: This is the correct label for images predicted with no problems.</li>\n<li>RLE Format: Your example for case 13977 correctly uses semicolons (;) to separate the distinct masks (channels).</li>\n</ul>\n<p>⚠️ If you are ever in doubt about the RLE submission file format, check the official implementation here: <a href=\"https://www.kaggle.com/code/metric/recodai-f1/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/recodai-f1/notebook</a></p>\n<p>2- Metric calculation:</p>\n<p>As this could also be a doubt from other participants, I will include some more info here:</p>\n<p>This is the side-by-side visualization of your prediction and the annotation for case 13977:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F5b7e7398ca257b925f92e3b5d35657d8%2Fvis.png?generation=1767533547596680&amp;alt=media\" alt=\"\"></p>\n<p>Note that your CSV/Prediction has 3 channels (visualized as Lime, Magenta, and Orange), whereas the Ground Truth (GT) has only 2 channels (visualized as Red and Cyan).</p>\n<p>The metric optimizes the maximum F1 score between the GT and Prediction channels. It matches each GT channel to the single best Prediction channel.</p>\n<p>In your example, the Magenta (Ch1) and Orange (Ch2) predictions appear to correspond to the same forgery instance (the Red GT channel).</p>\n<p>Because they are split into separate channels in your prediction, the metric likely matches only one of them (e.g., Orange) to the RED GT and ignores the other.</p>\n<p>In this case, your score would double if you merge the source/copies channels that belong to the same forgery instance.</p>\n<p>3- Overview Page Sample:</p>\n<p>You are correct. The sample on the overview page is missing a comma. It should be:</p>\n<pre><code>case_id,annotation\n1,authentic\n2,\"[123, 4]\"\n</code></pre>\n<p>Thanks for pointing that out!</p>\n<p>Hope that helps :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3385909": "Hi everyone,\nI’m having trouble confirming the correct RLE format for this competition’s submission.\n\nBelow is a snippet from my submission CSV as an example. It includes:\n\none sample with three connected components in the mask (so the RLE contains multiple segments), and\n\none authentic sample (no forged region).\n\nCould someone please help me verify whether my submission format is correct? Also, is it possible that the submission format description on the competition page is ambiguous or contains a mistake?\n\ncase_id,annotation\n1,authentic\n2,\"[123 4]\"\n\nHere is my CSV example (numbers are from my actual submission):\n\ncase_id,annotation\n13977,\"[1500356, 110, 1501738, 110, 1503120, 110, 1504502, 110, 1505884, 110, 1507266, 110, 1508648, 110, 1510030, 110, 1511412, 110, 1512794, 110, 1514176, 110, 1515558, 110, 1516940, 110, 1518322, 110, 1519704, 110, 1521086, 110, 1522468, 110, 1523850, 110];[1428492, 101, 1429874, 101, 1431256, 101, 1432638, 101, 1434020, 101, 1435402, 101, 1436784, 101, 1438166, 101, 1439548, 101, 1440930, 101, 1442312, 101, 1443694, 101, 1445076, 101, 1446458, 101, 1447840, 101, 1449222, 101, 1450604, 101, 1451986, 101, 1453368, 101, 1454750, 101, 1456132, 101, 1457514, 101, 1458896, 101, 1460278, 101, 1461660, 101, 1463042, 101, 1464424, 101, 1465806, 101, 1467188, 101, 1468570, 101, 1469952, 101, 1471334, 101, 1472716, 101, 1474098, 101, 1475480, 101];[627074, 99, 628456, 99, 629838, 99, 631220, 99, 632602, 99, 633984, 99, 635366, 99, 636748, 99, 638130, 99, 639512, 99, 640894, 99, 642276, 99, 643658, 99, 645040, 99, 646422, 99, 647804, 99, 649186, 99, 650568, 99, 651950, 99, 653332, 99, 654714, 99, 656096, 99, 657478, 99, 658860, 99, 660242, 99, 661624, 99, 663006, 99, 664388, 99, 665770, 99, 667152, 99, 668534, 99, 669916, 99, 671298, 99, 672680, 99, 674062, 99, 675444, 99, 676826, 99, 678208, 99]\"\n15257,authentic\n\nThanks a lot!",
    "3386049": "Hi @sch01ar ,\n\n1- Submission Format:\n\nWhen properly parsed with newlines, your submission format is correct \n- authentic: This is the correct label for images predicted with no problems.\n- RLE Format: Your example for case 13977 correctly uses semicolons (;) to separate the distinct masks (channels).\n\n⚠️ If you are ever in doubt about the RLE submission file format, check the official implementation here: https://www.kaggle.com/code/metric/recodai-f1/notebook\n\n2- Metric calculation:\n\nAs this could also be a doubt from other participants, I will include some more info here:\n\nThis is the side-by-side visualization of your prediction and the annotation for case 13977:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11373480%2F5b7e7398ca257b925f92e3b5d35657d8%2Fvis.png?generation=1767533547596680&alt=media)\n\nNote that your CSV/Prediction has 3 channels (visualized as Lime, Magenta, and Orange), whereas the Ground Truth (GT) has only 2 channels (visualized as Red and Cyan).\n\nThe metric optimizes the maximum F1 score between the GT and Prediction channels. It matches each GT channel to the single best Prediction channel.\n\nIn your example, the Magenta (Ch1) and Orange (Ch2) predictions appear to correspond to the same forgery instance (the Red GT channel).\n\nBecause they are split into separate channels in your prediction, the metric likely matches only one of them (e.g., Orange) to the RED GT and ignores the other.\n\nIn this case, your score would double if you merge the source/copies channels that belong to the same forgery instance.\n\n3- Overview Page Sample:\n\n You are correct. The sample on the overview page is missing a comma. It should be:\n\n```\ncase_id,annotation\n1,authentic\n2,\"[123, 4]\"\n```\n\nThanks for pointing that out!\n\nHope that helps :)"
  },
  "source": "meta"
}