{
  "id": 413664,
  "title": "Why did I get very low public LB(around 0.03) which was not consistent with the real performance?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/413664",
  "author_name": "",
  "post_date": "2023-05-29T17:09:20.671775700Z",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I submitted with this rle function, but got LB score 0.03, which was not consistent with the real performance of my model.</p>\n<pre><code> ():\n    \n    pixels = img.flatten()\n    \n\n    pixels = np.concatenate([[], pixels, []])\n    runs = np.where(pixels[:] != pixels[:-])[] + \n    runs[::] -= runs[::]\n     .join((x)  x  runs)\n</code></pre>\n<p>I have tried to transpose my mask_pred array before using the RLE function as well, but still got the same score of 0.03.  Is there anyone who encountered the same issue, and any suggestions? Thanks a lot for your help!!</p>",
  "messages": [
    {
      "id": "2279832",
      "postDate": "05/29/2023 17:09:20",
      "content": "<p>I submitted with this rle function, but got LB score 0.03, which was not consistent with the real performance of my model.</p>\n<pre><code> ():\n    \n    pixels = img.flatten()\n    \n\n    pixels = np.concatenate([[], pixels, []])\n    runs = np.where(pixels[:] != pixels[:-])[] + \n    runs[::] -= runs[::]\n     .join((x)  x  runs)\n</code></pre>\n<p>I have tried to transpose my mask_pred array before using the RLE function as well, but still got the same score of 0.03.  Is there anyone who encountered the same issue, and any suggestions? Thanks a lot for your help!!</p>",
      "rawMarkdown": "I submitted with this rle function, but got LB score 0.03, which was not consistent with the real performance of my model.\n```python\ndef rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels = img.flatten()\n    # pixels = (pixels >= thr).astype(int)\n    \n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n```\nI have tried to transpose my mask_pred array before using the RLE function as well, but still got the same score of 0.03.  Is there anyone who encountered the same issue, and any suggestions? Thanks a lot for your help!!",
      "votes": null
    },
    {
      "id": "2280169",
      "postDate": "05/29/2023 23:25:01",
      "content": "<p>What happens when you use training mask as an <code>img</code> and then submit it? Does you score is around 0.11?</p>",
      "rawMarkdown": "What happens when you use training mask as an `img` and then submit it? Does you score is around 0.11?",
      "votes": null
    },
    {
      "id": "2280194",
      "postDate": "05/30/2023 00:35:59",
      "content": "<p>What do you mean by training mask, you mean ground truth? Of which fragment?</p>",
      "rawMarkdown": "What do you mean by training mask, you mean ground truth? Of which fragment?",
      "votes": null
    },
    {
      "id": "2280206",
      "postDate": "05/30/2023 01:03:40",
      "content": "<p>Oh, I have tried my RLE function with the binary mask of test fragment and submitted, and got 0.09 score. So maybe the rle function I used is wrong? </p>",
      "rawMarkdown": "Oh, I have tried my RLE function with the binary mask of test fragment and submitted, and got 0.09 score. So maybe the rle function I used is wrong?",
      "votes": null
    },
    {
      "id": "2280461",
      "postDate": "05/30/2023 05:58:26",
      "content": "<p>Yes, looks like it. One from <a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask</a> must work. As well as this one: <a href=\"https://www.kaggle.com/code/yoyobar/3d-resnet-baseline-inference#kln-55\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/3d-resnet-baseline-inference#kln-55</a></p>",
      "rawMarkdown": "Yes, looks like it. One from https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask must work. As well as this one: https://www.kaggle.com/code/yoyobar/3d-resnet-baseline-inference#kln-55",
      "votes": null
    },
    {
      "id": "2280476",
      "postDate": "05/30/2023 06:06:39",
      "content": "<p>Oh, second one is exactly what you posted. So, something is wrong with the way you load and send data to rle function. It may be simple transposition or something more involved.</p>",
      "rawMarkdown": "Oh, second one is exactly what you posted. So, something is wrong with the way you load and send data to rle function. It may be simple transposition or something more involved.",
      "votes": null
    },
    {
      "id": "2280482",
      "postDate": "05/30/2023 06:11:43",
      "content": "<p>E.g. if you image has different dimensions or even different amount of dimensions rle function will not detect that.</p>",
      "rawMarkdown": "E.g. if you image has different dimensions or even different amount of dimensions rle function will not detect that.",
      "votes": null
    },
    {
      "id": "2280644",
      "postDate": "05/30/2023 08:33:08",
      "content": "<p>I will have a try on the code you shared with me, thanks for your help! </p>",
      "rawMarkdown": "I will have a try on the code you shared with me, thanks for your help!",
      "votes": null
    },
    {
      "id": "2281949",
      "postDate": "05/31/2023 08:54:44",
      "content": "<p>I found the problems! I forgot to drop the padding pixels off before doing the RLE. After I made it to the same size with the mask.png, the score made sense to me. Thank you for your kindly help anyway!</p>",
      "rawMarkdown": "I found the problems! I forgot to drop the padding pixels off before doing the RLE. After I made it to the same size with the mask.png, the score made sense to me. Thank you for your kindly help anyway!",
      "votes": null
    },
    {
      "id": "2282717",
      "postDate": "05/31/2023 19:17:07",
      "content": "<p>Glad that issue is resolved</p>",
      "rawMarkdown": "Glad that issue is resolved",
      "votes": null
    },
    {
      "id": "2283283",
      "postDate": "06/01/2023 07:21:01",
      "content": "<p>Its because of the cross validation strategy that you are using</p>",
      "rawMarkdown": "Its because of the cross validation strategy that you are using",
      "votes": null
    },
    {
      "id": "2294010",
      "postDate": "06/09/2023 16:50:15",
      "content": "<p>Dear All,</p>\n<p>I appreciate any feedback on this topic of \"visually good test results but low LB score\". </p>\n<p>I have been struggling for a long time on this problem but made no progress. </p>\n<p>I tried a couple of cross-validated models and my pipeline does not add padding.  </p>\n<p>I did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook (<a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask)\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask)</a>, my submission still received very low public score (0.06).</p>\n<p>On the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial). </p>\n<p>Comparing the attached image with the result generated by the 0-11 simplest submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?). </p>\n<p>Any suggestions please?</p>\n<p>Thanks very much in advance for your help and insights!</p>",
      "rawMarkdown": "Dear All,\n\nI appreciate any feedback on this topic of \"visually good test results but low LB score\". \n\nI have been struggling for a long time on this problem but made no progress. \n\nI tried a couple of cross-validated models and my pipeline does not add padding.  \n\nI did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook (https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask), my submission still received very low public score (0.06).\n\nOn the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial). \n\nComparing the attached image with the result generated by the 0-11 simplest submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?). \n\nAny suggestions please?\n\nThanks very much in advance for your help and insights!",
      "votes": null
    },
    {
      "id": "2294078",
      "postDate": "06/09/2023 18:01:21",
      "content": "<p>Does you model use only specific layers for making prediction? Test data can have relevant information on different layers then training data.</p>",
      "rawMarkdown": "Does you model use only specific layers for making prediction? Test data can have relevant information on different layers then training data.",
      "votes": null
    },
    {
      "id": "2294106",
      "postDate": "06/09/2023 18:29:17",
      "content": "<p>My model samples 11 to 15 of the 65 slices.</p>\n<p>Is it correct that the dummy \"test a\" and \"test b\" are substantially different than the real test data? </p>\n<p>I also wonder if the LB score is computed by using the mask decoded from the RLE submission, or calculates the LB score directly using the RLE string.</p>\n<p>Thanks so much for getting back, Serhii.</p>",
      "rawMarkdown": "My model samples 11 to 15 of the 65 slices.\n\nIs it correct that the dummy \"test a\" and \"test b\" are substantially different than the real test data? \n\nI also wonder if the LB score is computed by using the mask decoded from the RLE submission, or calculates the LB score directly using the RLE string.\n\nThanks so much for getting back, Serhii.",
      "votes": null
    },
    {
      "id": "2294113",
      "postDate": "06/09/2023 18:43:37",
      "content": "<p>I'm not sure if mask is getting subtracted. This actually would be a good experiment </p>",
      "rawMarkdown": "I'm not sure if mask is getting subtracted. This actually would be a good experiment",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2280169,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "05/29/2023 23:25:01",
      "content": "<p>What happens when you use training mask as an <code>img</code> and then submit it? Does you score is around 0.11?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2280194,
          "author_name": "keanehuang",
          "author_url": "",
          "post_date": "05/30/2023 00:35:59",
          "content": "<p>What do you mean by training mask, you mean ground truth? Of which fragment?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2280206,
          "author_name": "keanehuang",
          "author_url": "",
          "post_date": "05/30/2023 01:03:40",
          "content": "<p>Oh, I have tried my RLE function with the binary mask of test fragment and submitted, and got 0.09 score. So maybe the rle function I used is wrong? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2280461,
              "author_name": "elvenmonk",
              "author_url": "",
              "post_date": "05/30/2023 05:58:26",
              "content": "<p>Yes, looks like it. One from <a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask</a> must work. As well as this one: <a href=\"https://www.kaggle.com/code/yoyobar/3d-resnet-baseline-inference#kln-55\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/3d-resnet-baseline-inference#kln-55</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2280476,
                  "author_name": "elvenmonk",
                  "author_url": "",
                  "post_date": "05/30/2023 06:06:39",
                  "content": "<p>Oh, second one is exactly what you posted. So, something is wrong with the way you load and send data to rle function. It may be simple transposition or something more involved.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2280482,
                      "author_name": "elvenmonk",
                      "author_url": "",
                      "post_date": "05/30/2023 06:11:43",
                      "content": "<p>E.g. if you image has different dimensions or even different amount of dimensions rle function will not detect that.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2280644,
                          "author_name": "keanehuang",
                          "author_url": "",
                          "post_date": "05/30/2023 08:33:08",
                          "content": "<p>I will have a try on the code you shared with me, thanks for your help! </p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 2281949,
                          "author_name": "keanehuang",
                          "author_url": "",
                          "post_date": "05/31/2023 08:54:44",
                          "content": "<p>I found the problems! I forgot to drop the padding pixels off before doing the RLE. After I made it to the same size with the mask.png, the score made sense to me. Thank you for your kindly help anyway!</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2282717,
                              "author_name": "elvenmonk",
                              "author_url": "",
                              "post_date": "05/31/2023 19:17:07",
                              "content": "<p>Glad that issue is resolved</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2283283,
      "author_name": "vishakkbhat",
      "author_url": "",
      "post_date": "06/01/2023 07:21:01",
      "content": "<p>Its because of the cross validation strategy that you are using</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2294010,
      "author_name": "vcolliym",
      "author_url": "",
      "post_date": "06/09/2023 16:50:15",
      "content": "<p>Dear All,</p>\n<p>I appreciate any feedback on this topic of \"visually good test results but low LB score\". </p>\n<p>I have been struggling for a long time on this problem but made no progress. </p>\n<p>I tried a couple of cross-validated models and my pipeline does not add padding.  </p>\n<p>I did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook (<a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask)\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask)</a>, my submission still received very low public score (0.06).</p>\n<p>On the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial). </p>\n<p>Comparing the attached image with the result generated by the 0-11 simplest submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?). </p>\n<p>Any suggestions please?</p>\n<p>Thanks very much in advance for your help and insights!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2294078,
          "author_name": "elvenmonk",
          "author_url": "",
          "post_date": "06/09/2023 18:01:21",
          "content": "<p>Does you model use only specific layers for making prediction? Test data can have relevant information on different layers then training data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2294106,
              "author_name": "vcolliym",
              "author_url": "",
              "post_date": "06/09/2023 18:29:17",
              "content": "<p>My model samples 11 to 15 of the 65 slices.</p>\n<p>Is it correct that the dummy \"test a\" and \"test b\" are substantially different than the real test data? </p>\n<p>I also wonder if the LB score is computed by using the mask decoded from the RLE submission, or calculates the LB score directly using the RLE string.</p>\n<p>Thanks so much for getting back, Serhii.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2294113,
                  "author_name": "elvenmonk",
                  "author_url": "",
                  "post_date": "06/09/2023 18:43:37",
                  "content": "<p>I'm not sure if mask is getting subtracted. This actually would be a good experiment </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2279832": "I submitted with this rle function, but got LB score 0.03, which was not consistent with the real performance of my model.\n```python\ndef rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels = img.flatten()\n    # pixels = (pixels >= thr).astype(int)\n    \n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n```\nI have tried to transpose my mask_pred array before using the RLE function as well, but still got the same score of 0.03.  Is there anyone who encountered the same issue, and any suggestions? Thanks a lot for your help!!",
    "2280169": "What happens when you use training mask as an `img` and then submit it? Does you score is around 0.11?",
    "2280194": "What do you mean by training mask, you mean ground truth? Of which fragment?",
    "2280206": "Oh, I have tried my RLE function with the binary mask of test fragment and submitted, and got 0.09 score. So maybe the rle function I used is wrong?",
    "2280461": "Yes, looks like it. One from https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask must work. As well as this one: https://www.kaggle.com/code/yoyobar/3d-resnet-baseline-inference#kln-55",
    "2280476": "Oh, second one is exactly what you posted. So, something is wrong with the way you load and send data to rle function. It may be simple transposition or something more involved.",
    "2280482": "E.g. if you image has different dimensions or even different amount of dimensions rle function will not detect that.",
    "2280644": "I will have a try on the code you shared with me, thanks for your help!",
    "2281949": "I found the problems! I forgot to drop the padding pixels off before doing the RLE. After I made it to the same size with the mask.png, the score made sense to me. Thank you for your kindly help anyway!",
    "2282717": "Glad that issue is resolved",
    "2283283": "Its because of the cross validation strategy that you are using",
    "2294010": "Dear All,\n\nI appreciate any feedback on this topic of \"visually good test results but low LB score\". \n\nI have been struggling for a long time on this problem but made no progress. \n\nI tried a couple of cross-validated models and my pipeline does not add padding.  \n\nI did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook (https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask), my submission still received very low public score (0.06).\n\nOn the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial). \n\nComparing the attached image with the result generated by the 0-11 simplest submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?). \n\nAny suggestions please?\n\nThanks very much in advance for your help and insights!",
    "2294078": "Does you model use only specific layers for making prediction? Test data can have relevant information on different layers then training data.",
    "2294106": "My model samples 11 to 15 of the 65 slices.\n\nIs it correct that the dummy \"test a\" and \"test b\" are substantially different than the real test data? \n\nI also wonder if the LB score is computed by using the mask decoded from the RLE submission, or calculates the LB score directly using the RLE string.\n\nThanks so much for getting back, Serhii.",
    "2294113": "I'm not sure if mask is getting subtracted. This actually would be a good experiment"
  },
  "source": "meta"
}