{
  "id": 411105,
  "title": "Very low submission score, but good results otherwise. Thoughts?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/411105",
  "author_name": "",
  "post_date": "2023-05-17T18:03:31.107116100Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hey everyone, I seem to be having a problem with my Vesuvius notebook submission. This is my first competition so I suspect I'm making some common error(s).</p>\n<p>My model is producing output on the test fragments that looks pretty decent:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F7c06cb56c8528d9341adb80a44dbca78%2Ftest_frag_a.png?generation=1684345868028709&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F72f781ef8b8f58015a90c98893fcabe6%2Ftest_frag_b.png?generation=1684345892487577&amp;alt=media\" alt=\"\"></p>\n<p>However, the highest f-beta score I've been able to achieve on a submitted notebook is 0.12.</p>\n<p>To solve this problem, I've tried using both versions of the RLE code provided in the examples:</p>\n<p>`def rle(predictions_map):<br>\n    threshold= .5<br>\n    flat_img = predictions_map.flatten()<br>\n    flat_img = np.where(flat_img &gt; threshold, 1, 0).astype(np.uint8)</p>\n<pre><code>starts = np.array((flat_img[:-1] == 0) &amp; (flat_img[1:] == 1))\nends = np.array((flat_img[:-1] == 1) &amp; (flat_img[1:] == 0))\nstarts_ix = np.where(starts)[0] + 2\nends_ix = np.where(ends)[0] + 2\nlengths = ends_ix - starts_ix\nreturn \" \".join(map(str, sum(zip(starts_ix, lengths), ())))`\n</code></pre>\n<p>and </p>\n<p><code>def rle_2(input):\n    pixels = input.flatten()\n    pixels[0] = 0\n    pixels[-1] = 0\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n    runs[1::2] = runs[1::2] - runs[:-1:2]\n    return ' '.join(str(x) for x in runs)</code></p>\n<p>Both of these code blocks create identical results. I then create a simple submission csv:</p>\n<p><code>submission = pd.DataFrame(data=submission_dict)\nsubmission.to_csv('submission.csv', index=False)</code></p>\n<p>The results look normal to me:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F3333c304bde4841a992357b7f8e32f16%2Fpd_printout.png?generation=1684346192452552&amp;alt=media\" alt=\"\"></p>\n<p>When I run my  model on the first training set, I get results that look similar to the test set:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fec4322da75b4431952df536413b66c5c%2Ftrain_frag_1.png?generation=1684346431929857&amp;alt=media\" alt=\"\"></p>\n<p>And when I calculate the f-beta score on the training set I'm getting a 0.72:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fa436c3282a6103834331b94462e1b67d%2Ff-beta_calc.png?generation=1684346471456256&amp;alt=media\" alt=\"\"></p>\n<p>Does anyone have any thoughts on what I might be doing wrong?</p>",
  "messages": [
    {
      "id": "2263594",
      "postDate": "05/17/2023 18:03:31",
      "content": "<p>Hey everyone, I seem to be having a problem with my Vesuvius notebook submission. This is my first competition so I suspect I'm making some common error(s).</p>\n<p>My model is producing output on the test fragments that looks pretty decent:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F7c06cb56c8528d9341adb80a44dbca78%2Ftest_frag_a.png?generation=1684345868028709&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F72f781ef8b8f58015a90c98893fcabe6%2Ftest_frag_b.png?generation=1684345892487577&amp;alt=media\" alt=\"\"></p>\n<p>However, the highest f-beta score I've been able to achieve on a submitted notebook is 0.12.</p>\n<p>To solve this problem, I've tried using both versions of the RLE code provided in the examples:</p>\n<p>`def rle(predictions_map):<br>\n    threshold= .5<br>\n    flat_img = predictions_map.flatten()<br>\n    flat_img = np.where(flat_img &gt; threshold, 1, 0).astype(np.uint8)</p>\n<pre><code>starts = np.array((flat_img[:-1] == 0) &amp; (flat_img[1:] == 1))\nends = np.array((flat_img[:-1] == 1) &amp; (flat_img[1:] == 0))\nstarts_ix = np.where(starts)[0] + 2\nends_ix = np.where(ends)[0] + 2\nlengths = ends_ix - starts_ix\nreturn \" \".join(map(str, sum(zip(starts_ix, lengths), ())))`\n</code></pre>\n<p>and </p>\n<p><code>def rle_2(input):\n    pixels = input.flatten()\n    pixels[0] = 0\n    pixels[-1] = 0\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n    runs[1::2] = runs[1::2] - runs[:-1:2]\n    return ' '.join(str(x) for x in runs)</code></p>\n<p>Both of these code blocks create identical results. I then create a simple submission csv:</p>\n<p><code>submission = pd.DataFrame(data=submission_dict)\nsubmission.to_csv('submission.csv', index=False)</code></p>\n<p>The results look normal to me:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F3333c304bde4841a992357b7f8e32f16%2Fpd_printout.png?generation=1684346192452552&amp;alt=media\" alt=\"\"></p>\n<p>When I run my  model on the first training set, I get results that look similar to the test set:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fec4322da75b4431952df536413b66c5c%2Ftrain_frag_1.png?generation=1684346431929857&amp;alt=media\" alt=\"\"></p>\n<p>And when I calculate the f-beta score on the training set I'm getting a 0.72:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fa436c3282a6103834331b94462e1b67d%2Ff-beta_calc.png?generation=1684346471456256&amp;alt=media\" alt=\"\"></p>\n<p>Does anyone have any thoughts on what I might be doing wrong?</p>",
      "rawMarkdown": "Hey everyone, I seem to be having a problem with my Vesuvius notebook submission. This is my first competition so I suspect I'm making some common error(s).\n\nMy model is producing output on the test fragments that looks pretty decent:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F7c06cb56c8528d9341adb80a44dbca78%2Ftest_frag_a.png?generation=1684345868028709&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F72f781ef8b8f58015a90c98893fcabe6%2Ftest_frag_b.png?generation=1684345892487577&alt=media)\n\nHowever, the highest f-beta score I've been able to achieve on a submitted notebook is 0.12.\n\nTo solve this problem, I've tried using both versions of the RLE code provided in the examples:\n\n`def rle(predictions_map):\n    threshold= .5\n    flat_img = predictions_map.flatten()\n    flat_img = np.where(flat_img > threshold, 1, 0).astype(np.uint8)\n\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))`\n\nand \n\n`def rle_2(input):\n    pixels = input.flatten()\n    pixels[0] = 0\n    pixels[-1] = 0\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n    runs[1::2] = runs[1::2] - runs[:-1:2]\n    return ' '.join(str(x) for x in runs)`\n\nBoth of these code blocks create identical results. I then create a simple submission csv:\n\n`submission = pd.DataFrame(data=submission_dict)\nsubmission.to_csv('submission.csv', index=False)`\n\nThe results look normal to me:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F3333c304bde4841a992357b7f8e32f16%2Fpd_printout.png?generation=1684346192452552&alt=media)\n\nWhen I run my  model on the first training set, I get results that look similar to the test set:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fec4322da75b4431952df536413b66c5c%2Ftrain_frag_1.png?generation=1684346431929857&alt=media)\n\nAnd when I calculate the f-beta score on the training set I'm getting a 0.72:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fa436c3282a6103834331b94462e1b67d%2Ff-beta_calc.png?generation=1684346471456256&alt=media)\n\nDoes anyone have any thoughts on what I might be doing wrong?",
      "votes": null
    },
    {
      "id": "2263730",
      "postDate": "05/17/2023 20:50:53",
      "content": "<p>Note that data in the test folder is a dummy data and gets substituted with a real data when your notebook is submitted to competition. So good score in your local run does not mean good score on the leaderboard.<br>\nTry training on one of the training samples and then evaluate it on other training sample. This way you make sure your model is not overfitting. I usually train on sample 3 and validate on sample 1.</p>",
      "rawMarkdown": "Note that data in the test folder is a dummy data and gets substituted with a real data when your notebook is submitted to competition. So good score in your local run does not mean good score on the leaderboard.\nTry training on one of the training samples and then evaluate it on other training sample. This way you make sure your model is not overfitting. I usually train on sample 3 and validate on sample 1.",
      "votes": null
    },
    {
      "id": "2263736",
      "postDate": "05/17/2023 21:00:07",
      "content": "<p>Thanks, good idea. Doing that now, will let you know how the results turn out. </p>",
      "rawMarkdown": "Thanks, good idea. Doing that now, will let you know how the results turn out.",
      "votes": null
    },
    {
      "id": "2264624",
      "postDate": "05/18/2023 15:30:47",
      "content": "<p>Alright results are in. I trained on sets 1 and 2, and tested on 3. The results are less flattering but still reasonable. Here's how the image looks:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F493acb86a885f4ed90355e21a4aab580%2Ftrain-3.png?generation=1684423498933596&amp;alt=media\" alt=\"\"></p>\n<p>To my eye it doesn't look fantastic, but it's earning an f-beta score of 0.44:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F252d72647127a1c58d01664d870bedfc%2Ffbeta-3.png?generation=1684423568180034&amp;alt=media\" alt=\"\"></p>\n<p>That's still 4x what I'm earning on my submissions. Feels like perhaps there's an issue with the submission, no?</p>",
      "rawMarkdown": "Alright results are in. I trained on sets 1 and 2, and tested on 3. The results are less flattering but still reasonable. Here's how the image looks:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F493acb86a885f4ed90355e21a4aab580%2Ftrain-3.png?generation=1684423498933596&alt=media)\n\nTo my eye it doesn't look fantastic, but it's earning an f-beta score of 0.44:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F252d72647127a1c58d01664d870bedfc%2Ffbeta-3.png?generation=1684423568180034&alt=media)\n\nThat's still 4x what I'm earning on my submissions. Feels like perhaps there's an issue with the submission, no?",
      "votes": null
    },
    {
      "id": "2264652",
      "postDate": "05/18/2023 15:51:53",
      "content": "<p>I am facing the same issue. I train on 1 + 2 and use 3 for validation. The score on the validation looks good and the predicted mask seems also plausible. Instead of the original data I trained on the compressed png version (<a href=\"https://www.kaggle.com/code/yoyobar/compressed-image-lossless-6gb)\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/compressed-image-lossless-6gb)</a>. I suspected some normalization issues at first, but validating the code inside kaggle using the original .tif files lead to the same results…</p>",
      "rawMarkdown": "I am facing the same issue. I train on 1 + 2 and use 3 for validation. The score on the validation looks good and the predicted mask seems also plausible. Instead of the original data I trained on the compressed png version (https://www.kaggle.com/code/yoyobar/compressed-image-lossless-6gb). I suspected some normalization issues at first, but validating the code inside kaggle using the original .tif files lead to the same results...",
      "votes": null
    },
    {
      "id": "2264681",
      "postDate": "05/18/2023 16:30:51",
      "content": "<p>I faced the same problem (results are: 0, 0.11, 0.12) when there were empty tiles of mask (tile.max() == 0) in the training dataset.</p>",
      "rawMarkdown": "I faced the same problem (results are: 0, 0.11, 0.12) when there were empty tiles of mask (tile.max() == 0) in the training dataset.",
      "votes": null
    },
    {
      "id": "2264736",
      "postDate": "05/18/2023 17:10:18",
      "content": "<p>Hmm that's interesting. I'm assuming \"tile\" is the grid of output pixels you're looking at. When you produce the RLE, are you counting pixels that aren't part of the mask?</p>",
      "rawMarkdown": "Hmm that's interesting. I'm assuming \"tile\" is the grid of output pixels you're looking at. When you produce the RLE, are you counting pixels that aren't part of the mask?",
      "votes": null
    },
    {
      "id": "2264778",
      "postDate": "05/18/2023 17:47:10",
      "content": "<p>Glad to know I'm not the only one facing this issue, please let me know if you find a solution.</p>",
      "rawMarkdown": "Glad to know I'm not the only one facing this issue, please let me know if you find a solution.",
      "votes": null
    },
    {
      "id": "2264872",
      "postDate": "05/18/2023 19:51:52",
      "content": "<p>yeah, 2 common issues with submissions are:</p>\n<ul>\n<li>x and y coordinates messed up</li>\n<li>rle produces incorrect result</li>\n</ul>\n<p>Try following RLE function.</p>\n<p><code>def rle(img):\n    pixels = img.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    f = StringIO()\n    np.savetxt(f, runs.reshape(1, -1), delimiter=\" \", fmt=\"%d\")\n    predicted = f.getvalue().strip()\n    return predicted</code></p>",
      "rawMarkdown": "yeah, 2 common issues with submissions are:\n* x and y coordinates messed up\n* rle produces incorrect result\n\nTry following RLE function.\n\n```def rle(img):\n    pixels = img.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    f = StringIO()\n    np.savetxt(f, runs.reshape(1, -1), delimiter=\" \", fmt=\"%d\")\n    predicted = f.getvalue().strip()\n    return predicted```",
      "votes": null
    },
    {
      "id": "2266027",
      "postDate": "05/19/2023 16:49:25",
      "content": "<p>To get that to work I had to include: from io import StringIO. Please let me know if that was the correct package to import. </p>\n<p>If so, the score decreased from 0.12 to 0.07. </p>",
      "rawMarkdown": "To get that to work I had to include: from io import StringIO. Please let me know if that was the correct package to import. \n\nIf so, the score decreased from 0.12 to 0.07.",
      "votes": null
    },
    {
      "id": "2266118",
      "postDate": "05/19/2023 18:10:04",
      "content": "<p>Wow, then it is not an rle. Might be something with coordinates order or resolution. Out of curiosity can you try placing prediction for fragment a before fragment b? I see that your data frame has b fragment first, which might matter somehow for score evaluation</p>",
      "rawMarkdown": "Wow, then it is not an rle. Might be something with coordinates order or resolution. Out of curiosity can you try placing prediction for fragment a before fragment b? I see that your data frame has b fragment first, which might matter somehow for score evaluation",
      "votes": null
    },
    {
      "id": "2267035",
      "postDate": "05/20/2023 14:38:10",
      "content": "<p>Try transposing your image before doing RLE encoding. </p>",
      "rawMarkdown": "Try transposing your image before doing RLE encoding.",
      "votes": null
    },
    {
      "id": "2269566",
      "postDate": "05/22/2023 14:46:09",
      "content": "<p>Haha, yea I tried changing the order a few weeks ago. Originally it was a first, then b. I switched it because one of the example submissions seemed to have it backwards. No impact on the score from the change. </p>",
      "rawMarkdown": "Haha, yea I tried changing the order a few weeks ago. Originally it was a first, then b. I switched it because one of the example submissions seemed to have it backwards. No impact on the score from the change.",
      "votes": null
    },
    {
      "id": "2271057",
      "postDate": "05/23/2023 15:36:35",
      "content": "<p>Hi Dennis, thanks for the idea. I gave that a try and the notebook ran but produced a scoring error. The logs indicate the rle code ran and produced something that looks normal:</p>\n<p>1812.8s    4     ID                                          Predicted<br>\n1812.8s    5   0  b  168523 1 168525 1 168527 1 168529 1 168531 1 1…<br>\n1812.8s    6   1  a  2906806 1 2906808 1 2909533 1 2909535 1 291226…</p>\n<p>However, the output file was 153MB, compared to 498kB on files that were scored. What have the file sizes been on your successful submissions?</p>",
      "rawMarkdown": "Hi Dennis, thanks for the idea. I gave that a try and the notebook ran but produced a scoring error. The logs indicate the rle code ran and produced something that looks normal:\n\n1812.8s\t4\t  ID                                          Predicted\n1812.8s\t5\t0  b  168523 1 168525 1 168527 1 168529 1 168531 1 1...\n1812.8s\t6\t1  a  2906806 1 2906808 1 2909533 1 2909535 1 291226...\n\nHowever, the output file was 153MB, compared to 498kB on files that were scored. What have the file sizes been on your successful submissions?",
      "votes": null
    },
    {
      "id": "2282226",
      "postDate": "05/31/2023 13:15:29",
      "content": "<p>Quick update on this issue, which I hope will be of use to other competitors. My highest submission score is now 0.27, which is a significant improvement over my prior high score of 0.12. After much testing, I have concluded that the issue was not with my RLE algorithm. I tried a few different algos, which all produced similar results. Also importantly, I read in another post that if you simply submit the mask, your score will be 0.11. I submitted the mask using my RLE ago and received a score of 0.11, confirming both that my RLE code was working as expected, and my model was producing more or less garbage results. </p>\n<p>I also submitted a version of the mask where only a band of 20% of the pixels were labeled as ink, and this produced a score of 0. This result suggested that my model was invalidating too many pixels. To raise my score above .12, the key change I made was dramatically lowering my threshold. My best score was produced at a threshold of 0.30. </p>\n<p>I also experimented with the following code to try to dynamically set the threshold of each image so that x% of pixels would be labeled as ink:</p>\n<pre><code>    positives = \n    THRESHOLD = \n     positives &lt; :\n        THRESHOLD -= \n        threshold_test = np.where(model_output &gt; THRESHOLD, , )\n        positives = np.mean(threshold_test) \n</code></pre>\n<p>From the log files, I can see that my best results were achieved when I set 'positives' to 0.12 (meaning the model is selecting 12% of the pixels to be labeled), but both images settled on the same exact threshold of 0.28.  </p>\n<p>I addition to tinkering with the threshold, I have also experimented with augmenting/transforming the data, and regularization of the model. I'm hoping this will close the gap between my test and submission results, but so far no luck.</p>",
      "rawMarkdown": "Quick update on this issue, which I hope will be of use to other competitors. My highest submission score is now 0.27, which is a significant improvement over my prior high score of 0.12. After much testing, I have concluded that the issue was not with my RLE algorithm. I tried a few different algos, which all produced similar results. Also importantly, I read in another post that if you simply submit the mask, your score will be 0.11. I submitted the mask using my RLE ago and received a score of 0.11, confirming both that my RLE code was working as expected, and my model was producing more or less garbage results. \n\nI also submitted a version of the mask where only a band of 20% of the pixels were labeled as ink, and this produced a score of 0. This result suggested that my model was invalidating too many pixels. To raise my score above .12, the key change I made was dramatically lowering my threshold. My best score was produced at a threshold of 0.30. \n\nI also experimented with the following code to try to dynamically set the threshold of each image so that x% of pixels would be labeled as ink:\n\n```python\n    positives = 0\n    THRESHOLD = 1.00\n    while positives < 0.12:\n        THRESHOLD -= 0.01\n        threshold_test = np.where(model_output > THRESHOLD, 1, 0)\n        positives = np.mean(threshold_test) \n```\nFrom the log files, I can see that my best results were achieved when I set 'positives' to 0.12 (meaning the model is selecting 12% of the pixels to be labeled), but both images settled on the same exact threshold of 0.28.  \n\nI addition to tinkering with the threshold, I have also experimented with augmenting/transforming the data, and regularization of the model. I'm hoping this will close the gap between my test and submission results, but so far no luck.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2263730,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "05/17/2023 20:50:53",
      "content": "<p>Note that data in the test folder is a dummy data and gets substituted with a real data when your notebook is submitted to competition. So good score in your local run does not mean good score on the leaderboard.<br>\nTry training on one of the training samples and then evaluate it on other training sample. This way you make sure your model is not overfitting. I usually train on sample 3 and validate on sample 1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2263736,
          "author_name": "jeffborack",
          "author_url": "",
          "post_date": "05/17/2023 21:00:07",
          "content": "<p>Thanks, good idea. Doing that now, will let you know how the results turn out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2264624,
          "author_name": "jeffborack",
          "author_url": "",
          "post_date": "05/18/2023 15:30:47",
          "content": "<p>Alright results are in. I trained on sets 1 and 2, and tested on 3. The results are less flattering but still reasonable. Here's how the image looks:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F493acb86a885f4ed90355e21a4aab580%2Ftrain-3.png?generation=1684423498933596&amp;alt=media\" alt=\"\"></p>\n<p>To my eye it doesn't look fantastic, but it's earning an f-beta score of 0.44:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F252d72647127a1c58d01664d870bedfc%2Ffbeta-3.png?generation=1684423568180034&amp;alt=media\" alt=\"\"></p>\n<p>That's still 4x what I'm earning on my submissions. Feels like perhaps there's an issue with the submission, no?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2264681,
              "author_name": "mknzfr",
              "author_url": "",
              "post_date": "05/18/2023 16:30:51",
              "content": "<p>I faced the same problem (results are: 0, 0.11, 0.12) when there were empty tiles of mask (tile.max() == 0) in the training dataset.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2264736,
                  "author_name": "jeffborack",
                  "author_url": "",
                  "post_date": "05/18/2023 17:10:18",
                  "content": "<p>Hmm that's interesting. I'm assuming \"tile\" is the grid of output pixels you're looking at. When you produce the RLE, are you counting pixels that aren't part of the mask?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2264652,
      "author_name": "felixmneumann",
      "author_url": "",
      "post_date": "05/18/2023 15:51:53",
      "content": "<p>I am facing the same issue. I train on 1 + 2 and use 3 for validation. The score on the validation looks good and the predicted mask seems also plausible. Instead of the original data I trained on the compressed png version (<a href=\"https://www.kaggle.com/code/yoyobar/compressed-image-lossless-6gb)\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/compressed-image-lossless-6gb)</a>. I suspected some normalization issues at first, but validating the code inside kaggle using the original .tif files lead to the same results…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2264778,
          "author_name": "jeffborack",
          "author_url": "",
          "post_date": "05/18/2023 17:47:10",
          "content": "<p>Glad to know I'm not the only one facing this issue, please let me know if you find a solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2264872,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "05/18/2023 19:51:52",
      "content": "<p>yeah, 2 common issues with submissions are:</p>\n<ul>\n<li>x and y coordinates messed up</li>\n<li>rle produces incorrect result</li>\n</ul>\n<p>Try following RLE function.</p>\n<p><code>def rle(img):\n    pixels = img.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    f = StringIO()\n    np.savetxt(f, runs.reshape(1, -1), delimiter=\" \", fmt=\"%d\")\n    predicted = f.getvalue().strip()\n    return predicted</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 2266027,
          "author_name": "jeffborack",
          "author_url": "",
          "post_date": "05/19/2023 16:49:25",
          "content": "<p>To get that to work I had to include: from io import StringIO. Please let me know if that was the correct package to import. </p>\n<p>If so, the score decreased from 0.12 to 0.07. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2266118,
              "author_name": "elvenmonk",
              "author_url": "",
              "post_date": "05/19/2023 18:10:04",
              "content": "<p>Wow, then it is not an rle. Might be something with coordinates order or resolution. Out of curiosity can you try placing prediction for fragment a before fragment b? I see that your data frame has b fragment first, which might matter somehow for score evaluation</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2269566,
                  "author_name": "jeffborack",
                  "author_url": "",
                  "post_date": "05/22/2023 14:46:09",
                  "content": "<p>Haha, yea I tried changing the order a few weeks ago. Originally it was a first, then b. I switched it because one of the example submissions seemed to have it backwards. No impact on the score from the change. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2267035,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "05/20/2023 14:38:10",
      "content": "<p>Try transposing your image before doing RLE encoding. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2271057,
          "author_name": "jeffborack",
          "author_url": "",
          "post_date": "05/23/2023 15:36:35",
          "content": "<p>Hi Dennis, thanks for the idea. I gave that a try and the notebook ran but produced a scoring error. The logs indicate the rle code ran and produced something that looks normal:</p>\n<p>1812.8s    4     ID                                          Predicted<br>\n1812.8s    5   0  b  168523 1 168525 1 168527 1 168529 1 168531 1 1…<br>\n1812.8s    6   1  a  2906806 1 2906808 1 2909533 1 2909535 1 291226…</p>\n<p>However, the output file was 153MB, compared to 498kB on files that were scored. What have the file sizes been on your successful submissions?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2282226,
      "author_name": "jeffborack",
      "author_url": "",
      "post_date": "05/31/2023 13:15:29",
      "content": "<p>Quick update on this issue, which I hope will be of use to other competitors. My highest submission score is now 0.27, which is a significant improvement over my prior high score of 0.12. After much testing, I have concluded that the issue was not with my RLE algorithm. I tried a few different algos, which all produced similar results. Also importantly, I read in another post that if you simply submit the mask, your score will be 0.11. I submitted the mask using my RLE ago and received a score of 0.11, confirming both that my RLE code was working as expected, and my model was producing more or less garbage results. </p>\n<p>I also submitted a version of the mask where only a band of 20% of the pixels were labeled as ink, and this produced a score of 0. This result suggested that my model was invalidating too many pixels. To raise my score above .12, the key change I made was dramatically lowering my threshold. My best score was produced at a threshold of 0.30. </p>\n<p>I also experimented with the following code to try to dynamically set the threshold of each image so that x% of pixels would be labeled as ink:</p>\n<pre><code>    positives = \n    THRESHOLD = \n     positives &lt; :\n        THRESHOLD -= \n        threshold_test = np.where(model_output &gt; THRESHOLD, , )\n        positives = np.mean(threshold_test) \n</code></pre>\n<p>From the log files, I can see that my best results were achieved when I set 'positives' to 0.12 (meaning the model is selecting 12% of the pixels to be labeled), but both images settled on the same exact threshold of 0.28.  </p>\n<p>I addition to tinkering with the threshold, I have also experimented with augmenting/transforming the data, and regularization of the model. I'm hoping this will close the gap between my test and submission results, but so far no luck.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2263594": "Hey everyone, I seem to be having a problem with my Vesuvius notebook submission. This is my first competition so I suspect I'm making some common error(s).\n\nMy model is producing output on the test fragments that looks pretty decent:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F7c06cb56c8528d9341adb80a44dbca78%2Ftest_frag_a.png?generation=1684345868028709&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F72f781ef8b8f58015a90c98893fcabe6%2Ftest_frag_b.png?generation=1684345892487577&alt=media)\n\nHowever, the highest f-beta score I've been able to achieve on a submitted notebook is 0.12.\n\nTo solve this problem, I've tried using both versions of the RLE code provided in the examples:\n\n`def rle(predictions_map):\n    threshold= .5\n    flat_img = predictions_map.flatten()\n    flat_img = np.where(flat_img > threshold, 1, 0).astype(np.uint8)\n\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))`\n\nand \n\n`def rle_2(input):\n    pixels = input.flatten()\n    pixels[0] = 0\n    pixels[-1] = 0\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n    runs[1::2] = runs[1::2] - runs[:-1:2]\n    return ' '.join(str(x) for x in runs)`\n\nBoth of these code blocks create identical results. I then create a simple submission csv:\n\n`submission = pd.DataFrame(data=submission_dict)\nsubmission.to_csv('submission.csv', index=False)`\n\nThe results look normal to me:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F3333c304bde4841a992357b7f8e32f16%2Fpd_printout.png?generation=1684346192452552&alt=media)\n\nWhen I run my  model on the first training set, I get results that look similar to the test set:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fec4322da75b4431952df536413b66c5c%2Ftrain_frag_1.png?generation=1684346431929857&alt=media)\n\nAnd when I calculate the f-beta score on the training set I'm getting a 0.72:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2Fa436c3282a6103834331b94462e1b67d%2Ff-beta_calc.png?generation=1684346471456256&alt=media)\n\nDoes anyone have any thoughts on what I might be doing wrong?",
    "2263730": "Note that data in the test folder is a dummy data and gets substituted with a real data when your notebook is submitted to competition. So good score in your local run does not mean good score on the leaderboard.\nTry training on one of the training samples and then evaluate it on other training sample. This way you make sure your model is not overfitting. I usually train on sample 3 and validate on sample 1.",
    "2263736": "Thanks, good idea. Doing that now, will let you know how the results turn out.",
    "2264624": "Alright results are in. I trained on sets 1 and 2, and tested on 3. The results are less flattering but still reasonable. Here's how the image looks:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F493acb86a885f4ed90355e21a4aab580%2Ftrain-3.png?generation=1684423498933596&alt=media)\n\nTo my eye it doesn't look fantastic, but it's earning an f-beta score of 0.44:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F134770%2F252d72647127a1c58d01664d870bedfc%2Ffbeta-3.png?generation=1684423568180034&alt=media)\n\nThat's still 4x what I'm earning on my submissions. Feels like perhaps there's an issue with the submission, no?",
    "2264652": "I am facing the same issue. I train on 1 + 2 and use 3 for validation. The score on the validation looks good and the predicted mask seems also plausible. Instead of the original data I trained on the compressed png version (https://www.kaggle.com/code/yoyobar/compressed-image-lossless-6gb). I suspected some normalization issues at first, but validating the code inside kaggle using the original .tif files lead to the same results...",
    "2264681": "I faced the same problem (results are: 0, 0.11, 0.12) when there were empty tiles of mask (tile.max() == 0) in the training dataset.",
    "2264736": "Hmm that's interesting. I'm assuming \"tile\" is the grid of output pixels you're looking at. When you produce the RLE, are you counting pixels that aren't part of the mask?",
    "2264778": "Glad to know I'm not the only one facing this issue, please let me know if you find a solution.",
    "2264872": "yeah, 2 common issues with submissions are:\n* x and y coordinates messed up\n* rle produces incorrect result\n\nTry following RLE function.\n\n```def rle(img):\n    pixels = img.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    f = StringIO()\n    np.savetxt(f, runs.reshape(1, -1), delimiter=\" \", fmt=\"%d\")\n    predicted = f.getvalue().strip()\n    return predicted```",
    "2266027": "To get that to work I had to include: from io import StringIO. Please let me know if that was the correct package to import. \n\nIf so, the score decreased from 0.12 to 0.07.",
    "2266118": "Wow, then it is not an rle. Might be something with coordinates order or resolution. Out of curiosity can you try placing prediction for fragment a before fragment b? I see that your data frame has b fragment first, which might matter somehow for score evaluation",
    "2267035": "Try transposing your image before doing RLE encoding.",
    "2269566": "Haha, yea I tried changing the order a few weeks ago. Originally it was a first, then b. I switched it because one of the example submissions seemed to have it backwards. No impact on the score from the change.",
    "2271057": "Hi Dennis, thanks for the idea. I gave that a try and the notebook ran but produced a scoring error. The logs indicate the rle code ran and produced something that looks normal:\n\n1812.8s\t4\t  ID                                          Predicted\n1812.8s\t5\t0  b  168523 1 168525 1 168527 1 168529 1 168531 1 1...\n1812.8s\t6\t1  a  2906806 1 2906808 1 2909533 1 2909535 1 291226...\n\nHowever, the output file was 153MB, compared to 498kB on files that were scored. What have the file sizes been on your successful submissions?",
    "2282226": "Quick update on this issue, which I hope will be of use to other competitors. My highest submission score is now 0.27, which is a significant improvement over my prior high score of 0.12. After much testing, I have concluded that the issue was not with my RLE algorithm. I tried a few different algos, which all produced similar results. Also importantly, I read in another post that if you simply submit the mask, your score will be 0.11. I submitted the mask using my RLE ago and received a score of 0.11, confirming both that my RLE code was working as expected, and my model was producing more or less garbage results. \n\nI also submitted a version of the mask where only a band of 20% of the pixels were labeled as ink, and this produced a score of 0. This result suggested that my model was invalidating too many pixels. To raise my score above .12, the key change I made was dramatically lowering my threshold. My best score was produced at a threshold of 0.30. \n\nI also experimented with the following code to try to dynamically set the threshold of each image so that x% of pixels would be labeled as ink:\n\n```python\n    positives = 0\n    THRESHOLD = 1.00\n    while positives < 0.12:\n        THRESHOLD -= 0.01\n        threshold_test = np.where(model_output > THRESHOLD, 1, 0)\n        positives = np.mean(threshold_test) \n```\nFrom the log files, I can see that my best results were achieved when I set 'positives' to 0.12 (meaning the model is selecting 12% of the pixels to be labeled), but both images settled on the same exact threshold of 0.28.  \n\nI addition to tinkering with the threshold, I have also experimented with augmenting/transforming the data, and regularization of the model. I'm hoping this will close the gap between my test and submission results, but so far no luck."
  },
  "source": "meta"
}