{
  "id": 412404,
  "title": "RLE Help! and asking about the model overfitting",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/412404",
  "author_name": "",
  "post_date": "2023-05-23T16:06:19.915355700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>this is my RLE, and I can not see why this is happening, so please help!</p>\n<p>def rle(output):<br>\n    flat_img = output.flatten()<br>\n    flat_img = np.where(output &gt; 0.5, 1, 0).astype(np.uint8)<br>\n    starts = np.array((flat_img[:-1] == 0) &amp; (flat_img[1:] == 1))<br>\n    ends = np.array((flat_img[:-1] == 1) &amp; (flat_img[1:] == 0))<br>\n    starts_ix = np.where(starts)[0] + 2<br>\n    ends_ix = np.where(ends)[0] + 2<br>\n    lengths = ends_ix - starts_ix<br>\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))</p>\n<p>another qu, is this overfitting?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2Fb8ebc6d7a88071ec20bf1cfd0e08bc48%2F1.PNG?generation=1684857839791262&amp;alt=media\" alt=\" the error\"> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2F40d192528cf0e70088365f97bcf9fe83%2F2.PNG?generation=1684857889493002&amp;alt=media\" alt=\"the output\"></p>",
  "messages": [
    {
      "id": "2271091",
      "postDate": "05/23/2023 16:06:19",
      "content": "<p>this is my RLE, and I can not see why this is happening, so please help!</p>\n<p>def rle(output):<br>\n    flat_img = output.flatten()<br>\n    flat_img = np.where(output &gt; 0.5, 1, 0).astype(np.uint8)<br>\n    starts = np.array((flat_img[:-1] == 0) &amp; (flat_img[1:] == 1))<br>\n    ends = np.array((flat_img[:-1] == 1) &amp; (flat_img[1:] == 0))<br>\n    starts_ix = np.where(starts)[0] + 2<br>\n    ends_ix = np.where(ends)[0] + 2<br>\n    lengths = ends_ix - starts_ix<br>\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))</p>\n<p>another qu, is this overfitting?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2Fb8ebc6d7a88071ec20bf1cfd0e08bc48%2F1.PNG?generation=1684857839791262&amp;alt=media\" alt=\" the error\"> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2F40d192528cf0e70088365f97bcf9fe83%2F2.PNG?generation=1684857889493002&amp;alt=media\" alt=\"the output\"></p>",
      "rawMarkdown": "this is my RLE, and I can not see why this is happening, so please help!\n\ndef rle(output):\n    flat_img = output.flatten()\n    flat_img = np.where(output > 0.5, 1, 0).astype(np.uint8)\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))\n\n\nanother qu, is this overfitting?![ the error](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2Fb8ebc6d7a88071ec20bf1cfd0e08bc48%2F1.PNG?generation=1684857839791262&alt=media) \n![the output](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2F40d192528cf0e70088365f97bcf9fe83%2F2.PNG?generation=1684857889493002&alt=media)",
      "votes": null
    },
    {
      "id": "2272604",
      "postDate": "05/24/2023 15:51:24",
      "content": "<p>Are you training and validating/evaluating on the same data? Your output looks almost identical to the mask.</p>",
      "rawMarkdown": "Are you training and validating/evaluating on the same data? Your output looks almost identical to the mask.",
      "votes": null
    },
    {
      "id": "2272645",
      "postDate": "05/24/2023 16:11:24",
      "content": "<p>Sharing my working RLE code here. To test if it works for you, try submitting the mask as your prediction for each fragment and you should get a score of 0.11:</p>\n<pre><code>\nTHRESHOLD = \n\n ():\n    pixels = np.where(output.flatten() &gt; THRESHOLD, , ).astype(np.uint8)\n    pixels[] = \n    pixels[-] = \n    runs = np.where(pixels[:] != pixels[:-])[] + \n    runs[::] = runs[::] - runs[:-:]\n     .join((x)  x  runs)\n\n\nrle_output1 = rle(FIRST_PREDICTION)\nrle_output2 = rle(SECOND_PREDICTION)\n\n( + rle_output1 +  + rle_output2, file=(, ))\n</code></pre>",
      "rawMarkdown": "Sharing my working RLE code here. To test if it works for you, try submitting the mask as your prediction for each fragment and you should get a score of 0.11:\n```python\n# Set this to your prediction binarization threshold. I use 0.5 for my models.\nTHRESHOLD = 0.5\n\ndef rle(output):\n    pixels = np.where(output.flatten() > THRESHOLD, 1, 0).astype(np.uint8)\n    pixels[0] = 0\n    pixels[-1] = 0\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n    runs[1::2] = runs[1::2] - runs[:-1:2]\n    return ' '.join(str(x) for x in runs)\n\n# Substitute below with your predictions as numpy arrays.\nrle_output1 = rle(FIRST_PREDICTION)\nrle_output2 = rle(SECOND_PREDICTION)\n\nprint(\"Id,Predicted\\na,\" + rle_output1 + \"\\nb,\" + rle_output2, file=open('submission.csv', 'w'))\n```",
      "votes": null
    },
    {
      "id": "2272657",
      "postDate": "05/24/2023 16:19:56",
      "content": "<p>Could you please elaborate about this\"try submitting the mask as your prediction for each fragment and you should get a score of 0.11:\"</p>",
      "rawMarkdown": "Could you please elaborate about this\"try submitting the mask as your prediction for each fragment and you should get a score of 0.11:\"",
      "votes": null
    },
    {
      "id": "2272662",
      "postDate": "05/24/2023 16:22:21",
      "content": "<p>It's an overfitting problem, actually the model architecture is very good I think, but the problem is with the noise of the data and the lack of augmentation </p>",
      "rawMarkdown": "It's an overfitting problem, actually the model architecture is very good I think, but the problem is with the noise of the data and the lack of augmentation",
      "votes": null
    },
    {
      "id": "2274389",
      "postDate": "05/26/2023 00:10:09",
      "content": "<p>If you create a submission notebook that runs RLE on the masks for the test fragments (no model involved) you will earn a score of 0.11 on the leaderboard. That's one way to ensure your RLE code is working.</p>",
      "rawMarkdown": "If you create a submission notebook that runs RLE on the masks for the test fragments (no model involved) you will earn a score of 0.11 on the leaderboard. That's one way to ensure your RLE code is working.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2272604,
      "author_name": "stevenhewitt",
      "author_url": "",
      "post_date": "05/24/2023 15:51:24",
      "content": "<p>Are you training and validating/evaluating on the same data? Your output looks almost identical to the mask.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2272662,
          "author_name": "salaheldinelnabarawy",
          "author_url": "",
          "post_date": "05/24/2023 16:22:21",
          "content": "<p>It's an overfitting problem, actually the model architecture is very good I think, but the problem is with the noise of the data and the lack of augmentation </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2272645,
      "author_name": "stevenhewitt",
      "author_url": "",
      "post_date": "05/24/2023 16:11:24",
      "content": "<p>Sharing my working RLE code here. To test if it works for you, try submitting the mask as your prediction for each fragment and you should get a score of 0.11:</p>\n<pre><code>\nTHRESHOLD = \n\n ():\n    pixels = np.where(output.flatten() &gt; THRESHOLD, , ).astype(np.uint8)\n    pixels[] = \n    pixels[-] = \n    runs = np.where(pixels[:] != pixels[:-])[] + \n    runs[::] = runs[::] - runs[:-:]\n     .join((x)  x  runs)\n\n\nrle_output1 = rle(FIRST_PREDICTION)\nrle_output2 = rle(SECOND_PREDICTION)\n\n( + rle_output1 +  + rle_output2, file=(, ))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2272657,
          "author_name": "salaheldinelnabarawy",
          "author_url": "",
          "post_date": "05/24/2023 16:19:56",
          "content": "<p>Could you please elaborate about this\"try submitting the mask as your prediction for each fragment and you should get a score of 0.11:\"</p>",
          "votes": null,
          "replies": [
            {
              "id": 2274389,
              "author_name": "stevenhewitt",
              "author_url": "",
              "post_date": "05/26/2023 00:10:09",
              "content": "<p>If you create a submission notebook that runs RLE on the masks for the test fragments (no model involved) you will earn a score of 0.11 on the leaderboard. That's one way to ensure your RLE code is working.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2271091": "this is my RLE, and I can not see why this is happening, so please help!\n\ndef rle(output):\n    flat_img = output.flatten()\n    flat_img = np.where(output > 0.5, 1, 0).astype(np.uint8)\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))\n\n\nanother qu, is this overfitting?![ the error](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2Fb8ebc6d7a88071ec20bf1cfd0e08bc48%2F1.PNG?generation=1684857839791262&alt=media) \n![the output](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6004998%2F40d192528cf0e70088365f97bcf9fe83%2F2.PNG?generation=1684857889493002&alt=media)",
    "2272604": "Are you training and validating/evaluating on the same data? Your output looks almost identical to the mask.",
    "2272645": "Sharing my working RLE code here. To test if it works for you, try submitting the mask as your prediction for each fragment and you should get a score of 0.11:\n```python\n# Set this to your prediction binarization threshold. I use 0.5 for my models.\nTHRESHOLD = 0.5\n\ndef rle(output):\n    pixels = np.where(output.flatten() > THRESHOLD, 1, 0).astype(np.uint8)\n    pixels[0] = 0\n    pixels[-1] = 0\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n    runs[1::2] = runs[1::2] - runs[:-1:2]\n    return ' '.join(str(x) for x in runs)\n\n# Substitute below with your predictions as numpy arrays.\nrle_output1 = rle(FIRST_PREDICTION)\nrle_output2 = rle(SECOND_PREDICTION)\n\nprint(\"Id,Predicted\\na,\" + rle_output1 + \"\\nb,\" + rle_output2, file=open('submission.csv', 'w'))\n```",
    "2272657": "Could you please elaborate about this\"try submitting the mask as your prediction for each fragment and you should get a score of 0.11:\"",
    "2272662": "It's an overfitting problem, actually the model architecture is very good I think, but the problem is with the noise of the data and the lack of augmentation",
    "2274389": "If you create a submission notebook that runs RLE on the masks for the test fragments (no model involved) you will earn a score of 0.11 on the leaderboard. That's one way to ensure your RLE code is working."
  },
  "source": "meta"
}