{
  "id": 416120,
  "title": "Visually better than \"mask\" but much lower LB score?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/416120",
  "author_name": "",
  "post_date": "2023-06-09T16:58:20.607325100Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Dear All,</p>\n<p>I appreciate any feedback on this topic of <strong>\"visually good test results but low Leaderboard (LB) score\"</strong>.</p>\n<p>I have been struggling for a long time on this problem but made no progress.</p>\n<p>I tried a couple of cross-validated models and my pipeline does not add padding.</p>\n<p>I did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook <a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask</a> (thank you Lucas!), my submission still received very low public score (0.06).</p>\n<p>On the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial).</p>\n<p>Comparing the attached image with the \"mask\" used by the 0-11 simplest test submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?).</p>\n<p>Any suggestions please?</p>\n<p>Thanks very much in advance for your help and insights!</p>\n<p>===========================================</p>\n<p>References:</p>\n<p><a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask/comments\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask/comments</a></p>\n<p><a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/413664#2294078\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/413664#2294078</a></p>",
  "messages": [
    {
      "id": "2294019",
      "postDate": "06/09/2023 16:58:20",
      "content": "<p>Dear All,</p>\n<p>I appreciate any feedback on this topic of <strong>\"visually good test results but low Leaderboard (LB) score\"</strong>.</p>\n<p>I have been struggling for a long time on this problem but made no progress.</p>\n<p>I tried a couple of cross-validated models and my pipeline does not add padding.</p>\n<p>I did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook <a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask</a> (thank you Lucas!), my submission still received very low public score (0.06).</p>\n<p>On the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial).</p>\n<p>Comparing the attached image with the \"mask\" used by the 0-11 simplest test submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?).</p>\n<p>Any suggestions please?</p>\n<p>Thanks very much in advance for your help and insights!</p>\n<p>===========================================</p>\n<p>References:</p>\n<p><a href=\"https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask/comments\" target=\"_blank\">https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask/comments</a></p>\n<p><a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/413664#2294078\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/413664#2294078</a></p>",
      "rawMarkdown": "Dear All,\n\nI appreciate any feedback on this topic of **\"visually good test results but low Leaderboard (LB) score\"**.\n\nI have been struggling for a long time on this problem but made no progress.\n\nI tried a couple of cross-validated models and my pipeline does not add padding.\n\nI did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask (thank you Lucas!), my submission still received very low public score (0.06).\n\nOn the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial).\n\nComparing the attached image with the \"mask\" used by the 0-11 simplest test submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?).\n\nAny suggestions please?\n\nThanks very much in advance for your help and insights!\n\n===========================================\n\nReferences:\n\nhttps://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask/comments\n\nhttps://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\n\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/413664#2294078",
      "votes": null
    },
    {
      "id": "2294349",
      "postDate": "06/10/2023 02:33:56",
      "content": "<p>Di you use fragment 1 for training? In the competition data that is shared with us, real test data is replaced with copy of data from training fragment 1. If you model learned on fragment 1 then no wonder it can predict it. However it does not mean that it can predict anything else (due to potential overfitting)</p>",
      "rawMarkdown": "Di you use fragment 1 for training? In the competition data that is shared with us, real test data is replaced with copy of data from training fragment 1. If you model learned on fragment 1 then no wonder it can predict it. However it does not mean that it can predict anything else (due to potential overfitting)",
      "votes": null
    },
    {
      "id": "2294552",
      "postDate": "06/10/2023 06:39:47",
      "content": "<p>Ah, sure, I will explore the route of dropping training samples from fragment 1 to evaluate how much overfitting to fragment 1. </p>\n<p>However, I suspect it's huge lost of valuable data if the model does not get to learn from fragment 1 ever. I will try also increasing dropout rate and other regularization strategies.</p>\n<p>Thanks so much, Serhii, for the question/ idea!</p>",
      "rawMarkdown": "Ah, sure, I will explore the route of dropping training samples from fragment 1 to evaluate how much overfitting to fragment 1. \n\nHowever, I suspect it's huge lost of valuable data if the model does not get to learn from fragment 1 ever. I will try also increasing dropout rate and other regularization strategies.\n\nThanks so much, Serhii, for the question/ idea!",
      "votes": null
    },
    {
      "id": "2300683",
      "postDate": "06/13/2023 10:52:42",
      "content": "<p>That's what you can use ensembling for, if you train one model on frag 2+3 one on 1+3 and one on 1+2 for example, you can validate on unseen data for each model and then combine the prediction of the 3 models so that your final model incorporates information from all fragments. </p>\n<p>Keep in mind that this is not necessarily the best way for ensembling though. </p>\n<p>You can also split fragments into subfragments and validate only on 50% of frag 1 for example.</p>",
      "rawMarkdown": "That's what you can use ensembling for, if you train one model on frag 2+3 one on 1+3 and one on 1+2 for example, you can validate on unseen data for each model and then combine the prediction of the 3 models so that your final model incorporates information from all fragments. \n\nKeep in mind that this is not necessarily the best way for ensembling though. \n\nYou can also split fragments into subfragments and validate only on 50% of frag 1 for example.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2294349,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "06/10/2023 02:33:56",
      "content": "<p>Di you use fragment 1 for training? In the competition data that is shared with us, real test data is replaced with copy of data from training fragment 1. If you model learned on fragment 1 then no wonder it can predict it. However it does not mean that it can predict anything else (due to potential overfitting)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2294552,
          "author_name": "vcolliym",
          "author_url": "",
          "post_date": "06/10/2023 06:39:47",
          "content": "<p>Ah, sure, I will explore the route of dropping training samples from fragment 1 to evaluate how much overfitting to fragment 1. </p>\n<p>However, I suspect it's huge lost of valuable data if the model does not get to learn from fragment 1 ever. I will try also increasing dropout rate and other regularization strategies.</p>\n<p>Thanks so much, Serhii, for the question/ idea!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2300683,
              "author_name": "raki21",
              "author_url": "",
              "post_date": "06/13/2023 10:52:42",
              "content": "<p>That's what you can use ensembling for, if you train one model on frag 2+3 one on 1+3 and one on 1+2 for example, you can validate on unseen data for each model and then combine the prediction of the 3 models so that your final model incorporates information from all fragments. </p>\n<p>Keep in mind that this is not necessarily the best way for ensembling though. </p>\n<p>You can also split fragments into subfragments and validate only on 50% of frag 1 for example.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2294019": "Dear All,\n\nI appreciate any feedback on this topic of **\"visually good test results but low Leaderboard (LB) score\"**.\n\nI have been struggling for a long time on this problem but made no progress.\n\nI tried a couple of cross-validated models and my pipeline does not add padding.\n\nI did not use transpose because I have used the same Image Reader* (PIL) and RLE function as employed in a test submission notebook https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask (thank you Lucas!), my submission still received very low public score (0.06).\n\nOn the other hand, when the same models were evaluated on the training samples, the F1-beta score was 0.60-0.93 (for F1-beta score calculation, I used the implementation shown in the pinned tutorial).\n\nComparing the attached image with the \"mask\" used by the 0-11 simplest test submission, their corresponding LB scores do not seem to reflect readability of text materials (agreed?).\n\nAny suggestions please?\n\nThanks very much in advance for your help and insights!\n\n===========================================\n\nReferences:\n\nhttps://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask/comments\n\nhttps://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\n\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/413664#2294078",
    "2294349": "Di you use fragment 1 for training? In the competition data that is shared with us, real test data is replaced with copy of data from training fragment 1. If you model learned on fragment 1 then no wonder it can predict it. However it does not mean that it can predict anything else (due to potential overfitting)",
    "2294552": "Ah, sure, I will explore the route of dropping training samples from fragment 1 to evaluate how much overfitting to fragment 1. \n\nHowever, I suspect it's huge lost of valuable data if the model does not get to learn from fragment 1 ever. I will try also increasing dropout rate and other regularization strategies.\n\nThanks so much, Serhii, for the question/ idea!",
    "2300683": "That's what you can use ensembling for, if you train one model on frag 2+3 one on 1+3 and one on 1+2 for example, you can validate on unseen data for each model and then combine the prediction of the 3 models so that your final model incorporates information from all fragments. \n\nKeep in mind that this is not necessarily the best way for ensembling though. \n\nYou can also split fragments into subfragments and validate only on 50% of frag 1 for example."
  },
  "source": "meta"
}