{
  "id": 400451,
  "title": "CV Strategy and choosing confidence threshold?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/400451",
  "author_name": "",
  "post_date": "2023-04-08T12:36:38.818653600Z",
  "votes": 18,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I trained a 3D model using fragments 2 and 3, with fragment 1 as my test fragment. My best F0.5 score on this fragment is 0.6, using a confidence threshold of 0.75. But when I use the same model and conf threshold to submit, I get a LB score of 0.01 😂. Reducing the confidence threshold improves my LB score.</p>\n<p>Choosing the model and conf threshold based on 1 fragment seems to be risky, and not correlating with my best scores. Another option is to use random crops from all three fragments as validation, but that can lead to leaky folds, so this approach seems risky too.</p>\n<p>I'll try both these approaches and edit the results here later. Would love to hear other's experience regarding these topics!</p>",
  "messages": [
    {
      "id": "2214388",
      "postDate": "04/08/2023 12:36:38",
      "content": "<p>I trained a 3D model using fragments 2 and 3, with fragment 1 as my test fragment. My best F0.5 score on this fragment is 0.6, using a confidence threshold of 0.75. But when I use the same model and conf threshold to submit, I get a LB score of 0.01 😂. Reducing the confidence threshold improves my LB score.</p>\n<p>Choosing the model and conf threshold based on 1 fragment seems to be risky, and not correlating with my best scores. Another option is to use random crops from all three fragments as validation, but that can lead to leaky folds, so this approach seems risky too.</p>\n<p>I'll try both these approaches and edit the results here later. Would love to hear other's experience regarding these topics!</p>",
      "rawMarkdown": "I trained a 3D model using fragments 2 and 3, with fragment 1 as my test fragment. My best F0.5 score on this fragment is 0.6, using a confidence threshold of 0.75. But when I use the same model and conf threshold to submit, I get a LB score of 0.01 :joy:. Reducing the confidence threshold improves my LB score.\n\nChoosing the model and conf threshold based on 1 fragment seems to be risky, and not correlating with my best scores. Another option is to use random crops from all three fragments as validation, but that can lead to leaky folds, so this approach seems risky too.\n\nI'll try both these approaches and edit the results here later. Would love to hear other's experience regarding these topics!",
      "votes": null
    },
    {
      "id": "2214626",
      "postDate": "04/08/2023 16:19:39",
      "content": "<p>0.01 LB score  looks suspicious if you get 0.6 using 1st fragment for validation.</p>\n<p>I suggest using k-fold cross validation to get better idea about model's performance. </p>",
      "rawMarkdown": "0.01 LB score  looks suspicious if you get 0.6 using 1st fragment for validation.\n\nI suggest using k-fold cross validation to get better idea about model's performance.",
      "votes": null
    },
    {
      "id": "2214666",
      "postDate": "04/08/2023 16:40:06",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/danieliusk\" target=\"_blank\">@danieliusk</a>, 0.01 LB score looks suspicious, especially when a random number generator can get LB~0.1. Maybe there is some difference in data preparation if you have different data pipelines for training and testing.</p>",
      "rawMarkdown": "I agree with @danieliusk, 0.01 LB score looks suspicious, especially when a random number generator can get LB~0.1. Maybe there is some difference in data preparation if you have different data pipelines for training and testing.",
      "votes": null
    },
    {
      "id": "2215728",
      "postDate": "04/09/2023 14:43:04",
      "content": "<p>Updating my initial training results with 3-folds</p>\n<table>\n<thead>\n<tr>\n<th>Test fragment</th>\n<th>Conf threshold</th>\n<th>F0.5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fragment 1</td>\n<td>0.75</td>\n<td>0.597942</td>\n</tr>\n<tr>\n<td>Fragment 2</td>\n<td>0.4</td>\n<td>0.436897</td>\n</tr>\n<tr>\n<td>Fragment 3</td>\n<td>0.6</td>\n<td>0.632643</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Ensemble sub of all 3 models results in 0.07 LB score with 0.4 conf threshold, and 0.23 LB score with 0.1 conf threshold. Still way off than my CV.</li>\n<li>Although my training and sub's data pipelines are slightly different, I can reproduce these scores using my submission notebook. I've checked my code for hours, I cannot find what is wrong in my code which results in this difference.</li>\n<li>I'm suspecting there's something different in the test fragments, maybe the Z-DIM order is reversed. Will test these hypotheses and see if there's any improvements.</li>\n</ul>",
      "rawMarkdown": "Updating my initial training results with 3-folds\n\n| Test fragment | Conf threshold | F0.5|\n| --- | --- | --- |\n| Fragment 1 | 0.75 | 0.597942 |\n| Fragment 2 | 0.4 | 0.436897 |\n| Fragment 3 | 0.6 | 0.632643 |\n\n* Ensemble sub of all 3 models results in 0.07 LB score with 0.4 conf threshold, and 0.23 LB score with 0.1 conf threshold. Still way off than my CV.\n* Although my training and sub's data pipelines are slightly different, I can reproduce these scores using my submission notebook. I've checked my code for hours, I cannot find what is wrong in my code which results in this difference.\n* I'm suspecting there's something different in the test fragments, maybe the Z-DIM order is reversed. Will test these hypotheses and see if there's any improvements.",
      "votes": null
    },
    {
      "id": "2215819",
      "postDate": "04/09/2023 15:55:54",
      "content": "<p>I am pretty sure the Z-DIM order is not reversed. Is it possible you did RLE wrong?</p>",
      "rawMarkdown": "I am pretty sure the Z-DIM order is not reversed. Is it possible you did RLE wrong?",
      "votes": null
    },
    {
      "id": "2216160",
      "postDate": "04/09/2023 20:21:01",
      "content": "<p>I am facing the same exact problem, very strange</p>",
      "rawMarkdown": "I am facing the same exact problem, very strange",
      "votes": null
    },
    {
      "id": "2216704",
      "postDate": "04/10/2023 09:37:01",
      "content": "<p><a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> so I had the same issue and also my local CV numbers are close to yours, so I submitted all 3 separately and for me best model is when I have used id 1 as validation (it gave 0.31 lb score vs 0.07 lb score when I used id 3 as validation) - and if to add some augmentations it reached my current score of 0.44 - so I belive you have the same problem, try to submit separately all 3 folds - and also I think 3 fold CV is not a good way for local validation, becasue scores on LB jumps very high.</p>",
      "rawMarkdown": "samfc10 so I had the same issue and also my local CV numbers are close to yours, so I submitted all 3 separately and for me best model is when I have used id 1 as validation (it gave 0.31 lb score vs 0.07 lb score when I used id 3 as validation) - and if to add some augmentations it reached my current score of 0.44 - so I belive you have the same problem, try to submit separately all 3 folds - and also I think 3 fold CV is not a good way for local validation, becasue scores on LB jumps very high.",
      "votes": null
    },
    {
      "id": "2216755",
      "postDate": "04/10/2023 10:20:17",
      "content": "<p>Got it. I ran some more test submissions today, the problem is with my models only. And yes, I have trained these models with only H-Flip augmentation. I'll train with more augmentations and try again. Thanks for getting back and updating here!</p>",
      "rawMarkdown": "Got it. I ran some more test submissions today, the problem is with my models only. And yes, I have trained these models with only H-Flip augmentation. I'll train with more augmentations and try again. Thanks for getting back and updating here!",
      "votes": null
    },
    {
      "id": "2216941",
      "postDate": "04/10/2023 13:29:55",
      "content": "<p>Yeah, for me also 3fold samplewise validation doesn't work well.</p>",
      "rawMarkdown": "Yeah, for me also 3fold samplewise validation doesn't work well.",
      "votes": null
    },
    {
      "id": "2217523",
      "postDate": "04/11/2023 01:45:22",
      "content": "<p>One suggestion on choosing the confidence threshold:  sweep a range of confidence values for each fragment and choose it based on the distribution of resulting feature sizes.  Real ink is going to be on the order of millimeters to tens of millimeters in size.  If your resulting features are all single pixels, then (assuming your model is OK) your confidence threshold is too high and you need to drop it.  If your features are almost as big as the fragment, your confidence threshold is too low and you need to raise it.</p>",
      "rawMarkdown": "One suggestion on choosing the confidence threshold:  sweep a range of confidence values for each fragment and choose it based on the distribution of resulting feature sizes.  Real ink is going to be on the order of millimeters to tens of millimeters in size.  If your resulting features are all single pixels, then (assuming your model is OK) your confidence threshold is too high and you need to drop it.  If your features are almost as big as the fragment, your confidence threshold is too low and you need to raise it.",
      "votes": null
    },
    {
      "id": "2217741",
      "postDate": "04/11/2023 06:35:00",
      "content": "<p>Put a bit of work into this idea and it seems to work pretty well.  Tested it out with a (crummy) simple model and sweeping the threshold then picking the threshold that gives you the largest fraction of pixels within a given size range works pretty well.  See here: <a href=\"https://www.kaggle.com/code/brettolsen/adaptive-threshold-selection-using-feature-sizes\" target=\"_blank\">https://www.kaggle.com/code/brettolsen/adaptive-threshold-selection-using-feature-sizes</a></p>",
      "rawMarkdown": "Put a bit of work into this idea and it seems to work pretty well.  Tested it out with a (crummy) simple model and sweeping the threshold then picking the threshold that gives you the largest fraction of pixels within a given size range works pretty well.  See here: https://www.kaggle.com/code/brettolsen/adaptive-threshold-selection-using-feature-sizes",
      "votes": null
    },
    {
      "id": "2218158",
      "postDate": "04/11/2023 13:18:08",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> <br>\nCan you explain confidence levels for noobies? Thanks!</p>",
      "rawMarkdown": "Hey @samfc10 \nCan you explain confidence levels for noobies? Thanks!",
      "votes": null
    },
    {
      "id": "2220414",
      "postDate": "04/13/2023 11:29:26",
      "content": "<p>Dear sir. do we have any updated score ?.<br>\nI have tried your approach but no luck for me that i got worse score than you. I can't improve it too.<br>\nCan you improve your score ?</p>",
      "rawMarkdown": "Dear sir. do we have any updated score ?.\nI have tried your approach but no luck for me that i got worse score than you. I can't improve it too.\nCan you improve your score ?",
      "votes": null
    },
    {
      "id": "2221314",
      "postDate": "04/14/2023 06:21:53",
      "content": "<p>Not yet. Did not get much time this week to train models.</p>",
      "rawMarkdown": "Not yet. Did not get much time this week to train models.",
      "votes": null
    },
    {
      "id": "2228220",
      "postDate": "04/20/2023 11:15:39",
      "content": "<p><a href=\"https://www.kaggle.com/maksimovka\" target=\"_blank\">@maksimovka</a> <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> <br>\nThank you for the insightful discussion.<br>\nDoes changing the augmentation solve the problem in this topic?</p>\n<p>I have been experimenting with various augmentations, and although I'm getting a CV of around 0.6 with 5-fold, my LB score is not good.</p>",
      "rawMarkdown": "maksimovka @samfc10 \nThank you for the insightful discussion.\nDoes changing the augmentation solve the problem in this topic?\n\nI have been experimenting with various augmentations, and although I'm getting a CV of around 0.6 with 5-fold, my LB score is not good.",
      "votes": null
    },
    {
      "id": "2228328",
      "postDate": "04/20/2023 13:15:23",
      "content": "<p>Yeah, augmentations solved the issue for me.</p>",
      "rawMarkdown": "Yeah, augmentations solved the issue for me.",
      "votes": null
    },
    {
      "id": "2228335",
      "postDate": "04/20/2023 13:23:40",
      "content": "<p>Oh great job!<br>\nThank you for the information! Let's both do our best!</p>",
      "rawMarkdown": "Oh great job!\nThank you for the information! Let's both do our best!",
      "votes": null
    },
    {
      "id": "2232082",
      "postDate": "04/24/2023 02:12:27",
      "content": "<p>seems like you solved the problem. Can you share it? I face the same problem…</p>",
      "rawMarkdown": "seems like you solved the problem. Can you share it? I face the same problem...",
      "votes": null
    },
    {
      "id": "2234467",
      "postDate": "04/25/2023 08:16:21",
      "content": "<p>Indeed, also submitting the test mask gives a score of 0.11 on the LB</p>",
      "rawMarkdown": "Indeed, also submitting the test mask gives a score of 0.11 on the LB",
      "votes": null
    },
    {
      "id": "2234548",
      "postDate": "04/25/2023 09:32:12",
      "content": "<p><a href=\"https://www.kaggle.com/ktakita\" target=\"_blank\">@ktakita</a> <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> <a href=\"https://www.kaggle.com/maksimovka\" target=\"_blank\">@maksimovka</a> are you talking about adding test time augmentations or just use augmentations during training? I'm currently only doing the latter and when submitting my 3 models separately  (e.g. one model but trained 3 times on single folds), I'm seeing more or less Local 0.5 vs LB 0.2.</p>",
      "rawMarkdown": "ktakita @samfc10 @maksimovka are you talking about adding test time augmentations or just use augmentations during training? I'm currently only doing the latter and when submitting my 3 models separately  (e.g. one model but trained 3 times on single folds), I'm seeing more or less Local 0.5 vs LB 0.2.",
      "votes": null
    },
    {
      "id": "2234553",
      "postDate": "04/25/2023 09:35:35",
      "content": "<p>The confidence level describes the cut-off level to determine whether something is ink vs no-ink. In the end, the model will output probabilities of a pixel being ink. This value is between 0 and 1. But for submission we need to output whether we believe something is ink or not, and not the probability. So we use a threshold (confidence value) to say: \"if the probability of a pixel being ink is being estimated by our model to be larger then 0.6 (or 0.5, or 0.4 or whatever) we classify the pixel as being ink\"</p>",
      "rawMarkdown": "The confidence level describes the cut-off level to determine whether something is ink vs no-ink. In the end, the model will output probabilities of a pixel being ink. This value is between 0 and 1. But for submission we need to output whether we believe something is ink or not, and not the probability. So we use a threshold (confidence value) to say: \"if the probability of a pixel being ink is being estimated by our model to be larger then 0.6 (or 0.5, or 0.4 or whatever) we classify the pixel as being ink\"",
      "votes": null
    },
    {
      "id": "2234677",
      "postDate": "04/25/2023 11:58:20",
      "content": "<p>augmentations during training. Haven't tried TTA yet.</p>",
      "rawMarkdown": "augmentations during training. Haven't tried TTA yet.",
      "votes": null
    },
    {
      "id": "2240474",
      "postDate": "04/30/2023 15:11:15",
      "content": "<p>Jebastin, what was the issue in the end?</p>",
      "rawMarkdown": "Jebastin, what was the issue in the end?",
      "votes": null
    },
    {
      "id": "2242354",
      "postDate": "05/02/2023 07:51:16",
      "content": "<p>Thanks, Im doing the same at the moment</p>",
      "rawMarkdown": "Thanks, Im doing the same at the moment",
      "votes": null
    },
    {
      "id": "2282055",
      "postDate": "05/31/2023 10:27:10",
      "content": "<p>Did you get an error while submitting? Maybe it used a submission file that wasn't correctly filled?</p>",
      "rawMarkdown": "Did you get an error while submitting? Maybe it used a submission file that wasn't correctly filled?",
      "votes": null
    },
    {
      "id": "2282056",
      "postDate": "05/31/2023 10:27:16",
      "content": "<p>It seems that random crops from the 3 segments is a better strategy than a whole segment. </p>",
      "rawMarkdown": "It seems that random crops from the 3 segments is a better strategy than a whole segment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2214626,
      "author_name": "danieliusk",
      "author_url": "",
      "post_date": "04/08/2023 16:19:39",
      "content": "<p>0.01 LB score  looks suspicious if you get 0.6 using 1st fragment for validation.</p>\n<p>I suggest using k-fold cross validation to get better idea about model's performance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2214666,
      "author_name": "junxhuang",
      "author_url": "",
      "post_date": "04/08/2023 16:40:06",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/danieliusk\" target=\"_blank\">@danieliusk</a>, 0.01 LB score looks suspicious, especially when a random number generator can get LB~0.1. Maybe there is some difference in data preparation if you have different data pipelines for training and testing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2234467,
          "author_name": "lucasvw",
          "author_url": "",
          "post_date": "04/25/2023 08:16:21",
          "content": "<p>Indeed, also submitting the test mask gives a score of 0.11 on the LB</p>",
          "votes": null,
          "replies": [
            {
              "id": 2282055,
              "author_name": "yassinealouini",
              "author_url": "",
              "post_date": "05/31/2023 10:27:10",
              "content": "<p>Did you get an error while submitting? Maybe it used a submission file that wasn't correctly filled?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2215728,
      "author_name": "samfc10",
      "author_url": "",
      "post_date": "04/09/2023 14:43:04",
      "content": "<p>Updating my initial training results with 3-folds</p>\n<table>\n<thead>\n<tr>\n<th>Test fragment</th>\n<th>Conf threshold</th>\n<th>F0.5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fragment 1</td>\n<td>0.75</td>\n<td>0.597942</td>\n</tr>\n<tr>\n<td>Fragment 2</td>\n<td>0.4</td>\n<td>0.436897</td>\n</tr>\n<tr>\n<td>Fragment 3</td>\n<td>0.6</td>\n<td>0.632643</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Ensemble sub of all 3 models results in 0.07 LB score with 0.4 conf threshold, and 0.23 LB score with 0.1 conf threshold. Still way off than my CV.</li>\n<li>Although my training and sub's data pipelines are slightly different, I can reproduce these scores using my submission notebook. I've checked my code for hours, I cannot find what is wrong in my code which results in this difference.</li>\n<li>I'm suspecting there's something different in the test fragments, maybe the Z-DIM order is reversed. Will test these hypotheses and see if there's any improvements.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2215819,
          "author_name": "junxhuang",
          "author_url": "",
          "post_date": "04/09/2023 15:55:54",
          "content": "<p>I am pretty sure the Z-DIM order is not reversed. Is it possible you did RLE wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2216160,
          "author_name": "maksimovka",
          "author_url": "",
          "post_date": "04/09/2023 20:21:01",
          "content": "<p>I am facing the same exact problem, very strange</p>",
          "votes": null,
          "replies": [
            {
              "id": 2216704,
              "author_name": "maksimovka",
              "author_url": "",
              "post_date": "04/10/2023 09:37:01",
              "content": "<p><a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> so I had the same issue and also my local CV numbers are close to yours, so I submitted all 3 separately and for me best model is when I have used id 1 as validation (it gave 0.31 lb score vs 0.07 lb score when I used id 3 as validation) - and if to add some augmentations it reached my current score of 0.44 - so I belive you have the same problem, try to submit separately all 3 folds - and also I think 3 fold CV is not a good way for local validation, becasue scores on LB jumps very high.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2216755,
                  "author_name": "samfc10",
                  "author_url": "",
                  "post_date": "04/10/2023 10:20:17",
                  "content": "<p>Got it. I ran some more test submissions today, the problem is with my models only. And yes, I have trained these models with only H-Flip augmentation. I'll train with more augmentations and try again. Thanks for getting back and updating here!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2216941,
                      "author_name": "igorkrashenyi",
                      "author_url": "",
                      "post_date": "04/10/2023 13:29:55",
                      "content": "<p>Yeah, for me also 3fold samplewise validation doesn't work well.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2228220,
                      "author_name": "ktakita",
                      "author_url": "",
                      "post_date": "04/20/2023 11:15:39",
                      "content": "<p><a href=\"https://www.kaggle.com/maksimovka\" target=\"_blank\">@maksimovka</a> <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> <br>\nThank you for the insightful discussion.<br>\nDoes changing the augmentation solve the problem in this topic?</p>\n<p>I have been experimenting with various augmentations, and although I'm getting a CV of around 0.6 with 5-fold, my LB score is not good.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2228328,
                          "author_name": "samfc10",
                          "author_url": "",
                          "post_date": "04/20/2023 13:15:23",
                          "content": "<p>Yeah, augmentations solved the issue for me.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2228335,
                              "author_name": "ktakita",
                              "author_url": "",
                              "post_date": "04/20/2023 13:23:40",
                              "content": "<p>Oh great job!<br>\nThank you for the information! Let's both do our best!</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2234548,
                                  "author_name": "lucasvw",
                                  "author_url": "",
                                  "post_date": "04/25/2023 09:32:12",
                                  "content": "<p><a href=\"https://www.kaggle.com/ktakita\" target=\"_blank\">@ktakita</a> <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> <a href=\"https://www.kaggle.com/maksimovka\" target=\"_blank\">@maksimovka</a> are you talking about adding test time augmentations or just use augmentations during training? I'm currently only doing the latter and when submitting my 3 models separately  (e.g. one model but trained 3 times on single folds), I'm seeing more or less Local 0.5 vs LB 0.2.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2234677,
                                      "author_name": "samfc10",
                                      "author_url": "",
                                      "post_date": "04/25/2023 11:58:20",
                                      "content": "<p>augmentations during training. Haven't tried TTA yet.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2242354,
                                          "author_name": "lucasvw",
                                          "author_url": "",
                                          "post_date": "05/02/2023 07:51:16",
                                          "content": "<p>Thanks, Im doing the same at the moment</p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2220414,
          "author_name": "hutedu",
          "author_url": "",
          "post_date": "04/13/2023 11:29:26",
          "content": "<p>Dear sir. do we have any updated score ?.<br>\nI have tried your approach but no luck for me that i got worse score than you. I can't improve it too.<br>\nCan you improve your score ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2221314,
              "author_name": "samfc10",
              "author_url": "",
              "post_date": "04/14/2023 06:21:53",
              "content": "<p>Not yet. Did not get much time this week to train models.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2217523,
      "author_name": "brettolsen",
      "author_url": "",
      "post_date": "04/11/2023 01:45:22",
      "content": "<p>One suggestion on choosing the confidence threshold:  sweep a range of confidence values for each fragment and choose it based on the distribution of resulting feature sizes.  Real ink is going to be on the order of millimeters to tens of millimeters in size.  If your resulting features are all single pixels, then (assuming your model is OK) your confidence threshold is too high and you need to drop it.  If your features are almost as big as the fragment, your confidence threshold is too low and you need to raise it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2217741,
          "author_name": "brettolsen",
          "author_url": "",
          "post_date": "04/11/2023 06:35:00",
          "content": "<p>Put a bit of work into this idea and it seems to work pretty well.  Tested it out with a (crummy) simple model and sweeping the threshold then picking the threshold that gives you the largest fraction of pixels within a given size range works pretty well.  See here: <a href=\"https://www.kaggle.com/code/brettolsen/adaptive-threshold-selection-using-feature-sizes\" target=\"_blank\">https://www.kaggle.com/code/brettolsen/adaptive-threshold-selection-using-feature-sizes</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2218158,
      "author_name": "ppb00x",
      "author_url": "",
      "post_date": "04/11/2023 13:18:08",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> <br>\nCan you explain confidence levels for noobies? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2234553,
          "author_name": "lucasvw",
          "author_url": "",
          "post_date": "04/25/2023 09:35:35",
          "content": "<p>The confidence level describes the cut-off level to determine whether something is ink vs no-ink. In the end, the model will output probabilities of a pixel being ink. This value is between 0 and 1. But for submission we need to output whether we believe something is ink or not, and not the probability. So we use a threshold (confidence value) to say: \"if the probability of a pixel being ink is being estimated by our model to be larger then 0.6 (or 0.5, or 0.4 or whatever) we classify the pixel as being ink\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2232082,
      "author_name": "huyidao",
      "author_url": "",
      "post_date": "04/24/2023 02:12:27",
      "content": "<p>seems like you solved the problem. Can you share it? I face the same problem…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2240474,
      "author_name": "jozefsebestyen",
      "author_url": "",
      "post_date": "04/30/2023 15:11:15",
      "content": "<p>Jebastin, what was the issue in the end?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2282056,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "05/31/2023 10:27:16",
      "content": "<p>It seems that random crops from the 3 segments is a better strategy than a whole segment. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2214388": "I trained a 3D model using fragments 2 and 3, with fragment 1 as my test fragment. My best F0.5 score on this fragment is 0.6, using a confidence threshold of 0.75. But when I use the same model and conf threshold to submit, I get a LB score of 0.01 :joy:. Reducing the confidence threshold improves my LB score.\n\nChoosing the model and conf threshold based on 1 fragment seems to be risky, and not correlating with my best scores. Another option is to use random crops from all three fragments as validation, but that can lead to leaky folds, so this approach seems risky too.\n\nI'll try both these approaches and edit the results here later. Would love to hear other's experience regarding these topics!",
    "2214626": "0.01 LB score  looks suspicious if you get 0.6 using 1st fragment for validation.\n\nI suggest using k-fold cross validation to get better idea about model's performance.",
    "2214666": "I agree with @danieliusk, 0.01 LB score looks suspicious, especially when a random number generator can get LB~0.1. Maybe there is some difference in data preparation if you have different data pipelines for training and testing.",
    "2215728": "Updating my initial training results with 3-folds\n\n| Test fragment | Conf threshold | F0.5|\n| --- | --- | --- |\n| Fragment 1 | 0.75 | 0.597942 |\n| Fragment 2 | 0.4 | 0.436897 |\n| Fragment 3 | 0.6 | 0.632643 |\n\n* Ensemble sub of all 3 models results in 0.07 LB score with 0.4 conf threshold, and 0.23 LB score with 0.1 conf threshold. Still way off than my CV.\n* Although my training and sub's data pipelines are slightly different, I can reproduce these scores using my submission notebook. I've checked my code for hours, I cannot find what is wrong in my code which results in this difference.\n* I'm suspecting there's something different in the test fragments, maybe the Z-DIM order is reversed. Will test these hypotheses and see if there's any improvements.",
    "2215819": "I am pretty sure the Z-DIM order is not reversed. Is it possible you did RLE wrong?",
    "2216160": "I am facing the same exact problem, very strange",
    "2216704": "samfc10 so I had the same issue and also my local CV numbers are close to yours, so I submitted all 3 separately and for me best model is when I have used id 1 as validation (it gave 0.31 lb score vs 0.07 lb score when I used id 3 as validation) - and if to add some augmentations it reached my current score of 0.44 - so I belive you have the same problem, try to submit separately all 3 folds - and also I think 3 fold CV is not a good way for local validation, becasue scores on LB jumps very high.",
    "2216755": "Got it. I ran some more test submissions today, the problem is with my models only. And yes, I have trained these models with only H-Flip augmentation. I'll train with more augmentations and try again. Thanks for getting back and updating here!",
    "2216941": "Yeah, for me also 3fold samplewise validation doesn't work well.",
    "2217523": "One suggestion on choosing the confidence threshold:  sweep a range of confidence values for each fragment and choose it based on the distribution of resulting feature sizes.  Real ink is going to be on the order of millimeters to tens of millimeters in size.  If your resulting features are all single pixels, then (assuming your model is OK) your confidence threshold is too high and you need to drop it.  If your features are almost as big as the fragment, your confidence threshold is too low and you need to raise it.",
    "2217741": "Put a bit of work into this idea and it seems to work pretty well.  Tested it out with a (crummy) simple model and sweeping the threshold then picking the threshold that gives you the largest fraction of pixels within a given size range works pretty well.  See here: https://www.kaggle.com/code/brettolsen/adaptive-threshold-selection-using-feature-sizes",
    "2218158": "Hey @samfc10 \nCan you explain confidence levels for noobies? Thanks!",
    "2220414": "Dear sir. do we have any updated score ?.\nI have tried your approach but no luck for me that i got worse score than you. I can't improve it too.\nCan you improve your score ?",
    "2221314": "Not yet. Did not get much time this week to train models.",
    "2228220": "maksimovka @samfc10 \nThank you for the insightful discussion.\nDoes changing the augmentation solve the problem in this topic?\n\nI have been experimenting with various augmentations, and although I'm getting a CV of around 0.6 with 5-fold, my LB score is not good.",
    "2228328": "Yeah, augmentations solved the issue for me.",
    "2228335": "Oh great job!\nThank you for the information! Let's both do our best!",
    "2232082": "seems like you solved the problem. Can you share it? I face the same problem...",
    "2234467": "Indeed, also submitting the test mask gives a score of 0.11 on the LB",
    "2234548": "ktakita @samfc10 @maksimovka are you talking about adding test time augmentations or just use augmentations during training? I'm currently only doing the latter and when submitting my 3 models separately  (e.g. one model but trained 3 times on single folds), I'm seeing more or less Local 0.5 vs LB 0.2.",
    "2234553": "The confidence level describes the cut-off level to determine whether something is ink vs no-ink. In the end, the model will output probabilities of a pixel being ink. This value is between 0 and 1. But for submission we need to output whether we believe something is ink or not, and not the probability. So we use a threshold (confidence value) to say: \"if the probability of a pixel being ink is being estimated by our model to be larger then 0.6 (or 0.5, or 0.4 or whatever) we classify the pixel as being ink\"",
    "2234677": "augmentations during training. Haven't tried TTA yet.",
    "2240474": "Jebastin, what was the issue in the end?",
    "2242354": "Thanks, Im doing the same at the moment",
    "2282055": "Did you get an error while submitting? Maybe it used a submission file that wasn't correctly filled?",
    "2282056": "It seems that random crops from the 3 segments is a better strategy than a whole segment."
  },
  "source": "meta"
}