{
  "id": 401667,
  "title": "CV vs LB Thread",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/401667",
  "author_name": "",
  "post_date": "2023-04-14T10:28:54.981224800Z",
  "votes": 13,
  "comment_count": 15,
  "views": 0,
  "content": "<p>According to my most recent multiple experiments, LB varies significantly even though CV and optimal thresholds remain the same.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.648</td>\n<td>0.58</td>\n</tr>\n<tr>\n<td>0.644</td>\n<td>0.56</td>\n</tr>\n<tr>\n<td>0.650</td>\n<td>0.59</td>\n</tr>\n<tr>\n<td>0.645</td>\n<td>0.61</td>\n</tr>\n</tbody>\n</table>\n<p>I use is 5fold CV (fragment2 is divided into 3 parts in the height direction so that the number of pixels in the mask is equal).<br>\nAlso, the appropriate threshold for public LB score seems to be higher than the one done with CV.<br>\nWhat is the correlation between your CV and LB?</p>",
  "messages": [
    {
      "id": "2221515",
      "postDate": "04/14/2023 10:28:54",
      "content": "<p>According to my most recent multiple experiments, LB varies significantly even though CV and optimal thresholds remain the same.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.648</td>\n<td>0.58</td>\n</tr>\n<tr>\n<td>0.644</td>\n<td>0.56</td>\n</tr>\n<tr>\n<td>0.650</td>\n<td>0.59</td>\n</tr>\n<tr>\n<td>0.645</td>\n<td>0.61</td>\n</tr>\n</tbody>\n</table>\n<p>I use is 5fold CV (fragment2 is divided into 3 parts in the height direction so that the number of pixels in the mask is equal).<br>\nAlso, the appropriate threshold for public LB score seems to be higher than the one done with CV.<br>\nWhat is the correlation between your CV and LB?</p>",
      "rawMarkdown": "According to my most recent multiple experiments, LB varies significantly even though CV and optimal thresholds remain the same.\n| CV | LB |\n| --- | --- |\n|  0.648 | 0.58 |\n|  0.644 | 0.56 |\n|  0.650 | 0.59 |\n|  0.645 | 0.61 |\n\n\nI use is 5fold CV (fragment2 is divided into 3 parts in the height direction so that the number of pixels in the mask is equal).\nAlso, the appropriate threshold for public LB score seems to be higher than the one done with CV.\nWhat is the correlation between your CV and LB?",
      "votes": null
    },
    {
      "id": "2221539",
      "postDate": "04/14/2023 10:48:58",
      "content": "<p>Are these scores from a single model or after ensembling multiple models?  </p>",
      "rawMarkdown": "Are these scores from a single model or after ensembling multiple models?",
      "votes": null
    },
    {
      "id": "2221540",
      "postDate": "04/14/2023 10:49:38",
      "content": "<p>Mine is awful, I am talking about 0.6 CV and 0.07 LB… But this is surely an overfitting problem, still trying to debug it.</p>",
      "rawMarkdown": "Mine is awful, I am talking about 0.6 CV and 0.07 LB... But this is surely an overfitting problem, still trying to debug it.",
      "votes": null
    },
    {
      "id": "2221543",
      "postDate": "04/14/2023 10:51:45",
      "content": "<p>Simple average of 5 fold models.</p>",
      "rawMarkdown": "Simple average of 5 fold models.",
      "votes": null
    },
    {
      "id": "2221697",
      "postDate": "04/14/2023 13:30:56",
      "content": "<p>you train one class with sigmoid, your oprimal threshold should 0.5</p>",
      "rawMarkdown": "you train one class with sigmoid, your oprimal threshold should 0.5",
      "votes": null
    },
    {
      "id": "2221699",
      "postDate": "04/14/2023 13:31:48",
      "content": "<p>check this thread <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/400451\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/400451</a></p>",
      "rawMarkdown": "check this thread https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/400451",
      "votes": null
    },
    {
      "id": "2221830",
      "postDate": "04/14/2023 16:07:31",
      "content": "<p>Well, I have the same situation and I am pretty sure it is because public is around 10% of the test data, you will expect some deviation. You could train your model with 10 KFold and measure the difference of the score between folds, that gab should also be expected on the public lb. </p>\n<p>In my case:</p>\n<p>CV;LB<br>\n0.5928; 0.41<br>\n0.6064; 0.51<br>\n0.6183; 0.56<br>\n0.6298; 0.52</p>\n<p>Not very stable 🙃</p>",
      "rawMarkdown": "Well, I have the same situation and I am pretty sure it is because public is around 10% of the test data, you will expect some deviation. You could train your model with 10 KFold and measure the difference of the score between folds, that gab should also be expected on the public lb. \n\nIn my case:\n\nCV;LB\n0.5928; 0.41\n0.6064; 0.51\n0.6183; 0.56\n0.6298; 0.52\n\nNot very stable 🙃",
      "votes": null
    },
    {
      "id": "2222190",
      "postDate": "04/15/2023 02:55:29",
      "content": "<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6007</td>\n<td>0.59</td>\n</tr>\n<tr>\n<td>0.6024</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>0.6143</td>\n<td>0.61</td>\n</tr>\n<tr>\n<td>0.6269</td>\n<td>0.63</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "| CV | LB |\n| --- | --- |\n| 0.6007 | 0.59 |\n| 0.6024 | 0.6 |\n| 0.6143| 0.61 |\n| 0.6269| 0.63 |",
      "votes": null
    },
    {
      "id": "2222277",
      "postDate": "04/15/2023 05:33:12",
      "content": "<p>wow, that looks stable! Do you mind sharing your CV strategy?</p>",
      "rawMarkdown": "wow, that looks stable! Do you mind sharing your CV strategy?",
      "votes": null
    },
    {
      "id": "2222626",
      "postDate": "04/15/2023 12:33:22",
      "content": "<p>by ensemble 3 fold model</p>",
      "rawMarkdown": "by ensemble 3 fold model",
      "votes": null
    },
    {
      "id": "2229847",
      "postDate": "04/21/2023 18:48:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> , could you please share more about your 5fold CV. I also tried to divided fragment2 into 3 parts but it did not work well. Thank you.</p>",
      "rawMarkdown": "Hi @tattaka , could you please share more about your 5fold CV. I also tried to divided fragment2 into 3 parts but it did not work well. Thank you.",
      "votes": null
    },
    {
      "id": "2234512",
      "postDate": "04/25/2023 08:58:43",
      "content": "<p>Here are some of my results. I split fragment 2 into two parts and used the middle 4 z-slices.</p>\n<table>\n<thead>\n<tr>\n<th>Val Fragment</th>\n<th>F0.5</th>\n<th>Thresh</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.5626</td>\n<td>0.81</td>\n<td>0.51</td>\n</tr>\n<tr>\n<td>2a</td>\n<td>0.5673</td>\n<td>0.81</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>2b</td>\n<td>0.5129</td>\n<td>0.76</td>\n<td>0.44</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.6197</td>\n<td>0.91</td>\n<td>0.36</td>\n</tr>\n</tbody>\n</table>\n<p>Don't know why there are such differences between the folds</p>",
      "rawMarkdown": "Here are some of my results. I split fragment 2 into two parts and used the middle 4 z-slices.\n\n| Val Fragment | F0.5 | Thresh | LB |\n| --- | --- | --- | --- |\n| 1 | 0.5626 | 0.81 | 0.51 |\n| 2a | 0.5673 | 0.81 | 0.63 |\n| 2b | 0.5129 | 0.76 | 0.44 |\n| 3 | 0.6197 | 0.91 | 0.36 |\n\nDon't know why there are such differences between the folds",
      "votes": null
    },
    {
      "id": "2236399",
      "postDate": "04/26/2023 20:27:34",
      "content": "<p>CV - 0.58  Th - 0.5   LB - 0.6 - fragement 1<br>\nCV - 0.57  Th - 0.45   LB - 0.57 - fragement 3<br>\nCV - 0.56  Th - 0.5   LB - 0.53 - fragement 2</p>",
      "rawMarkdown": "CV - 0.58  Th - 0.5   LB - 0.6 - fragement 1\nCV - 0.57  Th - 0.45   LB - 0.57 - fragement 3\nCV - 0.56  Th - 0.5   LB - 0.53 - fragement 2",
      "votes": null
    },
    {
      "id": "2242886",
      "postDate": "05/02/2023 15:00:25",
      "content": "<p>what is model architecture you are using, I was using linknet with resnet50 a backbone I got </p>\n<table>\n<thead>\n<tr>\n<th>image used for CV</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>IMAGE 1</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>IMAGE 2</td>\n<td>0.38</td>\n</tr>\n<tr>\n<td>IMAGE 3</td>\n<td>0.55</td>\n</tr>\n</tbody>\n</table>\n<p>current LB score is 0.39 for this model, I am not able to improve any further. do you have some tips for this?<br>\nimage size using is 224</p>",
      "rawMarkdown": "what is model architecture you are using, I was using linknet with resnet50 a backbone I got \n| image used for CV | CV|\n| --- | --- |\n| IMAGE 1 | 0.48 |\n| IMAGE 2|  0.38|\n| IMAGE 3 |  0.55|\n\ncurrent LB score is 0.39 for this model, I am not able to improve any further. do you have some tips for this?\nimage size using is 224",
      "votes": null
    },
    {
      "id": "2271138",
      "postDate": "05/23/2023 16:35:31",
      "content": "<p>Can I ask you why your TH can be set so high and what kind of loss are you using?</p>",
      "rawMarkdown": "Can I ask you why your TH can be set so high and what kind of loss are you using?",
      "votes": null
    },
    {
      "id": "2271330",
      "postDate": "05/23/2023 19:27:15",
      "content": "<p>So just to confirm, the CV reported is the result of the average scores of the individual folds/fragments, right ?</p>",
      "rawMarkdown": "So just to confirm, the CV reported is the result of the average scores of the individual folds/fragments, right ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2221539,
      "author_name": "petersk20",
      "author_url": "",
      "post_date": "04/14/2023 10:48:58",
      "content": "<p>Are these scores from a single model or after ensembling multiple models?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2221543,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "04/14/2023 10:51:45",
          "content": "<p>Simple average of 5 fold models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2221540,
      "author_name": "fpeccia",
      "author_url": "",
      "post_date": "04/14/2023 10:49:38",
      "content": "<p>Mine is awful, I am talking about 0.6 CV and 0.07 LB… But this is surely an overfitting problem, still trying to debug it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2221699,
          "author_name": "maksimovka",
          "author_url": "",
          "post_date": "04/14/2023 13:31:48",
          "content": "<p>check this thread <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/400451\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/400451</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2221697,
      "author_name": "maksimovka",
      "author_url": "",
      "post_date": "04/14/2023 13:30:56",
      "content": "<p>you train one class with sigmoid, your oprimal threshold should 0.5</p>",
      "votes": null,
      "replies": [
        {
          "id": 2242886,
          "author_name": "lucario129",
          "author_url": "",
          "post_date": "05/02/2023 15:00:25",
          "content": "<p>what is model architecture you are using, I was using linknet with resnet50 a backbone I got </p>\n<table>\n<thead>\n<tr>\n<th>image used for CV</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>IMAGE 1</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>IMAGE 2</td>\n<td>0.38</td>\n</tr>\n<tr>\n<td>IMAGE 3</td>\n<td>0.55</td>\n</tr>\n</tbody>\n</table>\n<p>current LB score is 0.39 for this model, I am not able to improve any further. do you have some tips for this?<br>\nimage size using is 224</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2221830,
      "author_name": "ragnar123",
      "author_url": "",
      "post_date": "04/14/2023 16:07:31",
      "content": "<p>Well, I have the same situation and I am pretty sure it is because public is around 10% of the test data, you will expect some deviation. You could train your model with 10 KFold and measure the difference of the score between folds, that gab should also be expected on the public lb. </p>\n<p>In my case:</p>\n<p>CV;LB<br>\n0.5928; 0.41<br>\n0.6064; 0.51<br>\n0.6183; 0.56<br>\n0.6298; 0.52</p>\n<p>Not very stable 🙃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2222190,
      "author_name": "tanyong666",
      "author_url": "",
      "post_date": "04/15/2023 02:55:29",
      "content": "<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6007</td>\n<td>0.59</td>\n</tr>\n<tr>\n<td>0.6024</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>0.6143</td>\n<td>0.61</td>\n</tr>\n<tr>\n<td>0.6269</td>\n<td>0.63</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2222277,
          "author_name": "danieliusk",
          "author_url": "",
          "post_date": "04/15/2023 05:33:12",
          "content": "<p>wow, that looks stable! Do you mind sharing your CV strategy?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2222626,
              "author_name": "tanyong666",
              "author_url": "",
              "post_date": "04/15/2023 12:33:22",
              "content": "<p>by ensemble 3 fold model</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2271330,
                  "author_name": "imeintanis",
                  "author_url": "",
                  "post_date": "05/23/2023 19:27:15",
                  "content": "<p>So just to confirm, the CV reported is the result of the average scores of the individual folds/fragments, right ?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2229847,
      "author_name": "gianghus",
      "author_url": "",
      "post_date": "04/21/2023 18:48:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> , could you please share more about your 5fold CV. I also tried to divided fragment2 into 3 parts but it did not work well. Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234512,
      "author_name": "clemchris",
      "author_url": "",
      "post_date": "04/25/2023 08:58:43",
      "content": "<p>Here are some of my results. I split fragment 2 into two parts and used the middle 4 z-slices.</p>\n<table>\n<thead>\n<tr>\n<th>Val Fragment</th>\n<th>F0.5</th>\n<th>Thresh</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.5626</td>\n<td>0.81</td>\n<td>0.51</td>\n</tr>\n<tr>\n<td>2a</td>\n<td>0.5673</td>\n<td>0.81</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>2b</td>\n<td>0.5129</td>\n<td>0.76</td>\n<td>0.44</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.6197</td>\n<td>0.91</td>\n<td>0.36</td>\n</tr>\n</tbody>\n</table>\n<p>Don't know why there are such differences between the folds</p>",
      "votes": null,
      "replies": [
        {
          "id": 2271138,
          "author_name": "xiaoqinglong1996",
          "author_url": "",
          "post_date": "05/23/2023 16:35:31",
          "content": "<p>Can I ask you why your TH can be set so high and what kind of loss are you using?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2236399,
      "author_name": "arunodhayan",
      "author_url": "",
      "post_date": "04/26/2023 20:27:34",
      "content": "<p>CV - 0.58  Th - 0.5   LB - 0.6 - fragement 1<br>\nCV - 0.57  Th - 0.45   LB - 0.57 - fragement 3<br>\nCV - 0.56  Th - 0.5   LB - 0.53 - fragement 2</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2221515": "According to my most recent multiple experiments, LB varies significantly even though CV and optimal thresholds remain the same.\n| CV | LB |\n| --- | --- |\n|  0.648 | 0.58 |\n|  0.644 | 0.56 |\n|  0.650 | 0.59 |\n|  0.645 | 0.61 |\n\n\nI use is 5fold CV (fragment2 is divided into 3 parts in the height direction so that the number of pixels in the mask is equal).\nAlso, the appropriate threshold for public LB score seems to be higher than the one done with CV.\nWhat is the correlation between your CV and LB?",
    "2221539": "Are these scores from a single model or after ensembling multiple models?",
    "2221540": "Mine is awful, I am talking about 0.6 CV and 0.07 LB... But this is surely an overfitting problem, still trying to debug it.",
    "2221543": "Simple average of 5 fold models.",
    "2221697": "you train one class with sigmoid, your oprimal threshold should 0.5",
    "2221699": "check this thread https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/400451",
    "2221830": "Well, I have the same situation and I am pretty sure it is because public is around 10% of the test data, you will expect some deviation. You could train your model with 10 KFold and measure the difference of the score between folds, that gab should also be expected on the public lb. \n\nIn my case:\n\nCV;LB\n0.5928; 0.41\n0.6064; 0.51\n0.6183; 0.56\n0.6298; 0.52\n\nNot very stable 🙃",
    "2222190": "| CV | LB |\n| --- | --- |\n| 0.6007 | 0.59 |\n| 0.6024 | 0.6 |\n| 0.6143| 0.61 |\n| 0.6269| 0.63 |",
    "2222277": "wow, that looks stable! Do you mind sharing your CV strategy?",
    "2222626": "by ensemble 3 fold model",
    "2229847": "Hi @tattaka , could you please share more about your 5fold CV. I also tried to divided fragment2 into 3 parts but it did not work well. Thank you.",
    "2234512": "Here are some of my results. I split fragment 2 into two parts and used the middle 4 z-slices.\n\n| Val Fragment | F0.5 | Thresh | LB |\n| --- | --- | --- | --- |\n| 1 | 0.5626 | 0.81 | 0.51 |\n| 2a | 0.5673 | 0.81 | 0.63 |\n| 2b | 0.5129 | 0.76 | 0.44 |\n| 3 | 0.6197 | 0.91 | 0.36 |\n\nDon't know why there are such differences between the folds",
    "2236399": "CV - 0.58  Th - 0.5   LB - 0.6 - fragement 1\nCV - 0.57  Th - 0.45   LB - 0.57 - fragement 3\nCV - 0.56  Th - 0.5   LB - 0.53 - fragement 2",
    "2242886": "what is model architecture you are using, I was using linknet with resnet50 a backbone I got \n| image used for CV | CV|\n| --- | --- |\n| IMAGE 1 | 0.48 |\n| IMAGE 2|  0.38|\n| IMAGE 3 |  0.55|\n\ncurrent LB score is 0.39 for this model, I am not able to improve any further. do you have some tips for this?\nimage size using is 224",
    "2271138": "Can I ask you why your TH can be set so high and what kind of loss are you using?",
    "2271330": "So just to confirm, the CV reported is the result of the average scores of the individual folds/fragments, right ?"
  },
  "source": "meta"
}