{
  "id": 403557,
  "title": "Problem with public score vs local validation score",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/403557",
  "author_name": "",
  "post_date": "2023-04-23T19:04:48.477180100Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>I was stuck for the past two weeks with a problem with my models, which looked like overfitting at first, but now I am not that sure anymore. The problem is a typical one: a good (at least not completely bad) local validation score, and an awful public score. </p>\n<p>I trained 4 different models:</p>\n<ul>\n<li>\"1\": model was trained with the entire volume 1 as validation data</li>\n<li>\"2\": model was trained with the entire volume 2 as validation data</li>\n<li>\"3\": model was trained with the entire volume 3 as validation data</li>\n<li>\"CUSTOM\": Validation data is build from 3 random patches taken from the middle of each of the 3 original volumes. The rest of the volumes were used as training data</li>\n</ul>\n<p>These are my current scores:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Score volume 1</th>\n<th>Score volume 2</th>\n<th>Score volume 3</th>\n<th>Leaderboard score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.38</td>\n<td>0.59</td>\n<td>0.74</td>\n<td>0.02</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.77</td>\n<td>0.24</td>\n<td>0.7</td>\n<td>0.12</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.78</td>\n<td>0.54</td>\n<td>0.36</td>\n<td>0.02</td>\n</tr>\n<tr>\n<td>CUSTOM</td>\n<td>0.64</td>\n<td>0.44</td>\n<td>0.7</td>\n<td>0.07</td>\n</tr>\n</tbody>\n</table>\n<p>The code to generate these submissions, together with the 4 models I trained with different validation sets, are available in <a href=\"https://www.kaggle.com/code/fpeccia/efficientnet-b0-unet-submission/notebook\" target=\"_blank\">this notebook</a>. I am worried this is some code bug that I am not seeing. If someone is interested in taking a look and have some comments, I would really appreciate it!</p>",
  "messages": [
    {
      "id": "2231875",
      "postDate": "04/23/2023 19:04:48",
      "content": "<p>Hi all!</p>\n<p>I was stuck for the past two weeks with a problem with my models, which looked like overfitting at first, but now I am not that sure anymore. The problem is a typical one: a good (at least not completely bad) local validation score, and an awful public score. </p>\n<p>I trained 4 different models:</p>\n<ul>\n<li>\"1\": model was trained with the entire volume 1 as validation data</li>\n<li>\"2\": model was trained with the entire volume 2 as validation data</li>\n<li>\"3\": model was trained with the entire volume 3 as validation data</li>\n<li>\"CUSTOM\": Validation data is build from 3 random patches taken from the middle of each of the 3 original volumes. The rest of the volumes were used as training data</li>\n</ul>\n<p>These are my current scores:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Score volume 1</th>\n<th>Score volume 2</th>\n<th>Score volume 3</th>\n<th>Leaderboard score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.38</td>\n<td>0.59</td>\n<td>0.74</td>\n<td>0.02</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.77</td>\n<td>0.24</td>\n<td>0.7</td>\n<td>0.12</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.78</td>\n<td>0.54</td>\n<td>0.36</td>\n<td>0.02</td>\n</tr>\n<tr>\n<td>CUSTOM</td>\n<td>0.64</td>\n<td>0.44</td>\n<td>0.7</td>\n<td>0.07</td>\n</tr>\n</tbody>\n</table>\n<p>The code to generate these submissions, together with the 4 models I trained with different validation sets, are available in <a href=\"https://www.kaggle.com/code/fpeccia/efficientnet-b0-unet-submission/notebook\" target=\"_blank\">this notebook</a>. I am worried this is some code bug that I am not seeing. If someone is interested in taking a look and have some comments, I would really appreciate it!</p>",
      "rawMarkdown": "Hi all!\n\nI was stuck for the past two weeks with a problem with my models, which looked like overfitting at first, but now I am not that sure anymore. The problem is a typical one: a good (at least not completely bad) local validation score, and an awful public score. \n\nI trained 4 different models:\n\n* \"1\": model was trained with the entire volume 1 as validation data\n* \"2\": model was trained with the entire volume 2 as validation data\n* \"3\": model was trained with the entire volume 3 as validation data\n* \"CUSTOM\": Validation data is build from 3 random patches taken from the middle of each of the 3 original volumes. The rest of the volumes were used as training data\n\nThese are my current scores:\n\n| Model | Score volume 1 | Score volume 2 | Score volume 3 | Leaderboard score | \n| --- | --- | --- | --- | --- |\n| 1 | 0.38 | 0.59 | 0.74 | 0.02 |\n| 2 | 0.77 | 0.24 | 0.7 | 0.12 |\n| 3 | 0.78 | 0.54 | 0.36 | 0.02 |\n| CUSTOM | 0.64 | 0.44 | 0.7 | 0.07 |\n\nThe code to generate these submissions, together with the 4 models I trained with different validation sets, are available in [this notebook](https://www.kaggle.com/code/fpeccia/efficientnet-b0-unet-submission/notebook). I am worried this is some code bug that I am not seeing. If someone is interested in taking a look and have some comments, I would really appreciate it!",
      "votes": null
    },
    {
      "id": "2232114",
      "postDate": "04/24/2023 03:19:31",
      "content": "<p>be aware because my lack of knowledge, which result with bad intuition: <br>\nthere is already a topic for that and some advise more folds<br>\ni would say bad generalization/overfitting for the difference between cv and public score<br>\nwrong approach (had similar results) for the public score, the depth seems very important</p>",
      "rawMarkdown": "be aware because my lack of knowledge, which result with bad intuition: \nthere is already a topic for that and some advise more folds\ni would say bad generalization/overfitting for the difference between cv and public score\nwrong approach (had similar results) for the public score, the depth seems very important",
      "votes": null
    },
    {
      "id": "2234209",
      "postDate": "04/25/2023 02:27:23",
      "content": "<p>In case you do not know. Submitting the given test masks will give you 0.11</p>",
      "rawMarkdown": "In case you do not know. Submitting the given test masks will give you 0.11",
      "votes": null
    },
    {
      "id": "2234450",
      "postDate": "04/25/2023 08:08:22",
      "content": "<p>I would say there is something wrong with your setup. I didn't look into your notebooks, but getting a local score of 0.38, 0.24 and 0.36 (the values from the diagonal in your matrix reflect the validation losses of the 3 folds, correct?) and a LB score of 0.02, 0.12 and 0.02 look strange.</p>\n<p>I'm currently having local scores of around 0.5 and getting LB scores of around 0.4</p>",
      "rawMarkdown": "I would say there is something wrong with your setup. I didn't look into your notebooks, but getting a local score of 0.38, 0.24 and 0.36 (the values from the diagonal in your matrix reflect the validation losses of the 3 folds, correct?) and a LB score of 0.02, 0.12 and 0.02 look strange.\n\nI'm currently having local scores of around 0.5 and getting LB scores of around 0.4",
      "votes": null
    },
    {
      "id": "2241928",
      "postDate": "05/02/2023 00:49:55",
      "content": "<p>I am providing the code I am using to train, the same I used to create the models of these tests, here: <a href=\"https://www.kaggle.com/fpeccia/efficientnet-b0-unet-train\" target=\"_blank\">https://www.kaggle.com/fpeccia/efficientnet-b0-unet-train</a></p>",
      "rawMarkdown": "I am providing the code I am using to train, the same I used to create the models of these tests, here: https://www.kaggle.com/fpeccia/efficientnet-b0-unet-train",
      "votes": null
    },
    {
      "id": "2242362",
      "postDate": "05/02/2023 07:59:39",
      "content": "<p>Federico, have you tried different model archs? Check it out this: <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/406038#2242263\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/406038#2242263</a><br>\nBy the way I also have trained the efficientnet arch (b3) on 3rd fold and the result was 0.0, then decreased threshold, which led to 0.21.</p>",
      "rawMarkdown": "Federico, have you tried different model archs? Check it out this: https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/406038#2242263\nBy the way I also have trained the efficientnet arch (b3) on 3rd fold and the result was 0.0, then decreased threshold, which led to 0.21.",
      "votes": null
    },
    {
      "id": "2272522",
      "postDate": "05/24/2023 14:50:57",
      "content": "<p>Have you found a solution to this? I'm having the same issue.</p>",
      "rawMarkdown": "Have you found a solution to this? I'm having the same issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2232114,
      "author_name": "iraqbot",
      "author_url": "",
      "post_date": "04/24/2023 03:19:31",
      "content": "<p>be aware because my lack of knowledge, which result with bad intuition: <br>\nthere is already a topic for that and some advise more folds<br>\ni would say bad generalization/overfitting for the difference between cv and public score<br>\nwrong approach (had similar results) for the public score, the depth seems very important</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234209,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "04/25/2023 02:27:23",
      "content": "<p>In case you do not know. Submitting the given test masks will give you 0.11</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234450,
      "author_name": "lucasvw",
      "author_url": "",
      "post_date": "04/25/2023 08:08:22",
      "content": "<p>I would say there is something wrong with your setup. I didn't look into your notebooks, but getting a local score of 0.38, 0.24 and 0.36 (the values from the diagonal in your matrix reflect the validation losses of the 3 folds, correct?) and a LB score of 0.02, 0.12 and 0.02 look strange.</p>\n<p>I'm currently having local scores of around 0.5 and getting LB scores of around 0.4</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2241928,
      "author_name": "fpeccia",
      "author_url": "",
      "post_date": "05/02/2023 00:49:55",
      "content": "<p>I am providing the code I am using to train, the same I used to create the models of these tests, here: <a href=\"https://www.kaggle.com/fpeccia/efficientnet-b0-unet-train\" target=\"_blank\">https://www.kaggle.com/fpeccia/efficientnet-b0-unet-train</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2242362,
          "author_name": "mknzfr",
          "author_url": "",
          "post_date": "05/02/2023 07:59:39",
          "content": "<p>Federico, have you tried different model archs? Check it out this: <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/406038#2242263\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/406038#2242263</a><br>\nBy the way I also have trained the efficientnet arch (b3) on 3rd fold and the result was 0.0, then decreased threshold, which led to 0.21.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2272522,
      "author_name": "jeffborack",
      "author_url": "",
      "post_date": "05/24/2023 14:50:57",
      "content": "<p>Have you found a solution to this? I'm having the same issue.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2231875": "Hi all!\n\nI was stuck for the past two weeks with a problem with my models, which looked like overfitting at first, but now I am not that sure anymore. The problem is a typical one: a good (at least not completely bad) local validation score, and an awful public score. \n\nI trained 4 different models:\n\n* \"1\": model was trained with the entire volume 1 as validation data\n* \"2\": model was trained with the entire volume 2 as validation data\n* \"3\": model was trained with the entire volume 3 as validation data\n* \"CUSTOM\": Validation data is build from 3 random patches taken from the middle of each of the 3 original volumes. The rest of the volumes were used as training data\n\nThese are my current scores:\n\n| Model | Score volume 1 | Score volume 2 | Score volume 3 | Leaderboard score | \n| --- | --- | --- | --- | --- |\n| 1 | 0.38 | 0.59 | 0.74 | 0.02 |\n| 2 | 0.77 | 0.24 | 0.7 | 0.12 |\n| 3 | 0.78 | 0.54 | 0.36 | 0.02 |\n| CUSTOM | 0.64 | 0.44 | 0.7 | 0.07 |\n\nThe code to generate these submissions, together with the 4 models I trained with different validation sets, are available in [this notebook](https://www.kaggle.com/code/fpeccia/efficientnet-b0-unet-submission/notebook). I am worried this is some code bug that I am not seeing. If someone is interested in taking a look and have some comments, I would really appreciate it!",
    "2232114": "be aware because my lack of knowledge, which result with bad intuition: \nthere is already a topic for that and some advise more folds\ni would say bad generalization/overfitting for the difference between cv and public score\nwrong approach (had similar results) for the public score, the depth seems very important",
    "2234209": "In case you do not know. Submitting the given test masks will give you 0.11",
    "2234450": "I would say there is something wrong with your setup. I didn't look into your notebooks, but getting a local score of 0.38, 0.24 and 0.36 (the values from the diagonal in your matrix reflect the validation losses of the 3 folds, correct?) and a LB score of 0.02, 0.12 and 0.02 look strange.\n\nI'm currently having local scores of around 0.5 and getting LB scores of around 0.4",
    "2241928": "I am providing the code I am using to train, the same I used to create the models of these tests, here: https://www.kaggle.com/fpeccia/efficientnet-b0-unet-train",
    "2242362": "Federico, have you tried different model archs? Check it out this: https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/406038#2242263\nBy the way I also have trained the efficientnet arch (b3) on 3rd fold and the result was 0.0, then decreased threshold, which led to 0.21.",
    "2272522": "Have you found a solution to this? I'm having the same issue."
  },
  "source": "meta"
}