{
  "id": 412133,
  "title": "The right validation procedure and slpit",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/412133",
  "author_name": "",
  "post_date": "2023-05-22T11:27:17.927156700Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n<p>I used two types of validation procedures. In \"fold 5\", I split all three train images into five parts and used an 80/20 train/validation split [1]. I calculated the CV by averaging the slice scores. For \"fold 6\", I split the images into two slices [2]. I saw that the public is approximately 10% of the evaluation dataset. I have a large gap between my local CV and the public LB and a better CV/LB ratio than I. Do you know what I should change?</p>\n<p>Thanks for the help!</p>\n<table>\n<thead>\n<tr>\n<th>Step size</th>\n<th>Local CV</th>\n<th>Fold type</th>\n<th>TTA</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>stride / 4</td>\n<td>0.62095</td>\n<td>Fold 6</td>\n<td>yes</td>\n<td>0.37</td>\n</tr>\n<tr>\n<td>stride / 4</td>\n<td>0.605</td>\n<td>Fold 5</td>\n<td>no</td>\n<td>0.39</td>\n</tr>\n<tr>\n<td>stride / 4</td>\n<td>0.6559</td>\n<td>Fold 5</td>\n<td>yes</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>stride / 4</td>\n<td>0.6742</td>\n<td>Fold 5</td>\n<td>yes</td>\n<td>0.54</td>\n</tr>\n</tbody>\n</table>\n<p>[1]</p>\n<table>\n<thead>\n<tr>\n<th>path</th>\n<th>fold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2_slice_01_01</td>\n<td>3</td>\n</tr>\n<tr>\n<td>2_slice_02_01</td>\n<td>0</td>\n</tr>\n<tr>\n<td>2_slice_03_01</td>\n<td>4</td>\n</tr>\n<tr>\n<td>2_slice_04_01</td>\n<td>2</td>\n</tr>\n<tr>\n<td>2_slice_05_01</td>\n<td>1</td>\n</tr>\n<tr>\n<td>3_slice_01_01</td>\n<td>1</td>\n</tr>\n<tr>\n<td>3_slice_02_01</td>\n<td>2</td>\n</tr>\n<tr>\n<td>3_slice_03_01</td>\n<td>4</td>\n</tr>\n<tr>\n<td>3_slice_04_01</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3_slice_05_01</td>\n<td>3</td>\n</tr>\n<tr>\n<td>1_slice_01_01</td>\n<td>4</td>\n</tr>\n<tr>\n<td>1_slice_02_01</td>\n<td>3</td>\n</tr>\n<tr>\n<td>1_slice_03_01</td>\n<td>1</td>\n</tr>\n<tr>\n<td>1_slice_04_01</td>\n<td>2</td>\n</tr>\n<tr>\n<td>1_slice_05_01</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>[2]</p>\n<table>\n<thead>\n<tr>\n<th>path</th>\n<th>fold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1_slice_01_01</td>\n<td>valid_0</td>\n</tr>\n<tr>\n<td>1_slice_02_01</td>\n<td>train_0</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_0</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_0</td>\n</tr>\n<tr>\n<td>2_slice_01_01</td>\n<td>valid_1</td>\n</tr>\n<tr>\n<td>2_slice_02_01</td>\n<td>train_1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_1</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_1</td>\n</tr>\n<tr>\n<td>3_slice_01_01</td>\n<td>valid_2</td>\n</tr>\n<tr>\n<td>3_slice_02_01</td>\n<td>train_2</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_2</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_2</td>\n</tr>\n<tr>\n<td>1_slice_02_01</td>\n<td>valid_3</td>\n</tr>\n<tr>\n<td>1_slice_01_01</td>\n<td>train_3</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_3</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_3</td>\n</tr>\n<tr>\n<td>2_slice_02_01</td>\n<td>valid_4</td>\n</tr>\n<tr>\n<td>2_slice_01_01</td>\n<td>train_4</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_4</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_4</td>\n</tr>\n<tr>\n<td>3_slice_02_01</td>\n<td>valid_5</td>\n</tr>\n<tr>\n<td>3_slice_01_01</td>\n<td>train_5</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_5</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_5</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2269311",
      "postDate": "05/22/2023 11:27:17",
      "content": "<p>Hi everyone!</p>\n<p>I used two types of validation procedures. In \"fold 5\", I split all three train images into five parts and used an 80/20 train/validation split [1]. I calculated the CV by averaging the slice scores. For \"fold 6\", I split the images into two slices [2]. I saw that the public is approximately 10% of the evaluation dataset. I have a large gap between my local CV and the public LB and a better CV/LB ratio than I. Do you know what I should change?</p>\n<p>Thanks for the help!</p>\n<table>\n<thead>\n<tr>\n<th>Step size</th>\n<th>Local CV</th>\n<th>Fold type</th>\n<th>TTA</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>stride / 4</td>\n<td>0.62095</td>\n<td>Fold 6</td>\n<td>yes</td>\n<td>0.37</td>\n</tr>\n<tr>\n<td>stride / 4</td>\n<td>0.605</td>\n<td>Fold 5</td>\n<td>no</td>\n<td>0.39</td>\n</tr>\n<tr>\n<td>stride / 4</td>\n<td>0.6559</td>\n<td>Fold 5</td>\n<td>yes</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>stride / 4</td>\n<td>0.6742</td>\n<td>Fold 5</td>\n<td>yes</td>\n<td>0.54</td>\n</tr>\n</tbody>\n</table>\n<p>[1]</p>\n<table>\n<thead>\n<tr>\n<th>path</th>\n<th>fold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2_slice_01_01</td>\n<td>3</td>\n</tr>\n<tr>\n<td>2_slice_02_01</td>\n<td>0</td>\n</tr>\n<tr>\n<td>2_slice_03_01</td>\n<td>4</td>\n</tr>\n<tr>\n<td>2_slice_04_01</td>\n<td>2</td>\n</tr>\n<tr>\n<td>2_slice_05_01</td>\n<td>1</td>\n</tr>\n<tr>\n<td>3_slice_01_01</td>\n<td>1</td>\n</tr>\n<tr>\n<td>3_slice_02_01</td>\n<td>2</td>\n</tr>\n<tr>\n<td>3_slice_03_01</td>\n<td>4</td>\n</tr>\n<tr>\n<td>3_slice_04_01</td>\n<td>0</td>\n</tr>\n<tr>\n<td>3_slice_05_01</td>\n<td>3</td>\n</tr>\n<tr>\n<td>1_slice_01_01</td>\n<td>4</td>\n</tr>\n<tr>\n<td>1_slice_02_01</td>\n<td>3</td>\n</tr>\n<tr>\n<td>1_slice_03_01</td>\n<td>1</td>\n</tr>\n<tr>\n<td>1_slice_04_01</td>\n<td>2</td>\n</tr>\n<tr>\n<td>1_slice_05_01</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>[2]</p>\n<table>\n<thead>\n<tr>\n<th>path</th>\n<th>fold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1_slice_01_01</td>\n<td>valid_0</td>\n</tr>\n<tr>\n<td>1_slice_02_01</td>\n<td>train_0</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_0</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_0</td>\n</tr>\n<tr>\n<td>2_slice_01_01</td>\n<td>valid_1</td>\n</tr>\n<tr>\n<td>2_slice_02_01</td>\n<td>train_1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_1</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_1</td>\n</tr>\n<tr>\n<td>3_slice_01_01</td>\n<td>valid_2</td>\n</tr>\n<tr>\n<td>3_slice_02_01</td>\n<td>train_2</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_2</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_2</td>\n</tr>\n<tr>\n<td>1_slice_02_01</td>\n<td>valid_3</td>\n</tr>\n<tr>\n<td>1_slice_01_01</td>\n<td>train_3</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_3</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_3</td>\n</tr>\n<tr>\n<td>2_slice_02_01</td>\n<td>valid_4</td>\n</tr>\n<tr>\n<td>2_slice_01_01</td>\n<td>train_4</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_4</td>\n</tr>\n<tr>\n<td>3</td>\n<td>train_4</td>\n</tr>\n<tr>\n<td>3_slice_02_01</td>\n<td>valid_5</td>\n</tr>\n<tr>\n<td>3_slice_01_01</td>\n<td>train_5</td>\n</tr>\n<tr>\n<td>2</td>\n<td>train_5</td>\n</tr>\n<tr>\n<td>1</td>\n<td>train_5</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Hi everyone!\n\nI used two types of validation procedures. In \"fold 5\", I split all three train images into five parts and used an 80/20 train/validation split [1]. I calculated the CV by averaging the slice scores. For \"fold 6\", I split the images into two slices [2]. I saw that the public is approximately 10% of the evaluation dataset. I have a large gap between my local CV and the public LB and a better CV/LB ratio than I. Do you know what I should change?\n\nThanks for the help!\n\n\n| Step size  | Local CV | Fold type | TTA | LB   |\n| ---------- | -------- | --------- | --- | ---- |\n| stride / 4 | 0.62095  | Fold 6    | yes | 0.37 |\n| stride / 4 | 0.605    | Fold 5    | no  | 0.39 |\n| stride / 4 | 0.6559   | Fold 5    | yes | 0.48 |\n| stride / 4 | 0.6742   | Fold 5    | yes | 0.54 |\n\n\n[1]\n| path |    fold|\n| --- | --- |\n| 2_slice_01_01 |    3|\n| 2_slice_02_01 |    0|\n| 2_slice_03_01 |    4|\n| 2_slice_04_01 |    2|\n| 2_slice_05_01 |    1|\n| 3_slice_01_01 |    1|\n| 3_slice_02_01 |    2|\n| 3_slice_03_01 |    4|\n| 3_slice_04_01 |    0|\n| 3_slice_05_01 |    3|\n| 1_slice_01_01 |    4|\n| 1_slice_02_01 |    3|\n| 1_slice_03_01 |    1|\n| 1_slice_04_01 |    2|\n| 1_slice_05_01 |    0|\n\n[2]\n|path | fold|\n| --- | --- |\n|1_slice_01_01 | valid_0|\n|1_slice_02_01 | train_0|\n|2 | train_0|\n|3 | train_0|\n|2_slice_01_01 | valid_1|\n|2_slice_02_01 | train_1|\n|1 | train_1|\n|3 | train_1|\n|3_slice_01_01 | valid_2|\n|3_slice_02_01 | train_2|\n|2 | train_2|\n|1 | train_2|\n|1_slice_02_01 | valid_3|\n|1_slice_01_01 | train_3|\n|2 | train_3|\n|3 | train_3|\n|2_slice_02_01 | valid_4|\n|2_slice_01_01 | train_4|\n|1 | train_4|\n|3 | train_4|\n|3_slice_02_01 | valid_5|\n|3_slice_01_01 | train_5|\n|2 | train_5|\n|1 | train_5|",
      "votes": null
    },
    {
      "id": "2270378",
      "postDate": "05/23/2023 06:35:41",
      "content": "<p>Maybe the threshold needs to be modified during infer?</p>",
      "rawMarkdown": "Maybe the threshold needs to be modified during infer?",
      "votes": null
    },
    {
      "id": "2270553",
      "postDate": "05/23/2023 08:20:25",
      "content": "<p>Optimized the threshold for the best CV with bayesian optimization, so I used that threshold that has the best CV</p>",
      "rawMarkdown": "Optimized the threshold for the best CV with bayesian optimization, so I used that threshold that has the best CV",
      "votes": null
    },
    {
      "id": "2271412",
      "postDate": "05/23/2023 21:31:33",
      "content": "<p>I used to encounter this performance drop as well, here are a couple of things worth checking</p>\n<ul>\n<li>did you leak data? did you normalise the training images using the validation set images as well? in general do the val split and exclude the validation set from all preprocessing activities</li>\n<li>did you perform rle correctly? I ran into this issue when my LB doesn't match my local CV. Try saving out the rle predictions of the model and write an evaluation script to do local CV score from the saved rle. I found my bug this way so maybe it helps</li>\n</ul>",
      "rawMarkdown": "I used to encounter this performance drop as well, here are a couple of things worth checking\n- did you leak data? did you normalise the training images using the validation set images as well? in general do the val split and exclude the validation set from all preprocessing activities\n- did you perform rle correctly? I ran into this issue when my LB doesn't match my local CV. Try saving out the rle predictions of the model and write an evaluation script to do local CV score from the saved rle. I found my bug this way so maybe it helps",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2270378,
      "author_name": "xiaoqinglong1996",
      "author_url": "",
      "post_date": "05/23/2023 06:35:41",
      "content": "<p>Maybe the threshold needs to be modified during infer?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2270553,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "05/23/2023 08:20:25",
          "content": "<p>Optimized the threshold for the best CV with bayesian optimization, so I used that threshold that has the best CV</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2271412,
      "author_name": "sirapoabchaikunsaeng",
      "author_url": "",
      "post_date": "05/23/2023 21:31:33",
      "content": "<p>I used to encounter this performance drop as well, here are a couple of things worth checking</p>\n<ul>\n<li>did you leak data? did you normalise the training images using the validation set images as well? in general do the val split and exclude the validation set from all preprocessing activities</li>\n<li>did you perform rle correctly? I ran into this issue when my LB doesn't match my local CV. Try saving out the rle predictions of the model and write an evaluation script to do local CV score from the saved rle. I found my bug this way so maybe it helps</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2269311": "Hi everyone!\n\nI used two types of validation procedures. In \"fold 5\", I split all three train images into five parts and used an 80/20 train/validation split [1]. I calculated the CV by averaging the slice scores. For \"fold 6\", I split the images into two slices [2]. I saw that the public is approximately 10% of the evaluation dataset. I have a large gap between my local CV and the public LB and a better CV/LB ratio than I. Do you know what I should change?\n\nThanks for the help!\n\n\n| Step size  | Local CV | Fold type | TTA | LB   |\n| ---------- | -------- | --------- | --- | ---- |\n| stride / 4 | 0.62095  | Fold 6    | yes | 0.37 |\n| stride / 4 | 0.605    | Fold 5    | no  | 0.39 |\n| stride / 4 | 0.6559   | Fold 5    | yes | 0.48 |\n| stride / 4 | 0.6742   | Fold 5    | yes | 0.54 |\n\n\n[1]\n| path |    fold|\n| --- | --- |\n| 2_slice_01_01 |    3|\n| 2_slice_02_01 |    0|\n| 2_slice_03_01 |    4|\n| 2_slice_04_01 |    2|\n| 2_slice_05_01 |    1|\n| 3_slice_01_01 |    1|\n| 3_slice_02_01 |    2|\n| 3_slice_03_01 |    4|\n| 3_slice_04_01 |    0|\n| 3_slice_05_01 |    3|\n| 1_slice_01_01 |    4|\n| 1_slice_02_01 |    3|\n| 1_slice_03_01 |    1|\n| 1_slice_04_01 |    2|\n| 1_slice_05_01 |    0|\n\n[2]\n|path | fold|\n| --- | --- |\n|1_slice_01_01 | valid_0|\n|1_slice_02_01 | train_0|\n|2 | train_0|\n|3 | train_0|\n|2_slice_01_01 | valid_1|\n|2_slice_02_01 | train_1|\n|1 | train_1|\n|3 | train_1|\n|3_slice_01_01 | valid_2|\n|3_slice_02_01 | train_2|\n|2 | train_2|\n|1 | train_2|\n|1_slice_02_01 | valid_3|\n|1_slice_01_01 | train_3|\n|2 | train_3|\n|3 | train_3|\n|2_slice_02_01 | valid_4|\n|2_slice_01_01 | train_4|\n|1 | train_4|\n|3 | train_4|\n|3_slice_02_01 | valid_5|\n|3_slice_01_01 | train_5|\n|2 | train_5|\n|1 | train_5|",
    "2270378": "Maybe the threshold needs to be modified during infer?",
    "2270553": "Optimized the threshold for the best CV with bayesian optimization, so I used that threshold that has the best CV",
    "2271412": "I used to encounter this performance drop as well, here are a couple of things worth checking\n- did you leak data? did you normalise the training images using the validation set images as well? in general do the val split and exclude the validation set from all preprocessing activities\n- did you perform rle correctly? I ran into this issue when my LB doesn't match my local CV. Try saving out the rle predictions of the model and write an evaluation script to do local CV score from the saved rle. I found my bug this way so maybe it helps"
  },
  "source": "meta"
}