{
  "id": 398329,
  "title": "Question about training and validation strategy",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/398329",
  "author_name": "",
  "post_date": "2023-03-29T13:34:30.496046300Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So there are three labeled samples in the training set.  Most probably a lot of submissions are splitting those images into smaller parts for training.</p>\n<p>My question is, what is your validation strategy? Are you training on two samples and validate on rest or do you take specific parts of those samples and use them for validation (meaning all three samples are used while training)?</p>\n<p>In my case, the best scores comes from using all samples for training but it has a huge risk of overfitting.</p>",
  "messages": [
    {
      "id": "2201675",
      "postDate": "03/29/2023 13:34:30",
      "content": "<p>So there are three labeled samples in the training set.  Most probably a lot of submissions are splitting those images into smaller parts for training.</p>\n<p>My question is, what is your validation strategy? Are you training on two samples and validate on rest or do you take specific parts of those samples and use them for validation (meaning all three samples are used while training)?</p>\n<p>In my case, the best scores comes from using all samples for training but it has a huge risk of overfitting.</p>",
      "rawMarkdown": "So there are three labeled samples in the training set.  Most probably a lot of submissions are splitting those images into smaller parts for training.\n\nMy question is, what is your validation strategy? Are you training on two samples and validate on rest or do you take specific parts of those samples and use them for validation (meaning all three samples are used while training)?\n\nIn my case, the best scores comes from using all samples for training but it has a huge risk of overfitting.",
      "votes": null
    },
    {
      "id": "2201953",
      "postDate": "03/29/2023 16:37:15",
      "content": "<p>As fragment 2 is much larger than 1 and 3, I split it into three pieces (see attachment) which lead to 5-fold CV. I think this way gives more balanced fold split than plain 1/2/3.</p>",
      "rawMarkdown": "As fragment 2 is much larger than 1 and 3, I split it into three pieces (see attachment) which lead to 5-fold CV. I think this way gives more balanced fold split than plain 1/2/3.",
      "votes": null
    },
    {
      "id": "2202059",
      "postDate": "03/29/2023 17:47:37",
      "content": "<p>great idea! Do you ensemble those five models for final prediction?</p>",
      "rawMarkdown": "great idea! Do you ensemble those five models for final prediction?",
      "votes": null
    },
    {
      "id": "2202070",
      "postDate": "03/29/2023 18:00:24",
      "content": "<p>Not yet. It is a bit early and I am trying different architects/tricks with single model experiment to save time. I kind of believe patch transformer is usable in this problem, but somehow my current experiments with that lowered the score. Once a best single model is decided I plan to do ensemble.</p>",
      "rawMarkdown": "Not yet. It is a bit early and I am trying different architects/tricks with single model experiment to save time. I kind of believe patch transformer is usable in this problem, but somehow my current experiments with that lowered the score. Once a best single model is decided I plan to do ensemble.",
      "votes": null
    },
    {
      "id": "2202083",
      "postDate": "03/29/2023 18:11:06",
      "content": "<p>ok, thanks for honest response! For me, 3D convolutions ( I tried Unet like architecture with 3DCNN encoder and 3DCNN classifier) seems to give lower scores than simple Unet. Thats probably because of not optimal training procedure as benchmark is reached with 3d conv classifier.</p>",
      "rawMarkdown": "ok, thanks for honest response! For me, 3D convolutions ( I tried Unet like architecture with 3DCNN encoder and 3DCNN classifier) seems to give lower scores than simple Unet. Thats probably because of not optimal training procedure as benchmark is reached with 3d conv classifier.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2201953,
      "author_name": "junxhuang",
      "author_url": "",
      "post_date": "03/29/2023 16:37:15",
      "content": "<p>As fragment 2 is much larger than 1 and 3, I split it into three pieces (see attachment) which lead to 5-fold CV. I think this way gives more balanced fold split than plain 1/2/3.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2202059,
          "author_name": "danieliusk",
          "author_url": "",
          "post_date": "03/29/2023 17:47:37",
          "content": "<p>great idea! Do you ensemble those five models for final prediction?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2202070,
              "author_name": "junxhuang",
              "author_url": "",
              "post_date": "03/29/2023 18:00:24",
              "content": "<p>Not yet. It is a bit early and I am trying different architects/tricks with single model experiment to save time. I kind of believe patch transformer is usable in this problem, but somehow my current experiments with that lowered the score. Once a best single model is decided I plan to do ensemble.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2202083,
                  "author_name": "danieliusk",
                  "author_url": "",
                  "post_date": "03/29/2023 18:11:06",
                  "content": "<p>ok, thanks for honest response! For me, 3D convolutions ( I tried Unet like architecture with 3DCNN encoder and 3DCNN classifier) seems to give lower scores than simple Unet. Thats probably because of not optimal training procedure as benchmark is reached with 3d conv classifier.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2201675": "So there are three labeled samples in the training set.  Most probably a lot of submissions are splitting those images into smaller parts for training.\n\nMy question is, what is your validation strategy? Are you training on two samples and validate on rest or do you take specific parts of those samples and use them for validation (meaning all three samples are used while training)?\n\nIn my case, the best scores comes from using all samples for training but it has a huge risk of overfitting.",
    "2201953": "As fragment 2 is much larger than 1 and 3, I split it into three pieces (see attachment) which lead to 5-fold CV. I think this way gives more balanced fold split than plain 1/2/3.",
    "2202059": "great idea! Do you ensemble those five models for final prediction?",
    "2202070": "Not yet. It is a bit early and I am trying different architects/tricks with single model experiment to save time. I kind of believe patch transformer is usable in this problem, but somehow my current experiments with that lowered the score. Once a best single model is decided I plan to do ensemble.",
    "2202083": "ok, thanks for honest response! For me, 3D convolutions ( I tried Unet like architecture with 3DCNN encoder and 3DCNN classifier) seems to give lower scores than simple Unet. Thats probably because of not optimal training procedure as benchmark is reached with 3d conv classifier."
  },
  "source": "meta"
}