{
  "id": 470385,
  "title": "Pretrained vs Train From Scratch",
  "url": "/competitions/blood-vessel-segmentation/discussion/470385",
  "author_name": "chemdatafarmer",
  "post_date": "2024-01-24T03:38:23.881000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>This has been a really interesting competition. I'm learning a lot from it and I thought I'd reach out to you all with a question I've been having.</p>\n<p>My models have been based on a U-Net I've expanded upon on from some tutorials that got me started. They started out as simple conv nets, but then I slowly added more complex architectures in such as residual blocks from ResNet, grouped convolutions similar to those used in ResNext, squeeze and excite add ons to the residual blocks, as well as a few others. I built these models from scratch because I wanted to get a better sense for how they worked, which has been fun so far.</p>\n<p>However, because I built these blocks from \"scratch\" using Conv2D and Conv2DTranspose, I have been training from scratch and I'm curious, has anyone here noticed pretrained weights have been helpful over training from zero? It's on my list of things to look at, but my focus right now is figuring out the best way to normalize or standardize the data, so I thought I'd ask to see what others have done.</p>",
  "messages": [
    {
      "id": 2617333,
      "postDate": "2024-01-24T07:19:45.627Z",
      "content": "<p>For this competition, I have read Kagglers using both pretrained and not-pretrained models high in the leaderboard.<br>\nSo both should work.<br>\nBut I surely guess pretrained models are dominating. </p>\n<p>On the other hand: Getting a not-pretrained model to work is a much more rigourous and delicate process: Longer training hours, more data requirements, and a harder process to manage such as 'overfitting'.</p>\n<p>So if you have time and resources I would say go for it, but of course if you have both, why not taking both approaches? </p>\n<p>But for sure, Getting a non-pretrained model high into the leaderboard, would show you know it :-)</p>",
      "rawMarkdown": "For this competition, I have read Kagglers using both pretrained and not-pretrained models high in the leaderboard.\nSo both should work.\nBut I surely guess pretrained models are dominating. \n\nOn the other hand: Getting a not-pretrained model to work is a much more rigourous and delicate process: Longer training hours, more data requirements, and a harder process to manage such as 'overfitting'.\n\nSo if you have time and resources I would say go for it, but of course if you have both, why not taking both approaches? \n\nBut for sure, Getting a non-pretrained model high into the leaderboard, would show you know it :-)",
      "replies": [
        {
          "id": 2618704,
          "postDate": "2024-01-24T23:32:51.213Z",
          "content": "<p>It's been fun so far! I may not have time to try both, which is why I was curious how others were doing with pretrained weights. My models don't do so well on the leaderboard so far and I don't have a good answer for how to manage the resolution differences we may observe, but I'm happy to say the score is pretty stable regardless of threshold so hopefully they generalize somewhat well.</p>",
          "rawMarkdown": "It's been fun so far! I may not have time to try both, which is why I was curious how others were doing with pretrained weights. My models don't do so well on the leaderboard so far and I don't have a good answer for how to manage the resolution differences we may observe, but I'm happy to say the score is pretty stable regardless of threshold so hopefully they generalize somewhat well."
        }
      ]
    },
    {
      "id": 2617041,
      "postDate": "2024-01-24T03:38:23.883Z",
      "content": "<p>Hi Everyone,</p>\n<p>This has been a really interesting competition. I'm learning a lot from it and I thought I'd reach out to you all with a question I've been having.</p>\n<p>My models have been based on a U-Net I've expanded upon on from some tutorials that got me started. They started out as simple conv nets, but then I slowly added more complex architectures in such as residual blocks from ResNet, grouped convolutions similar to those used in ResNext, squeeze and excite add ons to the residual blocks, as well as a few others. I built these models from scratch because I wanted to get a better sense for how they worked, which has been fun so far.</p>\n<p>However, because I built these blocks from \"scratch\" using Conv2D and Conv2DTranspose, I have been training from scratch and I'm curious, has anyone here noticed pretrained weights have been helpful over training from zero? It's on my list of things to look at, but my focus right now is figuring out the best way to normalize or standardize the data, so I thought I'd ask to see what others have done.</p>",
      "rawMarkdown": "Hi Everyone,\n\nThis has been a really interesting competition. I'm learning a lot from it and I thought I'd reach out to you all with a question I've been having.\n\nMy models have been based on a U-Net I've expanded upon on from some tutorials that got me started. They started out as simple conv nets, but then I slowly added more complex architectures in such as residual blocks from ResNet, grouped convolutions similar to those used in ResNext, squeeze and excite add ons to the residual blocks, as well as a few others. I built these models from scratch because I wanted to get a better sense for how they worked, which has been fun so far.\n\nHowever, because I built these blocks from \"scratch\" using Conv2D and Conv2DTranspose, I have been training from scratch and I'm curious, has anyone here noticed pretrained weights have been helpful over training from zero? It's on my list of things to look at, but my focus right now is figuring out the best way to normalize or standardize the data, so I thought I'd ask to see what others have done."
    },
    {
      "id": 2618479,
      "postDate": "2024-01-24T18:40:57.787Z",
      "content": "<p>Hi! I agree with kagglers below. Pretrained models can be useful if you want to train model with same architecture for your purpose. You wouldn't start training from random state and quality of non-fitted model would better than random-initialized one. <br>\nHowever some models can be trained from scratch for few minutes. As you can see, UNet-like models train very fast. Stable Diffusion uses this architecture as a denoising autoencoder because it's easy to train even if you have a small dataset. Long residual connections do their work perfectly. Your last output layer is connected with input directly with residual connection so weights for the last convolution block will trained fast. The rest conv/deconv blocks just refine predictions of the last deconvolution block. <br>\nDon't forget about batch normalization technique. It could be useful for speed up your training process. It really works, I checked.<br>\nGood luck in competition)</p>",
      "rawMarkdown": "Hi! I agree with kagglers below. Pretrained models can be useful if you want to train model with same architecture for your purpose. You wouldn't start training from random state and quality of non-fitted model would better than random-initialized one. \nHowever some models can be trained from scratch for few minutes. As you can see, UNet-like models train very fast. Stable Diffusion uses this architecture as a denoising autoencoder because it's easy to train even if you have a small dataset. Long residual connections do their work perfectly. Your last output layer is connected with input directly with residual connection so weights for the last convolution block will trained fast. The rest conv/deconv blocks just refine predictions of the last deconvolution block. \nDon't forget about batch normalization technique. It could be useful for speed up your training process. It really works, I checked.\nGood luck in competition)\n",
      "replies": [
        {
          "id": 2618701,
          "postDate": "2024-01-24T23:29:59.743Z",
          "content": "<p>Thanks :) When I incorporated the residual blocks, I added in batch norm layers as well to keep things consistent with the paper. I've also found them to be helpful, even when using relatively small batch sizes (training on a RTX 3060, so not quite as powerful as the Kaggle GPUs)</p>",
          "rawMarkdown": "Thanks :) When I incorporated the residual blocks, I added in batch norm layers as well to keep things consistent with the paper. I've also found them to be helpful, even when using relatively small batch sizes (training on a RTX 3060, so not quite as powerful as the Kaggle GPUs)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2618010,
      "postDate": "2024-01-24T14:33:34.293Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2617080,
      "postDate": "2024-01-24T04:04:14.193Z",
      "rawMarkdown": "",
      "votes": -5,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2617333,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2024-01-24T07:19:45.627000",
      "content": "<p>For this competition, I have read Kagglers using both pretrained and not-pretrained models high in the leaderboard.<br>\nSo both should work.<br>\nBut I surely guess pretrained models are dominating. </p>\n<p>On the other hand: Getting a not-pretrained model to work is a much more rigourous and delicate process: Longer training hours, more data requirements, and a harder process to manage such as 'overfitting'.</p>\n<p>So if you have time and resources I would say go for it, but of course if you have both, why not taking both approaches? </p>\n<p>But for sure, Getting a non-pretrained model high into the leaderboard, would show you know it :-)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2618704,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-01-24T23:32:51.213000",
          "content": "<p>It's been fun so far! I may not have time to try both, which is why I was curious how others were doing with pretrained weights. My models don't do so well on the leaderboard so far and I don't have a good answer for how to manage the resolution differences we may observe, but I'm happy to say the score is pretty stable regardless of threshold so hopefully they generalize somewhat well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2618479,
      "author_name": "Stanislav Baratov",
      "author_url": "",
      "post_date": "2024-01-24T18:40:57.787000",
      "content": "<p>Hi! I agree with kagglers below. Pretrained models can be useful if you want to train model with same architecture for your purpose. You wouldn't start training from random state and quality of non-fitted model would better than random-initialized one. <br>\nHowever some models can be trained from scratch for few minutes. As you can see, UNet-like models train very fast. Stable Diffusion uses this architecture as a denoising autoencoder because it's easy to train even if you have a small dataset. Long residual connections do their work perfectly. Your last output layer is connected with input directly with residual connection so weights for the last convolution block will trained fast. The rest conv/deconv blocks just refine predictions of the last deconvolution block. <br>\nDon't forget about batch normalization technique. It could be useful for speed up your training process. It really works, I checked.<br>\nGood luck in competition)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2618701,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-01-24T23:29:59.743000",
          "content": "<p>Thanks :) When I incorporated the residual blocks, I added in batch norm layers as well to keep things consistent with the paper. I've also found them to be helpful, even when using relatively small batch sizes (training on a RTX 3060, so not quite as powerful as the Kaggle GPUs)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2618010,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-24T14:33:34.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2617080,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-24T04:04:14.193000",
      "content": "",
      "votes": -5,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2617333": "For this competition, I have read Kagglers using both pretrained and not-pretrained models high in the leaderboard.\nSo both should work.\nBut I surely guess pretrained models are dominating. \n\nOn the other hand: Getting a not-pretrained model to work is a much more rigourous and delicate process: Longer training hours, more data requirements, and a harder process to manage such as 'overfitting'.\n\nSo if you have time and resources I would say go for it, but of course if you have both, why not taking both approaches? \n\nBut for sure, Getting a non-pretrained model high into the leaderboard, would show you know it :-)",
    "2617041": "Hi Everyone,\n\nThis has been a really interesting competition. I'm learning a lot from it and I thought I'd reach out to you all with a question I've been having.\n\nMy models have been based on a U-Net I've expanded upon on from some tutorials that got me started. They started out as simple conv nets, but then I slowly added more complex architectures in such as residual blocks from ResNet, grouped convolutions similar to those used in ResNext, squeeze and excite add ons to the residual blocks, as well as a few others. I built these models from scratch because I wanted to get a better sense for how they worked, which has been fun so far.\n\nHowever, because I built these blocks from \"scratch\" using Conv2D and Conv2DTranspose, I have been training from scratch and I'm curious, has anyone here noticed pretrained weights have been helpful over training from zero? It's on my list of things to look at, but my focus right now is figuring out the best way to normalize or standardize the data, so I thought I'd ask to see what others have done.",
    "2618479": "Hi! I agree with kagglers below. Pretrained models can be useful if you want to train model with same architecture for your purpose. You wouldn't start training from random state and quality of non-fitted model would better than random-initialized one. \nHowever some models can be trained from scratch for few minutes. As you can see, UNet-like models train very fast. Stable Diffusion uses this architecture as a denoising autoencoder because it's easy to train even if you have a small dataset. Long residual connections do their work perfectly. Your last output layer is connected with input directly with residual connection so weights for the last convolution block will trained fast. The rest conv/deconv blocks just refine predictions of the last deconvolution block. \nDon't forget about batch normalization technique. It could be useful for speed up your training process. It really works, I checked.\nGood luck in competition)\n",
    "2618010": "",
    "2617080": ""
  }
}