{
  "id": 399429,
  "title": "Overfitting with UNet",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/399429",
  "author_name": "",
  "post_date": "2023-04-04T04:46:26.633911Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am using standard Unet with timm-B0 from (<a href=\"https://github.com/qubvel/segmentation_models.pytorch)\" target=\"_blank\">https://github.com/qubvel/segmentation_models.pytorch)</a>. It overfits easily on 256x256 crops (train-F05 is much better than val-F05). <br>\nI am converting 3d to 2d by Conv3d, then average and max pooling from depth.<br>\nI added random x,y,z(depth) flips and 90-rotations in xy-plane. Also added dropouts of x,y,z, slices. Little improvement.</p>\n<p>Has anyone tried random 3d-resizing? Has anyone tried LSTM in depth instead of Conv3d?</p>",
  "messages": [
    {
      "id": "2208461",
      "postDate": "04/04/2023 04:46:26",
      "content": "<p>I am using standard Unet with timm-B0 from (<a href=\"https://github.com/qubvel/segmentation_models.pytorch)\" target=\"_blank\">https://github.com/qubvel/segmentation_models.pytorch)</a>. It overfits easily on 256x256 crops (train-F05 is much better than val-F05). <br>\nI am converting 3d to 2d by Conv3d, then average and max pooling from depth.<br>\nI added random x,y,z(depth) flips and 90-rotations in xy-plane. Also added dropouts of x,y,z, slices. Little improvement.</p>\n<p>Has anyone tried random 3d-resizing? Has anyone tried LSTM in depth instead of Conv3d?</p>",
      "rawMarkdown": "I am using standard Unet with timm-B0 from (https://github.com/qubvel/segmentation_models.pytorch). It overfits easily on 256x256 crops (train-F05 is much better than val-F05). \nI am converting 3d to 2d by Conv3d, then average and max pooling from depth.\nI added random x,y,z(depth) flips and 90-rotations in xy-plane. Also added dropouts of x,y,z, slices. Little improvement.\n\nHas anyone tried random 3d-resizing? Has anyone tried LSTM in depth instead of Conv3d?",
      "votes": null
    },
    {
      "id": "2208803",
      "postDate": "04/04/2023 10:03:56",
      "content": "<p>UNets should be able to accept 65 channels - so very early 3D-&gt;2D conversion may lead to information loss.  In my experements, UNets also overfit but a few suggestions that may help would be:</p>\n<ul>\n<li>more augmentations (this is sometimes tricky as we have more channels but some transformations such as  smaller rotations can be easily implemented)</li>\n<li>lowering the complexity of the model (lowering number of filters, etc.)</li>\n</ul>",
      "rawMarkdown": "UNets should be able to accept 65 channels - so very early 3D->2D conversion may lead to information loss.  In my experements, UNets also overfit but a few suggestions that may help would be:\n- more augmentations (this is sometimes tricky as we have more channels but some transformations such as  smaller rotations can be easily implemented)\n- lowering the complexity of the model (lowering number of filters, etc.)",
      "votes": null
    },
    {
      "id": "2208810",
      "postDate": "04/04/2023 10:06:19",
      "content": "<p>Interesting question… If I do not use bn in 3d conv, everything is overfitting when randomly resizing 3d. But if I use BN and random resizing, it gets val better. Does anyone see the same? Only 3 people are going to win some money in this, the rest of us could learn something useful.</p>",
      "rawMarkdown": "Interesting question... If I do not use bn in 3d conv, everything is overfitting when randomly resizing 3d. But if I use BN and random resizing, it gets val better. Does anyone see the same? Only 3 people are going to win some money in this, the rest of us could learn something useful.",
      "votes": null
    },
    {
      "id": "2208814",
      "postDate": "04/04/2023 10:10:08",
      "content": "<p>true true. But the location of the signal in depth is random. It looks like a filter-only problem. Baseline is claiming just to use 3d-conv 4 times.</p>",
      "rawMarkdown": "true true. But the location of the signal in depth is random. It looks like a filter-only problem. Baseline is claiming just to use 3d-conv 4 times.",
      "votes": null
    },
    {
      "id": "2255742",
      "postDate": "05/12/2023 00:55:06",
      "content": "<p>I've tried several variations of UNet, and I've been having the same problem. I've also tried recreating the tutorial single-pixel model - my training accuracy gets results comparable to what is reported in the tutorial, but my validation accuracy shows little if any real progress. Making the model less complex just results in lower training accuracy with no improvement in the gross overfitting. Also, my training data is pulled from all three fragments. I'm running out of things to try.</p>",
      "rawMarkdown": "I've tried several variations of UNet, and I've been having the same problem. I've also tried recreating the tutorial single-pixel model - my training accuracy gets results comparable to what is reported in the tutorial, but my validation accuracy shows little if any real progress. Making the model less complex just results in lower training accuracy with no improvement in the gross overfitting. Also, my training data is pulled from all three fragments. I'm running out of things to try.",
      "votes": null
    },
    {
      "id": "2256844",
      "postDate": "05/12/2023 19:08:17",
      "content": "<p>Overfitting usually means either actual pattern in data is lost during transformations or you model is too big and remembers all the cases individually instead of finding common patterns.</p>",
      "rawMarkdown": "Overfitting usually means either actual pattern in data is lost during transformations or you model is too big and remembers all the cases individually instead of finding common patterns.",
      "votes": null
    },
    {
      "id": "2256849",
      "postDate": "05/12/2023 19:10:49",
      "content": "<p>Smaller models have less overfitting. (Just because they don't have space to store all the individual information)</p>",
      "rawMarkdown": "Smaller models have less overfitting. (Just because they don't have space to store all the individual information)",
      "votes": null
    },
    {
      "id": "2256852",
      "postDate": "05/12/2023 19:14:07",
      "content": "<p>I would not rely to baseline too much as a measure. From what I saw it tends to learn patterns of where ink is not present way more then patterns of actual ink presence. With current top public score I'm still not convinced that there are actual ink patterns, rather then bunch of exclusion patterns based on cracks in papyrus, etc.</p>",
      "rawMarkdown": "I would not rely to baseline too much as a measure. From what I saw it tends to learn patterns of where ink is not present way more then patterns of actual ink presence. With current top public score I'm still not convinced that there are actual ink patterns, rather then bunch of exclusion patterns based on cracks in papyrus, etc.",
      "votes": null
    },
    {
      "id": "2258258",
      "postDate": "05/14/2023 03:41:36",
      "content": "<p>With the reduced resolution we're working with, the ink only occupies about three or four scan layers. I've been thinking that that simply might not be enough to get a good look at the transition from \"not ink\" to ink.</p>",
      "rawMarkdown": "With the reduced resolution we're working with, the ink only occupies about three or four scan layers. I've been thinking that that simply might not be enough to get a good look at the transition from \"not ink\" to ink.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2208803,
      "author_name": "danieliusk",
      "author_url": "",
      "post_date": "04/04/2023 10:03:56",
      "content": "<p>UNets should be able to accept 65 channels - so very early 3D-&gt;2D conversion may lead to information loss.  In my experements, UNets also overfit but a few suggestions that may help would be:</p>\n<ul>\n<li>more augmentations (this is sometimes tricky as we have more channels but some transformations such as  smaller rotations can be easily implemented)</li>\n<li>lowering the complexity of the model (lowering number of filters, etc.)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2208814,
          "author_name": "dmitrykonovalov",
          "author_url": "",
          "post_date": "04/04/2023 10:10:08",
          "content": "<p>true true. But the location of the signal in depth is random. It looks like a filter-only problem. Baseline is claiming just to use 3d-conv 4 times.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2256852,
              "author_name": "elvenmonk",
              "author_url": "",
              "post_date": "05/12/2023 19:14:07",
              "content": "<p>I would not rely to baseline too much as a measure. From what I saw it tends to learn patterns of where ink is not present way more then patterns of actual ink presence. With current top public score I'm still not convinced that there are actual ink patterns, rather then bunch of exclusion patterns based on cracks in papyrus, etc.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2258258,
                  "author_name": "matthewdehaven",
                  "author_url": "",
                  "post_date": "05/14/2023 03:41:36",
                  "content": "<p>With the reduced resolution we're working with, the ink only occupies about three or four scan layers. I've been thinking that that simply might not be enough to get a good look at the transition from \"not ink\" to ink.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2208810,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "04/04/2023 10:06:19",
      "content": "<p>Interesting question… If I do not use bn in 3d conv, everything is overfitting when randomly resizing 3d. But if I use BN and random resizing, it gets val better. Does anyone see the same? Only 3 people are going to win some money in this, the rest of us could learn something useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2255742,
      "author_name": "matthewdehaven",
      "author_url": "",
      "post_date": "05/12/2023 00:55:06",
      "content": "<p>I've tried several variations of UNet, and I've been having the same problem. I've also tried recreating the tutorial single-pixel model - my training accuracy gets results comparable to what is reported in the tutorial, but my validation accuracy shows little if any real progress. Making the model less complex just results in lower training accuracy with no improvement in the gross overfitting. Also, my training data is pulled from all three fragments. I'm running out of things to try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2256849,
          "author_name": "elvenmonk",
          "author_url": "",
          "post_date": "05/12/2023 19:10:49",
          "content": "<p>Smaller models have less overfitting. (Just because they don't have space to store all the individual information)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2256844,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "05/12/2023 19:08:17",
      "content": "<p>Overfitting usually means either actual pattern in data is lost during transformations or you model is too big and remembers all the cases individually instead of finding common patterns.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2208461": "I am using standard Unet with timm-B0 from (https://github.com/qubvel/segmentation_models.pytorch). It overfits easily on 256x256 crops (train-F05 is much better than val-F05). \nI am converting 3d to 2d by Conv3d, then average and max pooling from depth.\nI added random x,y,z(depth) flips and 90-rotations in xy-plane. Also added dropouts of x,y,z, slices. Little improvement.\n\nHas anyone tried random 3d-resizing? Has anyone tried LSTM in depth instead of Conv3d?",
    "2208803": "UNets should be able to accept 65 channels - so very early 3D->2D conversion may lead to information loss.  In my experements, UNets also overfit but a few suggestions that may help would be:\n- more augmentations (this is sometimes tricky as we have more channels but some transformations such as  smaller rotations can be easily implemented)\n- lowering the complexity of the model (lowering number of filters, etc.)",
    "2208810": "Interesting question... If I do not use bn in 3d conv, everything is overfitting when randomly resizing 3d. But if I use BN and random resizing, it gets val better. Does anyone see the same? Only 3 people are going to win some money in this, the rest of us could learn something useful.",
    "2208814": "true true. But the location of the signal in depth is random. It looks like a filter-only problem. Baseline is claiming just to use 3d-conv 4 times.",
    "2255742": "I've tried several variations of UNet, and I've been having the same problem. I've also tried recreating the tutorial single-pixel model - my training accuracy gets results comparable to what is reported in the tutorial, but my validation accuracy shows little if any real progress. Making the model less complex just results in lower training accuracy with no improvement in the gross overfitting. Also, my training data is pulled from all three fragments. I'm running out of things to try.",
    "2256844": "Overfitting usually means either actual pattern in data is lost during transformations or you model is too big and remembers all the cases individually instead of finding common patterns.",
    "2256849": "Smaller models have less overfitting. (Just because they don't have space to store all the individual information)",
    "2256852": "I would not rely to baseline too much as a measure. From what I saw it tends to learn patterns of where ink is not present way more then patterns of actual ink presence. With current top public score I'm still not convinced that there are actual ink patterns, rather then bunch of exclusion patterns based on cracks in papyrus, etc.",
    "2258258": "With the reduced resolution we're working with, the ink only occupies about three or four scan layers. I've been thinking that that simply might not be enough to get a good look at the transition from \"not ink\" to ink."
  },
  "source": "meta"
}