{
  "id": 119072,
  "title": "[private/public-lb]  0.669/0.670 : efficient-net-b4-unet 5-fold",
  "url": "/competitions/understanding_cloud_organization/discussion/119072",
  "author_name": "hengck23",
  "post_date": "2019-11-26T12:45:12.440000",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4653396b40e08568c7939467de7eb7b5%2FSelection_057.png?generation=1574772217029465&amp;alt=media\" alt=\"\"></p>\n\n<p>input = 544,352\nnet = efficient-net-b4-unet with classifier\nensemble: 4 to 5 fold, each with SWA for 5 snapshot</p>\n\n<p>i manage to train for 80 epoch for deeper network without much overfitting</p>\n\n<p>```\ntraining method:</p>\n\n<p>750 iter = 1 epoch</p>\n\n<ol>\n<li>pretrain with wadam (lr = 0.001) to 5k iter</li>\n<li>train with sdg (lr = 0.05) to 12k iter</li>\n<li>train with cyclic sdg (lr = 0.05 to 0.01) to 54.75k iter</li>\n<li>average network weights (SWA Stochastic Weight Averaging):\nuse snapshot at iter = 51.75, 52.50, 53.25, 54.00, 54.75\nno bn retraining  (the bn statistics for each snapshot is very close. bn retraining leads to poorer results)</li>\n</ol>\n\n<p>```</p>",
  "messages": [
    {
      "id": 681702,
      "postDate": "2019-11-26T12:45:12.440Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4653396b40e08568c7939467de7eb7b5%2FSelection_057.png?generation=1574772217029465&amp;alt=media\" alt=\"\"></p>\n\n<p>input = 544,352\nnet = efficient-net-b4-unet with classifier\nensemble: 4 to 5 fold, each with SWA for 5 snapshot</p>\n\n<p>i manage to train for 80 epoch for deeper network without much overfitting</p>\n\n<p>```\ntraining method:</p>\n\n<p>750 iter = 1 epoch</p>\n\n<ol>\n<li>pretrain with wadam (lr = 0.001) to 5k iter</li>\n<li>train with sdg (lr = 0.05) to 12k iter</li>\n<li>train with cyclic sdg (lr = 0.05 to 0.01) to 54.75k iter</li>\n<li>average network weights (SWA Stochastic Weight Averaging):\nuse snapshot at iter = 51.75, 52.50, 53.25, 54.00, 54.75\nno bn retraining  (the bn statistics for each snapshot is very close. bn retraining leads to poorer results)</li>\n</ol>\n\n<p>```</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4653396b40e08568c7939467de7eb7b5%2FSelection_057.png?generation=1574772217029465&amp;alt=media)\n\ninput = 544,352\nnet = efficient-net-b4-unet with classifier\nensemble: 4 to 5 fold, each with SWA for 5 snapshot\n\ni manage to train for 80 epoch for deeper network without much overfitting\n\n```\ntraining method:\n\n750 iter = 1 epoch\n \n1. pretrain with wadam (lr = 0.001) to 5k iter\n2. train with sdg (lr = 0.05) to 12k iter\n3. train with cyclic sdg (lr = 0.05 to 0.01) to 54.75k iter\n4. average network weights (SWA Stochastic Weight Averaging):\n    use snapshot at iter = 51.75, 52.50, 53.25, 54.00, 54.75\n    no bn retraining  (the bn statistics for each snapshot is very close. bn retraining leads to poorer results)\n\n```",
      "votes": 15
    },
    {
      "id": 681763,
      "postDate": "2019-11-26T14:17:59.837Z",
      "content": "<p>note:\n- increasing the input size (e.g. 576x384) seems to improve results.\n- adjusting the loss weight at the later training also improve results: e.g. loss =0.2*loss_label +0.8*loss_mask</p>",
      "rawMarkdown": "note:\n- increasing the input size (e.g. 576x384) seems to improve results.\n- adjusting the loss weight at the later training also improve results: e.g. loss =0.2\\*loss\\_label +0.8\\*loss\\_mask",
      "votes": 1,
      "replies": [
        {
          "id": 682347,
          "postDate": "2019-11-27T10:00:34.907Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffef0feb1b4f5f1de200c309f8a066b8a%2FSelection_067.png?generation=1574848832529132&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffef0feb1b4f5f1de200c309f8a066b8a%2FSelection_067.png?generation=1574848832529132&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 683204,
          "postDate": "2019-11-28T07:36:30.670Z",
          "content": "<p>i have probably set the pixel threshold to low. 0.40 gives better results\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5e06c09c54a5bf5c7349e1de544066e%2FSelection_078.png?generation=1574926588237580&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "i have probably set the pixel threshold to low. 0.40 gives better results\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5e06c09c54a5bf5c7349e1de544066e%2FSelection_078.png?generation=1574926588237580&amp;alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 681910,
      "postDate": "2019-11-26T17:08:02.933Z",
      "content": "<p>Hello Heng! Why you don't use full count of train iterations for one epoch?</p>",
      "rawMarkdown": "Hello Heng! Why you don't use full count of train iterations for one epoch?"
    },
    {
      "id": 681752,
      "postDate": "2019-11-26T14:06:21.790Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 682173,
      "postDate": "2019-11-27T03:12:30.597Z",
      "content": "<p>gold solution, thanks for sharing</p>",
      "rawMarkdown": "gold solution, thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 681763,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-26T14:17:59.837000",
      "content": "<p>note:\n- increasing the input size (e.g. 576x384) seems to improve results.\n- adjusting the loss weight at the later training also improve results: e.g. loss =0.2*loss_label +0.8*loss_mask</p>",
      "votes": 1,
      "replies": [
        {
          "id": 682347,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-27T10:00:34.907000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffef0feb1b4f5f1de200c309f8a066b8a%2FSelection_067.png?generation=1574848832529132&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 683204,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-28T07:36:30.670000",
          "content": "<p>i have probably set the pixel threshold to low. 0.40 gives better results\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5e06c09c54a5bf5c7349e1de544066e%2FSelection_078.png?generation=1574926588237580&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 681910,
      "author_name": "Rinat",
      "author_url": "",
      "post_date": "2019-11-26T17:08:02.933000",
      "content": "<p>Hello Heng! Why you don't use full count of train iterations for one epoch?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 681752,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-26T14:06:21.790000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 682173,
      "author_name": "liuze",
      "author_url": "",
      "post_date": "2019-11-27T03:12:30.597000",
      "content": "<p>gold solution, thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "681702": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4653396b40e08568c7939467de7eb7b5%2FSelection_057.png?generation=1574772217029465&amp;alt=media)\n\ninput = 544,352\nnet = efficient-net-b4-unet with classifier\nensemble: 4 to 5 fold, each with SWA for 5 snapshot\n\ni manage to train for 80 epoch for deeper network without much overfitting\n\n```\ntraining method:\n\n750 iter = 1 epoch\n \n1. pretrain with wadam (lr = 0.001) to 5k iter\n2. train with sdg (lr = 0.05) to 12k iter\n3. train with cyclic sdg (lr = 0.05 to 0.01) to 54.75k iter\n4. average network weights (SWA Stochastic Weight Averaging):\n    use snapshot at iter = 51.75, 52.50, 53.25, 54.00, 54.75\n    no bn retraining  (the bn statistics for each snapshot is very close. bn retraining leads to poorer results)\n\n```",
    "681763": "note:\n- increasing the input size (e.g. 576x384) seems to improve results.\n- adjusting the loss weight at the later training also improve results: e.g. loss =0.2\\*loss\\_label +0.8\\*loss\\_mask",
    "681910": "Hello Heng! Why you don't use full count of train iterations for one epoch?",
    "681752": "",
    "682173": "gold solution, thanks for sharing"
  }
}