{
  "id": 172975,
  "title": "XXL Average model experiment",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172975",
  "author_name": "Kirderf",
  "post_date": "2020-08-07T08:00:50.865000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Interesting to use TPU to the limit and build one large model of other large models.<br>\n I'll also concatenate GlobalAveragePooling2D and GlobalMaxPooling2D, and replace with SyncBatchNormalization in models architectures.<br>\nThen do the training with different image dimensions and finally average weights from some of the last checkpoints to create one final model. Experiment with the same model for all resolutions and also with different and instead ensemble them.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fce6b2b2a79d71d787fddb21b537680d3%2Fmodel3.png?generation=1596786840600209&amp;alt=media\" alt=\"\"></p>\n<p>I got this training result with the lowest 128x128 resolution with cdeotte's kernels without tuning LR:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fccd5006881e6da852fd27143c02f04e7%2F__results___16_1.png?generation=1596786749946002&amp;alt=media\" alt=\"\"></p>\n<p>And if I also add B7 with Imagenet to the other two and with SGD:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F47bd5a9d79e249473098a887d32ad180%2F__results___16_1%20(1).png?generation=1596786822821360&amp;alt=media\" alt=\"\"></p>\n<p>We will see how the rest goes.</p>\n<p>Have anyone tried similar with success?</p>",
  "messages": [
    {
      "id": 961479,
      "postDate": "2020-08-07T08:00:50.867Z",
      "content": "<p>Interesting to use TPU to the limit and build one large model of other large models.<br>\n I'll also concatenate GlobalAveragePooling2D and GlobalMaxPooling2D, and replace with SyncBatchNormalization in models architectures.<br>\nThen do the training with different image dimensions and finally average weights from some of the last checkpoints to create one final model. Experiment with the same model for all resolutions and also with different and instead ensemble them.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fce6b2b2a79d71d787fddb21b537680d3%2Fmodel3.png?generation=1596786840600209&amp;alt=media\" alt=\"\"></p>\n<p>I got this training result with the lowest 128x128 resolution with cdeotte's kernels without tuning LR:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fccd5006881e6da852fd27143c02f04e7%2F__results___16_1.png?generation=1596786749946002&amp;alt=media\" alt=\"\"></p>\n<p>And if I also add B7 with Imagenet to the other two and with SGD:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F47bd5a9d79e249473098a887d32ad180%2F__results___16_1%20(1).png?generation=1596786822821360&amp;alt=media\" alt=\"\"></p>\n<p>We will see how the rest goes.</p>\n<p>Have anyone tried similar with success?</p>",
      "rawMarkdown": "Interesting to use TPU to the limit and build one large model of other large models.\n I'll also concatenate GlobalAveragePooling2D and GlobalMaxPooling2D, and replace with SyncBatchNormalization in models architectures.\nThen do the training with different image dimensions and finally average weights from some of the last checkpoints to create one final model. Experiment with the same model for all resolutions and also with different and instead ensemble them.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fce6b2b2a79d71d787fddb21b537680d3%2Fmodel3.png?generation=1596786840600209&amp;alt=media)\n\nI got this training result with the lowest 128x128 resolution with cdeotte's kernels without tuning LR:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fccd5006881e6da852fd27143c02f04e7%2F__results___16_1.png?generation=1596786749946002&amp;alt=media)\n\nAnd if I also add B7 with Imagenet to the other two and with SGD:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F47bd5a9d79e249473098a887d32ad180%2F__results___16_1%20(1).png?generation=1596786822821360&amp;alt=media)\n\nWe will see how the rest goes.\n\n Have anyone tried similar with success?\n",
      "votes": 4
    },
    {
      "id": 961726,
      "postDate": "2020-08-07T13:02:40.853Z",
      "content": "<p>How do you average the weights of checkpoints? I also tried concatenating global avg/max pooling. I've notice a similar dip in the plot most likely due to training instability.</p>",
      "rawMarkdown": "How do you average the weights of checkpoints? I also tried concatenating global avg/max pooling. I've notice a similar dip in the plot most likely due to training instability.",
      "replies": [
        {
          "id": 961732,
          "postDate": "2020-08-07T13:08:40.083Z",
          "content": "<p>many ways, SWA is good way, add it to the training with some tuning or do it after with the same code/method. Always do it with the same fold/training not with different folds etc. And if you do it after, don't just pick the best epochs/ckp, risk for overfitting, learned all this the hard way ;)</p>",
          "rawMarkdown": "many ways, SWA is good way, add it to the training with some tuning or do it after with the same code/method. Always do it with the same fold/training not with different folds etc. And if you do it after, don't just pick the best epochs/ckp, risk for overfitting, learned all this the hard way ;)"
        }
      ]
    },
    {
      "id": 961702,
      "postDate": "2020-08-07T12:36:13.747Z",
      "content": "<p>be prepared with a side activity, it takes time to load into TPU memory initially. And with colab, which I think uses v2, it's no fun ;) better use smaller models with colab. I use this instead:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3f7198c304f5c1e9f7dbc15dc35b3be8%2Fd9ec0497-2262-413a-9350-602fe852f18b.png?generation=1596803565036799&amp;alt=media\" alt=\"\"></p>\n<p>Total params: 82,624,598<br>\nTrainable params: 82,210,902<br>\nNon-trainable params: 413,696</p>",
      "rawMarkdown": "be prepared with a side activity, it takes time to load into TPU memory initially. And with colab, which I think uses v2, it's no fun ;) better use smaller models with colab. I use this instead:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3f7198c304f5c1e9f7dbc15dc35b3be8%2Fd9ec0497-2262-413a-9350-602fe852f18b.png?generation=1596803565036799&amp;alt=media)\n\nTotal params: 82,624,598\nTrainable params: 82,210,902\nNon-trainable params: 413,696"
    },
    {
      "id": 961693,
      "postDate": "2020-08-07T12:23:17.350Z",
      "content": "<p>Sounds very promising, how about the others folds ? </p>",
      "rawMarkdown": "Sounds very promising, how about the others folds ? ",
      "replies": [
        {
          "id": 961700,
          "postDate": "2020-08-07T12:33:10.767Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 961703,
          "postDate": "2020-08-07T12:37:28.333Z",
          "content": "<p>Havn't got there yet, just experiment phase right now</p>",
          "rawMarkdown": "Havn't got there yet, just experiment phase right now"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 961726,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-08-07T13:02:40.853000",
      "content": "<p>How do you average the weights of checkpoints? I also tried concatenating global avg/max pooling. I've notice a similar dip in the plot most likely due to training instability.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 961732,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-08-07T13:08:40.083000",
          "content": "<p>many ways, SWA is good way, add it to the training with some tuning or do it after with the same code/method. Always do it with the same fold/training not with different folds etc. And if you do it after, don't just pick the best epochs/ckp, risk for overfitting, learned all this the hard way ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 961702,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-08-07T12:36:13.747000",
      "content": "<p>be prepared with a side activity, it takes time to load into TPU memory initially. And with colab, which I think uses v2, it's no fun ;) better use smaller models with colab. I use this instead:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3f7198c304f5c1e9f7dbc15dc35b3be8%2Fd9ec0497-2262-413a-9350-602fe852f18b.png?generation=1596803565036799&amp;alt=media\" alt=\"\"></p>\n<p>Total params: 82,624,598<br>\nTrainable params: 82,210,902<br>\nNon-trainable params: 413,696</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 961693,
      "author_name": "Seifeddine Fezzani",
      "author_url": "",
      "post_date": "2020-08-07T12:23:17.350000",
      "content": "<p>Sounds very promising, how about the others folds ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 961700,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-07T12:33:10.767000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961703,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-08-07T12:37:28.333000",
          "content": "<p>Havn't got there yet, just experiment phase right now</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "961479": "Interesting to use TPU to the limit and build one large model of other large models.\n I'll also concatenate GlobalAveragePooling2D and GlobalMaxPooling2D, and replace with SyncBatchNormalization in models architectures.\nThen do the training with different image dimensions and finally average weights from some of the last checkpoints to create one final model. Experiment with the same model for all resolutions and also with different and instead ensemble them.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fce6b2b2a79d71d787fddb21b537680d3%2Fmodel3.png?generation=1596786840600209&amp;alt=media)\n\nI got this training result with the lowest 128x128 resolution with cdeotte's kernels without tuning LR:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fccd5006881e6da852fd27143c02f04e7%2F__results___16_1.png?generation=1596786749946002&amp;alt=media)\n\nAnd if I also add B7 with Imagenet to the other two and with SGD:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F47bd5a9d79e249473098a887d32ad180%2F__results___16_1%20(1).png?generation=1596786822821360&amp;alt=media)\n\nWe will see how the rest goes.\n\n Have anyone tried similar with success?\n",
    "961726": "How do you average the weights of checkpoints? I also tried concatenating global avg/max pooling. I've notice a similar dip in the plot most likely due to training instability.",
    "961702": "be prepared with a side activity, it takes time to load into TPU memory initially. And with colab, which I think uses v2, it's no fun ;) better use smaller models with colab. I use this instead:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3f7198c304f5c1e9f7dbc15dc35b3be8%2Fd9ec0497-2262-413a-9350-602fe852f18b.png?generation=1596803565036799&amp;alt=media)\n\nTotal params: 82,624,598\nTrainable params: 82,210,902\nNon-trainable params: 413,696",
    "961693": "Sounds very promising, how about the others folds ? "
  }
}