{
  "id": 173410,
  "title": "Training multiple big models on TPU with metadata",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173410",
  "author_name": "",
  "post_date": "2020-08-09T06:28:39.058793200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Last weeks I saw the work of <a href=\"https://www.kaggle.com/agentauers\" target=\"_blank\">@agentauers</a> doing<a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\" target=\"_blank\"> Incredible TPU</a> notebook that lets you train multiple EF net models at one run using TPU. There are some questions that came to my mind following other posts.</p>\n<p>Does using blend of small models like B0 or B1 with big models b5, or b6 improve or decrease score? </p>\n<p>Is it better to use multiple bigger models like two b4 and two b5 trained on the image net and noisy student? </p>\n<p>What is the gain of adding dense layers with metadata to the top? </p>\n<p>Is center crop better than padding (padding might be better to capture changes that are on the sides of the images? </p>\n<p>I also prepared <a href=\"https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata\" target=\"_blank\">notebook</a> combining this approaches: <br>\nused models: B4 and B5 (4 models in total)<br>\nsmall size images 192px with padding<br>\nadded dense layers to the metadata. </p>\n<p>The result is good as for such small images (0.9233 LB) though I hoped using metadata would give better lift.</p>\n<p>What are results of your experiments in that area? <br>\nDoes training blend of the same models on image net and noisy student make sense? </p>",
  "messages": [
    {
      "id": "963578",
      "postDate": "08/09/2020 06:28:39",
      "content": "<p>Last weeks I saw the work of <a href=\"https://www.kaggle.com/agentauers\" target=\"_blank\">@agentauers</a> doing<a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\" target=\"_blank\"> Incredible TPU</a> notebook that lets you train multiple EF net models at one run using TPU. There are some questions that came to my mind following other posts.</p>\n<p>Does using blend of small models like B0 or B1 with big models b5, or b6 improve or decrease score? </p>\n<p>Is it better to use multiple bigger models like two b4 and two b5 trained on the image net and noisy student? </p>\n<p>What is the gain of adding dense layers with metadata to the top? </p>\n<p>Is center crop better than padding (padding might be better to capture changes that are on the sides of the images? </p>\n<p>I also prepared <a href=\"https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata\" target=\"_blank\">notebook</a> combining this approaches: <br>\nused models: B4 and B5 (4 models in total)<br>\nsmall size images 192px with padding<br>\nadded dense layers to the metadata. </p>\n<p>The result is good as for such small images (0.9233 LB) though I hoped using metadata would give better lift.</p>\n<p>What are results of your experiments in that area? <br>\nDoes training blend of the same models on image net and noisy student make sense? </p>",
      "rawMarkdown": "Last weeks I saw the work of @agentauers doing[ Incredible TPU](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once) notebook that lets you train multiple EF net models at one run using TPU. There are some questions that came to my mind following other posts.\n\nDoes using blend of small models like B0 or B1 with big models b5, or b6 improve or decrease score? \n\nIs it better to use multiple bigger models like two b4 and two b5 trained on the image net and noisy student? \n\nWhat is the gain of adding dense layers with metadata to the top? \n\nIs center crop better than padding (padding might be better to capture changes that are on the sides of the images? \n\nI also prepared [notebook](https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata) combining this approaches: \nused models: B4 and B5 (4 models in total)\nsmall size images 192px with padding\nadded dense layers to the metadata. \n\nThe result is good as for such small images (0.9233 LB) though I hoped using metadata would give better lift.\n\nWhat are results of your experiments in that area? \nDoes training blend of the same models on image net and noisy student make sense?",
      "votes": null
    },
    {
      "id": "963717",
      "postDate": "08/09/2020 08:24:26",
      "content": "<p>I have exactly the same question that you post here, and I am experiment on Incredible TPU notebook with different sizes of images and different combinations of effnet models. I'll try your notebook after I finish my one experiment.</p>",
      "rawMarkdown": "I have exactly the same question that you post here, and I am experiment on Incredible TPU notebook with different sizes of images and different combinations of effnet models. I'll try your notebook after I finish my one experiment.",
      "votes": null
    },
    {
      "id": "964052",
      "postDate": "08/09/2020 14:49:51",
      "content": "<p><a href=\"https://www.kaggle.com/guagugu\" target=\"_blank\">@guagugu</a> What are your results so far? What worked best for you? </p>",
      "rawMarkdown": "guagugu What are your results so far? What worked best for you?",
      "votes": null
    },
    {
      "id": "967714",
      "postDate": "08/12/2020 13:01:08",
      "content": "<p>got some progress on that. It seems that the models That I am training have very simmilar behaviour and are able to find similar patterns. </p>\n<p>If I train B4 imagenet and B4 noisy, the results that they achieve are very strongly correlated. <br>\nFor now I think it is no point on loosing time and training same model size on noisy and imgnet as there is not much gain in the score. </p>\n<p>Here is the loss of both:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F293395%2F9e55956d6d6216feb4e4cfac72f3f88a%2FSelection_069.png?generation=1597237247618313&amp;alt=media\" alt=\"loss for B$ image net and B4 noisy\"></p>",
      "rawMarkdown": "got some progress on that. It seems that the models That I am training have very simmilar behaviour and are able to find similar patterns. \n\nIf I train B4 imagenet and B4 noisy, the results that they achieve are very strongly correlated. \nFor now I think it is no point on loosing time and training same model size on noisy and imgnet as there is not much gain in the score. \n\nHere is the loss of both:\n\n![loss for B$ image net and B4 noisy](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F293395%2F9e55956d6d6216feb4e4cfac72f3f88a%2FSelection_069.png?generation=1597237247618313&amp;alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 963717,
      "author_name": "guagugu",
      "author_url": "",
      "post_date": "08/09/2020 08:24:26",
      "content": "<p>I have exactly the same question that you post here, and I am experiment on Incredible TPU notebook with different sizes of images and different combinations of effnet models. I'll try your notebook after I finish my one experiment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 964052,
          "author_name": "janidziak",
          "author_url": "",
          "post_date": "08/09/2020 14:49:51",
          "content": "<p><a href=\"https://www.kaggle.com/guagugu\" target=\"_blank\">@guagugu</a> What are your results so far? What worked best for you? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 967714,
      "author_name": "janidziak",
      "author_url": "",
      "post_date": "08/12/2020 13:01:08",
      "content": "<p>got some progress on that. It seems that the models That I am training have very simmilar behaviour and are able to find similar patterns. </p>\n<p>If I train B4 imagenet and B4 noisy, the results that they achieve are very strongly correlated. <br>\nFor now I think it is no point on loosing time and training same model size on noisy and imgnet as there is not much gain in the score. </p>\n<p>Here is the loss of both:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F293395%2F9e55956d6d6216feb4e4cfac72f3f88a%2FSelection_069.png?generation=1597237247618313&amp;alt=media\" alt=\"loss for B$ image net and B4 noisy\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "963578": "Last weeks I saw the work of @agentauers doing[ Incredible TPU](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once) notebook that lets you train multiple EF net models at one run using TPU. There are some questions that came to my mind following other posts.\n\nDoes using blend of small models like B0 or B1 with big models b5, or b6 improve or decrease score? \n\nIs it better to use multiple bigger models like two b4 and two b5 trained on the image net and noisy student? \n\nWhat is the gain of adding dense layers with metadata to the top? \n\nIs center crop better than padding (padding might be better to capture changes that are on the sides of the images? \n\nI also prepared [notebook](https://www.kaggle.com/janidziak/crazy-incredible-tpu-noisy-and-imgn-192-metadata) combining this approaches: \nused models: B4 and B5 (4 models in total)\nsmall size images 192px with padding\nadded dense layers to the metadata. \n\nThe result is good as for such small images (0.9233 LB) though I hoped using metadata would give better lift.\n\nWhat are results of your experiments in that area? \nDoes training blend of the same models on image net and noisy student make sense?",
    "963717": "I have exactly the same question that you post here, and I am experiment on Incredible TPU notebook with different sizes of images and different combinations of effnet models. I'll try your notebook after I finish my one experiment.",
    "964052": "guagugu What are your results so far? What worked best for you?",
    "967714": "got some progress on that. It seems that the models That I am training have very simmilar behaviour and are able to find similar patterns. \n\nIf I train B4 imagenet and B4 noisy, the results that they achieve are very strongly correlated. \nFor now I think it is no point on loosing time and training same model size on noisy and imgnet as there is not much gain in the score. \n\nHere is the loss of both:\n\n![loss for B$ image net and B4 noisy](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F293395%2F9e55956d6d6216feb4e4cfac72f3f88a%2FSelection_069.png?generation=1597237247618313&amp;alt=media)"
  },
  "source": "meta"
}