{
  "id": 238442,
  "title": "pretrained models and normalization",
  "url": "/competitions/birdclef-2021/discussion/238442",
  "author_name": "",
  "post_date": "2021-05-12T07:39:34.100880Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear all,</p>\n<p>From <a href=\"https://pytorch.org/vision/stable/models.html:\" target=\"_blank\">https://pytorch.org/vision/stable/models.html:</a></p>\n<p>\"All pre-trained models expect input images normalized in the same way, i.e. mini-batches of 3-channel RGB images of shape (3 x H x W), where H and W are expected to be at least 224. The images have to be loaded in to a range of [0, 1] and then normalized using mean = [0.485, 0.456, 0.406] and std = [0.229, 0.224, 0.225]\"</p>\n<p>Does this normalization and sizing also hold for spectrogram images?</p>\n<p>Thank you very much.<br>\nKind regards,<br>\nKoen.</p>",
  "messages": [
    {
      "id": "1303683",
      "postDate": "05/12/2021 07:39:34",
      "content": "<p>Dear all,</p>\n<p>From <a href=\"https://pytorch.org/vision/stable/models.html:\" target=\"_blank\">https://pytorch.org/vision/stable/models.html:</a></p>\n<p>\"All pre-trained models expect input images normalized in the same way, i.e. mini-batches of 3-channel RGB images of shape (3 x H x W), where H and W are expected to be at least 224. The images have to be loaded in to a range of [0, 1] and then normalized using mean = [0.485, 0.456, 0.406] and std = [0.229, 0.224, 0.225]\"</p>\n<p>Does this normalization and sizing also hold for spectrogram images?</p>\n<p>Thank you very much.<br>\nKind regards,<br>\nKoen.</p>",
      "rawMarkdown": "Dear all,\n\nFrom https://pytorch.org/vision/stable/models.html:\n\n\"All pre-trained models expect input images normalized in the same way, i.e. mini-batches of 3-channel RGB images of shape (3 x H x W), where H and W are expected to be at least 224. The images have to be loaded in to a range of [0, 1] and then normalized using mean = [0.485, 0.456, 0.406] and std = [0.229, 0.224, 0.225]\"\n\nDoes this normalization and sizing also hold for spectrogram images?\n\nThank you very much.\nKind regards,\nKoen.",
      "votes": null
    },
    {
      "id": "1305785",
      "postDate": "05/13/2021 13:30:23",
      "content": "<p>Hi botkop, I have found that using a ResNet-50 pretrained model with imagenet normalization (the means/stds you posted) allows them to train faster (0.60+ F1 on train_soundscape_labels within 10 epochs or so) on this task than with unnormalized inputs.</p>",
      "rawMarkdown": "Hi botkop, I have found that using a ResNet-50 pretrained model with imagenet normalization (the means/stds you posted) allows them to train faster (0.60+ F1 on train_soundscape_labels within 10 epochs or so) on this task than with unnormalized inputs.",
      "votes": null
    },
    {
      "id": "1306512",
      "postDate": "05/13/2021 20:35:56",
      "content": "<p><a href=\"https://www.kaggle.com/brentspell\" target=\"_blank\">@brentspell</a> Most kind of you. Thank you.</p>",
      "rawMarkdown": "brentspell Most kind of you. Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1305785,
      "author_name": "brentspell",
      "author_url": "",
      "post_date": "05/13/2021 13:30:23",
      "content": "<p>Hi botkop, I have found that using a ResNet-50 pretrained model with imagenet normalization (the means/stds you posted) allows them to train faster (0.60+ F1 on train_soundscape_labels within 10 epochs or so) on this task than with unnormalized inputs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1306512,
      "author_name": "botkop",
      "author_url": "",
      "post_date": "05/13/2021 20:35:56",
      "content": "<p><a href=\"https://www.kaggle.com/brentspell\" target=\"_blank\">@brentspell</a> Most kind of you. Thank you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1303683": "Dear all,\n\nFrom https://pytorch.org/vision/stable/models.html:\n\n\"All pre-trained models expect input images normalized in the same way, i.e. mini-batches of 3-channel RGB images of shape (3 x H x W), where H and W are expected to be at least 224. The images have to be loaded in to a range of [0, 1] and then normalized using mean = [0.485, 0.456, 0.406] and std = [0.229, 0.224, 0.225]\"\n\nDoes this normalization and sizing also hold for spectrogram images?\n\nThank you very much.\nKind regards,\nKoen.",
    "1305785": "Hi botkop, I have found that using a ResNet-50 pretrained model with imagenet normalization (the means/stds you posted) allows them to train faster (0.60+ F1 on train_soundscape_labels within 10 epochs or so) on this task than with unnormalized inputs.",
    "1306512": "brentspell Most kind of you. Thank you."
  },
  "source": "meta"
}