{
  "id": 89640,
  "title": "Approach pre-trained deep learning models with caution",
  "url": "/competitions/imet-2019-fgvc6/discussion/89640",
  "author_name": "",
  "post_date": "2019-04-16T11:59:32.937540300Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As a newbie, I wanna share with you this nice article for reusing any pre-trained deep learning models.</p>\n\n<p>Source: <a href=\"https://medium.com/comet-ml/approach-pre-trained-deep-learning-models-with-caution-9f0ff739010c\">https://medium.com/comet-ml/approach-pre-trained-deep-learning-models-with-caution-9f0ff739010c</a></p>\n\n<p>Interesting points:\n1. resnet architectures perform better in PyTorch and inception architectures perform better in Keras \n2. The published benchmarks on Keras Applications cannot be reproduced, even when exactly copying the example code. In fact, their reported accuracies (as of Feb. 2019) are usually higher than the actual accuracies \n3. Some pre-trained Keras models yield inconsistent or lower accuracies when deployed on a server or run in sequence with other Keras models \n4. Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time. </p>\n\n<p>Hope that it is helpful to you.</p>",
  "messages": [
    {
      "id": "517704",
      "postDate": "04/16/2019 11:59:32",
      "content": "<p>As a newbie, I wanna share with you this nice article for reusing any pre-trained deep learning models.</p>\n\n<p>Source: <a href=\"https://medium.com/comet-ml/approach-pre-trained-deep-learning-models-with-caution-9f0ff739010c\">https://medium.com/comet-ml/approach-pre-trained-deep-learning-models-with-caution-9f0ff739010c</a></p>\n\n<p>Interesting points:\n1. resnet architectures perform better in PyTorch and inception architectures perform better in Keras \n2. The published benchmarks on Keras Applications cannot be reproduced, even when exactly copying the example code. In fact, their reported accuracies (as of Feb. 2019) are usually higher than the actual accuracies \n3. Some pre-trained Keras models yield inconsistent or lower accuracies when deployed on a server or run in sequence with other Keras models \n4. Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time. </p>\n\n<p>Hope that it is helpful to you.</p>",
      "rawMarkdown": "As a newbie, I wanna share with you this nice article for reusing any pre-trained deep learning models.\n\nSource: https://medium.com/comet-ml/approach-pre-trained-deep-learning-models-with-caution-9f0ff739010c\n\nInteresting points:\n1. resnet architectures perform better in PyTorch and inception architectures perform better in Keras \n2. The published benchmarks on Keras Applications cannot be reproduced, even when exactly copying the example code. In fact, their reported accuracies (as of Feb. 2019) are usually higher than the actual accuracies \n3. Some pre-trained Keras models yield inconsistent or lower accuracies when deployed on a server or run in sequence with other Keras models \n4. Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time. \n\nHope that it is helpful to you.",
      "votes": null
    },
    {
      "id": "518209",
      "postDate": "04/17/2019 01:20:01",
      "content": "<blockquote>\n  <ol>\n  <li>Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time.</li>\n  </ol>\n</blockquote>\n\n<p>Try the sequential API with keras when using models with BN layers... It works fine with me but be careful at some point you will be suffering from exploding gradient. I didn't find any explanation for that but it only works with a limited number of epochs (between 20-30).\nWhen using the Model API the BN layer will be using the mean w variance estimated from imagenet in the test mode. That's why the train loss will be much lower than the validation loss.\nI switech to pytorch with fastai for this reason.</p>",
      "rawMarkdown": "&gt; 4.  Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time.\n\nTry the sequential API with keras when using models with BN layers... It works fine with me but be careful at some point you will be suffering from exploding gradient. I didn't find any explanation for that but it only works with a limited number of epochs (between 20-30).\nWhen using the Model API the BN layer will be using the mean w variance estimated from imagenet in the test mode. That's why the train loss will be much lower than the validation loss.\nI switech to pytorch with fastai for this reason.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 518209,
      "author_name": "rinnqd",
      "author_url": "",
      "post_date": "04/17/2019 01:20:01",
      "content": "<blockquote>\n  <ol>\n  <li>Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time.</li>\n  </ol>\n</blockquote>\n\n<p>Try the sequential API with keras when using models with BN layers... It works fine with me but be careful at some point you will be suffering from exploding gradient. I didn't find any explanation for that but it only works with a limited number of epochs (between 20-30).\nWhen using the Model API the BN layer will be using the mean w variance estimated from imagenet in the test mode. That's why the train loss will be much lower than the validation loss.\nI switech to pytorch with fastai for this reason.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "517704": "As a newbie, I wanna share with you this nice article for reusing any pre-trained deep learning models.\n\nSource: https://medium.com/comet-ml/approach-pre-trained-deep-learning-models-with-caution-9f0ff739010c\n\nInteresting points:\n1. resnet architectures perform better in PyTorch and inception architectures perform better in Keras \n2. The published benchmarks on Keras Applications cannot be reproduced, even when exactly copying the example code. In fact, their reported accuracies (as of Feb. 2019) are usually higher than the actual accuracies \n3. Some pre-trained Keras models yield inconsistent or lower accuracies when deployed on a server or run in sequence with other Keras models \n4. Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time. \n\nHope that it is helpful to you.",
    "518209": "&gt; 4.  Keras models using batch normalization can be unreliable. For some models, forward-pass evaluations (with gradients supposedly off) still result in weights changing at inference time.\n\nTry the sequential API with keras when using models with BN layers... It works fine with me but be careful at some point you will be suffering from exploding gradient. I didn't find any explanation for that but it only works with a limited number of epochs (between 20-30).\nWhen using the Model API the BN layer will be using the mean w variance estimated from imagenet in the test mode. That's why the train loss will be much lower than the validation loss.\nI switech to pytorch with fastai for this reason."
  },
  "source": "meta"
}