{
  "id": 125420,
  "title": "Debugging the model",
  "url": "/competitions/deepfake-detection-challenge/discussion/125420",
  "author_name": "Carlos Souza",
  "post_date": "2020-01-10T13:15:05.426000",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Here's some helpful articles about debugging the model:\n- <a href=\"https://missinglink.ai/blog/computer-vision/most-common-neural-net-pytorch-mistakes/\">Most Common Neural Net PyTorch Mistakes</a>\n- <a href=\"https://karpathy.github.io/2019/04/25/recipe/\">A Recipe for Training Neural Networks</a>\n- <a href=\"https://blog.slavv.com/37-reasons-why-your-neural-network-is-not-working-4020854bd607\">37 Reasons why your Neural Network is not working</a>\n- <a href=\"https://gist.github.com/XinDongol/fe066cb76e1c5238ecbc0cb729806410\">How to profile your pytorch codes</a></p>",
  "messages": [
    {
      "id": 715374,
      "postDate": "2020-01-10T13:15:05.427Z",
      "content": "<p>Here's some helpful articles about debugging the model:\n- <a href=\"https://missinglink.ai/blog/computer-vision/most-common-neural-net-pytorch-mistakes/\">Most Common Neural Net PyTorch Mistakes</a>\n- <a href=\"https://karpathy.github.io/2019/04/25/recipe/\">A Recipe for Training Neural Networks</a>\n- <a href=\"https://blog.slavv.com/37-reasons-why-your-neural-network-is-not-working-4020854bd607\">37 Reasons why your Neural Network is not working</a>\n- <a href=\"https://gist.github.com/XinDongol/fe066cb76e1c5238ecbc0cb729806410\">How to profile your pytorch codes</a></p>",
      "rawMarkdown": "Here's some helpful articles about debugging the model:\n- [Most Common Neural Net PyTorch Mistakes](https://missinglink.ai/blog/computer-vision/most-common-neural-net-pytorch-mistakes/)\n- [A Recipe for Training Neural Networks](https://karpathy.github.io/2019/04/25/recipe/)\n- [37 Reasons why your Neural Network is not working](https://blog.slavv.com/37-reasons-why-your-neural-network-is-not-working-4020854bd607)\n- [How to profile your pytorch codes](https://gist.github.com/XinDongol/fe066cb76e1c5238ecbc0cb729806410)",
      "votes": 12
    },
    {
      "id": 715403,
      "postDate": "2020-01-10T13:48:19.597Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 715443,
          "postDate": "2020-01-10T14:43:15.767Z",
          "content": "<p>:)\nI found a huge bug in my code... as it's not described in the articles, let me register here, so others can be aware and avoid.\nWhen doing transfer learning/using ConvNet as fixed feature extractor (<a href=\"https://pytorch.org/tutorials/beginner/transfer_learning_tutorial.html\">tutorial here</a>),  we need to freeze all the network except the final layer. To implement it, we do something like:\n```\nexceptions = ['last_linear']\nfor name, param in model.named_parameters():\n        for e in exceptions:\n            if e in name:\n                param.requires_grad = True\n            else:\n                param.requires_grad = False</p>\n\n<p>optimizer = optim.SGD(model.last_linear.parameters(), lr=0.001, momentum=0.9)\n```\nWith that, we tell the optimizer which tensors should be optimized.</p>\n\n<p>FaceForensics++ training protocol is:\n1) pre-load NN with ImageNet weights\n2) pre-train NN, fixing all parameters except the ones from the last layer for 3 epochs\n3) train the NN for 15 epochs, now optimizing all parameters</p>\n\n<p>From 2 to 3, we have to set <code>param.requires_grad = True</code> to all parameters, and make sure optimizer to optimize all tensors, i.e. use <code>model.parameters()</code> instead of <code>model.last_linear.parameters()</code>. Well, I forgot the latter.. :)</p>",
          "rawMarkdown": ":)\nI found a huge bug in my code... as it's not described in the articles, let me register here, so others can be aware and avoid.\nWhen doing transfer learning/using ConvNet as fixed feature extractor ([tutorial here](https://pytorch.org/tutorials/beginner/transfer_learning_tutorial.html)),  we need to freeze all the network except the final layer. To implement it, we do something like:\n```\nexceptions = ['last_linear']\nfor name, param in model.named_parameters():\n        for e in exceptions:\n            if e in name:\n                param.requires_grad = True\n            else:\n                param.requires_grad = False\n\noptimizer = optim.SGD(model.last_linear.parameters(), lr=0.001, momentum=0.9)\n```\nWith that, we tell the optimizer which tensors should be optimized.\n\nFaceForensics++ training protocol is:\n1) pre-load NN with ImageNet weights\n2) pre-train NN, fixing all parameters except the ones from the last layer for 3 epochs\n3) train the NN for 15 epochs, now optimizing all parameters\n\nFrom 2 to 3, we have to set `param.requires_grad = True` to all parameters, and make sure optimizer to optimize all tensors, i.e. use `model.parameters()` instead of `model.last_linear.parameters()`. Well, I forgot the latter.. :)\n",
          "votes": 1
        },
        {
          "id": 715471,
          "postDate": "2020-01-10T15:12:32.983Z",
          "content": "<p>If you want to make things more generic you can do this on the optimizer (after setting requires_grad):</p>\n\n<p><code>optimizer = optim.SGD(filter(lambda p: p.requires_grad, model.parameters()), lr=0.001, momentum=0.9)</code></p>",
          "rawMarkdown": "If you want to make things more generic you can do this on the optimizer (after setting requires_grad):\n\n`optimizer = optim.SGD(filter(lambda p: p.requires_grad, model.parameters()), lr=0.001, momentum=0.9)`",
          "votes": 1
        },
        {
          "id": 715496,
          "postDate": "2020-01-10T15:37:38.143Z",
          "content": "<p>&gt; <strong>Carlos Souza wrote:</strong>\n&gt; \n&gt; FaceForensics++ training protocol is:\n&gt; 1) pre-load NN with ImageNet weights\n&gt; 2) pre-train NN, fixing all parameters except the ones from the last layer for 3 epochs\n&gt; 3) train the NN for 15 epochs, now optimizing all parameters\n&gt; \n&gt; From 2 to 3, we have to set <code>param.requires_grad = True</code> to all parameters, and make sure optimizer to optimize all tensors, i.e. use <code>model.parameters()</code> instead of <code>model.last_linear.parameters()</code>. Well, I forgot the latter.. :)\n&gt; </p>\n\n<p>Keras/TF man here, so can't help with pytorch stuff 😿 \nbut I have a question regarding step 3 here, did they just train it for 15 epochs without any early stopping? </p>\n\n<p>great tips by the way, thanks!</p>",
          "rawMarkdown": "&gt; **Carlos Souza wrote:**\n&gt; \n&gt; FaceForensics++ training protocol is:\n&gt; 1) pre-load NN with ImageNet weights\n&gt; 2) pre-train NN, fixing all parameters except the ones from the last layer for 3 epochs\n&gt; 3) train the NN for 15 epochs, now optimizing all parameters\n&gt; \n&gt; From 2 to 3, we have to set `param.requires_grad = True` to all parameters, and make sure optimizer to optimize all tensors, i.e. use `model.parameters()` instead of `model.last_linear.parameters()`. Well, I forgot the latter.. :)\n&gt; \n\nKeras/TF man here, so can't help with pytorch stuff 😿 \nbut I have a question regarding step 3 here, did they just train it for 15 epochs without any early stopping? \n\ngreat tips by the way, thanks!\n"
        },
        {
          "id": 715628,
          "postDate": "2020-01-10T17:07:55.440Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> , great trick, thanks!\n<a href=\"/yifanxie\">@yifanxie</a> , the paper is not clear, but in my understanding, no, they don't implement early stopping. Actually 15 epochs is not that long, so imho it doesn't make sense to implement early stopping anyway. Cheers!</p>",
          "rawMarkdown": "@pedromb , great trick, thanks!\n@yifanxie , the paper is not clear, but in my understanding, no, they don't implement early stopping. Actually 15 epochs is not that long, so imho it doesn't make sense to implement early stopping anyway. Cheers!"
        },
        {
          "id": 715637,
          "postDate": "2020-01-10T17:18:18.687Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> <a href=\"/yifanxie\">@yifanxie</a> They do implement early stopping, it's described at the very end of the appendix on the paper:</p>\n\n<blockquote>\n  <p>We compute validation accuracies ten times per epoch and stop the training process if the validation accuracy does not change for 10 consecutive checks.</p>\n</blockquote>\n\n<p>Reference: <a href=\"https://arxiv.org/pdf/1901.08971.pdf\">https://arxiv.org/pdf/1901.08971.pdf</a></p>",
          "rawMarkdown": "@carlossouza @yifanxie They do implement early stopping, it's described at the very end of the appendix on the paper:\n\n&gt; We compute validation accuracies ten times per epoch and stop the training process if the validation accuracy does not change for 10 consecutive checks.\n\nReference: https://arxiv.org/pdf/1901.08971.pdf",
          "votes": 2
        },
        {
          "id": 715638,
          "postDate": "2020-01-10T17:20:11.540Z",
          "content": "<blockquote>\n  <p><a href=\"/yifanxie\">@yifanxie</a> , the paper is not clear, but in my understanding, no, they don't implement early stopping. Actually 15 epochs is not that long, so imho it doesn't make sense to implement early stopping anyway. Cheers!</p>\n</blockquote>\n\n<p>Thanks for getting back to me again. I have another point about steps 2 &amp; 3.\nI don't have extensive experience in using pre-train model, just projects here and there, so I can't say for sure, but I haven't come across any \"formal method of potential steps to be taken regarding using pre-train models (their architect &amp; weight) </p>\n\n<p>For instance, is it a recognised procedure that you would do something like step 2 (i.e. training with bottleneck features with top N layer frozen), and then step 3 with full training?</p>\n\n<p>At least in my experience, I have had success doing step 3 first, and then retrain with step 2 - but then was it just by chance? </p>\n\n<p>a couple of years ago, when <a href=\"/alexisbcook\">@alexisbcook</a> was still with Udacity, she gave a great tutorial on how many layers of bottleneck to use in different situations. If I remember/understand correctly, for large dataset that is different to source data of pre-train weighted (i.e. imagenet) you should just retrain the network from the ground-up. </p>\n\n<p>anyway, just sharing my thought, perhaps someone of the experienced guys around here can share some tips? </p>",
          "rawMarkdown": "&gt; @yifanxie , the paper is not clear, but in my understanding, no, they don't implement early stopping. Actually 15 epochs is not that long, so imho it doesn't make sense to implement early stopping anyway. Cheers!\n\nThanks for getting back to me again. I have another point about steps 2 &amp; 3.\nI don't have extensive experience in using pre-train model, just projects here and there, so I can't say for sure, but I haven't come across any \"formal method of potential steps to be taken regarding using pre-train models (their architect &amp; weight) \n\nFor instance, is it a recognised procedure that you would do something like step 2 (i.e. training with bottleneck features with top N layer frozen), and then step 3 with full training?\n\nAt least in my experience, I have had success doing step 3 first, and then retrain with step 2 - but then was it just by chance? \n\na couple of years ago, when @alexisbcook was still with Udacity, she gave a great tutorial on how many layers of bottleneck to use in different situations. If I remember/understand correctly, for large dataset that is different to source data of pre-train weighted (i.e. imagenet) you should just retrain the network from the ground-up. \n\nanyway, just sharing my thought, perhaps someone of the experienced guys around here can share some tips? "
        },
        {
          "id": 715745,
          "postDate": "2020-01-10T19:18:11.763Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F97fe592837be08cd8d164ddbdbed43a4%2F6B6D338E-0A11-4688-9136-A36B817CC69B.jpeg?generation=1578683870088691&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://towardsdatascience.com/transfer-learning-from-pre-trained-models-f2393f124751\">https://towardsdatascience.com/transfer-learning-from-pre-trained-models-f2393f124751</a></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F97fe592837be08cd8d164ddbdbed43a4%2F6B6D338E-0A11-4688-9136-A36B817CC69B.jpeg?generation=1578683870088691&amp;alt=media)\n\nhttps://towardsdatascience.com/transfer-learning-from-pre-trained-models-f2393f124751\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 715403,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-10T13:48:19.597000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 715443,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-10T14:43:15.767000",
          "content": "<p>:)\nI found a huge bug in my code... as it's not described in the articles, let me register here, so others can be aware and avoid.\nWhen doing transfer learning/using ConvNet as fixed feature extractor (<a href=\"https://pytorch.org/tutorials/beginner/transfer_learning_tutorial.html\">tutorial here</a>),  we need to freeze all the network except the final layer. To implement it, we do something like:\n```\nexceptions = ['last_linear']\nfor name, param in model.named_parameters():\n        for e in exceptions:\n            if e in name:\n                param.requires_grad = True\n            else:\n                param.requires_grad = False</p>\n\n<p>optimizer = optim.SGD(model.last_linear.parameters(), lr=0.001, momentum=0.9)\n```\nWith that, we tell the optimizer which tensors should be optimized.</p>\n\n<p>FaceForensics++ training protocol is:\n1) pre-load NN with ImageNet weights\n2) pre-train NN, fixing all parameters except the ones from the last layer for 3 epochs\n3) train the NN for 15 epochs, now optimizing all parameters</p>\n\n<p>From 2 to 3, we have to set <code>param.requires_grad = True</code> to all parameters, and make sure optimizer to optimize all tensors, i.e. use <code>model.parameters()</code> instead of <code>model.last_linear.parameters()</code>. Well, I forgot the latter.. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715471,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-10T15:12:32.983000",
          "content": "<p>If you want to make things more generic you can do this on the optimizer (after setting requires_grad):</p>\n\n<p><code>optimizer = optim.SGD(filter(lambda p: p.requires_grad, model.parameters()), lr=0.001, momentum=0.9)</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715496,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T15:37:38.143000",
          "content": "<p>&gt; <strong>Carlos Souza wrote:</strong>\n&gt; \n&gt; FaceForensics++ training protocol is:\n&gt; 1) pre-load NN with ImageNet weights\n&gt; 2) pre-train NN, fixing all parameters except the ones from the last layer for 3 epochs\n&gt; 3) train the NN for 15 epochs, now optimizing all parameters\n&gt; \n&gt; From 2 to 3, we have to set <code>param.requires_grad = True</code> to all parameters, and make sure optimizer to optimize all tensors, i.e. use <code>model.parameters()</code> instead of <code>model.last_linear.parameters()</code>. Well, I forgot the latter.. :)\n&gt; </p>\n\n<p>Keras/TF man here, so can't help with pytorch stuff 😿 \nbut I have a question regarding step 3 here, did they just train it for 15 epochs without any early stopping? </p>\n\n<p>great tips by the way, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715628,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-10T17:07:55.440000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> , great trick, thanks!\n<a href=\"/yifanxie\">@yifanxie</a> , the paper is not clear, but in my understanding, no, they don't implement early stopping. Actually 15 epochs is not that long, so imho it doesn't make sense to implement early stopping anyway. Cheers!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715637,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-10T17:18:18.687000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> <a href=\"/yifanxie\">@yifanxie</a> They do implement early stopping, it's described at the very end of the appendix on the paper:</p>\n\n<blockquote>\n  <p>We compute validation accuracies ten times per epoch and stop the training process if the validation accuracy does not change for 10 consecutive checks.</p>\n</blockquote>\n\n<p>Reference: <a href=\"https://arxiv.org/pdf/1901.08971.pdf\">https://arxiv.org/pdf/1901.08971.pdf</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 715638,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T17:20:11.540000",
          "content": "<blockquote>\n  <p><a href=\"/yifanxie\">@yifanxie</a> , the paper is not clear, but in my understanding, no, they don't implement early stopping. Actually 15 epochs is not that long, so imho it doesn't make sense to implement early stopping anyway. Cheers!</p>\n</blockquote>\n\n<p>Thanks for getting back to me again. I have another point about steps 2 &amp; 3.\nI don't have extensive experience in using pre-train model, just projects here and there, so I can't say for sure, but I haven't come across any \"formal method of potential steps to be taken regarding using pre-train models (their architect &amp; weight) </p>\n\n<p>For instance, is it a recognised procedure that you would do something like step 2 (i.e. training with bottleneck features with top N layer frozen), and then step 3 with full training?</p>\n\n<p>At least in my experience, I have had success doing step 3 first, and then retrain with step 2 - but then was it just by chance? </p>\n\n<p>a couple of years ago, when <a href=\"/alexisbcook\">@alexisbcook</a> was still with Udacity, she gave a great tutorial on how many layers of bottleneck to use in different situations. If I remember/understand correctly, for large dataset that is different to source data of pre-train weighted (i.e. imagenet) you should just retrain the network from the ground-up. </p>\n\n<p>anyway, just sharing my thought, perhaps someone of the experienced guys around here can share some tips? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715745,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-10T19:18:11.763000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F97fe592837be08cd8d164ddbdbed43a4%2F6B6D338E-0A11-4688-9136-A36B817CC69B.jpeg?generation=1578683870088691&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://towardsdatascience.com/transfer-learning-from-pre-trained-models-f2393f124751\">https://towardsdatascience.com/transfer-learning-from-pre-trained-models-f2393f124751</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "715374": "Here's some helpful articles about debugging the model:\n- [Most Common Neural Net PyTorch Mistakes](https://missinglink.ai/blog/computer-vision/most-common-neural-net-pytorch-mistakes/)\n- [A Recipe for Training Neural Networks](https://karpathy.github.io/2019/04/25/recipe/)\n- [37 Reasons why your Neural Network is not working](https://blog.slavv.com/37-reasons-why-your-neural-network-is-not-working-4020854bd607)\n- [How to profile your pytorch codes](https://gist.github.com/XinDongol/fe066cb76e1c5238ecbc0cb729806410)",
    "715403": ""
  }
}