{
  "id": 210226,
  "title": "SAM Optimizer",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210226",
  "author_name": "Bibek",
  "post_date": "2021-01-10T06:25:47.529000",
  "votes": 20,
  "comment_count": 9,
  "views": 0,
  "content": "<p><strong>Sharpness-Aware Minimization for Efficiently Improving Generalization</strong></p>\n<p><img src=\"https://raw.githubusercontent.com/davda54/sam/main/img/loss_landscape.png\" alt=\"\"></p>\n<pre><code>In today’s heavily overparameterized models, the value of the training loss provides few guarantees on model generalization ability. Indeed, optimizing only the training loss value, as is commonly done, can easily lead to suboptimal model quality. Motivated by the connection between\ngeometry of the loss landscape and generalization—including a generalization bound that we prove\nhere—we introduce a novel, effective procedure for instead simultaneously minimizing loss value\nand loss sharpness. In particular, our procedure, Sharpness-Aware Minimization (SAM), seeks\nparameters that lie in neighborhoods having uniformly low loss; this formulation results in a minmax optimization problem on which gradient descent can be performed efficiently. We present\nempirical results showing that SAM improves model generalization across a variety of benchmark\ndatasets (e.g., CIFAR-{10, 100}, ImageNet, finetuning tasks) and models, yielding novel state-ofthe-art performance for several. Additionally, we find that SAM natively provides robustness to\nlabel noise on par with that provided by state-of-the-art procedures that specifically target learning with noisy labels\n</code></pre>\n<p>This looks interesting and might be worth trying given that it claims to be robust to label noise.</p>\n<p>Pytorch Implementation: <a href=\"https://github.com/davda54/sam\" target=\"_blank\">https://github.com/davda54/sam</a></p>",
  "messages": [
    {
      "id": 1146888,
      "postDate": "2021-01-10T06:25:47.530Z",
      "content": "<p><strong>Sharpness-Aware Minimization for Efficiently Improving Generalization</strong></p>\n<p><img src=\"https://raw.githubusercontent.com/davda54/sam/main/img/loss_landscape.png\" alt=\"\"></p>\n<pre><code>In today’s heavily overparameterized models, the value of the training loss provides few guarantees on model generalization ability. Indeed, optimizing only the training loss value, as is commonly done, can easily lead to suboptimal model quality. Motivated by the connection between\ngeometry of the loss landscape and generalization—including a generalization bound that we prove\nhere—we introduce a novel, effective procedure for instead simultaneously minimizing loss value\nand loss sharpness. In particular, our procedure, Sharpness-Aware Minimization (SAM), seeks\nparameters that lie in neighborhoods having uniformly low loss; this formulation results in a minmax optimization problem on which gradient descent can be performed efficiently. We present\nempirical results showing that SAM improves model generalization across a variety of benchmark\ndatasets (e.g., CIFAR-{10, 100}, ImageNet, finetuning tasks) and models, yielding novel state-ofthe-art performance for several. Additionally, we find that SAM natively provides robustness to\nlabel noise on par with that provided by state-of-the-art procedures that specifically target learning with noisy labels\n</code></pre>\n<p>This looks interesting and might be worth trying given that it claims to be robust to label noise.</p>\n<p>Pytorch Implementation: <a href=\"https://github.com/davda54/sam\" target=\"_blank\">https://github.com/davda54/sam</a></p>",
      "rawMarkdown": "**Sharpness-Aware Minimization for Efficiently Improving Generalization**\n\n![](https://raw.githubusercontent.com/davda54/sam/main/img/loss_landscape.png)\n\n```\nIn today’s heavily overparameterized models, the value of the training loss provides few guarantees on model generalization ability. Indeed, optimizing only the training loss value, as is commonly done, can easily lead to suboptimal model quality. Motivated by the connection between\ngeometry of the loss landscape and generalization—including a generalization bound that we prove\nhere—we introduce a novel, effective procedure for instead simultaneously minimizing loss value\nand loss sharpness. In particular, our procedure, Sharpness-Aware Minimization (SAM), seeks\nparameters that lie in neighborhoods having uniformly low loss; this formulation results in a minmax optimization problem on which gradient descent can be performed efficiently. We present\nempirical results showing that SAM improves model generalization across a variety of benchmark\ndatasets (e.g., CIFAR-{10, 100}, ImageNet, finetuning tasks) and models, yielding novel state-ofthe-art performance for several. Additionally, we find that SAM natively provides robustness to\nlabel noise on par with that provided by state-of-the-art procedures that specifically target learning with noisy labels\n```\nThis looks interesting and might be worth trying given that it claims to be robust to label noise.\n\nPytorch Implementation: https://github.com/davda54/sam",
      "votes": 20
    },
    {
      "id": 1152830,
      "postDate": "2021-01-14T13:33:28.113Z",
      "content": "<p>Hey, I came up with this notebook adding shapeness-aware minimization, you can check it out:<br>\n<a href=\"https://www.kaggle.com/alexanderriedel/cassava-resnext50-32x4d-sam\" target=\"_blank\">https://www.kaggle.com/alexanderriedel/cassava-resnext50-32x4d-sam</a></p>",
      "rawMarkdown": "Hey, I came up with this notebook adding shapeness-aware minimization, you can check it out:\nhttps://www.kaggle.com/alexanderriedel/cassava-resnext50-32x4d-sam",
      "votes": 1
    },
    {
      "id": 1149012,
      "postDate": "2021-01-11T14:32:07.027Z",
      "content": "<p>Have you tried any experiments with this?<br>\nIf yes, could you please share them</p>",
      "rawMarkdown": "Have you tried any experiments with this?\nIf yes, could you please share them"
    },
    {
      "id": 1147877,
      "postDate": "2021-01-10T19:11:34.547Z",
      "content": "<p>This seems extremely useful, hope to give it a go soon. Thanks for posting this.<br>\nEdit: On the second backward step, I get a RuntimeError: Trying to backward through the graph a second time, but the saved intermediate results have already been freed. Specify retain_graph=True when calling backward the first time.</p>\n<p>When I do retain_graph = True, I get one of the variables needed for gradient computation has been modified by an inplace operation: </p>\n<p>Maybe something is wrong with my model.</p>",
      "rawMarkdown": "This seems extremely useful, hope to give it a go soon. Thanks for posting this.\nEdit: On the second backward step, I get a RuntimeError: Trying to backward through the graph a second time, but the saved intermediate results have already been freed. Specify retain_graph=True when calling backward the first time.\n\nWhen I do retain_graph = True, I get one of the variables needed for gradient computation has been modified by an inplace operation: \n\nMaybe something is wrong with my model.",
      "replies": [
        {
          "id": 1149458,
          "postDate": "2021-01-11T21:10:49.190Z",
          "content": "<p>You have to use it as shown in the github repository:</p>\n<pre><code>...\n# first forward-backward pass\n  loss = loss_function(output, model(input))  # use this loss for any training statistics\n  loss.backward()\n  optimizer.first_step(zero_grad=True)\n\n  # second forward-backward pass\n  loss_function(output, model(input)).backward()\n  optimizer.second_step(zero_grad=True)\n...\n</code></pre>\n<p>Dont create a second loss variable, you have to do it exactly how it is shown here, then it should work.</p>",
          "rawMarkdown": "You have to use it as shown in the github repository:\n\n```\n...\n# first forward-backward pass\n  loss = loss_function(output, model(input))  # use this loss for any training statistics\n  loss.backward()\n  optimizer.first_step(zero_grad=True)\n  \n  # second forward-backward pass\n  loss_function(output, model(input)).backward()\n  optimizer.second_step(zero_grad=True)\n...\n```\n\nDont create a second loss variable, you have to do it exactly how it is shown here, then it should work.",
          "votes": 2
        },
        {
          "id": 1149486,
          "postDate": "2021-01-11T21:43:23.817Z",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I tried using this optimizer but I need to reduce the batch size to 8 instead of 16 or else Cuda memory gets exhausted<br>\nAny idea why?</p>",
          "rawMarkdown": "@aliabdin1 I tried using this optimizer but I need to reduce the batch size to 8 instead of 16 or else Cuda memory gets exhausted\nAny idea why?",
          "votes": 1
        },
        {
          "id": 1149495,
          "postDate": "2021-01-11T21:58:08.017Z",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> You are foward-backwarding twice instead of once with this optimizer therefore you have to save more parameters in your iteration. Training will also take longer due to that.</p>",
          "rawMarkdown": "@debarshichanda You are foward-backwarding twice instead of once with this optimizer therefore you have to save more parameters in your iteration. Training will also take longer due to that.",
          "votes": 2
        },
        {
          "id": 1149520,
          "postDate": "2021-01-11T22:49:36.780Z",
          "content": "<p>Thanks for the reply Ali. I thought I was following the example without creating second loss variable. But then I realized that instead of criterion(model(input), ground_truth).backward(), I was using criterion(output, ground_truth).backward(), i.e. I wasn't going through a second forward pass. Thanks for leading me down the right path.</p>",
          "rawMarkdown": "Thanks for the reply Ali. I thought I was following the example without creating second loss variable. But then I realized that instead of criterion(model(input), ground_truth).backward(), I was using criterion(output, ground_truth).backward(), i.e. I wasn't going through a second forward pass. Thanks for leading me down the right path."
        },
        {
          "id": 1149521,
          "postDate": "2021-01-11T22:52:13.877Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1149013,
      "postDate": "2021-01-11T14:32:07.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1152830,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-01-14T13:33:28.113000",
      "content": "<p>Hey, I came up with this notebook adding shapeness-aware minimization, you can check it out:<br>\n<a href=\"https://www.kaggle.com/alexanderriedel/cassava-resnext50-32x4d-sam\" target=\"_blank\">https://www.kaggle.com/alexanderriedel/cassava-resnext50-32x4d-sam</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1149012,
      "author_name": "Debarshi Chanda",
      "author_url": "",
      "post_date": "2021-01-11T14:32:07.027000",
      "content": "<p>Have you tried any experiments with this?<br>\nIf yes, could you please share them</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1147877,
      "author_name": "Michael Bolton",
      "author_url": "",
      "post_date": "2021-01-10T19:11:34.547000",
      "content": "<p>This seems extremely useful, hope to give it a go soon. Thanks for posting this.<br>\nEdit: On the second backward step, I get a RuntimeError: Trying to backward through the graph a second time, but the saved intermediate results have already been freed. Specify retain_graph=True when calling backward the first time.</p>\n<p>When I do retain_graph = True, I get one of the variables needed for gradient computation has been modified by an inplace operation: </p>\n<p>Maybe something is wrong with my model.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1149458,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2021-01-11T21:10:49.190000",
          "content": "<p>You have to use it as shown in the github repository:</p>\n<pre><code>...\n# first forward-backward pass\n  loss = loss_function(output, model(input))  # use this loss for any training statistics\n  loss.backward()\n  optimizer.first_step(zero_grad=True)\n\n  # second forward-backward pass\n  loss_function(output, model(input)).backward()\n  optimizer.second_step(zero_grad=True)\n...\n</code></pre>\n<p>Dont create a second loss variable, you have to do it exactly how it is shown here, then it should work.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1149486,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2021-01-11T21:43:23.817000",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I tried using this optimizer but I need to reduce the batch size to 8 instead of 16 or else Cuda memory gets exhausted<br>\nAny idea why?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1149495,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2021-01-11T21:58:08.017000",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> You are foward-backwarding twice instead of once with this optimizer therefore you have to save more parameters in your iteration. Training will also take longer due to that.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1149520,
          "author_name": "Michael Bolton",
          "author_url": "",
          "post_date": "2021-01-11T22:49:36.780000",
          "content": "<p>Thanks for the reply Ali. I thought I was following the example without creating second loss variable. But then I realized that instead of criterion(model(input), ground_truth).backward(), I was using criterion(output, ground_truth).backward(), i.e. I wasn't going through a second forward pass. Thanks for leading me down the right path.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1149521,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-11T22:52:13.877000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1149013,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-11T14:32:07.027000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1146888": "**Sharpness-Aware Minimization for Efficiently Improving Generalization**\n\n![](https://raw.githubusercontent.com/davda54/sam/main/img/loss_landscape.png)\n\n```\nIn today’s heavily overparameterized models, the value of the training loss provides few guarantees on model generalization ability. Indeed, optimizing only the training loss value, as is commonly done, can easily lead to suboptimal model quality. Motivated by the connection between\ngeometry of the loss landscape and generalization—including a generalization bound that we prove\nhere—we introduce a novel, effective procedure for instead simultaneously minimizing loss value\nand loss sharpness. In particular, our procedure, Sharpness-Aware Minimization (SAM), seeks\nparameters that lie in neighborhoods having uniformly low loss; this formulation results in a minmax optimization problem on which gradient descent can be performed efficiently. We present\nempirical results showing that SAM improves model generalization across a variety of benchmark\ndatasets (e.g., CIFAR-{10, 100}, ImageNet, finetuning tasks) and models, yielding novel state-ofthe-art performance for several. Additionally, we find that SAM natively provides robustness to\nlabel noise on par with that provided by state-of-the-art procedures that specifically target learning with noisy labels\n```\nThis looks interesting and might be worth trying given that it claims to be robust to label noise.\n\nPytorch Implementation: https://github.com/davda54/sam",
    "1152830": "Hey, I came up with this notebook adding shapeness-aware minimization, you can check it out:\nhttps://www.kaggle.com/alexanderriedel/cassava-resnext50-32x4d-sam",
    "1149012": "Have you tried any experiments with this?\nIf yes, could you please share them",
    "1147877": "This seems extremely useful, hope to give it a go soon. Thanks for posting this.\nEdit: On the second backward step, I get a RuntimeError: Trying to backward through the graph a second time, but the saved intermediate results have already been freed. Specify retain_graph=True when calling backward the first time.\n\nWhen I do retain_graph = True, I get one of the variables needed for gradient computation has been modified by an inplace operation: \n\nMaybe something is wrong with my model.",
    "1149013": ""
  }
}