{
  "id": 312327,
  "title": "[Info]: SOTA on ImageNet (2022): Model Soups (ViT)",
  "url": "/competitions/happy-whale-and-dolphin/discussion/312327",
  "author_name": "",
  "post_date": "2022-03-11T12:55:58.825711100Z",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time</h2>\n<p><a href=\"https://arxiv.org/pdf/2203.05482v1.pdf\" target=\"_blank\">Paper.</a><br>\n<a href=\"https://www.linkedin.com/posts/papers-with-code_machinelearning-deeplearning-activity-6908011230762778625-lBLG\" target=\"_blank\">References.</a></p>\n<p><strong>TL;DR:</strong> A new paper by Wortsman et al. proposes a simple approach (coined model soup) that often improves accuracy by averaging the weights of models fine-tuned independently, with no additional training and no additional cost at inference time. This approach is often more effective than the conventional recipe for maximizing accuracy and selecting the best model which involves: 1) training models with various hyperparameters and 2) picking the best model on the held-out validation set when fine-tuning. </p>\n<p>Results: <strong>A fine-tuned ViT-G model attains 90.94% top-1 accuracy on ImageNet (a new state-of-the-art).</strong></p>",
  "messages": [
    {
      "id": "1719071",
      "postDate": "03/11/2022 12:55:58",
      "content": "<h2>Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time</h2>\n<p><a href=\"https://arxiv.org/pdf/2203.05482v1.pdf\" target=\"_blank\">Paper.</a><br>\n<a href=\"https://www.linkedin.com/posts/papers-with-code_machinelearning-deeplearning-activity-6908011230762778625-lBLG\" target=\"_blank\">References.</a></p>\n<p><strong>TL;DR:</strong> A new paper by Wortsman et al. proposes a simple approach (coined model soup) that often improves accuracy by averaging the weights of models fine-tuned independently, with no additional training and no additional cost at inference time. This approach is often more effective than the conventional recipe for maximizing accuracy and selecting the best model which involves: 1) training models with various hyperparameters and 2) picking the best model on the held-out validation set when fine-tuning. </p>\n<p>Results: <strong>A fine-tuned ViT-G model attains 90.94% top-1 accuracy on ImageNet (a new state-of-the-art).</strong></p>",
      "rawMarkdown": "## Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time\n\n[Paper.](https://arxiv.org/pdf/2203.05482v1.pdf)\n[References.](https://www.linkedin.com/posts/papers-with-code_machinelearning-deeplearning-activity-6908011230762778625-lBLG)\n\n**TL;DR:** A new paper by Wortsman et al. proposes a simple approach (coined model soup) that often improves accuracy by averaging the weights of models fine-tuned independently, with no additional training and no additional cost at inference time. This approach is often more effective than the conventional recipe for maximizing accuracy and selecting the best model which involves: 1) training models with various hyperparameters and 2) picking the best model on the held-out validation set when fine-tuning. \n\nResults: **A fine-tuned ViT-G model attains 90.94% top-1 accuracy on ImageNet (a new state-of-the-art).**",
      "votes": null
    },
    {
      "id": "1719221",
      "postDate": "03/11/2022 15:15:41",
      "content": "<p>Have you tried it?</p>",
      "rawMarkdown": "Have you tried it?",
      "votes": null
    },
    {
      "id": "1719237",
      "postDate": "03/11/2022 15:32:21",
      "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Thanks for sharing. How often we overlook the power of simplistic approaches in our obsession with something 'advanced'.</p>",
      "rawMarkdown": "ipythonx Thanks for sharing. How often we overlook the power of simplistic approaches in our obsession with something 'advanced'.",
      "votes": null
    },
    {
      "id": "1719247",
      "postDate": "03/11/2022 15:39:45",
      "content": "<p>Have you tried it?</p>",
      "rawMarkdown": "Have you tried it?",
      "votes": null
    },
    {
      "id": "1719333",
      "postDate": "03/11/2022 16:31:36",
      "content": "<p>paper with no code.</p>",
      "rawMarkdown": "paper with no code.",
      "votes": null
    },
    {
      "id": "1720156",
      "postDate": "03/12/2022 14:15:45",
      "content": "<p>To my understanding this does not really have any affect on kaggle competitions since we don't care about inference time and it does not beat ensembles. </p>",
      "rawMarkdown": "To my understanding this does not really have any affect on kaggle competitions since we don't care about inference time and it does not beat ensembles.",
      "votes": null
    },
    {
      "id": "1720368",
      "postDate": "03/12/2022 18:15:28",
      "content": "<blockquote>\n  <p>on kaggle competitions since we don't care about inference time</p>\n</blockquote>\n<p>that maybe holds here but it's not true for code competitions</p>",
      "rawMarkdown": "> on kaggle competitions since we don't care about inference time\n\nthat maybe holds here but it's not true for code competitions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1719221,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "03/11/2022 15:15:41",
      "content": "<p>Have you tried it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1719237,
      "author_name": "gianetan",
      "author_url": "",
      "post_date": "03/11/2022 15:32:21",
      "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Thanks for sharing. How often we overlook the power of simplistic approaches in our obsession with something 'advanced'.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1719247,
      "author_name": "yongdawoon",
      "author_url": "",
      "post_date": "03/11/2022 15:39:45",
      "content": "<p>Have you tried it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1719333,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "03/11/2022 16:31:36",
      "content": "<p>paper with no code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1720156,
      "author_name": "tvaranka",
      "author_url": "",
      "post_date": "03/12/2022 14:15:45",
      "content": "<p>To my understanding this does not really have any affect on kaggle competitions since we don't care about inference time and it does not beat ensembles. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1720368,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "03/12/2022 18:15:28",
          "content": "<blockquote>\n  <p>on kaggle competitions since we don't care about inference time</p>\n</blockquote>\n<p>that maybe holds here but it's not true for code competitions</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1719071": "## Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time\n\n[Paper.](https://arxiv.org/pdf/2203.05482v1.pdf)\n[References.](https://www.linkedin.com/posts/papers-with-code_machinelearning-deeplearning-activity-6908011230762778625-lBLG)\n\n**TL;DR:** A new paper by Wortsman et al. proposes a simple approach (coined model soup) that often improves accuracy by averaging the weights of models fine-tuned independently, with no additional training and no additional cost at inference time. This approach is often more effective than the conventional recipe for maximizing accuracy and selecting the best model which involves: 1) training models with various hyperparameters and 2) picking the best model on the held-out validation set when fine-tuning. \n\nResults: **A fine-tuned ViT-G model attains 90.94% top-1 accuracy on ImageNet (a new state-of-the-art).**",
    "1719221": "Have you tried it?",
    "1719237": "ipythonx Thanks for sharing. How often we overlook the power of simplistic approaches in our obsession with something 'advanced'.",
    "1719247": "Have you tried it?",
    "1719333": "paper with no code.",
    "1720156": "To my understanding this does not really have any affect on kaggle competitions since we don't care about inference time and it does not beat ensembles.",
    "1720368": "> on kaggle competitions since we don't care about inference time\n\nthat maybe holds here but it's not true for code competitions"
  },
  "source": "meta"
}