{
  "id": 398741,
  "title": "Model soup is not working here",
  "url": "/competitions/asl-signs/discussion/398741",
  "author_name": "",
  "post_date": "2023-03-31T13:58:45.761041300Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello, guys. </p>\n<p>I've tried to cook a model soup. I've taken models trained using K-Fold validation. The result I obtain is horrible. Having a ~0.70 val acc model soup from 5 models get 0.01 accuracy. The first thing I checked was my recipe, maybe something went wrong in my kitchen, but it seems there is no place for mistakes.<br>\nTo be sure, I've created a minimalistic example of classification model training and soup cooking. <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow/custom-receipt-for-tiny-fc-model-soup\" target=\"_blank\">Here is my minimalistic notebook</a>, if someone is interested. Model soup always gives a worse score.</p>\n<p>I believe the original paper authors provide real results, so I reproduced an example from this <a href=\"https://github.com/Burf/ModelSoups\" target=\"_blank\">repository</a>, and classical soup with ResNet50 works as expected.<br>\nDoes anyone know why model soup performs worse than an individual model in our task or doesn't work for FC layers generally? Is the model size a key issue here? </p>\n<p>Would love to hear your thoughts. Thank you in advance </p>",
  "messages": [
    {
      "id": "2204357",
      "postDate": "03/31/2023 13:58:45",
      "content": "<p>Hello, guys. </p>\n<p>I've tried to cook a model soup. I've taken models trained using K-Fold validation. The result I obtain is horrible. Having a ~0.70 val acc model soup from 5 models get 0.01 accuracy. The first thing I checked was my recipe, maybe something went wrong in my kitchen, but it seems there is no place for mistakes.<br>\nTo be sure, I've created a minimalistic example of classification model training and soup cooking. <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow/custom-receipt-for-tiny-fc-model-soup\" target=\"_blank\">Here is my minimalistic notebook</a>, if someone is interested. Model soup always gives a worse score.</p>\n<p>I believe the original paper authors provide real results, so I reproduced an example from this <a href=\"https://github.com/Burf/ModelSoups\" target=\"_blank\">repository</a>, and classical soup with ResNet50 works as expected.<br>\nDoes anyone know why model soup performs worse than an individual model in our task or doesn't work for FC layers generally? Is the model size a key issue here? </p>\n<p>Would love to hear your thoughts. Thank you in advance </p>",
      "rawMarkdown": "Hello, guys. \n\nI've tried to cook a model soup. I've taken models trained using K-Fold validation. The result I obtain is horrible. Having a ~0.70 val acc model soup from 5 models get 0.01 accuracy. The first thing I checked was my recipe, maybe something went wrong in my kitchen, but it seems there is no place for mistakes.\nTo be sure, I've created a minimalistic example of classification model training and soup cooking. [Here is my minimalistic notebook](https://www.kaggle.com/meowmeowmeowmeowmeow/custom-receipt-for-tiny-fc-model-soup), if someone is interested. Model soup always gives a worse score.\n\nI believe the original paper authors provide real results, so I reproduced an example from this [repository](https://github.com/Burf/ModelSoups), and classical soup with ResNet50 works as expected.\nDoes anyone know why model soup performs worse than an individual model in our task or doesn't work for FC layers generally? Is the model size a key issue here? \n\nWould love to hear your thoughts. Thank you in advance",
      "votes": null
    },
    {
      "id": "2205034",
      "postDate": "04/01/2023 07:29:56",
      "content": "<p>In the paper they say:</p>\n<blockquote>\n  <p>While the most straightforward approach to making a model soup is to average all the weights uniformly, we find that <strong>greedy soups</strong>, where models are sequentially added to the soup if they improve accuracy on held-out data, <strong>outperforms uniform averaging</strong>. <strong>Greedy soups avoid adding in models which may lie in a different basin of the error landscape</strong>, which could happen if, for example, models are fine-tuned with high learning rates.</p>\n</blockquote>\n<p>So it seems to merge different basins of the error landscape negatively impacts the performance of the resulting model. In the paper, they fine-tune from a pre-trained checkpoint and as shown in the paper <a href=\"https://arxiv.org/abs/1912.02757\" target=\"_blank\">Deep Ensembles: A Loss Landscape Perspective</a> this is very likely to produce the model from the same locally-optimal basin. So I would suggest training the K-Fold not from scratch but from some reasonably trained common checkpoint.</p>",
      "rawMarkdown": "In the paper they say:\n\n>While the most straightforward approach to making a model soup is to average all the weights uniformly, we find that **greedy soups**, where models are sequentially added to the soup if they improve accuracy on held-out data, **outperforms uniform averaging**. **Greedy soups avoid adding in models which may lie in a different basin of the error landscape**, which could happen if, for example, models are fine-tuned with high learning rates.\n\nSo it seems to merge different basins of the error landscape negatively impacts the performance of the resulting model. In the paper, they fine-tune from a pre-trained checkpoint and as shown in the paper [Deep Ensembles: A Loss Landscape Perspective](https://arxiv.org/abs/1912.02757) this is very likely to produce the model from the same locally-optimal basin. So I would suggest training the K-Fold not from scratch but from some reasonably trained common checkpoint.",
      "votes": null
    },
    {
      "id": "2205236",
      "postDate": "04/01/2023 11:13:39",
      "content": "<p>Thank you so much. That does make sense to fine-tune the model from the pretrained checkpoint. Regarding greedy soups - I've tried it as well but that also is not working if the model is traiend from scratch.</p>",
      "rawMarkdown": "Thank you so much. That does make sense to fine-tune the model from the pretrained checkpoint. Regarding greedy soups - I've tried it as well but that also is not working if the model is traiend from scratch.",
      "votes": null
    },
    {
      "id": "2207179",
      "postDate": "04/03/2023 07:12:38",
      "content": "<p>Danylo's explanation was spot on. For me, model souping k-last/ k-best checkpoints from same fold improved val. acc a lot. </p>",
      "rawMarkdown": "Danylo's explanation was spot on. For me, model souping k-last/ k-best checkpoints from same fold improved val. acc a lot.",
      "votes": null
    },
    {
      "id": "2207297",
      "postDate": "04/03/2023 10:13:38",
      "content": "<p>That is a beautiful idea. Thank you for sharing it. I will try it out</p>",
      "rawMarkdown": "That is a beautiful idea. Thank you for sharing it. I will try it out",
      "votes": null
    },
    {
      "id": "2212818",
      "postDate": "04/07/2023 05:43:05",
      "content": "<p>!!!<br>\nyou can try this paper<br>\n<img src=\"https://i.ibb.co/hCFMT1v/Selection-999-1716.png\" alt=\"https://i.ibb.co/hCFMT1v/Selection-999-1716.png\"></p>",
      "rawMarkdown": "!!!\nyou can try this paper\n![https://i.ibb.co/hCFMT1v/Selection-999-1716.png](https://i.ibb.co/hCFMT1v/Selection-999-1716.png)",
      "votes": null
    },
    {
      "id": "2229086",
      "postDate": "04/21/2023 04:31:31",
      "content": "<p>Model soups are for finetuning large pre-trained models. If your model is neither pre-trained nor large making soup will not help. Models trained with different random initialization will lie in different error basins. See Git Re-Basin: Merging Models modulo Permutation Symmetries </p>",
      "rawMarkdown": "Model soups are for finetuning large pre-trained models. If your model is neither pre-trained nor large making soup will not help. Models trained with different random initialization will lie in different error basins. See Git Re-Basin: Merging Models modulo Permutation Symmetries",
      "votes": null
    },
    {
      "id": "2229172",
      "postDate": "04/21/2023 06:01:30",
      "content": "<p><img src=\"https://i.ibb.co/f9WPCQW/Selection-999-1854.png\" alt=\"https://i.ibb.co/f9WPCQW/Selection-999-1854.png\"></p>\n<p>this diagram from the paper is interesting.<br>\ncombined data beats all</p>",
      "rawMarkdown": "![https://i.ibb.co/f9WPCQW/Selection-999-1854.png](https://i.ibb.co/f9WPCQW/Selection-999-1854.png)\n\nthis diagram from the paper is interesting.\ncombined data beats all",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2205034,
      "author_name": "kasyanenko",
      "author_url": "",
      "post_date": "04/01/2023 07:29:56",
      "content": "<p>In the paper they say:</p>\n<blockquote>\n  <p>While the most straightforward approach to making a model soup is to average all the weights uniformly, we find that <strong>greedy soups</strong>, where models are sequentially added to the soup if they improve accuracy on held-out data, <strong>outperforms uniform averaging</strong>. <strong>Greedy soups avoid adding in models which may lie in a different basin of the error landscape</strong>, which could happen if, for example, models are fine-tuned with high learning rates.</p>\n</blockquote>\n<p>So it seems to merge different basins of the error landscape negatively impacts the performance of the resulting model. In the paper, they fine-tune from a pre-trained checkpoint and as shown in the paper <a href=\"https://arxiv.org/abs/1912.02757\" target=\"_blank\">Deep Ensembles: A Loss Landscape Perspective</a> this is very likely to produce the model from the same locally-optimal basin. So I would suggest training the K-Fold not from scratch but from some reasonably trained common checkpoint.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2205236,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "04/01/2023 11:13:39",
          "content": "<p>Thank you so much. That does make sense to fine-tune the model from the pretrained checkpoint. Regarding greedy soups - I've tried it as well but that also is not working if the model is traiend from scratch.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2207179,
              "author_name": "andy2709",
              "author_url": "",
              "post_date": "04/03/2023 07:12:38",
              "content": "<p>Danylo's explanation was spot on. For me, model souping k-last/ k-best checkpoints from same fold improved val. acc a lot. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2207297,
                  "author_name": "meowmeowmeowmeowmeow",
                  "author_url": "",
                  "post_date": "04/03/2023 10:13:38",
                  "content": "<p>That is a beautiful idea. Thank you for sharing it. I will try it out</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2212818,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/07/2023 05:43:05",
      "content": "<p>!!!<br>\nyou can try this paper<br>\n<img src=\"https://i.ibb.co/hCFMT1v/Selection-999-1716.png\" alt=\"https://i.ibb.co/hCFMT1v/Selection-999-1716.png\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2229086,
      "author_name": "readoc",
      "author_url": "",
      "post_date": "04/21/2023 04:31:31",
      "content": "<p>Model soups are for finetuning large pre-trained models. If your model is neither pre-trained nor large making soup will not help. Models trained with different random initialization will lie in different error basins. See Git Re-Basin: Merging Models modulo Permutation Symmetries </p>",
      "votes": null,
      "replies": [
        {
          "id": 2229172,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/21/2023 06:01:30",
          "content": "<p><img src=\"https://i.ibb.co/f9WPCQW/Selection-999-1854.png\" alt=\"https://i.ibb.co/f9WPCQW/Selection-999-1854.png\"></p>\n<p>this diagram from the paper is interesting.<br>\ncombined data beats all</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2204357": "Hello, guys. \n\nI've tried to cook a model soup. I've taken models trained using K-Fold validation. The result I obtain is horrible. Having a ~0.70 val acc model soup from 5 models get 0.01 accuracy. The first thing I checked was my recipe, maybe something went wrong in my kitchen, but it seems there is no place for mistakes.\nTo be sure, I've created a minimalistic example of classification model training and soup cooking. [Here is my minimalistic notebook](https://www.kaggle.com/meowmeowmeowmeowmeow/custom-receipt-for-tiny-fc-model-soup), if someone is interested. Model soup always gives a worse score.\n\nI believe the original paper authors provide real results, so I reproduced an example from this [repository](https://github.com/Burf/ModelSoups), and classical soup with ResNet50 works as expected.\nDoes anyone know why model soup performs worse than an individual model in our task or doesn't work for FC layers generally? Is the model size a key issue here? \n\nWould love to hear your thoughts. Thank you in advance",
    "2205034": "In the paper they say:\n\n>While the most straightforward approach to making a model soup is to average all the weights uniformly, we find that **greedy soups**, where models are sequentially added to the soup if they improve accuracy on held-out data, **outperforms uniform averaging**. **Greedy soups avoid adding in models which may lie in a different basin of the error landscape**, which could happen if, for example, models are fine-tuned with high learning rates.\n\nSo it seems to merge different basins of the error landscape negatively impacts the performance of the resulting model. In the paper, they fine-tune from a pre-trained checkpoint and as shown in the paper [Deep Ensembles: A Loss Landscape Perspective](https://arxiv.org/abs/1912.02757) this is very likely to produce the model from the same locally-optimal basin. So I would suggest training the K-Fold not from scratch but from some reasonably trained common checkpoint.",
    "2205236": "Thank you so much. That does make sense to fine-tune the model from the pretrained checkpoint. Regarding greedy soups - I've tried it as well but that also is not working if the model is traiend from scratch.",
    "2207179": "Danylo's explanation was spot on. For me, model souping k-last/ k-best checkpoints from same fold improved val. acc a lot.",
    "2207297": "That is a beautiful idea. Thank you for sharing it. I will try it out",
    "2212818": "!!!\nyou can try this paper\n![https://i.ibb.co/hCFMT1v/Selection-999-1716.png](https://i.ibb.co/hCFMT1v/Selection-999-1716.png)",
    "2229086": "Model soups are for finetuning large pre-trained models. If your model is neither pre-trained nor large making soup will not help. Models trained with different random initialization will lie in different error basins. See Git Re-Basin: Merging Models modulo Permutation Symmetries",
    "2229172": "![https://i.ibb.co/f9WPCQW/Selection-999-1854.png](https://i.ibb.co/f9WPCQW/Selection-999-1854.png)\n\nthis diagram from the paper is interesting.\ncombined data beats all"
  },
  "source": "meta"
}