{
  "id": 430020,
  "title": "254th Place Solution for HuBMAP - Hacking the Human Vasculature Competition (Greedy Model Soup Code)",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/430020",
  "author_name": "Aaryam Sharma",
  "post_date": "2023-08-08T02:35:36.448000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi Everyone! I'm just posting my solution here, in case it helps someone, and I can also learn to write good solutions!</p>\n<p><strong>Competition Context</strong><br>\nThe context of this solution is the recent HuBMAP - Hacking the Human Vasculature Competition:<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview</a></p>\n<p>The data of this competition can be found here:<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></p>\n<p><strong>My Approach</strong><br>\nIn short, I implemented 2 models:</p>\n<ol>\n<li>A plain Mask-RCNN using PyTorch</li>\n<li>A model soup of Mask-RCNNs using PyTorch</li>\n</ol>\n<p>Both scored around the same amount (surprisingly, the model soup scored 0.001 less).</p>\n<p>The models I used for Mask-RCNN were taken from PyTorch.<br>\nTo implement a Model Soup, I implemented a Greedy Soup (Hill-Climbing method).<br>\nEssentially, in this, we start with some model as our original model, and a pool of candidate models. Usually, our first model chosen is our best model. We then pick models from the candidate models. <br>\nThe parameters of the selected models are averaged, and we check validation scores.<br>\nIf the score improves, we keep all the current models for our final model soup. If it doesn't we discard this model from the candidate list.</p>\n<p>For my candidate models, I simply took the models that were saved after each epoch (I had 10 epochs only).<br>\nMy first chosen model was the model from the final epoch.</p>\n<p>My further polished code can be found below:</p>\n<pre><code>greedy_soup = {}\nnum_souped = \nbest_score=  \n\n e, val_loss  (validation_mask_losses):\n    \n    model_chk = \n    (, model_chk)\n    model = get_model((cell_type_dict), model_chk)\n    curr_state_dict = torch.load(model_chk)\n    temp_soup = {}\n    curr_score = \n    (e == ): \n        greedy_soup = {k : v  k, v  curr_state_dict.items()}\n        num_souped = \n        model.load_state_dict(greedy_soup)\n        model.to(device)\n        curr_score = get_score(ds_val, model)\n        best_score = curr_score\n    :\n        \n        temp_soup = {k : (v + greedy_soup[k] * num_souped) / (num_souped + )  k, v  curr_state_dict.items()}\n        \n        model.load_state_dict(temp_soup)\n        model = model.to(device)\n        curr_score = get_score(ds_val, model) \n\n    (e &gt;   curr_score &gt; best_score):\n        \n        ()\n        greedy_soup = temp_soup\n        num_souped += \n        best_score = curr_score\n</code></pre>\n<p>When this competition started, I had little experience in this field. As a result, I didn't participate too much. I knew of Mask-RCNN as a model, and just decided to implement that to improve my knowledge and understanding.</p>\n<p>To go a few steps beyond, I wanted to implement Weighted Box Fusion for making an ensemble, but I couldn't get to it in time.</p>\n<p>I had many ideas and could have tried a lot more. But I got disappointed in my low public leaderboard score. I thought that I could work harder at other competitions. </p>\n<p><em>Validation</em><br>\nDue to my inexperience in larger Kaggle competitions, I did not implement any good sort of validation. I've learnt better now, seeing that validation is one of the main components of any Kaggle competition.</p>\n<p><strong>Further Improvements</strong><br>\nI could've tried larger models, such as OneFormer and MaskFormer. I could've tried YOLOv8 as well. <br>\nI could also try more sophisticated ensembles using Weighted Box Fusion.</p>\n<p>Honestly, I wish this competition was held now, and not earlier. I definitely have a lot more confidence and experience, and would hopefully have tried a lot more!!!</p>\n<p>In any case, I really appreciate any feedback! I'm just a student and I want to learn a lot from all the experts here!</p>\n<p>If you want to further support me, please upvote this notebook!<br>\n(Ok, enough with the YouTuber outro)</p>\n<p>Thank you for reading till here!</p>",
  "messages": [
    {
      "id": 2379121,
      "postDate": "2023-08-08T02:35:36.447Z",
      "content": "<p>Hi Everyone! I'm just posting my solution here, in case it helps someone, and I can also learn to write good solutions!</p>\n<p><strong>Competition Context</strong><br>\nThe context of this solution is the recent HuBMAP - Hacking the Human Vasculature Competition:<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview</a></p>\n<p>The data of this competition can be found here:<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></p>\n<p><strong>My Approach</strong><br>\nIn short, I implemented 2 models:</p>\n<ol>\n<li>A plain Mask-RCNN using PyTorch</li>\n<li>A model soup of Mask-RCNNs using PyTorch</li>\n</ol>\n<p>Both scored around the same amount (surprisingly, the model soup scored 0.001 less).</p>\n<p>The models I used for Mask-RCNN were taken from PyTorch.<br>\nTo implement a Model Soup, I implemented a Greedy Soup (Hill-Climbing method).<br>\nEssentially, in this, we start with some model as our original model, and a pool of candidate models. Usually, our first model chosen is our best model. We then pick models from the candidate models. <br>\nThe parameters of the selected models are averaged, and we check validation scores.<br>\nIf the score improves, we keep all the current models for our final model soup. If it doesn't we discard this model from the candidate list.</p>\n<p>For my candidate models, I simply took the models that were saved after each epoch (I had 10 epochs only).<br>\nMy first chosen model was the model from the final epoch.</p>\n<p>My further polished code can be found below:</p>\n<pre><code>greedy_soup = {}\nnum_souped = \nbest_score=  \n\n e, val_loss  (validation_mask_losses):\n    \n    model_chk = \n    (, model_chk)\n    model = get_model((cell_type_dict), model_chk)\n    curr_state_dict = torch.load(model_chk)\n    temp_soup = {}\n    curr_score = \n    (e == ): \n        greedy_soup = {k : v  k, v  curr_state_dict.items()}\n        num_souped = \n        model.load_state_dict(greedy_soup)\n        model.to(device)\n        curr_score = get_score(ds_val, model)\n        best_score = curr_score\n    :\n        \n        temp_soup = {k : (v + greedy_soup[k] * num_souped) / (num_souped + )  k, v  curr_state_dict.items()}\n        \n        model.load_state_dict(temp_soup)\n        model = model.to(device)\n        curr_score = get_score(ds_val, model) \n\n    (e &gt;   curr_score &gt; best_score):\n        \n        ()\n        greedy_soup = temp_soup\n        num_souped += \n        best_score = curr_score\n</code></pre>\n<p>When this competition started, I had little experience in this field. As a result, I didn't participate too much. I knew of Mask-RCNN as a model, and just decided to implement that to improve my knowledge and understanding.</p>\n<p>To go a few steps beyond, I wanted to implement Weighted Box Fusion for making an ensemble, but I couldn't get to it in time.</p>\n<p>I had many ideas and could have tried a lot more. But I got disappointed in my low public leaderboard score. I thought that I could work harder at other competitions. </p>\n<p><em>Validation</em><br>\nDue to my inexperience in larger Kaggle competitions, I did not implement any good sort of validation. I've learnt better now, seeing that validation is one of the main components of any Kaggle competition.</p>\n<p><strong>Further Improvements</strong><br>\nI could've tried larger models, such as OneFormer and MaskFormer. I could've tried YOLOv8 as well. <br>\nI could also try more sophisticated ensembles using Weighted Box Fusion.</p>\n<p>Honestly, I wish this competition was held now, and not earlier. I definitely have a lot more confidence and experience, and would hopefully have tried a lot more!!!</p>\n<p>In any case, I really appreciate any feedback! I'm just a student and I want to learn a lot from all the experts here!</p>\n<p>If you want to further support me, please upvote this notebook!<br>\n(Ok, enough with the YouTuber outro)</p>\n<p>Thank you for reading till here!</p>",
      "rawMarkdown": "Hi Everyone! I'm just posting my solution here, in case it helps someone, and I can also learn to write good solutions!\n\n**Competition Context**\nThe context of this solution is the recent HuBMAP - Hacking the Human Vasculature Competition:\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\n\nThe data of this competition can be found here:\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n\n**My Approach**\nIn short, I implemented 2 models:\n1. A plain Mask-RCNN using PyTorch\n2. A model soup of Mask-RCNNs using PyTorch\n\nBoth scored around the same amount (surprisingly, the model soup scored 0.001 less).\n\nThe models I used for Mask-RCNN were taken from PyTorch.\nTo implement a Model Soup, I implemented a Greedy Soup (Hill-Climbing method).\nEssentially, in this, we start with some model as our original model, and a pool of candidate models. Usually, our first model chosen is our best model. We then pick models from the candidate models. \nThe parameters of the selected models are averaged, and we check validation scores.\nIf the score improves, we keep all the current models for our final model soup. If it doesn't we discard this model from the candidate list.\n\nFor my candidate models, I simply took the models that were saved after each epoch (I had 10 epochs only).\nMy first chosen model was the model from the final epoch.\n\nMy further polished code can be found below:\n\n```python\ngreedy_soup = {}\nnum_souped = 0\nbest_score= 0 # Track the best validation mAP\n# validation_mask_losses was computed during training and it's not relevant here\nfor e, val_loss in enumerate(validation_mask_losses):\n    # Load models Last Epoch first\n    model_chk = f\"mask-rcnn_{num_epochs - e}.bin\"\n    print(\"Loading:\", model_chk)\n    model = get_model(len(cell_type_dict), model_chk)\n    curr_state_dict = torch.load(model_chk)\n    temp_soup = {}\n    curr_score = 0\n    if(e == 0): # First model\n        greedy_soup = {k : v for k, v in curr_state_dict.items()}\n        num_souped = 1\n        model.load_state_dict(greedy_soup)\n        model.to(device)\n        curr_score = get_score(ds_val, model)\n        best_score = curr_score\n    else:\n        # Make a temporary model soup for testing the next candidate model\n        temp_soup = {k : (v + greedy_soup[k] * num_souped) / (num_souped + 1) for k, v in curr_state_dict.items()}\n        # Load the model with the parameter weights from the soup to be tested\n        model.load_state_dict(temp_soup)\n        model = model.to(device)\n        curr_score = get_score(ds_val, model) # Get validation mAP score\n    \n    if(e > 0 and curr_score > best_score):\n        # If the current model soup's score is better than the previous one, then we save this as our soup.\n        print(f\"Adding model {num_epochs - e}\")\n        greedy_soup = temp_soup\n        num_souped += 1\n        best_score = curr_score\n```\n\n\nWhen this competition started, I had little experience in this field. As a result, I didn't participate too much. I knew of Mask-RCNN as a model, and just decided to implement that to improve my knowledge and understanding.\n\nTo go a few steps beyond, I wanted to implement Weighted Box Fusion for making an ensemble, but I couldn't get to it in time.\n\nI had many ideas and could have tried a lot more. But I got disappointed in my low public leaderboard score. I thought that I could work harder at other competitions. \n\n*Validation*\nDue to my inexperience in larger Kaggle competitions, I did not implement any good sort of validation. I've learnt better now, seeing that validation is one of the main components of any Kaggle competition.\n\n**Further Improvements**\nI could've tried larger models, such as OneFormer and MaskFormer. I could've tried YOLOv8 as well. \nI could also try more sophisticated ensembles using Weighted Box Fusion.\n\nHonestly, I wish this competition was held now, and not earlier. I definitely have a lot more confidence and experience, and would hopefully have tried a lot more!!!\n\nIn any case, I really appreciate any feedback! I'm just a student and I want to learn a lot from all the experts here!\n\nIf you want to further support me, please upvote this notebook!\n(Ok, enough with the YouTuber outro)\n\nThank you for reading till here!",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2379121": "Hi Everyone! I'm just posting my solution here, in case it helps someone, and I can also learn to write good solutions!\n\n**Competition Context**\nThe context of this solution is the recent HuBMAP - Hacking the Human Vasculature Competition:\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\n\nThe data of this competition can be found here:\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n\n**My Approach**\nIn short, I implemented 2 models:\n1. A plain Mask-RCNN using PyTorch\n2. A model soup of Mask-RCNNs using PyTorch\n\nBoth scored around the same amount (surprisingly, the model soup scored 0.001 less).\n\nThe models I used for Mask-RCNN were taken from PyTorch.\nTo implement a Model Soup, I implemented a Greedy Soup (Hill-Climbing method).\nEssentially, in this, we start with some model as our original model, and a pool of candidate models. Usually, our first model chosen is our best model. We then pick models from the candidate models. \nThe parameters of the selected models are averaged, and we check validation scores.\nIf the score improves, we keep all the current models for our final model soup. If it doesn't we discard this model from the candidate list.\n\nFor my candidate models, I simply took the models that were saved after each epoch (I had 10 epochs only).\nMy first chosen model was the model from the final epoch.\n\nMy further polished code can be found below:\n\n```python\ngreedy_soup = {}\nnum_souped = 0\nbest_score= 0 # Track the best validation mAP\n# validation_mask_losses was computed during training and it's not relevant here\nfor e, val_loss in enumerate(validation_mask_losses):\n    # Load models Last Epoch first\n    model_chk = f\"mask-rcnn_{num_epochs - e}.bin\"\n    print(\"Loading:\", model_chk)\n    model = get_model(len(cell_type_dict), model_chk)\n    curr_state_dict = torch.load(model_chk)\n    temp_soup = {}\n    curr_score = 0\n    if(e == 0): # First model\n        greedy_soup = {k : v for k, v in curr_state_dict.items()}\n        num_souped = 1\n        model.load_state_dict(greedy_soup)\n        model.to(device)\n        curr_score = get_score(ds_val, model)\n        best_score = curr_score\n    else:\n        # Make a temporary model soup for testing the next candidate model\n        temp_soup = {k : (v + greedy_soup[k] * num_souped) / (num_souped + 1) for k, v in curr_state_dict.items()}\n        # Load the model with the parameter weights from the soup to be tested\n        model.load_state_dict(temp_soup)\n        model = model.to(device)\n        curr_score = get_score(ds_val, model) # Get validation mAP score\n    \n    if(e > 0 and curr_score > best_score):\n        # If the current model soup's score is better than the previous one, then we save this as our soup.\n        print(f\"Adding model {num_epochs - e}\")\n        greedy_soup = temp_soup\n        num_souped += 1\n        best_score = curr_score\n```\n\n\nWhen this competition started, I had little experience in this field. As a result, I didn't participate too much. I knew of Mask-RCNN as a model, and just decided to implement that to improve my knowledge and understanding.\n\nTo go a few steps beyond, I wanted to implement Weighted Box Fusion for making an ensemble, but I couldn't get to it in time.\n\nI had many ideas and could have tried a lot more. But I got disappointed in my low public leaderboard score. I thought that I could work harder at other competitions. \n\n*Validation*\nDue to my inexperience in larger Kaggle competitions, I did not implement any good sort of validation. I've learnt better now, seeing that validation is one of the main components of any Kaggle competition.\n\n**Further Improvements**\nI could've tried larger models, such as OneFormer and MaskFormer. I could've tried YOLOv8 as well. \nI could also try more sophisticated ensembles using Weighted Box Fusion.\n\nHonestly, I wish this competition was held now, and not earlier. I definitely have a lot more confidence and experience, and would hopefully have tried a lot more!!!\n\nIn any case, I really appreciate any feedback! I'm just a student and I want to learn a lot from all the experts here!\n\nIf you want to further support me, please upvote this notebook!\n(Ok, enough with the YouTuber outro)\n\nThank you for reading till here!"
  }
}