{
  "id": 238443,
  "title": "12th solutions (0.948/0.949 strageties)",
  "url": "/competitions/hubmap-kidney-segmentation/writeups/12th-solutions-0-948-0-949-strageties",
  "author_name": "",
  "post_date": "2021-05-14T01:29:47.083Z",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We used a simple ensemble of three types of models. (three b7 models<strong>[1]</strong> from 5-folds, one b7 models<strong>[2]</strong> best score from 5-folds using other seeds, four b5 models<strong>[3]</strong> using d488 labels).<br>\nIn this competition, pseudo labels do have big effect on our submissions. Following I will introduce my strategies.</p>\n<p>First, we trained a b7 model using Zhao's labels and LB reached to 0.934. Then I combined this model with models<strong>[2]</strong> and reached to 0.935. I drawed some lost targets depended on the results in d488 until we reached to 0.937, which is the pseudo labels using in models<strong>[3]</strong>.</p>\n<p>According to hosts' comments, I strongly believe the usage of pseudo labels, but in our experiment on public LB, I found using d488 could hurt the prediction in 575. Thus I should deal with it carefully since TN is more fatal than FN. I like it but also fear it, thus,  I used them(models[3]) as part of model ensembles and tune its weight to lessen its influence, divided submission into two parts to reduce marginal mistakes. (overlap = 230, overlap = 32).</p>\n<p>Three models<strong>[1]</strong> are ensembled in the first part. Due to tech limitation, I choose TTA strategies as TTA[0] -&gt; model[0], TTA[1] -&gt; model[1], TTA[2] -&gt; model[2]. Then combine their results using weighted average to make the results of first part more stable.</p>\n<p>Five models (models<strong>[2]</strong>, models<strong>[3]</strong>) ensembled in the second part. I used models<strong>[2]</strong> as a Guider in this part. TTA strategies as TTA[0] -&gt; model[0]-(Guider), TTA[1] -&gt; [model[1], model[2]] , TTA[2] -&gt; [model[3], model[4]]. since the overlap is very small, I use average of the results with the results in the first part.</p>\n<p><strong>Compared to using pseudo with not using, private score could increase about +0.003 ~ +0.004 in our strategies(0.945 - 0.949). The parameters in my models aren't well adjusted therefore the single model didn't get a good score (0.942).</strong></p>\n<p>This is my first competition in CV and the second one in kaggle ,my teammates are busy with graduation with no time in the later stage due to this long-term competition. Very thanks to <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a> and other competitors for their excelent notebooks, which made me familiar with the process of CV competition in few months. Some of my logic and ideas might look funny since the strageties are sometimes by my intuition and guess and I still need to learn and improve in the further life. </p>\n<p>Notebook link: <a href=\"https://www.kaggle.com/puyuzhou/test-xx?scriptVersionId=62471723\" target=\"_blank\">https://www.kaggle.com/puyuzhou/test-xx?scriptVersionId=62471723</a><br>\nAugmentation link:  <a href=\"https://www.kaggle.com/southsakura/kaggle\" target=\"_blank\">https://www.kaggle.com/southsakura/kaggle</a></p>",
  "messages": [
    {
      "id": "1303685",
      "postDate": "05/12/2021 07:39:48",
      "content": "<p>We used a simple ensemble of three types of models. (three b7 models<strong>[1]</strong> from 5-folds, one b7 models<strong>[2]</strong> best score from 5-folds using other seeds, four b5 models<strong>[3]</strong> using d488 labels).<br>\nIn this competition, pseudo labels do have big effect on our submissions. Following I will introduce my strategies.</p>\n<p>First, we trained a b7 model using Zhao's labels and LB reached to 0.934. Then I combined this model with models<strong>[2]</strong> and reached to 0.935. I drawed some lost targets depended on the results in d488 until we reached to 0.937, which is the pseudo labels using in models<strong>[3]</strong>.</p>\n<p>According to hosts' comments, I strongly believe the usage of pseudo labels, but in our experiment on public LB, I found using d488 could hurt the prediction in 575. Thus I should deal with it carefully since TN is more fatal than FN. I like it but also fear it, thus,  I used them(models[3]) as part of model ensembles and tune its weight to lessen its influence, divided submission into two parts to reduce marginal mistakes. (overlap = 230, overlap = 32).</p>\n<p>Three models<strong>[1]</strong> are ensembled in the first part. Due to tech limitation, I choose TTA strategies as TTA[0] -&gt; model[0], TTA[1] -&gt; model[1], TTA[2] -&gt; model[2]. Then combine their results using weighted average to make the results of first part more stable.</p>\n<p>Five models (models<strong>[2]</strong>, models<strong>[3]</strong>) ensembled in the second part. I used models<strong>[2]</strong> as a Guider in this part. TTA strategies as TTA[0] -&gt; model[0]-(Guider), TTA[1] -&gt; [model[1], model[2]] , TTA[2] -&gt; [model[3], model[4]]. since the overlap is very small, I use average of the results with the results in the first part.</p>\n<p><strong>Compared to using pseudo with not using, private score could increase about +0.003 ~ +0.004 in our strategies(0.945 - 0.949). The parameters in my models aren't well adjusted therefore the single model didn't get a good score (0.942).</strong></p>\n<p>This is my first competition in CV and the second one in kaggle ,my teammates are busy with graduation with no time in the later stage due to this long-term competition. Very thanks to <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a> and other competitors for their excelent notebooks, which made me familiar with the process of CV competition in few months. Some of my logic and ideas might look funny since the strageties are sometimes by my intuition and guess and I still need to learn and improve in the further life. </p>\n<p>Notebook link: <a href=\"https://www.kaggle.com/puyuzhou/test-xx?scriptVersionId=62471723\" target=\"_blank\">https://www.kaggle.com/puyuzhou/test-xx?scriptVersionId=62471723</a><br>\nAugmentation link:  <a href=\"https://www.kaggle.com/southsakura/kaggle\" target=\"_blank\">https://www.kaggle.com/southsakura/kaggle</a></p>",
      "rawMarkdown": "We used a simple ensemble of three types of models. (three b7 models**[1]** from 5-folds, one b7 models**[2]** best score from 5-folds using other seeds, four b5 models**[3]** using d488 labels).\nIn this competition, pseudo labels do have big effect on our submissions. Following I will introduce my strategies.\n\nFirst, we trained a b7 model using Zhao's labels and LB reached to 0.934. Then I combined this model with models**[2]** and reached to 0.935. I drawed some lost targets depended on the results in d488 until we reached to 0.937, which is the pseudo labels using in models**[3]**.\n\nAccording to hosts' comments, I strongly believe the usage of pseudo labels, but in our experiment on public LB, I found using d488 could hurt the prediction in 575. Thus I should deal with it carefully since TN is more fatal than FN. I like it but also fear it, thus,  I used them(models[3]) as part of model ensembles and tune its weight to lessen its influence, divided submission into two parts to reduce marginal mistakes. (overlap = 230, overlap = 32).\n\nThree models**[1]** are ensembled in the first part. Due to tech limitation, I choose TTA strategies as TTA[0] -> model[0], TTA[1] -> model[1], TTA[2] -> model[2]. Then combine their results using weighted average to make the results of first part more stable.\n\nFive models (models**[2]**, models**[3]**) ensembled in the second part. I used models**[2]** as a Guider in this part. TTA strategies as TTA[0] -> model[0]-(Guider), TTA[1] -> [model[1], model[2]] , TTA[2] -> [model[3], model[4]]. since the overlap is very small, I use average of the results with the results in the first part.\n\n**Compared to using pseudo with not using, private score could increase about +0.003 ~ +0.004 in our strategies(0.945 - 0.949). The parameters in my models aren't well adjusted therefore the single model didn't get a good score (0.942).**\n\nThis is my first competition in CV and the second one in kaggle ,my teammates are busy with graduation with no time in the later stage due to this long-term competition. Very thanks to @iafoss @wrrosa and other competitors for their excelent notebooks, which made me familiar with the process of CV competition in few months. Some of my logic and ideas might look funny since the strageties are sometimes by my intuition and guess and I still need to learn and improve in the further life. \n\nNotebook link: https://www.kaggle.com/puyuzhou/test-xx?scriptVersionId=62471723\nAugmentation link:  https://www.kaggle.com/southsakura/kaggle",
      "votes": null
    },
    {
      "id": "1305828",
      "postDate": "05/13/2021 13:50:20",
      "content": "<p><a href=\"https://www.kaggle.com/southsakura\" target=\"_blank\">@southsakura</a> Congratulations and Thanks for sharing the approach</p>",
      "rawMarkdown": "southsakura Congratulations and Thanks for sharing the approach",
      "votes": null
    },
    {
      "id": "1306126",
      "postDate": "05/13/2021 16:09:41",
      "content": "<p>You are very welcome!</p>",
      "rawMarkdown": "You are very welcome!",
      "votes": null
    },
    {
      "id": "1907424",
      "postDate": "08/20/2022 19:16:58",
      "content": "<p>Hi, could you please share your code for training if it is possible?</p>",
      "rawMarkdown": "Hi, could you please share your code for training if it is possible?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1305828,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "05/13/2021 13:50:20",
      "content": "<p><a href=\"https://www.kaggle.com/southsakura\" target=\"_blank\">@southsakura</a> Congratulations and Thanks for sharing the approach</p>",
      "votes": null,
      "replies": [
        {
          "id": 1306126,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "05/13/2021 16:09:41",
          "content": "<p>You are very welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1907424,
      "author_name": "nurkhanlaiyk",
      "author_url": "",
      "post_date": "08/20/2022 19:16:58",
      "content": "<p>Hi, could you please share your code for training if it is possible?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1303685": "We used a simple ensemble of three types of models. (three b7 models**[1]** from 5-folds, one b7 models**[2]** best score from 5-folds using other seeds, four b5 models**[3]** using d488 labels).\nIn this competition, pseudo labels do have big effect on our submissions. Following I will introduce my strategies.\n\nFirst, we trained a b7 model using Zhao's labels and LB reached to 0.934. Then I combined this model with models**[2]** and reached to 0.935. I drawed some lost targets depended on the results in d488 until we reached to 0.937, which is the pseudo labels using in models**[3]**.\n\nAccording to hosts' comments, I strongly believe the usage of pseudo labels, but in our experiment on public LB, I found using d488 could hurt the prediction in 575. Thus I should deal with it carefully since TN is more fatal than FN. I like it but also fear it, thus,  I used them(models[3]) as part of model ensembles and tune its weight to lessen its influence, divided submission into two parts to reduce marginal mistakes. (overlap = 230, overlap = 32).\n\nThree models**[1]** are ensembled in the first part. Due to tech limitation, I choose TTA strategies as TTA[0] -> model[0], TTA[1] -> model[1], TTA[2] -> model[2]. Then combine their results using weighted average to make the results of first part more stable.\n\nFive models (models**[2]**, models**[3]**) ensembled in the second part. I used models**[2]** as a Guider in this part. TTA strategies as TTA[0] -> model[0]-(Guider), TTA[1] -> [model[1], model[2]] , TTA[2] -> [model[3], model[4]]. since the overlap is very small, I use average of the results with the results in the first part.\n\n**Compared to using pseudo with not using, private score could increase about +0.003 ~ +0.004 in our strategies(0.945 - 0.949). The parameters in my models aren't well adjusted therefore the single model didn't get a good score (0.942).**\n\nThis is my first competition in CV and the second one in kaggle ,my teammates are busy with graduation with no time in the later stage due to this long-term competition. Very thanks to @iafoss @wrrosa and other competitors for their excelent notebooks, which made me familiar with the process of CV competition in few months. Some of my logic and ideas might look funny since the strageties are sometimes by my intuition and guess and I still need to learn and improve in the further life. \n\nNotebook link: https://www.kaggle.com/puyuzhou/test-xx?scriptVersionId=62471723\nAugmentation link:  https://www.kaggle.com/southsakura/kaggle",
    "1305828": "southsakura Congratulations and Thanks for sharing the approach",
    "1306126": "You are very welcome!",
    "1907424": "Hi, could you please share your code for training if it is possible?"
  },
  "source": "meta"
}