{
  "id": 94790,
  "title": "0.669 Public LB Solution",
  "url": "/competitions/imet-2019-fgvc6/writeups/evgeny-kononenko-0-669-public-lb-solution",
  "author_name": "",
  "post_date": "2019-06-06T22:16:35.322514200Z",
  "votes": 39,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers for an interesting competition, and also to congratulate all the participants who finished in high positions.</p>\n\n<p>Oddly enough, but I would like to thank <a href=\"/bestfitting\">@bestfitting</a>  for his detailed description of his solution for <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913\">Human Protein Atlas Image Classification</a>, which inspired me at the beginning of this competition.</p>\n\n<p><strong>SOLUTION</strong></p>\n\n<p><strong>First model</strong>\nThe first thing that caught my eye in this competition was that the target can be divided into two groups: tags and culture.  It seems that the target is heterogeneous, since culture is responsible for the color spectrum and style, and the tag for the details and content.  It also means that some augmentations can be good for one group of target but bad for another group. The first idea that comes to mind is try to train two models for culture and for tag...</p>\n\n<p>When I traind my first model (with all classes), I simultaneously measured the following metrics:\nF2, F2 for only tags, F2 for only culture.  Then I decided to compare the metrics for the model trained for one group of target.  But as you know multitask learning is a very great thing\n(and the relationship between culture and tag is definitely exists), so I did the following:</p>\n\n<p><strong>1)</strong>  Instead of the last fully connected layer, I made two a fully connected layers, one for the culture and one for the tag.\n<strong>2)</strong> To emphasize the importance of one of group of targets, I balanced the loss as follows:\n<code>final_loss = alpha*culture_loss + (1 - alpha)*tag_loss</code>\nwhere <code>alpha = 0.2</code> for tag model and <code>alpha = 0.8</code> for culture model\n<strong>3)</strong> Then i just concatenate two parts of predictions from two models (i also have tried blend, almost the same result)\n<strong>4)</strong> As I said before, I used different types of augmentations for such models.</p>\n\n<p>I trained SeResNext50 for culture and Senet154 for tags, both models was trained on 288x288 crops. This model really gave a much better result than the standard model for all classes. But there is one drawback - it is time consuming to train 2 models for one prediction. Therefore, for the following models, I used the entire target</p>\n\n<p><strong>Second model</strong>\nThe second model was quite simple. I just trained one more Senet154  on 288x288 crops</p>\n\n<p><strong>Third model</strong>\nThis model was based on <a href=\"https://forums.fast.ai/t/mixup-data-augmentation/22764\">MixUp</a> augmentation. At first I used the same crop size (288x288).  But it seems strange to mix one defective (сrop does not describe the entire target) picture with another defective picture.  So here I switched to resize. (352x352). I trained SeResNext101 here. And it was really good as a solo model and in the blend too.</p>\n\n<p><strong>OTHER SETUP:</strong></p>\n\n<p><strong>Hardware</strong>\n4 x 1080Ti</p>\n\n<p><strong>Validation</strong>\nI used 5 folds CV with Multilabel Iterative Stratification</p>\n\n<p><strong>Loss</strong>\nAnd i mentioned the post from <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913\">Human Protein Atlas Image Classification</a> for a reason. For me, Folcal loss + Lovasz loss was the best choiсe. I used it for all model.</p>\n\n<p><strong>Training procedure</strong>\nI used Adam optimizer.\nIn my case, i got a big improvement of my validation metric when i drop my learning rate, but it works only once during training(\nSo i started from big enough learning rate, then decreased it for five times and continue training.</p>\n\n<p><strong>Data augmentation</strong>\nThe augmentations looked like this (sometimes changed a little):\n<code>Compose([ <br>\n                    OneOf([ <br>\n                    RandomBrightness(limit=0.2, p=1), <br>\n                    RandomContrast(limit=0.2, p=1), <br>\n                    ], p=1),\n                    RandomCrop() or Resize(),\n                    GaussNoise(var_limit=(0, 20), p=0.4),\n                    HorizontalFlip() ])\n</code></p>\n\n<p><strong>Prediction</strong></p>\n\n<p>I used 3tta for the first and the second models:\nIt was simple random crop,  but I made it so that the crops were equally distributed on the picture. I pre-determined the places for crops and only added a random offset. Also i randomly flip the crops horizontally to increase diversity of the models</p>\n\n<p>For the third model i used 2tta: simple horizontal flip</p>\n\n<p><strong>Blend</strong>\nFor blending my models i used simple two-layer fully connected network. It takes concatenated predictions from four my models and gave aggregated prediction. Then i also blend this aggregated prediction with mean-prediction to get final prediction</p>\n\n<p>To be honest, I spent little time on this network, so i think it’s possible to make it better.</p>\n\n<p><strong>OTHER COMMENTS</strong>\nI used batch accumulation (about 300-500 total batch), but did not notice any global improvement.</p>\n\n<p>I decided to not use pseudo labeling, because we already had 100k samples in our train and additional 7k are not critical. Yes, it will increase my score on public LB, but i'm not sure about private (Maybe I'm wrong)</p>\n\n<p>I hope this description was useful for you, thanks!</p>",
  "messages": [
    {
      "id": "546795",
      "postDate": "06/06/2019 22:16:35",
      "content": "<p>First of all, I would like to thank the organizers for an interesting competition, and also to congratulate all the participants who finished in high positions.</p>\n\n<p>Oddly enough, but I would like to thank <a href=\"/bestfitting\">@bestfitting</a>  for his detailed description of his solution for <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913\">Human Protein Atlas Image Classification</a>, which inspired me at the beginning of this competition.</p>\n\n<p><strong>SOLUTION</strong></p>\n\n<p><strong>First model</strong>\nThe first thing that caught my eye in this competition was that the target can be divided into two groups: tags and culture.  It seems that the target is heterogeneous, since culture is responsible for the color spectrum and style, and the tag for the details and content.  It also means that some augmentations can be good for one group of target but bad for another group. The first idea that comes to mind is try to train two models for culture and for tag...</p>\n\n<p>When I traind my first model (with all classes), I simultaneously measured the following metrics:\nF2, F2 for only tags, F2 for only culture.  Then I decided to compare the metrics for the model trained for one group of target.  But as you know multitask learning is a very great thing\n(and the relationship between culture and tag is definitely exists), so I did the following:</p>\n\n<p><strong>1)</strong>  Instead of the last fully connected layer, I made two a fully connected layers, one for the culture and one for the tag.\n<strong>2)</strong> To emphasize the importance of one of group of targets, I balanced the loss as follows:\n<code>final_loss = alpha*culture_loss + (1 - alpha)*tag_loss</code>\nwhere <code>alpha = 0.2</code> for tag model and <code>alpha = 0.8</code> for culture model\n<strong>3)</strong> Then i just concatenate two parts of predictions from two models (i also have tried blend, almost the same result)\n<strong>4)</strong> As I said before, I used different types of augmentations for such models.</p>\n\n<p>I trained SeResNext50 for culture and Senet154 for tags, both models was trained on 288x288 crops. This model really gave a much better result than the standard model for all classes. But there is one drawback - it is time consuming to train 2 models for one prediction. Therefore, for the following models, I used the entire target</p>\n\n<p><strong>Second model</strong>\nThe second model was quite simple. I just trained one more Senet154  on 288x288 crops</p>\n\n<p><strong>Third model</strong>\nThis model was based on <a href=\"https://forums.fast.ai/t/mixup-data-augmentation/22764\">MixUp</a> augmentation. At first I used the same crop size (288x288).  But it seems strange to mix one defective (сrop does not describe the entire target) picture with another defective picture.  So here I switched to resize. (352x352). I trained SeResNext101 here. And it was really good as a solo model and in the blend too.</p>\n\n<p><strong>OTHER SETUP:</strong></p>\n\n<p><strong>Hardware</strong>\n4 x 1080Ti</p>\n\n<p><strong>Validation</strong>\nI used 5 folds CV with Multilabel Iterative Stratification</p>\n\n<p><strong>Loss</strong>\nAnd i mentioned the post from <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913\">Human Protein Atlas Image Classification</a> for a reason. For me, Folcal loss + Lovasz loss was the best choiсe. I used it for all model.</p>\n\n<p><strong>Training procedure</strong>\nI used Adam optimizer.\nIn my case, i got a big improvement of my validation metric when i drop my learning rate, but it works only once during training(\nSo i started from big enough learning rate, then decreased it for five times and continue training.</p>\n\n<p><strong>Data augmentation</strong>\nThe augmentations looked like this (sometimes changed a little):\n<code>Compose([ <br>\n                    OneOf([ <br>\n                    RandomBrightness(limit=0.2, p=1), <br>\n                    RandomContrast(limit=0.2, p=1), <br>\n                    ], p=1),\n                    RandomCrop() or Resize(),\n                    GaussNoise(var_limit=(0, 20), p=0.4),\n                    HorizontalFlip() ])\n</code></p>\n\n<p><strong>Prediction</strong></p>\n\n<p>I used 3tta for the first and the second models:\nIt was simple random crop,  but I made it so that the crops were equally distributed on the picture. I pre-determined the places for crops and only added a random offset. Also i randomly flip the crops horizontally to increase diversity of the models</p>\n\n<p>For the third model i used 2tta: simple horizontal flip</p>\n\n<p><strong>Blend</strong>\nFor blending my models i used simple two-layer fully connected network. It takes concatenated predictions from four my models and gave aggregated prediction. Then i also blend this aggregated prediction with mean-prediction to get final prediction</p>\n\n<p>To be honest, I spent little time on this network, so i think it’s possible to make it better.</p>\n\n<p><strong>OTHER COMMENTS</strong>\nI used batch accumulation (about 300-500 total batch), but did not notice any global improvement.</p>\n\n<p>I decided to not use pseudo labeling, because we already had 100k samples in our train and additional 7k are not critical. Yes, it will increase my score on public LB, but i'm not sure about private (Maybe I'm wrong)</p>\n\n<p>I hope this description was useful for you, thanks!</p>",
      "rawMarkdown": "First of all, I would like to thank the organizers for an interesting competition, and also to congratulate all the participants who finished in high positions.\n\nOddly enough, but I would like to thank @bestfitting  for his detailed description of his solution for [Human Protein Atlas Image Classification](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913), which inspired me at the beginning of this competition.\n\n**SOLUTION**\n\n**First model**\nThe first thing that caught my eye in this competition was that the target can be divided into two groups: tags and culture.  It seems that the target is heterogeneous, since culture is responsible for the color spectrum and style, and the tag for the details and content.  It also means that some augmentations can be good for one group of target but bad for another group. The first idea that comes to mind is try to train two models for culture and for tag...\n\nWhen I traind my first model (with all classes), I simultaneously measured the following metrics:\nF2, F2 for only tags, F2 for only culture.  Then I decided to compare the metrics for the model trained for one group of target.  But as you know multitask learning is a very great thing\n(and the relationship between culture and tag is definitely exists), so I did the following:\n\n**1)**  Instead of the last fully connected layer, I made two a fully connected layers, one for the culture and one for the tag.\n**2)** To emphasize the importance of one of group of targets, I balanced the loss as follows:\n`final_loss = alpha*culture_loss + (1 - alpha)*tag_loss`\nwhere `alpha = 0.2` for tag model and `alpha = 0.8` for culture model\n**3)** Then i just concatenate two parts of predictions from two models (i also have tried blend, almost the same result)\n**4)** As I said before, I used different types of augmentations for such models.\n\nI trained SeResNext50 for culture and Senet154 for tags, both models was trained on 288x288 crops. This model really gave a much better result than the standard model for all classes. But there is one drawback - it is time consuming to train 2 models for one prediction. Therefore, for the following models, I used the entire target\n\n**Second model**\nThe second model was quite simple. I just trained one more Senet154  on 288x288 crops\n\n**Third model**\nThis model was based on [MixUp](https://forums.fast.ai/t/mixup-data-augmentation/22764) augmentation. At first I used the same crop size (288x288).  But it seems strange to mix one defective (сrop does not describe the entire target) picture with another defective picture.  So here I switched to resize. (352x352). I trained SeResNext101 here. And it was really good as a solo model and in the blend too.\n\n**OTHER SETUP:**\n\n**Hardware**\n4 x 1080Ti\n\n**Validation**\nI used 5 folds CV with Multilabel Iterative Stratification\n\n**Loss**\nAnd i mentioned the post from [Human Protein Atlas Image Classification](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913) for a reason. For me, Folcal loss + Lovasz loss was the best choiсe. I used it for all model.\n\n**Training procedure**\nI used Adam optimizer.\nIn my case, i got a big improvement of my validation metric when i drop my learning rate, but it works only once during training(\nSo i started from big enough learning rate, then decreased it for five times and continue training.\n\n**Data augmentation**\nThe augmentations looked like this (sometimes changed a little):\n`    Compose([              \n                    OneOf([    \n                    RandomBrightness(limit=0.2, p=1),      \n                    RandomContrast(limit=0.2, p=1),      \n                    ], p=1),\n                    RandomCrop() or Resize(),\n                    GaussNoise(var_limit=(0, 20), p=0.4),\n                    HorizontalFlip() ])\n`\n\n**Prediction**\n\nI used 3tta for the first and the second models:\nIt was simple random crop,  but I made it so that the crops were equally distributed on the picture. I pre-determined the places for crops and only added a random offset. Also i randomly flip the crops horizontally to increase diversity of the models\n\nFor the third model i used 2tta: simple horizontal flip\n\n**Blend**\nFor blending my models i used simple two-layer fully connected network. It takes concatenated predictions from four my models and gave aggregated prediction. Then i also blend this aggregated prediction with mean-prediction to get final prediction\n\nTo be honest, I spent little time on this network, so i think it’s possible to make it better.\n\n**OTHER COMMENTS**\nI used batch accumulation (about 300-500 total batch), but did not notice any global improvement.\n\nI decided to not use pseudo labeling, because we already had 100k samples in our train and additional 7k are not critical. Yes, it will increase my score on public LB, but i'm not sure about private (Maybe I'm wrong)\n\nI hope this description was useful for you, thanks!",
      "votes": null
    },
    {
      "id": "546863",
      "postDate": "06/07/2019 01:27:33",
      "content": "<p>Thanks for sharing, Evgeny! Could you also share the specific public/validation result for each of your model? : ) </p>",
      "rawMarkdown": "Thanks for sharing, Evgeny! Could you also share the specific public/validation result for each of your model? : )",
      "votes": null
    },
    {
      "id": "546937",
      "postDate": "06/07/2019 03:22:49",
      "content": "<p>Thanks for sharing this cool solution!</p>\n\n<p>Correct me if I'm wrong... It seems to me that you trained two models(model_1, model_2) each with two fully connected layers(fc_culture, fc_tag). When predicting, you then concatenate output of fc_culture from model_1 and output of fc_tag from model_2 as the final prediction?</p>\n\n<p>```\nmodel_1 = se_resnext50() <br>\nmodel_1.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_1.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )</p>\n\n<p>model_2 =Senet154()\nmodel_2.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_2.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )\n```</p>",
      "rawMarkdown": "Thanks for sharing this cool solution!\n\nCorrect me if I'm wrong... It seems to me that you trained two models(model\\_1, model\\_2) each with two fully connected layers(fc\\_culture, fc\\_tag). When predicting, you then concatenate output of fc\\_culture from model\\_1 and output of fc\\_tag from model\\_2 as the final prediction?\n\n```\nmodel_1 = se_resnext50()          \nmodel_1.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_1.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )\n\nmodel_2 =Senet154()\nmodel_2.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_2.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )\n```",
      "votes": null
    },
    {
      "id": "547024",
      "postDate": "06/07/2019 06:50:23",
      "content": "<p>Yes, you are right</p>",
      "rawMarkdown": "Yes, you are right",
      "votes": null
    },
    {
      "id": "547025",
      "postDate": "06/07/2019 06:52:43",
      "content": "<p>(CV / LB)\nFirst model:  0.623 / 0.652 \nSecond model: 0.618 / 0.646\nThird model: 0.618 / 0.648</p>",
      "rawMarkdown": "(CV / LB)\nFirst model:  0.623 / 0.652 \nSecond model: 0.618 / 0.646\nThird model: 0.618 / 0.648",
      "votes": null
    },
    {
      "id": "547059",
      "postDate": "06/07/2019 08:03:02",
      "content": "<p>Thanks for confirming :)</p>",
      "rawMarkdown": "Thanks for confirming :)",
      "votes": null
    },
    {
      "id": "549283",
      "postDate": "06/10/2019 13:59:20",
      "content": "<p>it seems that many top teams seperate tags and culture... Thanks for sharing</p>",
      "rawMarkdown": "it seems that many top teams seperate tags and culture... Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 546863,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "06/07/2019 01:27:33",
      "content": "<p>Thanks for sharing, Evgeny! Could you also share the specific public/validation result for each of your model? : ) </p>",
      "votes": null,
      "replies": [
        {
          "id": 547025,
          "author_name": "lenny27",
          "author_url": "",
          "post_date": "06/07/2019 06:52:43",
          "content": "<p>(CV / LB)\nFirst model:  0.623 / 0.652 \nSecond model: 0.618 / 0.646\nThird model: 0.618 / 0.648</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546937,
      "author_name": "a31314431",
      "author_url": "",
      "post_date": "06/07/2019 03:22:49",
      "content": "<p>Thanks for sharing this cool solution!</p>\n\n<p>Correct me if I'm wrong... It seems to me that you trained two models(model_1, model_2) each with two fully connected layers(fc_culture, fc_tag). When predicting, you then concatenate output of fc_culture from model_1 and output of fc_tag from model_2 as the final prediction?</p>\n\n<p>```\nmodel_1 = se_resnext50() <br>\nmodel_1.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_1.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )</p>\n\n<p>model_2 =Senet154()\nmodel_2.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_2.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )\n```</p>",
      "votes": null,
      "replies": [
        {
          "id": 547024,
          "author_name": "lenny27",
          "author_url": "",
          "post_date": "06/07/2019 06:50:23",
          "content": "<p>Yes, you are right</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547059,
          "author_name": "a31314431",
          "author_url": "",
          "post_date": "06/07/2019 08:03:02",
          "content": "<p>Thanks for confirming :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 549283,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "06/10/2019 13:59:20",
      "content": "<p>it seems that many top teams seperate tags and culture... Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "546795": "First of all, I would like to thank the organizers for an interesting competition, and also to congratulate all the participants who finished in high positions.\n\nOddly enough, but I would like to thank @bestfitting  for his detailed description of his solution for [Human Protein Atlas Image Classification](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913), which inspired me at the beginning of this competition.\n\n**SOLUTION**\n\n**First model**\nThe first thing that caught my eye in this competition was that the target can be divided into two groups: tags and culture.  It seems that the target is heterogeneous, since culture is responsible for the color spectrum and style, and the tag for the details and content.  It also means that some augmentations can be good for one group of target but bad for another group. The first idea that comes to mind is try to train two models for culture and for tag...\n\nWhen I traind my first model (with all classes), I simultaneously measured the following metrics:\nF2, F2 for only tags, F2 for only culture.  Then I decided to compare the metrics for the model trained for one group of target.  But as you know multitask learning is a very great thing\n(and the relationship between culture and tag is definitely exists), so I did the following:\n\n**1)**  Instead of the last fully connected layer, I made two a fully connected layers, one for the culture and one for the tag.\n**2)** To emphasize the importance of one of group of targets, I balanced the loss as follows:\n`final_loss = alpha*culture_loss + (1 - alpha)*tag_loss`\nwhere `alpha = 0.2` for tag model and `alpha = 0.8` for culture model\n**3)** Then i just concatenate two parts of predictions from two models (i also have tried blend, almost the same result)\n**4)** As I said before, I used different types of augmentations for such models.\n\nI trained SeResNext50 for culture and Senet154 for tags, both models was trained on 288x288 crops. This model really gave a much better result than the standard model for all classes. But there is one drawback - it is time consuming to train 2 models for one prediction. Therefore, for the following models, I used the entire target\n\n**Second model**\nThe second model was quite simple. I just trained one more Senet154  on 288x288 crops\n\n**Third model**\nThis model was based on [MixUp](https://forums.fast.ai/t/mixup-data-augmentation/22764) augmentation. At first I used the same crop size (288x288).  But it seems strange to mix one defective (сrop does not describe the entire target) picture with another defective picture.  So here I switched to resize. (352x352). I trained SeResNext101 here. And it was really good as a solo model and in the blend too.\n\n**OTHER SETUP:**\n\n**Hardware**\n4 x 1080Ti\n\n**Validation**\nI used 5 folds CV with Multilabel Iterative Stratification\n\n**Loss**\nAnd i mentioned the post from [Human Protein Atlas Image Classification](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-539913) for a reason. For me, Folcal loss + Lovasz loss was the best choiсe. I used it for all model.\n\n**Training procedure**\nI used Adam optimizer.\nIn my case, i got a big improvement of my validation metric when i drop my learning rate, but it works only once during training(\nSo i started from big enough learning rate, then decreased it for five times and continue training.\n\n**Data augmentation**\nThe augmentations looked like this (sometimes changed a little):\n`    Compose([              \n                    OneOf([    \n                    RandomBrightness(limit=0.2, p=1),      \n                    RandomContrast(limit=0.2, p=1),      \n                    ], p=1),\n                    RandomCrop() or Resize(),\n                    GaussNoise(var_limit=(0, 20), p=0.4),\n                    HorizontalFlip() ])\n`\n\n**Prediction**\n\nI used 3tta for the first and the second models:\nIt was simple random crop,  but I made it so that the crops were equally distributed on the picture. I pre-determined the places for crops and only added a random offset. Also i randomly flip the crops horizontally to increase diversity of the models\n\nFor the third model i used 2tta: simple horizontal flip\n\n**Blend**\nFor blending my models i used simple two-layer fully connected network. It takes concatenated predictions from four my models and gave aggregated prediction. Then i also blend this aggregated prediction with mean-prediction to get final prediction\n\nTo be honest, I spent little time on this network, so i think it’s possible to make it better.\n\n**OTHER COMMENTS**\nI used batch accumulation (about 300-500 total batch), but did not notice any global improvement.\n\nI decided to not use pseudo labeling, because we already had 100k samples in our train and additional 7k are not critical. Yes, it will increase my score on public LB, but i'm not sure about private (Maybe I'm wrong)\n\nI hope this description was useful for you, thanks!",
    "546863": "Thanks for sharing, Evgeny! Could you also share the specific public/validation result for each of your model? : )",
    "546937": "Thanks for sharing this cool solution!\n\nCorrect me if I'm wrong... It seems to me that you trained two models(model\\_1, model\\_2) each with two fully connected layers(fc\\_culture, fc\\_tag). When predicting, you then concatenate output of fc\\_culture from model\\_1 and output of fc\\_tag from model\\_2 as the final prediction?\n\n```\nmodel_1 = se_resnext50()          \nmodel_1.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_1.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )\n\nmodel_2 =Senet154()\nmodel_2.fc_culture = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_culture\n                        )\nmodel_2.fc_tag = torch.nn.Linear(\n                            in_features=2048,\n                            out_features=num_tag\n                        )\n```",
    "547024": "Yes, you are right",
    "547025": "(CV / LB)\nFirst model:  0.623 / 0.652 \nSecond model: 0.618 / 0.646\nThird model: 0.618 / 0.648",
    "547059": "Thanks for confirming :)",
    "549283": "it seems that many top teams seperate tags and culture... Thanks for sharing"
  },
  "source": "meta"
}