{
  "id": 155000,
  "title": "10th place: instant collaboration",
  "url": "/competitions/herbarium-2020-fgvc7/writeups/lilelife-emil-bogomolov-10th-place-instant-collabo",
  "author_name": "",
  "post_date": "2020-05-31T09:50:18.497Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, community! </p>\n\n<p>First of all I want to say thanks to organizers and participants for such interesting competition.</p>\n\n<p>Here I'll present the details of our model. This model is weighted ensemble of predictions of three. First two introduces my teammate, the third one me.</p>\n\n<p>&gt; I am Len(Leheng) Li, i introduce the part of network and my team mate Emil introduce model ensemble.\nOur main work is based on <a href=\"https://arxiv.org/pdf/1912.02413.pdf\">BBN</a>(CVPR2020 oral, <a href=\"https://github.com/Megvii-Nanjing/BBN\">code</a>), it's main contribution is dynamically paying attention to the tailed data. we than replace resnet50 with resnest101 and Efficient net b3.\n<img src=\"https://i.ibb.co/nfYTxVn/1.png\" width=\"800\">\nFor resnest101, we first center-crop 650*650 image and resize to 256*256. It got 0.61094.\nFor Efficient net b3, we first center-crop 650*650 image and resize to 300*300. It got 0.60282.\nI am using 4*2080ti and 1*v100 16g, which is a drop in the bucket for this challenge. I only train my model in 20 epochs. I think performance will higher if I train more times.</p>\n\n<p>&gt; At last, many thanks to my team mate Emil, it’s the first time i collaborate with foreign people(China-Russia). I really enjoy the experience with Emil. And thanks for organizer host this challenge.</p>\n\n<p>And my part:</p>\n\n<p>&gt; ## 1. Model description\nI've tried to squeeze from se_resnext50_32x4d feature generator as much as possible. The main source of ideas was this A Bag of Freebies <a href=\"https://youtu.be/iRrDwzcxg08\">video</a>. There were 32k of classes that hierarchically arranged in superclasses with smaller number of distinct labels. In my model I've pursued them all with different weights before loss. The loss was a weighted FocalLoss from this paper. I've analyzed in detail how much that improved performance, but it seems obvious that one should try to weight such unbalanced data. As info says:</p>\n\n<p>&gt; Each category has at least 1 instance in both the training and test datasets. Note that the test set distribution is slightly different from the training set distribution. The training set contains species with hundreds of examples, but the test set has the number of examples per species capped at a maximum of 10.</p>\n\n<p>&gt; I've collected statistics of mean and variance of whole dataset to normalize input images and that has given a little boost. The next improvement I've introduced was decaying optimizer learning rate. It starts from 30e-5 and divided by two each time every 5 or 10 steps. I've used Adam optimizator. There was an attempt to switch oprimizators during training from Adam to SGD and vice versa. I called this paradigm AdamEva optimizer. The attempt crashed against the constraints of pytorch-lightning framework that allow you to use several optimizers but each optimizer has it's own train call.</p>\n\n<p>&gt; Finally, about augmentations and TTA. Random crop, random flip of herbarium image improves performance but not a Color Jitter, that's my insight. For the TTA I've sampled 4 inputs. I've done 3 times train augmentation than 1 more time just normalization and resize. After that logits of these four samples have averaged to produce output logits.</p>\n\n<p>&gt; ## 2. What also hasn't work\n&gt; * Metric Learning\nI've not figured out how to do metric learning for such a big dataset with such amount of unique labels. I've engineered how to store this amount of data but haven't engineered a way to use it</p>\n\n<p>&gt; * Descriminator Network\nThe main idea is to train a network to predict whether image from train or test set. This might help to create robust validation set as a part of train set. Good validation set gives you an ability to know how good your machine learning model is. Sounds obvious, but it isn't easy to find such when your dataset very disbalanced. By the way, I've used a random subset of a train as val set and it isn't right. In my opinion this is an extreme problem (problem of validation set choosing) to solve in first place in the future!</p>\n\n<p>&gt; ## 3. What went into production (best submissions)\nI've mixed our models with two approaches: averaging of probabilities and averaging of logits. The first hasn't worked well while the second showed better results than both of our networks.</p>",
  "messages": [
    {
      "id": "867969",
      "postDate": "05/30/2020 19:30:12",
      "content": "<p>Hi, community! </p>\n\n<p>First of all I want to say thanks to organizers and participants for such interesting competition.</p>\n\n<p>Here I'll present the details of our model. This model is weighted ensemble of predictions of three. First two introduces my teammate, the third one me.</p>\n\n<p>&gt; I am Len(Leheng) Li, i introduce the part of network and my team mate Emil introduce model ensemble.\nOur main work is based on <a href=\"https://arxiv.org/pdf/1912.02413.pdf\">BBN</a>(CVPR2020 oral, <a href=\"https://github.com/Megvii-Nanjing/BBN\">code</a>), it's main contribution is dynamically paying attention to the tailed data. we than replace resnet50 with resnest101 and Efficient net b3.\n<img src=\"https://i.ibb.co/nfYTxVn/1.png\" width=\"800\">\nFor resnest101, we first center-crop 650*650 image and resize to 256*256. It got 0.61094.\nFor Efficient net b3, we first center-crop 650*650 image and resize to 300*300. It got 0.60282.\nI am using 4*2080ti and 1*v100 16g, which is a drop in the bucket for this challenge. I only train my model in 20 epochs. I think performance will higher if I train more times.</p>\n\n<p>&gt; At last, many thanks to my team mate Emil, it’s the first time i collaborate with foreign people(China-Russia). I really enjoy the experience with Emil. And thanks for organizer host this challenge.</p>\n\n<p>And my part:</p>\n\n<p>&gt; ## 1. Model description\nI've tried to squeeze from se_resnext50_32x4d feature generator as much as possible. The main source of ideas was this A Bag of Freebies <a href=\"https://youtu.be/iRrDwzcxg08\">video</a>. There were 32k of classes that hierarchically arranged in superclasses with smaller number of distinct labels. In my model I've pursued them all with different weights before loss. The loss was a weighted FocalLoss from this paper. I've analyzed in detail how much that improved performance, but it seems obvious that one should try to weight such unbalanced data. As info says:</p>\n\n<p>&gt; Each category has at least 1 instance in both the training and test datasets. Note that the test set distribution is slightly different from the training set distribution. The training set contains species with hundreds of examples, but the test set has the number of examples per species capped at a maximum of 10.</p>\n\n<p>&gt; I've collected statistics of mean and variance of whole dataset to normalize input images and that has given a little boost. The next improvement I've introduced was decaying optimizer learning rate. It starts from 30e-5 and divided by two each time every 5 or 10 steps. I've used Adam optimizator. There was an attempt to switch oprimizators during training from Adam to SGD and vice versa. I called this paradigm AdamEva optimizer. The attempt crashed against the constraints of pytorch-lightning framework that allow you to use several optimizers but each optimizer has it's own train call.</p>\n\n<p>&gt; Finally, about augmentations and TTA. Random crop, random flip of herbarium image improves performance but not a Color Jitter, that's my insight. For the TTA I've sampled 4 inputs. I've done 3 times train augmentation than 1 more time just normalization and resize. After that logits of these four samples have averaged to produce output logits.</p>\n\n<p>&gt; ## 2. What also hasn't work\n&gt; * Metric Learning\nI've not figured out how to do metric learning for such a big dataset with such amount of unique labels. I've engineered how to store this amount of data but haven't engineered a way to use it</p>\n\n<p>&gt; * Descriminator Network\nThe main idea is to train a network to predict whether image from train or test set. This might help to create robust validation set as a part of train set. Good validation set gives you an ability to know how good your machine learning model is. Sounds obvious, but it isn't easy to find such when your dataset very disbalanced. By the way, I've used a random subset of a train as val set and it isn't right. In my opinion this is an extreme problem (problem of validation set choosing) to solve in first place in the future!</p>\n\n<p>&gt; ## 3. What went into production (best submissions)\nI've mixed our models with two approaches: averaging of probabilities and averaging of logits. The first hasn't worked well while the second showed better results than both of our networks.</p>",
      "rawMarkdown": "Hi, community! \n\nFirst of all I want to say thanks to organizers and participants for such interesting competition.\n \nHere I'll present the details of our model. This model is weighted ensemble of predictions of three. First two introduces my teammate, the third one me.\n\n&gt; I am Len(Leheng) Li, i introduce the part of network and my team mate Emil introduce model ensemble.\nOur main work is based on [BBN](https://arxiv.org/pdf/1912.02413.pdf)(CVPR2020 oral, [code](https://github.com/Megvii-Nanjing/BBN)), it's main contribution is dynamically paying attention to the tailed data. we than replace resnet50 with resnest101 and Efficient net b3.\n<img src=\"https://i.ibb.co/nfYTxVn/1.png\" width=\"800\">\nFor resnest101, we first center-crop 650*650 image and resize to 256*256. It got 0.61094.\nFor Efficient net b3, we first center-crop 650*650 image and resize to 300*300. It got 0.60282.\nI am using 4*2080ti and 1*v100 16g, which is a drop in the bucket for this challenge. I only train my model in 20 epochs. I think performance will higher if I train more times.\n\n&gt; At last, many thanks to my team mate Emil, it’s the first time i collaborate with foreign people(China-Russia). I really enjoy the experience with Emil. And thanks for organizer host this challenge.\n\nAnd my part:\n\n&gt; ## 1. Model description\nI've tried to squeeze from se_resnext50_32x4d feature generator as much as possible. The main source of ideas was this A Bag of Freebies [video](https://youtu.be/iRrDwzcxg08). There were 32k of classes that hierarchically arranged in superclasses with smaller number of distinct labels. In my model I've pursued them all with different weights before loss. The loss was a weighted FocalLoss from this paper. I've analyzed in detail how much that improved performance, but it seems obvious that one should try to weight such unbalanced data. As info says:\n\n&gt; Each category has at least 1 instance in both the training and test datasets. Note that the test set distribution is slightly different from the training set distribution. The training set contains species with hundreds of examples, but the test set has the number of examples per species capped at a maximum of 10.\n\n&gt; I've collected statistics of mean and variance of whole dataset to normalize input images and that has given a little boost. The next improvement I've introduced was decaying optimizer learning rate. It starts from 30e-5 and divided by two each time every 5 or 10 steps. I've used Adam optimizator. There was an attempt to switch oprimizators during training from Adam to SGD and vice versa. I called this paradigm AdamEva optimizer. The attempt crashed against the constraints of pytorch-lightning framework that allow you to use several optimizers but each optimizer has it's own train call.\n\n&gt; Finally, about augmentations and TTA. Random crop, random flip of herbarium image improves performance but not a Color Jitter, that's my insight. For the TTA I've sampled 4 inputs. I've done 3 times train augmentation than 1 more time just normalization and resize. After that logits of these four samples have averaged to produce output logits.\n\n&gt; ## 2. What also hasn't work\n&gt; * Metric Learning\nI've not figured out how to do metric learning for such a big dataset with such amount of unique labels. I've engineered how to store this amount of data but haven't engineered a way to use it\n\n&gt; * Descriminator Network\nThe main idea is to train a network to predict whether image from train or test set. This might help to create robust validation set as a part of train set. Good validation set gives you an ability to know how good your machine learning model is. Sounds obvious, but it isn't easy to find such when your dataset very disbalanced. By the way, I've used a random subset of a train as val set and it isn't right. In my opinion this is an extreme problem (problem of validation set choosing) to solve in first place in the future!\n\n&gt; ## 3. What went into production (best submissions)\nI've mixed our models with two approaches: averaging of probabilities and averaging of logits. The first hasn't worked well while the second showed better results than both of our networks.",
      "votes": null
    },
    {
      "id": "867974",
      "postDate": "05/30/2020 19:35:06",
      "content": "<p>Oh, yeah.\nAnd here's the my code \n<a href=\"https://github.com/zetyquickly/herbarium-fgvc7-2020/blob/master/DESCRIPTION.md\">https://github.com/zetyquickly/herbarium-fgvc7-2020/blob/master/DESCRIPTION.md</a></p>",
      "rawMarkdown": "Oh, yeah.\nAnd here's the my code \nhttps://github.com/zetyquickly/herbarium-fgvc7-2020/blob/master/DESCRIPTION.md",
      "votes": null
    },
    {
      "id": "868050",
      "postDate": "05/30/2020 21:04:08",
      "content": "<p>\"A Bag of Freebies video\"  Can you share the link of the video pls</p>",
      "rawMarkdown": "\"A Bag of Freebies video\"  Can you share the link of the video pls",
      "votes": null
    },
    {
      "id": "868421",
      "postDate": "05/31/2020 08:02:42",
      "content": "<p>NICE!</p>",
      "rawMarkdown": "NICE!",
      "votes": null
    },
    {
      "id": "868510",
      "postDate": "05/31/2020 09:51:01",
      "content": "<p>Yes, I've missed it.\nAdded link to the text of my post</p>",
      "rawMarkdown": "Yes, I've missed it.\nAdded link to the text of my post",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 867974,
      "author_name": "discoholic",
      "author_url": "",
      "post_date": "05/30/2020 19:35:06",
      "content": "<p>Oh, yeah.\nAnd here's the my code \n<a href=\"https://github.com/zetyquickly/herbarium-fgvc7-2020/blob/master/DESCRIPTION.md\">https://github.com/zetyquickly/herbarium-fgvc7-2020/blob/master/DESCRIPTION.md</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 868050,
      "author_name": "marcelosanchezortega",
      "author_url": "",
      "post_date": "05/30/2020 21:04:08",
      "content": "<p>\"A Bag of Freebies video\"  Can you share the link of the video pls</p>",
      "votes": null,
      "replies": [
        {
          "id": 868510,
          "author_name": "discoholic",
          "author_url": "",
          "post_date": "05/31/2020 09:51:01",
          "content": "<p>Yes, I've missed it.\nAdded link to the text of my post</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 868421,
      "author_name": "lilelife",
      "author_url": "",
      "post_date": "05/31/2020 08:02:42",
      "content": "<p>NICE!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "867969": "Hi, community! \n\nFirst of all I want to say thanks to organizers and participants for such interesting competition.\n \nHere I'll present the details of our model. This model is weighted ensemble of predictions of three. First two introduces my teammate, the third one me.\n\n&gt; I am Len(Leheng) Li, i introduce the part of network and my team mate Emil introduce model ensemble.\nOur main work is based on [BBN](https://arxiv.org/pdf/1912.02413.pdf)(CVPR2020 oral, [code](https://github.com/Megvii-Nanjing/BBN)), it's main contribution is dynamically paying attention to the tailed data. we than replace resnet50 with resnest101 and Efficient net b3.\n<img src=\"https://i.ibb.co/nfYTxVn/1.png\" width=\"800\">\nFor resnest101, we first center-crop 650*650 image and resize to 256*256. It got 0.61094.\nFor Efficient net b3, we first center-crop 650*650 image and resize to 300*300. It got 0.60282.\nI am using 4*2080ti and 1*v100 16g, which is a drop in the bucket for this challenge. I only train my model in 20 epochs. I think performance will higher if I train more times.\n\n&gt; At last, many thanks to my team mate Emil, it’s the first time i collaborate with foreign people(China-Russia). I really enjoy the experience with Emil. And thanks for organizer host this challenge.\n\nAnd my part:\n\n&gt; ## 1. Model description\nI've tried to squeeze from se_resnext50_32x4d feature generator as much as possible. The main source of ideas was this A Bag of Freebies [video](https://youtu.be/iRrDwzcxg08). There were 32k of classes that hierarchically arranged in superclasses with smaller number of distinct labels. In my model I've pursued them all with different weights before loss. The loss was a weighted FocalLoss from this paper. I've analyzed in detail how much that improved performance, but it seems obvious that one should try to weight such unbalanced data. As info says:\n\n&gt; Each category has at least 1 instance in both the training and test datasets. Note that the test set distribution is slightly different from the training set distribution. The training set contains species with hundreds of examples, but the test set has the number of examples per species capped at a maximum of 10.\n\n&gt; I've collected statistics of mean and variance of whole dataset to normalize input images and that has given a little boost. The next improvement I've introduced was decaying optimizer learning rate. It starts from 30e-5 and divided by two each time every 5 or 10 steps. I've used Adam optimizator. There was an attempt to switch oprimizators during training from Adam to SGD and vice versa. I called this paradigm AdamEva optimizer. The attempt crashed against the constraints of pytorch-lightning framework that allow you to use several optimizers but each optimizer has it's own train call.\n\n&gt; Finally, about augmentations and TTA. Random crop, random flip of herbarium image improves performance but not a Color Jitter, that's my insight. For the TTA I've sampled 4 inputs. I've done 3 times train augmentation than 1 more time just normalization and resize. After that logits of these four samples have averaged to produce output logits.\n\n&gt; ## 2. What also hasn't work\n&gt; * Metric Learning\nI've not figured out how to do metric learning for such a big dataset with such amount of unique labels. I've engineered how to store this amount of data but haven't engineered a way to use it\n\n&gt; * Descriminator Network\nThe main idea is to train a network to predict whether image from train or test set. This might help to create robust validation set as a part of train set. Good validation set gives you an ability to know how good your machine learning model is. Sounds obvious, but it isn't easy to find such when your dataset very disbalanced. By the way, I've used a random subset of a train as val set and it isn't right. In my opinion this is an extreme problem (problem of validation set choosing) to solve in first place in the future!\n\n&gt; ## 3. What went into production (best submissions)\nI've mixed our models with two approaches: averaging of probabilities and averaging of logits. The first hasn't worked well while the second showed better results than both of our networks.",
    "867974": "Oh, yeah.\nAnd here's the my code \nhttps://github.com/zetyquickly/herbarium-fgvc7-2020/blob/master/DESCRIPTION.md",
    "868050": "\"A Bag of Freebies video\"  Can you share the link of the video pls",
    "868421": "NICE!",
    "868510": "Yes, I've missed it.\nAdded link to the text of my post"
  },
  "source": "meta"
}