{
  "id": 94687,
  "title": "Solution | Private Top 1",
  "url": "/competitions/imet-2019-fgvc6/writeups/ods-ai-konstantin-gavrilchik-solution-private-top-",
  "author_name": "",
  "post_date": "2019-06-10T22:14:14.573Z",
  "votes": 155,
  "comment_count": 64,
  "views": 0,
  "content": "<p>Congrats to all who finished in the gold zone. Now, while we are waiting for second stage results let me share my solution part. I do not expect shake up, so I think leaderboard will be stay the same.</p>\n\n<p>First of all, what is the main challenge in this competition? Of course it is noisy data and we should find a way to work with it. \nSo, i can divide my solution into several stages:</p>\n\n<p>Hardware: 6x 1080 ti with 36 cores and 120 gb RAM in total.</p>\n\n<p><strong>Stage 0. The same part for all stages:</strong>\nModels: SENet154, PNasNet-5, SE-ResNext101 (all pretrained from cadene repo)\nCV: 5 folds with Multilabel Iterative Stratification\n<em>Data Augmentations:</em>\n<code>\n        HorizontalFlip(p=0.5),\n        OneOf([\n            RandomBrightness(0.1, p=1),\n            RandomContrast(0.1, p=1),\n        ], p=0.3),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.0, rotate_limit=15, p=0.3),\n        IAAAdditiveGaussianNoise(p=0.3),\n</code>\n<em>Data Preprocessing:</em>\nThis part was a great finding which speed up convergence and increased score very good. So, analysis of different models with Crops(224), Crops(288), Crops(300) and so on shows that it influence on tags a lots. Actually, lets imagine that you have picture 500x300 and this image labeled with tag «person» if there are persons somewhere at this image. There is a not so low probability that you crop it =&gt; you data becomes more and more noisy. \nSo, I decided to use CropIfNedeed + Resize. The main Idea is perform crop if it possible (for example crop 600x600 from 300x500 image — 300x500. from 300x900 — 300x600). Code below:</p>\n\n<p>```\nclass RandomCropIfNeeded(RandomCrop):\n    def <strong>init</strong>(self, height, width, always_apply=False, p=1.0):\n        super(RandomCrop, self).<strong>init</strong>(always_apply, p)\n        self.height = height\n        self.width = width</p>\n\n<pre><code>def apply(self, img, h_start=0, w_start=0, **params):\n    h, w, _ = img.shape\n    return F.random_crop(img, min(self.height, h), min(self.width, w), h_start, w_start)\n</code></pre>\n\n<p><code>\nSo, I used the following couple:\n</code>\nRandomCropIfNeeded(SIZE * 2, SIZE * 2),\nResize(SIZE, SIZE)\n```\nWith SIZE = 320 for SEResNext101 / SENet154 and SIZE = 331 for PNasNet-5</p>\n\n<p><em>TTA</em>: Using the previous approach for data preprocessing we can just use TTA2 (original + hflip image)</p>\n\n<p><em>Scheduler</em>: manual, the correct scheduler increased accuracy a lot. I found it by analyzing plots with metrics / losses. \nI started with LR: 0.005 then train 15 epochs, drop it by 5 times and train some epochs more.</p>\n\n<p><strong>Stage 1. Training the zoo:</strong></p>\n\n<p><em>Loss</em>: Focal\n<em>Sampling</em> with logarithmic weights\n<em>Batch size</em>: 1000-1500 (accumulation 10-20 times for different models)</p>\n\n<p><strong>Stage 2. Filtering predictions:</strong>\nDrop images from train with very high error between OOF predictions and labels (I consider them as very noisy and incorrect)</p>\n\n<p><em>Loss</em>: Focal\n+ Hard negative mining (Sample 5% of hardest samples each epoch)</p>\n\n<p>Re-train models from scratch.</p>\n\n<p><strong>Stage 3. Pseudo labeling:</strong></p>\n\n<p><em>Loss</em>: Focal</p>\n\n<p>The simplest version of pseudo labeling, add the most confident predictions (highest  np.mean(np.abs(probabilities - 0.5)) ) to the training dataset.</p>\n\n<p>Re-train models from scratch.</p>\n\n<p><strong>Stage 4. Culture and tags separately:</strong></p>\n\n<p>I found 2 main things about cultures and tags:\n1. tags are less noisy than cultures (and starting from some epoch the culture accuracy did not increased very much)\n2. some tags classes are very similar to ImageNet classes</p>\n\n<p>So, using already trained weights as pretrain I continue training model for 705 classes (tags only) </p>\n\n<p><em>Loss: Focal</em> -&gt; BCE (yes, here I switched the loss, I helped too)</p>\n\n<p><strong>Stage 5. Second-level model</strong></p>\n\n<p>I construct the binary classification dataset: I took 1103 (number of classes) rows per each image and trying to predict that this class relates to this image (0 or 1). So it means the length of my train data becomes <code>len(data) * 1103</code>.</p>\n\n<p>I extract next features:\n- probabilities of each models, sum / division / multiplication of each pair / triple / .. of models \n- mean / median / std / max / min of each channel\n- brightness / colorness of each image (you can say me that NN can easily detect it — yes, but here i can do it without cropping and resizing — it is less noisy)\n- Max side size and binary flag — height more than width or no (it is a little bit better for tree boosting than just height + width in case of lower side == 300)\n- Aaaaand the secret sauce: ImageNet predictions ;) As I already mentioned — some tags classes similar to ImageNet classes, but ImageNet much bigger, pretrained models much more generalized. So, I add all 1000 (number of ImageNet classes) predictions to this dataset</p>\n\n<p>So, then I trained LightGBM on all this data.</p>\n\n<p><strong>Hints / Postprocessing:</strong>\n- Different threshold for cultures and tags models. \n- EDA shows that tags are fully labeled and cultures may be not. So, I binarize predictions using the following code:</p>\n\n<p><code>\nculture_predictions = binarize_predictions(predictions[:, :398], threshold=0.28, min_samples=0, max_samples=3)\ntags_predictions = binarize_predictions(predictions[:, 398:], threshold=0.1, min_samples=1, max_samples=8)\npredictions = np.hstack([culture_predictions, tags_predictions])\n</code></p>\n\n<p>(thank you to <a href=\"/lopuhin\">@lopuhin</a> for the binarize_predictions code in his kernel)</p>\n\n<p>Total training time: ~10-15 days (that's why I submitted not very often ;) )</p>\n\n<p>P.S. I did not submit all this ensemble to the private stage… Due to I faced with kernel limit time :facepalm: (for example I used LGBM for pseudo-labeling but did not use it at the next cycle of test predictions)</p>",
  "messages": [
    {
      "id": "546074",
      "postDate": "06/06/2019 08:22:16",
      "content": "<p>Congrats to all who finished in the gold zone. Now, while we are waiting for second stage results let me share my solution part. I do not expect shake up, so I think leaderboard will be stay the same.</p>\n\n<p>First of all, what is the main challenge in this competition? Of course it is noisy data and we should find a way to work with it. \nSo, i can divide my solution into several stages:</p>\n\n<p>Hardware: 6x 1080 ti with 36 cores and 120 gb RAM in total.</p>\n\n<p><strong>Stage 0. The same part for all stages:</strong>\nModels: SENet154, PNasNet-5, SE-ResNext101 (all pretrained from cadene repo)\nCV: 5 folds with Multilabel Iterative Stratification\n<em>Data Augmentations:</em>\n<code>\n        HorizontalFlip(p=0.5),\n        OneOf([\n            RandomBrightness(0.1, p=1),\n            RandomContrast(0.1, p=1),\n        ], p=0.3),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.0, rotate_limit=15, p=0.3),\n        IAAAdditiveGaussianNoise(p=0.3),\n</code>\n<em>Data Preprocessing:</em>\nThis part was a great finding which speed up convergence and increased score very good. So, analysis of different models with Crops(224), Crops(288), Crops(300) and so on shows that it influence on tags a lots. Actually, lets imagine that you have picture 500x300 and this image labeled with tag «person» if there are persons somewhere at this image. There is a not so low probability that you crop it =&gt; you data becomes more and more noisy. \nSo, I decided to use CropIfNedeed + Resize. The main Idea is perform crop if it possible (for example crop 600x600 from 300x500 image — 300x500. from 300x900 — 300x600). Code below:</p>\n\n<p>```\nclass RandomCropIfNeeded(RandomCrop):\n    def <strong>init</strong>(self, height, width, always_apply=False, p=1.0):\n        super(RandomCrop, self).<strong>init</strong>(always_apply, p)\n        self.height = height\n        self.width = width</p>\n\n<pre><code>def apply(self, img, h_start=0, w_start=0, **params):\n    h, w, _ = img.shape\n    return F.random_crop(img, min(self.height, h), min(self.width, w), h_start, w_start)\n</code></pre>\n\n<p><code>\nSo, I used the following couple:\n</code>\nRandomCropIfNeeded(SIZE * 2, SIZE * 2),\nResize(SIZE, SIZE)\n```\nWith SIZE = 320 for SEResNext101 / SENet154 and SIZE = 331 for PNasNet-5</p>\n\n<p><em>TTA</em>: Using the previous approach for data preprocessing we can just use TTA2 (original + hflip image)</p>\n\n<p><em>Scheduler</em>: manual, the correct scheduler increased accuracy a lot. I found it by analyzing plots with metrics / losses. \nI started with LR: 0.005 then train 15 epochs, drop it by 5 times and train some epochs more.</p>\n\n<p><strong>Stage 1. Training the zoo:</strong></p>\n\n<p><em>Loss</em>: Focal\n<em>Sampling</em> with logarithmic weights\n<em>Batch size</em>: 1000-1500 (accumulation 10-20 times for different models)</p>\n\n<p><strong>Stage 2. Filtering predictions:</strong>\nDrop images from train with very high error between OOF predictions and labels (I consider them as very noisy and incorrect)</p>\n\n<p><em>Loss</em>: Focal\n+ Hard negative mining (Sample 5% of hardest samples each epoch)</p>\n\n<p>Re-train models from scratch.</p>\n\n<p><strong>Stage 3. Pseudo labeling:</strong></p>\n\n<p><em>Loss</em>: Focal</p>\n\n<p>The simplest version of pseudo labeling, add the most confident predictions (highest  np.mean(np.abs(probabilities - 0.5)) ) to the training dataset.</p>\n\n<p>Re-train models from scratch.</p>\n\n<p><strong>Stage 4. Culture and tags separately:</strong></p>\n\n<p>I found 2 main things about cultures and tags:\n1. tags are less noisy than cultures (and starting from some epoch the culture accuracy did not increased very much)\n2. some tags classes are very similar to ImageNet classes</p>\n\n<p>So, using already trained weights as pretrain I continue training model for 705 classes (tags only) </p>\n\n<p><em>Loss: Focal</em> -&gt; BCE (yes, here I switched the loss, I helped too)</p>\n\n<p><strong>Stage 5. Second-level model</strong></p>\n\n<p>I construct the binary classification dataset: I took 1103 (number of classes) rows per each image and trying to predict that this class relates to this image (0 or 1). So it means the length of my train data becomes <code>len(data) * 1103</code>.</p>\n\n<p>I extract next features:\n- probabilities of each models, sum / division / multiplication of each pair / triple / .. of models \n- mean / median / std / max / min of each channel\n- brightness / colorness of each image (you can say me that NN can easily detect it — yes, but here i can do it without cropping and resizing — it is less noisy)\n- Max side size and binary flag — height more than width or no (it is a little bit better for tree boosting than just height + width in case of lower side == 300)\n- Aaaaand the secret sauce: ImageNet predictions ;) As I already mentioned — some tags classes similar to ImageNet classes, but ImageNet much bigger, pretrained models much more generalized. So, I add all 1000 (number of ImageNet classes) predictions to this dataset</p>\n\n<p>So, then I trained LightGBM on all this data.</p>\n\n<p><strong>Hints / Postprocessing:</strong>\n- Different threshold for cultures and tags models. \n- EDA shows that tags are fully labeled and cultures may be not. So, I binarize predictions using the following code:</p>\n\n<p><code>\nculture_predictions = binarize_predictions(predictions[:, :398], threshold=0.28, min_samples=0, max_samples=3)\ntags_predictions = binarize_predictions(predictions[:, 398:], threshold=0.1, min_samples=1, max_samples=8)\npredictions = np.hstack([culture_predictions, tags_predictions])\n</code></p>\n\n<p>(thank you to <a href=\"/lopuhin\">@lopuhin</a> for the binarize_predictions code in his kernel)</p>\n\n<p>Total training time: ~10-15 days (that's why I submitted not very often ;) )</p>\n\n<p>P.S. I did not submit all this ensemble to the private stage… Due to I faced with kernel limit time :facepalm: (for example I used LGBM for pseudo-labeling but did not use it at the next cycle of test predictions)</p>",
      "rawMarkdown": "Congrats to all who finished in the gold zone. Now, while we are waiting for second stage results let me share my solution part. I do not expect shake up, so I think leaderboard will be stay the same.\n\nFirst of all, what is the main challenge in this competition? Of course it is noisy data and we should find a way to work with it. \nSo, i can divide my solution into several stages:\n\nHardware: 6x 1080 ti with 36 cores and 120 gb RAM in total.\n\n**Stage 0. The same part for all stages:**\nModels: SENet154, PNasNet-5, SE-ResNext101 (all pretrained from cadene repo)\nCV: 5 folds with Multilabel Iterative Stratification\n*Data Augmentations:*\n```\n        HorizontalFlip(p=0.5),\n        OneOf([\n            RandomBrightness(0.1, p=1),\n            RandomContrast(0.1, p=1),\n        ], p=0.3),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.0, rotate_limit=15, p=0.3),\n        IAAAdditiveGaussianNoise(p=0.3),\n```\n*Data Preprocessing:*\nThis part was a great finding which speed up convergence and increased score very good. So, analysis of different models with Crops(224), Crops(288), Crops(300) and so on shows that it influence on tags a lots. Actually, lets imagine that you have picture 500x300 and this image labeled with tag «person» if there are persons somewhere at this image. There is a not so low probability that you crop it =&gt; you data becomes more and more noisy. \nSo, I decided to use CropIfNedeed + Resize. The main Idea is perform crop if it possible (for example crop 600x600 from 300x500 image — 300x500. from 300x900 — 300x600). Code below:\n\n```\nclass RandomCropIfNeeded(RandomCrop):\n    def __init__(self, height, width, always_apply=False, p=1.0):\n        super(RandomCrop, self).__init__(always_apply, p)\n        self.height = height\n        self.width = width\n\n    def apply(self, img, h_start=0, w_start=0, **params):\n        h, w, _ = img.shape\n        return F.random_crop(img, min(self.height, h), min(self.width, w), h_start, w_start)\n``` \nSo, I used the following couple:\n```\nRandomCropIfNeeded(SIZE * 2, SIZE * 2),\nResize(SIZE, SIZE)\n```\nWith SIZE = 320 for SEResNext101 / SENet154 and SIZE = 331 for PNasNet-5\n\n*TTA*: Using the previous approach for data preprocessing we can just use TTA2 (original + hflip image)\n\n*Scheduler*: manual, the correct scheduler increased accuracy a lot. I found it by analyzing plots with metrics / losses. \nI started with LR: 0.005 then train 15 epochs, drop it by 5 times and train some epochs more.\n\n**Stage 1. Training the zoo:**\n\n*Loss*: Focal\n*Sampling* with logarithmic weights\n*Batch size*: 1000-1500 (accumulation 10-20 times for different models)\n\n**Stage 2. Filtering predictions:**\nDrop images from train with very high error between OOF predictions and labels (I consider them as very noisy and incorrect)\n\n*Loss*: Focal\n+ Hard negative mining (Sample 5% of hardest samples each epoch)\n\n\nRe-train models from scratch.\n\n**Stage 3. Pseudo labeling:**\n\n*Loss*: Focal\n\nThe simplest version of pseudo labeling, add the most confident predictions (highest  np.mean(np.abs(probabilities - 0.5)) ) to the training dataset.\n\nRe-train models from scratch.\n\n**Stage 4. Culture and tags separately:**\n\nI found 2 main things about cultures and tags:\n1. tags are less noisy than cultures (and starting from some epoch the culture accuracy did not increased very much)\n2. some tags classes are very similar to ImageNet classes\n\nSo, using already trained weights as pretrain I continue training model for 705 classes (tags only) \n\n*Loss: Focal* -&gt; BCE (yes, here I switched the loss, I helped too)\n\n**Stage 5. Second-level model**\n\nI construct the binary classification dataset: I took 1103 (number of classes) rows per each image and trying to predict that this class relates to this image (0 or 1). So it means the length of my train data becomes `len(data) * 1103`.\n\nI extract next features:\n- probabilities of each models, sum / division / multiplication of each pair / triple / .. of models \n- mean / median / std / max / min of each channel\n- brightness / colorness of each image (you can say me that NN can easily detect it — yes, but here i can do it without cropping and resizing — it is less noisy)\n- Max side size and binary flag — height more than width or no (it is a little bit better for tree boosting than just height + width in case of lower side == 300)\n- Aaaaand the secret sauce: ImageNet predictions ;) As I already mentioned — some tags classes similar to ImageNet classes, but ImageNet much bigger, pretrained models much more generalized. So, I add all 1000 (number of ImageNet classes) predictions to this dataset\n\nSo, then I trained LightGBM on all this data.\n\n**Hints / Postprocessing:**\n- Different threshold for cultures and tags models. \n- EDA shows that tags are fully labeled and cultures may be not. So, I binarize predictions using the following code:\n\n``` \nculture_predictions = binarize_predictions(predictions[:, :398], threshold=0.28, min_samples=0, max_samples=3)\ntags_predictions = binarize_predictions(predictions[:, 398:], threshold=0.1, min_samples=1, max_samples=8)\npredictions = np.hstack([culture_predictions, tags_predictions])\n```\n\n(thank you to @lopuhin for the binarize_predictions code in his kernel)\n\nTotal training time: ~10-15 days (that's why I submitted not very often ;) )\n\n\nP.S. I did not submit all this ensemble to the private stage… Due to I faced with kernel limit time :facepalm: (for example I used LGBM for pseudo-labeling but did not use it at the next cycle of test predictions)",
      "votes": null
    },
    {
      "id": "546090",
      "postDate": "06/06/2019 08:46:05",
      "content": "<p>Well done!</p>",
      "rawMarkdown": "Well done!",
      "votes": null
    },
    {
      "id": "546097",
      "postDate": "06/06/2019 08:54:43",
      "content": "<p>Thanks for sharing! I have learned a lot from this discussion.</p>",
      "rawMarkdown": "Thanks for sharing! I have learned a lot from this discussion.",
      "votes": null
    },
    {
      "id": "546098",
      "postDate": "06/06/2019 08:54:43",
      "content": "<p>Amazing!</p>",
      "rawMarkdown": "Amazing!",
      "votes": null
    },
    {
      "id": "546099",
      "postDate": "06/06/2019 08:55:34",
      "content": "<p>Congratulations <a href=\"/dempton\">@dempton</a> ! And thanks a lot for sharing!</p>\n\n<p>Did you use the fastai library?</p>\n\n<p>Regarding the network structure, did you use any GAPnet or altered the structure for the feature extraction for any network?</p>\n\n<p>Did you use the original network heads? Appart from changing the number of output classes :-) and the multitask learner</p>",
      "rawMarkdown": "Congratulations @dempton ! And thanks a lot for sharing!\n\nDid you use the fastai library?\n\nRegarding the network structure, did you use any GAPnet or altered the structure for the feature extraction for any network?\n\nDid you use the original network heads? Appart from changing the number of output classes :-) and the multitask learner",
      "votes": null
    },
    {
      "id": "546101",
      "postDate": "06/06/2019 08:56:37",
      "content": "<p>Thanks for sharing, a nice solution.\nCould you tell me which point is the key point and How much LB has been improved?</p>",
      "rawMarkdown": "Thanks for sharing, a nice solution.\nCould you tell me which point is the key point and How much LB has been improved?",
      "votes": null
    },
    {
      "id": "546104",
      "postDate": "06/06/2019 08:58:35",
      "content": "<p>Thanks, really learn a lot.</p>",
      "rawMarkdown": "Thanks, really learn a lot.",
      "votes": null
    },
    {
      "id": "546106",
      "postDate": "06/06/2019 08:59:40",
      "content": "<p>No, I used just pure pytorch as a DL framework\nNo, I did not use GAPNet, Just resnet, DenseNet, NasNet like architectures </p>",
      "rawMarkdown": "No, I used just pure pytorch as a DL framework\nNo, I did not use GAPNet, Just resnet, DenseNet, NasNet like architectures",
      "votes": null
    },
    {
      "id": "546107",
      "postDate": "06/06/2019 09:03:53",
      "content": "<p>It is very hard to say how much each step influenced on the LB because I did not make a lot of submissions (validation was very good).\nAccording to validation:\n- filtering (~0.005 - 0.01, do not remember exactly)\n- pseudo labeling (it increased (0.01)\n- second level (it increased 0.005 - 0.007)</p>",
      "rawMarkdown": "It is very hard to say how much each step influenced on the LB because I did not make a lot of submissions (validation was very good).\nAccording to validation:\n- filtering (~0.005 - 0.01, do not remember exactly)\n- pseudo labeling (it increased (0.01)\n- second level (it increased 0.005 - 0.007)",
      "votes": null
    },
    {
      "id": "546109",
      "postDate": "06/06/2019 09:05:46",
      "content": "<p>But anyway the most important for this competitions was a correct LR scheduler (which I set manually for each epoch) + preprocessing (CropIfNeeded + Resize)\nPreprocessing speed up convergence which helped me to perform a lot of experiments</p>",
      "rawMarkdown": "But anyway the most important for this competitions was a correct LR scheduler (which I set manually for each epoch) + preprocessing (CropIfNeeded + Resize)\nPreprocessing speed up convergence which helped me to perform a lot of experiments",
      "votes": null
    },
    {
      "id": "546110",
      "postDate": "06/06/2019 09:09:02",
      "content": "<p>Thanks, you deserve the decent result! By the way, how do you use 1000+ batch size with 6 1080ti (66GB GPU memory)</p>",
      "rawMarkdown": "Thanks, you deserve the decent result! By the way, how do you use 1000+ batch size with 6 1080ti (66GB GPU memory)",
      "votes": null
    },
    {
      "id": "546111",
      "postDate": "06/06/2019 09:09:26",
      "content": "<p>Good job ! ! </p>",
      "rawMarkdown": "Good job ! !",
      "votes": null
    },
    {
      "id": "546112",
      "postDate": "06/06/2019 09:10:06",
      "content": "<p>Got it, thank you.</p>",
      "rawMarkdown": "Got it, thank you.",
      "votes": null
    },
    {
      "id": "546114",
      "postDate": "06/06/2019 09:11:24",
      "content": "<p>Batch size accumulation\n(you can perform optimizer step each N batches)</p>",
      "rawMarkdown": "Batch size accumulation\n(you can perform optimizer step each N batches)",
      "votes": null
    },
    {
      "id": "546118",
      "postDate": "06/06/2019 09:14:59",
      "content": "<p>I see</p>",
      "rawMarkdown": "I see",
      "votes": null
    },
    {
      "id": "546120",
      "postDate": "06/06/2019 09:17:51",
      "content": "<p>you deserve the first place! </p>",
      "rawMarkdown": "you deserve the first place!",
      "votes": null
    },
    {
      "id": "546122",
      "postDate": "06/06/2019 09:18:52",
      "content": "<p>And, one more question, why did you use the large batch size(1k-1.5k). Is that better than 32/64/128?</p>",
      "rawMarkdown": "And, one more question, why did you use the large batch size(1k-1.5k). Is that better than 32/64/128?",
      "votes": null
    },
    {
      "id": "546125",
      "postDate": "06/06/2019 09:21:31",
      "content": "<p>good work! will you Open source code？</p>",
      "rawMarkdown": "good work! will you Open source code？",
      "votes": null
    },
    {
      "id": "546126",
      "postDate": "06/06/2019 09:22:26",
      "content": "<p>yep, I found that moving from 64 to 64*4 increased my performance very good\nSo, I decided choose batch size ~ 64 * 20 (it increased a little bit more and I stopped tuning it)</p>",
      "rawMarkdown": "yep, I found that moving from 64 to 64*4 increased my performance very good\nSo, I decided choose batch size ~ 64 * 20 (it increased a little bit more and I stopped tuning it)",
      "votes": null
    },
    {
      "id": "546127",
      "postDate": "06/06/2019 09:23:16",
      "content": "<p>Probably after stage 2</p>",
      "rawMarkdown": "Probably after stage 2",
      "votes": null
    },
    {
      "id": "546141",
      "postDate": "06/06/2019 09:42:23",
      "content": "<p>Congrats and thank you for the detailed information! I think it's hard to surpass your work in this competition and you did it solo!</p>",
      "rawMarkdown": "Congrats and thank you for the detailed information! I think it's hard to surpass your work in this competition and you did it solo!",
      "votes": null
    },
    {
      "id": "546159",
      "postDate": "06/06/2019 10:01:18",
      "content": "<p>Really great solution! </p>",
      "rawMarkdown": "Really great solution!",
      "votes": null
    },
    {
      "id": "546166",
      "postDate": "06/06/2019 10:08:21",
      "content": "<p>Did you change the NN input size during the training?</p>\n\n<p>Did you use mixed FP32/FP16 ?</p>\n\n<p>Thanks again <a href=\"/dempton\">@dempton</a> !</p>",
      "rawMarkdown": "Did you change the NN input size during the training?\n\nDid you use mixed FP32/FP16 ?\n\nThanks again @dempton !",
      "votes": null
    },
    {
      "id": "546199",
      "postDate": "06/06/2019 10:57:00",
      "content": "<p>No, nothing from this</p>",
      "rawMarkdown": "No, nothing from this",
      "votes": null
    },
    {
      "id": "546218",
      "postDate": "06/06/2019 11:14:52",
      "content": "<p>Really nice solution and a great description. Thanks for sharing. Can you also share how did you use ImageNet predictions ? Did you generate for both culture &amp; tag images? And how did you use them in you final solution?</p>",
      "rawMarkdown": "Really nice solution and a great description. Thanks for sharing. Can you also share how did you use ImageNet predictions ? Did you generate for both culture &amp; tag images? And how did you use them in you final solution?",
      "votes": null
    },
    {
      "id": "546226",
      "postDate": "06/06/2019 11:18:51",
      "content": "<p>Yes, for both of them.\nI generated it just as a features for LightGBM.</p>",
      "rawMarkdown": "Yes, for both of them.\nI generated it just as a features for LightGBM.",
      "votes": null
    },
    {
      "id": "546241",
      "postDate": "06/06/2019 11:38:21",
      "content": "<blockquote>\n  <p><strong>Konstantin Gavrilchik wrote</strong></p>\n  \n  <blockquote>\n    <p>Batch size accumulation\n    (you can perform optimizer step each N batches)</p>\n  </blockquote>\n</blockquote>\n\n<p>Have you done anything special regarding the Batch Normalization layer in the network? I tried the same thing and found there is around 0.004 difference in CV comparing the not using the accumulation. </p>",
      "rawMarkdown": "&gt; **Konstantin Gavrilchik wrote**\n&gt; \n&gt; &gt; Batch size accumulation\n&gt; (you can perform optimizer step each N batches)\n\n\nHave you done anything special regarding the Batch Normalization layer in the network? I tried the same thing and found there is around 0.004 difference in CV comparing the not using the accumulation.",
      "votes": null
    },
    {
      "id": "546249",
      "postDate": "06/06/2019 11:48:08",
      "content": "<p>No, but I increased LR with increasing batch size (1e-4 -&gt; 5e-4), with the same LR performance was worse</p>",
      "rawMarkdown": "No, but I increased LR with increasing batch size (1e-4 -&gt; 5e-4), with the same LR performance was worse",
      "votes": null
    },
    {
      "id": "546250",
      "postDate": "06/06/2019 11:48:08",
      "content": "<p>Interesting solution!!</p>",
      "rawMarkdown": "Interesting solution!!",
      "votes": null
    },
    {
      "id": "546259",
      "postDate": "06/06/2019 11:58:09",
      "content": "<p>cool! </p>",
      "rawMarkdown": "cool!",
      "votes": null
    },
    {
      "id": "546363",
      "postDate": "06/06/2019 14:10:47",
      "content": "<p>Really great solution! </p>",
      "rawMarkdown": "Really great solution!",
      "votes": null
    },
    {
      "id": "546394",
      "postDate": "06/06/2019 14:40:17",
      "content": "<p>Thanks for sharing! Do you have plan to present it at the FGVC workshop?</p>",
      "rawMarkdown": "Thanks for sharing! Do you have plan to present it at the FGVC workshop?",
      "votes": null
    },
    {
      "id": "546409",
      "postDate": "06/06/2019 14:55:50",
      "content": "<p>I have no visa ;(</p>",
      "rawMarkdown": "I have no visa ;(",
      "votes": null
    },
    {
      "id": "546527",
      "postDate": "06/06/2019 16:35:01",
      "content": "<p>Congrats <a href=\"/dempton\">@dempton</a> ! Thanks for sharing. How much  did your score improve with CropIfNedeed + Resize ?</p>",
      "rawMarkdown": "Congrats @dempton ! Thanks for sharing. How much  did your score improve with CropIfNedeed + Resize ?",
      "votes": null
    },
    {
      "id": "546549",
      "postDate": "06/06/2019 17:16:57",
      "content": "<p>I can't say it exactly because I found it improvements at the early stage of my trainings. At the beginning I got boost smth like 0.59-&gt;0.605 on CV (yes, it is great, but I did not compare the last models)</p>",
      "rawMarkdown": "I can't say it exactly because I found it improvements at the early stage of my trainings. At the beginning I got boost smth like 0.59-&gt;0.605 on CV (yes, it is great, but I did not compare the last models)",
      "votes": null
    },
    {
      "id": "546560",
      "postDate": "06/06/2019 17:23:57",
      "content": "<p>You really deserve the first place! That's an amazing solution, thanks for sharing.</p>",
      "rawMarkdown": "You really deserve the first place! That's an amazing solution, thanks for sharing.",
      "votes": null
    },
    {
      "id": "546606",
      "postDate": "06/06/2019 18:15:17",
      "content": "<p>Suppose we do Batch (of batch size k) accumulation for every N batches, and we define 'effective batch size' as (k x N). We get a decent score with this approach. </p>\n\n<p>If someone has a super powerful GPU and he can put the a batch of (k x N) images into it. If we use batch size = (k x N) with this powerful GPU, do you think we need to do anything to LR to reach the same score? </p>\n\n<p>I am asking because I am wondering whether there is any performance decrease due to the synchronization of Batch Normalization. </p>",
      "rawMarkdown": "Suppose we do Batch (of batch size k) accumulation for every N batches, and we define 'effective batch size' as (k x N). We get a decent score with this approach. \n\nIf someone has a super powerful GPU and he can put the a batch of (k x N) images into it. If we use batch size = (k x N) with this powerful GPU, do you think we need to do anything to LR to reach the same score? \n\nI am asking because I am wondering whether there is any performance decrease due to the synchronization of Batch Normalization.",
      "votes": null
    },
    {
      "id": "546637",
      "postDate": "06/06/2019 19:07:29",
      "content": "<p>Thanks <a href=\"/dempton\">@dempton</a>. Did augmentation help in this competition much ? I was using Kaggle kernel to train and I wanted to train SE-ResNext50 as fast as I can. So, I used just Horizontal flip and no other augmentation. Trained for 12 epochs. I did 6-fold CV and got LB 0.617</p>\n\n<p>I found learning rate crucial. What does your findings say ?</p>",
      "rawMarkdown": "Thanks @dempton. Did augmentation help in this competition much ? I was using Kaggle kernel to train and I wanted to train SE-ResNext50 as fast as I can. So, I used just Horizontal flip and no other augmentation. Trained for 12 epochs. I did 6-fold CV and got LB 0.617\n\nI found learning rate crucial. What does your findings say ?",
      "votes": null
    },
    {
      "id": "546672",
      "postDate": "06/06/2019 19:54:06",
      "content": "<p>Thanks for sharing! Learnt a lot from you!! I may can't have computation resources as you.. But your idea to deal with the data really inspired me, you really deserve the first place! </p>",
      "rawMarkdown": "Thanks for sharing! Learnt a lot from you!! I may can't have computation resources as you.. But your idea to deal with the data really inspired me, you really deserve the first place!",
      "votes": null
    },
    {
      "id": "546686",
      "postDate": "06/06/2019 20:09:09",
      "content": "<p>Got a question about the gradient accumulation. Does it mean you did loss.backward() for every batch but the optimizer.step(), and left it to be done after several batches? Is this natively supported in PyTorch?</p>",
      "rawMarkdown": "Got a question about the gradient accumulation. Does it mean you did loss.backward() for every batch but the optimizer.step(), and left it to be done after several batches? Is this natively supported in PyTorch?",
      "votes": null
    },
    {
      "id": "546724",
      "postDate": "06/06/2019 20:58:05",
      "content": "<p>US visa is so hard to obtain= = even for short visit sometimes US need to check your record for many weeks = =</p>",
      "rawMarkdown": "US visa is so hard to obtain= = even for short visit sometimes US need to check your record for many weeks = =",
      "votes": null
    },
    {
      "id": "546806",
      "postDate": "06/06/2019 22:49:45",
      "content": "<p>Great job and thanks for sharing, this seems to be a very well constructed process, congrats.</p>",
      "rawMarkdown": "Great job and thanks for sharing, this seems to be a very well constructed process, congrats.",
      "votes": null
    },
    {
      "id": "546917",
      "postDate": "06/07/2019 02:36:31",
      "content": "<p>This is awesome! Thanks for sharing.</p>",
      "rawMarkdown": "This is awesome! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "547330",
      "postDate": "06/07/2019 15:29:39",
      "content": "<p>Thanks for sharing the great solution.\nI have several questions.\n1. What does \"Sampling with logarithmic weights\" in stage1 mean? \n2. What does \"Hard negative mining (Sample 5% of hardest samples each epoch)\" in stage2 mean? You drop such hardest data from train data in each epoch? <br>\nThank you.</p>",
      "rawMarkdown": "Thanks for sharing the great solution.\nI have several questions.\n1. What does \"Sampling with logarithmic weights\" in stage1 mean? \n2. What does \"Hard negative mining (Sample 5% of hardest samples each epoch)\" in stage2 mean? You drop such hardest data from train data in each epoch?  \nThank you.",
      "votes": null
    },
    {
      "id": "547417",
      "postDate": "06/07/2019 17:13:33",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "547462",
      "postDate": "06/07/2019 18:45:10",
      "content": "<p>It is discussed <a href=\"https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/23\">here</a>. There are also some discussions on fast.ai forum regarding BN synchronization by implementing this way. </p>\n\n<p>I tested it for batch = 64 (no gradient accumulation) vs batch = 32 x 2 (gradient accumulation every 2 steps), everything else the same. The gradient accumulation case has a CV ~0.001 difference between the two cases. </p>",
      "rawMarkdown": "It is discussed [here](https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/23). There are also some discussions on fast.ai forum regarding BN synchronization by implementing this way. \n\nI tested it for batch = 64 (no gradient accumulation) vs batch = 32 x 2 (gradient accumulation every 2 steps), everything else the same. The gradient accumulation case has a CV ~0.001 difference between the two cases.",
      "votes": null
    },
    {
      "id": "547634",
      "postDate": "06/08/2019 02:21:11",
      "content": "<p>You are the champion in my heart!!!Thanks for sharing.</p>",
      "rawMarkdown": "You are the champion in my heart!!!Thanks for sharing.",
      "votes": null
    },
    {
      "id": "549060",
      "postDate": "06/10/2019 08:58:34",
      "content": "<p>nice job</p>",
      "rawMarkdown": "nice job",
      "votes": null
    },
    {
      "id": "549088",
      "postDate": "06/10/2019 09:32:50",
      "content": "<p><em>Batch size: 1000-1500</em>  Unbelievable. \nHow big is your gpu memory?</p>",
      "rawMarkdown": "*Batch size: 1000-1500*  Unbelievable. \nHow big is your gpu memory?",
      "votes": null
    },
    {
      "id": "549091",
      "postDate": "06/10/2019 09:36:01",
      "content": "<p>Can you tell me why the big batch size is so effective in this model?</p>",
      "rawMarkdown": "Can you tell me why the big batch size is so effective in this model?",
      "votes": null
    },
    {
      "id": "549669",
      "postDate": "06/10/2019 23:24:47",
      "content": "<p>Congrats! Well done! You deserve the private top 1! </p>",
      "rawMarkdown": "Congrats! Well done! You deserve the private top 1!",
      "votes": null
    },
    {
      "id": "549826",
      "postDate": "06/11/2019 04:14:52",
      "content": "<p>Congratulations! You deserved the top spot from the very beginning! Is it possible for you to open source the code? Many thanks! :-)</p>",
      "rawMarkdown": "Congratulations! You deserved the top spot from the very beginning! Is it possible for you to open source the code? Many thanks! :-)",
      "votes": null
    },
    {
      "id": "549837",
      "postDate": "06/11/2019 04:33:08",
      "content": "<p>Congrats! You deserve the first place 😄 </p>",
      "rawMarkdown": "Congrats! You deserve the first place 😄",
      "votes": null
    },
    {
      "id": "550182",
      "postDate": "06/11/2019 11:39:59",
      "content": "<p>Sorry for late answer\n1. I used WeightedSampler where weights was set to log of probability of classes in dataset\n2. No, I add hardest data in each epoch (from previous). I sampled 5% of samples with biggest loss on the previous epoch to current. (number of images in epoch != max number of images)</p>",
      "rawMarkdown": "Sorry for late answer\n1. I used WeightedSampler where weights was set to log of probability of classes in dataset\n2. No, I add hardest data in each epoch (from previous). I sampled 5% of samples with biggest loss on the previous epoch to current. (number of images in epoch != max number of images)",
      "votes": null
    },
    {
      "id": "550193",
      "postDate": "06/11/2019 11:50:42",
      "content": "<p>It's batch accumulation, you can do it with any GPU.</p>",
      "rawMarkdown": "It's batch accumulation, you can do it with any GPU.",
      "votes": null
    },
    {
      "id": "550528",
      "postDate": "06/11/2019 17:46:50",
      "content": "<p>Hey Konstantin,</p>\n\n<p>Do you mind putting your summary into 1 slide so that I can share it in the workshop on your behalf?</p>",
      "rawMarkdown": "Hey Konstantin,\n\nDo you mind putting your summary into 1 slide so that I can share it in the workshop on your behalf?",
      "votes": null
    },
    {
      "id": "550534",
      "postDate": "06/11/2019 17:50:19",
      "content": "<p>Yep, I can</p>",
      "rawMarkdown": "Yep, I can",
      "votes": null
    },
    {
      "id": "550560",
      "postDate": "06/11/2019 18:39:16",
      "content": "<p>Cool! Feel free to ping me when you are ready! Thanks</p>",
      "rawMarkdown": "Cool! Feel free to ping me when you are ready! Thanks",
      "votes": null
    },
    {
      "id": "550844",
      "postDate": "06/12/2019 04:43:49",
      "content": "<p>Cool!</p>",
      "rawMarkdown": "Cool!",
      "votes": null
    },
    {
      "id": "556973",
      "postDate": "06/20/2019 23:35:02",
      "content": "<p>Thank you for your kind replying!</p>",
      "rawMarkdown": "Thank you for your kind replying!",
      "votes": null
    },
    {
      "id": "564904",
      "postDate": "06/30/2019 06:29:08",
      "content": "<p>I saw the method you used again, it is really amazing.</p>",
      "rawMarkdown": "I saw the method you used again, it is really amazing.",
      "votes": null
    },
    {
      "id": "566044",
      "postDate": "07/01/2019 18:26:29",
      "content": "<p>Thanks for the description!\nAs the final results are released, could you share the code? I'm interested in Hard-Negative part of code (and also others)</p>",
      "rawMarkdown": "Thanks for the description!\nAs the final results are released, could you share the code? I'm interested in Hard-Negative part of code (and also others)",
      "votes": null
    },
    {
      "id": "570986",
      "postDate": "07/09/2019 02:25:57",
      "content": "<p>Hi, Gavrilchik,\nMany days have passed, I still didn't wait until stage 2 finished,  and I have a little question.\nMy understanding of \"number of images in epoch != max number of images\" is：\nFirst, use a Initial training set (keep the negative sample the same size as the positive sample) train model, secondly use the model to classify the samples, and the hard negative samples which have the biggest loss in the negative samples are put into the initial training set. Then train the next epoch. so the classifier is trained again and again. Is that right？Thank you.</p>",
      "rawMarkdown": "Hi, Gavrilchik,\nMany days have passed, I still didn't wait until stage 2 finished,  and I have a little question.\nMy understanding of \"number of images in epoch != max number of images\" is：\nFirst, use a Initial training set (keep the negative sample the same size as the positive sample) train model, secondly use the model to classify the samples, and the hard negative samples which have the biggest loss in the negative samples are put into the initial training set. Then train the next epoch. so the classifier is trained again and again. Is that right？Thank you.",
      "votes": null
    },
    {
      "id": "672034",
      "postDate": "11/13/2019 13:13:04",
      "content": "<p>Instead of cropping the images, why not resize them? Won't that be better for the validity of the labels?</p>",
      "rawMarkdown": "Instead of cropping the images, why not resize them? Won't that be better for the validity of the labels?",
      "votes": null
    },
    {
      "id": "963699",
      "postDate": "08/09/2020 08:12:12",
      "content": "<p>Thank for sharing, Great solution!</p>",
      "rawMarkdown": "Thank for sharing, Great solution!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 546090,
      "author_name": "qiaojian",
      "author_url": "",
      "post_date": "06/06/2019 08:46:05",
      "content": "<p>Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546097,
      "author_name": "",
      "author_url": "",
      "post_date": "06/06/2019 08:54:43",
      "content": "<p>Thanks for sharing! I have learned a lot from this discussion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546098,
      "author_name": "nekoyyy",
      "author_url": "",
      "post_date": "06/06/2019 08:54:43",
      "content": "<p>Amazing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546099,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "06/06/2019 08:55:34",
      "content": "<p>Congratulations <a href=\"/dempton\">@dempton</a> ! And thanks a lot for sharing!</p>\n\n<p>Did you use the fastai library?</p>\n\n<p>Regarding the network structure, did you use any GAPnet or altered the structure for the feature extraction for any network?</p>\n\n<p>Did you use the original network heads? Appart from changing the number of output classes :-) and the multitask learner</p>",
      "votes": null,
      "replies": [
        {
          "id": 546106,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 08:59:40",
          "content": "<p>No, I used just pure pytorch as a DL framework\nNo, I did not use GAPNet, Just resnet, DenseNet, NasNet like architectures </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546101,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "06/06/2019 08:56:37",
      "content": "<p>Thanks for sharing, a nice solution.\nCould you tell me which point is the key point and How much LB has been improved?</p>",
      "votes": null,
      "replies": [
        {
          "id": 546107,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 09:03:53",
          "content": "<p>It is very hard to say how much each step influenced on the LB because I did not make a lot of submissions (validation was very good).\nAccording to validation:\n- filtering (~0.005 - 0.01, do not remember exactly)\n- pseudo labeling (it increased (0.01)\n- second level (it increased 0.005 - 0.007)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546109,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 09:05:46",
          "content": "<p>But anyway the most important for this competitions was a correct LR scheduler (which I set manually for each epoch) + preprocessing (CropIfNeeded + Resize)\nPreprocessing speed up convergence which helped me to perform a lot of experiments</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546112,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "06/06/2019 09:10:06",
          "content": "<p>Got it, thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546122,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "06/06/2019 09:18:52",
          "content": "<p>And, one more question, why did you use the large batch size(1k-1.5k). Is that better than 32/64/128?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546126,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 09:22:26",
          "content": "<p>yep, I found that moving from 64 to 64*4 increased my performance very good\nSo, I decided choose batch size ~ 64 * 20 (it increased a little bit more and I stopped tuning it)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546141,
          "author_name": "shentao",
          "author_url": "",
          "post_date": "06/06/2019 09:42:23",
          "content": "<p>Congrats and thank you for the detailed information! I think it's hard to surpass your work in this competition and you did it solo!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546104,
      "author_name": "fengari",
      "author_url": "",
      "post_date": "06/06/2019 08:58:35",
      "content": "<p>Thanks, really learn a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546110,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "06/06/2019 09:09:02",
      "content": "<p>Thanks, you deserve the decent result! By the way, how do you use 1000+ batch size with 6 1080ti (66GB GPU memory)</p>",
      "votes": null,
      "replies": [
        {
          "id": 546114,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 09:11:24",
          "content": "<p>Batch size accumulation\n(you can perform optimizer step each N batches)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546118,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "06/06/2019 09:14:59",
          "content": "<p>I see</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546241,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "06/06/2019 11:38:21",
          "content": "<blockquote>\n  <p><strong>Konstantin Gavrilchik wrote</strong></p>\n  \n  <blockquote>\n    <p>Batch size accumulation\n    (you can perform optimizer step each N batches)</p>\n  </blockquote>\n</blockquote>\n\n<p>Have you done anything special regarding the Batch Normalization layer in the network? I tried the same thing and found there is around 0.004 difference in CV comparing the not using the accumulation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546249,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 11:48:08",
          "content": "<p>No, but I increased LR with increasing batch size (1e-4 -&gt; 5e-4), with the same LR performance was worse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546606,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "06/06/2019 18:15:17",
          "content": "<p>Suppose we do Batch (of batch size k) accumulation for every N batches, and we define 'effective batch size' as (k x N). We get a decent score with this approach. </p>\n\n<p>If someone has a super powerful GPU and he can put the a batch of (k x N) images into it. If we use batch size = (k x N) with this powerful GPU, do you think we need to do anything to LR to reach the same score? </p>\n\n<p>I am asking because I am wondering whether there is any performance decrease due to the synchronization of Batch Normalization. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546686,
          "author_name": "manyfoldcv",
          "author_url": "",
          "post_date": "06/06/2019 20:09:09",
          "content": "<p>Got a question about the gradient accumulation. Does it mean you did loss.backward() for every batch but the optimizer.step(), and left it to be done after several batches? Is this natively supported in PyTorch?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547462,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "06/07/2019 18:45:10",
          "content": "<p>It is discussed <a href=\"https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/23\">here</a>. There are also some discussions on fast.ai forum regarding BN synchronization by implementing this way. </p>\n\n<p>I tested it for batch = 64 (no gradient accumulation) vs batch = 32 x 2 (gradient accumulation every 2 steps), everything else the same. The gradient accumulation case has a CV ~0.001 difference between the two cases. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546111,
      "author_name": "luyujia",
      "author_url": "",
      "post_date": "06/06/2019 09:09:26",
      "content": "<p>Good job ! ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546120,
      "author_name": "",
      "author_url": "",
      "post_date": "06/06/2019 09:17:51",
      "content": "<p>you deserve the first place! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546125,
      "author_name": "",
      "author_url": "",
      "post_date": "06/06/2019 09:21:31",
      "content": "<p>good work! will you Open source code？</p>",
      "votes": null,
      "replies": [
        {
          "id": 546127,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 09:23:16",
          "content": "<p>Probably after stage 2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 566044,
          "author_name": "melgor",
          "author_url": "",
          "post_date": "07/01/2019 18:26:29",
          "content": "<p>Thanks for the description!\nAs the final results are released, could you share the code? I'm interested in Hard-Negative part of code (and also others)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546159,
      "author_name": "izmaylov",
      "author_url": "",
      "post_date": "06/06/2019 10:01:18",
      "content": "<p>Really great solution! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546166,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "06/06/2019 10:08:21",
      "content": "<p>Did you change the NN input size during the training?</p>\n\n<p>Did you use mixed FP32/FP16 ?</p>\n\n<p>Thanks again <a href=\"/dempton\">@dempton</a> !</p>",
      "votes": null,
      "replies": [
        {
          "id": 546199,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 10:57:00",
          "content": "<p>No, nothing from this</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546218,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "06/06/2019 11:14:52",
      "content": "<p>Really nice solution and a great description. Thanks for sharing. Can you also share how did you use ImageNet predictions ? Did you generate for both culture &amp; tag images? And how did you use them in you final solution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 546226,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 11:18:51",
          "content": "<p>Yes, for both of them.\nI generated it just as a features for LightGBM.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546250,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "06/06/2019 11:48:08",
      "content": "<p>Interesting solution!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546259,
      "author_name": "martinid",
      "author_url": "",
      "post_date": "06/06/2019 11:58:09",
      "content": "<p>cool! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546363,
      "author_name": "angelecarre",
      "author_url": "",
      "post_date": "06/06/2019 14:10:47",
      "content": "<p>Really great solution! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546394,
      "author_name": "codingcliff",
      "author_url": "",
      "post_date": "06/06/2019 14:40:17",
      "content": "<p>Thanks for sharing! Do you have plan to present it at the FGVC workshop?</p>",
      "votes": null,
      "replies": [
        {
          "id": 546409,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 14:55:50",
          "content": "<p>I have no visa ;(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546724,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "06/06/2019 20:58:05",
          "content": "<p>US visa is so hard to obtain= = even for short visit sometimes US need to check your record for many weeks = =</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550528,
          "author_name": "codingcliff",
          "author_url": "",
          "post_date": "06/11/2019 17:46:50",
          "content": "<p>Hey Konstantin,</p>\n\n<p>Do you mind putting your summary into 1 slide so that I can share it in the workshop on your behalf?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550534,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/11/2019 17:50:19",
          "content": "<p>Yep, I can</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550560,
          "author_name": "codingcliff",
          "author_url": "",
          "post_date": "06/11/2019 18:39:16",
          "content": "<p>Cool! Feel free to ping me when you are ready! Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546527,
      "author_name": "harshthaker",
      "author_url": "",
      "post_date": "06/06/2019 16:35:01",
      "content": "<p>Congrats <a href=\"/dempton\">@dempton</a> ! Thanks for sharing. How much  did your score improve with CropIfNedeed + Resize ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 546549,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/06/2019 17:16:57",
          "content": "<p>I can't say it exactly because I found it improvements at the early stage of my trainings. At the beginning I got boost smth like 0.59-&gt;0.605 on CV (yes, it is great, but I did not compare the last models)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546637,
          "author_name": "harshthaker",
          "author_url": "",
          "post_date": "06/06/2019 19:07:29",
          "content": "<p>Thanks <a href=\"/dempton\">@dempton</a>. Did augmentation help in this competition much ? I was using Kaggle kernel to train and I wanted to train SE-ResNext50 as fast as I can. So, I used just Horizontal flip and no other augmentation. Trained for 12 epochs. I did 6-fold CV and got LB 0.617</p>\n\n<p>I found learning rate crucial. What does your findings say ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547417,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "06/07/2019 17:13:33",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546560,
      "author_name": "manyfoldcv",
      "author_url": "",
      "post_date": "06/06/2019 17:23:57",
      "content": "<p>You really deserve the first place! That's an amazing solution, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546672,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "06/06/2019 19:54:06",
      "content": "<p>Thanks for sharing! Learnt a lot from you!! I may can't have computation resources as you.. But your idea to deal with the data really inspired me, you really deserve the first place! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546806,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "06/06/2019 22:49:45",
      "content": "<p>Great job and thanks for sharing, this seems to be a very well constructed process, congrats.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546917,
      "author_name": "a31314431",
      "author_url": "",
      "post_date": "06/07/2019 02:36:31",
      "content": "<p>This is awesome! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547330,
      "author_name": "tanakin",
      "author_url": "",
      "post_date": "06/07/2019 15:29:39",
      "content": "<p>Thanks for sharing the great solution.\nI have several questions.\n1. What does \"Sampling with logarithmic weights\" in stage1 mean? \n2. What does \"Hard negative mining (Sample 5% of hardest samples each epoch)\" in stage2 mean? You drop such hardest data from train data in each epoch? <br>\nThank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 550182,
          "author_name": "dempton",
          "author_url": "",
          "post_date": "06/11/2019 11:39:59",
          "content": "<p>Sorry for late answer\n1. I used WeightedSampler where weights was set to log of probability of classes in dataset\n2. No, I add hardest data in each epoch (from previous). I sampled 5% of samples with biggest loss on the previous epoch to current. (number of images in epoch != max number of images)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 556973,
          "author_name": "tanakin",
          "author_url": "",
          "post_date": "06/20/2019 23:35:02",
          "content": "<p>Thank you for your kind replying!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 570986,
          "author_name": "guanzhongzeng",
          "author_url": "",
          "post_date": "07/09/2019 02:25:57",
          "content": "<p>Hi, Gavrilchik,\nMany days have passed, I still didn't wait until stage 2 finished,  and I have a little question.\nMy understanding of \"number of images in epoch != max number of images\" is：\nFirst, use a Initial training set (keep the negative sample the same size as the positive sample) train model, secondly use the model to classify the samples, and the hard negative samples which have the biggest loss in the negative samples are put into the initial training set. Then train the next epoch. so the classifier is trained again and again. Is that right？Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 547634,
      "author_name": "guiyuan320",
      "author_url": "",
      "post_date": "06/08/2019 02:21:11",
      "content": "<p>You are the champion in my heart!!!Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549060,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "06/10/2019 08:58:34",
      "content": "<p>nice job</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549088,
      "author_name": "",
      "author_url": "",
      "post_date": "06/10/2019 09:32:50",
      "content": "<p><em>Batch size: 1000-1500</em>  Unbelievable. \nHow big is your gpu memory?</p>",
      "votes": null,
      "replies": [
        {
          "id": 550193,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "06/11/2019 11:50:42",
          "content": "<p>It's batch accumulation, you can do it with any GPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 549091,
      "author_name": "",
      "author_url": "",
      "post_date": "06/10/2019 09:36:01",
      "content": "<p>Can you tell me why the big batch size is so effective in this model?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549669,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "06/10/2019 23:24:47",
      "content": "<p>Congrats! Well done! You deserve the private top 1! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549826,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "06/11/2019 04:14:52",
      "content": "<p>Congratulations! You deserved the top spot from the very beginning! Is it possible for you to open source the code? Many thanks! :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549837,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "06/11/2019 04:33:08",
      "content": "<p>Congrats! You deserve the first place 😄 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550844,
      "author_name": "kevin0401",
      "author_url": "",
      "post_date": "06/12/2019 04:43:49",
      "content": "<p>Cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 564904,
      "author_name": "",
      "author_url": "",
      "post_date": "06/30/2019 06:29:08",
      "content": "<p>I saw the method you used again, it is really amazing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 672034,
      "author_name": "imenhaddad",
      "author_url": "",
      "post_date": "11/13/2019 13:13:04",
      "content": "<p>Instead of cropping the images, why not resize them? Won't that be better for the validity of the labels?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 963699,
      "author_name": "",
      "author_url": "",
      "post_date": "08/09/2020 08:12:12",
      "content": "<p>Thank for sharing, Great solution!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "546074": "Congrats to all who finished in the gold zone. Now, while we are waiting for second stage results let me share my solution part. I do not expect shake up, so I think leaderboard will be stay the same.\n\nFirst of all, what is the main challenge in this competition? Of course it is noisy data and we should find a way to work with it. \nSo, i can divide my solution into several stages:\n\nHardware: 6x 1080 ti with 36 cores and 120 gb RAM in total.\n\n**Stage 0. The same part for all stages:**\nModels: SENet154, PNasNet-5, SE-ResNext101 (all pretrained from cadene repo)\nCV: 5 folds with Multilabel Iterative Stratification\n*Data Augmentations:*\n```\n        HorizontalFlip(p=0.5),\n        OneOf([\n            RandomBrightness(0.1, p=1),\n            RandomContrast(0.1, p=1),\n        ], p=0.3),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.0, rotate_limit=15, p=0.3),\n        IAAAdditiveGaussianNoise(p=0.3),\n```\n*Data Preprocessing:*\nThis part was a great finding which speed up convergence and increased score very good. So, analysis of different models with Crops(224), Crops(288), Crops(300) and so on shows that it influence on tags a lots. Actually, lets imagine that you have picture 500x300 and this image labeled with tag «person» if there are persons somewhere at this image. There is a not so low probability that you crop it =&gt; you data becomes more and more noisy. \nSo, I decided to use CropIfNedeed + Resize. The main Idea is perform crop if it possible (for example crop 600x600 from 300x500 image — 300x500. from 300x900 — 300x600). Code below:\n\n```\nclass RandomCropIfNeeded(RandomCrop):\n    def __init__(self, height, width, always_apply=False, p=1.0):\n        super(RandomCrop, self).__init__(always_apply, p)\n        self.height = height\n        self.width = width\n\n    def apply(self, img, h_start=0, w_start=0, **params):\n        h, w, _ = img.shape\n        return F.random_crop(img, min(self.height, h), min(self.width, w), h_start, w_start)\n``` \nSo, I used the following couple:\n```\nRandomCropIfNeeded(SIZE * 2, SIZE * 2),\nResize(SIZE, SIZE)\n```\nWith SIZE = 320 for SEResNext101 / SENet154 and SIZE = 331 for PNasNet-5\n\n*TTA*: Using the previous approach for data preprocessing we can just use TTA2 (original + hflip image)\n\n*Scheduler*: manual, the correct scheduler increased accuracy a lot. I found it by analyzing plots with metrics / losses. \nI started with LR: 0.005 then train 15 epochs, drop it by 5 times and train some epochs more.\n\n**Stage 1. Training the zoo:**\n\n*Loss*: Focal\n*Sampling* with logarithmic weights\n*Batch size*: 1000-1500 (accumulation 10-20 times for different models)\n\n**Stage 2. Filtering predictions:**\nDrop images from train with very high error between OOF predictions and labels (I consider them as very noisy and incorrect)\n\n*Loss*: Focal\n+ Hard negative mining (Sample 5% of hardest samples each epoch)\n\n\nRe-train models from scratch.\n\n**Stage 3. Pseudo labeling:**\n\n*Loss*: Focal\n\nThe simplest version of pseudo labeling, add the most confident predictions (highest  np.mean(np.abs(probabilities - 0.5)) ) to the training dataset.\n\nRe-train models from scratch.\n\n**Stage 4. Culture and tags separately:**\n\nI found 2 main things about cultures and tags:\n1. tags are less noisy than cultures (and starting from some epoch the culture accuracy did not increased very much)\n2. some tags classes are very similar to ImageNet classes\n\nSo, using already trained weights as pretrain I continue training model for 705 classes (tags only) \n\n*Loss: Focal* -&gt; BCE (yes, here I switched the loss, I helped too)\n\n**Stage 5. Second-level model**\n\nI construct the binary classification dataset: I took 1103 (number of classes) rows per each image and trying to predict that this class relates to this image (0 or 1). So it means the length of my train data becomes `len(data) * 1103`.\n\nI extract next features:\n- probabilities of each models, sum / division / multiplication of each pair / triple / .. of models \n- mean / median / std / max / min of each channel\n- brightness / colorness of each image (you can say me that NN can easily detect it — yes, but here i can do it without cropping and resizing — it is less noisy)\n- Max side size and binary flag — height more than width or no (it is a little bit better for tree boosting than just height + width in case of lower side == 300)\n- Aaaaand the secret sauce: ImageNet predictions ;) As I already mentioned — some tags classes similar to ImageNet classes, but ImageNet much bigger, pretrained models much more generalized. So, I add all 1000 (number of ImageNet classes) predictions to this dataset\n\nSo, then I trained LightGBM on all this data.\n\n**Hints / Postprocessing:**\n- Different threshold for cultures and tags models. \n- EDA shows that tags are fully labeled and cultures may be not. So, I binarize predictions using the following code:\n\n``` \nculture_predictions = binarize_predictions(predictions[:, :398], threshold=0.28, min_samples=0, max_samples=3)\ntags_predictions = binarize_predictions(predictions[:, 398:], threshold=0.1, min_samples=1, max_samples=8)\npredictions = np.hstack([culture_predictions, tags_predictions])\n```\n\n(thank you to @lopuhin for the binarize_predictions code in his kernel)\n\nTotal training time: ~10-15 days (that's why I submitted not very often ;) )\n\n\nP.S. I did not submit all this ensemble to the private stage… Due to I faced with kernel limit time :facepalm: (for example I used LGBM for pseudo-labeling but did not use it at the next cycle of test predictions)",
    "546090": "Well done!",
    "546097": "Thanks for sharing! I have learned a lot from this discussion.",
    "546098": "Amazing!",
    "546099": "Congratulations @dempton ! And thanks a lot for sharing!\n\nDid you use the fastai library?\n\nRegarding the network structure, did you use any GAPnet or altered the structure for the feature extraction for any network?\n\nDid you use the original network heads? Appart from changing the number of output classes :-) and the multitask learner",
    "546101": "Thanks for sharing, a nice solution.\nCould you tell me which point is the key point and How much LB has been improved?",
    "546104": "Thanks, really learn a lot.",
    "546106": "No, I used just pure pytorch as a DL framework\nNo, I did not use GAPNet, Just resnet, DenseNet, NasNet like architectures",
    "546107": "It is very hard to say how much each step influenced on the LB because I did not make a lot of submissions (validation was very good).\nAccording to validation:\n- filtering (~0.005 - 0.01, do not remember exactly)\n- pseudo labeling (it increased (0.01)\n- second level (it increased 0.005 - 0.007)",
    "546109": "But anyway the most important for this competitions was a correct LR scheduler (which I set manually for each epoch) + preprocessing (CropIfNeeded + Resize)\nPreprocessing speed up convergence which helped me to perform a lot of experiments",
    "546110": "Thanks, you deserve the decent result! By the way, how do you use 1000+ batch size with 6 1080ti (66GB GPU memory)",
    "546111": "Good job ! !",
    "546112": "Got it, thank you.",
    "546114": "Batch size accumulation\n(you can perform optimizer step each N batches)",
    "546118": "I see",
    "546120": "you deserve the first place!",
    "546122": "And, one more question, why did you use the large batch size(1k-1.5k). Is that better than 32/64/128?",
    "546125": "good work! will you Open source code？",
    "546126": "yep, I found that moving from 64 to 64*4 increased my performance very good\nSo, I decided choose batch size ~ 64 * 20 (it increased a little bit more and I stopped tuning it)",
    "546127": "Probably after stage 2",
    "546141": "Congrats and thank you for the detailed information! I think it's hard to surpass your work in this competition and you did it solo!",
    "546159": "Really great solution!",
    "546166": "Did you change the NN input size during the training?\n\nDid you use mixed FP32/FP16 ?\n\nThanks again @dempton !",
    "546199": "No, nothing from this",
    "546218": "Really nice solution and a great description. Thanks for sharing. Can you also share how did you use ImageNet predictions ? Did you generate for both culture &amp; tag images? And how did you use them in you final solution?",
    "546226": "Yes, for both of them.\nI generated it just as a features for LightGBM.",
    "546241": "&gt; **Konstantin Gavrilchik wrote**\n&gt; \n&gt; &gt; Batch size accumulation\n&gt; (you can perform optimizer step each N batches)\n\n\nHave you done anything special regarding the Batch Normalization layer in the network? I tried the same thing and found there is around 0.004 difference in CV comparing the not using the accumulation.",
    "546249": "No, but I increased LR with increasing batch size (1e-4 -&gt; 5e-4), with the same LR performance was worse",
    "546250": "Interesting solution!!",
    "546259": "cool!",
    "546363": "Really great solution!",
    "546394": "Thanks for sharing! Do you have plan to present it at the FGVC workshop?",
    "546409": "I have no visa ;(",
    "546527": "Congrats @dempton ! Thanks for sharing. How much  did your score improve with CropIfNedeed + Resize ?",
    "546549": "I can't say it exactly because I found it improvements at the early stage of my trainings. At the beginning I got boost smth like 0.59-&gt;0.605 on CV (yes, it is great, but I did not compare the last models)",
    "546560": "You really deserve the first place! That's an amazing solution, thanks for sharing.",
    "546606": "Suppose we do Batch (of batch size k) accumulation for every N batches, and we define 'effective batch size' as (k x N). We get a decent score with this approach. \n\nIf someone has a super powerful GPU and he can put the a batch of (k x N) images into it. If we use batch size = (k x N) with this powerful GPU, do you think we need to do anything to LR to reach the same score? \n\nI am asking because I am wondering whether there is any performance decrease due to the synchronization of Batch Normalization.",
    "546637": "Thanks @dempton. Did augmentation help in this competition much ? I was using Kaggle kernel to train and I wanted to train SE-ResNext50 as fast as I can. So, I used just Horizontal flip and no other augmentation. Trained for 12 epochs. I did 6-fold CV and got LB 0.617\n\nI found learning rate crucial. What does your findings say ?",
    "546672": "Thanks for sharing! Learnt a lot from you!! I may can't have computation resources as you.. But your idea to deal with the data really inspired me, you really deserve the first place!",
    "546686": "Got a question about the gradient accumulation. Does it mean you did loss.backward() for every batch but the optimizer.step(), and left it to be done after several batches? Is this natively supported in PyTorch?",
    "546724": "US visa is so hard to obtain= = even for short visit sometimes US need to check your record for many weeks = =",
    "546806": "Great job and thanks for sharing, this seems to be a very well constructed process, congrats.",
    "546917": "This is awesome! Thanks for sharing.",
    "547330": "Thanks for sharing the great solution.\nI have several questions.\n1. What does \"Sampling with logarithmic weights\" in stage1 mean? \n2. What does \"Hard negative mining (Sample 5% of hardest samples each epoch)\" in stage2 mean? You drop such hardest data from train data in each epoch?  \nThank you.",
    "547417": "",
    "547462": "It is discussed [here](https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/23). There are also some discussions on fast.ai forum regarding BN synchronization by implementing this way. \n\nI tested it for batch = 64 (no gradient accumulation) vs batch = 32 x 2 (gradient accumulation every 2 steps), everything else the same. The gradient accumulation case has a CV ~0.001 difference between the two cases.",
    "547634": "You are the champion in my heart!!!Thanks for sharing.",
    "549060": "nice job",
    "549088": "*Batch size: 1000-1500*  Unbelievable. \nHow big is your gpu memory?",
    "549091": "Can you tell me why the big batch size is so effective in this model?",
    "549669": "Congrats! Well done! You deserve the private top 1!",
    "549826": "Congratulations! You deserved the top spot from the very beginning! Is it possible for you to open source the code? Many thanks! :-)",
    "549837": "Congrats! You deserve the first place 😄",
    "550182": "Sorry for late answer\n1. I used WeightedSampler where weights was set to log of probability of classes in dataset\n2. No, I add hardest data in each epoch (from previous). I sampled 5% of samples with biggest loss on the previous epoch to current. (number of images in epoch != max number of images)",
    "550193": "It's batch accumulation, you can do it with any GPU.",
    "550528": "Hey Konstantin,\n\nDo you mind putting your summary into 1 slide so that I can share it in the workshop on your behalf?",
    "550534": "Yep, I can",
    "550560": "Cool! Feel free to ping me when you are ready! Thanks",
    "550844": "Cool!",
    "556973": "Thank you for your kind replying!",
    "564904": "I saw the method you used again, it is really amazing.",
    "566044": "Thanks for the description!\nAs the final results are released, could you share the code? I'm interested in Hard-Negative part of code (and also others)",
    "570986": "Hi, Gavrilchik,\nMany days have passed, I still didn't wait until stage 2 finished,  and I have a little question.\nMy understanding of \"number of images in epoch != max number of images\" is：\nFirst, use a Initial training set (keep the negative sample the same size as the positive sample) train model, secondly use the model to classify the samples, and the hard negative samples which have the biggest loss in the negative samples are put into the initial training set. Then train the next epoch. so the classifier is trained again and again. Is that right？Thank you.",
    "672034": "Instead of cropping the images, why not resize them? Won't that be better for the validity of the labels?",
    "963699": "Thank for sharing, Great solution!"
  },
  "source": "meta"
}