{
  "id": 134905,
  "title": "GeM pooling convergence slow?",
  "url": "/competitions/bengaliai-cv19/discussion/134905",
  "author_name": "",
  "post_date": "2020-03-11T03:27:12.318468800Z",
  "votes": 8,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi Guys:\nJust joined the competition 2 days ago, first of all, thanks for all the great knowledge sharing of you. Since the information in discussion section just too much, thus I am not able to go through them very carefully. So I'm not sure anyone discussed about this before, but it is just me or the GeM pooling really make the model converge slow? I switched to GeM pooling yesterday, and after 120 epochs, the model still not convergence. Normally my model converge in around 60 epochs, so just wondering if there are same situation as mine right now.  </p>\n\n<p>Also, I started this competition after <a href=\"/seesee\">@seesee</a> created the tfrec training data. But I knew that we can't use augmentation modules like albumentations or ImageDataGenerator with tfrec format. Just want to share you guys some links about how to do augmentation with tfrec format.</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">Rotation Augmentation GPU/TPU shared by Chris</a>\n<a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">CutMix and MixUp on GPU/TPU shared by Chris</a>\n<a href=\"https://www.kaggle.com/xiejialun/gridmask-data-augmentation-with-tensorflow\">GridMask data augmentation with tensorflow shared by me</a>\n<a href=\"https://www.kaggle.com/xiejialun/customize-data-augmentation-with-tensorflow\">Gaussian blur and Cutout augmentation shared by me</a>\n<a href=\"https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations\">Batch implementation of above augmentation shared by Yih-Dar SHIEH</a></p>",
  "messages": [
    {
      "id": "768627",
      "postDate": "03/11/2020 03:27:12",
      "content": "<p>Hi Guys:\nJust joined the competition 2 days ago, first of all, thanks for all the great knowledge sharing of you. Since the information in discussion section just too much, thus I am not able to go through them very carefully. So I'm not sure anyone discussed about this before, but it is just me or the GeM pooling really make the model converge slow? I switched to GeM pooling yesterday, and after 120 epochs, the model still not convergence. Normally my model converge in around 60 epochs, so just wondering if there are same situation as mine right now.  </p>\n\n<p>Also, I started this competition after <a href=\"/seesee\">@seesee</a> created the tfrec training data. But I knew that we can't use augmentation modules like albumentations or ImageDataGenerator with tfrec format. Just want to share you guys some links about how to do augmentation with tfrec format.</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">Rotation Augmentation GPU/TPU shared by Chris</a>\n<a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">CutMix and MixUp on GPU/TPU shared by Chris</a>\n<a href=\"https://www.kaggle.com/xiejialun/gridmask-data-augmentation-with-tensorflow\">GridMask data augmentation with tensorflow shared by me</a>\n<a href=\"https://www.kaggle.com/xiejialun/customize-data-augmentation-with-tensorflow\">Gaussian blur and Cutout augmentation shared by me</a>\n<a href=\"https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations\">Batch implementation of above augmentation shared by Yih-Dar SHIEH</a></p>",
      "rawMarkdown": "Hi Guys:\nJust joined the competition 2 days ago, first of all, thanks for all the great knowledge sharing of you. Since the information in discussion section just too much, thus I am not able to go through them very carefully. So I'm not sure anyone discussed about this before, but it is just me or the GeM pooling really make the model converge slow? I switched to GeM pooling yesterday, and after 120 epochs, the model still not convergence. Normally my model converge in around 60 epochs, so just wondering if there are same situation as mine right now.  \n\nAlso, I started this competition after @seesee created the tfrec training data. But I knew that we can't use augmentation modules like albumentations or ImageDataGenerator with tfrec format. Just want to share you guys some links about how to do augmentation with tfrec format.\n\n[Rotation Augmentation GPU/TPU shared by Chris](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96)\n[CutMix and MixUp on GPU/TPU shared by Chris](https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu)\n[GridMask data augmentation with tensorflow shared by me](https://www.kaggle.com/xiejialun/gridmask-data-augmentation-with-tensorflow)\n[Gaussian blur and Cutout augmentation shared by me](https://www.kaggle.com/xiejialun/customize-data-augmentation-with-tensorflow)\n[Batch implementation of above augmentation shared by Yih-Dar SHIEH](https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations)",
      "votes": null
    },
    {
      "id": "768684",
      "postDate": "03/11/2020 04:56:20",
      "content": "<p>Awesome. Thank you for sharing such great notebooks. Are you using tfrec format data set? </p>",
      "rawMarkdown": "Awesome. Thank you for sharing such great notebooks. Are you using tfrec format data set?",
      "votes": null
    },
    {
      "id": "768697",
      "postDate": "03/11/2020 05:13:56",
      "content": "<p>Yes, since the competition is going to end, so I need to use TPU for fast model training.</p>",
      "rawMarkdown": "Yes, since the competition is going to end, so I need to use TPU for fast model training.",
      "votes": null
    },
    {
      "id": "768721",
      "postDate": "03/11/2020 05:43:17",
      "content": "<p>I was planning to use TPU since hearing when its available but couldn't make it :( But now, See-- had made an awesome starter but because of the last moment of the competition, I stick with my set up. I was planning to use CpasNet on TPU to get the leverage of faster training. </p>",
      "rawMarkdown": "I was planning to use TPU since hearing when its available but couldn't make it :( But now, See-- had made an awesome starter but because of the last moment of the competition, I stick with my set up. I was planning to use CpasNet on TPU to get the leverage of faster training.",
      "votes": null
    },
    {
      "id": "768748",
      "postDate": "03/11/2020 06:25:15",
      "content": "<p>Actually, the model setup is pretty much the same in TPU as GPU. As long as the model is built in strategy scope, then everything should be fine. But the data flow/preprocessing/augmentation must to be changed for tfrec format.</p>",
      "rawMarkdown": "Actually, the model setup is pretty much the same in TPU as GPU. As long as the model is built in strategy scope, then everything should be fine. But the data flow/preprocessing/augmentation must to be changed for tfrec format.",
      "votes": null
    },
    {
      "id": "768788",
      "postDate": "03/11/2020 07:32:16",
      "content": "<p>I tried GeM, and in the first few epoch loss functions have been nan, so I gave up...</p>",
      "rawMarkdown": "I tried GeM, and in the first few epoch loss functions have been nan, so I gave up...",
      "votes": null
    },
    {
      "id": "768791",
      "postDate": "03/11/2020 07:36:23",
      "content": "<p>It did happen to me at the first time, but after I modify the initial value, the problem disappeared. \nMy implementation in tensorflow/keras\n```\nclass Generalized_mean_pooling2D(tf.keras.layers.Layer):\n    def <strong>init</strong>(self, p=3, epsilon=1e-6, name='', **kwargs):\n      super(Generalized_mean_pooling2D, self).<strong>init</strong>(name, **kwargs)</p>\n\n<pre><code>  self.init_p = p\n  self.epsilon = epsilon\n\ndef build(self, input_shape):\n\n  if isinstance(input_shape, list) or len(input_shape) != 4:\n    raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n\n  self.build_shape = input_shape\n\n  self.p = self.add_weight(\n          name='p',\n          shape=[1,],\n          initializer=tf.keras.initializers.Constant(value=self.init_p),\n          regularizer=None,\n          trainable=True,\n          dtype=tf.float32\n          )\n\n  self.built=True\n\ndef call(self, inputs):\n  input_shape = inputs.get_shape()\n  if isinstance(inputs, list) or len(input_shape) != 4:\n    raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n  return (tf.reduce_mean(tf.abs(inputs**self.p), axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "It did happen to me at the first time, but after I modify the initial value, the problem disappeared. \nMy implementation in tensorflow/keras\n```\nclass Generalized_mean_pooling2D(tf.keras.layers.Layer):\n    def __init__(self, p=3, epsilon=1e-6, name='', **kwargs):\n      super(Generalized_mean_pooling2D, self).__init__(name, **kwargs)\n\n      self.init_p = p\n      self.epsilon = epsilon\n\n    def build(self, input_shape):\n\n      if isinstance(input_shape, list) or len(input_shape) != 4:\n        raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n\n      self.build_shape = input_shape\n\n      self.p = self.add_weight(\n              name='p',\n              shape=[1,],\n              initializer=tf.keras.initializers.Constant(value=self.init_p),\n              regularizer=None,\n              trainable=True,\n              dtype=tf.float32\n              )\n\n      self.built=True\n\n    def call(self, inputs):\n      input_shape = inputs.get_shape()\n      if isinstance(inputs, list) or len(input_shape) != 4:\n        raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n      return (tf.reduce_mean(tf.abs(inputs**self.p), axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n```",
      "votes": null
    },
    {
      "id": "768792",
      "postDate": "03/11/2020 07:38:20",
      "content": "<p>Thanks, i will try again😁 </p>",
      "rawMarkdown": "Thanks, i will try again😁",
      "votes": null
    },
    {
      "id": "768794",
      "postDate": "03/11/2020 07:40:25",
      "content": "<p>My model finished convergence  after 180 epochs, and the CV score decreased...😂 </p>",
      "rawMarkdown": "My model finished convergence  after 180 epochs, and the CV score decreased...😂",
      "votes": null
    },
    {
      "id": "768824",
      "postDate": "03/11/2020 08:17:09",
      "content": "<p>Thanks for sharing these augmentations. Have you tried implementing RICAP augmentation?</p>\n\n<p>Also another concern for me is the high score for TPU modelling with fork of <a href=\"/seesee\">@seesee</a> kernel may suffer from overfitting as test set is just 36 images</p>",
      "rawMarkdown": "Thanks for sharing these augmentations. Have you tried implementing RICAP augmentation?\n\nAlso another concern for me is the high score for TPU modelling with fork of @seesee kernel may suffer from overfitting as test set is just 36 images",
      "votes": null
    },
    {
      "id": "768828",
      "postDate": "03/11/2020 08:21:21",
      "content": "<p>Certain augmentations like Rotation and blur doesn't make sense for this competition, as rotation images can change the meaning of Bengali characters. Cutmix, Mixup and all works as mentioned here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127902\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127902</a></p>",
      "rawMarkdown": "Certain augmentations like Rotation and blur doesn't make sense for this competition, as rotation images can change the meaning of Bengali characters. Cutmix, Mixup and all works as mentioned here: https://www.kaggle.com/c/bengaliai-cv19/discussion/127902",
      "votes": null
    },
    {
      "id": "769118",
      "postDate": "03/11/2020 14:45:53",
      "content": "<p>Great! your gridmask implementation was very useful.\nHave you found any improvement with GeM??</p>",
      "rawMarkdown": "Great! your gridmask implementation was very useful.\nHave you found any improvement with GeM??",
      "votes": null
    },
    {
      "id": "769391",
      "postDate": "03/11/2020 21:26:53",
      "content": "<p>For me, it was almost same as GAP.</p>",
      "rawMarkdown": "For me, it was almost same as GAP.",
      "votes": null
    },
    {
      "id": "769484",
      "postDate": "03/12/2020 00:36:50",
      "content": "<p>Thanks a lot for your remainder and sharing! I will give them a try!</p>",
      "rawMarkdown": "Thanks a lot for your remainder and sharing! I will give them a try!",
      "votes": null
    },
    {
      "id": "769485",
      "postDate": "03/12/2020 00:37:41",
      "content": "<p>The validation dataset in his tfrec is 20% of total training dataset. So I think it should be okay?</p>",
      "rawMarkdown": "The validation dataset in his tfrec is 20% of total training dataset. So I think it should be okay?",
      "votes": null
    },
    {
      "id": "769486",
      "postDate": "03/12/2020 00:38:35",
      "content": "<p>Thanks! GeM seems didn't work for me either.. Still searching the magic😂 </p>",
      "rawMarkdown": "Thanks! GeM seems didn't work for me either.. Still searching the magic😂",
      "votes": null
    },
    {
      "id": "770401",
      "postDate": "03/12/2020 22:00:49",
      "content": "<p>GeM worked for me as in it works, but i didnt see clear improvement.</p>",
      "rawMarkdown": "GeM worked for me as in it works, but i didnt see clear improvement.",
      "votes": null
    },
    {
      "id": "770478",
      "postDate": "03/13/2020 01:54:47",
      "content": "<p>Yes, same results here, very little improvement. But I need to increase the learning rate for faster convergence with GeM pooling tail. </p>",
      "rawMarkdown": "Yes, same results here, very little improvement. But I need to increase the learning rate for faster convergence with GeM pooling tail.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768684,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "03/11/2020 04:56:20",
      "content": "<p>Awesome. Thank you for sharing such great notebooks. Are you using tfrec format data set? </p>",
      "votes": null,
      "replies": [
        {
          "id": 768697,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/11/2020 05:13:56",
          "content": "<p>Yes, since the competition is going to end, so I need to use TPU for fast model training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768721,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "03/11/2020 05:43:17",
          "content": "<p>I was planning to use TPU since hearing when its available but couldn't make it :( But now, See-- had made an awesome starter but because of the last moment of the competition, I stick with my set up. I was planning to use CpasNet on TPU to get the leverage of faster training. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768748,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/11/2020 06:25:15",
          "content": "<p>Actually, the model setup is pretty much the same in TPU as GPU. As long as the model is built in strategy scope, then everything should be fine. But the data flow/preprocessing/augmentation must to be changed for tfrec format.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768788,
      "author_name": "noxuslol",
      "author_url": "",
      "post_date": "03/11/2020 07:32:16",
      "content": "<p>I tried GeM, and in the first few epoch loss functions have been nan, so I gave up...</p>",
      "votes": null,
      "replies": [
        {
          "id": 768791,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/11/2020 07:36:23",
          "content": "<p>It did happen to me at the first time, but after I modify the initial value, the problem disappeared. \nMy implementation in tensorflow/keras\n```\nclass Generalized_mean_pooling2D(tf.keras.layers.Layer):\n    def <strong>init</strong>(self, p=3, epsilon=1e-6, name='', **kwargs):\n      super(Generalized_mean_pooling2D, self).<strong>init</strong>(name, **kwargs)</p>\n\n<pre><code>  self.init_p = p\n  self.epsilon = epsilon\n\ndef build(self, input_shape):\n\n  if isinstance(input_shape, list) or len(input_shape) != 4:\n    raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n\n  self.build_shape = input_shape\n\n  self.p = self.add_weight(\n          name='p',\n          shape=[1,],\n          initializer=tf.keras.initializers.Constant(value=self.init_p),\n          regularizer=None,\n          trainable=True,\n          dtype=tf.float32\n          )\n\n  self.built=True\n\ndef call(self, inputs):\n  input_shape = inputs.get_shape()\n  if isinstance(inputs, list) or len(input_shape) != 4:\n    raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n  return (tf.reduce_mean(tf.abs(inputs**self.p), axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768792,
          "author_name": "noxuslol",
          "author_url": "",
          "post_date": "03/11/2020 07:38:20",
          "content": "<p>Thanks, i will try again😁 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768794,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/11/2020 07:40:25",
          "content": "<p>My model finished convergence  after 180 epochs, and the CV score decreased...😂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768824,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "03/11/2020 08:17:09",
      "content": "<p>Thanks for sharing these augmentations. Have you tried implementing RICAP augmentation?</p>\n\n<p>Also another concern for me is the high score for TPU modelling with fork of <a href=\"/seesee\">@seesee</a> kernel may suffer from overfitting as test set is just 36 images</p>",
      "votes": null,
      "replies": [
        {
          "id": 769485,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/12/2020 00:37:41",
          "content": "<p>The validation dataset in his tfrec is 20% of total training dataset. So I think it should be okay?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768828,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "03/11/2020 08:21:21",
      "content": "<p>Certain augmentations like Rotation and blur doesn't make sense for this competition, as rotation images can change the meaning of Bengali characters. Cutmix, Mixup and all works as mentioned here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127902\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127902</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 769484,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/12/2020 00:36:50",
          "content": "<p>Thanks a lot for your remainder and sharing! I will give them a try!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 769118,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "03/11/2020 14:45:53",
      "content": "<p>Great! your gridmask implementation was very useful.\nHave you found any improvement with GeM??</p>",
      "votes": null,
      "replies": [
        {
          "id": 769391,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "03/11/2020 21:26:53",
          "content": "<p>For me, it was almost same as GAP.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 769486,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/12/2020 00:38:35",
          "content": "<p>Thanks! GeM seems didn't work for me either.. Still searching the magic😂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 770401,
      "author_name": "yl1202",
      "author_url": "",
      "post_date": "03/12/2020 22:00:49",
      "content": "<p>GeM worked for me as in it works, but i didnt see clear improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 770478,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "03/13/2020 01:54:47",
          "content": "<p>Yes, same results here, very little improvement. But I need to increase the learning rate for faster convergence with GeM pooling tail. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "768627": "Hi Guys:\nJust joined the competition 2 days ago, first of all, thanks for all the great knowledge sharing of you. Since the information in discussion section just too much, thus I am not able to go through them very carefully. So I'm not sure anyone discussed about this before, but it is just me or the GeM pooling really make the model converge slow? I switched to GeM pooling yesterday, and after 120 epochs, the model still not convergence. Normally my model converge in around 60 epochs, so just wondering if there are same situation as mine right now.  \n\nAlso, I started this competition after @seesee created the tfrec training data. But I knew that we can't use augmentation modules like albumentations or ImageDataGenerator with tfrec format. Just want to share you guys some links about how to do augmentation with tfrec format.\n\n[Rotation Augmentation GPU/TPU shared by Chris](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96)\n[CutMix and MixUp on GPU/TPU shared by Chris](https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu)\n[GridMask data augmentation with tensorflow shared by me](https://www.kaggle.com/xiejialun/gridmask-data-augmentation-with-tensorflow)\n[Gaussian blur and Cutout augmentation shared by me](https://www.kaggle.com/xiejialun/customize-data-augmentation-with-tensorflow)\n[Batch implementation of above augmentation shared by Yih-Dar SHIEH](https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations)",
    "768684": "Awesome. Thank you for sharing such great notebooks. Are you using tfrec format data set?",
    "768697": "Yes, since the competition is going to end, so I need to use TPU for fast model training.",
    "768721": "I was planning to use TPU since hearing when its available but couldn't make it :( But now, See-- had made an awesome starter but because of the last moment of the competition, I stick with my set up. I was planning to use CpasNet on TPU to get the leverage of faster training.",
    "768748": "Actually, the model setup is pretty much the same in TPU as GPU. As long as the model is built in strategy scope, then everything should be fine. But the data flow/preprocessing/augmentation must to be changed for tfrec format.",
    "768788": "I tried GeM, and in the first few epoch loss functions have been nan, so I gave up...",
    "768791": "It did happen to me at the first time, but after I modify the initial value, the problem disappeared. \nMy implementation in tensorflow/keras\n```\nclass Generalized_mean_pooling2D(tf.keras.layers.Layer):\n    def __init__(self, p=3, epsilon=1e-6, name='', **kwargs):\n      super(Generalized_mean_pooling2D, self).__init__(name, **kwargs)\n\n      self.init_p = p\n      self.epsilon = epsilon\n\n    def build(self, input_shape):\n\n      if isinstance(input_shape, list) or len(input_shape) != 4:\n        raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n\n      self.build_shape = input_shape\n\n      self.p = self.add_weight(\n              name='p',\n              shape=[1,],\n              initializer=tf.keras.initializers.Constant(value=self.init_p),\n              regularizer=None,\n              trainable=True,\n              dtype=tf.float32\n              )\n\n      self.built=True\n\n    def call(self, inputs):\n      input_shape = inputs.get_shape()\n      if isinstance(inputs, list) or len(input_shape) != 4:\n        raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n\n      return (tf.reduce_mean(tf.abs(inputs**self.p), axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n```",
    "768792": "Thanks, i will try again😁",
    "768794": "My model finished convergence  after 180 epochs, and the CV score decreased...😂",
    "768824": "Thanks for sharing these augmentations. Have you tried implementing RICAP augmentation?\n\nAlso another concern for me is the high score for TPU modelling with fork of @seesee kernel may suffer from overfitting as test set is just 36 images",
    "768828": "Certain augmentations like Rotation and blur doesn't make sense for this competition, as rotation images can change the meaning of Bengali characters. Cutmix, Mixup and all works as mentioned here: https://www.kaggle.com/c/bengaliai-cv19/discussion/127902",
    "769118": "Great! your gridmask implementation was very useful.\nHave you found any improvement with GeM??",
    "769391": "For me, it was almost same as GAP.",
    "769484": "Thanks a lot for your remainder and sharing! I will give them a try!",
    "769485": "The validation dataset in his tfrec is 20% of total training dataset. So I think it should be okay?",
    "769486": "Thanks! GeM seems didn't work for me either.. Still searching the magic😂",
    "770401": "GeM worked for me as in it works, but i didnt see clear improvement.",
    "770478": "Yes, same results here, very little improvement. But I need to increase the learning rate for faster convergence with GeM pooling tail."
  },
  "source": "meta"
}