{
  "id": 135985,
  "title": "9th Solution",
  "url": "/competitions/bengaliai-cv19/discussion/135985",
  "author_name": "Weimin Wang",
  "post_date": "2020-03-17T01:40:17.070000",
  "votes": 45,
  "comment_count": 14,
  "views": 0,
  "content": "<p><strong>Model</strong> \nMy model backbone is simply based on seresnext50, using CV of randomly split 4 folds for training, no stratification. However, i have two extra loss heads which I found quite helpful to regularize the model - 1292-grapheme head CE loss and a Arcface loss of the grapheme head. </p>\n\n<p>However, the 1292 head is on top of the combined 3 tasks heads with binary mapping - </p>\n\n<p><code>self.logit_to_1292 = nn.Linear(168+11+7, 1292, bias=None)</code></p>\n\n<p><code>out_1292 = self.logit_to_1292(out_3tasks)</code></p>\n\n<p>where I initialize the <code>self.logit_to_1292</code> layer with a frozen binary weight that has dimension of (168+11+7, 1292), where it correctly maps each three combination of (grapheme_root, vowel_diacritic, consonant_diacritic) to its corresponding grapheme. I don't train this layer, however.</p>\n\n<p>During back-propagation, I also used a very low weighting for the 3 tasks head and arcface head (0.02), but have a very high weight for grapheme head (1), which is quite counter-intuitive. </p>\n\n<p><strong>Augmentation</strong></p>\n\n<p>I used 75% of CutMix, sample-wise. Also I gradually add in grid mask as training goes on, but I used 'grid-mix' instead - so instead of removing those grids, I randomly shuffle them within each image, so the per image stats won't change comparing to if you crop out those grids from image.  I found this slightly more helpful. </p>\n\n<p>For sample-wise cut mix (this is the most helpful data aug) - each image will randomly draw a binary cutmix mask during pytorch's data loading function, and I only shuffle and pair up the images within each batch during training, using this pre-computed mask. </p>\n\n<p>I also add in 5% of augmentation proposed by XingJian Lyu. </p>\n\n<p>Overall cutmix is the most useful part. </p>\n\n<p>*<em>Self-training with external dataset \n*</em>\nI would say this is the part that helps my model much in generalizing to the test set. I only discovered the external dataset of ekush disclosed on the external data disclosure thread on the last week of the competition. </p>\n\n<p>Basically I downloaded and resize those 300K unlabeled images, and used my best trained 4 folds model to generate soft targets for the grapheme (1292) and 3 tasks heads (168+11+7). Then combined with original training data, I re-train my above model in 4 folds, with soft targets as labels. In each fold, I combine the entire 300K external data plus the training folds to train each model, and do this 4 times. </p>\n\n<p>This approach is similar to the idea of noisy students self-training paper shared earlier on in the forum, where I generated soft targets without noise, and retrain the model with noise. The student and teacher models are exactly same as above. </p>\n\n<p>The key thing is how you do sampling with those unlabeled dataset - the one that works best is first to generate hard labels for graphemes, and then combined with our train dataset, to use Heng's balanced sampler to do sampling based on all graphemes. Other sampling strategies simply don't work. </p>\n\n<p>This finally pushes my model from 994 to 996 on Local, and on LB to improved from 985 to 990, where ensembling of the 4 folds further boosts it to my current score of 916 (all based on public score) </p>\n\n<p><strong>Things didn't work for me</strong></p>\n\n<p>OHEM\nEfficientnet (slightly worse than seresnext50)\nsenet152\ndropblock\nattentiondrop\nshake-drop\nintensive-augmentation or autoaugment\nusing flip or 180 rotate as addition identities (strategies that worked best for humpback competition)</p>",
  "messages": [
    {
      "id": 775843,
      "postDate": "2020-03-17T01:40:17.070Z",
      "content": "<p><strong>Model</strong> \nMy model backbone is simply based on seresnext50, using CV of randomly split 4 folds for training, no stratification. However, i have two extra loss heads which I found quite helpful to regularize the model - 1292-grapheme head CE loss and a Arcface loss of the grapheme head. </p>\n\n<p>However, the 1292 head is on top of the combined 3 tasks heads with binary mapping - </p>\n\n<p><code>self.logit_to_1292 = nn.Linear(168+11+7, 1292, bias=None)</code></p>\n\n<p><code>out_1292 = self.logit_to_1292(out_3tasks)</code></p>\n\n<p>where I initialize the <code>self.logit_to_1292</code> layer with a frozen binary weight that has dimension of (168+11+7, 1292), where it correctly maps each three combination of (grapheme_root, vowel_diacritic, consonant_diacritic) to its corresponding grapheme. I don't train this layer, however.</p>\n\n<p>During back-propagation, I also used a very low weighting for the 3 tasks head and arcface head (0.02), but have a very high weight for grapheme head (1), which is quite counter-intuitive. </p>\n\n<p><strong>Augmentation</strong></p>\n\n<p>I used 75% of CutMix, sample-wise. Also I gradually add in grid mask as training goes on, but I used 'grid-mix' instead - so instead of removing those grids, I randomly shuffle them within each image, so the per image stats won't change comparing to if you crop out those grids from image.  I found this slightly more helpful. </p>\n\n<p>For sample-wise cut mix (this is the most helpful data aug) - each image will randomly draw a binary cutmix mask during pytorch's data loading function, and I only shuffle and pair up the images within each batch during training, using this pre-computed mask. </p>\n\n<p>I also add in 5% of augmentation proposed by XingJian Lyu. </p>\n\n<p>Overall cutmix is the most useful part. </p>\n\n<p>*<em>Self-training with external dataset \n*</em>\nI would say this is the part that helps my model much in generalizing to the test set. I only discovered the external dataset of ekush disclosed on the external data disclosure thread on the last week of the competition. </p>\n\n<p>Basically I downloaded and resize those 300K unlabeled images, and used my best trained 4 folds model to generate soft targets for the grapheme (1292) and 3 tasks heads (168+11+7). Then combined with original training data, I re-train my above model in 4 folds, with soft targets as labels. In each fold, I combine the entire 300K external data plus the training folds to train each model, and do this 4 times. </p>\n\n<p>This approach is similar to the idea of noisy students self-training paper shared earlier on in the forum, where I generated soft targets without noise, and retrain the model with noise. The student and teacher models are exactly same as above. </p>\n\n<p>The key thing is how you do sampling with those unlabeled dataset - the one that works best is first to generate hard labels for graphemes, and then combined with our train dataset, to use Heng's balanced sampler to do sampling based on all graphemes. Other sampling strategies simply don't work. </p>\n\n<p>This finally pushes my model from 994 to 996 on Local, and on LB to improved from 985 to 990, where ensembling of the 4 folds further boosts it to my current score of 916 (all based on public score) </p>\n\n<p><strong>Things didn't work for me</strong></p>\n\n<p>OHEM\nEfficientnet (slightly worse than seresnext50)\nsenet152\ndropblock\nattentiondrop\nshake-drop\nintensive-augmentation or autoaugment\nusing flip or 180 rotate as addition identities (strategies that worked best for humpback competition)</p>",
      "rawMarkdown": "**Model** \nMy model backbone is simply based on seresnext50, using CV of randomly split 4 folds for training, no stratification. However, i have two extra loss heads which I found quite helpful to regularize the model - 1292-grapheme head CE loss and a Arcface loss of the grapheme head. \n\nHowever, the 1292 head is on top of the combined 3 tasks heads with binary mapping - \n\n ```self.logit_to_1292 = nn.Linear(168+11+7, 1292, bias=None)```\n \n```out_1292 = self.logit_to_1292(out_3tasks)```\n\nwhere I initialize the `self.logit_to_1292` layer with a frozen binary weight that has dimension of (168+11+7, 1292), where it correctly maps each three combination of (grapheme_root, vowel_diacritic, consonant_diacritic) to its corresponding grapheme. I don't train this layer, however.\n\nDuring back-propagation, I also used a very low weighting for the 3 tasks head and arcface head (0.02), but have a very high weight for grapheme head (1), which is quite counter-intuitive. \n\n\n**Augmentation**\n\nI used 75% of CutMix, sample-wise. Also I gradually add in grid mask as training goes on, but I used 'grid-mix' instead - so instead of removing those grids, I randomly shuffle them within each image, so the per image stats won't change comparing to if you crop out those grids from image.  I found this slightly more helpful. \n\nFor sample-wise cut mix (this is the most helpful data aug) - each image will randomly draw a binary cutmix mask during pytorch's data loading function, and I only shuffle and pair up the images within each batch during training, using this pre-computed mask. \n\nI also add in 5% of augmentation proposed by XingJian Lyu. \n\nOverall cutmix is the most useful part. \n\n\n**Self-training with external dataset \n**\nI would say this is the part that helps my model much in generalizing to the test set. I only discovered the external dataset of ekush disclosed on the external data disclosure thread on the last week of the competition. \n\nBasically I downloaded and resize those 300K unlabeled images, and used my best trained 4 folds model to generate soft targets for the grapheme (1292) and 3 tasks heads (168+11+7). Then combined with original training data, I re-train my above model in 4 folds, with soft targets as labels. In each fold, I combine the entire 300K external data plus the training folds to train each model, and do this 4 times. \n\nThis approach is similar to the idea of noisy students self-training paper shared earlier on in the forum, where I generated soft targets without noise, and retrain the model with noise. The student and teacher models are exactly same as above. \n\nThe key thing is how you do sampling with those unlabeled dataset - the one that works best is first to generate hard labels for graphemes, and then combined with our train dataset, to use Heng's balanced sampler to do sampling based on all graphemes. Other sampling strategies simply don't work. \n\nThis finally pushes my model from 994 to 996 on Local, and on LB to improved from 985 to 990, where ensembling of the 4 folds further boosts it to my current score of 916 (all based on public score) \n\n**Things didn't work for me**\n\nOHEM\nEfficientnet (slightly worse than seresnext50)\nsenet152\ndropblock\nattentiondrop\nshake-drop\nintensive-augmentation or autoaugment\nusing flip or 180 rotate as addition identities (strategies that worked best for humpback competition)\n\n\n",
      "votes": 44
    },
    {
      "id": 782509,
      "postDate": "2020-03-22T11:20:05.277Z",
      "content": "<p>Question for anyone reading this: When using k-folds for training, do you average the predictions from each fold's model during inference?</p>",
      "rawMarkdown": "Question for anyone reading this: When using k-folds for training, do you average the predictions from each fold's model during inference?"
    },
    {
      "id": 776552,
      "postDate": "2020-03-17T13:21:02.833Z",
      "content": "<p>Congrats.  Interesting that ekush data helped, we thought the low resolution of images would not be helpful.</p>",
      "rawMarkdown": "Congrats.  Interesting that ekush data helped, we thought the low resolution of images would not be helpful.\n",
      "replies": [
        {
          "id": 776729,
          "postDate": "2020-03-17T15:39:49.513Z",
          "content": "<p>Thanks. I thought that as well orignally, but experiments showed differently  :)</p>",
          "rawMarkdown": "Thanks. I thought that as well orignally, but experiments showed differently  :)"
        }
      ]
    },
    {
      "id": 776227,
      "postDate": "2020-03-17T08:05:25.550Z",
      "content": "<p>Congrats. If I don't miss anything, the way you're using the unlabeled dataset is much like psuedo-labeling.</p>\n\n<p>&gt; The key thing is how you do sampling with those unlabeled dataset </p>\n\n<p>Could you explain more about this? I'm confused about what the 'sampling' is doing. Is this used to generate different training folds with unlabeled dataset?</p>",
      "rawMarkdown": "Congrats. If I don't miss anything, the way you're using the unlabeled dataset is much like psuedo-labeling.\n\n&gt; The key thing is how you do sampling with those unlabeled dataset \n\nCould you explain more about this? I'm confused about what the 'sampling' is doing. Is this used to generate different training folds with unlabeled dataset?",
      "replies": [
        {
          "id": 776724,
          "postDate": "2020-03-17T15:37:50.133Z",
          "content": "<p>yes, like PL. since I have inference for 'grapheme' as well, I will do argmax to get the hard label for 'grapheme' for all external dataset. Then after I combined with original training data to retrain the model, I will do balanced sampling during training based on the 'grapheme' for all data. Let me know if this is not clear</p>",
          "rawMarkdown": "yes, like PL. since I have inference for 'grapheme' as well, I will do argmax to get the hard label for 'grapheme' for all external dataset. Then after I combined with original training data to retrain the model, I will do balanced sampling during training based on the 'grapheme' for all data. Let me know if this is not clear"
        },
        {
          "id": 782526,
          "postDate": "2020-03-22T11:43:39.930Z",
          "content": "<p>Isn't there a decent chance that the predicted grapheme label might be wrong for the external data since the model can only predict 1292 combinations?</p>",
          "rawMarkdown": "Isn't there a decent chance that the predicted grapheme label might be wrong for the external data since the model can only predict 1292 combinations?"
        }
      ]
    },
    {
      "id": 776105,
      "postDate": "2020-03-17T05:52:30.230Z",
      "content": "<p>Congrats! Thanks for sharing! You've done plenty of experiments! Impresive!!! May I ask how do you choose your back-propagation weights? </p>",
      "rawMarkdown": "Congrats! Thanks for sharing! You've done plenty of experiments! Impresive!!! May I ask how do you choose your back-propagation weights? ",
      "replies": [
        {
          "id": 776726,
          "postDate": "2020-03-17T15:38:56.883Z",
          "content": "<p>I tried 3 - 1*grapheme + 0.1 * three tasks, 1 * grapheme + 1 * three tasks, 0.1 * grapheme + 1 * three tasks, and found the first way works best </p>",
          "rawMarkdown": "I tried 3 - 1*grapheme + 0.1 * three tasks, 1 * grapheme + 1 * three tasks, 0.1 * grapheme + 1 * three tasks, and found the first way works best "
        }
      ]
    },
    {
      "id": 775906,
      "postDate": "2020-03-17T02:41:41.500Z",
      "content": "<p>Congratulation and thank you for sharing.</p>\n\n<p>I wonder if there are any special tricks with cutmix because I could not get any good result with cutmix.\n- what should be the correct preprocessing operations for cutmix?  (I assume you want to crop and resize the image properly so that when doing random box cut, the box contain significant information). <br>\n- What alpha you choose for the beta distribution. The way I see it is that if the alpha is too small, most of time your cut portion will be very small (or very large, in this case first image will be mostly cropped out).  Depending how you crop/resize your image at beginning, the cut portion can be useless to make any prediction.    If the alpha is too large,   it may underfit the model. </p>",
      "rawMarkdown": "Congratulation and thank you for sharing.\n\nI wonder if there are any special tricks with cutmix because I could not get any good result with cutmix.\n- what should be the correct preprocessing operations for cutmix?  (I assume you want to crop and resize the image properly so that when doing random box cut, the box contain significant information).   \n- What alpha you choose for the beta distribution. The way I see it is that if the alpha is too small, most of time your cut portion will be very small (or very large, in this case first image will be mostly cropped out).  Depending how you crop/resize your image at beginning, the cut portion can be useless to make any prediction.    If the alpha is too large,   it may underfit the model. \n\n\n",
      "replies": [
        {
          "id": 775932,
          "postDate": "2020-03-17T03:03:03.010Z",
          "content": "<p>I used a simplified way to sample the cutmix box: </p>\n\n<p><code>shorter = min(img_shape[0], img_shape[1])</code>\n<code>cut_size = np.random.randint(int(shorter*0.1), int(shorter*0.8))</code>\n<code>bbx1, bby1, bbx2, bby2 = rand_bbox_center_square(img_shape, cut_size)</code></p>\n\n<p>I found the hyper params for cut mix are not sensitive. beta distribution is not needed, just cut the square box uniformly random</p>",
          "rawMarkdown": "I used a simplified way to sample the cutmix box: \n\n`shorter = min(img_shape[0], img_shape[1])`\n`cut_size = np.random.randint(int(shorter*0.1), int(shorter*0.8))`\n`bbx1, bby1, bbx2, bby2 = rand_bbox_center_square(img_shape, cut_size)`\n\nI found the hyper params for cut mix are not sensitive. beta distribution is not needed, just cut the square box uniformly random\n\n"
        }
      ]
    },
    {
      "id": 779369,
      "postDate": "2020-03-19T09:17:02.973Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 775846,
      "postDate": "2020-03-17T01:46:58.023Z",
      "content": "<p>Thank you for sharing. </p>",
      "rawMarkdown": "Thank you for sharing. ",
      "votes": 1
    },
    {
      "id": 780443,
      "postDate": "2020-03-20T09:17:25.010Z",
      "content": "<p>Congrats, and thanks for sharing. 🙂 </p>",
      "rawMarkdown": "Congrats, and thanks for sharing. 🙂 "
    },
    {
      "id": 777848,
      "postDate": "2020-03-18T01:49:46.083Z",
      "content": "<p>Congrats, thank you for sharing!</p>",
      "rawMarkdown": "Congrats, thank you for sharing!"
    }
  ],
  "comments": [
    {
      "id": 782509,
      "author_name": "Syed Saad",
      "author_url": "",
      "post_date": "2020-03-22T11:20:05.277000",
      "content": "<p>Question for anyone reading this: When using k-folds for training, do you average the predictions from each fold's model during inference?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776552,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-03-17T13:21:02.833000",
      "content": "<p>Congrats.  Interesting that ekush data helped, we thought the low resolution of images would not be helpful.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776729,
          "author_name": "Weimin Wang",
          "author_url": "",
          "post_date": "2020-03-17T15:39:49.513000",
          "content": "<p>Thanks. I thought that as well orignally, but experiments showed differently  :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776227,
      "author_name": "syoya",
      "author_url": "",
      "post_date": "2020-03-17T08:05:25.550000",
      "content": "<p>Congrats. If I don't miss anything, the way you're using the unlabeled dataset is much like psuedo-labeling.</p>\n\n<p>&gt; The key thing is how you do sampling with those unlabeled dataset </p>\n\n<p>Could you explain more about this? I'm confused about what the 'sampling' is doing. Is this used to generate different training folds with unlabeled dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776724,
          "author_name": "Weimin Wang",
          "author_url": "",
          "post_date": "2020-03-17T15:37:50.133000",
          "content": "<p>yes, like PL. since I have inference for 'grapheme' as well, I will do argmax to get the hard label for 'grapheme' for all external dataset. Then after I combined with original training data to retrain the model, I will do balanced sampling during training based on the 'grapheme' for all data. Let me know if this is not clear</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 782526,
          "author_name": "Syed Saad",
          "author_url": "",
          "post_date": "2020-03-22T11:43:39.930000",
          "content": "<p>Isn't there a decent chance that the predicted grapheme label might be wrong for the external data since the model can only predict 1292 combinations?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776105,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-17T05:52:30.230000",
      "content": "<p>Congrats! Thanks for sharing! You've done plenty of experiments! Impresive!!! May I ask how do you choose your back-propagation weights? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 776726,
          "author_name": "Weimin Wang",
          "author_url": "",
          "post_date": "2020-03-17T15:38:56.883000",
          "content": "<p>I tried 3 - 1*grapheme + 0.1 * three tasks, 1 * grapheme + 1 * three tasks, 0.1 * grapheme + 1 * three tasks, and found the first way works best </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 775906,
      "author_name": "sandiudiu",
      "author_url": "",
      "post_date": "2020-03-17T02:41:41.500000",
      "content": "<p>Congratulation and thank you for sharing.</p>\n\n<p>I wonder if there are any special tricks with cutmix because I could not get any good result with cutmix.\n- what should be the correct preprocessing operations for cutmix?  (I assume you want to crop and resize the image properly so that when doing random box cut, the box contain significant information). <br>\n- What alpha you choose for the beta distribution. The way I see it is that if the alpha is too small, most of time your cut portion will be very small (or very large, in this case first image will be mostly cropped out).  Depending how you crop/resize your image at beginning, the cut portion can be useless to make any prediction.    If the alpha is too large,   it may underfit the model. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 775932,
          "author_name": "Weimin Wang",
          "author_url": "",
          "post_date": "2020-03-17T03:03:03.010000",
          "content": "<p>I used a simplified way to sample the cutmix box: </p>\n\n<p><code>shorter = min(img_shape[0], img_shape[1])</code>\n<code>cut_size = np.random.randint(int(shorter*0.1), int(shorter*0.8))</code>\n<code>bbx1, bby1, bbx2, bby2 = rand_bbox_center_square(img_shape, cut_size)</code></p>\n\n<p>I found the hyper params for cut mix are not sensitive. beta distribution is not needed, just cut the square box uniformly random</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 779369,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-19T09:17:02.973000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 775846,
      "author_name": "MachineLP",
      "author_url": "",
      "post_date": "2020-03-17T01:46:58.023000",
      "content": "<p>Thank you for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 780443,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-20T09:17:25.010000",
      "content": "<p>Congrats, and thanks for sharing. 🙂 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 777848,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T01:49:46.083000",
      "content": "<p>Congrats, thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775843": "**Model** \nMy model backbone is simply based on seresnext50, using CV of randomly split 4 folds for training, no stratification. However, i have two extra loss heads which I found quite helpful to regularize the model - 1292-grapheme head CE loss and a Arcface loss of the grapheme head. \n\nHowever, the 1292 head is on top of the combined 3 tasks heads with binary mapping - \n\n ```self.logit_to_1292 = nn.Linear(168+11+7, 1292, bias=None)```\n \n```out_1292 = self.logit_to_1292(out_3tasks)```\n\nwhere I initialize the `self.logit_to_1292` layer with a frozen binary weight that has dimension of (168+11+7, 1292), where it correctly maps each three combination of (grapheme_root, vowel_diacritic, consonant_diacritic) to its corresponding grapheme. I don't train this layer, however.\n\nDuring back-propagation, I also used a very low weighting for the 3 tasks head and arcface head (0.02), but have a very high weight for grapheme head (1), which is quite counter-intuitive. \n\n\n**Augmentation**\n\nI used 75% of CutMix, sample-wise. Also I gradually add in grid mask as training goes on, but I used 'grid-mix' instead - so instead of removing those grids, I randomly shuffle them within each image, so the per image stats won't change comparing to if you crop out those grids from image.  I found this slightly more helpful. \n\nFor sample-wise cut mix (this is the most helpful data aug) - each image will randomly draw a binary cutmix mask during pytorch's data loading function, and I only shuffle and pair up the images within each batch during training, using this pre-computed mask. \n\nI also add in 5% of augmentation proposed by XingJian Lyu. \n\nOverall cutmix is the most useful part. \n\n\n**Self-training with external dataset \n**\nI would say this is the part that helps my model much in generalizing to the test set. I only discovered the external dataset of ekush disclosed on the external data disclosure thread on the last week of the competition. \n\nBasically I downloaded and resize those 300K unlabeled images, and used my best trained 4 folds model to generate soft targets for the grapheme (1292) and 3 tasks heads (168+11+7). Then combined with original training data, I re-train my above model in 4 folds, with soft targets as labels. In each fold, I combine the entire 300K external data plus the training folds to train each model, and do this 4 times. \n\nThis approach is similar to the idea of noisy students self-training paper shared earlier on in the forum, where I generated soft targets without noise, and retrain the model with noise. The student and teacher models are exactly same as above. \n\nThe key thing is how you do sampling with those unlabeled dataset - the one that works best is first to generate hard labels for graphemes, and then combined with our train dataset, to use Heng's balanced sampler to do sampling based on all graphemes. Other sampling strategies simply don't work. \n\nThis finally pushes my model from 994 to 996 on Local, and on LB to improved from 985 to 990, where ensembling of the 4 folds further boosts it to my current score of 916 (all based on public score) \n\n**Things didn't work for me**\n\nOHEM\nEfficientnet (slightly worse than seresnext50)\nsenet152\ndropblock\nattentiondrop\nshake-drop\nintensive-augmentation or autoaugment\nusing flip or 180 rotate as addition identities (strategies that worked best for humpback competition)\n\n\n",
    "782509": "Question for anyone reading this: When using k-folds for training, do you average the predictions from each fold's model during inference?",
    "776552": "Congrats.  Interesting that ekush data helped, we thought the low resolution of images would not be helpful.\n",
    "776227": "Congrats. If I don't miss anything, the way you're using the unlabeled dataset is much like psuedo-labeling.\n\n&gt; The key thing is how you do sampling with those unlabeled dataset \n\nCould you explain more about this? I'm confused about what the 'sampling' is doing. Is this used to generate different training folds with unlabeled dataset?",
    "776105": "Congrats! Thanks for sharing! You've done plenty of experiments! Impresive!!! May I ask how do you choose your back-propagation weights? ",
    "775906": "Congratulation and thank you for sharing.\n\nI wonder if there are any special tricks with cutmix because I could not get any good result with cutmix.\n- what should be the correct preprocessing operations for cutmix?  (I assume you want to crop and resize the image properly so that when doing random box cut, the box contain significant information).   \n- What alpha you choose for the beta distribution. The way I see it is that if the alpha is too small, most of time your cut portion will be very small (or very large, in this case first image will be mostly cropped out).  Depending how you crop/resize your image at beginning, the cut portion can be useless to make any prediction.    If the alpha is too large,   it may underfit the model. \n\n\n",
    "779369": "",
    "775846": "Thank you for sharing. ",
    "780443": "Congrats, and thanks for sharing. 🙂 ",
    "777848": "Congrats, thank you for sharing!"
  }
}