{
  "id": 136791,
  "title": "14th Place Solution",
  "url": "/competitions/bengaliai-cv19/writeups/14th-place-solution",
  "author_name": "",
  "post_date": "2020-03-17T21:25:05.416448100Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>Congratulations and Thank You</h2>\n\n<p>This has been an interesting, challenging, and very educational experience for us. We would like to congratulate all the winners and everyone else who has worked so hard over the past three months on this competition. Hope that even those of you who have been burned by the shakeup at the end can come away from this competition with a lot of new things that you have learned. We certainly are.</p>\n\n<p>We had lots of fun teaming up with each other. Keeping each other abreast of all that we were trying to do, and encouraging each other, especially in the last phase of this competition, has been an incredible experience for us. </p>\n\n<p>We would like to thank the organizers, Bengali.ai and Kaggle, for putting this competition together. We hope that the results of this competition can be used in a meaningful and productive way for the advancement of the handwritten Bengali language. </p>\n\n<p>Chris and Bojan would like to thank Nvidia for its ongoing generous support in terms of time and resources. </p>\n\n<h2>Frameworks</h2>\n\n<p>For our final solution, we have used both PyTorch and TensorFlow. We have also worked with Rapids on an alternative solution, which showed lots of promise but was not possible to incorporate with our best solution in the allotted submission time limits. </p>\n\n<h2>Datasets</h2>\n\n<p>Most of our models used just the competition dataset. One of our models (EfficientNet B7) trained partially with some of the external data. We had also tried using synthetic data, but that did not help our local models. It is possible that it did better on the private dataset, but we have not made any submissions with it yet.</p>\n\n<h2>Image Resolutions</h2>\n\n<p>Like everyone else, throughout the competition, we had experimented with different resolutions and aspect ratios. Our final solution was a combination of the following different resolutions: 128x128, 224x224, 137x236, 128x256. Our single best model used the 256x256 resolution, but the inference with it was taking too long, so we had to make a difficult decision to leave it out of the final ensemble in favor of a more diverse blend. </p>\n\n<h2>Validation Scheme</h2>\n\n<p>Our main validation scheme has been a 90/10 train/validation split. For most of the models, we have used the same 90/10 split, but in the latter stages of the competition, we had added a few with different splits for the sake of diversity.</p>\n\n<h2>Network Architectures</h2>\n\n<p>We have tried several different architectures and discovered that for the cutmix augmentation as we had implemented it, bigger networks were better, but they required a longer time to train and larger batch sizes. All but one network were trained with a single head. In hindsight, this might have seriously limited our post-processing as we had implemented it. Our final blend consisted of the following seven network architectures:</p>\n\n<ul>\n<li>EfficientNet B4</li>\n<li>EfficentNet B6</li>\n<li>EfficientNet B7</li>\n<li>DenseNet 201</li>\n<li>SeNet 154</li>\n<li>2 SE-ResNext 101</li>\n</ul>\n\n<h2>Augmentations</h2>\n\n<p>Two main augmentations that we used were CutMix and CoarseDropout. We had also used a mild shift, scale, rotate augmentations. We also used CAM CutMix, Chris describes it <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">here</a>.</p>\n\n<h2>Training Regiment(s)</h2>\n\n<p>Our main training consisted of one cycle with starting lr of 0.0001, and a decrease at the plateau. We would train between 100 and 200 epochs, with batch sizes between 150 and 300. After the one main cycle, we would retrain for another 70 epochs, which would usually bump our validation for another 0.001. Depending on the machine, network, and batch sizes, one epoch would take between 10 and 30 minutes to train. Our single best network achieved 0.9980 for local validation, and 0.9892 on public leaderboard. </p>\n\n<h2>Ensembling</h2>\n\n<p>We’ve tried to build some advanced second level models, but our best ensembling ended up being just a weighted average of our models. Since our models were trained to predict logits and not probabilities, finding a good blend was tricky. In the end, we managed to get a pretty decent set of weights, that was the best for both oof predictions as well as the public leaderboard. </p>\n\n<h2>Post Processing</h2>\n\n<p>This was one of the most important steps in our model, and most of the models that finished at the top in this competition. It prevented us from dropping like a stone after the final shakeup. In a nutshell, it turns out that the recall metric is very sensitive to the distribution of various classes. In order to exploit this insight, it was necessary to scale various predictions in such a way that the probabilities of less frequent items got proportionally higher weight. The most mathematically reasonable way of doing this is to multiply various probabilities by their inverse frequencies. In practice, we had to deviate slightly from the perfect inverse law. We had spent a lot of time in the last days trying to improve this post processing, but in the end, a very simple set of exponent worked best for our models. You can read more about this and other aspects of our solution in <a href=\"https://www.kaggle.com/cdeotte\">Chris Deotte</a>’s <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">wonderful writeup</a>. </p>\n\n<h2>Hardware</h2>\n\n<p>The main workhorses for our model building have been two Nvidia DGX Stations which have been generously allocated by Nvidia for the purposes of Kaggle competitions. We have also relied on our own private resources, including a dual RTX Titan workstation, and several 2080 Ti workstations. </p>",
  "messages": [
    {
      "id": "777695",
      "postDate": "03/17/2020 21:25:05",
      "content": "<h2>Congratulations and Thank You</h2>\n\n<p>This has been an interesting, challenging, and very educational experience for us. We would like to congratulate all the winners and everyone else who has worked so hard over the past three months on this competition. Hope that even those of you who have been burned by the shakeup at the end can come away from this competition with a lot of new things that you have learned. We certainly are.</p>\n\n<p>We had lots of fun teaming up with each other. Keeping each other abreast of all that we were trying to do, and encouraging each other, especially in the last phase of this competition, has been an incredible experience for us. </p>\n\n<p>We would like to thank the organizers, Bengali.ai and Kaggle, for putting this competition together. We hope that the results of this competition can be used in a meaningful and productive way for the advancement of the handwritten Bengali language. </p>\n\n<p>Chris and Bojan would like to thank Nvidia for its ongoing generous support in terms of time and resources. </p>\n\n<h2>Frameworks</h2>\n\n<p>For our final solution, we have used both PyTorch and TensorFlow. We have also worked with Rapids on an alternative solution, which showed lots of promise but was not possible to incorporate with our best solution in the allotted submission time limits. </p>\n\n<h2>Datasets</h2>\n\n<p>Most of our models used just the competition dataset. One of our models (EfficientNet B7) trained partially with some of the external data. We had also tried using synthetic data, but that did not help our local models. It is possible that it did better on the private dataset, but we have not made any submissions with it yet.</p>\n\n<h2>Image Resolutions</h2>\n\n<p>Like everyone else, throughout the competition, we had experimented with different resolutions and aspect ratios. Our final solution was a combination of the following different resolutions: 128x128, 224x224, 137x236, 128x256. Our single best model used the 256x256 resolution, but the inference with it was taking too long, so we had to make a difficult decision to leave it out of the final ensemble in favor of a more diverse blend. </p>\n\n<h2>Validation Scheme</h2>\n\n<p>Our main validation scheme has been a 90/10 train/validation split. For most of the models, we have used the same 90/10 split, but in the latter stages of the competition, we had added a few with different splits for the sake of diversity.</p>\n\n<h2>Network Architectures</h2>\n\n<p>We have tried several different architectures and discovered that for the cutmix augmentation as we had implemented it, bigger networks were better, but they required a longer time to train and larger batch sizes. All but one network were trained with a single head. In hindsight, this might have seriously limited our post-processing as we had implemented it. Our final blend consisted of the following seven network architectures:</p>\n\n<ul>\n<li>EfficientNet B4</li>\n<li>EfficentNet B6</li>\n<li>EfficientNet B7</li>\n<li>DenseNet 201</li>\n<li>SeNet 154</li>\n<li>2 SE-ResNext 101</li>\n</ul>\n\n<h2>Augmentations</h2>\n\n<p>Two main augmentations that we used were CutMix and CoarseDropout. We had also used a mild shift, scale, rotate augmentations. We also used CAM CutMix, Chris describes it <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">here</a>.</p>\n\n<h2>Training Regiment(s)</h2>\n\n<p>Our main training consisted of one cycle with starting lr of 0.0001, and a decrease at the plateau. We would train between 100 and 200 epochs, with batch sizes between 150 and 300. After the one main cycle, we would retrain for another 70 epochs, which would usually bump our validation for another 0.001. Depending on the machine, network, and batch sizes, one epoch would take between 10 and 30 minutes to train. Our single best network achieved 0.9980 for local validation, and 0.9892 on public leaderboard. </p>\n\n<h2>Ensembling</h2>\n\n<p>We’ve tried to build some advanced second level models, but our best ensembling ended up being just a weighted average of our models. Since our models were trained to predict logits and not probabilities, finding a good blend was tricky. In the end, we managed to get a pretty decent set of weights, that was the best for both oof predictions as well as the public leaderboard. </p>\n\n<h2>Post Processing</h2>\n\n<p>This was one of the most important steps in our model, and most of the models that finished at the top in this competition. It prevented us from dropping like a stone after the final shakeup. In a nutshell, it turns out that the recall metric is very sensitive to the distribution of various classes. In order to exploit this insight, it was necessary to scale various predictions in such a way that the probabilities of less frequent items got proportionally higher weight. The most mathematically reasonable way of doing this is to multiply various probabilities by their inverse frequencies. In practice, we had to deviate slightly from the perfect inverse law. We had spent a lot of time in the last days trying to improve this post processing, but in the end, a very simple set of exponent worked best for our models. You can read more about this and other aspects of our solution in <a href=\"https://www.kaggle.com/cdeotte\">Chris Deotte</a>’s <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">wonderful writeup</a>. </p>\n\n<h2>Hardware</h2>\n\n<p>The main workhorses for our model building have been two Nvidia DGX Stations which have been generously allocated by Nvidia for the purposes of Kaggle competitions. We have also relied on our own private resources, including a dual RTX Titan workstation, and several 2080 Ti workstations. </p>",
      "rawMarkdown": "## Congratulations and Thank You\n\nThis has been an interesting, challenging, and very educational experience for us. We would like to congratulate all the winners and everyone else who has worked so hard over the past three months on this competition. Hope that even those of you who have been burned by the shakeup at the end can come away from this competition with a lot of new things that you have learned. We certainly are.\n\nWe had lots of fun teaming up with each other. Keeping each other abreast of all that we were trying to do, and encouraging each other, especially in the last phase of this competition, has been an incredible experience for us. \n\nWe would like to thank the organizers, Bengali.ai and Kaggle, for putting this competition together. We hope that the results of this competition can be used in a meaningful and productive way for the advancement of the handwritten Bengali language. \n\nChris and Bojan would like to thank Nvidia for its ongoing generous support in terms of time and resources. \n\n## Frameworks\n\nFor our final solution, we have used both PyTorch and TensorFlow. We have also worked with Rapids on an alternative solution, which showed lots of promise but was not possible to incorporate with our best solution in the allotted submission time limits. \n\n## Datasets\n\nMost of our models used just the competition dataset. One of our models (EfficientNet B7) trained partially with some of the external data. We had also tried using synthetic data, but that did not help our local models. It is possible that it did better on the private dataset, but we have not made any submissions with it yet.\n\n## Image Resolutions\n\nLike everyone else, throughout the competition, we had experimented with different resolutions and aspect ratios. Our final solution was a combination of the following different resolutions: 128x128, 224x224, 137x236, 128x256. Our single best model used the 256x256 resolution, but the inference with it was taking too long, so we had to make a difficult decision to leave it out of the final ensemble in favor of a more diverse blend. \n\n\n## Validation Scheme\n\nOur main validation scheme has been a 90/10 train/validation split. For most of the models, we have used the same 90/10 split, but in the latter stages of the competition, we had added a few with different splits for the sake of diversity.\n\n## Network Architectures\n\nWe have tried several different architectures and discovered that for the cutmix augmentation as we had implemented it, bigger networks were better, but they required a longer time to train and larger batch sizes. All but one network were trained with a single head. In hindsight, this might have seriously limited our post-processing as we had implemented it. Our final blend consisted of the following seven network architectures:\n\n\n- EfficientNet B4\n- EfficentNet B6\n- EfficientNet B7\n- DenseNet 201\n- SeNet 154\n- 2 SE-ResNext 101\n\n## Augmentations\n\nTwo main augmentations that we used were CutMix and CoarseDropout. We had also used a mild shift, scale, rotate augmentations. We also used CAM CutMix, Chris describes it [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021).\n\n## Training Regiment(s)\n\nOur main training consisted of one cycle with starting lr of 0.0001, and a decrease at the plateau. We would train between 100 and 200 epochs, with batch sizes between 150 and 300. After the one main cycle, we would retrain for another 70 epochs, which would usually bump our validation for another 0.001. Depending on the machine, network, and batch sizes, one epoch would take between 10 and 30 minutes to train. Our single best network achieved 0.9980 for local validation, and 0.9892 on public leaderboard. \n\n## Ensembling\n\nWe’ve tried to build some advanced second level models, but our best ensembling ended up being just a weighted average of our models. Since our models were trained to predict logits and not probabilities, finding a good blend was tricky. In the end, we managed to get a pretty decent set of weights, that was the best for both oof predictions as well as the public leaderboard. \n\n## Post Processing\n\nThis was one of the most important steps in our model, and most of the models that finished at the top in this competition. It prevented us from dropping like a stone after the final shakeup. In a nutshell, it turns out that the recall metric is very sensitive to the distribution of various classes. In order to exploit this insight, it was necessary to scale various predictions in such a way that the probabilities of less frequent items got proportionally higher weight. The most mathematically reasonable way of doing this is to multiply various probabilities by their inverse frequencies. In practice, we had to deviate slightly from the perfect inverse law. We had spent a lot of time in the last days trying to improve this post processing, but in the end, a very simple set of exponent worked best for our models. You can read more about this and other aspects of our solution in [Chris Deotte](https://www.kaggle.com/cdeotte)’s [wonderful writeup](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021). \n\n## Hardware\n\nThe main workhorses for our model building have been two Nvidia DGX Stations which have been generously allocated by Nvidia for the purposes of Kaggle competitions. We have also relied on our own private resources, including a dual RTX Titan workstation, and several 2080 Ti workstations.",
      "votes": null
    },
    {
      "id": "777707",
      "postDate": "03/17/2020 21:42:37",
      "content": "<p>Congratulations Bojan, Shai, Yasin, Jahmed ( <a href=\"/tunguz\">@tunguz</a> <a href=\"/sgalib\">@sgalib</a> <a href=\"/mykttu\">@mykttu</a> <a href=\"/jasemahmed\">@jasemahmed</a> ) on a job well done! It was fun working together and chatting over Slack. I learned a lot from your code, watching you all work, and our discussions.</p>\n\n<p>Before I teamed up my personal model only had CV 0.992 LB 0.983 (with PP LB 0.988) and I couldn't understand how teams were building CV 0.997 and scoring LB 0.99 without PP.</p>\n\n<p>After teaming up I discovered how. First CAM CutMix raised me 0.001. Then I added your coarse dropout to my augmentations of CAM CutMix, rotation, scale, and shift and that increased me by 0.001. Then I used your reduce on plateau LR schedule and that increased me by 0.001. Then I did your retrain a second cycle trick for another 0.001 gain. And then training on 90% fold instead of 80% fold increased another 0.001. I was happy to see my model reach CV 0.997 and LB 0.989 without PP and LB 0.917 with PP !! Thanks guys!</p>",
      "rawMarkdown": "Congratulations Bojan, Shai, Yasin, Jahmed ( @tunguz @sgalib @mykttu @jasemahmed ) on a job well done! It was fun working together and chatting over Slack. I learned a lot from your code, watching you all work, and our discussions.\n\nBefore I teamed up my personal model only had CV 0.992 LB 0.983 (with PP LB 0.988) and I couldn't understand how teams were building CV 0.997 and scoring LB 0.99 without PP.\n\nAfter teaming up I discovered how. First CAM CutMix raised me 0.001. Then I added your coarse dropout to my augmentations of CAM CutMix, rotation, scale, and shift and that increased me by 0.001. Then I used your reduce on plateau LR schedule and that increased me by 0.001. Then I did your retrain a second cycle trick for another 0.001 gain. And then training on 90% fold instead of 80% fold increased another 0.001. I was happy to see my model reach CV 0.997 and LB 0.989 without PP and LB 0.917 with PP !! Thanks guys!",
      "votes": null
    },
    {
      "id": "777719",
      "postDate": "03/17/2020 21:53:00",
      "content": "<p>Congratulations Yasin <a href=\"/mykttu\">@mykttu</a> becoming competition master! </p>",
      "rawMarkdown": "Congratulations Yasin @mykttu becoming competition master!",
      "votes": null
    },
    {
      "id": "777960",
      "postDate": "03/18/2020 03:33:05",
      "content": "<p>Congratulations and thanks for sharing! I was wondering what are your public LB and private LB scores for single model efficientnet. They are usually pretty low on public LB but seem to perform well on private LB. Thanks!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! I was wondering what are your public LB and private LB scores for single model efficientnet. They are usually pretty low on public LB but seem to perform well on private LB. Thanks!",
      "votes": null
    },
    {
      "id": "778239",
      "postDate": "03/18/2020 09:09:04",
      "content": "<p>great job!</p>",
      "rawMarkdown": "great job!",
      "votes": null
    },
    {
      "id": "778455",
      "postDate": "03/18/2020 13:10:30",
      "content": "<p>Congratulations and Thank you for sharing!</p>",
      "rawMarkdown": "Congratulations and Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 777707,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/17/2020 21:42:37",
      "content": "<p>Congratulations Bojan, Shai, Yasin, Jahmed ( <a href=\"/tunguz\">@tunguz</a> <a href=\"/sgalib\">@sgalib</a> <a href=\"/mykttu\">@mykttu</a> <a href=\"/jasemahmed\">@jasemahmed</a> ) on a job well done! It was fun working together and chatting over Slack. I learned a lot from your code, watching you all work, and our discussions.</p>\n\n<p>Before I teamed up my personal model only had CV 0.992 LB 0.983 (with PP LB 0.988) and I couldn't understand how teams were building CV 0.997 and scoring LB 0.99 without PP.</p>\n\n<p>After teaming up I discovered how. First CAM CutMix raised me 0.001. Then I added your coarse dropout to my augmentations of CAM CutMix, rotation, scale, and shift and that increased me by 0.001. Then I used your reduce on plateau LR schedule and that increased me by 0.001. Then I did your retrain a second cycle trick for another 0.001 gain. And then training on 90% fold instead of 80% fold increased another 0.001. I was happy to see my model reach CV 0.997 and LB 0.989 without PP and LB 0.917 with PP !! Thanks guys!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 777719,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/17/2020 21:53:00",
      "content": "<p>Congratulations Yasin <a href=\"/mykttu\">@mykttu</a> becoming competition master! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 777960,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "03/18/2020 03:33:05",
      "content": "<p>Congratulations and thanks for sharing! I was wondering what are your public LB and private LB scores for single model efficientnet. They are usually pretty low on public LB but seem to perform well on private LB. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 778239,
      "author_name": "davidesantangelo",
      "author_url": "",
      "post_date": "03/18/2020 09:09:04",
      "content": "<p>great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 778455,
      "author_name": "parmarsuraj99",
      "author_url": "",
      "post_date": "03/18/2020 13:10:30",
      "content": "<p>Congratulations and Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "777695": "## Congratulations and Thank You\n\nThis has been an interesting, challenging, and very educational experience for us. We would like to congratulate all the winners and everyone else who has worked so hard over the past three months on this competition. Hope that even those of you who have been burned by the shakeup at the end can come away from this competition with a lot of new things that you have learned. We certainly are.\n\nWe had lots of fun teaming up with each other. Keeping each other abreast of all that we were trying to do, and encouraging each other, especially in the last phase of this competition, has been an incredible experience for us. \n\nWe would like to thank the organizers, Bengali.ai and Kaggle, for putting this competition together. We hope that the results of this competition can be used in a meaningful and productive way for the advancement of the handwritten Bengali language. \n\nChris and Bojan would like to thank Nvidia for its ongoing generous support in terms of time and resources. \n\n## Frameworks\n\nFor our final solution, we have used both PyTorch and TensorFlow. We have also worked with Rapids on an alternative solution, which showed lots of promise but was not possible to incorporate with our best solution in the allotted submission time limits. \n\n## Datasets\n\nMost of our models used just the competition dataset. One of our models (EfficientNet B7) trained partially with some of the external data. We had also tried using synthetic data, but that did not help our local models. It is possible that it did better on the private dataset, but we have not made any submissions with it yet.\n\n## Image Resolutions\n\nLike everyone else, throughout the competition, we had experimented with different resolutions and aspect ratios. Our final solution was a combination of the following different resolutions: 128x128, 224x224, 137x236, 128x256. Our single best model used the 256x256 resolution, but the inference with it was taking too long, so we had to make a difficult decision to leave it out of the final ensemble in favor of a more diverse blend. \n\n\n## Validation Scheme\n\nOur main validation scheme has been a 90/10 train/validation split. For most of the models, we have used the same 90/10 split, but in the latter stages of the competition, we had added a few with different splits for the sake of diversity.\n\n## Network Architectures\n\nWe have tried several different architectures and discovered that for the cutmix augmentation as we had implemented it, bigger networks were better, but they required a longer time to train and larger batch sizes. All but one network were trained with a single head. In hindsight, this might have seriously limited our post-processing as we had implemented it. Our final blend consisted of the following seven network architectures:\n\n\n- EfficientNet B4\n- EfficentNet B6\n- EfficientNet B7\n- DenseNet 201\n- SeNet 154\n- 2 SE-ResNext 101\n\n## Augmentations\n\nTwo main augmentations that we used were CutMix and CoarseDropout. We had also used a mild shift, scale, rotate augmentations. We also used CAM CutMix, Chris describes it [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021).\n\n## Training Regiment(s)\n\nOur main training consisted of one cycle with starting lr of 0.0001, and a decrease at the plateau. We would train between 100 and 200 epochs, with batch sizes between 150 and 300. After the one main cycle, we would retrain for another 70 epochs, which would usually bump our validation for another 0.001. Depending on the machine, network, and batch sizes, one epoch would take between 10 and 30 minutes to train. Our single best network achieved 0.9980 for local validation, and 0.9892 on public leaderboard. \n\n## Ensembling\n\nWe’ve tried to build some advanced second level models, but our best ensembling ended up being just a weighted average of our models. Since our models were trained to predict logits and not probabilities, finding a good blend was tricky. In the end, we managed to get a pretty decent set of weights, that was the best for both oof predictions as well as the public leaderboard. \n\n## Post Processing\n\nThis was one of the most important steps in our model, and most of the models that finished at the top in this competition. It prevented us from dropping like a stone after the final shakeup. In a nutshell, it turns out that the recall metric is very sensitive to the distribution of various classes. In order to exploit this insight, it was necessary to scale various predictions in such a way that the probabilities of less frequent items got proportionally higher weight. The most mathematically reasonable way of doing this is to multiply various probabilities by their inverse frequencies. In practice, we had to deviate slightly from the perfect inverse law. We had spent a lot of time in the last days trying to improve this post processing, but in the end, a very simple set of exponent worked best for our models. You can read more about this and other aspects of our solution in [Chris Deotte](https://www.kaggle.com/cdeotte)’s [wonderful writeup](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021). \n\n## Hardware\n\nThe main workhorses for our model building have been two Nvidia DGX Stations which have been generously allocated by Nvidia for the purposes of Kaggle competitions. We have also relied on our own private resources, including a dual RTX Titan workstation, and several 2080 Ti workstations.",
    "777707": "Congratulations Bojan, Shai, Yasin, Jahmed ( @tunguz @sgalib @mykttu @jasemahmed ) on a job well done! It was fun working together and chatting over Slack. I learned a lot from your code, watching you all work, and our discussions.\n\nBefore I teamed up my personal model only had CV 0.992 LB 0.983 (with PP LB 0.988) and I couldn't understand how teams were building CV 0.997 and scoring LB 0.99 without PP.\n\nAfter teaming up I discovered how. First CAM CutMix raised me 0.001. Then I added your coarse dropout to my augmentations of CAM CutMix, rotation, scale, and shift and that increased me by 0.001. Then I used your reduce on plateau LR schedule and that increased me by 0.001. Then I did your retrain a second cycle trick for another 0.001 gain. And then training on 90% fold instead of 80% fold increased another 0.001. I was happy to see my model reach CV 0.997 and LB 0.989 without PP and LB 0.917 with PP !! Thanks guys!",
    "777719": "Congratulations Yasin @mykttu becoming competition master!",
    "777960": "Congratulations and thanks for sharing! I was wondering what are your public LB and private LB scores for single model efficientnet. They are usually pretty low on public LB but seem to perform well on private LB. Thanks!",
    "778239": "great job!",
    "778455": "Congratulations and Thank you for sharing!"
  },
  "source": "meta"
}