{
  "id": 159080,
  "title": "1st Place - TPU Leaderboard Prizes",
  "url": "/competitions/flower-classification-with-tpus/writeups/manuel-campos-1st-place-tpu-leaderboard-prizes",
  "author_name": "",
  "post_date": "2020-06-16T19:28:58.717Z",
  "votes": 23,
  "comment_count": 9,
  "views": 0,
  "content": "<h3>Submission Notebooks</h3>\n\n<p>I am delighted to share with you the winning submission in these notebooks,</p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v2\">Training-Densenet201-Bagging1</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v3\">Training-Densenet201-Bagging2</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v2\">Training-EfficientnetB7-Bagging1</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v3\">Training-EfficientnetB7-Bagging2</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-classification-tpu-inference-v3\">Inference-Flower-Classification-TPU</a></li>\n</ol>\n\n<p>Your links are pinned to the key versions. Although you will probably find some parts of the code redundant or not quite efficient, I didn't want to touch a comma to show the composition of the winning solution as it is here. This way, you can check both the clean inputs included in the inference notebooks and the automatic generation of the final solution. </p>\n\n<p>I just want to add 'a little' to my summary of the competition below and highlight the small details that perhaps made my score stronger.</p>\n\n<h3>Acknowledgements</h3>\n\n<p>Finally the TPU Leaderboard Prizes were settled and as you can imagine I am very happy with the results. Thank you very much to everyone. Thanks to all the Kaggle team and Google-Cloud-TPU for making this competition possible and thanks to the 848 teams that participated in this challenge.</p>\n\n<p>Whether it is explicitly (sharing notebooks and posting comments and discussions) or implicitly (pushing the leaderboard up with high scores) I think Kagglers are the ones that give a unique value to any challenge.</p>\n\n<h3>What you can't miss</h3>\n\n<p>I believe that more than a few Data Scientists will come to review the material in 'Flower Classification with TPUs' to get initiated or indeed improve the use of this powerful technology in their projects. If you are one of them here I detail, without order of preference or priority, a series of authors (sorry because I surely left some behind) of kernels whose work in this competition you should review (all Tensorflow).</p>\n\n<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> (Martin Görner), <a href=\"/xiejialun\">@xiejialun</a> (Xie29), <a href=\"/yihdarshieh\">@yihdarshieh</a> (Yih-Dar SHIEH), <a href=\"/cdeotte\">@cdeotte</a> (Chris Deotte), <a href=\"/calebeverett\">@calebeverett</a> (Caleb), <a href=\"/romanweilguny\">@romanweilguny</a> (Roman Weilguny) but also don't forget to look at everything else if needed. </p>\n\n<p>I would state that the playground objective, which was aimed at facilitating the use of TPU technology, has been more than achieved.</p>\n\n<h3>The Funny Part</h3>\n\n<p>There are three key kernels at the beinning of the competition,\n- <a href=\"https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus\">ratan123/densenet201-flower-classification-with-tpus</a>\n- <a href=\"https://www.kaggle.com/xhlulu/flowers-tpu-concise-efficientnet-b7\">xhlulu/flowers-tpu-concise-efficientnet-b7</a>\n- <a href=\"https://www.kaggle.com/wrrosa/tpu-enet-b7-densenet\">wrrosa/tpu-enet-b7-densenet</a></p>\n\n<p>The baseline of these notebooks is enough, for with the allowed competition data, to obtain the limit in the range of 0.9700 score. I am referring to the limit because of what we have seen in practice later on.</p>\n\n<p>On how to beat the limit of that score: the use of external data. <a href=\"/cdeotte\">@cdeotte</a> (Chris Deotte) in this discussions (1) <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140964\">Oxford Flowers - Kaggle Dataset for Everyone</a>. (2) <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329\">How to Score Over LB 0.970+</a> and <a href=\"/hengck23\">@hengck23</a> (Heng CherKeng) in <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140866\">[placeholder] external data and how to use them</a> made public the manner.    </p>\n\n<p>In my opinion this was the exciting part, I mean the rush to include these new datasets in training and see how the performance of the models improved very fast.</p>\n\n<p>I have to make a super special mention here for <a href=\"/kirillblinov\">@kirillblinov</a> (Kirill Blinov) and its shared external dataset <a href=\"https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec\">tf_flower_photo_tfrec</a>, great work. In this way, <a href=\"/szacho\">@szacho</a> (Michal Szachniewicz) y <a href=\"/calebeverett\">@calebeverett</a> (Caleb) made an interesting contribution.</p>\n\n<p>By the way, one of the external dataset did not work for me, <code>openimage</code>.</p>\n\n<h3>The Controversial Part</h3>\n\n<p>I am referring here to all the final polemics about whether or not to use external datasets. The host never tried to hide the composition of the images in the test set and their origin. In my particular opinion I would have let the competition run until the end by allowing any source of data. </p>\n\n<p>It was also a very hard challenge to work on these new data and in any case the final obstacle is not so much the difficulty of the objective itself but the competition between the participants to overcome a score. The later, however, as long as everyone can access any advantage as happened here.</p>\n\n<p>You can be more or less in line with the direction that was finally decided, but you don't have to lose your commitment to respect and follow the rules of the competition for this reason. </p>\n\n<p>I expressed my disagreement with the final treatment of the Private Leaderboard and explained it <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494#84954\">here</a>.</p>\n\n<h3>Training Topics</h3>\n\n<p>If you have enjoyed the competition from the beginning, participating in the published notebooks and discussions I think you won't find anything new in the kernels I have shared but you will probably be interested in knowing what could be missing or left over from your solution to achieve a higher score in the Public or Private Leaderboard.</p>\n\n<ul>\n<li><p><strong>Augmentation</strong> \nI just used the files for an image size of  <code>512x512</code>, I did not preprocess original images. On them, <code>Random Blockout</code> y <code>Horizontal Random Flip</code> were the best choices. </p></li>\n<li><p><strong>Validation</strong> \nSince the initial validation dataset was very extensive and its information very useful in terms of performance, I divided the validation files into two subsets generating an additional Fold. Something like that looking for cross validation,</p></li>\n</ul>\n\n<pre><code>Fold1&gt;&gt;&gt;\nVALIDATIONFILENAMES1=VALIDATIONFILENAMES[:int(len(VALIDATIONFILENAMES)/2)] \n</code></pre>\n\n<pre><code>Fold2&gt;&gt;&gt; VALIDATIONFILENAMES2=VALIDATIONFILENAMES[int(len(VALIDATIONFILENAMES)/2):]\n</code></pre>\n\n<pre><code>TRAININGFILENAMES1=TRAININGFILENAMES+VALIDATIONFILENAMES1  TRAININGFILENAMES2=TRAININGFILENAMES+VALIDATIONFILENAMES2\n</code></pre>\n\n<ul>\n<li><p><strong>DenseNet-201</strong>\n<code>Imagenet Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. Little or no adjustment here, I kept BATCH_SIZE=128 and I brought to 4 the LR-RAMPUP-EPOCHS and LR-SUSTAIN-EPOCHS parameters of the well-tuned 'CustomLRSchedule' from <a href=\"/mgornergoogle\">@mgornergoogle</a> (Martin Görner) in his ' Five 'Flowers' Notebook.</p></li>\n<li><p><strong>EfficientNet-B7</strong>\n<code>Noisy-Student Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. The rest is the same as Densenet-201 except for 5 the epochs number ramp-up and sustain. </p></li>\n<li><p><strong>Loss y Optimizer</strong> <br>\nShared by both models, <code>Loss</code> with <code>'sparse_categorical_crossentropy'</code> and <code>Optimizer</code> with <code>Adam(lr='CustomLRSchedule')</code>.  </p></li>\n<li><p><strong>Model Checkpoints</strong>\nThe following callback was used to save the best weights according to the loss in validation,   </p></li>\n</ul>\n\n<pre><code>mc=tf.keras.callbacks.ModelCheckpoint('name.h5',verbose=1,monitor='val_loss',               \n    mode='min',save_best_only=True,save_weights_only=True)\n</code></pre>\n\n<ul>\n<li><strong>2-Baggings</strong>\nWhile the use of the Blockout made the training results more stable, in general the performance of the models varied significantly (+/- 0.001) from run to run. Considering this problem of randomness I trained each model twice with the same basic setting. You can presume that the random augmentations are relevant here but even without them the variability of the TPU workouts is something important to consider. <code>NUM_EPOCHS</code> number <code>35, 50</code> to DenseNet and <code>35, 40</code> to EfficientNet.</li>\n</ul>\n\n<h3>Inference Topics</h3>\n\n<p>Therefore with the details above I actually had 8 trained models available with which to perform the classification inference.\n1. DenseNet (4), 2xFold with <code>best-checkpoints</code> epochs <code>24-18, Fold1</code> and <code>25-24, Fold2</code>.\n2. EfficientNet (4),  2xFold with <code>best-checkpoints</code> on epochs <code>17-16, Fold1</code> and <code>14-13, Fold2</code>.</p>\n\n<ul>\n<li><p><strong>Test Time Augmentation</strong>\nThe strategy here was to explode the training augmentated with 2 types of prediction. One of them (1) with the original 512x512 images without any transformation and the other (2) with 100% of the images flipped using the function <code>tf.image.flip_left_right(image)</code> as the only augment. </p></li>\n<li><p><strong>Averaging Out of Folds</strong>\nI join the oofs of each bagging and model with a <code>simple average</code> of the <code>logits</code>, not probabilities. Following TTA strategy I perform the same operation with the full-flip images and let's say I get the following sets,</p></li>\n</ul>\n\n<pre><code>1 oof-non-flip-densenet\n1 oof-non-flip-efficientnet\n1 oof-full-flip-densenet\n1 oof-full-flip-efficientnet\n</code></pre>\n\n<ul>\n<li><strong>Weighting Models OOFs Non-Flip</strong>\nWith the logit transformation still applied, I use a simple optimizer to linearly maximize the weight of each model in the non-flip context.</li>\n</ul>\n\n<pre><code>Best weighting to EfficientNet-B7: 0.71\nValidation F1-Score: 0.969903\n</code></pre>\n\n<ul>\n<li><strong>Weighting Models OOFs Full-Flip</strong>\nFollowing the same procedure I get now,</li>\n</ul>\n\n<pre><code>Best weighting for EfficientNet-B7:  0.65\nValidation F1-Score: 0.970014\n</code></pre>\n\n<ul>\n<li><strong>One Last Optimization</strong>\nA final step was to obtain a new weighting but now to calibrate the power of the TTA strategy between full-flip and non-flip derived classifications.</li>\n</ul>\n\n<pre><code>Best weighting for validation Non-Flip: 0.05 (any value 0.05 to 0.21)\nValidation F1-Score: 0.970367\n</code></pre>\n\n<h3>That's all folks !!!</h3>",
  "messages": [
    {
      "id": "888451",
      "postDate": "06/16/2020 11:15:37",
      "content": "<h3>Submission Notebooks</h3>\n\n<p>I am delighted to share with you the winning submission in these notebooks,</p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v2\">Training-Densenet201-Bagging1</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v3\">Training-Densenet201-Bagging2</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v2\">Training-EfficientnetB7-Bagging1</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v3\">Training-EfficientnetB7-Bagging2</a></li>\n<li><a href=\"https://www.kaggle.com/coreacasa/flower-classification-tpu-inference-v3\">Inference-Flower-Classification-TPU</a></li>\n</ol>\n\n<p>Your links are pinned to the key versions. Although you will probably find some parts of the code redundant or not quite efficient, I didn't want to touch a comma to show the composition of the winning solution as it is here. This way, you can check both the clean inputs included in the inference notebooks and the automatic generation of the final solution. </p>\n\n<p>I just want to add 'a little' to my summary of the competition below and highlight the small details that perhaps made my score stronger.</p>\n\n<h3>Acknowledgements</h3>\n\n<p>Finally the TPU Leaderboard Prizes were settled and as you can imagine I am very happy with the results. Thank you very much to everyone. Thanks to all the Kaggle team and Google-Cloud-TPU for making this competition possible and thanks to the 848 teams that participated in this challenge.</p>\n\n<p>Whether it is explicitly (sharing notebooks and posting comments and discussions) or implicitly (pushing the leaderboard up with high scores) I think Kagglers are the ones that give a unique value to any challenge.</p>\n\n<h3>What you can't miss</h3>\n\n<p>I believe that more than a few Data Scientists will come to review the material in 'Flower Classification with TPUs' to get initiated or indeed improve the use of this powerful technology in their projects. If you are one of them here I detail, without order of preference or priority, a series of authors (sorry because I surely left some behind) of kernels whose work in this competition you should review (all Tensorflow).</p>\n\n<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> (Martin Görner), <a href=\"/xiejialun\">@xiejialun</a> (Xie29), <a href=\"/yihdarshieh\">@yihdarshieh</a> (Yih-Dar SHIEH), <a href=\"/cdeotte\">@cdeotte</a> (Chris Deotte), <a href=\"/calebeverett\">@calebeverett</a> (Caleb), <a href=\"/romanweilguny\">@romanweilguny</a> (Roman Weilguny) but also don't forget to look at everything else if needed. </p>\n\n<p>I would state that the playground objective, which was aimed at facilitating the use of TPU technology, has been more than achieved.</p>\n\n<h3>The Funny Part</h3>\n\n<p>There are three key kernels at the beinning of the competition,\n- <a href=\"https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus\">ratan123/densenet201-flower-classification-with-tpus</a>\n- <a href=\"https://www.kaggle.com/xhlulu/flowers-tpu-concise-efficientnet-b7\">xhlulu/flowers-tpu-concise-efficientnet-b7</a>\n- <a href=\"https://www.kaggle.com/wrrosa/tpu-enet-b7-densenet\">wrrosa/tpu-enet-b7-densenet</a></p>\n\n<p>The baseline of these notebooks is enough, for with the allowed competition data, to obtain the limit in the range of 0.9700 score. I am referring to the limit because of what we have seen in practice later on.</p>\n\n<p>On how to beat the limit of that score: the use of external data. <a href=\"/cdeotte\">@cdeotte</a> (Chris Deotte) in this discussions (1) <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140964\">Oxford Flowers - Kaggle Dataset for Everyone</a>. (2) <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329\">How to Score Over LB 0.970+</a> and <a href=\"/hengck23\">@hengck23</a> (Heng CherKeng) in <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140866\">[placeholder] external data and how to use them</a> made public the manner.    </p>\n\n<p>In my opinion this was the exciting part, I mean the rush to include these new datasets in training and see how the performance of the models improved very fast.</p>\n\n<p>I have to make a super special mention here for <a href=\"/kirillblinov\">@kirillblinov</a> (Kirill Blinov) and its shared external dataset <a href=\"https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec\">tf_flower_photo_tfrec</a>, great work. In this way, <a href=\"/szacho\">@szacho</a> (Michal Szachniewicz) y <a href=\"/calebeverett\">@calebeverett</a> (Caleb) made an interesting contribution.</p>\n\n<p>By the way, one of the external dataset did not work for me, <code>openimage</code>.</p>\n\n<h3>The Controversial Part</h3>\n\n<p>I am referring here to all the final polemics about whether or not to use external datasets. The host never tried to hide the composition of the images in the test set and their origin. In my particular opinion I would have let the competition run until the end by allowing any source of data. </p>\n\n<p>It was also a very hard challenge to work on these new data and in any case the final obstacle is not so much the difficulty of the objective itself but the competition between the participants to overcome a score. The later, however, as long as everyone can access any advantage as happened here.</p>\n\n<p>You can be more or less in line with the direction that was finally decided, but you don't have to lose your commitment to respect and follow the rules of the competition for this reason. </p>\n\n<p>I expressed my disagreement with the final treatment of the Private Leaderboard and explained it <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494#84954\">here</a>.</p>\n\n<h3>Training Topics</h3>\n\n<p>If you have enjoyed the competition from the beginning, participating in the published notebooks and discussions I think you won't find anything new in the kernels I have shared but you will probably be interested in knowing what could be missing or left over from your solution to achieve a higher score in the Public or Private Leaderboard.</p>\n\n<ul>\n<li><p><strong>Augmentation</strong> \nI just used the files for an image size of  <code>512x512</code>, I did not preprocess original images. On them, <code>Random Blockout</code> y <code>Horizontal Random Flip</code> were the best choices. </p></li>\n<li><p><strong>Validation</strong> \nSince the initial validation dataset was very extensive and its information very useful in terms of performance, I divided the validation files into two subsets generating an additional Fold. Something like that looking for cross validation,</p></li>\n</ul>\n\n<pre><code>Fold1&gt;&gt;&gt;\nVALIDATIONFILENAMES1=VALIDATIONFILENAMES[:int(len(VALIDATIONFILENAMES)/2)] \n</code></pre>\n\n<pre><code>Fold2&gt;&gt;&gt; VALIDATIONFILENAMES2=VALIDATIONFILENAMES[int(len(VALIDATIONFILENAMES)/2):]\n</code></pre>\n\n<pre><code>TRAININGFILENAMES1=TRAININGFILENAMES+VALIDATIONFILENAMES1  TRAININGFILENAMES2=TRAININGFILENAMES+VALIDATIONFILENAMES2\n</code></pre>\n\n<ul>\n<li><p><strong>DenseNet-201</strong>\n<code>Imagenet Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. Little or no adjustment here, I kept BATCH_SIZE=128 and I brought to 4 the LR-RAMPUP-EPOCHS and LR-SUSTAIN-EPOCHS parameters of the well-tuned 'CustomLRSchedule' from <a href=\"/mgornergoogle\">@mgornergoogle</a> (Martin Görner) in his ' Five 'Flowers' Notebook.</p></li>\n<li><p><strong>EfficientNet-B7</strong>\n<code>Noisy-Student Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. The rest is the same as Densenet-201 except for 5 the epochs number ramp-up and sustain. </p></li>\n<li><p><strong>Loss y Optimizer</strong> <br>\nShared by both models, <code>Loss</code> with <code>'sparse_categorical_crossentropy'</code> and <code>Optimizer</code> with <code>Adam(lr='CustomLRSchedule')</code>.  </p></li>\n<li><p><strong>Model Checkpoints</strong>\nThe following callback was used to save the best weights according to the loss in validation,   </p></li>\n</ul>\n\n<pre><code>mc=tf.keras.callbacks.ModelCheckpoint('name.h5',verbose=1,monitor='val_loss',               \n    mode='min',save_best_only=True,save_weights_only=True)\n</code></pre>\n\n<ul>\n<li><strong>2-Baggings</strong>\nWhile the use of the Blockout made the training results more stable, in general the performance of the models varied significantly (+/- 0.001) from run to run. Considering this problem of randomness I trained each model twice with the same basic setting. You can presume that the random augmentations are relevant here but even without them the variability of the TPU workouts is something important to consider. <code>NUM_EPOCHS</code> number <code>35, 50</code> to DenseNet and <code>35, 40</code> to EfficientNet.</li>\n</ul>\n\n<h3>Inference Topics</h3>\n\n<p>Therefore with the details above I actually had 8 trained models available with which to perform the classification inference.\n1. DenseNet (4), 2xFold with <code>best-checkpoints</code> epochs <code>24-18, Fold1</code> and <code>25-24, Fold2</code>.\n2. EfficientNet (4),  2xFold with <code>best-checkpoints</code> on epochs <code>17-16, Fold1</code> and <code>14-13, Fold2</code>.</p>\n\n<ul>\n<li><p><strong>Test Time Augmentation</strong>\nThe strategy here was to explode the training augmentated with 2 types of prediction. One of them (1) with the original 512x512 images without any transformation and the other (2) with 100% of the images flipped using the function <code>tf.image.flip_left_right(image)</code> as the only augment. </p></li>\n<li><p><strong>Averaging Out of Folds</strong>\nI join the oofs of each bagging and model with a <code>simple average</code> of the <code>logits</code>, not probabilities. Following TTA strategy I perform the same operation with the full-flip images and let's say I get the following sets,</p></li>\n</ul>\n\n<pre><code>1 oof-non-flip-densenet\n1 oof-non-flip-efficientnet\n1 oof-full-flip-densenet\n1 oof-full-flip-efficientnet\n</code></pre>\n\n<ul>\n<li><strong>Weighting Models OOFs Non-Flip</strong>\nWith the logit transformation still applied, I use a simple optimizer to linearly maximize the weight of each model in the non-flip context.</li>\n</ul>\n\n<pre><code>Best weighting to EfficientNet-B7: 0.71\nValidation F1-Score: 0.969903\n</code></pre>\n\n<ul>\n<li><strong>Weighting Models OOFs Full-Flip</strong>\nFollowing the same procedure I get now,</li>\n</ul>\n\n<pre><code>Best weighting for EfficientNet-B7:  0.65\nValidation F1-Score: 0.970014\n</code></pre>\n\n<ul>\n<li><strong>One Last Optimization</strong>\nA final step was to obtain a new weighting but now to calibrate the power of the TTA strategy between full-flip and non-flip derived classifications.</li>\n</ul>\n\n<pre><code>Best weighting for validation Non-Flip: 0.05 (any value 0.05 to 0.21)\nValidation F1-Score: 0.970367\n</code></pre>\n\n<h3>That's all folks !!!</h3>",
      "rawMarkdown": "### Submission Notebooks    \n   \nI am delighted to share with you the winning submission in these notebooks,\n    \n1. [Training-Densenet201-Bagging1](https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v2)\n2. [Training-Densenet201-Bagging2](https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v3)\n3. [Training-EfficientnetB7-Bagging1](https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v2)\n4. [Training-EfficientnetB7-Bagging2](https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v3)\n5. [Inference-Flower-Classification-TPU](https://www.kaggle.com/coreacasa/flower-classification-tpu-inference-v3)\n        \nYour links are pinned to the key versions. Although you will probably find some parts of the code redundant or not quite efficient, I didn't want to touch a comma to show the composition of the winning solution as it is here. This way, you can check both the clean inputs included in the inference notebooks and the automatic generation of the final solution. \n    \nI just want to add 'a little' to my summary of the competition below and highlight the small details that perhaps made my score stronger.\n\n### Acknowledgements\nFinally the TPU Leaderboard Prizes were settled and as you can imagine I am very happy with the results. Thank you very much to everyone. Thanks to all the Kaggle team and Google-Cloud-TPU for making this competition possible and thanks to the 848 teams that participated in this challenge.\n    \nWhether it is explicitly (sharing notebooks and posting comments and discussions) or implicitly (pushing the leaderboard up with high scores) I think Kagglers are the ones that give a unique value to any challenge.\n    \n### What you can't miss \nI believe that more than a few Data Scientists will come to review the material in 'Flower Classification with TPUs' to get initiated or indeed improve the use of this powerful technology in their projects. If you are one of them here I detail, without order of preference or priority, a series of authors (sorry because I surely left some behind) of kernels whose work in this competition you should review (all Tensorflow).\n    \nThanks @mgornergoogle (Martin Görner), @xiejialun (Xie29), @yihdarshieh (Yih-Dar SHIEH), @cdeotte (Chris Deotte), @calebeverett (Caleb), @romanweilguny (Roman Weilguny) but also don't forget to look at everything else if needed. \n\nI would state that the playground objective, which was aimed at facilitating the use of TPU technology, has been more than achieved.\n\n### The Funny Part \nThere are three key kernels at the beinning of the competition,\n- [ratan123/densenet201-flower-classification-with-tpus](https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus)\n- [xhlulu/flowers-tpu-concise-efficientnet-b7](https://www.kaggle.com/xhlulu/flowers-tpu-concise-efficientnet-b7)\n- [wrrosa/tpu-enet-b7-densenet](https://www.kaggle.com/wrrosa/tpu-enet-b7-densenet)\n        \nThe baseline of these notebooks is enough, for with the allowed competition data, to obtain the limit in the range of 0.9700 score. I am referring to the limit because of what we have seen in practice later on.\n        \nOn how to beat the limit of that score: the use of external data. @cdeotte (Chris Deotte) in this discussions (1) [Oxford Flowers - Kaggle Dataset for Everyone](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140964). (2) [How to Score Over LB 0.970+](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329) and @hengck23 (Heng CherKeng) in [[placeholder] external data and how to use them](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140866) made public the manner.    \n           \nIn my opinion this was the exciting part, I mean the rush to include these new datasets in training and see how the performance of the models improved very fast.\n        \nI have to make a super special mention here for @kirillblinov (Kirill Blinov) and its shared external dataset [tf\\_flower\\_photo\\_tfrec](https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec), great work. In this way, @szacho (Michal Szachniewicz) y @calebeverett (Caleb) made an interesting contribution.\n        \nBy the way, one of the external dataset did not work for me, <code>openimage</code>.\n\n### The Controversial Part\nI am referring here to all the final polemics about whether or not to use external datasets. The host never tried to hide the composition of the images in the test set and their origin. In my particular opinion I would have let the competition run until the end by allowing any source of data. \n        \nIt was also a very hard challenge to work on these new data and in any case the final obstacle is not so much the difficulty of the objective itself but the competition between the participants to overcome a score. The later, however, as long as everyone can access any advantage as happened here.\n        \nYou can be more or less in line with the direction that was finally decided, but you don't have to lose your commitment to respect and follow the rules of the competition for this reason. \n    \nI expressed my disagreement with the final treatment of the Private Leaderboard and explained it [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494#84954).\n\n### Training Topics\nIf you have enjoyed the competition from the beginning, participating in the published notebooks and discussions I think you won't find anything new in the kernels I have shared but you will probably be interested in knowing what could be missing or left over from your solution to achieve a higher score in the Public or Private Leaderboard.\n     \n- **Augmentation** \nI just used the files for an image size of  <code>512x512</code>, I did not preprocess original images. On them, <code>Random Blockout</code> y <code>Horizontal Random Flip</code> were the best choices. \n    \n- **Validation** \nSince the initial validation dataset was very extensive and its information very useful in terms of performance, I divided the validation files into two subsets generating an additional Fold. Something like that looking for cross validation,\n\n\n<pre><code>Fold1&gt;&gt;&gt;\nVALIDATIONFILENAMES1=VALIDATIONFILENAMES[:int(len(VALIDATIONFILENAMES)/2)] \n</code></pre>\n<pre><code>Fold2&gt;&gt;&gt; VALIDATIONFILENAMES2=VALIDATIONFILENAMES[int(len(VALIDATIONFILENAMES)/2):]\n</code></pre>\n<pre><code>TRAININGFILENAMES1=TRAININGFILENAMES+VALIDATIONFILENAMES1  TRAININGFILENAMES2=TRAININGFILENAMES+VALIDATIONFILENAMES2\n</code></pre>\n        \n- **DenseNet-201**\n<code>Imagenet Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. Little or no adjustment here, I kept BATCH_SIZE=128 and I brought to 4 the LR-RAMPUP-EPOCHS and LR-SUSTAIN-EPOCHS parameters of the well-tuned 'CustomLRSchedule' from @mgornergoogle (Martin Görner) in his ' Five 'Flowers' Notebook.\n        \n- **EfficientNet-B7**\n<code>Noisy-Student Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. The rest is the same as Densenet-201 except for 5 the epochs number ramp-up and sustain. \n     \n- **Loss y Optimizer**    \nShared by both models, <code>Loss</code> with <code>'sparse_categorical_crossentropy'</code> and <code>Optimizer</code> with <code>Adam(lr='CustomLRSchedule')</code>.  \n    \n- **Model Checkpoints**\nThe following callback was used to save the best weights according to the loss in validation,   \n<pre><code>mc=tf.keras.callbacks.ModelCheckpoint('name.h5',verbose=1,monitor='val_loss',               \n    mode='min',save_best_only=True,save_weights_only=True)\n</code></pre>\n                    \n- **2-Baggings**\nWhile the use of the Blockout made the training results more stable, in general the performance of the models varied significantly (+/- 0.001) from run to run. Considering this problem of randomness I trained each model twice with the same basic setting. You can presume that the random augmentations are relevant here but even without them the variability of the TPU workouts is something important to consider. <code>NUM_EPOCHS</code> number <code>35, 50</code> to DenseNet and <code>35, 40</code> to EfficientNet.\n    \n### Inference Topics\nTherefore with the details above I actually had 8 trained models available with which to perform the classification inference.\n1. DenseNet (4), 2xFold with <code>best-checkpoints</code> epochs <code>24-18, Fold1</code> and <code>25-24, Fold2</code>.\n2. EfficientNet (4),  2xFold with <code>best-checkpoints</code> on epochs <code>17-16, Fold1</code> and <code>14-13, Fold2</code>.\n    \n- **Test Time Augmentation**\nThe strategy here was to explode the training augmentated with 2 types of prediction. One of them (1) with the original 512x512 images without any transformation and the other (2) with 100% of the images flipped using the function <code>tf.image.flip_left_right(image)</code> as the only augment. \n    \n- **Averaging Out of Folds**\nI join the oofs of each bagging and model with a <code>simple average</code> of the <code>logits</code>, not probabilities. Following TTA strategy I perform the same operation with the full-flip images and let's say I get the following sets,\n    \n<pre><code>1 oof-non-flip-densenet\n1 oof-non-flip-efficientnet\n1 oof-full-flip-densenet\n1 oof-full-flip-efficientnet\n</code></pre>\n        \n- **Weighting Models OOFs Non-Flip**\nWith the logit transformation still applied, I use a simple optimizer to linearly maximize the weight of each model in the non-flip context.\n    \n<pre><code>Best weighting to EfficientNet-B7: 0.71\nValidation F1-Score: 0.969903\n</code></pre>\n    \n- **Weighting Models OOFs Full-Flip**\nFollowing the same procedure I get now,\n    \n<pre><code>Best weighting for EfficientNet-B7:  0.65\nValidation F1-Score: 0.970014\n</code></pre>\n        \n- **One Last Optimization**\nA final step was to obtain a new weighting but now to calibrate the power of the TTA strategy between full-flip and non-flip derived classifications.\n    \n<pre><code>Best weighting for validation Non-Flip: 0.05 (any value 0.05 to 0.21)\nValidation F1-Score: 0.970367\n</code></pre>\n    \n### That's all folks !!!",
      "votes": null
    },
    {
      "id": "888919",
      "postDate": "06/16/2020 16:41:39",
      "content": "<p>Congratulations! and thanks for the mention.</p>",
      "rawMarkdown": "Congratulations! and thanks for the mention.",
      "votes": null
    },
    {
      "id": "888931",
      "postDate": "06/16/2020 16:51:18",
      "content": "<p>Thank you for sharing a great job!</p>",
      "rawMarkdown": "Thank you for sharing a great job!",
      "votes": null
    },
    {
      "id": "888993",
      "postDate": "06/16/2020 17:47:31",
      "content": "<p>Thank you for sharing the solution. I like the fact that you ensemble at the logit level instead of ensembling the probabilities. (Not quite sure how you get the logits from the probabilities though.)</p>",
      "rawMarkdown": "Thank you for sharing the solution. I like the fact that you ensemble at the logit level instead of ensembling the probabilities. (Not quite sure how you get the logits from the probabilities though.)",
      "votes": null
    },
    {
      "id": "889142",
      "postDate": "06/16/2020 19:47:35",
      "content": "<p>Hi <a href=\"/mgornergoogle\">@mgornergoogle</a> , thanks to you too. I used the <code>logit</code> and <code>expit</code> (inverse logit) functions of <code>scipy.special</code> to get these transformations. To avoid -inf and +inf values in the logit output (for subsequent manipulations) I use a simple <code>np.nan_to_num</code>. </p>\n\n<p>About the performance of these transformations, but with a focus on TTA, <a href=\"/calebeverett\">@calebeverett</a> shared a huge notebook  <a href=\"https://www.kaggle.com/calebeverett/comparison-of-tta-prediction-procedures\">here</a> as a benchmark.</p>",
      "rawMarkdown": "Hi @mgornergoogle , thanks to you too. I used the <code>logit</code> and <code>expit</code> (inverse logit) functions of <code>scipy.special</code> to get these transformations. To avoid -inf and +inf values in the logit output (for subsequent manipulations) I use a simple <code>np.nan_to_num</code>. \n    \nAbout the performance of these transformations, but with a focus on TTA, @calebeverett shared a huge notebook  [here](https://www.kaggle.com/calebeverett/comparison-of-tta-prediction-procedures) as a benchmark.",
      "votes": null
    },
    {
      "id": "969843",
      "postDate": "08/14/2020 01:54:05",
      "content": "<p>Congratulations Manuel. You achieved a fantastic score using only competition train data. Well done!</p>",
      "rawMarkdown": "Congratulations Manuel. You achieved a fantastic score using only competition train data. Well done!",
      "votes": null
    },
    {
      "id": "970319",
      "postDate": "08/14/2020 11:04:10",
      "content": "<p>This solution will be better with your post here… Thanks a lot!</p>",
      "rawMarkdown": "This solution will be better with your post here... Thanks a lot!",
      "votes": null
    },
    {
      "id": "987849",
      "postDate": "08/27/2020 15:13:38",
      "content": "<p>Thanks for sharing, I learned a lot from looking through your notebooks!</p>",
      "rawMarkdown": "Thanks for sharing, I learned a lot from looking through your notebooks!",
      "votes": null
    },
    {
      "id": "987853",
      "postDate": "08/27/2020 15:16:55",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/benghertner\" target=\"_blank\">@benghertner</a>, that's great for me!</p>",
      "rawMarkdown": "Thanks @benghertner, that's great for me!",
      "votes": null
    },
    {
      "id": "1050391",
      "postDate": "10/15/2020 11:22:08",
      "content": "<p>thank you very much very insightful even for newbie (me)</p>",
      "rawMarkdown": "thank you very much very insightful even for newbie (me)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969843,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/14/2020 01:54:05",
      "content": "<p>Congratulations Manuel. You achieved a fantastic score using only competition train data. Well done!</p>",
      "votes": null,
      "replies": [
        {
          "id": 970319,
          "author_name": "coreacasa",
          "author_url": "",
          "post_date": "08/14/2020 11:04:10",
          "content": "<p>This solution will be better with your post here… Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 987849,
      "author_name": "benghertner",
      "author_url": "",
      "post_date": "08/27/2020 15:13:38",
      "content": "<p>Thanks for sharing, I learned a lot from looking through your notebooks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 987853,
          "author_name": "coreacasa",
          "author_url": "",
          "post_date": "08/27/2020 15:16:55",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/benghertner\" target=\"_blank\">@benghertner</a>, that's great for me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1050391,
      "author_name": "rizkioktafianto",
      "author_url": "",
      "post_date": "10/15/2020 11:22:08",
      "content": "<p>thank you very much very insightful even for newbie (me)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888919,
      "author_name": "calebeverett",
      "author_url": "",
      "post_date": "06/16/2020 16:41:39",
      "content": "<p>Congratulations! and thanks for the mention.</p>",
      "votes": null,
      "replies": [
        {
          "id": 888931,
          "author_name": "coreacasa",
          "author_url": "",
          "post_date": "06/16/2020 16:51:18",
          "content": "<p>Thank you for sharing a great job!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 888993,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "06/16/2020 17:47:31",
      "content": "<p>Thank you for sharing the solution. I like the fact that you ensemble at the logit level instead of ensembling the probabilities. (Not quite sure how you get the logits from the probabilities though.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 889142,
          "author_name": "coreacasa",
          "author_url": "",
          "post_date": "06/16/2020 19:47:35",
          "content": "<p>Hi <a href=\"/mgornergoogle\">@mgornergoogle</a> , thanks to you too. I used the <code>logit</code> and <code>expit</code> (inverse logit) functions of <code>scipy.special</code> to get these transformations. To avoid -inf and +inf values in the logit output (for subsequent manipulations) I use a simple <code>np.nan_to_num</code>. </p>\n\n<p>About the performance of these transformations, but with a focus on TTA, <a href=\"/calebeverett\">@calebeverett</a> shared a huge notebook  <a href=\"https://www.kaggle.com/calebeverett/comparison-of-tta-prediction-procedures\">here</a> as a benchmark.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "888451": "### Submission Notebooks    \n   \nI am delighted to share with you the winning submission in these notebooks,\n    \n1. [Training-Densenet201-Bagging1](https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v2)\n2. [Training-Densenet201-Bagging2](https://www.kaggle.com/coreacasa/flower-class-tpu-densenet-training-v3)\n3. [Training-EfficientnetB7-Bagging1](https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v2)\n4. [Training-EfficientnetB7-Bagging2](https://www.kaggle.com/coreacasa/flower-class-tpu-efficientnet-training-v3)\n5. [Inference-Flower-Classification-TPU](https://www.kaggle.com/coreacasa/flower-classification-tpu-inference-v3)\n        \nYour links are pinned to the key versions. Although you will probably find some parts of the code redundant or not quite efficient, I didn't want to touch a comma to show the composition of the winning solution as it is here. This way, you can check both the clean inputs included in the inference notebooks and the automatic generation of the final solution. \n    \nI just want to add 'a little' to my summary of the competition below and highlight the small details that perhaps made my score stronger.\n\n### Acknowledgements\nFinally the TPU Leaderboard Prizes were settled and as you can imagine I am very happy with the results. Thank you very much to everyone. Thanks to all the Kaggle team and Google-Cloud-TPU for making this competition possible and thanks to the 848 teams that participated in this challenge.\n    \nWhether it is explicitly (sharing notebooks and posting comments and discussions) or implicitly (pushing the leaderboard up with high scores) I think Kagglers are the ones that give a unique value to any challenge.\n    \n### What you can't miss \nI believe that more than a few Data Scientists will come to review the material in 'Flower Classification with TPUs' to get initiated or indeed improve the use of this powerful technology in their projects. If you are one of them here I detail, without order of preference or priority, a series of authors (sorry because I surely left some behind) of kernels whose work in this competition you should review (all Tensorflow).\n    \nThanks @mgornergoogle (Martin Görner), @xiejialun (Xie29), @yihdarshieh (Yih-Dar SHIEH), @cdeotte (Chris Deotte), @calebeverett (Caleb), @romanweilguny (Roman Weilguny) but also don't forget to look at everything else if needed. \n\nI would state that the playground objective, which was aimed at facilitating the use of TPU technology, has been more than achieved.\n\n### The Funny Part \nThere are three key kernels at the beinning of the competition,\n- [ratan123/densenet201-flower-classification-with-tpus](https://www.kaggle.com/ratan123/densenet201-flower-classification-with-tpus)\n- [xhlulu/flowers-tpu-concise-efficientnet-b7](https://www.kaggle.com/xhlulu/flowers-tpu-concise-efficientnet-b7)\n- [wrrosa/tpu-enet-b7-densenet](https://www.kaggle.com/wrrosa/tpu-enet-b7-densenet)\n        \nThe baseline of these notebooks is enough, for with the allowed competition data, to obtain the limit in the range of 0.9700 score. I am referring to the limit because of what we have seen in practice later on.\n        \nOn how to beat the limit of that score: the use of external data. @cdeotte (Chris Deotte) in this discussions (1) [Oxford Flowers - Kaggle Dataset for Everyone](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140964). (2) [How to Score Over LB 0.970+](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329) and @hengck23 (Heng CherKeng) in [[placeholder] external data and how to use them](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/140866) made public the manner.    \n           \nIn my opinion this was the exciting part, I mean the rush to include these new datasets in training and see how the performance of the models improved very fast.\n        \nI have to make a super special mention here for @kirillblinov (Kirill Blinov) and its shared external dataset [tf\\_flower\\_photo\\_tfrec](https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec), great work. In this way, @szacho (Michal Szachniewicz) y @calebeverett (Caleb) made an interesting contribution.\n        \nBy the way, one of the external dataset did not work for me, <code>openimage</code>.\n\n### The Controversial Part\nI am referring here to all the final polemics about whether or not to use external datasets. The host never tried to hide the composition of the images in the test set and their origin. In my particular opinion I would have let the competition run until the end by allowing any source of data. \n        \nIt was also a very hard challenge to work on these new data and in any case the final obstacle is not so much the difficulty of the objective itself but the competition between the participants to overcome a score. The later, however, as long as everyone can access any advantage as happened here.\n        \nYou can be more or less in line with the direction that was finally decided, but you don't have to lose your commitment to respect and follow the rules of the competition for this reason. \n    \nI expressed my disagreement with the final treatment of the Private Leaderboard and explained it [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494#84954).\n\n### Training Topics\nIf you have enjoyed the competition from the beginning, participating in the published notebooks and discussions I think you won't find anything new in the kernels I have shared but you will probably be interested in knowing what could be missing or left over from your solution to achieve a higher score in the Public or Private Leaderboard.\n     \n- **Augmentation** \nI just used the files for an image size of  <code>512x512</code>, I did not preprocess original images. On them, <code>Random Blockout</code> y <code>Horizontal Random Flip</code> were the best choices. \n    \n- **Validation** \nSince the initial validation dataset was very extensive and its information very useful in terms of performance, I divided the validation files into two subsets generating an additional Fold. Something like that looking for cross validation,\n\n\n<pre><code>Fold1&gt;&gt;&gt;\nVALIDATIONFILENAMES1=VALIDATIONFILENAMES[:int(len(VALIDATIONFILENAMES)/2)] \n</code></pre>\n<pre><code>Fold2&gt;&gt;&gt; VALIDATIONFILENAMES2=VALIDATIONFILENAMES[int(len(VALIDATIONFILENAMES)/2):]\n</code></pre>\n<pre><code>TRAININGFILENAMES1=TRAININGFILENAMES+VALIDATIONFILENAMES1  TRAININGFILENAMES2=TRAININGFILENAMES+VALIDATIONFILENAMES2\n</code></pre>\n        \n- **DenseNet-201**\n<code>Imagenet Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. Little or no adjustment here, I kept BATCH_SIZE=128 and I brought to 4 the LR-RAMPUP-EPOCHS and LR-SUSTAIN-EPOCHS parameters of the well-tuned 'CustomLRSchedule' from @mgornergoogle (Martin Görner) in his ' Five 'Flowers' Notebook.\n        \n- **EfficientNet-B7**\n<code>Noisy-Student Weights</code> with <code>tf.keras.layers.GlobalAveragePooling2D()</code>. The rest is the same as Densenet-201 except for 5 the epochs number ramp-up and sustain. \n     \n- **Loss y Optimizer**    \nShared by both models, <code>Loss</code> with <code>'sparse_categorical_crossentropy'</code> and <code>Optimizer</code> with <code>Adam(lr='CustomLRSchedule')</code>.  \n    \n- **Model Checkpoints**\nThe following callback was used to save the best weights according to the loss in validation,   \n<pre><code>mc=tf.keras.callbacks.ModelCheckpoint('name.h5',verbose=1,monitor='val_loss',               \n    mode='min',save_best_only=True,save_weights_only=True)\n</code></pre>\n                    \n- **2-Baggings**\nWhile the use of the Blockout made the training results more stable, in general the performance of the models varied significantly (+/- 0.001) from run to run. Considering this problem of randomness I trained each model twice with the same basic setting. You can presume that the random augmentations are relevant here but even without them the variability of the TPU workouts is something important to consider. <code>NUM_EPOCHS</code> number <code>35, 50</code> to DenseNet and <code>35, 40</code> to EfficientNet.\n    \n### Inference Topics\nTherefore with the details above I actually had 8 trained models available with which to perform the classification inference.\n1. DenseNet (4), 2xFold with <code>best-checkpoints</code> epochs <code>24-18, Fold1</code> and <code>25-24, Fold2</code>.\n2. EfficientNet (4),  2xFold with <code>best-checkpoints</code> on epochs <code>17-16, Fold1</code> and <code>14-13, Fold2</code>.\n    \n- **Test Time Augmentation**\nThe strategy here was to explode the training augmentated with 2 types of prediction. One of them (1) with the original 512x512 images without any transformation and the other (2) with 100% of the images flipped using the function <code>tf.image.flip_left_right(image)</code> as the only augment. \n    \n- **Averaging Out of Folds**\nI join the oofs of each bagging and model with a <code>simple average</code> of the <code>logits</code>, not probabilities. Following TTA strategy I perform the same operation with the full-flip images and let's say I get the following sets,\n    \n<pre><code>1 oof-non-flip-densenet\n1 oof-non-flip-efficientnet\n1 oof-full-flip-densenet\n1 oof-full-flip-efficientnet\n</code></pre>\n        \n- **Weighting Models OOFs Non-Flip**\nWith the logit transformation still applied, I use a simple optimizer to linearly maximize the weight of each model in the non-flip context.\n    \n<pre><code>Best weighting to EfficientNet-B7: 0.71\nValidation F1-Score: 0.969903\n</code></pre>\n    \n- **Weighting Models OOFs Full-Flip**\nFollowing the same procedure I get now,\n    \n<pre><code>Best weighting for EfficientNet-B7:  0.65\nValidation F1-Score: 0.970014\n</code></pre>\n        \n- **One Last Optimization**\nA final step was to obtain a new weighting but now to calibrate the power of the TTA strategy between full-flip and non-flip derived classifications.\n    \n<pre><code>Best weighting for validation Non-Flip: 0.05 (any value 0.05 to 0.21)\nValidation F1-Score: 0.970367\n</code></pre>\n    \n### That's all folks !!!",
    "888919": "Congratulations! and thanks for the mention.",
    "888931": "Thank you for sharing a great job!",
    "888993": "Thank you for sharing the solution. I like the fact that you ensemble at the logit level instead of ensembling the probabilities. (Not quite sure how you get the logits from the probabilities though.)",
    "889142": "Hi @mgornergoogle , thanks to you too. I used the <code>logit</code> and <code>expit</code> (inverse logit) functions of <code>scipy.special</code> to get these transformations. To avoid -inf and +inf values in the logit output (for subsequent manipulations) I use a simple <code>np.nan_to_num</code>. \n    \nAbout the performance of these transformations, but with a focus on TTA, @calebeverett shared a huge notebook  [here](https://www.kaggle.com/calebeverett/comparison-of-tta-prediction-procedures) as a benchmark.",
    "969843": "Congratulations Manuel. You achieved a fantastic score using only competition train data. Well done!",
    "970319": "This solution will be better with your post here... Thanks a lot!",
    "987849": "Thanks for sharing, I learned a lot from looking through your notebooks!",
    "987853": "Thanks @benghertner, that's great for me!",
    "1050391": "thank you very much very insightful even for newbie (me)"
  },
  "source": "meta"
}