{
  "id": 175891,
  "title": "TF vs Pyorch mystery solved",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175891",
  "author_name": "CPMP",
  "post_date": "2020-08-19T18:42:40.222000",
  "votes": 29,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Many were puzzled that TF models with a rather low CV had such great public LB.  We now know that on private LB Pytorch models were competitive.</p>\n<p>So, what's the trick here?  It seems that Chris Deotte baseline was overfitting public LB, see for instance the scores of his most popular notebook: <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F059179f975ed612d17b23ec901e5b2b6%2FScreenshot_2020-08-19%20Triple%20Stratified%20KFold%20with%20TFRecords.png?generation=1597862437376532&amp;alt=media\" alt=\"\"></p>\n<p>The relatively low CV was not the anomaly.  The score was the anomaly.</p>",
  "messages": [
    {
      "id": 977848,
      "postDate": "2020-08-19T18:42:40.223Z",
      "content": "<p>Many were puzzled that TF models with a rather low CV had such great public LB.  We now know that on private LB Pytorch models were competitive.</p>\n<p>So, what's the trick here?  It seems that Chris Deotte baseline was overfitting public LB, see for instance the scores of his most popular notebook: <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F059179f975ed612d17b23ec901e5b2b6%2FScreenshot_2020-08-19%20Triple%20Stratified%20KFold%20with%20TFRecords.png?generation=1597862437376532&amp;alt=media\" alt=\"\"></p>\n<p>The relatively low CV was not the anomaly.  The score was the anomaly.</p>",
      "rawMarkdown": "Many were puzzled that TF models with a rather low CV had such great public LB.  We now know that on private LB Pytorch models were competitive.\n\nSo, what's the trick here?  It seems that Chris Deotte baseline was overfitting public LB, see for instance the scores of his most popular notebook: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F059179f975ed612d17b23ec901e5b2b6%2FScreenshot_2020-08-19%20Triple%20Stratified%20KFold%20with%20TFRecords.png?generation=1597862437376532&alt=media)\n\nThe relatively low CV was not the anomaly.  The score was the anomaly.\n",
      "votes": 28
    },
    {
      "id": 977872,
      "postDate": "2020-08-19T19:06:09.340Z",
      "content": "<p>I'm not sure if that solves the mystery. Here are the two highest scoring PyTorch notebooks. Both TF and PyTorch notebooks have Private LB score that is 0.02 lower than Public LB score. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Faf2f8629b5ce574f5f02af2e1d082a71%2Fc1.png?generation=1597863858079347&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F3f566653c07a21e85da1ebbf07bc71e7%2Fc2.png?generation=1597863868755546&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'm not sure if that solves the mystery. Here are the two highest scoring PyTorch notebooks. Both TF and PyTorch notebooks have Private LB score that is 0.02 lower than Public LB score. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Faf2f8629b5ce574f5f02af2e1d082a71%2Fc1.png?generation=1597863858079347&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F3f566653c07a21e85da1ebbf07bc71e7%2Fc2.png?generation=1597863868755546&alt=media)",
      "votes": 22,
      "replies": [
        {
          "id": 977969,
          "postDate": "2020-08-19T20:30:45.120Z",
          "content": "<p>Fai enough,.  My point is that these notebooks overfit.   Yours is also better overall.  People trying to replicate your notebook without overfitting could only conclude that pytorch had an issue compared to TF.</p>\n<p>For example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01.  If I wanted to replicate your results then I needed to have a CV 0.02 higher than yours. </p>\n<p>That's why I say mystery is solved.  To be clear: the mystery is why you got such a high LB compared to the CV of your notebook.</p>",
          "rawMarkdown": "Fai enough,.  My point is that these notebooks overfit.   Yours is also better overall.  People trying to replicate your notebook without overfitting could only conclude that pytorch had an issue compared to TF.\n\nFor example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01.  If I wanted to replicate your results then I needed to have a CV 0.02 higher than yours. \n\nThat's why I say mystery is solved.  To be clear: the mystery is why you got such a high LB compared to the CV of your notebook."
        },
        {
          "id": 978294,
          "postDate": "2020-08-20T04:59:58.353Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 977937,
      "postDate": "2020-08-19T20:01:23.873Z",
      "content": "<p>I shared my thoughts <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175461\" target=\"_blank\">here</a> on why I think PyTorch performs better than TF. I'll paraphase here: </p>\n<p>In short, I believe PyTorch has better ecosystem/support for augmentations. The reason for a huge gap between Public LB and CV on Chris' notebook was likely due to luck. When I added <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175372\" target=\"_blank\">heavy augmentations</a> to Chris' triple stratified notebook, I increased my CV from 0.91 to 0.9379. The public LB and private LB were 0.9394 and 0.9325 respectively. The public LB and CV gap was a pretty big indicator that something wasn't adding up.</p>\n<p>I don't think there's a comparable library for TF like <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/66643\" target=\"_blank\">Albumenations</a> for PyTorch. I refer to Albumentations because that's what first place used. Second place used RandAugment. But the idea is still the same - better support for augmentation library.</p>",
      "rawMarkdown": "I shared my thoughts [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175461) on why I think PyTorch performs better than TF. I'll paraphase here: \n\nIn short, I believe PyTorch has better ecosystem/support for augmentations. The reason for a huge gap between Public LB and CV on Chris' notebook was likely due to luck. When I added [heavy augmentations](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175372) to Chris' triple stratified notebook, I increased my CV from 0.91 to 0.9379. The public LB and private LB were 0.9394 and 0.9325 respectively. The public LB and CV gap was a pretty big indicator that something wasn't adding up.\n\nI don't think there's a comparable library for TF like [Albumenations](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/66643) for PyTorch. I refer to Albumentations because that's what first place used. Second place used RandAugment. But the idea is still the same - better support for augmentation library.",
      "votes": 15,
      "replies": [
        {
          "id": 979603,
          "postDate": "2020-08-21T01:45:23.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> Would you say that heavy augmentations was the key to improve the CV?</p>",
          "rawMarkdown": "@teeyee314 Would you say that heavy augmentations was the key to improve the CV?",
          "votes": 1
        },
        {
          "id": 979658,
          "postDate": "2020-08-21T03:26:10.663Z",
          "content": "<p>for this competition, single model - yes. in general, better augmentations should lead to better cv score. when you read people's write ups, you will notice mention to specific augmentations used.</p>",
          "rawMarkdown": "for this competition, single model - yes. in general, better augmentations should lead to better cv score. when you read people's write ups, you will notice mention to specific augmentations used.",
          "votes": 2
        },
        {
          "id": 980692,
          "postDate": "2020-08-21T19:23:48.450Z",
          "content": "<p>Tim, it's my first competition and I'm still discovering things. Albumentations looks fantastic, and on my simple benchmark, it was about 3x faster. Their documentation says that you can use it with TF too, more here - <a href=\"https://albumentations.ai/docs/api_reference/augmentations/transforms/\" target=\"_blank\">https://albumentations.ai/docs/api_reference/augmentations/transforms/</a></p>\n<p>Is there a reason why you think this library only works with Pytorch?</p>",
          "rawMarkdown": "Tim, it's my first competition and I'm still discovering things. Albumentations looks fantastic, and on my simple benchmark, it was about 3x faster. Their documentation says that you can use it with TF too, more here - https://albumentations.ai/docs/api_reference/augmentations/transforms/\n\nIs there a reason why you think this library only works with Pytorch?"
        },
        {
          "id": 981554,
          "postDate": "2020-08-22T14:33:47.280Z",
          "content": "<p><a href=\"https://www.kaggle.com/sudhanshuraheja\" target=\"_blank\">@sudhanshuraheja</a> Regarding external library usage with TPU - see this <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132066\" target=\"_blank\">discussion</a>. Maybe this will change in the future. </p>",
          "rawMarkdown": "@sudhanshuraheja Regarding external library usage with TPU - see this [discussion](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132066). Maybe this will change in the future. ",
          "votes": 1
        },
        {
          "id": 981568,
          "postDate": "2020-08-22T14:51:27.040Z",
          "content": "<p>Numpy/opencv preprocessing functions, works with the tf.data.dataset as well, but only on GPU so far, there is still a bug for TPU, see: <a href=\"https://github.com/tensorflow/tensorflow/issues/30818\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/30818</a></p>\n<pre><code>    def applyAug(x):\n        x = x.numpy()\n        return aug(image=x)\n\n    def _preprocess_for_train(filename, label):\n        image_bytes = tf.io.read_file(filename)\n        image = tf.cond(\n          tf.image.is_jpeg(image_bytes),\n          lambda: tf.image.decode_jpeg(image_bytes, channels=3),\n          lambda: tf.image.decode_png(image_bytes, channels=3))\n\n\n        image = tf.py_function(func=applyAug, inp=[image], Tout=tf.float32)\n        image.set_shape(tf.TensorShape([None, None, None]))  \n\n        image = tf.image.resize(image, im_shape)\n        image = tf.cast(image, dtype=tf.float32) / 255.0\n        label = tf.one_hot(label, num_classes)\n        label = tf.cast(label, tf.float32)\n        return image, label\n\n    dataset = tf.data.Dataset.from_tensor_slices((filenames, labels))\n    autotune = tf.data.experimental.AUTOTUNE\n    dataset = dataset.map(_preprocess_for_train, num_parallel_calls=autotune)\n    dataset = dataset.batch(batch_size, drop_remainder=True) \n</code></pre>\n<p>Keep in mind, using this with GPU will slow down the training a little bit, I can only assume even if they fix it for TPU, it will cause horrible throttling for TPU.</p>",
          "rawMarkdown": "Numpy/opencv preprocessing functions, works with the tf.data.dataset as well, but only on GPU so far, there is still a bug for TPU, see: https://github.com/tensorflow/tensorflow/issues/30818\n\n```\n    def applyAug(x):\n        x = x.numpy()\n        return aug(image=x)\n\n    def _preprocess_for_train(filename, label):\n        image_bytes = tf.io.read_file(filename)\n        image = tf.cond(\n          tf.image.is_jpeg(image_bytes),\n          lambda: tf.image.decode_jpeg(image_bytes, channels=3),\n          lambda: tf.image.decode_png(image_bytes, channels=3))\n      \n        \n        image = tf.py_function(func=applyAug, inp=[image], Tout=tf.float32)\n        image.set_shape(tf.TensorShape([None, None, None]))  \n\n        image = tf.image.resize(image, im_shape)\n        image = tf.cast(image, dtype=tf.float32) / 255.0\n        label = tf.one_hot(label, num_classes)\n        label = tf.cast(label, tf.float32)\n        return image, label\n\n    dataset = tf.data.Dataset.from_tensor_slices((filenames, labels))\n    autotune = tf.data.experimental.AUTOTUNE\n    dataset = dataset.map(_preprocess_for_train, num_parallel_calls=autotune)\n    dataset = dataset.batch(batch_size, drop_remainder=True) \n```\n\nKeep in mind, using this with GPU will slow down the training a little bit, I can only assume even if they fix it for TPU, it will cause horrible throttling for TPU.",
          "votes": 2
        }
      ]
    },
    {
      "id": 978477,
      "postDate": "2020-08-20T07:47:54Z",
      "content": "<p>There is no difference between the accuracy of the PyTorch vs TensorFlow. A few months ago I had to switch to TensorFlow and replicate my PyTorch results with it. At first, I was baffled how can TF be worse, then I figured out, the difference was my lack of skills to implement the same training process. After dialling it in, it was the exact same flow, and I got almost the same results (+- 0.0001-0.0002 F1) as in PyTorch (non-kaggle dataset).</p>\n<p>Everybody is saying that PyTorch is better for research, I believe this is only true if you don't know the framework inside and out. After 2 months with TF, I feel the ceiling is so much higher. The transition was very hard especially because I had to learn both tf 1.15 and tf 2.2&gt; at the same time, but it was worth it, I'm not looking back.</p>",
      "rawMarkdown": "There is no difference between the accuracy of the PyTorch vs TensorFlow. A few months ago I had to switch to TensorFlow and replicate my PyTorch results with it. At first, I was baffled how can TF be worse, then I figured out, the difference was my lack of skills to implement the same training process. After dialling it in, it was the exact same flow, and I got almost the same results (+- 0.0001-0.0002 F1) as in PyTorch (non-kaggle dataset).\n\nEverybody is saying that PyTorch is better for research, I believe this is only true if you don't know the framework inside and out. After 2 months with TF, I feel the ceiling is so much higher. The transition was very hard especially because I had to learn both tf 1.15 and tf 2.2> at the same time, but it was worth it, I'm not looking back.",
      "votes": 7,
      "replies": [
        {
          "id": 978585,
          "postDate": "2020-08-20T09:14:37.337Z",
          "content": "<p>Great reply. Just to add something, I would include the following article here about <a href=\"https://towardsdatascience.com/pytorch-vs-tensorflow-in-2020-fe237862fae1\" target=\"_blank\">Pytorch vs Tensorflow in 2020</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F6afbcc31ee52007e844bdc6d4720f416%2F1.png?generation=1597914771146704&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Great reply. Just to add something, I would include the following article here about [Pytorch vs Tensorflow in 2020](https://towardsdatascience.com/pytorch-vs-tensorflow-in-2020-fe237862fae1)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F6afbcc31ee52007e844bdc6d4720f416%2F1.png?generation=1597914771146704&alt=media)"
        },
        {
          "id": 980699,
          "postDate": "2020-08-21T19:30:13.543Z",
          "content": "<p>Except they are not \"exactly the same\" at all.</p>",
          "rawMarkdown": "Except they are not \"exactly the same\" at all."
        },
        {
          "id": 981448,
          "postDate": "2020-08-22T12:56:22.283Z",
          "content": "<p>Nope, they're now. 😀</p>",
          "rawMarkdown": "Nope, they're now. 😀"
        },
        {
          "id": 981466,
          "postDate": "2020-08-22T13:15:32.130Z",
          "content": "<p>Give Pytorch a try and we'll discuss again.  On paper TF2x looks close to Pytorch indeed, but there is a difference between the look and the reality. </p>",
          "rawMarkdown": "Give Pytorch a try and we'll discuss again.  On paper TF2x looks close to Pytorch indeed, but there is a difference between the look and the reality. "
        },
        {
          "id": 981520,
          "postDate": "2020-08-22T14:00:22.670Z",
          "content": "<p>Obviously the coding workflow is not the same, not talking about that. I'm talking about the results, the math behind the adam optimizer etc. <br>\nKeep in mind, most of the TF notebooks used TPU s with much higher batch sizes, compared to the pytorch ones with GPUs. Batch size does affect the accuracy, but as I said if you compare the exact training flow, you should have very-very close results.</p>",
          "rawMarkdown": "Obviously the coding workflow is not the same, not talking about that. I'm talking about the results, the math behind the adam optimizer etc. \nKeep in mind, most of the TF notebooks used TPU s with much higher batch sizes, compared to the pytorch ones with GPUs. Batch size does affect the accuracy, but as I said if you compare the exact training flow, you should have very-very close results.",
          "votes": 3
        },
        {
          "id": 981570,
          "postDate": "2020-08-22T14:54:31.743Z",
          "content": "<p>Suppose if there is a difference in result between 2 same training settings, which factors cause it? </p>",
          "rawMarkdown": "Suppose if there is a difference in result between 2 same training settings, which factors cause it? "
        },
        {
          "id": 981624,
          "postDate": "2020-08-22T15:37:08.710Z",
          "content": "<p>The reality is,  it takes much less time and headache to code advanced workflow on Pytorch than tensorflow.  Otherwise TF 2.x wouldn't try to be as close as possible to Pytorch. </p>\n<p>OpenAI (now ClosedAI) has switched form TF to Pytorch for this reason. </p>\n<p>I'm saying this while I don't like Facebook at all. And Google has contributed much more on DL than them. </p>\n<p>P.S. : I was much more familiar with TF (both 1.X and 2.X ) than Pytorch 9 months ago. </p>",
          "rawMarkdown": "The reality is,  it takes much less time and headache to code advanced workflow on Pytorch than tensorflow.  Otherwise TF 2.x wouldn't try to be as close as possible to Pytorch. \n\n OpenAI (now ClosedAI) has switched form TF to Pytorch for this reason. \n\nI'm saying this while I don't like Facebook at all. And Google has contributed much more on DL than them. \n\nP.S. : I was much more familiar with TF (both 1.X and 2.X ) than Pytorch 9 months ago. "
        },
        {
          "id": 981647,
          "postDate": "2020-08-22T15:50:51.707Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> san, I don't find any offensive words in your response, don't get it why it's been downvoted! </p>\n<p>However, I should say, we don't have exact same comprehensive working experience in both TF and PyTorch. Compare to TF, we've very little experience in PyTorch (until now). PyTorch is awesome, no doubt and the coding experience in TF was much pain. My point only was, the coding experience between these two is almost similar now (AFAIK). <a href=\"https://www.kaggle.com/doncalculator\" target=\"_blank\">@doncalculator</a> san points some great stuff. But I would request you to explain these two more in detail (from an experimental perspective, with the latest TF). </p>",
          "rawMarkdown": "@cpmpml san, I don't find any offensive words in your response, don't get it why it's been downvoted! \n\nHowever, I should say, we don't have exact same comprehensive working experience in both TF and PyTorch. Compare to TF, we've very little experience in PyTorch (until now). PyTorch is awesome, no doubt and the coding experience in TF was much pain. My point only was, the coding experience between these two is almost similar now (AFAIK). @doncalculator san points some great stuff. But I would request you to explain these two more in detail (from an experimental perspective, with the latest TF). ",
          "votes": 1
        },
        {
          "id": 982077,
          "postDate": "2020-08-23T03:41:07.603Z",
          "content": "<p>Is there a way to do something like Funcional API of Keras with PyTorch? </p>",
          "rawMarkdown": "Is there a way to do something like Funcional API of Keras with PyTorch? "
        },
        {
          "id": 982211,
          "postDate": "2020-08-23T06:46:12.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Alright, I created identical training flow for both tensorflow and pytorch, I will post the notebooks today I think. Tbh, I'm not sure about the different Efficientnet implementations qubvel's vs lukemelas, but that's gonna be the only difference. Will see.</p>",
          "rawMarkdown": "@ipythonx Alright, I created identical training flow for both tensorflow and pytorch, I will post the notebooks today I think. Tbh, I'm not sure about the different Efficientnet implementations qubvel's vs lukemelas, but that's gonna be the only difference. Will see.",
          "votes": 1
        },
        {
          "id": 982404,
          "postDate": "2020-08-23T10:50:04.227Z",
          "content": "<p>Awesome. 👍</p>",
          "rawMarkdown": "Awesome. 👍"
        },
        {
          "id": 982665,
          "postDate": "2020-08-23T15:07:10.340Z",
          "content": "<p>I published the two notebooks:<br>\n<a href=\"https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-1-tensorflow\" target=\"_blank\">https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-1-tensorflow</a><br>\n<a href=\"https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-2-pytorch\" target=\"_blank\">https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-2-pytorch</a>    </p>\n<p>Here are the results:<br>\n-----------CV,--------------------PRIVATE, PUBLIC:<br>\nPytorch:    0.9820760704110646, 0.9077, 0.9196<br>\nTensorflow: 0.982470040003284   0.9043  0.9203</p>\n<p>Because of the low amount of positive images, I chose to include the external <br>\nmaligns into the validation set to have a more reliable CV score.</p>\n<p>Although these scores are close enough to support my statement, due to the low amount of<br>\nmalignant images, there is a high variance between runs, so if we want to prove that<br>\nthe results are there same across the two frameworks, this experiment should be rerun 5-10 times.</p>",
          "rawMarkdown": "I published the two notebooks:\nhttps://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-1-tensorflow\nhttps://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-2-pytorch    \n\nHere are the results:\n-----------CV,--------------------PRIVATE, PUBLIC:\nPytorch:    0.9820760704110646, 0.9077, 0.9196\nTensorflow: 0.982470040003284   0.9043  0.9203\n    \nBecause of the low amount of positive images, I chose to include the external \nmaligns into the validation set to have a more reliable CV score.\n\nAlthough these scores are close enough to support my statement, due to the low amount of\nmalignant images, there is a high variance between runs, so if we want to prove that\nthe results are there same across the two frameworks, this experiment should be rerun 5-10 times.",
          "votes": 2
        },
        {
          "id": 982707,
          "postDate": "2020-08-23T15:42:18.617Z",
          "content": "<blockquote>\n  <p>There is no difference between the accuracy of the PyTorch vs TensorFlow</p>\n</blockquote>\n<p>I agree with that statement in general and I assumed it was true here too.  It is why there was a mystery:  no one could reproduce Chris high LB score when using Pytorch.  Now that we know that the high public Lb score was an overfit.</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> san, I don't find any offensive words in your response, don't get it why it's been downvoted! </p>\n</blockquote>\n<p>I didn't even notice, LOL.  Please meet my serial downvoter ;)</p>",
          "rawMarkdown": "> There is no difference between the accuracy of the PyTorch vs TensorFlow\n\nI agree with that statement in general and I assumed it was true here too.  It is why there was a mystery:  no one could reproduce Chris high LB score when using Pytorch.  Now that we know that the high public Lb score was an overfit.\n\n> @cpmpml san, I don't find any offensive words in your response, don't get it why it's been downvoted! \n\nI didn't even notice, LOL.  Please meet my serial downvoter ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 977962,
      "postDate": "2020-08-19T20:20:10Z",
      "content": "<p>The sub with the high score was definitely lucky. I reran it two times, once it was 0.01 worse than the LB score of the NB, second time I removed the 2018 data and it was higher than my first run, but slightly lower than the LB run.</p>\n<p>I think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.</p>\n<p>I still thought for the longest time there is something in that NB that I need to figure out, and it cost me way too many brain cells because I have such a hard time understanding all the tensorflow augmentation code 😁</p>\n<p>I just checked the results:</p>\n<ul>\n<li>No extra data on that kernel: 0.9428 public LB, 0.9104 private LB</li>\n<li>Extra data as provided: 0.9342 public LB, 0.9256 private LB</li>\n</ul>",
      "rawMarkdown": "The sub with the high score was definitely lucky. I reran it two times, once it was 0.01 worse than the LB score of the NB, second time I removed the 2018 data and it was higher than my first run, but slightly lower than the LB run.\n\nI think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.\n\nI still thought for the longest time there is something in that NB that I need to figure out, and it cost me way too many brain cells because I have such a hard time understanding all the tensorflow augmentation code 😁\n\nI just checked the results:\n\n- No extra data on that kernel: 0.9428 public LB, 0.9104 private LB\n- Extra data as provided: 0.9342 public LB, 0.9256 private LB",
      "votes": 7,
      "replies": [
        {
          "id": 977973,
          "postDate": "2020-08-19T20:34:18.980Z",
          "content": "<p>This is interesting, I also think that early stopping is not the best way to do experiments, so your suggestions would be to run a few times and see at which epoch the model converges (let's say epoch 15), then write a schedule to end in 15 epochs to compare models even while changing parameters?</p>",
          "rawMarkdown": "This is interesting, I also think that early stopping is not the best way to do experiments, so your suggestions would be to run a few times and see at which epoch the model converges (let's say epoch 15), then write a schedule to end in 15 epochs to compare models even while changing parameters?",
          "votes": 2
        },
        {
          "id": 977977,
          "postDate": "2020-08-19T20:36:18.703Z",
          "content": "<p>Yes exactly, if your best epoch is let's say between 13-15 out of 15 epochs it can be fine, but in this kernel it was sometimes 3 or 4 out of 10 and in some other folds 10, which is definitely way more random.</p>\n<p>The only way imho is to take fixed epochs and multiple bags. We always fitted for 3 bags 5-fold cv with fixed epochs. Correlation was OK, I have seen worse, albeit not perfect.</p>",
          "rawMarkdown": "Yes exactly, if your best epoch is let's say between 13-15 out of 15 epochs it can be fine, but in this kernel it was sometimes 3 or 4 out of 10 and in some other folds 10, which is definitely way more random.\n\nThe only way imho is to take fixed epochs and multiple bags. We always fitted for 3 bags 5-fold cv with fixed epochs. Correlation was OK, I have seen worse, albeit not perfect.",
          "votes": 1
        },
        {
          "id": 977978,
          "postDate": "2020-08-19T20:38:20.123Z",
          "content": "<blockquote>\n  <p>I think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.</p>\n</blockquote>\n<p>I think multiple checkpoints averaging is the only way to stabilize CV.   I did it in all my experiments and  stopped worrying about public LB. Given it was not really always correlated with my CV</p>",
          "rawMarkdown": "> I think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.\n\nI think multiple checkpoints averaging is the only way to stabilize CV.   I did it in all my experiments and  stopped worrying about public LB. Given it was not really always correlated with my CV",
          "votes": 1
        },
        {
          "id": 977980,
          "postDate": "2020-08-19T20:40:21.813Z",
          "content": "<p>Agreed, I can definitely see the issue, thanks!</p>",
          "rawMarkdown": "Agreed, I can definitely see the issue, thanks!"
        },
        {
          "id": 978000,
          "postDate": "2020-08-19T20:59:06.647Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> There is a difference between choosing the same epochs for checkpoint ensembling (can make sense) vs. picking random epochs (makes it random).</p>",
          "rawMarkdown": "@serigne There is a difference between choosing the same epochs for checkpoint ensembling (can make sense) vs. picking random epochs (makes it random)."
        },
        {
          "id": 978421,
          "postDate": "2020-08-20T07:00:26.460Z",
          "content": "<p>What I did was to run relatively large number of epochs and chose 3 or 4  to average based on multiple criteria.  For instance for some folds I could chose Epochs : 12, 25, 30 while on others:  Epochs 31, 34, 40 etc.</p>",
          "rawMarkdown": "What I did was to run relatively large number of epochs and chose 3 or 4  to average based on multiple criteria.  For instance for some folds I could chose Epochs : 12, 25, 30 while on others:  Epochs 31, 34, 40 etc.\n\n"
        }
      ]
    },
    {
      "id": 980076,
      "postDate": "2020-08-21T10:16:20.193Z",
      "content": "<p>This comp has a very high randomness in score. Unusually high for an NN competition. </p>",
      "rawMarkdown": "This comp has a very high randomness in score. Unusually high for an NN competition. ",
      "votes": 3,
      "replies": [
        {
          "id": 980095,
          "postDate": "2020-08-21T10:25:00.800Z",
          "content": "<p>Indeed, probably because of the very low number of positive samples.  The metric is very sensitive to one of them being misclassified.</p>",
          "rawMarkdown": "Indeed, probably because of the very low number of positive samples.  The metric is very sensitive to one of them being misclassified.",
          "votes": 1
        }
      ]
    },
    {
      "id": 979737,
      "postDate": "2020-08-21T05:11:46.040Z",
      "content": "<p>My submitted custom EfficientNet-B6 384x384 noisy student (based on Chris popular kernel) + effective Focal Loss parameter computation has CV: 0.942, Public LB: 0.9446, and Private LB: 0.943 though it was hard to recreate the result may be because of reproducibility concern in TPU TensorFlow. The gap is between 0.004 - 0.007</p>",
      "rawMarkdown": "My submitted custom EfficientNet-B6 384x384 noisy student (based on Chris popular kernel) + effective Focal Loss parameter computation has CV: 0.942, Public LB: 0.9446, and Private LB: 0.943 though it was hard to recreate the result may be because of reproducibility concern in TPU TensorFlow. The gap is between 0.004 - 0.007",
      "votes": 3,
      "replies": [
        {
          "id": 979768,
          "postDate": "2020-08-21T05:35:54.083Z",
          "content": "<p>So, looks like the difference comes from using effective loss function! Is it? </p>",
          "rawMarkdown": "So, looks like the difference comes from using effective loss function! Is it? "
        },
        {
          "id": 979798,
          "postDate": "2020-08-21T05:58:22.203Z",
          "content": "<p>My model architecture consists of the following: 1. Use base model output as input to convolutional block attention modules (CBAM), 2. Create 2 outputs and average it: 2a. use CBAM output as input to Global Average Pooling (GAP) and create a final dense layer with sigmoid activation, 2b. use CBAM output as input to Attention Weighted Average Pooling and create a final dense layer with sigmoid activation. Some of my models have large gaps in CV and LB (0.008, 0.012) given I used that said loss function. Maybe because of the random generated processed output of attention modules.</p>",
          "rawMarkdown": "My model architecture consists of the following: 1. Use base model output as input to convolutional block attention modules (CBAM), 2. Create 2 outputs and average it: 2a. use CBAM output as input to Global Average Pooling (GAP) and create a final dense layer with sigmoid activation, 2b. use CBAM output as input to Attention Weighted Average Pooling and create a final dense layer with sigmoid activation. Some of my models have large gaps in CV and LB (0.008, 0.012) given I used that said loss function. Maybe because of the random generated processed output of attention modules."
        },
        {
          "id": 979826,
          "postDate": "2020-08-21T06:22:14.107Z",
          "content": "<p>Interesting. I've posted my solution <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175721\" target=\"_blank\">here</a>. May I ask to know is the Attention Weighted Avg Pooling that you've used are the same as mine. </p>",
          "rawMarkdown": "Interesting. I've posted my solution [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175721). May I ask to know is the Attention Weighted Avg Pooling that you've used are the same as mine. "
        },
        {
          "id": 979827,
          "postDate": "2020-08-21T06:22:19.397Z",
          "content": "<p>Loss Function: </p>\n<pre><code>pos_weight = ((1 / total number of positives per fold)*total number of samples per fold) / 2\nneg_weight = ((1 / total number of negatives per fold)*total number of samples per fold) / 2\n\nBinary Focal Loss where pos_weight=pos_weight and gamma=pos_weight*neg_weight\n</code></pre>\n<p><a href=\"https://github.com/artemmavrin/focal-loss\" target=\"_blank\">https://github.com/artemmavrin/focal-loss</a></p>",
          "rawMarkdown": "Loss Function: \n\n\tpos_weight = ((1 / total number of positives per fold)*total number of samples per fold) / 2\n\tneg_weight = ((1 / total number of negatives per fold)*total number of samples per fold) / 2\n\n\tBinary Focal Loss where pos_weight=pos_weight and gamma=pos_weight*neg_weight\n\nhttps://github.com/artemmavrin/focal-loss"
        },
        {
          "id": 979828,
          "postDate": "2020-08-21T06:24:15.127Z",
          "content": "<p>I refer to this link for Attention Weighted Average: <a href=\"https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\" target=\"_blank\">https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py</a></p>",
          "rawMarkdown": "I refer to this link for Attention Weighted Average: https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py",
          "votes": 1
        },
        {
          "id": 979856,
          "postDate": "2020-08-21T06:51:17.617Z",
          "content": "<p>oh, I see. I've used the Guanshuo Xu modified version of it. As far as I can remember he also used the same repo. -)</p>",
          "rawMarkdown": "oh, I see. I've used the Guanshuo Xu modified version of it. As far as I can remember he also used the same repo. -)"
        }
      ]
    },
    {
      "id": 977861,
      "postDate": "2020-08-19T18:51:54.510Z",
      "content": "<p>Chris Deotte needs to share notebooks with a Disclaimer 😝</p>",
      "rawMarkdown": "Chris Deotte needs to share notebooks with a Disclaimer 😝",
      "votes": 3
    },
    {
      "id": 978400,
      "postDate": "2020-08-20T06:41:21.240Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I would like to learn from a Grandmaster such as yourself about creating a stable and LB correlated CV.</p>\n<p><code>For example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01.</code></p>\n<p>Even Chris's Triple Stratified Leak-Free KFold CV still has a large gap, but <strong>your CV has gap of less than 0.01</strong>, so could you please share your techniques on this accomplishment? Do you have different techniques for different types of competitions?</p>",
      "rawMarkdown": "@cpmpml I would like to learn from a Grandmaster such as yourself about creating a stable and LB correlated CV.\n\n`For example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01. `\n\nEven Chris's Triple Stratified Leak-Free KFold CV still has a large gap, but **your CV has gap of less than 0.01**, so could you please share your techniques on this accomplishment? Do you have different techniques for different types of competitions?",
      "votes": 4,
      "replies": [
        {
          "id": 978521,
          "postDate": "2020-08-20T08:25:02.207Z",
          "content": "<p>Thanks for asking.</p>\n<p>You can read my previous writeups as I will not create one here.  Here is some info still.</p>\n<p>I always focus on getting a reliable CV.  I used the same approach here as in Tweet Sentiment: bagging several runs to average out CV variance. </p>\n<p>I was surprised to get a small gap.  Having a large CV LB gap isn't an issue in itself if CV and LB are correlated.  Here, my CV-LB difference was 0 += 0.002.</p>\n<p>When I manage to get a reliable CV in a competition, which is not always the case, then I keep a change only if it improves both CV and public LB.  </p>\n<p>In this competition I had few outliers when it comes to CV LB gap.  A first one was when I added 2019 positive examples only.  CV improved and LB sunk.  Second outlier was last day when ensembling.  I overfit to public LB.  I wish I had more time to fix it.</p>\n<p>TL;DR  I use the same overall approach across competitions: reduce CV variance with bagging if need be.  It means runs take longer as they need to be repeated, but progress is steady.</p>",
          "rawMarkdown": "Thanks for asking.\n\nYou can read my previous writeups as I will not create one here.  Here is some info still.\n\nI always focus on getting a reliable CV.  I used the same approach here as in Tweet Sentiment: bagging several runs to average out CV variance. \n\nI was surprised to get a small gap.  Having a large CV LB gap isn't an issue in itself if CV and LB are correlated.  Here, my CV-LB difference was 0 += 0.002.\n\nWhen I manage to get a reliable CV in a competition, which is not always the case, then I keep a change only if it improves both CV and public LB.  \n\nIn this competition I had few outliers when it comes to CV LB gap.  A first one was when I added 2019 positive examples only.  CV improved and LB sunk.  Second outlier was last day when ensembling.  I overfit to public LB.  I wish I had more time to fix it.\n\nTL;DR  I use the same overall approach across competitions: reduce CV variance with bagging if need be.  It means runs take longer as they need to be repeated, but progress is steady.\n\n",
          "votes": 5
        },
        {
          "id": 984826,
          "postDate": "2020-08-25T10:15:40.167Z",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> Is the above what you asked for?</p>",
          "rawMarkdown": "@sirishks Is the above what you asked for?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 977872,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-19T19:06:09.340000",
      "content": "<p>I'm not sure if that solves the mystery. Here are the two highest scoring PyTorch notebooks. Both TF and PyTorch notebooks have Private LB score that is 0.02 lower than Public LB score. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Faf2f8629b5ce574f5f02af2e1d082a71%2Fc1.png?generation=1597863858079347&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F3f566653c07a21e85da1ebbf07bc71e7%2Fc2.png?generation=1597863868755546&amp;alt=media\" alt=\"\"></p>",
      "votes": 22,
      "replies": [
        {
          "id": 977969,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-19T20:30:45.120000",
          "content": "<p>Fai enough,.  My point is that these notebooks overfit.   Yours is also better overall.  People trying to replicate your notebook without overfitting could only conclude that pytorch had an issue compared to TF.</p>\n<p>For example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01.  If I wanted to replicate your results then I needed to have a CV 0.02 higher than yours. </p>\n<p>That's why I say mystery is solved.  To be clear: the mystery is why you got such a high LB compared to the CV of your notebook.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978294,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-20T04:59:58.353000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977937,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-08-19T20:01:23.873000",
      "content": "<p>I shared my thoughts <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175461\" target=\"_blank\">here</a> on why I think PyTorch performs better than TF. I'll paraphase here: </p>\n<p>In short, I believe PyTorch has better ecosystem/support for augmentations. The reason for a huge gap between Public LB and CV on Chris' notebook was likely due to luck. When I added <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175372\" target=\"_blank\">heavy augmentations</a> to Chris' triple stratified notebook, I increased my CV from 0.91 to 0.9379. The public LB and private LB were 0.9394 and 0.9325 respectively. The public LB and CV gap was a pretty big indicator that something wasn't adding up.</p>\n<p>I don't think there's a comparable library for TF like <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/66643\" target=\"_blank\">Albumenations</a> for PyTorch. I refer to Albumentations because that's what first place used. Second place used RandAugment. But the idea is still the same - better support for augmentation library.</p>",
      "votes": 15,
      "replies": [
        {
          "id": 979603,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-21T01:45:23.080000",
          "content": "<p><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> Would you say that heavy augmentations was the key to improve the CV?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979658,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-08-21T03:26:10.663000",
          "content": "<p>for this competition, single model - yes. in general, better augmentations should lead to better cv score. when you read people's write ups, you will notice mention to specific augmentations used.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 980692,
          "author_name": "Sudhanshu Raheja",
          "author_url": "",
          "post_date": "2020-08-21T19:23:48.450000",
          "content": "<p>Tim, it's my first competition and I'm still discovering things. Albumentations looks fantastic, and on my simple benchmark, it was about 3x faster. Their documentation says that you can use it with TF too, more here - <a href=\"https://albumentations.ai/docs/api_reference/augmentations/transforms/\" target=\"_blank\">https://albumentations.ai/docs/api_reference/augmentations/transforms/</a></p>\n<p>Is there a reason why you think this library only works with Pytorch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981554,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-08-22T14:33:47.280000",
          "content": "<p><a href=\"https://www.kaggle.com/sudhanshuraheja\" target=\"_blank\">@sudhanshuraheja</a> Regarding external library usage with TPU - see this <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132066\" target=\"_blank\">discussion</a>. Maybe this will change in the future. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 981568,
          "author_name": "Krisztián Fekete",
          "author_url": "",
          "post_date": "2020-08-22T14:51:27.040000",
          "content": "<p>Numpy/opencv preprocessing functions, works with the tf.data.dataset as well, but only on GPU so far, there is still a bug for TPU, see: <a href=\"https://github.com/tensorflow/tensorflow/issues/30818\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/30818</a></p>\n<pre><code>    def applyAug(x):\n        x = x.numpy()\n        return aug(image=x)\n\n    def _preprocess_for_train(filename, label):\n        image_bytes = tf.io.read_file(filename)\n        image = tf.cond(\n          tf.image.is_jpeg(image_bytes),\n          lambda: tf.image.decode_jpeg(image_bytes, channels=3),\n          lambda: tf.image.decode_png(image_bytes, channels=3))\n\n\n        image = tf.py_function(func=applyAug, inp=[image], Tout=tf.float32)\n        image.set_shape(tf.TensorShape([None, None, None]))  \n\n        image = tf.image.resize(image, im_shape)\n        image = tf.cast(image, dtype=tf.float32) / 255.0\n        label = tf.one_hot(label, num_classes)\n        label = tf.cast(label, tf.float32)\n        return image, label\n\n    dataset = tf.data.Dataset.from_tensor_slices((filenames, labels))\n    autotune = tf.data.experimental.AUTOTUNE\n    dataset = dataset.map(_preprocess_for_train, num_parallel_calls=autotune)\n    dataset = dataset.batch(batch_size, drop_remainder=True) \n</code></pre>\n<p>Keep in mind, using this with GPU will slow down the training a little bit, I can only assume even if they fix it for TPU, it will cause horrible throttling for TPU.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 978477,
      "author_name": "Krisztián Fekete",
      "author_url": "",
      "post_date": "2020-08-20T07:47:54",
      "content": "<p>There is no difference between the accuracy of the PyTorch vs TensorFlow. A few months ago I had to switch to TensorFlow and replicate my PyTorch results with it. At first, I was baffled how can TF be worse, then I figured out, the difference was my lack of skills to implement the same training process. After dialling it in, it was the exact same flow, and I got almost the same results (+- 0.0001-0.0002 F1) as in PyTorch (non-kaggle dataset).</p>\n<p>Everybody is saying that PyTorch is better for research, I believe this is only true if you don't know the framework inside and out. After 2 months with TF, I feel the ceiling is so much higher. The transition was very hard especially because I had to learn both tf 1.15 and tf 2.2&gt; at the same time, but it was worth it, I'm not looking back.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 978585,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-20T09:14:37.337000",
          "content": "<p>Great reply. Just to add something, I would include the following article here about <a href=\"https://towardsdatascience.com/pytorch-vs-tensorflow-in-2020-fe237862fae1\" target=\"_blank\">Pytorch vs Tensorflow in 2020</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F6afbcc31ee52007e844bdc6d4720f416%2F1.png?generation=1597914771146704&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980699,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-21T19:30:13.543000",
          "content": "<p>Except they are not \"exactly the same\" at all.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981448,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-22T12:56:22.283000",
          "content": "<p>Nope, they're now. 😀</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981466,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-22T13:15:32.130000",
          "content": "<p>Give Pytorch a try and we'll discuss again.  On paper TF2x looks close to Pytorch indeed, but there is a difference between the look and the reality. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981520,
          "author_name": "Krisztián Fekete",
          "author_url": "",
          "post_date": "2020-08-22T14:00:22.670000",
          "content": "<p>Obviously the coding workflow is not the same, not talking about that. I'm talking about the results, the math behind the adam optimizer etc. <br>\nKeep in mind, most of the TF notebooks used TPU s with much higher batch sizes, compared to the pytorch ones with GPUs. Batch size does affect the accuracy, but as I said if you compare the exact training flow, you should have very-very close results.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 981570,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2020-08-22T14:54:31.743000",
          "content": "<p>Suppose if there is a difference in result between 2 same training settings, which factors cause it? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981624,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-22T15:37:08.710000",
          "content": "<p>The reality is,  it takes much less time and headache to code advanced workflow on Pytorch than tensorflow.  Otherwise TF 2.x wouldn't try to be as close as possible to Pytorch. </p>\n<p>OpenAI (now ClosedAI) has switched form TF to Pytorch for this reason. </p>\n<p>I'm saying this while I don't like Facebook at all. And Google has contributed much more on DL than them. </p>\n<p>P.S. : I was much more familiar with TF (both 1.X and 2.X ) than Pytorch 9 months ago. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981647,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-22T15:50:51.707000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> san, I don't find any offensive words in your response, don't get it why it's been downvoted! </p>\n<p>However, I should say, we don't have exact same comprehensive working experience in both TF and PyTorch. Compare to TF, we've very little experience in PyTorch (until now). PyTorch is awesome, no doubt and the coding experience in TF was much pain. My point only was, the coding experience between these two is almost similar now (AFAIK). <a href=\"https://www.kaggle.com/doncalculator\" target=\"_blank\">@doncalculator</a> san points some great stuff. But I would request you to explain these two more in detail (from an experimental perspective, with the latest TF). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 982077,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-23T03:41:07.603000",
          "content": "<p>Is there a way to do something like Funcional API of Keras with PyTorch? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 982211,
          "author_name": "Krisztián Fekete",
          "author_url": "",
          "post_date": "2020-08-23T06:46:12.170000",
          "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Alright, I created identical training flow for both tensorflow and pytorch, I will post the notebooks today I think. Tbh, I'm not sure about the different Efficientnet implementations qubvel's vs lukemelas, but that's gonna be the only difference. Will see.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 982404,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-23T10:50:04.227000",
          "content": "<p>Awesome. 👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 982665,
          "author_name": "Krisztián Fekete",
          "author_url": "",
          "post_date": "2020-08-23T15:07:10.340000",
          "content": "<p>I published the two notebooks:<br>\n<a href=\"https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-1-tensorflow\" target=\"_blank\">https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-1-tensorflow</a><br>\n<a href=\"https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-2-pytorch\" target=\"_blank\">https://www.kaggle.com/doncalculator/tensorflow-vs-pytorch-part-2-pytorch</a>    </p>\n<p>Here are the results:<br>\n-----------CV,--------------------PRIVATE, PUBLIC:<br>\nPytorch:    0.9820760704110646, 0.9077, 0.9196<br>\nTensorflow: 0.982470040003284   0.9043  0.9203</p>\n<p>Because of the low amount of positive images, I chose to include the external <br>\nmaligns into the validation set to have a more reliable CV score.</p>\n<p>Although these scores are close enough to support my statement, due to the low amount of<br>\nmalignant images, there is a high variance between runs, so if we want to prove that<br>\nthe results are there same across the two frameworks, this experiment should be rerun 5-10 times.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 982707,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-23T15:42:18.617000",
          "content": "<blockquote>\n  <p>There is no difference between the accuracy of the PyTorch vs TensorFlow</p>\n</blockquote>\n<p>I agree with that statement in general and I assumed it was true here too.  It is why there was a mystery:  no one could reproduce Chris high LB score when using Pytorch.  Now that we know that the high public Lb score was an overfit.</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> san, I don't find any offensive words in your response, don't get it why it's been downvoted! </p>\n</blockquote>\n<p>I didn't even notice, LOL.  Please meet my serial downvoter ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 977962,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-19T20:20:10",
      "content": "<p>The sub with the high score was definitely lucky. I reran it two times, once it was 0.01 worse than the LB score of the NB, second time I removed the 2018 data and it was higher than my first run, but slightly lower than the LB run.</p>\n<p>I think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.</p>\n<p>I still thought for the longest time there is something in that NB that I need to figure out, and it cost me way too many brain cells because I have such a hard time understanding all the tensorflow augmentation code 😁</p>\n<p>I just checked the results:</p>\n<ul>\n<li>No extra data on that kernel: 0.9428 public LB, 0.9104 private LB</li>\n<li>Extra data as provided: 0.9342 public LB, 0.9256 private LB</li>\n</ul>",
      "votes": 7,
      "replies": [
        {
          "id": 977973,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-19T20:34:18.980000",
          "content": "<p>This is interesting, I also think that early stopping is not the best way to do experiments, so your suggestions would be to run a few times and see at which epoch the model converges (let's say epoch 15), then write a schedule to end in 15 epochs to compare models even while changing parameters?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 977977,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-19T20:36:18.703000",
          "content": "<p>Yes exactly, if your best epoch is let's say between 13-15 out of 15 epochs it can be fine, but in this kernel it was sometimes 3 or 4 out of 10 and in some other folds 10, which is definitely way more random.</p>\n<p>The only way imho is to take fixed epochs and multiple bags. We always fitted for 3 bags 5-fold cv with fixed epochs. Correlation was OK, I have seen worse, albeit not perfect.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 977978,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-19T20:38:20.123000",
          "content": "<blockquote>\n  <p>I think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.</p>\n</blockquote>\n<p>I think multiple checkpoints averaging is the only way to stabilize CV.   I did it in all my experiments and  stopped worrying about public LB. Given it was not really always correlated with my CV</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 977980,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-19T20:40:21.813000",
          "content": "<p>Agreed, I can definitely see the issue, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978000,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-19T20:59:06.647000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> There is a difference between choosing the same epochs for checkpoint ensembling (can make sense) vs. picking random epochs (makes it random).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978421,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-20T07:00:26.460000",
          "content": "<p>What I did was to run relatively large number of epochs and chose 3 or 4  to average based on multiple criteria.  For instance for some folds I could chose Epochs : 12, 25, 30 while on others:  Epochs 31, 34, 40 etc.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 980076,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-08-21T10:16:20.193000",
      "content": "<p>This comp has a very high randomness in score. Unusually high for an NN competition. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 980095,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-21T10:25:00.800000",
          "content": "<p>Indeed, probably because of the very low number of positive samples.  The metric is very sensitive to one of them being misclassified.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 979737,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2020-08-21T05:11:46.040000",
      "content": "<p>My submitted custom EfficientNet-B6 384x384 noisy student (based on Chris popular kernel) + effective Focal Loss parameter computation has CV: 0.942, Public LB: 0.9446, and Private LB: 0.943 though it was hard to recreate the result may be because of reproducibility concern in TPU TensorFlow. The gap is between 0.004 - 0.007</p>",
      "votes": 3,
      "replies": [
        {
          "id": 979768,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-21T05:35:54.083000",
          "content": "<p>So, looks like the difference comes from using effective loss function! Is it? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979798,
          "author_name": "FGPC",
          "author_url": "",
          "post_date": "2020-08-21T05:58:22.203000",
          "content": "<p>My model architecture consists of the following: 1. Use base model output as input to convolutional block attention modules (CBAM), 2. Create 2 outputs and average it: 2a. use CBAM output as input to Global Average Pooling (GAP) and create a final dense layer with sigmoid activation, 2b. use CBAM output as input to Attention Weighted Average Pooling and create a final dense layer with sigmoid activation. Some of my models have large gaps in CV and LB (0.008, 0.012) given I used that said loss function. Maybe because of the random generated processed output of attention modules.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979826,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-21T06:22:14.107000",
          "content": "<p>Interesting. I've posted my solution <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175721\" target=\"_blank\">here</a>. May I ask to know is the Attention Weighted Avg Pooling that you've used are the same as mine. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979827,
          "author_name": "FGPC",
          "author_url": "",
          "post_date": "2020-08-21T06:22:19.397000",
          "content": "<p>Loss Function: </p>\n<pre><code>pos_weight = ((1 / total number of positives per fold)*total number of samples per fold) / 2\nneg_weight = ((1 / total number of negatives per fold)*total number of samples per fold) / 2\n\nBinary Focal Loss where pos_weight=pos_weight and gamma=pos_weight*neg_weight\n</code></pre>\n<p><a href=\"https://github.com/artemmavrin/focal-loss\" target=\"_blank\">https://github.com/artemmavrin/focal-loss</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979828,
          "author_name": "FGPC",
          "author_url": "",
          "post_date": "2020-08-21T06:24:15.127000",
          "content": "<p>I refer to this link for Attention Weighted Average: <a href=\"https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\" target=\"_blank\">https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979856,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-21T06:51:17.617000",
          "content": "<p>oh, I see. I've used the Guanshuo Xu modified version of it. As far as I can remember he also used the same repo. -)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977861,
      "author_name": "Mo Fahad",
      "author_url": "",
      "post_date": "2020-08-19T18:51:54.510000",
      "content": "<p>Chris Deotte needs to share notebooks with a Disclaimer 😝</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 978400,
      "author_name": "Sirish Somanchi",
      "author_url": "",
      "post_date": "2020-08-20T06:41:21.240000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I would like to learn from a Grandmaster such as yourself about creating a stable and LB correlated CV.</p>\n<p><code>For example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01.</code></p>\n<p>Even Chris's Triple Stratified Leak-Free KFold CV still has a large gap, but <strong>your CV has gap of less than 0.01</strong>, so could you please share your techniques on this accomplishment? Do you have different techniques for different types of competitions?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 978521,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-20T08:25:02.207000",
          "content": "<p>Thanks for asking.</p>\n<p>You can read my previous writeups as I will not create one here.  Here is some info still.</p>\n<p>I always focus on getting a reliable CV.  I used the same approach here as in Tweet Sentiment: bagging several runs to average out CV variance. </p>\n<p>I was surprised to get a small gap.  Having a large CV LB gap isn't an issue in itself if CV and LB are correlated.  Here, my CV-LB difference was 0 += 0.002.</p>\n<p>When I manage to get a reliable CV in a competition, which is not always the case, then I keep a change only if it improves both CV and public LB.  </p>\n<p>In this competition I had few outliers when it comes to CV LB gap.  A first one was when I added 2019 positive examples only.  CV improved and LB sunk.  Second outlier was last day when ensembling.  I overfit to public LB.  I wish I had more time to fix it.</p>\n<p>TL;DR  I use the same overall approach across competitions: reduce CV variance with bagging if need be.  It means runs take longer as they need to be repeated, but progress is steady.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 984826,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-25T10:15:40.167000",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> Is the above what you asked for?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "977848": "Many were puzzled that TF models with a rather low CV had such great public LB.  We now know that on private LB Pytorch models were competitive.\n\nSo, what's the trick here?  It seems that Chris Deotte baseline was overfitting public LB, see for instance the scores of his most popular notebook: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F059179f975ed612d17b23ec901e5b2b6%2FScreenshot_2020-08-19%20Triple%20Stratified%20KFold%20with%20TFRecords.png?generation=1597862437376532&alt=media)\n\nThe relatively low CV was not the anomaly.  The score was the anomaly.\n",
    "977872": "I'm not sure if that solves the mystery. Here are the two highest scoring PyTorch notebooks. Both TF and PyTorch notebooks have Private LB score that is 0.02 lower than Public LB score. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Faf2f8629b5ce574f5f02af2e1d082a71%2Fc1.png?generation=1597863858079347&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F3f566653c07a21e85da1ebbf07bc71e7%2Fc2.png?generation=1597863868755546&alt=media)",
    "977937": "I shared my thoughts [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175461) on why I think PyTorch performs better than TF. I'll paraphase here: \n\nIn short, I believe PyTorch has better ecosystem/support for augmentations. The reason for a huge gap between Public LB and CV on Chris' notebook was likely due to luck. When I added [heavy augmentations](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175372) to Chris' triple stratified notebook, I increased my CV from 0.91 to 0.9379. The public LB and private LB were 0.9394 and 0.9325 respectively. The public LB and CV gap was a pretty big indicator that something wasn't adding up.\n\nI don't think there's a comparable library for TF like [Albumenations](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/66643) for PyTorch. I refer to Albumentations because that's what first place used. Second place used RandAugment. But the idea is still the same - better support for augmentation library.",
    "978477": "There is no difference between the accuracy of the PyTorch vs TensorFlow. A few months ago I had to switch to TensorFlow and replicate my PyTorch results with it. At first, I was baffled how can TF be worse, then I figured out, the difference was my lack of skills to implement the same training process. After dialling it in, it was the exact same flow, and I got almost the same results (+- 0.0001-0.0002 F1) as in PyTorch (non-kaggle dataset).\n\nEverybody is saying that PyTorch is better for research, I believe this is only true if you don't know the framework inside and out. After 2 months with TF, I feel the ceiling is so much higher. The transition was very hard especially because I had to learn both tf 1.15 and tf 2.2> at the same time, but it was worth it, I'm not looking back.",
    "977962": "The sub with the high score was definitely lucky. I reran it two times, once it was 0.01 worse than the LB score of the NB, second time I removed the 2018 data and it was higher than my first run, but slightly lower than the LB run.\n\nI think early stopping makes it super random, it was picking different epochs based on choosing best val loss, which I personally dont believe to be a good idea here if you dont sharpen your learning routine.\n\nI still thought for the longest time there is something in that NB that I need to figure out, and it cost me way too many brain cells because I have such a hard time understanding all the tensorflow augmentation code 😁\n\nI just checked the results:\n\n- No extra data on that kernel: 0.9428 public LB, 0.9104 private LB\n- Extra data as provided: 0.9342 public LB, 0.9256 private LB",
    "980076": "This comp has a very high randomness in score. Unusually high for an NN competition. ",
    "979737": "My submitted custom EfficientNet-B6 384x384 noisy student (based on Chris popular kernel) + effective Focal Loss parameter computation has CV: 0.942, Public LB: 0.9446, and Private LB: 0.943 though it was hard to recreate the result may be because of reproducibility concern in TPU TensorFlow. The gap is between 0.004 - 0.007",
    "977861": "Chris Deotte needs to share notebooks with a Disclaimer 😝",
    "978400": "@cpmpml I would like to learn from a Grandmaster such as yourself about creating a stable and LB correlated CV.\n\n`For example, my CV - public gap is almost 0 across all my subs, and my cv - private LB is stable at 0.01. `\n\nEven Chris's Triple Stratified Leak-Free KFold CV still has a large gap, but **your CV has gap of less than 0.01**, so could you please share your techniques on this accomplishment? Do you have different techniques for different types of competitions?"
  }
}