{
  "id": 206663,
  "title": "gap between predictions from GPU / TPU",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206663",
  "author_name": "",
  "post_date": "2020-12-25T19:50:34.561716700Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>(In particular, for anyone using TPU for this competition)</p>\n<p>I have a TensorFlow TPU training / validation pipeline for this competition <a href=\"https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students\" target=\"_blank\">TPU - Track knowledge states of 1M+ students</a>.</p>\n<p>As I mentioned in an thread <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206488\" target=\"_blank\">Does the gap between CV and public LB score need to be very small?</a>, the gap between my CV and LB is large, ranging from <code>0.007</code> to <code>0.01</code>. Therefore I spent some time to look deeper. Nothing wrong being found, but there is an observation.</p>\n<p>I actually have 2 validation pipelines:</p>\n<ul>\n<li><p>one is using validation dataset stored in tensorflow record file, and with this, the validation is done with TPU.</p></li>\n<li><p>another one is using a pipeline similar to the submission process, so small batches are fed into the model running on GPU.</p></li>\n</ul>\n<p>I observed  that the prediction from TPU / GPU have some differences.</p>\n<p>For example, with TPU, one prediction is like</p>\n<pre><code>```\n&lt;tf.Tensor: shape=(128,), dtype=float32, numpy=\narray([0.01927836, 0.01927836, 0.01927836, 0.01927836, 0.01927836,\n        ......\n       0.3612144 , 0.9305125 , 0.44552624], dtype=float32)&gt;\n```\n</code></pre>\n<p>while with GPU, it gives</p>\n<pre><code>&lt;tf.Tensor: shape=(128,), dtype=float32, numpy=\narray([0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 ,\n        ......\n       0.36048502, 0.93047607, 0.46845183], dtype=float32)&gt;\n```\n</code></pre>\n<p>For some validation examples, the gaps between the probabilities from these 2 pipeline can be &gt; <code>0.01</code>, but the final AUC score doesn't change much.</p>\n<p>Anyway, this situation is not the cause of the gap between my CV / LB, however I hope it won't give me a bad surprise for the private LB …</p>",
  "messages": [
    {
      "id": "1126674",
      "postDate": "12/25/2020 19:50:34",
      "content": "<p>(In particular, for anyone using TPU for this competition)</p>\n<p>I have a TensorFlow TPU training / validation pipeline for this competition <a href=\"https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students\" target=\"_blank\">TPU - Track knowledge states of 1M+ students</a>.</p>\n<p>As I mentioned in an thread <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206488\" target=\"_blank\">Does the gap between CV and public LB score need to be very small?</a>, the gap between my CV and LB is large, ranging from <code>0.007</code> to <code>0.01</code>. Therefore I spent some time to look deeper. Nothing wrong being found, but there is an observation.</p>\n<p>I actually have 2 validation pipelines:</p>\n<ul>\n<li><p>one is using validation dataset stored in tensorflow record file, and with this, the validation is done with TPU.</p></li>\n<li><p>another one is using a pipeline similar to the submission process, so small batches are fed into the model running on GPU.</p></li>\n</ul>\n<p>I observed  that the prediction from TPU / GPU have some differences.</p>\n<p>For example, with TPU, one prediction is like</p>\n<pre><code>```\n&lt;tf.Tensor: shape=(128,), dtype=float32, numpy=\narray([0.01927836, 0.01927836, 0.01927836, 0.01927836, 0.01927836,\n        ......\n       0.3612144 , 0.9305125 , 0.44552624], dtype=float32)&gt;\n```\n</code></pre>\n<p>while with GPU, it gives</p>\n<pre><code>&lt;tf.Tensor: shape=(128,), dtype=float32, numpy=\narray([0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 ,\n        ......\n       0.36048502, 0.93047607, 0.46845183], dtype=float32)&gt;\n```\n</code></pre>\n<p>For some validation examples, the gaps between the probabilities from these 2 pipeline can be &gt; <code>0.01</code>, but the final AUC score doesn't change much.</p>\n<p>Anyway, this situation is not the cause of the gap between my CV / LB, however I hope it won't give me a bad surprise for the private LB …</p>",
      "rawMarkdown": "(In particular, for anyone using TPU for this competition)\n\nI have a TensorFlow TPU training / validation pipeline for this competition [TPU - Track knowledge states of 1M+ students](https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students).\n\nAs I mentioned in an thread [Does the gap between CV and public LB score need to be very small?](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206488), the gap between my CV and LB is large, ranging from `0.007` to `0.01`. Therefore I spent some time to look deeper. Nothing wrong being found, but there is an observation.\n\nI actually have 2 validation pipelines:\n\n  - one is using validation dataset stored in tensorflow record file, and with this, the validation is done with TPU.\n\n  - another one is using a pipeline similar to the submission process, so small batches are fed into the model running on GPU.\n\nI observed  that the prediction from TPU / GPU have some differences.\n\nFor example, with TPU, one prediction is like\n\n    ```\n    <tf.Tensor: shape=(128,), dtype=float32, numpy=\n    array([0.01927836, 0.01927836, 0.01927836, 0.01927836, 0.01927836,\n            ......\n           0.3612144 , 0.9305125 , 0.44552624], dtype=float32)>\n    ```\n\nwhile with GPU, it gives\n\n    <tf.Tensor: shape=(128,), dtype=float32, numpy=\n    array([0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 ,\n            ......\n           0.36048502, 0.93047607, 0.46845183], dtype=float32)>\n    ```\n\nFor some validation examples, the gaps between the probabilities from these 2 pipeline can be > `0.01`, but the final AUC score doesn't change much.\n\nAnyway, this situation is not the cause of the gap between my CV / LB, however I hope it won't give me a bad surprise for the private LB ...",
      "votes": null
    },
    {
      "id": "1126933",
      "postDate": "12/26/2020 05:16:23",
      "content": "<p>Well since computations are done using bfloat16 for TPUs, there would be small differences in both model outputs even if the weights of the models are same. (Weights are still saved as float32 in both cases)</p>\n<p>I think this slight differences would become smaller and smaller as we increase the number of interactions to check the AUC on.</p>",
      "rawMarkdown": "Well since computations are done using bfloat16 for TPUs, there would be small differences in both model outputs even if the weights of the models are same. (Weights are still saved as float32 in both cases)\n\nI think this slight differences would become smaller and smaller as we increase the number of interactions to check the AUC on.",
      "votes": null
    },
    {
      "id": "1127106",
      "postDate": "12/26/2020 08:43:42",
      "content": "<p>Yes, I am aware of <code>bfloat16 for TPUs</code>, but didn't know that the gap (between probabilities) could be higher &gt; 0.005 …</p>",
      "rawMarkdown": "Yes, I am aware of `bfloat16 for TPUs`, but didn't know that the gap (between probabilities) could be higher > 0.005 ...",
      "votes": null
    },
    {
      "id": "1127108",
      "postDate": "12/26/2020 08:47:33",
      "content": "<p>Are you using TPU for submission as well? I'm training on TPU but submitting on GPU since my inference pipeline without model is ~1 hour and built on np arrays so can't shift to TPU now.</p>\n<p>Btw if you have pipeline setup for TPU inference. Submit using GPU and TPU both on the same trained model and check which one gives higher Public LB. </p>\n<p>Do let me know the result of this experiment :)</p>",
      "rawMarkdown": "Are you using TPU for submission as well? I'm training on TPU but submitting on GPU since my inference pipeline without model is ~1 hour and built on np arrays so can't shift to TPU now.\n\nBtw if you have pipeline setup for TPU inference. Submit using GPU and TPU both on the same trained model and check which one gives higher Public LB. \n\nDo let me know the result of this experiment :)",
      "votes": null
    },
    {
      "id": "1127111",
      "postDate": "12/26/2020 08:52:55",
      "content": "<p>TPU for fast validation, not for submission. For submission, only GPU - can't use internet connection.<br>\nI took extra effort to setup TPU for validation, because it is way too much faster (to see the CV) than using GPU.</p>",
      "rawMarkdown": "TPU for fast validation, not for submission. For submission, only GPU - can't use internet connection.\nI took extra effort to setup TPU for validation, because it is way too much faster (to see the CV) than using GPU.",
      "votes": null
    },
    {
      "id": "1127120",
      "postDate": "12/26/2020 09:02:07",
      "content": "<p>Alright. Let's hope that the results are better from GPU on LB then :D</p>",
      "rawMarkdown": "Alright. Let's hope that the results are better from GPU on LB then :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1126933,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "12/26/2020 05:16:23",
      "content": "<p>Well since computations are done using bfloat16 for TPUs, there would be small differences in both model outputs even if the weights of the models are same. (Weights are still saved as float32 in both cases)</p>\n<p>I think this slight differences would become smaller and smaller as we increase the number of interactions to check the AUC on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1127106,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/26/2020 08:43:42",
          "content": "<p>Yes, I am aware of <code>bfloat16 for TPUs</code>, but didn't know that the gap (between probabilities) could be higher &gt; 0.005 …</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127108,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/26/2020 08:47:33",
          "content": "<p>Are you using TPU for submission as well? I'm training on TPU but submitting on GPU since my inference pipeline without model is ~1 hour and built on np arrays so can't shift to TPU now.</p>\n<p>Btw if you have pipeline setup for TPU inference. Submit using GPU and TPU both on the same trained model and check which one gives higher Public LB. </p>\n<p>Do let me know the result of this experiment :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127111,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/26/2020 08:52:55",
          "content": "<p>TPU for fast validation, not for submission. For submission, only GPU - can't use internet connection.<br>\nI took extra effort to setup TPU for validation, because it is way too much faster (to see the CV) than using GPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127120,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/26/2020 09:02:07",
          "content": "<p>Alright. Let's hope that the results are better from GPU on LB then :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1126674": "(In particular, for anyone using TPU for this competition)\n\nI have a TensorFlow TPU training / validation pipeline for this competition [TPU - Track knowledge states of 1M+ students](https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students).\n\nAs I mentioned in an thread [Does the gap between CV and public LB score need to be very small?](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206488), the gap between my CV and LB is large, ranging from `0.007` to `0.01`. Therefore I spent some time to look deeper. Nothing wrong being found, but there is an observation.\n\nI actually have 2 validation pipelines:\n\n  - one is using validation dataset stored in tensorflow record file, and with this, the validation is done with TPU.\n\n  - another one is using a pipeline similar to the submission process, so small batches are fed into the model running on GPU.\n\nI observed  that the prediction from TPU / GPU have some differences.\n\nFor example, with TPU, one prediction is like\n\n    ```\n    <tf.Tensor: shape=(128,), dtype=float32, numpy=\n    array([0.01927836, 0.01927836, 0.01927836, 0.01927836, 0.01927836,\n            ......\n           0.3612144 , 0.9305125 , 0.44552624], dtype=float32)>\n    ```\n\nwhile with GPU, it gives\n\n    <tf.Tensor: shape=(128,), dtype=float32, numpy=\n    array([0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 , 0.0193305 ,\n            ......\n           0.36048502, 0.93047607, 0.46845183], dtype=float32)>\n    ```\n\nFor some validation examples, the gaps between the probabilities from these 2 pipeline can be > `0.01`, but the final AUC score doesn't change much.\n\nAnyway, this situation is not the cause of the gap between my CV / LB, however I hope it won't give me a bad surprise for the private LB ...",
    "1126933": "Well since computations are done using bfloat16 for TPUs, there would be small differences in both model outputs even if the weights of the models are same. (Weights are still saved as float32 in both cases)\n\nI think this slight differences would become smaller and smaller as we increase the number of interactions to check the AUC on.",
    "1127106": "Yes, I am aware of `bfloat16 for TPUs`, but didn't know that the gap (between probabilities) could be higher > 0.005 ...",
    "1127108": "Are you using TPU for submission as well? I'm training on TPU but submitting on GPU since my inference pipeline without model is ~1 hour and built on np arrays so can't shift to TPU now.\n\nBtw if you have pipeline setup for TPU inference. Submit using GPU and TPU both on the same trained model and check which one gives higher Public LB. \n\nDo let me know the result of this experiment :)",
    "1127111": "TPU for fast validation, not for submission. For submission, only GPU - can't use internet connection.\nI took extra effort to setup TPU for validation, because it is way too much faster (to see the CV) than using GPU.",
    "1127120": "Alright. Let's hope that the results are better from GPU on LB then :D"
  },
  "source": "meta"
}