{
  "id": 131213,
  "title": "what's the correct place to call model.compile() ",
  "url": "/competitions/flower-classification-with-tpus/discussion/131213",
  "author_name": "",
  "post_date": "2020-02-18T21:37:02.198940100Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>(I hope <a href=\"/mgornergoogle\">@mgornergoogle</a> can answer this question, it's kinda important, not only for me.)</p>\n\n<p>In <a href=\"/mgornergoogle\">@mgornergoogle</a> starter kernel, we have </p>\n\n<pre><code>with strategy.scope():\n    model = ....\n\nmodel.compile(...)\nhistory = model.fit(....)\n</code></pre>\n\n<p>So <code>compile</code> and <code>fit</code> is not in  <code>strategy.scope</code> (??)</p>\n\n<p>In the doc provided in competition <code>Overview</code> tab <a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a>, <code>compile</code> is inside strategy.scope, but not <code>fit</code>. This is the same in the a tf doc <a href=\"https://www.tensorflow.org/guide/distributed_training\">Distributed training with TensorFlow</a>. In the last link, it says also:</p>\n\n<pre><code>Creating a model inside this scope allows us to create mirrored variables instead of regular variables. Compiling under the scope allows us to know that the user intends to train this model using this strategy. Once this is set up, you can fit your model like you would normally. MirroredStrategy takes care of replicating the model's training on the available GPUs, aggregating gradients, and more.\n</code></pre>\n\n<p>Furthermore, In this notebook <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">five-flowers-with-keras-and-xception-on-tpu</a> created by <a href=\"/mgornergoogle\">@mgornergoogle</a> , the <code>compile</code> is inside a method called <code>create_model</code>, which is called inside  strategy.scope().</p>\n\n<p>So <strong>question 1</strong>, should we <code>model.compile()</code> inside <code>strategy.scope</code>??</p>\n\n<p>I did try to compile inside strategy.scope, no error, but I observed the validation accuracy is always a bit less than compile outside strategy.scope. For example, with <code>Xception</code> model, I can get 91% val acc easily if compile outside scope. But with compiling inside scope, it's almost &lt; 89.5.\nI tried with different models and different runs, and it always like this. (I hope some of you can test it on Kaggle with TPU. I have no quota, and I tried with GCP + TPU).</p>\n\n<p>So <strong>question 2</strong>,  why the accuracy is always a bit lower when model.compile() outside scope.</p>\n\n<p>Then <strong>question 3</strong>,  suppose the correct way is to compile inside scope,  should we compile outside scope to have a better validation accuracy? Or even for leaderboard?</p>\n\n<hr>\n\n<p>Remark: Actually, I observed my subclassed model and custom training loop lead to lower validation accuracy than other public kernels, so I spent a lot of time to do experiment. And I finally observed this compile issue.</p>",
  "messages": [
    {
      "id": "749705",
      "postDate": "02/18/2020 21:37:02",
      "content": "<p>(I hope <a href=\"/mgornergoogle\">@mgornergoogle</a> can answer this question, it's kinda important, not only for me.)</p>\n\n<p>In <a href=\"/mgornergoogle\">@mgornergoogle</a> starter kernel, we have </p>\n\n<pre><code>with strategy.scope():\n    model = ....\n\nmodel.compile(...)\nhistory = model.fit(....)\n</code></pre>\n\n<p>So <code>compile</code> and <code>fit</code> is not in  <code>strategy.scope</code> (??)</p>\n\n<p>In the doc provided in competition <code>Overview</code> tab <a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a>, <code>compile</code> is inside strategy.scope, but not <code>fit</code>. This is the same in the a tf doc <a href=\"https://www.tensorflow.org/guide/distributed_training\">Distributed training with TensorFlow</a>. In the last link, it says also:</p>\n\n<pre><code>Creating a model inside this scope allows us to create mirrored variables instead of regular variables. Compiling under the scope allows us to know that the user intends to train this model using this strategy. Once this is set up, you can fit your model like you would normally. MirroredStrategy takes care of replicating the model's training on the available GPUs, aggregating gradients, and more.\n</code></pre>\n\n<p>Furthermore, In this notebook <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu\">five-flowers-with-keras-and-xception-on-tpu</a> created by <a href=\"/mgornergoogle\">@mgornergoogle</a> , the <code>compile</code> is inside a method called <code>create_model</code>, which is called inside  strategy.scope().</p>\n\n<p>So <strong>question 1</strong>, should we <code>model.compile()</code> inside <code>strategy.scope</code>??</p>\n\n<p>I did try to compile inside strategy.scope, no error, but I observed the validation accuracy is always a bit less than compile outside strategy.scope. For example, with <code>Xception</code> model, I can get 91% val acc easily if compile outside scope. But with compiling inside scope, it's almost &lt; 89.5.\nI tried with different models and different runs, and it always like this. (I hope some of you can test it on Kaggle with TPU. I have no quota, and I tried with GCP + TPU).</p>\n\n<p>So <strong>question 2</strong>,  why the accuracy is always a bit lower when model.compile() outside scope.</p>\n\n<p>Then <strong>question 3</strong>,  suppose the correct way is to compile inside scope,  should we compile outside scope to have a better validation accuracy? Or even for leaderboard?</p>\n\n<hr>\n\n<p>Remark: Actually, I observed my subclassed model and custom training loop lead to lower validation accuracy than other public kernels, so I spent a lot of time to do experiment. And I finally observed this compile issue.</p>",
      "rawMarkdown": "(I hope @mgornergoogle can answer this question, it's kinda important, not only for me.)\n\nIn @mgornergoogle starter kernel, we have \n\n    with strategy.scope():\n        model = ....\n\n    model.compile(...)\n    history = model.fit(....)\n\nSo `compile` and `fit` is not in  `strategy.scope` (??)\n\nIn the doc provided in competition `Overview` tab [https://www.kaggle.com/docs/tpu](https://www.kaggle.com/docs/tpu), `compile` is inside strategy.scope, but not `fit`. This is the same in the a tf doc [Distributed training with TensorFlow](https://www.tensorflow.org/guide/distributed_training). In the last link, it says also:\n\n\n    Creating a model inside this scope allows us to create mirrored variables instead of regular variables. Compiling under the scope allows us to know that the user intends to train this model using this strategy. Once this is set up, you can fit your model like you would normally. MirroredStrategy takes care of replicating the model's training on the available GPUs, aggregating gradients, and more.\n\nFurthermore, In this notebook [five-flowers-with-keras-and-xception-on-tpu](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu) created by @mgornergoogle , the `compile` is inside a method called `create_model`, which is called inside  strategy.scope().\n\nSo **question 1**, should we `model.compile()` inside `strategy.scope`??\n\nI did try to compile inside strategy.scope, no error, but I observed the validation accuracy is always a bit less than compile outside strategy.scope. For example, with `Xception` model, I can get 91% val acc easily if compile outside scope. But with compiling inside scope, it's almost &lt; 89.5.\nI tried with different models and different runs, and it always like this. (I hope some of you can test it on Kaggle with TPU. I have no quota, and I tried with GCP + TPU).\n\nSo **question 2**,  why the accuracy is always a bit lower when model.compile() outside scope.\n\nThen **question 3**,  suppose the correct way is to compile inside scope,  should we compile outside scope to have a better validation accuracy? Or even for leaderboard?\n\n-----\nRemark: Actually, I observed my subclassed model and custom training loop lead to lower validation accuracy than other public kernels, so I spent a lot of time to do experiment. And I finally observed this compile issue.",
      "votes": null
    },
    {
      "id": "749848",
      "postDate": "02/18/2020 23:59:12",
      "content": "<p>The rule is that anything that creates variables should be instantiated in the scope. So for example all tf.keras.layers.* layer creation code must be in the scope.</p>\n\n<p>The one exception we made was that even though .compile creates variables, being just a function call, people may not realize that. So the model object remembers which strategy scope it was defined in and uses that in .compile automatically.</p>\n\n<p>.fit(), .evaluate(), .predict() do not need to be called in the strategy scope. If the model was created in the strategy scope, they will do the right thing.</p>\n\n<p>Calling .compile() in the strategy scope or outside should give exactly the same result if the model was created in the scope. If you do not see that in practice, you can file a bug (and give me the link so that I can surface it to the TF team).</p>",
      "rawMarkdown": "The rule is that anything that creates variables should be instantiated in the scope. So for example all tf.keras.layers.* layer creation code must be in the scope.\n\nThe one exception we made was that even though .compile creates variables, being just a function call, people may not realize that. So the model object remembers which strategy scope it was defined in and uses that in .compile automatically.\n\n.fit(), .evaluate(), .predict() do not need to be called in the strategy scope. If the model was created in the strategy scope, they will do the right thing.\n\nCalling .compile() in the strategy scope or outside should give exactly the same result if the model was created in the scope. If you do not see that in practice, you can file a bug (and give me the link so that I can surface it to the TF team).",
      "votes": null
    },
    {
      "id": "749853",
      "postDate": "02/19/2020 00:03:26",
      "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> . I will do a few more test and file the bug if necessary.</p>",
      "rawMarkdown": "Thanks @mgornergoogle . I will do a few more test and file the bug if necessary.",
      "votes": null
    },
    {
      "id": "751056",
      "postDate": "02/20/2020 00:50:34",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> </p>\n\n<p>I filed a bug here <a href=\"https://github.com/tensorflow/tensorflow/issues/36912\">lower accuracy when model.compile() inside strategy.scope() (TPU)</a></p>\n\n<p>I hope it could be investigated and finally be solved.</p>",
      "rawMarkdown": "mgornergoogle \n\nI filed a bug here [lower accuracy when model.compile() inside strategy.scope() (TPU)](https://github.com/tensorflow/tensorflow/issues/36912)\n\nI hope it could be investigated and finally be solved.",
      "votes": null
    },
    {
      "id": "751103",
      "postDate": "02/20/2020 02:00:14",
      "content": "<p>Thank you for filing a bug on this. This behavior is indeed unexpected.</p>",
      "rawMarkdown": "Thank you for filing a bug on this. This behavior is indeed unexpected.",
      "votes": null
    },
    {
      "id": "752287",
      "postDate": "02/20/2020 20:47:18",
      "content": "<p>Looks like the issue doesn't exist in tf-nightly. I don't know if it's fixed after I reported it or before it.</p>",
      "rawMarkdown": "Looks like the issue doesn't exist in tf-nightly. I don't know if it's fixed after I reported it or before it.",
      "votes": null
    },
    {
      "id": "753784",
      "postDate": "02/22/2020 16:41:46",
      "content": "<p>This behavior is interesting. Perhaps loss is calculated with 16 bit precision when it is inside strategy scope and 32 bit precision when it is outside. And maybe in the case of 16 bit, we are observing overflow and/or underflow from 16 bit's restricted dynamic range thus reducing our validation accuracy.</p>",
      "rawMarkdown": "This behavior is interesting. Perhaps loss is calculated with 16 bit precision when it is inside strategy scope and 32 bit precision when it is outside. And maybe in the case of 16 bit, we are observing overflow and/or underflow from 16 bit's restricted dynamic range thus reducing our validation accuracy.",
      "votes": null
    },
    {
      "id": "753833",
      "postDate": "02/22/2020 17:48:09",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , I am not sure, I filed a bug on github.</p>\n\n<p>However, if the loss is not calculated in strategy scope, I think something is fundamentally wrong. By training on TPU, we expect each replica receive a partial of batch training examples, calculate forward pass, the gradients are calculated on each replica, then synced across the replicas by summing them.</p>\n\n<p>If loss is not calculated in strategy scope, how the above process can be done and give the correct results?</p>",
      "rawMarkdown": "cdeotte , I am not sure, I filed a bug on github.\n\nHowever, if the loss is not calculated in strategy scope, I think something is fundamentally wrong. By training on TPU, we expect each replica receive a partial of batch training examples, calculate forward pass, the gradients are calculated on each replica, then synced across the replicas by summing them.\n\nIf loss is not calculated in strategy scope, how the above process can be done and give the correct results?",
      "votes": null
    },
    {
      "id": "947088",
      "postDate": "07/27/2020 04:26:56",
      "content": "<p>您好！我遇到了一些问题，我需要在kaggle中运行这样的命令\n“ cd libs.box utils.cython utils</p>\n\n<p>python setup.py安装</p>\n\n<p>python setup.py build_ext --inplace”</p>\n\n<p>但我不知道该怎么做。直接运行将报告错误。</p>",
      "rawMarkdown": "您好！我遇到了一些问题，我需要在kaggle中运行这样的命令\n“ cd libs.box utils.cython utils\n\npython setup.py安装\n\npython setup.py build_ext --inplace”\n\n但我不知道该怎么做。直接运行将报告错误。",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749848,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/18/2020 23:59:12",
      "content": "<p>The rule is that anything that creates variables should be instantiated in the scope. So for example all tf.keras.layers.* layer creation code must be in the scope.</p>\n\n<p>The one exception we made was that even though .compile creates variables, being just a function call, people may not realize that. So the model object remembers which strategy scope it was defined in and uses that in .compile automatically.</p>\n\n<p>.fit(), .evaluate(), .predict() do not need to be called in the strategy scope. If the model was created in the strategy scope, they will do the right thing.</p>\n\n<p>Calling .compile() in the strategy scope or outside should give exactly the same result if the model was created in the scope. If you do not see that in practice, you can file a bug (and give me the link so that I can surface it to the TF team).</p>",
      "votes": null,
      "replies": [
        {
          "id": 749853,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/19/2020 00:03:26",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> . I will do a few more test and file the bug if necessary.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751056,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/20/2020 00:50:34",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> </p>\n\n<p>I filed a bug here <a href=\"https://github.com/tensorflow/tensorflow/issues/36912\">lower accuracy when model.compile() inside strategy.scope() (TPU)</a></p>\n\n<p>I hope it could be investigated and finally be solved.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751103,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/20/2020 02:00:14",
          "content": "<p>Thank you for filing a bug on this. This behavior is indeed unexpected.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752287,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/20/2020 20:47:18",
          "content": "<p>Looks like the issue doesn't exist in tf-nightly. I don't know if it's fixed after I reported it or before it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947088,
          "author_name": "vvvshazhihun",
          "author_url": "",
          "post_date": "07/27/2020 04:26:56",
          "content": "<p>您好！我遇到了一些问题，我需要在kaggle中运行这样的命令\n“ cd libs.box utils.cython utils</p>\n\n<p>python setup.py安装</p>\n\n<p>python setup.py build_ext --inplace”</p>\n\n<p>但我不知道该怎么做。直接运行将报告错误。</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 753784,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/22/2020 16:41:46",
      "content": "<p>This behavior is interesting. Perhaps loss is calculated with 16 bit precision when it is inside strategy scope and 32 bit precision when it is outside. And maybe in the case of 16 bit, we are observing overflow and/or underflow from 16 bit's restricted dynamic range thus reducing our validation accuracy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 753833,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/22/2020 17:48:09",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , I am not sure, I filed a bug on github.</p>\n\n<p>However, if the loss is not calculated in strategy scope, I think something is fundamentally wrong. By training on TPU, we expect each replica receive a partial of batch training examples, calculate forward pass, the gradients are calculated on each replica, then synced across the replicas by summing them.</p>\n\n<p>If loss is not calculated in strategy scope, how the above process can be done and give the correct results?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "749705": "(I hope @mgornergoogle can answer this question, it's kinda important, not only for me.)\n\nIn @mgornergoogle starter kernel, we have \n\n    with strategy.scope():\n        model = ....\n\n    model.compile(...)\n    history = model.fit(....)\n\nSo `compile` and `fit` is not in  `strategy.scope` (??)\n\nIn the doc provided in competition `Overview` tab [https://www.kaggle.com/docs/tpu](https://www.kaggle.com/docs/tpu), `compile` is inside strategy.scope, but not `fit`. This is the same in the a tf doc [Distributed training with TensorFlow](https://www.tensorflow.org/guide/distributed_training). In the last link, it says also:\n\n\n    Creating a model inside this scope allows us to create mirrored variables instead of regular variables. Compiling under the scope allows us to know that the user intends to train this model using this strategy. Once this is set up, you can fit your model like you would normally. MirroredStrategy takes care of replicating the model's training on the available GPUs, aggregating gradients, and more.\n\nFurthermore, In this notebook [five-flowers-with-keras-and-xception-on-tpu](https://www.kaggle.com/mgornergoogle/five-flowers-with-keras-and-xception-on-tpu) created by @mgornergoogle , the `compile` is inside a method called `create_model`, which is called inside  strategy.scope().\n\nSo **question 1**, should we `model.compile()` inside `strategy.scope`??\n\nI did try to compile inside strategy.scope, no error, but I observed the validation accuracy is always a bit less than compile outside strategy.scope. For example, with `Xception` model, I can get 91% val acc easily if compile outside scope. But with compiling inside scope, it's almost &lt; 89.5.\nI tried with different models and different runs, and it always like this. (I hope some of you can test it on Kaggle with TPU. I have no quota, and I tried with GCP + TPU).\n\nSo **question 2**,  why the accuracy is always a bit lower when model.compile() outside scope.\n\nThen **question 3**,  suppose the correct way is to compile inside scope,  should we compile outside scope to have a better validation accuracy? Or even for leaderboard?\n\n-----\nRemark: Actually, I observed my subclassed model and custom training loop lead to lower validation accuracy than other public kernels, so I spent a lot of time to do experiment. And I finally observed this compile issue.",
    "749848": "The rule is that anything that creates variables should be instantiated in the scope. So for example all tf.keras.layers.* layer creation code must be in the scope.\n\nThe one exception we made was that even though .compile creates variables, being just a function call, people may not realize that. So the model object remembers which strategy scope it was defined in and uses that in .compile automatically.\n\n.fit(), .evaluate(), .predict() do not need to be called in the strategy scope. If the model was created in the strategy scope, they will do the right thing.\n\nCalling .compile() in the strategy scope or outside should give exactly the same result if the model was created in the scope. If you do not see that in practice, you can file a bug (and give me the link so that I can surface it to the TF team).",
    "749853": "Thanks @mgornergoogle . I will do a few more test and file the bug if necessary.",
    "751056": "mgornergoogle \n\nI filed a bug here [lower accuracy when model.compile() inside strategy.scope() (TPU)](https://github.com/tensorflow/tensorflow/issues/36912)\n\nI hope it could be investigated and finally be solved.",
    "751103": "Thank you for filing a bug on this. This behavior is indeed unexpected.",
    "752287": "Looks like the issue doesn't exist in tf-nightly. I don't know if it's fixed after I reported it or before it.",
    "753784": "This behavior is interesting. Perhaps loss is calculated with 16 bit precision when it is inside strategy scope and 32 bit precision when it is outside. And maybe in the case of 16 bit, we are observing overflow and/or underflow from 16 bit's restricted dynamic range thus reducing our validation accuracy.",
    "753833": "cdeotte , I am not sure, I filed a bug on github.\n\nHowever, if the loss is not calculated in strategy scope, I think something is fundamentally wrong. By training on TPU, we expect each replica receive a partial of batch training examples, calculate forward pass, the gradients are calculated on each replica, then synced across the replicas by summing them.\n\nIf loss is not calculated in strategy scope, how the above process can be done and give the correct results?",
    "947088": "您好！我遇到了一些问题，我需要在kaggle中运行这样的命令\n“ cd libs.box utils.cython utils\n\npython setup.py安装\n\npython setup.py build_ext --inplace”\n\n但我不知道该怎么做。直接运行将报告错误。"
  },
  "source": "meta"
}