{
  "id": 140706,
  "title": "How to make TPU get the deterministic result(tf version)",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/140706",
  "author_name": "",
  "post_date": "2020-04-03T00:55:45.001214300Z",
  "votes": 10,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I tried to run <a href=\"/xhlulu\">@xhlulu</a> 's tensorflow kernel.  But after running  the kernel many times, I always got the different or random results  for every time running.  <br> My kernel: <a href=\"https://www.kaggle.com/qinhui1999/jigsaw-tpu-xlm-roberta?scriptVersionId=31385845\">https://www.kaggle.com/qinhui1999/jigsaw-tpu-xlm-roberta?scriptVersionId=31385845</a> <br>\nIn pytorch, if I fix the random seeds, then I always get the same result for every time running.  Hence I tried to fix the random seeds like pytorch as follow:</p>\n\n<p>SEED=2020<br>\nimport random<br>\ndef set_seed(seed):<br>\n    random.seed(seed)<br>\n    np.random.seed(seed)<br>\n    torch.manual_seed(seed)<br>\n    torch.cuda.manual_seed(seed)<br>\nset_seed(SEED) <br></p>\n\n<p>To my dispointed,  it does not work. <br>\nIs it anyway to make TPU  get the deterministic result?</p>\n\n<p>--------updated 04-03----\nFor tf 2.1 version, I use the following config and get the deterministic result in GPU.  But it can not work in TPU.  </p>\n\n<p>SEED=2020<br>\nos.environ['PYTHONHASHSEED']=str(SEED)<br>\nrandom.seed(SEED)<br>\nnp.random.seed(SEED)<br>\ntf.random.set_seed(SEED)<br></p>\n\n<p>os.environ['TF_DETERMINISTIC_OPS'] = '1'<br></p>\n\n<p>---------- other reference--------\nNvidia TensorFlow Determinism ---- <a href=\"https://github.com/NVIDIA/tensorflow-determinism\">https://github.com/NVIDIA/tensorflow-determinism</a></p>",
  "messages": [
    {
      "id": "795702",
      "postDate": "04/03/2020 00:55:45",
      "content": "<p>Hi,</p>\n\n<p>I tried to run <a href=\"/xhlulu\">@xhlulu</a> 's tensorflow kernel.  But after running  the kernel many times, I always got the different or random results  for every time running.  <br> My kernel: <a href=\"https://www.kaggle.com/qinhui1999/jigsaw-tpu-xlm-roberta?scriptVersionId=31385845\">https://www.kaggle.com/qinhui1999/jigsaw-tpu-xlm-roberta?scriptVersionId=31385845</a> <br>\nIn pytorch, if I fix the random seeds, then I always get the same result for every time running.  Hence I tried to fix the random seeds like pytorch as follow:</p>\n\n<p>SEED=2020<br>\nimport random<br>\ndef set_seed(seed):<br>\n    random.seed(seed)<br>\n    np.random.seed(seed)<br>\n    torch.manual_seed(seed)<br>\n    torch.cuda.manual_seed(seed)<br>\nset_seed(SEED) <br></p>\n\n<p>To my dispointed,  it does not work. <br>\nIs it anyway to make TPU  get the deterministic result?</p>\n\n<p>--------updated 04-03----\nFor tf 2.1 version, I use the following config and get the deterministic result in GPU.  But it can not work in TPU.  </p>\n\n<p>SEED=2020<br>\nos.environ['PYTHONHASHSEED']=str(SEED)<br>\nrandom.seed(SEED)<br>\nnp.random.seed(SEED)<br>\ntf.random.set_seed(SEED)<br></p>\n\n<p>os.environ['TF_DETERMINISTIC_OPS'] = '1'<br></p>\n\n<p>---------- other reference--------\nNvidia TensorFlow Determinism ---- <a href=\"https://github.com/NVIDIA/tensorflow-determinism\">https://github.com/NVIDIA/tensorflow-determinism</a></p>",
      "rawMarkdown": "Hi,\n\n\n\n I tried to run @xhlulu 's tensorflow kernel.  But after running  the kernel many times, I always got the different or random results  for every time running.  <br> My kernel: https://www.kaggle.com/qinhui1999/jigsaw-tpu-xlm-roberta?scriptVersionId=31385845 <br>\nIn pytorch, if I fix the random seeds, then I always get the same result for every time running.  Hence I tried to fix the random seeds like pytorch as follow:\n\nSEED=2020<br>\nimport random<br>\ndef set_seed(seed):<br>\n    random.seed(seed)<br>\n    np.random.seed(seed)<br>\n    torch.manual_seed(seed)<br>\n    torch.cuda.manual_seed(seed)<br>\nset_seed(SEED) <br>\n\nTo my dispointed,  it does not work.  \nIs it anyway to make TPU  get the deterministic result?\n\n--------updated 04-03----\nFor tf 2.1 version, I use the following config and get the deterministic result in GPU.  But it can not work in TPU.  \n\nSEED=2020<br>\nos.environ['PYTHONHASHSEED']=str(SEED)<br>\nrandom.seed(SEED)<br>\nnp.random.seed(SEED)<br>\ntf.random.set_seed(SEED)<br>\n\nos.environ['TF\\_DETERMINISTIC\\_OPS'] = '1'<br>\n\n---------- other reference--------\nNvidia TensorFlow Determinism ---- https://github.com/NVIDIA/tensorflow-determinism",
      "votes": null
    },
    {
      "id": "795841",
      "postDate": "04/03/2020 04:46:02",
      "content": "<p>I thought random results is a tensorflow problem, TPU or not?</p>",
      "rawMarkdown": "I thought random results is a tensorflow problem, TPU or not?",
      "votes": null
    },
    {
      "id": "795850",
      "postDate": "04/03/2020 04:51:32",
      "content": "<p><a href=\"/suicaokhoailang\">@suicaokhoailang</a> Yeah. But I still think tensorflow can do the same thing like pytorch. </p>",
      "rawMarkdown": "suicaokhoailang Yeah. But I still think tensorflow can do the same thing like pytorch.",
      "votes": null
    },
    {
      "id": "796028",
      "postDate": "04/03/2020 08:22:26",
      "content": "<p>It's tensorflow features/bugs(it's hard to say what is it really). You can obtain the same results in two independent running only if would not be used parallelism (used 1 core of cpu only), but and this not guaranteed the same results. P.s. This <a href=\"https://www.youtube.com/watch?v=Ys8ofBeR2kA&amp;feature=youtu.be\">video </a>can help you understand why it's happening.</p>",
      "rawMarkdown": "It's tensorflow features/bugs(it's hard to say what is it really). You can obtain the same results in two independent running only if would not be used parallelism (used 1 core of cpu only), but and this not guaranteed the same results. P.s. This [video ](https://www.youtube.com/watch?v=Ys8ofBeR2kA&amp;feature=youtu.be)can help you understand why it's happening.",
      "votes": null
    },
    {
      "id": "796790",
      "postDate": "04/03/2020 23:14:49",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> </p>\n\n<p>In my custom version, something like the following helps, but not completely deterministic (results are close in different run, but still having some small difference).</p>\n\n<p>In tensorflow, there are global seed and operation seed. You have to make sure both of them are set.</p>\n\n<p><code>\n    self.dense_layer = tf.keras.layers.Dense(1, activation='sigmoid', kernel_initializer=tf.keras.initializers.GlorotUniform(seed=OP_SEED), bias_initializer='zeros')\n</code>\nalong with </p>\n\n<p>```\n    NO_RANDOM = True</p>\n\n<pre><code>GLOBAL_SEED = None\nOP_SEED = None\n\nif NO_RANDOM:\n    GLOBAL_SEED = 127\n    OP_SEED = 127\n    tf.random.set_seed(seed=GLOBAL_SEED)\n</code></pre>",
      "rawMarkdown": "qinhui1999 \n\nIn my custom version, something like the following helps, but not completely deterministic (results are close in different run, but still having some small difference).\n\nIn tensorflow, there are global seed and operation seed. You have to make sure both of them are set.\n\n\n```\n    self.dense_layer = tf.keras.layers.Dense(1, activation='sigmoid', kernel_initializer=tf.keras.initializers.GlorotUniform(seed=OP_SEED), bias_initializer='zeros')\n```\nalong with \n\n```\n    NO_RANDOM = True\n\n    GLOBAL_SEED = None\n    OP_SEED = None\n\n    if NO_RANDOM:\n        GLOBAL_SEED = 127\n        OP_SEED = 127\n        tf.random.set_seed(seed=GLOBAL_SEED)",
      "votes": null
    },
    {
      "id": "796807",
      "postDate": "04/04/2020 00:07:41",
      "content": "<p>I also saw </p>\n\n<pre><code>os.environ['TF_DETERMINISTIC_OPS'] = '1'\n</code></pre>\n\n<p>from </p>\n\n<p><a href=\"https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops\">https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops</a></p>",
      "rawMarkdown": "I also saw \n\n    os.environ['TF_DETERMINISTIC_OPS'] = '1'\n\nfrom \n\n[https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops](https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops)",
      "votes": null
    },
    {
      "id": "796852",
      "postDate": "04/04/2020 02:06:00",
      "content": "<p><a href=\"/miklgr500\">@miklgr500</a> <a href=\"/yihdarshieh\">@yihdarshieh</a>  Thanks for your help.  I tested it again, now GPU is ok ,but TPU is not ok.  I still got  random results.   : (</p>",
      "rawMarkdown": "miklgr500 @yihdarshieh  Thanks for your help.  I tested it again, now GPU is ok ,but TPU is not ok.  I still got  random results.   : (",
      "votes": null
    },
    {
      "id": "802012",
      "postDate": "04/09/2020 03:32:27",
      "content": "<p><a href=\"https://medium.com/datadriveninvestor/getting-reproducible-results-in-tensorflow-3705536aa185\">https://medium.com/datadriveninvestor/getting-reproducible-results-in-tensorflow-3705536aa185</a>, you can always check this out if this helps.</p>",
      "rawMarkdown": "https://medium.com/datadriveninvestor/getting-reproducible-results-in-tensorflow-3705536aa185, you can always check this out if this helps.",
      "votes": null
    },
    {
      "id": "805264",
      "postDate": "04/12/2020 14:13:36",
      "content": "<p>Check out this help section about Deterministic Training on TPUs: <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training\">https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training</a></p>\n\n<p>Main takeaway: Single core TPU training can be deterministic, multi-core TPU training may lead to varying results (Due to data sharding)</p>",
      "rawMarkdown": "Check out this help section about Deterministic Training on TPUs: https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training\n\nMain takeaway: Single core TPU training can be deterministic, multi-core TPU training may lead to varying results (Due to data sharding)",
      "votes": null
    },
    {
      "id": "805305",
      "postDate": "04/12/2020 14:56:17",
      "content": "<p><a href=\"/jerryqu\">@jerryqu</a>  Thanks for your help.  So we always get the random results for multi-core TPU.  : (</p>",
      "rawMarkdown": "jerryqu  Thanks for your help.  So we always get the random results for multi-core TPU.  : (",
      "votes": null
    },
    {
      "id": "805309",
      "postDate": "04/12/2020 15:01:38",
      "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a>  Thanks for your help.  But it can not work in multi-core TPU. </p>",
      "rawMarkdown": "harshitsheoran  Thanks for your help.  But it can not work in multi-core TPU.",
      "votes": null
    },
    {
      "id": "806798",
      "postDate": "04/14/2020 04:39:22",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> In my understanding, if we turn on os.environ['TF_DETERMINISTIC_OPS'] = '1', training speed is way slower. Have you experienced that?</p>",
      "rawMarkdown": "qinhui1999 In my understanding, if we turn on os.environ['TF_DETERMINISTIC_OPS'] = '1', training speed is way slower. Have you experienced that?",
      "votes": null
    },
    {
      "id": "806805",
      "postDate": "04/14/2020 04:52:33",
      "content": "<p><a href=\"/bamps53\">@bamps53</a>  No, I do not test the speed.  When we use the deterministic function, we just want to check if the new features or changes can bring some improvements in kpi. </p>",
      "rawMarkdown": "bamps53  No, I do not test the speed.  When we use the deterministic function, we just want to check if the new features or changes can bring some improvements in kpi.",
      "votes": null
    },
    {
      "id": "807257",
      "postDate": "04/14/2020 14:15:33",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> I know what you want to do, but I saw someone reported if you turn on TFDETERMINISTICOPS, it makes training speed very slower. (he said more than 1.5 times) So I gave up to do even I eagerly want to deterministic result! </p>",
      "rawMarkdown": "qinhui1999 I know what you want to do, but I saw someone reported if you turn on TFDETERMINISTICOPS, it makes training speed very slower. (he said more than 1.5 times) So I gave up to do even I eagerly want to deterministic result!",
      "votes": null
    },
    {
      "id": "808226",
      "postDate": "04/15/2020 08:20:06",
      "content": "<p>yes, this is really cruel. Did you find any solution? <a href=\"/qinhui1999\">@qinhui1999</a> </p>",
      "rawMarkdown": "yes, this is really cruel. Did you find any solution? @qinhui1999",
      "votes": null
    },
    {
      "id": "808381",
      "postDate": "04/15/2020 11:24:16",
      "content": "<p><a href=\"/shahules\">@shahules</a> No solution. You can check this page: <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training\">https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training</a> </p>",
      "rawMarkdown": "shahules No solution. You can check this page: https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training",
      "votes": null
    },
    {
      "id": "808417",
      "postDate": "04/15/2020 12:13:52",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> Thanks for your reply, I was getting good performece when Running But it decreases whenever I commit.Any hacks for this issue ?</p>",
      "rawMarkdown": "qinhui1999 Thanks for your reply, I was getting good performece when Running But it decreases whenever I commit.Any hacks for this issue ?",
      "votes": null
    },
    {
      "id": "1106434",
      "postDate": "12/08/2020 20:44:59",
      "content": "<p>While we can solve the variation in GPU setup easily by setting the random seeds… Therefore it's more seen in TPU setup as the model is trained on multi core. The sharding and loss calculation techniques leads to different results each time and also different from a version that is trained on GPU. <br>\nFollowing documentation explains the reasons behind the variations very well along with debugging methods. <br>\n<a href=\"https://cloud.google.com/tpu/docs/troubleshooting#model-accuracy\" target=\"_blank\">https://cloud.google.com/tpu/docs/troubleshooting#model-accuracy</a></p>",
      "rawMarkdown": "While we can solve the variation in GPU setup easily by setting the random seeds... Therefore it's more seen in TPU setup as the model is trained on multi core. The sharding and loss calculation techniques leads to different results each time and also different from a version that is trained on GPU. \nFollowing documentation explains the reasons behind the variations very well along with debugging methods. \nhttps://cloud.google.com/tpu/docs/troubleshooting#model-accuracy",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 795841,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "04/03/2020 04:46:02",
      "content": "<p>I thought random results is a tensorflow problem, TPU or not?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1106434,
          "author_name": "akagarw",
          "author_url": "",
          "post_date": "12/08/2020 20:44:59",
          "content": "<p>While we can solve the variation in GPU setup easily by setting the random seeds… Therefore it's more seen in TPU setup as the model is trained on multi core. The sharding and loss calculation techniques leads to different results each time and also different from a version that is trained on GPU. <br>\nFollowing documentation explains the reasons behind the variations very well along with debugging methods. <br>\n<a href=\"https://cloud.google.com/tpu/docs/troubleshooting#model-accuracy\" target=\"_blank\">https://cloud.google.com/tpu/docs/troubleshooting#model-accuracy</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 795850,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "04/03/2020 04:51:32",
      "content": "<p><a href=\"/suicaokhoailang\">@suicaokhoailang</a> Yeah. But I still think tensorflow can do the same thing like pytorch. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 796028,
      "author_name": "miklgr500",
      "author_url": "",
      "post_date": "04/03/2020 08:22:26",
      "content": "<p>It's tensorflow features/bugs(it's hard to say what is it really). You can obtain the same results in two independent running only if would not be used parallelism (used 1 core of cpu only), but and this not guaranteed the same results. P.s. This <a href=\"https://www.youtube.com/watch?v=Ys8ofBeR2kA&amp;feature=youtu.be\">video </a>can help you understand why it's happening.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 796790,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "04/03/2020 23:14:49",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> </p>\n\n<p>In my custom version, something like the following helps, but not completely deterministic (results are close in different run, but still having some small difference).</p>\n\n<p>In tensorflow, there are global seed and operation seed. You have to make sure both of them are set.</p>\n\n<p><code>\n    self.dense_layer = tf.keras.layers.Dense(1, activation='sigmoid', kernel_initializer=tf.keras.initializers.GlorotUniform(seed=OP_SEED), bias_initializer='zeros')\n</code>\nalong with </p>\n\n<p>```\n    NO_RANDOM = True</p>\n\n<pre><code>GLOBAL_SEED = None\nOP_SEED = None\n\nif NO_RANDOM:\n    GLOBAL_SEED = 127\n    OP_SEED = 127\n    tf.random.set_seed(seed=GLOBAL_SEED)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 796807,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "04/04/2020 00:07:41",
      "content": "<p>I also saw </p>\n\n<pre><code>os.environ['TF_DETERMINISTIC_OPS'] = '1'\n</code></pre>\n\n<p>from </p>\n\n<p><a href=\"https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops\">https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 796852,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "04/04/2020 02:06:00",
      "content": "<p><a href=\"/miklgr500\">@miklgr500</a> <a href=\"/yihdarshieh\">@yihdarshieh</a>  Thanks for your help.  I tested it again, now GPU is ok ,but TPU is not ok.  I still got  random results.   : (</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 802012,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "04/09/2020 03:32:27",
      "content": "<p><a href=\"https://medium.com/datadriveninvestor/getting-reproducible-results-in-tensorflow-3705536aa185\">https://medium.com/datadriveninvestor/getting-reproducible-results-in-tensorflow-3705536aa185</a>, you can always check this out if this helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 805309,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "04/12/2020 15:01:38",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a>  Thanks for your help.  But it can not work in multi-core TPU. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 805264,
      "author_name": "jerryqu",
      "author_url": "",
      "post_date": "04/12/2020 14:13:36",
      "content": "<p>Check out this help section about Deterministic Training on TPUs: <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training\">https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training</a></p>\n\n<p>Main takeaway: Single core TPU training can be deterministic, multi-core TPU training may lead to varying results (Due to data sharding)</p>",
      "votes": null,
      "replies": [
        {
          "id": 805305,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "04/12/2020 14:56:17",
          "content": "<p><a href=\"/jerryqu\">@jerryqu</a>  Thanks for your help.  So we always get the random results for multi-core TPU.  : (</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 806798,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "04/14/2020 04:39:22",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> In my understanding, if we turn on os.environ['TF_DETERMINISTIC_OPS'] = '1', training speed is way slower. Have you experienced that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 806805,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "04/14/2020 04:52:33",
          "content": "<p><a href=\"/bamps53\">@bamps53</a>  No, I do not test the speed.  When we use the deterministic function, we just want to check if the new features or changes can bring some improvements in kpi. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 807257,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/14/2020 14:15:33",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> I know what you want to do, but I saw someone reported if you turn on TFDETERMINISTICOPS, it makes training speed very slower. (he said more than 1.5 times) So I gave up to do even I eagerly want to deterministic result! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 808226,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "04/15/2020 08:20:06",
      "content": "<p>yes, this is really cruel. Did you find any solution? <a href=\"/qinhui1999\">@qinhui1999</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 808381,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "04/15/2020 11:24:16",
          "content": "<p><a href=\"/shahules\">@shahules</a> No solution. You can check this page: <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training\">https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 808417,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "04/15/2020 12:13:52",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> Thanks for your reply, I was getting good performece when Running But it decreases whenever I commit.Any hacks for this issue ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "795702": "Hi,\n\n\n\n I tried to run @xhlulu 's tensorflow kernel.  But after running  the kernel many times, I always got the different or random results  for every time running.  <br> My kernel: https://www.kaggle.com/qinhui1999/jigsaw-tpu-xlm-roberta?scriptVersionId=31385845 <br>\nIn pytorch, if I fix the random seeds, then I always get the same result for every time running.  Hence I tried to fix the random seeds like pytorch as follow:\n\nSEED=2020<br>\nimport random<br>\ndef set_seed(seed):<br>\n    random.seed(seed)<br>\n    np.random.seed(seed)<br>\n    torch.manual_seed(seed)<br>\n    torch.cuda.manual_seed(seed)<br>\nset_seed(SEED) <br>\n\nTo my dispointed,  it does not work.  \nIs it anyway to make TPU  get the deterministic result?\n\n--------updated 04-03----\nFor tf 2.1 version, I use the following config and get the deterministic result in GPU.  But it can not work in TPU.  \n\nSEED=2020<br>\nos.environ['PYTHONHASHSEED']=str(SEED)<br>\nrandom.seed(SEED)<br>\nnp.random.seed(SEED)<br>\ntf.random.set_seed(SEED)<br>\n\nos.environ['TF\\_DETERMINISTIC\\_OPS'] = '1'<br>\n\n---------- other reference--------\nNvidia TensorFlow Determinism ---- https://github.com/NVIDIA/tensorflow-determinism",
    "795841": "I thought random results is a tensorflow problem, TPU or not?",
    "795850": "suicaokhoailang Yeah. But I still think tensorflow can do the same thing like pytorch.",
    "796028": "It's tensorflow features/bugs(it's hard to say what is it really). You can obtain the same results in two independent running only if would not be used parallelism (used 1 core of cpu only), but and this not guaranteed the same results. P.s. This [video ](https://www.youtube.com/watch?v=Ys8ofBeR2kA&amp;feature=youtu.be)can help you understand why it's happening.",
    "796790": "qinhui1999 \n\nIn my custom version, something like the following helps, but not completely deterministic (results are close in different run, but still having some small difference).\n\nIn tensorflow, there are global seed and operation seed. You have to make sure both of them are set.\n\n\n```\n    self.dense_layer = tf.keras.layers.Dense(1, activation='sigmoid', kernel_initializer=tf.keras.initializers.GlorotUniform(seed=OP_SEED), bias_initializer='zeros')\n```\nalong with \n\n```\n    NO_RANDOM = True\n\n    GLOBAL_SEED = None\n    OP_SEED = None\n\n    if NO_RANDOM:\n        GLOBAL_SEED = 127\n        OP_SEED = 127\n        tf.random.set_seed(seed=GLOBAL_SEED)",
    "796807": "I also saw \n\n    os.environ['TF_DETERMINISTIC_OPS'] = '1'\n\nfrom \n\n[https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops](https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops)",
    "796852": "miklgr500 @yihdarshieh  Thanks for your help.  I tested it again, now GPU is ok ,but TPU is not ok.  I still got  random results.   : (",
    "802012": "https://medium.com/datadriveninvestor/getting-reproducible-results-in-tensorflow-3705536aa185, you can always check this out if this helps.",
    "805264": "Check out this help section about Deterministic Training on TPUs: https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training\n\nMain takeaway: Single core TPU training can be deterministic, multi-core TPU training may lead to varying results (Due to data sharding)",
    "805305": "jerryqu  Thanks for your help.  So we always get the random results for multi-core TPU.  : (",
    "805309": "harshitsheoran  Thanks for your help.  But it can not work in multi-core TPU.",
    "806798": "qinhui1999 In my understanding, if we turn on os.environ['TF_DETERMINISTIC_OPS'] = '1', training speed is way slower. Have you experienced that?",
    "806805": "bamps53  No, I do not test the speed.  When we use the deterministic function, we just want to check if the new features or changes can bring some improvements in kpi.",
    "807257": "qinhui1999 I know what you want to do, but I saw someone reported if you turn on TFDETERMINISTICOPS, it makes training speed very slower. (he said more than 1.5 times) So I gave up to do even I eagerly want to deterministic result!",
    "808226": "yes, this is really cruel. Did you find any solution? @qinhui1999",
    "808381": "shahules No solution. You can check this page: https://cloud.google.com/tpu/docs/troubleshooting#deterministic-training",
    "808417": "qinhui1999 Thanks for your reply, I was getting good performece when Running But it decreases whenever I commit.Any hacks for this issue ?",
    "1106434": "While we can solve the variation in GPU setup easily by setting the random seeds... Therefore it's more seen in TPU setup as the model is trained on multi core. The sharding and loss calculation techniques leads to different results each time and also different from a version that is trained on GPU. \nFollowing documentation explains the reasons behind the variations very well along with debugging methods. \nhttps://cloud.google.com/tpu/docs/troubleshooting#model-accuracy"
  },
  "source": "meta"
}