{
  "id": 72040,
  "title": "Public LB variation",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72040",
  "author_name": "",
  "post_date": "2018-11-19T18:26:17.928403900Z",
  "votes": 31,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I accidentally ran the same version of my model twice, the results were surprising.</p>\n\n<p>At first, I got 0.693 (this is my current best)\nWhen I ran the exact same model (same random-seed, same everything) again, I got 0.685.\nNow, I tried it for the 3rd time and the public LB is: 0.684</p>\n\n<p>btw, the execution time varied a lot too:</p>\n\n<ol>\n<li>0.693 - 6970.2s</li>\n<li>0.685 - 7124.3s</li>\n<li>0.684 - 7213.0s (yes, kaggle accepted it!)</li>\n</ol>\n\n<p>Here is a screenshot:</p>\n\n<p><img src=\"https://albumizr.com/ia/fa292aaf3f12ce79bf8734e3343af18c.jpg\" alt=\"enter image description here\"></p>\n\n<p>Anyone can confirm (or disprove) this?\n(I used keras)</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "424208",
      "postDate": "11/19/2018 18:26:17",
      "content": "<p>I accidentally ran the same version of my model twice, the results were surprising.</p>\n\n<p>At first, I got 0.693 (this is my current best)\nWhen I ran the exact same model (same random-seed, same everything) again, I got 0.685.\nNow, I tried it for the 3rd time and the public LB is: 0.684</p>\n\n<p>btw, the execution time varied a lot too:</p>\n\n<ol>\n<li>0.693 - 6970.2s</li>\n<li>0.685 - 7124.3s</li>\n<li>0.684 - 7213.0s (yes, kaggle accepted it!)</li>\n</ol>\n\n<p>Here is a screenshot:</p>\n\n<p><img src=\"https://albumizr.com/ia/fa292aaf3f12ce79bf8734e3343af18c.jpg\" alt=\"enter image description here\"></p>\n\n<p>Anyone can confirm (or disprove) this?\n(I used keras)</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "I accidentally ran the same version of my model twice, the results were surprising.\n\nAt first, I got 0.693 (this is my current best)\nWhen I ran the exact same model (same random-seed, same everything) again, I got 0.685.\nNow, I tried it for the 3rd time and the public LB is: 0.684\n\nbtw, the execution time varied a lot too:\n\n 1.  0.693 - 6970.2s\n 2. 0.685 - 7124.3s\n 3. 0.684 - 7213.0s (yes, kaggle accepted it!)\n\nHere is a screenshot:\n\n![enter image description here][1]\n\n\nAnyone can confirm (or disprove) this?\n(I used keras)\n\nThanks\n\n\n  [1]: https://albumizr.com/ia/fa292aaf3f12ce79bf8734e3343af18c.jpg",
      "votes": null
    },
    {
      "id": "424230",
      "postDate": "11/19/2018 19:01:59",
      "content": "<p>I also use keras and I cant reproduce my results. I think it is due to using of cudnnGru or cudnnLstm. With these you cant reproduce the result's. Are you using CUDNN layers ?</p>\n\n<p>Link:\n<a href=\"https://github.com/keras-team/keras/issues/2479#issuecomment-213987747\">https://github.com/keras-team/keras/issues/2479#issuecomment-213987747</a></p>",
      "rawMarkdown": "I also use keras and I cant reproduce my results. I think it is due to using of cudnnGru or cudnnLstm. With these you cant reproduce the result's. Are you using CUDNN layers ?\n\n\nLink:\nhttps://github.com/keras-team/keras/issues/2479#issuecomment-213987747",
      "votes": null
    },
    {
      "id": "424245",
      "postDate": "11/19/2018 19:40:34",
      "content": "<p><a href=\"/suchith0312\">@suchith0312</a> yes, I use cudnngru. I thought keras (or the cudnnXX implementation) might be the problem. \nI think we can't afford that much difference, because kaggle will re-run our scripts. There will be lots of up/down movement on the private lb.. anyway I am switching to pytorch.</p>\n\n<p>thanks for your reply.</p>",
      "rawMarkdown": "suchith0312 yes, I use cudnngru. I thought keras (or the cudnnXX implementation) might be the problem. \nI think we can't afford that much difference, because kaggle will re-run our scripts. There will be lots of up/down movement on the private lb.. anyway I am switching to pytorch.\n\nthanks for your reply.",
      "votes": null
    },
    {
      "id": "424264",
      "postDate": "11/19/2018 20:15:12",
      "content": "<p>This happened due to randomness which is even more in CUDNN as compared to other keras model layers. Reproducing results might get difficult in such cases, as it is hard to go at each code snippet (not exposed easily) and fix the value for random seed.</p>",
      "rawMarkdown": "This happened due to randomness which is even more in CUDNN as compared to other keras model layers. Reproducing results might get difficult in such cases, as it is hard to go at each code snippet (not exposed easily) and fix the value for random seed.",
      "votes": null
    },
    {
      "id": "424281",
      "postDate": "11/19/2018 20:41:13",
      "content": "<p><a href=\"/nitinaggarwal008\">@nitinaggarwal008</a> I rarely use keras, I only using it because it was easier to copy SRK's starter code :) I was lazy to do it in pytorch. My crrent best is 14th on the public LB, the same model could end up at ~250-260 (current position for 0.684). That is way too extreme difference.  </p>",
      "rawMarkdown": "nitinaggarwal008 I rarely use keras, I only using it because it was easier to copy SRK's starter code :) I was lazy to do it in pytorch. My crrent best is 14th on the public LB, the same model could end up at ~250-260 (current position for 0.684). That is way too extreme difference.",
      "votes": null
    },
    {
      "id": "424282",
      "postDate": "11/19/2018 20:43:26",
      "content": "<p>I truly hope in stage 2, Kaggle can run our models multiple times and take an average. Too much randomness in CUDNN.</p>",
      "rawMarkdown": "I truly hope in stage 2, Kaggle can run our models multiple times and take an average. Too much randomness in CUDNN.",
      "votes": null
    },
    {
      "id": "424288",
      "postDate": "11/19/2018 21:02:44",
      "content": "<p><a href=\"/shujian\">@shujian</a> I don't think so. </p>\n\n<blockquote>\n  <p>The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;</p>\n</blockquote>\n\n<p>I don't think they will bother to run multiple times and average the results. The rule is clear: must be capable of generating the winning submission.</p>\n\n<p>Nobody forced me to use keras and/or cudnn. I knew it has this kind of behaviour, but I did not expect that much. </p>\n\n<p>btw, my model was a 5-fold cv and still have ~1% difference.</p>\n\n<p>The annoying part is that in all of my experiments could have the same difference.. I will have to start over everything.</p>\n\n<p>Lesson learned.</p>",
      "rawMarkdown": "shujian I don't think so. \n\n&gt; The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;\n\nI don't think they will bother to run multiple times and average the results. The rule is clear: must be capable of generating the winning submission.\n\nNobody forced me to use keras and/or cudnn. I knew it has this kind of behaviour, but I did not expect that much. \n\nbtw, my model was a 5-fold cv and still have ~1% difference.\n\nThe annoying part is that in all of my experiments could have the same difference.. I will have to start over everything.\n\nLesson learned.",
      "votes": null
    },
    {
      "id": "424291",
      "postDate": "11/19/2018 21:10:01",
      "content": "<p>I remember during Mercari competition ( kernel only ) We had 2 DL models : one in pure tensorflow and another in keras . While the tensorflow model gave stable results, there was some variation in the keras model (even without CuDNN)  ..But not such huge. </p>\n\n<p>Thanksfully it didn't have much impact on 2nd stage LB. </p>",
      "rawMarkdown": "I remember during Mercari competition ( kernel only ) We had 2 DL models : one in pure tensorflow and another in keras . While the tensorflow model gave stable results, there was some variation in the keras model (even without CuDNN)  ..But not such huge. \n\nThanksfully it didn't have much impact on 2nd stage LB.",
      "votes": null
    },
    {
      "id": "424294",
      "postDate": "11/19/2018 21:15:28",
      "content": "<p>AFAIK, Keras, Tensorflow and PyTorch just provide a wrapper of NVIDIA CuDNN implementation. So I'm wondering would it be any difference in randomness if you switch from one framework to another.</p>",
      "rawMarkdown": "AFAIK, Keras, Tensorflow and PyTorch just provide a wrapper of NVIDIA CuDNN implementation. So I'm wondering would it be any difference in randomness if you switch from one framework to another.",
      "votes": null
    },
    {
      "id": "424319",
      "postDate": "11/19/2018 22:12:17",
      "content": "<p>Hi Peter! Thank you for also confirming this on your end. I saw the same on my end in the past few days - I burned a few submissions to look at the impact to variations. I've also heard similar about CuDNN variants; thank you for posting this. (upvoted)</p>",
      "rawMarkdown": "Hi Peter! Thank you for also confirming this on your end. I saw the same on my end in the past few days - I burned a few submissions to look at the impact to variations. I've also heard similar about CuDNN variants; thank you for posting this. (upvoted)",
      "votes": null
    },
    {
      "id": "424321",
      "postDate": "11/19/2018 22:15:03",
      "content": "<p>You are welcome, and thanks for the upvote!</p>",
      "rawMarkdown": "You are welcome, and thanks for the upvote!",
      "votes": null
    },
    {
      "id": "424322",
      "postDate": "11/19/2018 22:17:17",
      "content": "<p><a href=\"/thinline72\">@thinline72</a>, it will be a good practice to implement everything in pytorch too. I'll post my results here. I hope it has not the same randomness, we'll see.</p>",
      "rawMarkdown": "thinline72, it will be a good practice to implement everything in pytorch too. I'll post my results here. I hope it has not the same randomness, we'll see.",
      "votes": null
    },
    {
      "id": "424335",
      "postDate": "11/19/2018 23:24:47",
      "content": "<p>The time variation is standard Kaggle kernel issues, for Mercari there were pretty significant runtime differences depending on time of day. For Keras, they just don't seem to have enough random state setting for the random parts of neural network to be \"deterministic\"</p>",
      "rawMarkdown": "The time variation is standard Kaggle kernel issues, for Mercari there were pretty significant runtime differences depending on time of day. For Keras, they just don't seem to have enough random state setting for the random parts of neural network to be \"deterministic\"",
      "votes": null
    },
    {
      "id": "425833",
      "postDate": "11/22/2018 07:27:49",
      "content": "<p>haha this is pretty cool, maybe the random seed value is differnet ?</p>",
      "rawMarkdown": "haha this is pretty cool, maybe the random seed value is differnet ?",
      "votes": null
    },
    {
      "id": "426618",
      "postDate": "11/23/2018 14:45:09",
      "content": "<p>Just to add another datapoint - I am unable to achieve deterministic results using PyTorch either.  I am (unsuccessfully) attempting to achieve determinism using:\n<code>\n    seed = 0\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    random.seed(seed)\n    torch.manual_seed(seed)\n    cuda.manual_seed(seed)\n    np.random.seed(seed)\n</code>\nEven if I disable CuDNN (change top setting above to False) I STILL am not getting deterministic results (in offline testing - have not tried in a kernel with CuDNN disabled due to extra compute time that requires).\nMy model is basically a BiLSTM with some extra structure around combining multiple embeddings to its input (which are just regularized FC layers at the end of the day)</p>",
      "rawMarkdown": "Just to add another datapoint - I am unable to achieve deterministic results using PyTorch either.  I am (unsuccessfully) attempting to achieve determinism using:\n```\n    seed = 0\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    random.seed(seed)\n    torch.manual_seed(seed)\n    cuda.manual_seed(seed)\n    np.random.seed(seed)\n```\nEven if I disable CuDNN (change top setting above to False) I STILL am not getting deterministic results (in offline testing - have not tried in a kernel with CuDNN disabled due to extra compute time that requires).\nMy model is basically a BiLSTM with some extra structure around combining multiple embeddings to its input (which are just regularized FC layers at the end of the day)",
      "votes": null
    },
    {
      "id": "426631",
      "postDate": "11/23/2018 15:10:19",
      "content": "<p><a href=\"/stevedraper\">@stevedraper</a>, that is interesting. I've just finished my initial experiments with PyTorch, for me it is working. I've got the exact same score/results (local) for multiple runs. I log and check everything almost line-by-line (including fold generation, tokenization, etc), to make sure it is working. So far, seems ok. </p>\n\n<p>I am using these configs:</p>\n\n<blockquote>\n  <p>torch.manual_seed(SEED)</p>\n  \n  <p>torch.cuda.manual_seed(SEED)</p>\n  \n  <p>torch.backends.cudnn.deterministic = True</p>\n</blockquote>",
      "rawMarkdown": "stevedraper, that is interesting. I've just finished my initial experiments with PyTorch, for me it is working. I've got the exact same score/results (local) for multiple runs. I log and check everything almost line-by-line (including fold generation, tokenization, etc), to make sure it is working. So far, seems ok. \n\nI am using these configs:\n\n&gt; torch.manual_seed(SEED)\n\n&gt; torch.cuda.manual_seed(SEED)\n\n&gt; torch.backends.cudnn.deterministic = True",
      "votes": null
    },
    {
      "id": "426759",
      "postDate": "11/23/2018 19:34:57",
      "content": "<p>After adding gradient clipping (max_norm) I recovered determinism.  I don't have a good theory as to why except perhaps some non-deterministic rounding going on somewhere.</p>",
      "rawMarkdown": "After adding gradient clipping (max_norm) I recovered determinism.  I don't have a good theory as to why except perhaps some non-deterministic rounding going on somewhere.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 424230,
      "author_name": "suchith0312",
      "author_url": "",
      "post_date": "11/19/2018 19:01:59",
      "content": "<p>I also use keras and I cant reproduce my results. I think it is due to using of cudnnGru or cudnnLstm. With these you cant reproduce the result's. Are you using CUDNN layers ?</p>\n\n<p>Link:\n<a href=\"https://github.com/keras-team/keras/issues/2479#issuecomment-213987747\">https://github.com/keras-team/keras/issues/2479#issuecomment-213987747</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 424245,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/19/2018 19:40:34",
          "content": "<p><a href=\"/suchith0312\">@suchith0312</a> yes, I use cudnngru. I thought keras (or the cudnnXX implementation) might be the problem. \nI think we can't afford that much difference, because kaggle will re-run our scripts. There will be lots of up/down movement on the private lb.. anyway I am switching to pytorch.</p>\n\n<p>thanks for your reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424264,
      "author_name": "nitinaggarwal008",
      "author_url": "",
      "post_date": "11/19/2018 20:15:12",
      "content": "<p>This happened due to randomness which is even more in CUDNN as compared to other keras model layers. Reproducing results might get difficult in such cases, as it is hard to go at each code snippet (not exposed easily) and fix the value for random seed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 424281,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/19/2018 20:41:13",
          "content": "<p><a href=\"/nitinaggarwal008\">@nitinaggarwal008</a> I rarely use keras, I only using it because it was easier to copy SRK's starter code :) I was lazy to do it in pytorch. My crrent best is 14th on the public LB, the same model could end up at ~250-260 (current position for 0.684). That is way too extreme difference.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424282,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/19/2018 20:43:26",
      "content": "<p>I truly hope in stage 2, Kaggle can run our models multiple times and take an average. Too much randomness in CUDNN.</p>",
      "votes": null,
      "replies": [
        {
          "id": 424288,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/19/2018 21:02:44",
          "content": "<p><a href=\"/shujian\">@shujian</a> I don't think so. </p>\n\n<blockquote>\n  <p>The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;</p>\n</blockquote>\n\n<p>I don't think they will bother to run multiple times and average the results. The rule is clear: must be capable of generating the winning submission.</p>\n\n<p>Nobody forced me to use keras and/or cudnn. I knew it has this kind of behaviour, but I did not expect that much. </p>\n\n<p>btw, my model was a 5-fold cv and still have ~1% difference.</p>\n\n<p>The annoying part is that in all of my experiments could have the same difference.. I will have to start over everything.</p>\n\n<p>Lesson learned.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424291,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "11/19/2018 21:10:01",
      "content": "<p>I remember during Mercari competition ( kernel only ) We had 2 DL models : one in pure tensorflow and another in keras . While the tensorflow model gave stable results, there was some variation in the keras model (even without CuDNN)  ..But not such huge. </p>\n\n<p>Thanksfully it didn't have much impact on 2nd stage LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 424294,
      "author_name": "thinline72",
      "author_url": "",
      "post_date": "11/19/2018 21:15:28",
      "content": "<p>AFAIK, Keras, Tensorflow and PyTorch just provide a wrapper of NVIDIA CuDNN implementation. So I'm wondering would it be any difference in randomness if you switch from one framework to another.</p>",
      "votes": null,
      "replies": [
        {
          "id": 424322,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/19/2018 22:17:17",
          "content": "<p><a href=\"/thinline72\">@thinline72</a>, it will be a good practice to implement everything in pytorch too. I'll post my results here. I hope it has not the same randomness, we'll see.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424319,
      "author_name": "learnmower",
      "author_url": "",
      "post_date": "11/19/2018 22:12:17",
      "content": "<p>Hi Peter! Thank you for also confirming this on your end. I saw the same on my end in the past few days - I burned a few submissions to look at the impact to variations. I've also heard similar about CuDNN variants; thank you for posting this. (upvoted)</p>",
      "votes": null,
      "replies": [
        {
          "id": 424321,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/19/2018 22:15:03",
          "content": "<p>You are welcome, and thanks for the upvote!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 424335,
      "author_name": "msf908",
      "author_url": "",
      "post_date": "11/19/2018 23:24:47",
      "content": "<p>The time variation is standard Kaggle kernel issues, for Mercari there were pretty significant runtime differences depending on time of day. For Keras, they just don't seem to have enough random state setting for the random parts of neural network to be \"deterministic\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 425833,
      "author_name": "",
      "author_url": "",
      "post_date": "11/22/2018 07:27:49",
      "content": "<p>haha this is pretty cool, maybe the random seed value is differnet ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 426618,
      "author_name": "stevedraper",
      "author_url": "",
      "post_date": "11/23/2018 14:45:09",
      "content": "<p>Just to add another datapoint - I am unable to achieve deterministic results using PyTorch either.  I am (unsuccessfully) attempting to achieve determinism using:\n<code>\n    seed = 0\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    random.seed(seed)\n    torch.manual_seed(seed)\n    cuda.manual_seed(seed)\n    np.random.seed(seed)\n</code>\nEven if I disable CuDNN (change top setting above to False) I STILL am not getting deterministic results (in offline testing - have not tried in a kernel with CuDNN disabled due to extra compute time that requires).\nMy model is basically a BiLSTM with some extra structure around combining multiple embeddings to its input (which are just regularized FC layers at the end of the day)</p>",
      "votes": null,
      "replies": [
        {
          "id": 426631,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/23/2018 15:10:19",
          "content": "<p><a href=\"/stevedraper\">@stevedraper</a>, that is interesting. I've just finished my initial experiments with PyTorch, for me it is working. I've got the exact same score/results (local) for multiple runs. I log and check everything almost line-by-line (including fold generation, tokenization, etc), to make sure it is working. So far, seems ok. </p>\n\n<p>I am using these configs:</p>\n\n<blockquote>\n  <p>torch.manual_seed(SEED)</p>\n  \n  <p>torch.cuda.manual_seed(SEED)</p>\n  \n  <p>torch.backends.cudnn.deterministic = True</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 426759,
          "author_name": "stevedraper",
          "author_url": "",
          "post_date": "11/23/2018 19:34:57",
          "content": "<p>After adding gradient clipping (max_norm) I recovered determinism.  I don't have a good theory as to why except perhaps some non-deterministic rounding going on somewhere.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "424208": "I accidentally ran the same version of my model twice, the results were surprising.\n\nAt first, I got 0.693 (this is my current best)\nWhen I ran the exact same model (same random-seed, same everything) again, I got 0.685.\nNow, I tried it for the 3rd time and the public LB is: 0.684\n\nbtw, the execution time varied a lot too:\n\n 1.  0.693 - 6970.2s\n 2. 0.685 - 7124.3s\n 3. 0.684 - 7213.0s (yes, kaggle accepted it!)\n\nHere is a screenshot:\n\n![enter image description here][1]\n\n\nAnyone can confirm (or disprove) this?\n(I used keras)\n\nThanks\n\n\n  [1]: https://albumizr.com/ia/fa292aaf3f12ce79bf8734e3343af18c.jpg",
    "424230": "I also use keras and I cant reproduce my results. I think it is due to using of cudnnGru or cudnnLstm. With these you cant reproduce the result's. Are you using CUDNN layers ?\n\n\nLink:\nhttps://github.com/keras-team/keras/issues/2479#issuecomment-213987747",
    "424245": "suchith0312 yes, I use cudnngru. I thought keras (or the cudnnXX implementation) might be the problem. \nI think we can't afford that much difference, because kaggle will re-run our scripts. There will be lots of up/down movement on the private lb.. anyway I am switching to pytorch.\n\nthanks for your reply.",
    "424264": "This happened due to randomness which is even more in CUDNN as compared to other keras model layers. Reproducing results might get difficult in such cases, as it is hard to go at each code snippet (not exposed easily) and fix the value for random seed.",
    "424281": "nitinaggarwal008 I rarely use keras, I only using it because it was easier to copy SRK's starter code :) I was lazy to do it in pytorch. My crrent best is 14th on the public LB, the same model could end up at ~250-260 (current position for 0.684). That is way too extreme difference.",
    "424282": "I truly hope in stage 2, Kaggle can run our models multiple times and take an average. Too much randomness in CUDNN.",
    "424288": "shujian I don't think so. \n\n&gt; The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;\n\nI don't think they will bother to run multiple times and average the results. The rule is clear: must be capable of generating the winning submission.\n\nNobody forced me to use keras and/or cudnn. I knew it has this kind of behaviour, but I did not expect that much. \n\nbtw, my model was a 5-fold cv and still have ~1% difference.\n\nThe annoying part is that in all of my experiments could have the same difference.. I will have to start over everything.\n\nLesson learned.",
    "424291": "I remember during Mercari competition ( kernel only ) We had 2 DL models : one in pure tensorflow and another in keras . While the tensorflow model gave stable results, there was some variation in the keras model (even without CuDNN)  ..But not such huge. \n\nThanksfully it didn't have much impact on 2nd stage LB.",
    "424294": "AFAIK, Keras, Tensorflow and PyTorch just provide a wrapper of NVIDIA CuDNN implementation. So I'm wondering would it be any difference in randomness if you switch from one framework to another.",
    "424319": "Hi Peter! Thank you for also confirming this on your end. I saw the same on my end in the past few days - I burned a few submissions to look at the impact to variations. I've also heard similar about CuDNN variants; thank you for posting this. (upvoted)",
    "424321": "You are welcome, and thanks for the upvote!",
    "424322": "thinline72, it will be a good practice to implement everything in pytorch too. I'll post my results here. I hope it has not the same randomness, we'll see.",
    "424335": "The time variation is standard Kaggle kernel issues, for Mercari there were pretty significant runtime differences depending on time of day. For Keras, they just don't seem to have enough random state setting for the random parts of neural network to be \"deterministic\"",
    "425833": "haha this is pretty cool, maybe the random seed value is differnet ?",
    "426618": "Just to add another datapoint - I am unable to achieve deterministic results using PyTorch either.  I am (unsuccessfully) attempting to achieve determinism using:\n```\n    seed = 0\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    random.seed(seed)\n    torch.manual_seed(seed)\n    cuda.manual_seed(seed)\n    np.random.seed(seed)\n```\nEven if I disable CuDNN (change top setting above to False) I STILL am not getting deterministic results (in offline testing - have not tried in a kernel with CuDNN disabled due to extra compute time that requires).\nMy model is basically a BiLSTM with some extra structure around combining multiple embeddings to its input (which are just regularized FC layers at the end of the day)",
    "426631": "stevedraper, that is interesting. I've just finished my initial experiments with PyTorch, for me it is working. I've got the exact same score/results (local) for multiple runs. I log and check everything almost line-by-line (including fold generation, tokenization, etc), to make sure it is working. So far, seems ok. \n\nI am using these configs:\n\n&gt; torch.manual_seed(SEED)\n\n&gt; torch.cuda.manual_seed(SEED)\n\n&gt; torch.backends.cudnn.deterministic = True",
    "426759": "After adding gradient clipping (max_norm) I recovered determinism.  I don't have a good theory as to why except perhaps some non-deterministic rounding going on somewhere."
  },
  "source": "meta"
}