{
  "id": 72509,
  "title": "Large variations in results in identical kernel runs",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72509",
  "author_name": "",
  "post_date": "2018-11-24T02:00:06.022999100Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I'm using pytorch for this competition and getting huge variations in loss and score when committing <a href=\"https://www.kaggle.com/bkkaggle/fork-of-quora-other-attention-relu-95k-words-spacy/notebook\">my kernel</a>  several times while not changing the code at all. There are no differences between the two versions of my kernel but the second version gets stuck at a loss of 0.23 and a f1 score of 0 while the first version gets a public score of 0.677. I added the pytorch code snippet from <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040#426631\">this post</a> which also talks about problems with determinism in pytorch to a third version of my kernel, but it also gets stuck.</p>\n\n<p>Has anyone else had problems like this?</p>",
  "messages": [
    {
      "id": "426846",
      "postDate": "11/24/2018 02:00:06",
      "content": "<p>Hi,</p>\n\n<p>I'm using pytorch for this competition and getting huge variations in loss and score when committing <a href=\"https://www.kaggle.com/bkkaggle/fork-of-quora-other-attention-relu-95k-words-spacy/notebook\">my kernel</a>  several times while not changing the code at all. There are no differences between the two versions of my kernel but the second version gets stuck at a loss of 0.23 and a f1 score of 0 while the first version gets a public score of 0.677. I added the pytorch code snippet from <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040#426631\">this post</a> which also talks about problems with determinism in pytorch to a third version of my kernel, but it also gets stuck.</p>\n\n<p>Has anyone else had problems like this?</p>",
      "rawMarkdown": "Hi,\n\nI'm using pytorch for this competition and getting huge variations in loss and score when committing [my kernel][1]  several times while not changing the code at all. There are no differences between the two versions of my kernel but the second version gets stuck at a loss of 0.23 and a f1 score of 0 while the first version gets a public score of 0.677. I added the pytorch code snippet from [this post][2] which also talks about problems with determinism in pytorch to a third version of my kernel, but it also gets stuck.\n\nHas anyone else had problems like this?\n\n  [1]: https://www.kaggle.com/bkkaggle/fork-of-quora-other-attention-relu-95k-words-spacy/notebook\n  [2]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040#426631",
      "votes": null
    },
    {
      "id": "427621",
      "postDate": "11/25/2018 21:33:56",
      "content": "<p>Yes, most of us have had issues like this. Apparently cudnn is a stochastic library and it's impossible to set its seeds. It is possible to start form random initial parameters which will get you into some kind of weird local minimum for which F1 = 0. Yet another challenge in this competition. :)</p>",
      "rawMarkdown": "Yes, most of us have had issues like this. Apparently cudnn is a stochastic library and it's impossible to set its seeds. It is possible to start form random initial parameters which will get you into some kind of weird local minimum for which F1 = 0. Yet another challenge in this competition. :)",
      "votes": null
    },
    {
      "id": "427669",
      "postDate": "11/25/2018 23:49:16",
      "content": "<p>Probably we should only use LR and XGB.</p>",
      "rawMarkdown": "Probably we should only use LR and XGB.",
      "votes": null
    },
    {
      "id": "427673",
      "postDate": "11/25/2018 23:57:20",
      "content": "<p>If you can figure out how to do it <strong>and</strong> get a competitive score, please share your kernel. ;)</p>",
      "rawMarkdown": "If you can figure out how to do it **and** get a competitive score, please share your kernel. ;)",
      "votes": null
    },
    {
      "id": "427730",
      "postDate": "11/26/2018 03:27:11",
      "content": "<p><a href=\"/bkkaggle\">@bkkaggle</a>, this is an issue we are all facing as mentioned by @Bojan here. It is good to know this variation is not unique to the keras DL library. In addition to setting the main seed, I have also tried setting the CuDNNLSTM, CuDNNGRU seeds to no avail. If anyone finds a way, please let us know :-)</p>",
      "rawMarkdown": "bkkaggle, this is an issue we are all facing as mentioned by @Bojan here. It is good to know this variation is not unique to the keras DL library. In addition to setting the main seed, I have also tried setting the CuDNNLSTM, CuDNNGRU seeds to no avail. If anyone finds a way, please let us know :-)",
      "votes": null
    },
    {
      "id": "427780",
      "postDate": "11/26/2018 06:06:50",
      "content": "<p>Interesting blog by Two Sigma: <a href=\"https://www.twosigma.com/insights/article/a-workaround-for-non-determinism-in-tensorflow/\">https://www.twosigma.com/insights/article/a-workaround-for-non-determinism-in-tensorflow/</a></p>",
      "rawMarkdown": "Interesting blog by Two Sigma: https://www.twosigma.com/insights/article/a-workaround-for-non-determinism-in-tensorflow/",
      "votes": null
    },
    {
      "id": "428302",
      "postDate": "11/27/2018 03:46:36",
      "content": "<p>Update:</p>\n\n<p>My kernel's variations in loss were due to having a relu layer after the attention layer which would zero-out some of the rescaled activations from the LSTM. After removing the relu layer, changing the initializer in the attention layer, and clipping the gradients of the model, different kernel runs no longer return large variations in results.</p>",
      "rawMarkdown": "Update:\n\nMy kernel's variations in loss were due to having a relu layer after the attention layer which would zero-out some of the rescaled activations from the LSTM. After removing the relu layer, changing the initializer in the attention layer, and clipping the gradients of the model, different kernel runs no longer return large variations in results.",
      "votes": null
    },
    {
      "id": "428942",
      "postDate": "11/28/2018 05:37:17",
      "content": "<p>I published here XGBoost baseline model :<a href=\"https://www.kaggle.com/jaguar00/xgboost-baseline\">https://www.kaggle.com/jaguar00/xgboost-baseline</a> \nIt uses character and word vectors. The model can be enhanced with more features and params tuning.</p>",
      "rawMarkdown": "I published here XGBoost baseline model :https://www.kaggle.com/jaguar00/xgboost-baseline \nIt uses character and word vectors. The model can be enhanced with more features and params tuning.",
      "votes": null
    },
    {
      "id": "431012",
      "postDate": "12/01/2018 12:33:11",
      "content": "<p>Lately, I am on this plateau too. All of my kernels' 0.00[3-8] scores are unstable on consecutive runs.</p>",
      "rawMarkdown": "Lately, I am on this plateau too. All of my kernels' 0.00[3-8] scores are unstable on consecutive runs.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 427621,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "11/25/2018 21:33:56",
      "content": "<p>Yes, most of us have had issues like this. Apparently cudnn is a stochastic library and it's impossible to set its seeds. It is possible to start form random initial parameters which will get you into some kind of weird local minimum for which F1 = 0. Yet another challenge in this competition. :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 427669,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "11/25/2018 23:49:16",
          "content": "<p>Probably we should only use LR and XGB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 427673,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "11/25/2018 23:57:20",
          "content": "<p>If you can figure out how to do it <strong>and</strong> get a competitive score, please share your kernel. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428942,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "11/28/2018 05:37:17",
          "content": "<p>I published here XGBoost baseline model :<a href=\"https://www.kaggle.com/jaguar00/xgboost-baseline\">https://www.kaggle.com/jaguar00/xgboost-baseline</a> \nIt uses character and word vectors. The model can be enhanced with more features and params tuning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 427730,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/26/2018 03:27:11",
      "content": "<p><a href=\"/bkkaggle\">@bkkaggle</a>, this is an issue we are all facing as mentioned by @Bojan here. It is good to know this variation is not unique to the keras DL library. In addition to setting the main seed, I have also tried setting the CuDNNLSTM, CuDNNGRU seeds to no avail. If anyone finds a way, please let us know :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 427780,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/26/2018 06:06:50",
      "content": "<p>Interesting blog by Two Sigma: <a href=\"https://www.twosigma.com/insights/article/a-workaround-for-non-determinism-in-tensorflow/\">https://www.twosigma.com/insights/article/a-workaround-for-non-determinism-in-tensorflow/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 428302,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "11/27/2018 03:46:36",
      "content": "<p>Update:</p>\n\n<p>My kernel's variations in loss were due to having a relu layer after the attention layer which would zero-out some of the rescaled activations from the LSTM. After removing the relu layer, changing the initializer in the attention layer, and clipping the gradients of the model, different kernel runs no longer return large variations in results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 431012,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "12/01/2018 12:33:11",
      "content": "<p>Lately, I am on this plateau too. All of my kernels' 0.00[3-8] scores are unstable on consecutive runs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "426846": "Hi,\n\nI'm using pytorch for this competition and getting huge variations in loss and score when committing [my kernel][1]  several times while not changing the code at all. There are no differences between the two versions of my kernel but the second version gets stuck at a loss of 0.23 and a f1 score of 0 while the first version gets a public score of 0.677. I added the pytorch code snippet from [this post][2] which also talks about problems with determinism in pytorch to a third version of my kernel, but it also gets stuck.\n\nHas anyone else had problems like this?\n\n  [1]: https://www.kaggle.com/bkkaggle/fork-of-quora-other-attention-relu-95k-words-spacy/notebook\n  [2]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040#426631",
    "427621": "Yes, most of us have had issues like this. Apparently cudnn is a stochastic library and it's impossible to set its seeds. It is possible to start form random initial parameters which will get you into some kind of weird local minimum for which F1 = 0. Yet another challenge in this competition. :)",
    "427669": "Probably we should only use LR and XGB.",
    "427673": "If you can figure out how to do it **and** get a competitive score, please share your kernel. ;)",
    "427730": "bkkaggle, this is an issue we are all facing as mentioned by @Bojan here. It is good to know this variation is not unique to the keras DL library. In addition to setting the main seed, I have also tried setting the CuDNNLSTM, CuDNNGRU seeds to no avail. If anyone finds a way, please let us know :-)",
    "427780": "Interesting blog by Two Sigma: https://www.twosigma.com/insights/article/a-workaround-for-non-determinism-in-tensorflow/",
    "428302": "Update:\n\nMy kernel's variations in loss were due to having a relu layer after the attention layer which would zero-out some of the rescaled activations from the LSTM. After removing the relu layer, changing the initializer in the attention layer, and clipping the gradients of the model, different kernel runs no longer return large variations in results.",
    "428942": "I published here XGBoost baseline model :https://www.kaggle.com/jaguar00/xgboost-baseline \nIt uses character and word vectors. The model can be enhanced with more features and params tuning.",
    "431012": "Lately, I am on this plateau too. All of my kernels' 0.00[3-8] scores are unstable on consecutive runs."
  },
  "source": "meta"
}