{
  "id": 72770,
  "title": "Pytorch and determinism",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72770",
  "author_name": "",
  "post_date": "2018-11-27T03:47:55.120594900Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I made a <a href=\"https://www.kaggle.com/bkkaggle/pytorch-determinism-test\">kernel</a> showing how making a pytorch kernel more deterministic affects the result of the kernel. I ran the kernel 3 times with and without determinism. Versions 4-6 are with determinism and versions 7-9 are without determinism.</p>\n\n<h3>Model:</h3>\n\n<p>I use a simple 2 layer bidirectional LSTM with an attention layer and a batch size of 512</p>\n\n<h3>Training:</h3>\n\n<p>I train for 7 epochs and choose the saved weights of the model with the lowest validation loss</p>\n\n<h3>Determinism:</h3>\n\n<p>For my determinism experiments I set torch's random seeds and set deterministic=True:</p>\n\n<pre><code>SEED = 1337\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\ntorch.cuda.manual_seed(SEED)\ntorch.backends.cudnn.deterministic = True\n</code></pre>\n\n<p>For my non-deterministic experiments I only set numpy's random seeds:</p>\n\n<pre><code>SEED = 1337\nnp.random.seed(SEED)\n</code></pre>\n\n<h3>Results:</h3>\n\n<p>| Deterministic? | Train F1 | Val F1 --| Train Loss | Val Loss | Best Epoch | Best threshold </p>\n\n<p>| Yes -------------|  0.781 -- | 0.6696 | 0.0842 ----| 0.1000 --| 5 ------------| 0.37</p>\n\n<p>| Yes -------------|  0.7684 -| 0.6773 | 0.0840 ----| 0.0992 --| 5 ------------| 0.33</p>\n\n<p>| Yes -------------| 0.7784 --| 0.6701 | 0.0957 -----| 0.1000 --| 2 ------------| 0.37</p>\n\n<hr>\n\n<p>| No ------------- | 0.7963 --| 0.6718 | 0.0930 -----| 0.0976 --| 2 ------------| 0.32</p>\n\n<p>| No ------------- | 0.7928 --| 0.6724 | 0.0930 ----| 0.0982 --| 2 ------------| 0.3</p>\n\n<p>| No ------------- | 0.7997 --| 0.6751 -| 0.0868 ----| 0.0974 --| 3 ------------| 0.35</p>\n\n<p>Non-deterministic runs have higher and more consistent Val F1s and have to be trained for less epochs to reach convergence. These results might change if the kernels are run more times and non-determinism might not give more stable results when using tensorflow or keras.</p>",
  "messages": [
    {
      "id": "428303",
      "postDate": "11/27/2018 03:47:55",
      "content": "<p>Hi,</p>\n\n<p>I made a <a href=\"https://www.kaggle.com/bkkaggle/pytorch-determinism-test\">kernel</a> showing how making a pytorch kernel more deterministic affects the result of the kernel. I ran the kernel 3 times with and without determinism. Versions 4-6 are with determinism and versions 7-9 are without determinism.</p>\n\n<h3>Model:</h3>\n\n<p>I use a simple 2 layer bidirectional LSTM with an attention layer and a batch size of 512</p>\n\n<h3>Training:</h3>\n\n<p>I train for 7 epochs and choose the saved weights of the model with the lowest validation loss</p>\n\n<h3>Determinism:</h3>\n\n<p>For my determinism experiments I set torch's random seeds and set deterministic=True:</p>\n\n<pre><code>SEED = 1337\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\ntorch.cuda.manual_seed(SEED)\ntorch.backends.cudnn.deterministic = True\n</code></pre>\n\n<p>For my non-deterministic experiments I only set numpy's random seeds:</p>\n\n<pre><code>SEED = 1337\nnp.random.seed(SEED)\n</code></pre>\n\n<h3>Results:</h3>\n\n<p>| Deterministic? | Train F1 | Val F1 --| Train Loss | Val Loss | Best Epoch | Best threshold </p>\n\n<p>| Yes -------------|  0.781 -- | 0.6696 | 0.0842 ----| 0.1000 --| 5 ------------| 0.37</p>\n\n<p>| Yes -------------|  0.7684 -| 0.6773 | 0.0840 ----| 0.0992 --| 5 ------------| 0.33</p>\n\n<p>| Yes -------------| 0.7784 --| 0.6701 | 0.0957 -----| 0.1000 --| 2 ------------| 0.37</p>\n\n<hr>\n\n<p>| No ------------- | 0.7963 --| 0.6718 | 0.0930 -----| 0.0976 --| 2 ------------| 0.32</p>\n\n<p>| No ------------- | 0.7928 --| 0.6724 | 0.0930 ----| 0.0982 --| 2 ------------| 0.3</p>\n\n<p>| No ------------- | 0.7997 --| 0.6751 -| 0.0868 ----| 0.0974 --| 3 ------------| 0.35</p>\n\n<p>Non-deterministic runs have higher and more consistent Val F1s and have to be trained for less epochs to reach convergence. These results might change if the kernels are run more times and non-determinism might not give more stable results when using tensorflow or keras.</p>",
      "rawMarkdown": "Hi,\n\nI made a [kernel] [1] showing how making a pytorch kernel more deterministic affects the result of the kernel. I ran the kernel 3 times with and without determinism. Versions 4-6 are with determinism and versions 7-9 are without determinism.\n\n### Model:\nI use a simple 2 layer bidirectional LSTM with an attention layer and a batch size of 512\n\n### Training:\nI train for 7 epochs and choose the saved weights of the model with the lowest validation loss\n\n### Determinism:\n\nFor my determinism experiments I set torch's random seeds and set deterministic=True:\n\n    SEED = 1337\n    np.random.seed(SEED)\n    torch.manual_seed(SEED)\n    torch.cuda.manual_seed(SEED)\n    torch.backends.cudnn.deterministic = True\n\nFor my non-deterministic experiments I only set numpy's random seeds:\n\n    SEED = 1337\n    np.random.seed(SEED)\n\n### Results:\n\n| Deterministic? | Train F1 | Val F1 --| Train Loss | Val Loss | Best Epoch | Best threshold \n\n| Yes -------------|  0.781 -- | 0.6696 | 0.0842 ----| 0.1000 --| 5 ------------| 0.37\n\n| Yes -------------|  0.7684 -| 0.6773 | 0.0840 ----| 0.0992 --| 5 ------------| 0.33\n\n| Yes -------------| 0.7784 --| 0.6701 | 0.0957 -----| 0.1000 --| 2 ------------| 0.37\n\n----------\n\n| No ------------- | 0.7963 --| 0.6718 | 0.0930 -----| 0.0976 --| 2 ------------| 0.32\n\n| No ------------- | 0.7928 --| 0.6724 | 0.0930 ----| 0.0982 --| 2 ------------| 0.3\n\n| No ------------- | 0.7997 --| 0.6751 -| 0.0868 ----| 0.0974 --| 3 ------------| 0.35\n\nNon-deterministic runs have higher and more consistent Val F1s and have to be trained for less epochs to reach convergence. These results might change if the kernels are run more times and non-determinism might not give more stable results when using tensorflow or keras.\n\n\n[1]: https://www.kaggle.com/bkkaggle/pytorch-determinism-test",
      "votes": null
    },
    {
      "id": "428476",
      "postDate": "11/27/2018 10:48:10",
      "content": "<p>I've got very similar results, thanks for sharing this <a href=\"/bkkaggle\">@bkkaggle</a></p>",
      "rawMarkdown": "I've got very similar results, thanks for sharing this @bkkaggle",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 428476,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "11/27/2018 10:48:10",
      "content": "<p>I've got very similar results, thanks for sharing this <a href=\"/bkkaggle\">@bkkaggle</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "428303": "Hi,\n\nI made a [kernel] [1] showing how making a pytorch kernel more deterministic affects the result of the kernel. I ran the kernel 3 times with and without determinism. Versions 4-6 are with determinism and versions 7-9 are without determinism.\n\n### Model:\nI use a simple 2 layer bidirectional LSTM with an attention layer and a batch size of 512\n\n### Training:\nI train for 7 epochs and choose the saved weights of the model with the lowest validation loss\n\n### Determinism:\n\nFor my determinism experiments I set torch's random seeds and set deterministic=True:\n\n    SEED = 1337\n    np.random.seed(SEED)\n    torch.manual_seed(SEED)\n    torch.cuda.manual_seed(SEED)\n    torch.backends.cudnn.deterministic = True\n\nFor my non-deterministic experiments I only set numpy's random seeds:\n\n    SEED = 1337\n    np.random.seed(SEED)\n\n### Results:\n\n| Deterministic? | Train F1 | Val F1 --| Train Loss | Val Loss | Best Epoch | Best threshold \n\n| Yes -------------|  0.781 -- | 0.6696 | 0.0842 ----| 0.1000 --| 5 ------------| 0.37\n\n| Yes -------------|  0.7684 -| 0.6773 | 0.0840 ----| 0.0992 --| 5 ------------| 0.33\n\n| Yes -------------| 0.7784 --| 0.6701 | 0.0957 -----| 0.1000 --| 2 ------------| 0.37\n\n----------\n\n| No ------------- | 0.7963 --| 0.6718 | 0.0930 -----| 0.0976 --| 2 ------------| 0.32\n\n| No ------------- | 0.7928 --| 0.6724 | 0.0930 ----| 0.0982 --| 2 ------------| 0.3\n\n| No ------------- | 0.7997 --| 0.6751 -| 0.0868 ----| 0.0974 --| 3 ------------| 0.35\n\nNon-deterministic runs have higher and more consistent Val F1s and have to be trained for less epochs to reach convergence. These results might change if the kernels are run more times and non-determinism might not give more stable results when using tensorflow or keras.\n\n\n[1]: https://www.kaggle.com/bkkaggle/pytorch-determinism-test",
    "428476": "I've got very similar results, thanks for sharing this @bkkaggle"
  },
  "source": "meta"
}