{
  "id": 231961,
  "title": "Pytorch + Numpy bug that you should be aware. ",
  "url": "/competitions/bms-molecular-translation/discussion/231961",
  "author_name": "",
  "post_date": "2021-04-11T13:02:04.085852400Z",
  "votes": 26,
  "comment_count": 7,
  "views": 0,
  "content": "<p>If you are using Pytorch <code>Dataset</code> and have inside a function that uses random number generated please consider adding following modification to <code>Dataloader</code> (it has shown that <code>random</code> generated is not random and the number always repeat) </p>\n<pre><code>def worker_init_fn(worker_id):                                                          \n    np.random.seed(np.random.get_state()[1][0] + worker_id)\n\ndataset = RandomDataset()\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4, \n                        worker_init_fn=worker_init_fn)\n\nfor batch in dataloader:\n    print(batch)\n</code></pre>\n<p>credits:<br>\n<a href=\"https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\" target=\"_blank\">https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/</a><br>\n<a href=\"https://twitter.com/karpathy/status/1381121537164513288?s=20\" target=\"_blank\">https://twitter.com/karpathy/status/1381121537164513288?s=20</a></p>",
  "messages": [
    {
      "id": "1270234",
      "postDate": "04/11/2021 13:02:04",
      "content": "<p>If you are using Pytorch <code>Dataset</code> and have inside a function that uses random number generated please consider adding following modification to <code>Dataloader</code> (it has shown that <code>random</code> generated is not random and the number always repeat) </p>\n<pre><code>def worker_init_fn(worker_id):                                                          \n    np.random.seed(np.random.get_state()[1][0] + worker_id)\n\ndataset = RandomDataset()\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4, \n                        worker_init_fn=worker_init_fn)\n\nfor batch in dataloader:\n    print(batch)\n</code></pre>\n<p>credits:<br>\n<a href=\"https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\" target=\"_blank\">https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/</a><br>\n<a href=\"https://twitter.com/karpathy/status/1381121537164513288?s=20\" target=\"_blank\">https://twitter.com/karpathy/status/1381121537164513288?s=20</a></p>",
      "rawMarkdown": "If you are using Pytorch `Dataset` and have inside a function that uses random number generated please consider adding following modification to `Dataloader` (it has shown that `random` generated is not random and the number always repeat) \n\n```\ndef worker_init_fn(worker_id):                                                          \n    np.random.seed(np.random.get_state()[1][0] + worker_id)\n\ndataset = RandomDataset()\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4, \n                        worker_init_fn=worker_init_fn)\n\nfor batch in dataloader:\n    print(batch)\n```\ncredits:\nhttps://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\nhttps://twitter.com/karpathy/status/1381121537164513288?s=20",
      "votes": null
    },
    {
      "id": "1270237",
      "postDate": "04/11/2021 13:05:38",
      "content": "<p>i face this problem before.<br>\nit is always important to write check function.</p>\n<p>it is actually not a bug. this is because np random function is not thread safe.</p>",
      "rawMarkdown": "i face this problem before.\nit is always important to write check function.\n\nit is actually not a bug. this is because np random function is not thread safe.",
      "votes": null
    },
    {
      "id": "1270240",
      "postDate": "04/11/2021 13:11:49",
      "content": "<p>I guess you are right its not a bug, but a feature which is probably not well-known and can cause a lot of issue and very hard to debug problems when overlooked.</p>",
      "rawMarkdown": "I guess you are right its not a bug, but a feature which is probably not well-known and can cause a lot of issue and very hard to debug problems when overlooked.",
      "votes": null
    },
    {
      "id": "1270400",
      "postDate": "04/11/2021 16:01:58",
      "content": "<p>I was amazed to see this today; so easy to miss it even if you do look at the data! Emphasizes how important it is to carefully inspect what goes into the model from one batch to another. </p>",
      "rawMarkdown": "I was amazed to see this today; so easy to miss it even if you do look at the data! Emphasizes how important it is to carefully inspect what goes into the model from one batch to another.",
      "votes": null
    },
    {
      "id": "1271254",
      "postDate": "04/12/2021 12:37:07",
      "content": "<p>here is more information<br>\n<a href=\"https://www.reddit.com/r/MachineLearning/comments/mocpgj/p_using_pytorch_numpy_a_bug_that_plagues/\" target=\"_blank\">https://www.reddit.com/r/MachineLearning/comments/mocpgj/p_using_pytorch_numpy_a_bug_that_plagues/</a></p>\n<p>Using NumPy’s random number generator with multi-process data loading in PyTorch causes identical augmentations unless you specifically set seeds using the worker_init_fn option in the DataLoader. I didn’t and this bug silently regressed my model’s accuracy.</p>\n<p>\". Out of these, over 95% of the repositories are plagued by this problem. It’s inside PyTorch's official tutorial, OpenAI’s code, and NVIDIA’s projects. Even Karpathy admitted falling prey to it.\"</p>\n<p><a href=\"https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\" target=\"_blank\">https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/</a></p>",
      "rawMarkdown": "here is more information\nhttps://www.reddit.com/r/MachineLearning/comments/mocpgj/p_using_pytorch_numpy_a_bug_that_plagues/\n\nUsing NumPy’s random number generator with multi-process data loading in PyTorch causes identical augmentations unless you specifically set seeds using the worker_init_fn option in the DataLoader. I didn’t and this bug silently regressed my model’s accuracy.\n\n\". Out of these, over 95% of the repositories are plagued by this problem. It’s inside PyTorch's official tutorial, OpenAI’s code, and NVIDIA’s projects. Even Karpathy admitted falling prey to it.\"\n\nhttps://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/",
      "votes": null
    },
    {
      "id": "1273535",
      "postDate": "04/14/2021 12:56:07",
      "content": "<p>Wow, thanks for that..I only recently started using Pytorch and just yesterday implemented exactly that.</p>",
      "rawMarkdown": "Wow, thanks for that..I only recently started using Pytorch and just yesterday implemented exactly that.",
      "votes": null
    },
    {
      "id": "1274498",
      "postDate": "04/15/2021 11:08:16",
      "content": "<p>terrified*</p>",
      "rawMarkdown": "terrified*",
      "votes": null
    },
    {
      "id": "1275822",
      "postDate": "04/16/2021 18:24:45",
      "content": "<p>I read this on Reddit when they posted it, but it isn't clear to me yet if this happen even if you dont set the seed manually before instantiating the Dataloader.</p>",
      "rawMarkdown": "I read this on Reddit when they posted it, but it isn't clear to me yet if this happen even if you dont set the seed manually before instantiating the Dataloader.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1270237,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/11/2021 13:05:38",
      "content": "<p>i face this problem before.<br>\nit is always important to write check function.</p>\n<p>it is actually not a bug. this is because np random function is not thread safe.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1270240,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/11/2021 13:11:49",
          "content": "<p>I guess you are right its not a bug, but a feature which is probably not well-known and can cause a lot of issue and very hard to debug problems when overlooked.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1270400,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "04/11/2021 16:01:58",
      "content": "<p>I was amazed to see this today; so easy to miss it even if you do look at the data! Emphasizes how important it is to carefully inspect what goes into the model from one batch to another. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1274498,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/15/2021 11:08:16",
          "content": "<p>terrified*</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1271254,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/12/2021 12:37:07",
      "content": "<p>here is more information<br>\n<a href=\"https://www.reddit.com/r/MachineLearning/comments/mocpgj/p_using_pytorch_numpy_a_bug_that_plagues/\" target=\"_blank\">https://www.reddit.com/r/MachineLearning/comments/mocpgj/p_using_pytorch_numpy_a_bug_that_plagues/</a></p>\n<p>Using NumPy’s random number generator with multi-process data loading in PyTorch causes identical augmentations unless you specifically set seeds using the worker_init_fn option in the DataLoader. I didn’t and this bug silently regressed my model’s accuracy.</p>\n<p>\". Out of these, over 95% of the repositories are plagued by this problem. It’s inside PyTorch's official tutorial, OpenAI’s code, and NVIDIA’s projects. Even Karpathy admitted falling prey to it.\"</p>\n<p><a href=\"https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\" target=\"_blank\">https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1273535,
      "author_name": "kalfasyan",
      "author_url": "",
      "post_date": "04/14/2021 12:56:07",
      "content": "<p>Wow, thanks for that..I only recently started using Pytorch and just yesterday implemented exactly that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1275822,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "04/16/2021 18:24:45",
      "content": "<p>I read this on Reddit when they posted it, but it isn't clear to me yet if this happen even if you dont set the seed manually before instantiating the Dataloader.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1270234": "If you are using Pytorch `Dataset` and have inside a function that uses random number generated please consider adding following modification to `Dataloader` (it has shown that `random` generated is not random and the number always repeat) \n\n```\ndef worker_init_fn(worker_id):                                                          \n    np.random.seed(np.random.get_state()[1][0] + worker_id)\n\ndataset = RandomDataset()\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4, \n                        worker_init_fn=worker_init_fn)\n\nfor batch in dataloader:\n    print(batch)\n```\ncredits:\nhttps://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\nhttps://twitter.com/karpathy/status/1381121537164513288?s=20",
    "1270237": "i face this problem before.\nit is always important to write check function.\n\nit is actually not a bug. this is because np random function is not thread safe.",
    "1270240": "I guess you are right its not a bug, but a feature which is probably not well-known and can cause a lot of issue and very hard to debug problems when overlooked.",
    "1270400": "I was amazed to see this today; so easy to miss it even if you do look at the data! Emphasizes how important it is to carefully inspect what goes into the model from one batch to another.",
    "1271254": "here is more information\nhttps://www.reddit.com/r/MachineLearning/comments/mocpgj/p_using_pytorch_numpy_a_bug_that_plagues/\n\nUsing NumPy’s random number generator with multi-process data loading in PyTorch causes identical augmentations unless you specifically set seeds using the worker_init_fn option in the DataLoader. I didn’t and this bug silently regressed my model’s accuracy.\n\n\". Out of these, over 95% of the repositories are plagued by this problem. It’s inside PyTorch's official tutorial, OpenAI’s code, and NVIDIA’s projects. Even Karpathy admitted falling prey to it.\"\n\nhttps://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/",
    "1273535": "Wow, thanks for that..I only recently started using Pytorch and just yesterday implemented exactly that.",
    "1274498": "terrified*",
    "1275822": "I read this on Reddit when they posted it, but it isn't clear to me yet if this happen even if you dont set the seed manually before instantiating the Dataloader."
  },
  "source": "meta"
}