{
  "id": 233359,
  "title": "Reminder of a bug that could happen when using Pytorch+Numpy",
  "url": "/competitions/birdclef-2021/discussion/233359",
  "author_name": "hide on bread",
  "post_date": "2021-04-18T18:47:50.149000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Be careful when using NumPy to generate random numbers (eg. np.random.randint) in the <strong>getitem</strong> method when doing data augmentation. </p>\n<p>A minimal  example taken from <a href=\"https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\" target=\"_blank\">https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/</a> :</p>\n<p>import numpy as np<br>\nfrom torch.utils.data import Dataset, DataLoader</p>\n<p>class RandomDataset(Dataset):</p>\n<pre><code>def __getitem__(self, index):\n    return np.random.randint(0, 1000, 3)\n\ndef __len__(self):\n    return 16\n</code></pre>\n<p>dataset = RandomDataset()<br>\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4)<br>\nfor batch in dataloader:<br>\n    print(batch)</p>\n<p>This will output:<br>\ntensor([[116, 760, 679],   # 1st batch, returned by process 0<br>\n        [754, 897, 764]])<br>\ntensor([[116, 760, 679],   # 2nd batch, returned by process 1<br>\n        [754, 897, 764]])<br>\ntensor([[116, 760, 679],   # 3rd batch, returned by process 2<br>\n        [754, 897, 764]])<br>\ntensor([[116, 760, 679],   # 4th batch, returned by process 3<br>\n        [754, 897, 764]])</p>\n<p>tensor([[866, 919, 441],   # 5th batch, returned by process 0<br>\n        [ 20, 727, 680]])<br>\ntensor([[866, 919, 441],   # 6th batch, returned by process 1<br>\n        [ 20, 727, 680]])<br>\ntensor([[866, 919, 441],   # 7th batch, returned by process 2<br>\n        [ 20, 727, 680]])<br>\ntensor([[866, 919, 441],   # 8th batch, returned by process 3<br>\n        [ 20, 727, 680]])</p>\n<p>An easy way to fix this minimal example problem is to use torch.randint(0, 1000, (3,)).<br>\nIn some cases, this issue will have minimal effect on the model, whereas in other cases, this issue will severely impact the model performance! It's always good to keep in mind of this buggy interaction that people often overlooked!</p>\n<p>Cheers!</p>",
  "messages": [
    {
      "id": 1277461,
      "postDate": "2021-04-18T18:47:50.150Z",
      "content": "<p>Be careful when using NumPy to generate random numbers (eg. np.random.randint) in the <strong>getitem</strong> method when doing data augmentation. </p>\n<p>A minimal  example taken from <a href=\"https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/\" target=\"_blank\">https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/</a> :</p>\n<p>import numpy as np<br>\nfrom torch.utils.data import Dataset, DataLoader</p>\n<p>class RandomDataset(Dataset):</p>\n<pre><code>def __getitem__(self, index):\n    return np.random.randint(0, 1000, 3)\n\ndef __len__(self):\n    return 16\n</code></pre>\n<p>dataset = RandomDataset()<br>\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4)<br>\nfor batch in dataloader:<br>\n    print(batch)</p>\n<p>This will output:<br>\ntensor([[116, 760, 679],   # 1st batch, returned by process 0<br>\n        [754, 897, 764]])<br>\ntensor([[116, 760, 679],   # 2nd batch, returned by process 1<br>\n        [754, 897, 764]])<br>\ntensor([[116, 760, 679],   # 3rd batch, returned by process 2<br>\n        [754, 897, 764]])<br>\ntensor([[116, 760, 679],   # 4th batch, returned by process 3<br>\n        [754, 897, 764]])</p>\n<p>tensor([[866, 919, 441],   # 5th batch, returned by process 0<br>\n        [ 20, 727, 680]])<br>\ntensor([[866, 919, 441],   # 6th batch, returned by process 1<br>\n        [ 20, 727, 680]])<br>\ntensor([[866, 919, 441],   # 7th batch, returned by process 2<br>\n        [ 20, 727, 680]])<br>\ntensor([[866, 919, 441],   # 8th batch, returned by process 3<br>\n        [ 20, 727, 680]])</p>\n<p>An easy way to fix this minimal example problem is to use torch.randint(0, 1000, (3,)).<br>\nIn some cases, this issue will have minimal effect on the model, whereas in other cases, this issue will severely impact the model performance! It's always good to keep in mind of this buggy interaction that people often overlooked!</p>\n<p>Cheers!</p>",
      "rawMarkdown": "Be careful when using NumPy to generate random numbers (eg. np.random.randint) in the __getitem__ method when doing data augmentation. \n\nA minimal  example taken from https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/ :\n\nimport numpy as np\nfrom torch.utils.data import Dataset, DataLoader\n\nclass RandomDataset(Dataset):\n\n    def __getitem__(self, index):\n        return np.random.randint(0, 1000, 3)\n\n    def __len__(self):\n        return 16\n    \ndataset = RandomDataset()\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4)\nfor batch in dataloader:\n    print(batch)\n\n\nThis will output:\ntensor([[116, 760, 679],   # 1st batch, returned by process 0\n        [754, 897, 764]])\ntensor([[116, 760, 679],   # 2nd batch, returned by process 1\n        [754, 897, 764]])\ntensor([[116, 760, 679],   # 3rd batch, returned by process 2\n        [754, 897, 764]])\ntensor([[116, 760, 679],   # 4th batch, returned by process 3\n        [754, 897, 764]])\n\ntensor([[866, 919, 441],   # 5th batch, returned by process 0\n        [ 20, 727, 680]])\ntensor([[866, 919, 441],   # 6th batch, returned by process 1\n        [ 20, 727, 680]])\ntensor([[866, 919, 441],   # 7th batch, returned by process 2\n        [ 20, 727, 680]])\ntensor([[866, 919, 441],   # 8th batch, returned by process 3\n        [ 20, 727, 680]])\n\nAn easy way to fix this minimal example problem is to use torch.randint(0, 1000, (3,)).\nIn some cases, this issue will have minimal effect on the model, whereas in other cases, this issue will severely impact the model performance! It's always good to keep in mind of this buggy interaction that people often overlooked!\n\nCheers!\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1277461": "Be careful when using NumPy to generate random numbers (eg. np.random.randint) in the __getitem__ method when doing data augmentation. \n\nA minimal  example taken from https://tanelp.github.io/posts/a-bug-that-plagues-thousands-of-open-source-ml-projects/ :\n\nimport numpy as np\nfrom torch.utils.data import Dataset, DataLoader\n\nclass RandomDataset(Dataset):\n\n    def __getitem__(self, index):\n        return np.random.randint(0, 1000, 3)\n\n    def __len__(self):\n        return 16\n    \ndataset = RandomDataset()\ndataloader = DataLoader(dataset, batch_size=2, num_workers=4)\nfor batch in dataloader:\n    print(batch)\n\n\nThis will output:\ntensor([[116, 760, 679],   # 1st batch, returned by process 0\n        [754, 897, 764]])\ntensor([[116, 760, 679],   # 2nd batch, returned by process 1\n        [754, 897, 764]])\ntensor([[116, 760, 679],   # 3rd batch, returned by process 2\n        [754, 897, 764]])\ntensor([[116, 760, 679],   # 4th batch, returned by process 3\n        [754, 897, 764]])\n\ntensor([[866, 919, 441],   # 5th batch, returned by process 0\n        [ 20, 727, 680]])\ntensor([[866, 919, 441],   # 6th batch, returned by process 1\n        [ 20, 727, 680]])\ntensor([[866, 919, 441],   # 7th batch, returned by process 2\n        [ 20, 727, 680]])\ntensor([[866, 919, 441],   # 8th batch, returned by process 3\n        [ 20, 727, 680]])\n\nAn easy way to fix this minimal example problem is to use torch.randint(0, 1000, (3,)).\nIn some cases, this issue will have minimal effect on the model, whereas in other cases, this issue will severely impact the model performance! It's always good to keep in mind of this buggy interaction that people often overlooked!\n\nCheers!\n"
  }
}