{
  "id": 71496,
  "title": "Reproducibility and the winner obligations",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/71496",
  "author_name": "",
  "post_date": "2018-11-14T06:31:22.298585600Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Not that I need to worry about this but I am curious. The rules say that the winner must hand over software that can reproduce the winning submission. I am using fastai(newbie) and in that package, as far as I know, it is not possible to reproduce a run in detail. I have tried all the suggestions I have found online including specifying numpy, pytorch and scikit random number seeds and get various levels of reproducibility but it always goes off the rails somewhere during the learning. It does work for \"precompute=True\" but that is hardly a solution.</p>\n\n<p>This can't be a real issue because there are, like, winners.</p>\n\n<p>How do the winners deal with this? Do they achieve detailed reproducibility?</p>",
  "messages": [
    {
      "id": "420804",
      "postDate": "11/14/2018 06:31:22",
      "content": "<p>Not that I need to worry about this but I am curious. The rules say that the winner must hand over software that can reproduce the winning submission. I am using fastai(newbie) and in that package, as far as I know, it is not possible to reproduce a run in detail. I have tried all the suggestions I have found online including specifying numpy, pytorch and scikit random number seeds and get various levels of reproducibility but it always goes off the rails somewhere during the learning. It does work for \"precompute=True\" but that is hardly a solution.</p>\n\n<p>This can't be a real issue because there are, like, winners.</p>\n\n<p>How do the winners deal with this? Do they achieve detailed reproducibility?</p>",
      "rawMarkdown": "Not that I need to worry about this but I am curious. The rules say that the winner must hand over software that can reproduce the winning submission. I am using fastai(newbie) and in that package, as far as I know, it is not possible to reproduce a run in detail. I have tried all the suggestions I have found online including specifying numpy, pytorch and scikit random number seeds and get various levels of reproducibility but it always goes off the rails somewhere during the learning. It does work for \"precompute=True\" but that is hardly a solution.\n\nThis can't be a real issue because there are, like, winners.\n\nHow do the winners deal with this? Do they achieve detailed reproducibility?",
      "votes": null
    },
    {
      "id": "420840",
      "postDate": "11/14/2018 07:50:58",
      "content": "<p>I think the rules mean inference, as in “you provide weights and model code that can be run and produce the same submission csv”. I don’t think the expectation is they should be able to train a model that produces the exact same outputs, as that would be nearly impossible.</p>",
      "rawMarkdown": "I think the rules mean inference, as in “you provide weights and model code that can be run and produce the same submission csv”. I don’t think the expectation is they should be able to train a model that produces the exact same outputs, as that would be nearly impossible.",
      "votes": null
    },
    {
      "id": "421028",
      "postDate": "11/14/2018 13:48:25",
      "content": "<p>I was also worried about that (before I learned that I'm faaar away from a probable winner haha)    </p>\n\n<p>But some people have written that some GPU computations are not deterministic. (I can't really deny or agree)</p>",
      "rawMarkdown": "I was also worried about that (before I learned that I'm faaar away from a probable winner haha)    \n\nBut some people have written that some GPU computations are not deterministic. (I can't really deny or agree)",
      "votes": null
    },
    {
      "id": "421141",
      "postDate": "11/14/2018 16:29:09",
      "content": "<p>I disagree with your GPU point. My knowledge on GPU is a little dated, but I used to write GPU hardware test that would heavily load the GPU for a fixed time and then do a checksum on the GPU memory at the end. Hardware and driver changes would change this behavior, but across thousands of machines with the same software, the result was deterministic. Using raw math operations instead of rendering graphics I would expect to be even more deterministic.</p>\n\n<p>I'm certain that the reproduction issues people are seeing can be attributed to variations in random generators. There has to be some parameter that is not being initialized the same. Maybe fast.ai isn't setting a seed, or there is another seed parameter that needs to be set somewhere.</p>",
      "rawMarkdown": "I disagree with your GPU point. My knowledge on GPU is a little dated, but I used to write GPU hardware test that would heavily load the GPU for a fixed time and then do a checksum on the GPU memory at the end. Hardware and driver changes would change this behavior, but across thousands of machines with the same software, the result was deterministic. Using raw math operations instead of rendering graphics I would expect to be even more deterministic.\n\nI'm certain that the reproduction issues people are seeing can be attributed to variations in random generators. There has to be some parameter that is not being initialized the same. Maybe fast.ai isn't setting a seed, or there is another seed parameter that needs to be set somewhere.",
      "votes": null
    },
    {
      "id": "421210",
      "postDate": "11/14/2018 18:14:51",
      "content": "<p>That’s not quite right, as some cuDNN operations are non deterministic. See the docs here <a href=\"https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility\">https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility</a> It affects certain convolution and pooling operations. If you use those operations, there’s no way to do a fully reproduceable training run on the GPU.</p>",
      "rawMarkdown": "That’s not quite right, as some cuDNN operations are non deterministic. See the docs here https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility It affects certain convolution and pooling operations. If you use those operations, there’s no way to do a fully reproduceable training run on the GPU.",
      "votes": null
    },
    {
      "id": "421219",
      "postDate": "11/14/2018 18:32:59",
      "content": "<p>Thanks for that link, definitely good to know. My expectation was wrong, I agree if you use these operations it is not perfectly reproducible.</p>",
      "rawMarkdown": "Thanks for that link, definitely good to know. My expectation was wrong, I agree if you use these operations it is not perfectly reproducible.",
      "votes": null
    },
    {
      "id": "427043",
      "postDate": "11/24/2018 12:10:38",
      "content": "<p>I was also looking into this problem concerning reproducibility, especially in fastai. Even if you put the same seed on random/numpy/pytorch/etc.. there will still be a problem with augmentation. If you look at <a href=\"https://github.com/fastai/fastai/blob/master/old/fastai/dataloader.py\">https://github.com/fastai/fastai/blob/master/old/fastai/dataloader.py</a>, there is a following code:</p>\n\n<pre><code>        with ThreadPoolExecutor(max_workers=self.num_workers) as e:\n            # avoid py3.6 issue where queue is infinite and can result in memory exhaustion\n            for c in chunk_iter(iter(self.batch_sampler), self.num_workers*10):\n                for batch in e.map(self.get_batch, c):\n                    yield get_tensor(batch, self.pin_memory, self.half)\n</code></pre>\n\n<p>When you use threads for data loading, the augmentation for each image is done inside different threads. So, even if you have set a random seed before, threads (since they share resources) will update the state of the random and share this state as they perform augmentations. <br>\nAs I understand, the solution would be to change to ProcessPoolExecutor, but it would require to change code too much (because some objects are not pickable currently). <br>\n(I am not sure if the same exists in the new v1 version of fastai). <br>\nFor me it would also make more sense to perform augmentation in separate processes, rather than threads, because of GIL, or?</p>",
      "rawMarkdown": "I was also looking into this problem concerning reproducibility, especially in fastai. Even if you put the same seed on random/numpy/pytorch/etc.. there will still be a problem with augmentation. If you look at https://github.com/fastai/fastai/blob/master/old/fastai/dataloader.py, there is a following code:\n\n            with ThreadPoolExecutor(max_workers=self.num_workers) as e:\n                # avoid py3.6 issue where queue is infinite and can result in memory exhaustion\n                for c in chunk_iter(iter(self.batch_sampler), self.num_workers*10):\n                    for batch in e.map(self.get_batch, c):\n                        yield get_tensor(batch, self.pin_memory, self.half)\n\nWhen you use threads for data loading, the augmentation for each image is done inside different threads. So, even if you have set a random seed before, threads (since they share resources) will update the state of the random and share this state as they perform augmentations.  \nAs I understand, the solution would be to change to ProcessPoolExecutor, but it would require to change code too much (because some objects are not pickable currently).  \n(I am not sure if the same exists in the new v1 version of fastai).  \nFor me it would also make more sense to perform augmentation in separate processes, rather than threads, because of GIL, or?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 420840,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "11/14/2018 07:50:58",
      "content": "<p>I think the rules mean inference, as in “you provide weights and model code that can be run and produce the same submission csv”. I don’t think the expectation is they should be able to train a model that produces the exact same outputs, as that would be nearly impossible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 421028,
      "author_name": "danmoller",
      "author_url": "",
      "post_date": "11/14/2018 13:48:25",
      "content": "<p>I was also worried about that (before I learned that I'm faaar away from a probable winner haha)    </p>\n\n<p>But some people have written that some GPU computations are not deterministic. (I can't really deny or agree)</p>",
      "votes": null,
      "replies": [
        {
          "id": 421141,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/14/2018 16:29:09",
          "content": "<p>I disagree with your GPU point. My knowledge on GPU is a little dated, but I used to write GPU hardware test that would heavily load the GPU for a fixed time and then do a checksum on the GPU memory at the end. Hardware and driver changes would change this behavior, but across thousands of machines with the same software, the result was deterministic. Using raw math operations instead of rendering graphics I would expect to be even more deterministic.</p>\n\n<p>I'm certain that the reproduction issues people are seeing can be attributed to variations in random generators. There has to be some parameter that is not being initialized the same. Maybe fast.ai isn't setting a seed, or there is another seed parameter that needs to be set somewhere.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 421210,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "11/14/2018 18:14:51",
          "content": "<p>That’s not quite right, as some cuDNN operations are non deterministic. See the docs here <a href=\"https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility\">https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility</a> It affects certain convolution and pooling operations. If you use those operations, there’s no way to do a fully reproduceable training run on the GPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 421219,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/14/2018 18:32:59",
          "content": "<p>Thanks for that link, definitely good to know. My expectation was wrong, I agree if you use these operations it is not perfectly reproducible.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 427043,
      "author_name": "bi0max",
      "author_url": "",
      "post_date": "11/24/2018 12:10:38",
      "content": "<p>I was also looking into this problem concerning reproducibility, especially in fastai. Even if you put the same seed on random/numpy/pytorch/etc.. there will still be a problem with augmentation. If you look at <a href=\"https://github.com/fastai/fastai/blob/master/old/fastai/dataloader.py\">https://github.com/fastai/fastai/blob/master/old/fastai/dataloader.py</a>, there is a following code:</p>\n\n<pre><code>        with ThreadPoolExecutor(max_workers=self.num_workers) as e:\n            # avoid py3.6 issue where queue is infinite and can result in memory exhaustion\n            for c in chunk_iter(iter(self.batch_sampler), self.num_workers*10):\n                for batch in e.map(self.get_batch, c):\n                    yield get_tensor(batch, self.pin_memory, self.half)\n</code></pre>\n\n<p>When you use threads for data loading, the augmentation for each image is done inside different threads. So, even if you have set a random seed before, threads (since they share resources) will update the state of the random and share this state as they perform augmentations. <br>\nAs I understand, the solution would be to change to ProcessPoolExecutor, but it would require to change code too much (because some objects are not pickable currently). <br>\n(I am not sure if the same exists in the new v1 version of fastai). <br>\nFor me it would also make more sense to perform augmentation in separate processes, rather than threads, because of GIL, or?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "420804": "Not that I need to worry about this but I am curious. The rules say that the winner must hand over software that can reproduce the winning submission. I am using fastai(newbie) and in that package, as far as I know, it is not possible to reproduce a run in detail. I have tried all the suggestions I have found online including specifying numpy, pytorch and scikit random number seeds and get various levels of reproducibility but it always goes off the rails somewhere during the learning. It does work for \"precompute=True\" but that is hardly a solution.\n\nThis can't be a real issue because there are, like, winners.\n\nHow do the winners deal with this? Do they achieve detailed reproducibility?",
    "420840": "I think the rules mean inference, as in “you provide weights and model code that can be run and produce the same submission csv”. I don’t think the expectation is they should be able to train a model that produces the exact same outputs, as that would be nearly impossible.",
    "421028": "I was also worried about that (before I learned that I'm faaar away from a probable winner haha)    \n\nBut some people have written that some GPU computations are not deterministic. (I can't really deny or agree)",
    "421141": "I disagree with your GPU point. My knowledge on GPU is a little dated, but I used to write GPU hardware test that would heavily load the GPU for a fixed time and then do a checksum on the GPU memory at the end. Hardware and driver changes would change this behavior, but across thousands of machines with the same software, the result was deterministic. Using raw math operations instead of rendering graphics I would expect to be even more deterministic.\n\nI'm certain that the reproduction issues people are seeing can be attributed to variations in random generators. There has to be some parameter that is not being initialized the same. Maybe fast.ai isn't setting a seed, or there is another seed parameter that needs to be set somewhere.",
    "421210": "That’s not quite right, as some cuDNN operations are non deterministic. See the docs here https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility It affects certain convolution and pooling operations. If you use those operations, there’s no way to do a fully reproduceable training run on the GPU.",
    "421219": "Thanks for that link, definitely good to know. My expectation was wrong, I agree if you use these operations it is not perfectly reproducible.",
    "427043": "I was also looking into this problem concerning reproducibility, especially in fastai. Even if you put the same seed on random/numpy/pytorch/etc.. there will still be a problem with augmentation. If you look at https://github.com/fastai/fastai/blob/master/old/fastai/dataloader.py, there is a following code:\n\n            with ThreadPoolExecutor(max_workers=self.num_workers) as e:\n                # avoid py3.6 issue where queue is infinite and can result in memory exhaustion\n                for c in chunk_iter(iter(self.batch_sampler), self.num_workers*10):\n                    for batch in e.map(self.get_batch, c):\n                        yield get_tensor(batch, self.pin_memory, self.half)\n\nWhen you use threads for data loading, the augmentation for each image is done inside different threads. So, even if you have set a random seed before, threads (since they share resources) will update the state of the random and share this state as they perform augmentations.  \nAs I understand, the solution would be to change to ProcessPoolExecutor, but it would require to change code too much (because some objects are not pickable currently).  \n(I am not sure if the same exists in the new v1 version of fastai).  \nFor me it would also make more sense to perform augmentation in separate processes, rather than threads, because of GIL, or?"
  },
  "source": "meta"
}