{
  "id": 58765,
  "title": "unpredictable keras and unreproducible results",
  "url": "/competitions/avito-demand-prediction/discussion/58765",
  "author_name": "Oleg Yaroshevskiy",
  "post_date": "2018-06-13T12:08:53.223000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I just wonder how you guys deal with random states and reproducing results in Keras. I can't reproduce any single keras run both on my machine or kaggle's one. The validation diff between two same-model-runs might be up to 0.006X (not keras fake numbers but my own val evaluation)</p>\n\n<p>The following code snippet doesn't work for me:</p>\n\n<pre><code>import numpy as np\nimport tensorflow as tf\nimport random as rn\n\nimport os\nos.environ['PYTHONHASHSEED'] = '0'\nnp.random.seed(42)\n\nrn.seed(42)\ntf.set_random_seed(42)\n\nsession_conf = tf.ConfigProto(intra_op_parallelism_threads=1, inter_op_parallelism_threads=1)\nfrom keras import backend as K\nsess = tf.Session(graph=tf.get_default_graph(), config=session_conf)\nK.set_session(sess)\n</code></pre>\n\n<p>How do you deal with that?</p>",
  "messages": [
    {
      "id": 342487,
      "postDate": "2018-06-13T15:28:39.160Z",
      "content": "<p>I have had that battle. <a href=\"https://twitter.com/fchollet/status/1003718133721387008\">This</a> from Chollet a few days ago gives context:</p>\n\n<p><em>If you're doing a Kaggle competition and you're evaluating your models / ideas according to a fixed validation split of the training data (plus the public leaderboard), you will consistently underperform on the private leaderboard. The same is true in research at large. Here is a very simple recommendation to help you overcome this: use a higher-entropy validation process, such as k-fold validation, or even better, iterated k-fold validation with shuffling. Only check your results on the official validation set at the very end. Yes, it's more expensive, but that cost itself is a regularization factor: it will force you to try fewer ideas instead of throwing spaghetti at the wall and seeing what sticks.</em></p>\n\n<p>I embraced it in the end and fit all NNs on 5-fold CV and average over several epoch's predictions.</p>",
      "rawMarkdown": "I have had that battle. [This][1] from Chollet a few days ago gives context:\n\n*If you're doing a Kaggle competition and you're evaluating your models / ideas according to a fixed validation split of the training data (plus the public leaderboard), you will consistently underperform on the private leaderboard. The same is true in research at large. Here is a very simple recommendation to help you overcome this: use a higher-entropy validation process, such as k-fold validation, or even better, iterated k-fold validation with shuffling. Only check your results on the official validation set at the very end. Yes, it's more expensive, but that cost itself is a regularization factor: it will force you to try fewer ideas instead of throwing spaghetti at the wall and seeing what sticks.*\n\nI embraced it in the end and fit all NNs on 5-fold CV and average over several epoch's predictions.\n\n  [1]: https://twitter.com/fchollet/status/1003718133721387008",
      "votes": 4
    },
    {
      "id": 342376,
      "postDate": "2018-06-13T12:08:53.223Z",
      "content": "<p>I just wonder how you guys deal with random states and reproducing results in Keras. I can't reproduce any single keras run both on my machine or kaggle's one. The validation diff between two same-model-runs might be up to 0.006X (not keras fake numbers but my own val evaluation)</p>\n\n<p>The following code snippet doesn't work for me:</p>\n\n<pre><code>import numpy as np\nimport tensorflow as tf\nimport random as rn\n\nimport os\nos.environ['PYTHONHASHSEED'] = '0'\nnp.random.seed(42)\n\nrn.seed(42)\ntf.set_random_seed(42)\n\nsession_conf = tf.ConfigProto(intra_op_parallelism_threads=1, inter_op_parallelism_threads=1)\nfrom keras import backend as K\nsess = tf.Session(graph=tf.get_default_graph(), config=session_conf)\nK.set_session(sess)\n</code></pre>\n\n<p>How do you deal with that?</p>",
      "rawMarkdown": "I just wonder how you guys deal with random states and reproducing results in Keras. I can't reproduce any single keras run both on my machine or kaggle's one. The validation diff between two same-model-runs might be up to 0.006X (not keras fake numbers but my own val evaluation)\n\nThe following code snippet doesn't work for me:\n\n    import numpy as np\n    import tensorflow as tf\n    import random as rn\n    \n    import os\n    os.environ['PYTHONHASHSEED'] = '0'\n    np.random.seed(42)\n    \n    rn.seed(42)\n    tf.set_random_seed(42)\n\n    session_conf = tf.ConfigProto(intra_op_parallelism_threads=1, inter_op_parallelism_threads=1)\n    from keras import backend as K\n    sess = tf.Session(graph=tf.get_default_graph(), config=session_conf)\n    K.set_session(sess)\n\nHow do you deal with that?",
      "votes": 2
    },
    {
      "id": 342591,
      "postDate": "2018-06-13T18:59:17.443Z",
      "content": "<p>If I recall correctly, you need to set your random seed after your final Keras import. Something in the import messes up the seeds. Move <code>from keras import backend as K</code> up with the other imports.</p>",
      "rawMarkdown": "If I recall correctly, you need to set your random seed after your final Keras import. Something in the import messes up the seeds. Move `from keras import backend as K` up with the other imports.",
      "replies": [
        {
          "id": 342880,
          "postDate": "2018-06-14T10:10:26.600Z",
          "content": "<p>What do you mean? I moved all keras import below and above seeds but that didn't help</p>",
          "rawMarkdown": "What do you mean? I moved all keras import below and above seeds but that didn't help"
        },
        {
          "id": 343021,
          "postDate": "2018-06-14T15:13:59.310Z",
          "content": "<p>Ah, that was the only thing I had to do when this happened to me. Sorry I can't be of more help!</p>",
          "rawMarkdown": "Ah, that was the only thing I had to do when this happened to me. Sorry I can't be of more help!"
        }
      ]
    },
    {
      "id": 342542,
      "postDate": "2018-06-13T17:14:04.083Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 342487,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2018-06-13T15:28:39.160000",
      "content": "<p>I have had that battle. <a href=\"https://twitter.com/fchollet/status/1003718133721387008\">This</a> from Chollet a few days ago gives context:</p>\n\n<p><em>If you're doing a Kaggle competition and you're evaluating your models / ideas according to a fixed validation split of the training data (plus the public leaderboard), you will consistently underperform on the private leaderboard. The same is true in research at large. Here is a very simple recommendation to help you overcome this: use a higher-entropy validation process, such as k-fold validation, or even better, iterated k-fold validation with shuffling. Only check your results on the official validation set at the very end. Yes, it's more expensive, but that cost itself is a regularization factor: it will force you to try fewer ideas instead of throwing spaghetti at the wall and seeing what sticks.</em></p>\n\n<p>I embraced it in the end and fit all NNs on 5-fold CV and average over several epoch's predictions.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 342591,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2018-06-13T18:59:17.443000",
      "content": "<p>If I recall correctly, you need to set your random seed after your final Keras import. Something in the import messes up the seeds. Move <code>from keras import backend as K</code> up with the other imports.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 342880,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2018-06-14T10:10:26.600000",
          "content": "<p>What do you mean? I moved all keras import below and above seeds but that didn't help</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343021,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-06-14T15:13:59.310000",
          "content": "<p>Ah, that was the only thing I had to do when this happened to me. Sorry I can't be of more help!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 342542,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-13T17:14:04.083000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "342487": "I have had that battle. [This][1] from Chollet a few days ago gives context:\n\n*If you're doing a Kaggle competition and you're evaluating your models / ideas according to a fixed validation split of the training data (plus the public leaderboard), you will consistently underperform on the private leaderboard. The same is true in research at large. Here is a very simple recommendation to help you overcome this: use a higher-entropy validation process, such as k-fold validation, or even better, iterated k-fold validation with shuffling. Only check your results on the official validation set at the very end. Yes, it's more expensive, but that cost itself is a regularization factor: it will force you to try fewer ideas instead of throwing spaghetti at the wall and seeing what sticks.*\n\nI embraced it in the end and fit all NNs on 5-fold CV and average over several epoch's predictions.\n\n  [1]: https://twitter.com/fchollet/status/1003718133721387008",
    "342376": "I just wonder how you guys deal with random states and reproducing results in Keras. I can't reproduce any single keras run both on my machine or kaggle's one. The validation diff between two same-model-runs might be up to 0.006X (not keras fake numbers but my own val evaluation)\n\nThe following code snippet doesn't work for me:\n\n    import numpy as np\n    import tensorflow as tf\n    import random as rn\n    \n    import os\n    os.environ['PYTHONHASHSEED'] = '0'\n    np.random.seed(42)\n    \n    rn.seed(42)\n    tf.set_random_seed(42)\n\n    session_conf = tf.ConfigProto(intra_op_parallelism_threads=1, inter_op_parallelism_threads=1)\n    from keras import backend as K\n    sess = tf.Session(graph=tf.get_default_graph(), config=session_conf)\n    K.set_session(sess)\n\nHow do you deal with that?",
    "342591": "If I recall correctly, you need to set your random seed after your final Keras import. Something in the import messes up the seeds. Move `from keras import backend as K` up with the other imports.",
    "342542": ""
  }
}