{
  "id": 100251,
  "title": "How to get deterministic results in Keras models",
  "url": "/competitions/aptos2019-blindness-detection/discussion/100251",
  "author_name": "",
  "post_date": "2019-07-17T12:06:26.523983800Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi folks,\nI have been playing around kernel based off <a href=\"https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter\">APTOS 2019: DenseNet Keras Starter</a> by <a href=\"https://www.kaggle.com/xhlulu\">xhlulu</a> and getting unstable CV scores +- 0.02 ish for each run.</p>\n\n<p>I have tried the following stuff (should be useful for Keras users), but still no luck.\nIf you have any ideas why it's still unstable please let me know!</p>\n\n<h3>seed_keras function</h3>\n\n<p><code>\ndef seed_keras(seed=1029):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    tf.set_random_seed(seed)\n    # hack: Even though we don't use torch, this can be used for Keras. \",\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code></p>\n\n<h3>random_transform</h3>\n\n<p><code>self.datagen.random_transform(X[i], seed=seed)</code></p>\n\n<h3>flow</h3>\n\n<p><code>create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed)</code></p>\n\n<h3>train_test_split</h3>\n\n<p><code>\nx_train, x_val, y_train, y_val = train_test_split(\n    x_train, y_train_multi, \n    test_size....\n    random_state=seed,\n)\n</code></p>",
  "messages": [
    {
      "id": "578155",
      "postDate": "07/17/2019 12:06:26",
      "content": "<p>Hi folks,\nI have been playing around kernel based off <a href=\"https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter\">APTOS 2019: DenseNet Keras Starter</a> by <a href=\"https://www.kaggle.com/xhlulu\">xhlulu</a> and getting unstable CV scores +- 0.02 ish for each run.</p>\n\n<p>I have tried the following stuff (should be useful for Keras users), but still no luck.\nIf you have any ideas why it's still unstable please let me know!</p>\n\n<h3>seed_keras function</h3>\n\n<p><code>\ndef seed_keras(seed=1029):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    tf.set_random_seed(seed)\n    # hack: Even though we don't use torch, this can be used for Keras. \",\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n</code></p>\n\n<h3>random_transform</h3>\n\n<p><code>self.datagen.random_transform(X[i], seed=seed)</code></p>\n\n<h3>flow</h3>\n\n<p><code>create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed)</code></p>\n\n<h3>train_test_split</h3>\n\n<p><code>\nx_train, x_val, y_train, y_val = train_test_split(\n    x_train, y_train_multi, \n    test_size....\n    random_state=seed,\n)\n</code></p>",
      "rawMarkdown": "Hi folks,\nI have been playing around kernel based off [APTOS 2019: DenseNet Keras Starter](https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter) by [xhlulu](https://www.kaggle.com/xhlulu) and getting unstable CV scores +- 0.02 ish for each run.\n\nI have tried the following stuff (should be useful for Keras users), but still no luck.\nIf you have any ideas why it's still unstable please let me know!\n\n### seed_keras function\n```\ndef seed_keras(seed=1029):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    tf.set_random_seed(seed)\n    # hack: Even though we don't use torch, this can be used for Keras. \",\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n```\n\n### random_transform\n`self.datagen.random_transform(X[i], seed=seed)`\n\n### flow\n`create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed)`\n\n### train_test_split\n```\nx_train, x_val, y_train, y_val = train_test_split(\n    x_train, y_train_multi, \n    test_size....\n    random_state=seed,\n)\n```",
      "votes": null
    },
    {
      "id": "578171",
      "postDate": "07/17/2019 12:37:36",
      "content": "<p>Can't say anything without looking at the exact code you are talking about. Like the kernel by xhlulu doesn't use mixup generator so your <code>self.datagen.random_transform(X[i], seed=seed)</code> is never used. </p>",
      "rawMarkdown": "Can't say anything without looking at the exact code you are talking about. Like the kernel by xhlulu doesn't use mixup generator so your `self.datagen.random_transform(X[i], seed=seed)` is never used.",
      "votes": null
    },
    {
      "id": "578185",
      "postDate": "07/17/2019 12:50:41",
      "content": "<p>I should have noticed. But you're absolutely right. Thanks for pointing that out.\nMuch appreciated.</p>",
      "rawMarkdown": "I should have noticed. But you're absolutely right. Thanks for pointing that out.\nMuch appreciated.",
      "votes": null
    },
    {
      "id": "579114",
      "postDate": "07/18/2019 14:05:58",
      "content": "<p>I have made a <a href=\"https://www.kaggle.com/dimitreoliveira/diabetic-retinopathy-shap-model-explainability\">kernel</a> and also seeded as much as I could, but I don't think it can guarantee consistent results.</p>",
      "rawMarkdown": "I have made a [kernel](https://www.kaggle.com/dimitreoliveira/diabetic-retinopathy-shap-model-explainability) and also seeded as much as I could, but I don't think it can guarantee consistent results.",
      "votes": null
    },
    {
      "id": "579519",
      "postDate": "07/18/2019 23:14:40",
      "content": "<p>To my surprise, I even got non-deterministic result with predict_generator not just fit 😹 Not sure if it is only my bug or anyone else having the same issue?</p>\n\n<p>BTW, <a href=\"/higepon\">@higepon</a> according to discussions of many past competitions, only pytorch not keras can make deterministic (training) results.</p>",
      "rawMarkdown": "To my surprise, I even got non-deterministic result with predict_generator not just fit 😹 Not sure if it is only my bug or anyone else having the same issue?\n\nBTW, @higepon according to discussions of many past competitions, only pytorch not keras can make deterministic (training) results.",
      "votes": null
    },
    {
      "id": "579560",
      "postDate": "07/19/2019 01:10:03",
      "content": "<p>Thanks for the info!\nI also have the issue with predict_generator :(</p>",
      "rawMarkdown": "Thanks for the info!\nI also have the issue with predict_generator :(",
      "votes": null
    },
    {
      "id": "579561",
      "postDate": "07/19/2019 01:10:35",
      "content": "<p>Maybe we should start using torch instead.</p>",
      "rawMarkdown": "Maybe we should start using torch instead.",
      "votes": null
    },
    {
      "id": "579956",
      "postDate": "07/19/2019 12:55:55",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> How are you dealing with non-deterministic CV score when making some model change?\nI run the change n times and then average them and see if it improves. </p>",
      "rawMarkdown": "ratthachat How are you dealing with non-deterministic CV score when making some model change?\nI run the change n times and then average them and see if it improves.",
      "votes": null
    },
    {
      "id": "579963",
      "postDate": "07/19/2019 13:18:06",
      "content": "<p>Hi <a href=\"/higepon\">@higepon</a> , thanks for asking. At the moment, I can live with a bit of randomness. As of now,  for each setting, I normally vary some hyperparameters to check my understanding, and I can see some small fluctuation. In contrast, when the model really improves, the performance improves quite significantly more than random fluctuation (e.g. from .73-.74 —&gt; .77-.78)</p>\n\n<p>I try to confirm that the model indeed make an improvement by visualizing the features that CNN really looks and makes a prediction decision (please refer to my kernel on Grad-CAM) ... Each time, I think the model is able to capture more significant features (scabs or bloods), not spurious ones.</p>\n\n<p>For pytorch, even though we really get deterministic results, but that deterministic can be bad (due to bad luck of parameter setting), so we still have to vary parameters to see average performance anyway.</p>\n\n<p><strong>EDIT:</strong> Another trick to stabilize more on your prediction is just to apply TTA, I think this help really (from 0.02 fluctuation down to 0.01).</p>",
      "rawMarkdown": "Hi @higepon , thanks for asking. At the moment, I can live with a bit of randomness. As of now,  for each setting, I normally vary some hyperparameters to check my understanding, and I can see some small fluctuation. In contrast, when the model really improves, the performance improves quite significantly more than random fluctuation (e.g. from .73-.74 —&gt; .77-.78)\n\nI try to confirm that the model indeed make an improvement by visualizing the features that CNN really looks and makes a prediction decision (please refer to my kernel on Grad-CAM) ... Each time, I think the model is able to capture more significant features (scabs or bloods), not spurious ones.\n\nFor pytorch, even though we really get deterministic results, but that deterministic can be bad (due to bad luck of parameter setting), so we still have to vary parameters to see average performance anyway.\n\n**EDIT:** Another trick to stabilize more on your prediction is just to apply TTA, I think this help really (from 0.02 fluctuation down to 0.01).",
      "votes": null
    },
    {
      "id": "579990",
      "postDate": "07/19/2019 14:11:05",
      "content": "<p>Thanks for the details and you insights. They are really helpful.\nI'll try to Grad-CAM!\nYeah TTA is basically averaging so it makes sense it has more stability.</p>",
      "rawMarkdown": "Thanks for the details and you insights. They are really helpful.\nI'll try to Grad-CAM!\nYeah TTA is basically averaging so it makes sense it has more stability.",
      "votes": null
    },
    {
      "id": "580313",
      "postDate": "07/20/2019 02:01:55",
      "content": "<p>I must admit that Pytorch has at least two main advantages anyway. First, many SOTA codes and pre-trained weights are implemented in Pytorch, and so as high quality kaggle kernels too. Second, I saw in the past that Pytorch usually runs faster and perhaps more flexible to optimize for running-time.</p>",
      "rawMarkdown": "I must admit that Pytorch has at least two main advantages anyway. First, many SOTA codes and pre-trained weights are implemented in Pytorch, and so as high quality kaggle kernels too. Second, I saw in the past that Pytorch usually runs faster and perhaps more flexible to optimize for running-time.",
      "votes": null
    },
    {
      "id": "580341",
      "postDate": "07/20/2019 02:50:31",
      "content": "<p>Yeah. I observed it in past competitions.\nI personally like Keras API and stick with it, but probably I should use torch...</p>",
      "rawMarkdown": "Yeah. I observed it in past competitions.\nI personally like Keras API and stick with it, but probably I should use torch...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 578171,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "07/17/2019 12:37:36",
      "content": "<p>Can't say anything without looking at the exact code you are talking about. Like the kernel by xhlulu doesn't use mixup generator so your <code>self.datagen.random_transform(X[i], seed=seed)</code> is never used. </p>",
      "votes": null,
      "replies": [
        {
          "id": 578185,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/17/2019 12:50:41",
          "content": "<p>I should have noticed. But you're absolutely right. Thanks for pointing that out.\nMuch appreciated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 579114,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "07/18/2019 14:05:58",
      "content": "<p>I have made a <a href=\"https://www.kaggle.com/dimitreoliveira/diabetic-retinopathy-shap-model-explainability\">kernel</a> and also seeded as much as I could, but I don't think it can guarantee consistent results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 579561,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/19/2019 01:10:35",
          "content": "<p>Maybe we should start using torch instead.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 579519,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "07/18/2019 23:14:40",
      "content": "<p>To my surprise, I even got non-deterministic result with predict_generator not just fit 😹 Not sure if it is only my bug or anyone else having the same issue?</p>\n\n<p>BTW, <a href=\"/higepon\">@higepon</a> according to discussions of many past competitions, only pytorch not keras can make deterministic (training) results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 579560,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/19/2019 01:10:03",
          "content": "<p>Thanks for the info!\nI also have the issue with predict_generator :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 579956,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/19/2019 12:55:55",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> How are you dealing with non-deterministic CV score when making some model change?\nI run the change n times and then average them and see if it improves. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 579963,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/19/2019 13:18:06",
          "content": "<p>Hi <a href=\"/higepon\">@higepon</a> , thanks for asking. At the moment, I can live with a bit of randomness. As of now,  for each setting, I normally vary some hyperparameters to check my understanding, and I can see some small fluctuation. In contrast, when the model really improves, the performance improves quite significantly more than random fluctuation (e.g. from .73-.74 —&gt; .77-.78)</p>\n\n<p>I try to confirm that the model indeed make an improvement by visualizing the features that CNN really looks and makes a prediction decision (please refer to my kernel on Grad-CAM) ... Each time, I think the model is able to capture more significant features (scabs or bloods), not spurious ones.</p>\n\n<p>For pytorch, even though we really get deterministic results, but that deterministic can be bad (due to bad luck of parameter setting), so we still have to vary parameters to see average performance anyway.</p>\n\n<p><strong>EDIT:</strong> Another trick to stabilize more on your prediction is just to apply TTA, I think this help really (from 0.02 fluctuation down to 0.01).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 579990,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/19/2019 14:11:05",
          "content": "<p>Thanks for the details and you insights. They are really helpful.\nI'll try to Grad-CAM!\nYeah TTA is basically averaging so it makes sense it has more stability.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 580313,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/20/2019 02:01:55",
          "content": "<p>I must admit that Pytorch has at least two main advantages anyway. First, many SOTA codes and pre-trained weights are implemented in Pytorch, and so as high quality kaggle kernels too. Second, I saw in the past that Pytorch usually runs faster and perhaps more flexible to optimize for running-time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 580341,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "07/20/2019 02:50:31",
          "content": "<p>Yeah. I observed it in past competitions.\nI personally like Keras API and stick with it, but probably I should use torch...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "578155": "Hi folks,\nI have been playing around kernel based off [APTOS 2019: DenseNet Keras Starter](https://www.kaggle.com/xhlulu/aptos-2019-densenet-keras-starter) by [xhlulu](https://www.kaggle.com/xhlulu) and getting unstable CV scores +- 0.02 ish for each run.\n\nI have tried the following stuff (should be useful for Keras users), but still no luck.\nIf you have any ideas why it's still unstable please let me know!\n\n### seed_keras function\n```\ndef seed_keras(seed=1029):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    tf.set_random_seed(seed)\n    # hack: Even though we don't use torch, this can be used for Keras. \",\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n```\n\n### random_transform\n`self.datagen.random_transform(X[i], seed=seed)`\n\n### flow\n`create_datagen().flow(x_train, y_train, batch_size=BATCH_SIZE, seed=seed)`\n\n### train_test_split\n```\nx_train, x_val, y_train, y_val = train_test_split(\n    x_train, y_train_multi, \n    test_size....\n    random_state=seed,\n)\n```",
    "578171": "Can't say anything without looking at the exact code you are talking about. Like the kernel by xhlulu doesn't use mixup generator so your `self.datagen.random_transform(X[i], seed=seed)` is never used.",
    "578185": "I should have noticed. But you're absolutely right. Thanks for pointing that out.\nMuch appreciated.",
    "579114": "I have made a [kernel](https://www.kaggle.com/dimitreoliveira/diabetic-retinopathy-shap-model-explainability) and also seeded as much as I could, but I don't think it can guarantee consistent results.",
    "579519": "To my surprise, I even got non-deterministic result with predict_generator not just fit 😹 Not sure if it is only my bug or anyone else having the same issue?\n\nBTW, @higepon according to discussions of many past competitions, only pytorch not keras can make deterministic (training) results.",
    "579560": "Thanks for the info!\nI also have the issue with predict_generator :(",
    "579561": "Maybe we should start using torch instead.",
    "579956": "ratthachat How are you dealing with non-deterministic CV score when making some model change?\nI run the change n times and then average them and see if it improves.",
    "579963": "Hi @higepon , thanks for asking. At the moment, I can live with a bit of randomness. As of now,  for each setting, I normally vary some hyperparameters to check my understanding, and I can see some small fluctuation. In contrast, when the model really improves, the performance improves quite significantly more than random fluctuation (e.g. from .73-.74 —&gt; .77-.78)\n\nI try to confirm that the model indeed make an improvement by visualizing the features that CNN really looks and makes a prediction decision (please refer to my kernel on Grad-CAM) ... Each time, I think the model is able to capture more significant features (scabs or bloods), not spurious ones.\n\nFor pytorch, even though we really get deterministic results, but that deterministic can be bad (due to bad luck of parameter setting), so we still have to vary parameters to see average performance anyway.\n\n**EDIT:** Another trick to stabilize more on your prediction is just to apply TTA, I think this help really (from 0.02 fluctuation down to 0.01).",
    "579990": "Thanks for the details and you insights. They are really helpful.\nI'll try to Grad-CAM!\nYeah TTA is basically averaging so it makes sense it has more stability.",
    "580313": "I must admit that Pytorch has at least two main advantages anyway. First, many SOTA codes and pre-trained weights are implemented in Pytorch, and so as high quality kaggle kernels too. Second, I saw in the past that Pytorch usually runs faster and perhaps more flexible to optimize for running-time.",
    "580341": "Yeah. I observed it in past competitions.\nI personally like Keras API and stick with it, but probably I should use torch..."
  },
  "source": "meta"
}