{
  "id": 74333,
  "title": "Keras reproduce your score ",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74333",
  "author_name": "",
  "post_date": "2018-12-11T10:33:36.586875900Z",
  "votes": 14,
  "comment_count": 27,
  "views": 0,
  "content": "<p>After many attempts I was able to submit a script with the same LB score 0.682 (WIP), just try to init your seed as below, you maybe can  reproduce your score, good luck and waiting for your feedback.</p>\n\n<p><strong>first line of your notebook</strong></p>\n\n<pre><code>seed_nb=14\nimport numpy as np \nnp.random.seed(seed_nb)\nimport tensorflow as tf\ntf.set_random_seed(seed_nb)\n</code></pre>\n\n<p><strong>Init all your Model seed</strong></p>\n\n<pre><code>def model_lstm_gru_atten(maxlen, max_features, embed_size,embedding_matrix):\n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(0.1,seed=seed_nb)(x)\n    x = Bidirectional(CuDNNLSTM(64, kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n    y = Bidirectional(CuDNNGRU(40,kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n\natten_1 = Attention(maxlen)(x) \natten_2 = Attention(maxlen)(y)\navg_pool = GlobalAveragePooling1D()(y)\nmax_pool = GlobalMaxPooling1D()(y)\n\nconc = concatenate([atten_1, atten_2, avg_pool, max_pool])\nconc = Dense(16,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"relu\")(conc)\nconc = Dropout(0.1,seed=seed_nb)(conc)\noutp = Dense(1,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"sigmoid\")(conc)    \n\nmodel = Model(inputs=inp, outputs=outp)\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=[f1])\nreturn model\n</code></pre>",
  "messages": [
    {
      "id": "437077",
      "postDate": "12/11/2018 10:33:36",
      "content": "<p>After many attempts I was able to submit a script with the same LB score 0.682 (WIP), just try to init your seed as below, you maybe can  reproduce your score, good luck and waiting for your feedback.</p>\n\n<p><strong>first line of your notebook</strong></p>\n\n<pre><code>seed_nb=14\nimport numpy as np \nnp.random.seed(seed_nb)\nimport tensorflow as tf\ntf.set_random_seed(seed_nb)\n</code></pre>\n\n<p><strong>Init all your Model seed</strong></p>\n\n<pre><code>def model_lstm_gru_atten(maxlen, max_features, embed_size,embedding_matrix):\n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(0.1,seed=seed_nb)(x)\n    x = Bidirectional(CuDNNLSTM(64, kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n    y = Bidirectional(CuDNNGRU(40,kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n\natten_1 = Attention(maxlen)(x) \natten_2 = Attention(maxlen)(y)\navg_pool = GlobalAveragePooling1D()(y)\nmax_pool = GlobalMaxPooling1D()(y)\n\nconc = concatenate([atten_1, atten_2, avg_pool, max_pool])\nconc = Dense(16,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"relu\")(conc)\nconc = Dropout(0.1,seed=seed_nb)(conc)\noutp = Dense(1,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"sigmoid\")(conc)    \n\nmodel = Model(inputs=inp, outputs=outp)\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=[f1])\nreturn model\n</code></pre>",
      "rawMarkdown": "After many attempts I was able to submit a script with the same LB score 0.682 (WIP), just try to init your seed as below, you maybe can  reproduce your score, good luck and waiting for your feedback.\n\n**first line of your notebook**\n\n\n    seed_nb=14\n    import numpy as np \n    np.random.seed(seed_nb)\n    import tensorflow as tf\n    tf.set_random_seed(seed_nb)\n\n\n**Init all your Model seed**\n\n\n\n    def model_lstm_gru_atten(maxlen, max_features, embed_size,embedding_matrix):\n        inp = Input(shape=(maxlen,))\n        x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n        x = SpatialDropout1D(0.1,seed=seed_nb)(x)\n        x = Bidirectional(CuDNNLSTM(64, kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n        y = Bidirectional(CuDNNGRU(40,kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n    \n    atten_1 = Attention(maxlen)(x) \n    atten_2 = Attention(maxlen)(y)\n    avg_pool = GlobalAveragePooling1D()(y)\n    max_pool = GlobalMaxPooling1D()(y)\n    \n    conc = concatenate([atten_1, atten_2, avg_pool, max_pool])\n    conc = Dense(16,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"relu\")(conc)\n    conc = Dropout(0.1,seed=seed_nb)(conc)\n    outp = Dense(1,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"sigmoid\")(conc)    \n\n    model = Model(inputs=inp, outputs=outp)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=[f1])\n    return model",
      "votes": null
    },
    {
      "id": "439302",
      "postDate": "12/15/2018 05:54:44",
      "content": "<p>have you check every line of keras output to ensure reproduce? it was too strange that no one realize this simple way before. Thank you, i will try</p>",
      "rawMarkdown": "have you check every line of keras output to ensure reproduce? it was too strange that no one realize this simple way before. Thank you, i will try",
      "votes": null
    },
    {
      "id": "439416",
      "postDate": "12/15/2018 12:16:28",
      "content": "<blockquote>\n  <p>it was too strange that no one realize this simple way before</p>\n</blockquote>\n\n<p>Actually some people were doing it way before :) <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424573\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424573</a></p>\n\n<p>But you can't have exact same results whatever you fix, as long as you're using CuDNN</p>",
      "rawMarkdown": "&gt; it was too strange that no one realize this simple way before\n\n\nActually some people were doing it way before :) https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424573\n\nBut you can't have exact same results whatever you fix, as long as you're using CuDNN",
      "votes": null
    },
    {
      "id": "439454",
      "postDate": "12/15/2018 13:51:32",
      "content": "<p>awesome, but LB and local CV differ great？</p>",
      "rawMarkdown": "awesome, but LB and local CV differ great？",
      "votes": null
    },
    {
      "id": "439478",
      "postDate": "12/15/2018 15:11:54",
      "content": "<p>I was  running the same kernel on kaggle VM with this config  and getting the same LB score, so I share it with the community, maybe it will inspire  kagglers or help to find out new tricks </p>",
      "rawMarkdown": "I was  running the same kernel on kaggle VM with this config  and getting the same LB score, so I share it with the community, maybe it will inspire  kagglers or help to find out new tricks",
      "votes": null
    },
    {
      "id": "439479",
      "postDate": "12/15/2018 15:13:55",
      "content": "<p>You just can't make the atomic operations of CUDNN deterministic, so your equal results were lucky :)</p>",
      "rawMarkdown": "You just can't make the atomic operations of CUDNN deterministic, so your equal results were lucky :)",
      "votes": null
    },
    {
      "id": "439482",
      "postDate": "12/15/2018 15:17:26",
      "content": "<p>so let's run the kernel again I'will report the result :)</p>",
      "rawMarkdown": "so let's run the kernel again I'will report the result :)",
      "votes": null
    },
    {
      "id": "439552",
      "postDate": "12/15/2018 19:05:33",
      "content": "<p>I agree. No keras or pytorch seems solving the randomness. </p>\n\n<p><a href=\"/serigne\">@serigne</a> Can you share how are you managing the score increments with randomness? Has your team  solved the infamous reproducability ? </p>",
      "rawMarkdown": "I agree. No keras or pytorch seems solving the randomness. \n\n@serigne Can you share how are you managing the score increments with randomness? Has your team  solved the infamous reproducability ?",
      "votes": null
    },
    {
      "id": "439572",
      "postDate": "12/15/2018 20:19:23",
      "content": "<p>5th run of the same notebook with the same score of 0.681, btw changing the seed_nb produce different LD score</p>",
      "rawMarkdown": "5th run of the same notebook with the same score of 0.681, btw changing the seed_nb produce different LD score",
      "votes": null
    },
    {
      "id": "439589",
      "postDate": "12/15/2018 21:13:15",
      "content": "<p>As soon as you go above 0.686 or something like that you will see more variations. I am doing exactly the same things as you and see large variations. See: <a href=\"https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility\">https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility</a></p>",
      "rawMarkdown": "As soon as you go above 0.686 or something like that you will see more variations. I am doing exactly the same things as you and see large variations. See: https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility",
      "votes": null
    },
    {
      "id": "439689",
      "postDate": "12/16/2018 04:54:42",
      "content": "<p>I second @Phillip. Progressing on higher scores made it more random. Also, I think its the leaderboard and competitive spirit we are wondering 0.699 to 0.700 as A BIG difference. But, in real scenario its entirely acceptable.</p>",
      "rawMarkdown": "I second @Phillip. Progressing on higher scores made it more random. Also, I think its the leaderboard and competitive spirit we are wondering 0.699 to 0.700 as A BIG difference. But, in real scenario its entirely acceptable.",
      "votes": null
    },
    {
      "id": "439782",
      "postDate": "12/16/2018 10:55:26",
      "content": "<p>We've not solved it yet..but plan to work on it.</p>",
      "rawMarkdown": "We've not solved it yet..but plan to work on it.",
      "votes": null
    },
    {
      "id": "439792",
      "postDate": "12/16/2018 11:19:45",
      "content": "<p>How do you consider F1 gap between LB score and local CV? i have no idea about which one should i trust</p>",
      "rawMarkdown": "How do you consider F1 gap between LB score and local CV? i have no idea about which one should i trust",
      "votes": null
    },
    {
      "id": "439812",
      "postDate": "12/16/2018 12:29:01",
      "content": "<p>You can trust nothing... I get lower F1 score on local CV and higher on public LB and the other way around...</p>",
      "rawMarkdown": "You can trust nothing... I get lower F1 score on local CV and higher on public LB and the other way around...",
      "votes": null
    },
    {
      "id": "441257",
      "postDate": "12/18/2018 13:34:59",
      "content": "<p>I think that some unreasonable random numbers can hinder the performance of the model. We can't just set all the random numbers for stable results. However, we can set some of them.</p>",
      "rawMarkdown": "I think that some unreasonable random numbers can hinder the performance of the model. We can't just set all the random numbers for stable results. However, we can set some of them.",
      "votes": null
    },
    {
      "id": "441452",
      "postDate": "12/18/2018 17:11:27",
      "content": "<p>Why can we not? If it is random there is no difference whether you set a lucky seed yourself or whether you are lucky otherwise. Except of Cudnn though, which we cannot make deterministic.</p>",
      "rawMarkdown": "Why can we not? If it is random there is no difference whether you set a lucky seed yourself or whether you are lucky otherwise. Except of Cudnn though, which we cannot make deterministic.",
      "votes": null
    },
    {
      "id": "441809",
      "postDate": "12/19/2018 05:22:28",
      "content": "<p>We have no way of knowing how random seeds affect random numbers, but the initial values of neural network parameters have a large impact on network performance.</p>",
      "rawMarkdown": "We have no way of knowing how random seeds affect random numbers, but the initial values of neural network parameters have a large impact on network performance.",
      "votes": null
    },
    {
      "id": "442052",
      "postDate": "12/19/2018 12:18:56",
      "content": "<p>Just curious. Why do people care about the randomness of the score among different runs? ML is stochastic by nature, isn't it? \nBy that, I mean, even if you could get reproducible result and choose the seed that scored highest in LB from running the same model with multiple different seeds, it does not guaranteed that the chosen seed will perform the best with the final data.</p>",
      "rawMarkdown": "Just curious. Why do people care about the randomness of the score among different runs? ML is stochastic by nature, isn't it? \nBy that, I mean, even if you could get reproducible result and choose the seed that scored highest in LB from running the same model with multiple different seeds, it does not guaranteed that the chosen seed will perform the best with the final data.",
      "votes": null
    },
    {
      "id": "442082",
      "postDate": "12/19/2018 13:14:49",
      "content": "<p>Agree with you. But I tend to set some seed so that the result is not so volatile.</p>",
      "rawMarkdown": "Agree with you. But I tend to set some seed so that the result is not so volatile.",
      "votes": null
    },
    {
      "id": "443534",
      "postDate": "12/21/2018 20:26:20",
      "content": "<p>Let's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.681! What!!</p>\n\n<p>Question: Did the dropout improve your model? Did the preprocessing?</p>\n\n<p>To be fully effective and scientific, you will have to orthogonalize your experiments. Meaning, in each experiment, you tweak one thing, and then see what that does. Repeat that and slowly work your way up. That's not possible if the natural noise is as high as it is here...</p>",
      "rawMarkdown": "Let's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.681! What!!\n\nQuestion: Did the dropout improve your model? Did the preprocessing?\n\nTo be fully effective and scientific, you will have to orthogonalize your experiments. Meaning, in each experiment, you tweak one thing, and then see what that does. Repeat that and slowly work your way up. That's not possible if the natural noise is as high as it is here...",
      "votes": null
    },
    {
      "id": "444071",
      "postDate": "12/23/2018 04:21:25",
      "content": "<p>But is it really rational to expect such behavior?  It sure would be nice if we can have what you described. But my understanding is that ML is stochastic. Even if you can reproduce the result with %100 certainty by fixing a random seed, it does not mean that the same random number sequence will be equally good/bad for a tweaked model.</p>\n\n<p>Let's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.689. What does that really mean? Does this really mean that the preprocessing step hurt your model? Is it not possble that the particular random number sequence you are getting with the particular random seed happens to give the best LB score your model without preprocessing will ever get, but give you the worst possible LB score with the preprocessing step added?</p>\n\n<p>To be really sure, you need to train the same model multiple times with different random number sequences and see on average which model gives you better LB score. Even then, the highest public LB score does not guarantee highest private LB score.</p>\n\n<p>I agree that it feels nice to have a repeatable training. But I'm afraid that it is more a false sense of comfort.</p>\n\n<p>Besides, even if you could \"tweak one thing, and then see what that does. Repeat that and slowly work your way up.\" What guarantee do you have that higher LB score you are getting by such method is not just overfitting to public LB test data?</p>\n\n<p>I'm no expert in the details of Keras/PyTorch implementation. But does PyTorch really provide more reproducible result? How does its performance compare to Keras? When things are happening in parallel as is the case with GPU, you can often get more performance when using relaxed consistency model. Maybe we are getting better performance because CuDNNGRU/CuDNNLSTM does not try to guarantee the same result when using the same random number seed. I'm not claiming that is the case. I'm suspecting that is the case. If so, then I'd rather squeeze in 1 more epoch of training than get a reproducible training.</p>\n\n<p>Just a thought.</p>",
      "rawMarkdown": "But is it really rational to expect such behavior?  It sure would be nice if we can have what you described. But my understanding is that ML is stochastic. Even if you can reproduce the result with %100 certainty by fixing a random seed, it does not mean that the same random number sequence will be equally good/bad for a tweaked model.\n\nLet's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.689. What does that really mean? Does this really mean that the preprocessing step hurt your model? Is it not possble that the particular random number sequence you are getting with the particular random seed happens to give the best LB score your model without preprocessing will ever get, but give you the worst possible LB score with the preprocessing step added?\n\nTo be really sure, you need to train the same model multiple times with different random number sequences and see on average which model gives you better LB score. Even then, the highest public LB score does not guarantee highest private LB score.\n\nI agree that it feels nice to have a repeatable training. But I'm afraid that it is more a false sense of comfort.\n\nBesides, even if you could \"tweak one thing, and then see what that does. Repeat that and slowly work your way up.\" What guarantee do you have that higher LB score you are getting by such method is not just overfitting to public LB test data?\n\nI'm no expert in the details of Keras/PyTorch implementation. But does PyTorch really provide more reproducible result? How does its performance compare to Keras? When things are happening in parallel as is the case with GPU, you can often get more performance when using relaxed consistency model. Maybe we are getting better performance because CuDNNGRU/CuDNNLSTM does not try to guarantee the same result when using the same random number seed. I'm not claiming that is the case. I'm suspecting that is the case. If so, then I'd rather squeeze in 1 more epoch of training than get a reproducible training.\n\nJust a thought.",
      "votes": null
    },
    {
      "id": "444109",
      "postDate": "12/23/2018 07:55:55",
      "content": "<p>In my opinion this is a very high noise dataset and every time you add something to improve your model and lb score doesn't improve, it doesn't guarantee that the thing you tried was not right..it is just that it didn't perform on the 25% test dataset chosen for lb public board. If you are confident about your model and have sound logics behind it then you should go with it...it will definitely work well on real world problems and real datasets in general. Trust your ML logics and ML will not disappoint you.</p>",
      "rawMarkdown": "In my opinion this is a very high noise dataset and every time you add something to improve your model and lb score doesn't improve, it doesn't guarantee that the thing you tried was not right..it is just that it didn't perform on the 25% test dataset chosen for lb public board. If you are confident about your model and have sound logics behind it then you should go with it...it will definitely work well on real world problems and real datasets in general. Trust your ML logics and ML will not disappoint you.",
      "votes": null
    },
    {
      "id": "444123",
      "postDate": "12/23/2018 09:21:58",
      "content": "<p>You are probably right, my method will mostly just serve to overfit to the public LB score. Thanks for grounding me back in reality :)</p>\n\n<p>Anyway, about PyTorch / Keras: PyTorch is much faster than Keras, and its results are 100% reproducible (look for a kernel promising this).</p>",
      "rawMarkdown": "You are probably right, my method will mostly just serve to overfit to the public LB score. Thanks for grounding me back in reality :)\n\nAnyway, about PyTorch / Keras: PyTorch is much faster than Keras, and its results are 100% reproducible (look for a kernel promising this).",
      "votes": null
    },
    {
      "id": "444129",
      "postDate": "12/23/2018 09:36:04",
      "content": "<p>I still do not believe that PyTorch is 100% reproducible, except they suddenly changed the atomic operations within CUDNN.</p>",
      "rawMarkdown": "I still do not believe that PyTorch is 100% reproducible, except they suddenly changed the atomic operations within CUDNN.",
      "votes": null
    },
    {
      "id": "444258",
      "postDate": "12/23/2018 17:12:39",
      "content": "<p>Well, I have a kernel that I can run multiple times and the loss numbers over the epochs are 100% the same each time. So I would say yes, it is 100% reproducible.</p>",
      "rawMarkdown": "Well, I have a kernel that I can run multiple times and the loss numbers over the epochs are 100% the same each time. So I would say yes, it is 100% reproducible.",
      "votes": null
    },
    {
      "id": "444264",
      "postDate": "12/23/2018 17:24:55",
      "content": "<p>Also F1 score and public lb score?</p>",
      "rawMarkdown": "Also F1 score and public lb score?",
      "votes": null
    },
    {
      "id": "444268",
      "postDate": "12/23/2018 17:37:30",
      "content": "<p>Yup</p>",
      "rawMarkdown": "Yup",
      "votes": null
    },
    {
      "id": "444285",
      "postDate": "12/23/2018 18:10:56",
      "content": "<p>Interesting, maybe I will try PyTorch then again, for me it didn't work.</p>",
      "rawMarkdown": "Interesting, maybe I will try PyTorch then again, for me it didn't work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 439302,
      "author_name": "liangdeli",
      "author_url": "",
      "post_date": "12/15/2018 05:54:44",
      "content": "<p>have you check every line of keras output to ensure reproduce? it was too strange that no one realize this simple way before. Thank you, i will try</p>",
      "votes": null,
      "replies": [
        {
          "id": 439416,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "12/15/2018 12:16:28",
          "content": "<blockquote>\n  <p>it was too strange that no one realize this simple way before</p>\n</blockquote>\n\n<p>Actually some people were doing it way before :) <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424573\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424573</a></p>\n\n<p>But you can't have exact same results whatever you fix, as long as you're using CuDNN</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439454,
          "author_name": "liangdeli",
          "author_url": "",
          "post_date": "12/15/2018 13:51:32",
          "content": "<p>awesome, but LB and local CV differ great？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439478,
          "author_name": "malekbadreddine",
          "author_url": "",
          "post_date": "12/15/2018 15:11:54",
          "content": "<p>I was  running the same kernel on kaggle VM with this config  and getting the same LB score, so I share it with the community, maybe it will inspire  kagglers or help to find out new tricks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439552,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/15/2018 19:05:33",
          "content": "<p>I agree. No keras or pytorch seems solving the randomness. </p>\n\n<p><a href=\"/serigne\">@serigne</a> Can you share how are you managing the score increments with randomness? Has your team  solved the infamous reproducability ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439782,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "12/16/2018 10:55:26",
          "content": "<p>We've not solved it yet..but plan to work on it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439792,
          "author_name": "liangdeli",
          "author_url": "",
          "post_date": "12/16/2018 11:19:45",
          "content": "<p>How do you consider F1 gap between LB score and local CV? i have no idea about which one should i trust</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439812,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/16/2018 12:29:01",
          "content": "<p>You can trust nothing... I get lower F1 score on local CV and higher on public LB and the other way around...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 439479,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "12/15/2018 15:13:55",
      "content": "<p>You just can't make the atomic operations of CUDNN deterministic, so your equal results were lucky :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 439482,
          "author_name": "malekbadreddine",
          "author_url": "",
          "post_date": "12/15/2018 15:17:26",
          "content": "<p>so let's run the kernel again I'will report the result :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439572,
          "author_name": "malekbadreddine",
          "author_url": "",
          "post_date": "12/15/2018 20:19:23",
          "content": "<p>5th run of the same notebook with the same score of 0.681, btw changing the seed_nb produce different LD score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439589,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/15/2018 21:13:15",
          "content": "<p>As soon as you go above 0.686 or something like that you will see more variations. I am doing exactly the same things as you and see large variations. See: <a href=\"https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility\">https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439689,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/16/2018 04:54:42",
          "content": "<p>I second @Phillip. Progressing on higher scores made it more random. Also, I think its the leaderboard and competitive spirit we are wondering 0.699 to 0.700 as A BIG difference. But, in real scenario its entirely acceptable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441257,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/18/2018 13:34:59",
      "content": "<p>I think that some unreasonable random numbers can hinder the performance of the model. We can't just set all the random numbers for stable results. However, we can set some of them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441452,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/18/2018 17:11:27",
          "content": "<p>Why can we not? If it is random there is no difference whether you set a lucky seed yourself or whether you are lucky otherwise. Except of Cudnn though, which we cannot make deterministic.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441809,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/19/2018 05:22:28",
          "content": "<p>We have no way of knowing how random seeds affect random numbers, but the initial values of neural network parameters have a large impact on network performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442052,
      "author_name": "bsp2020",
      "author_url": "",
      "post_date": "12/19/2018 12:18:56",
      "content": "<p>Just curious. Why do people care about the randomness of the score among different runs? ML is stochastic by nature, isn't it? \nBy that, I mean, even if you could get reproducible result and choose the seed that scored highest in LB from running the same model with multiple different seeds, it does not guaranteed that the chosen seed will perform the best with the final data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 442082,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/19/2018 13:14:49",
          "content": "<p>Agree with you. But I tend to set some seed so that the result is not so volatile.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443534,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/21/2018 20:26:20",
          "content": "<p>Let's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.681! What!!</p>\n\n<p>Question: Did the dropout improve your model? Did the preprocessing?</p>\n\n<p>To be fully effective and scientific, you will have to orthogonalize your experiments. Meaning, in each experiment, you tweak one thing, and then see what that does. Repeat that and slowly work your way up. That's not possible if the natural noise is as high as it is here...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444071,
          "author_name": "bsp2020",
          "author_url": "",
          "post_date": "12/23/2018 04:21:25",
          "content": "<p>But is it really rational to expect such behavior?  It sure would be nice if we can have what you described. But my understanding is that ML is stochastic. Even if you can reproduce the result with %100 certainty by fixing a random seed, it does not mean that the same random number sequence will be equally good/bad for a tweaked model.</p>\n\n<p>Let's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.689. What does that really mean? Does this really mean that the preprocessing step hurt your model? Is it not possble that the particular random number sequence you are getting with the particular random seed happens to give the best LB score your model without preprocessing will ever get, but give you the worst possible LB score with the preprocessing step added?</p>\n\n<p>To be really sure, you need to train the same model multiple times with different random number sequences and see on average which model gives you better LB score. Even then, the highest public LB score does not guarantee highest private LB score.</p>\n\n<p>I agree that it feels nice to have a repeatable training. But I'm afraid that it is more a false sense of comfort.</p>\n\n<p>Besides, even if you could \"tweak one thing, and then see what that does. Repeat that and slowly work your way up.\" What guarantee do you have that higher LB score you are getting by such method is not just overfitting to public LB test data?</p>\n\n<p>I'm no expert in the details of Keras/PyTorch implementation. But does PyTorch really provide more reproducible result? How does its performance compare to Keras? When things are happening in parallel as is the case with GPU, you can often get more performance when using relaxed consistency model. Maybe we are getting better performance because CuDNNGRU/CuDNNLSTM does not try to guarantee the same result when using the same random number seed. I'm not claiming that is the case. I'm suspecting that is the case. If so, then I'd rather squeeze in 1 more epoch of training than get a reproducible training.</p>\n\n<p>Just a thought.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444123,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/23/2018 09:21:58",
          "content": "<p>You are probably right, my method will mostly just serve to overfit to the public LB score. Thanks for grounding me back in reality :)</p>\n\n<p>Anyway, about PyTorch / Keras: PyTorch is much faster than Keras, and its results are 100% reproducible (look for a kernel promising this).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444129,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/23/2018 09:36:04",
          "content": "<p>I still do not believe that PyTorch is 100% reproducible, except they suddenly changed the atomic operations within CUDNN.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444258,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/23/2018 17:12:39",
          "content": "<p>Well, I have a kernel that I can run multiple times and the loss numbers over the epochs are 100% the same each time. So I would say yes, it is 100% reproducible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444264,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/23/2018 17:24:55",
          "content": "<p>Also F1 score and public lb score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444268,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/23/2018 17:37:30",
          "content": "<p>Yup</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444285,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/23/2018 18:10:56",
          "content": "<p>Interesting, maybe I will try PyTorch then again, for me it didn't work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 444109,
      "author_name": "raghav3490",
      "author_url": "",
      "post_date": "12/23/2018 07:55:55",
      "content": "<p>In my opinion this is a very high noise dataset and every time you add something to improve your model and lb score doesn't improve, it doesn't guarantee that the thing you tried was not right..it is just that it didn't perform on the 25% test dataset chosen for lb public board. If you are confident about your model and have sound logics behind it then you should go with it...it will definitely work well on real world problems and real datasets in general. Trust your ML logics and ML will not disappoint you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "437077": "After many attempts I was able to submit a script with the same LB score 0.682 (WIP), just try to init your seed as below, you maybe can  reproduce your score, good luck and waiting for your feedback.\n\n**first line of your notebook**\n\n\n    seed_nb=14\n    import numpy as np \n    np.random.seed(seed_nb)\n    import tensorflow as tf\n    tf.set_random_seed(seed_nb)\n\n\n**Init all your Model seed**\n\n\n\n    def model_lstm_gru_atten(maxlen, max_features, embed_size,embedding_matrix):\n        inp = Input(shape=(maxlen,))\n        x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n        x = SpatialDropout1D(0.1,seed=seed_nb)(x)\n        x = Bidirectional(CuDNNLSTM(64, kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n        y = Bidirectional(CuDNNGRU(40,kernel_initializer=glorot_uniform(seed=seed_nb), return_sequences=True))(x)\n    \n    atten_1 = Attention(maxlen)(x) \n    atten_2 = Attention(maxlen)(y)\n    avg_pool = GlobalAveragePooling1D()(y)\n    max_pool = GlobalMaxPooling1D()(y)\n    \n    conc = concatenate([atten_1, atten_2, avg_pool, max_pool])\n    conc = Dense(16,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"relu\")(conc)\n    conc = Dropout(0.1,seed=seed_nb)(conc)\n    outp = Dense(1,kernel_initializer=he_uniform(seed=seed_nb),  activation=\"sigmoid\")(conc)    \n\n    model = Model(inputs=inp, outputs=outp)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=[f1])\n    return model",
    "439302": "have you check every line of keras output to ensure reproduce? it was too strange that no one realize this simple way before. Thank you, i will try",
    "439416": "&gt; it was too strange that no one realize this simple way before\n\n\nActually some people were doing it way before :) https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424573\n\nBut you can't have exact same results whatever you fix, as long as you're using CuDNN",
    "439454": "awesome, but LB and local CV differ great？",
    "439478": "I was  running the same kernel on kaggle VM with this config  and getting the same LB score, so I share it with the community, maybe it will inspire  kagglers or help to find out new tricks",
    "439479": "You just can't make the atomic operations of CUDNN deterministic, so your equal results were lucky :)",
    "439482": "so let's run the kernel again I'will report the result :)",
    "439552": "I agree. No keras or pytorch seems solving the randomness. \n\n@serigne Can you share how are you managing the score increments with randomness? Has your team  solved the infamous reproducability ?",
    "439572": "5th run of the same notebook with the same score of 0.681, btw changing the seed_nb produce different LD score",
    "439589": "As soon as you go above 0.686 or something like that you will see more variations. I am doing exactly the same things as you and see large variations. See: https://docs.nvidia.com/deeplearning/sdk/cudnn-developer-guide/index.html#reproducibility",
    "439689": "I second @Phillip. Progressing on higher scores made it more random. Also, I think its the leaderboard and competitive spirit we are wondering 0.699 to 0.700 as A BIG difference. But, in real scenario its entirely acceptable.",
    "439782": "We've not solved it yet..but plan to work on it.",
    "439792": "How do you consider F1 gap between LB score and local CV? i have no idea about which one should i trust",
    "439812": "You can trust nothing... I get lower F1 score on local CV and higher on public LB and the other way around...",
    "441257": "I think that some unreasonable random numbers can hinder the performance of the model. We can't just set all the random numbers for stable results. However, we can set some of them.",
    "441452": "Why can we not? If it is random there is no difference whether you set a lucky seed yourself or whether you are lucky otherwise. Except of Cudnn though, which we cannot make deterministic.",
    "441809": "We have no way of knowing how random seeds affect random numbers, but the initial values of neural network parameters have a large impact on network performance.",
    "442052": "Just curious. Why do people care about the randomness of the score among different runs? ML is stochastic by nature, isn't it? \nBy that, I mean, even if you could get reproducible result and choose the seed that scored highest in LB from running the same model with multiple different seeds, it does not guaranteed that the chosen seed will perform the best with the final data.",
    "442082": "Agree with you. But I tend to set some seed so that the result is not so volatile.",
    "443534": "Let's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.681! What!!\n\nQuestion: Did the dropout improve your model? Did the preprocessing?\n\nTo be fully effective and scientific, you will have to orthogonalize your experiments. Meaning, in each experiment, you tweak one thing, and then see what that does. Repeat that and slowly work your way up. That's not possible if the natural noise is as high as it is here...",
    "444071": "But is it really rational to expect such behavior?  It sure would be nice if we can have what you described. But my understanding is that ML is stochastic. Even if you can reproduce the result with %100 certainty by fixing a random seed, it does not mean that the same random number sequence will be equally good/bad for a tweaked model.\n\nLet's say I run a baseline model. It gets 0.685 on the LB. Now I add some dropout. It gets 0.689! Yay! Next, I add some more preprocessing. It gets 0.684! Crap! Remove the preprocessing again. 0.689. What does that really mean? Does this really mean that the preprocessing step hurt your model? Is it not possble that the particular random number sequence you are getting with the particular random seed happens to give the best LB score your model without preprocessing will ever get, but give you the worst possible LB score with the preprocessing step added?\n\nTo be really sure, you need to train the same model multiple times with different random number sequences and see on average which model gives you better LB score. Even then, the highest public LB score does not guarantee highest private LB score.\n\nI agree that it feels nice to have a repeatable training. But I'm afraid that it is more a false sense of comfort.\n\nBesides, even if you could \"tweak one thing, and then see what that does. Repeat that and slowly work your way up.\" What guarantee do you have that higher LB score you are getting by such method is not just overfitting to public LB test data?\n\nI'm no expert in the details of Keras/PyTorch implementation. But does PyTorch really provide more reproducible result? How does its performance compare to Keras? When things are happening in parallel as is the case with GPU, you can often get more performance when using relaxed consistency model. Maybe we are getting better performance because CuDNNGRU/CuDNNLSTM does not try to guarantee the same result when using the same random number seed. I'm not claiming that is the case. I'm suspecting that is the case. If so, then I'd rather squeeze in 1 more epoch of training than get a reproducible training.\n\nJust a thought.",
    "444109": "In my opinion this is a very high noise dataset and every time you add something to improve your model and lb score doesn't improve, it doesn't guarantee that the thing you tried was not right..it is just that it didn't perform on the 25% test dataset chosen for lb public board. If you are confident about your model and have sound logics behind it then you should go with it...it will definitely work well on real world problems and real datasets in general. Trust your ML logics and ML will not disappoint you.",
    "444123": "You are probably right, my method will mostly just serve to overfit to the public LB score. Thanks for grounding me back in reality :)\n\nAnyway, about PyTorch / Keras: PyTorch is much faster than Keras, and its results are 100% reproducible (look for a kernel promising this).",
    "444129": "I still do not believe that PyTorch is 100% reproducible, except they suddenly changed the atomic operations within CUDNN.",
    "444258": "Well, I have a kernel that I can run multiple times and the loss numbers over the epochs are 100% the same each time. So I would say yes, it is 100% reproducible.",
    "444264": "Also F1 score and public lb score?",
    "444268": "Yup",
    "444285": "Interesting, maybe I will try PyTorch then again, for me it didn't work."
  },
  "source": "meta"
}