{
  "id": 80078,
  "title": "Any difference in CV of your pytorch kernel with Stage2 data?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/80078",
  "author_name": "",
  "post_date": "2019-02-10T10:39:45.373255Z",
  "votes": 1,
  "comment_count": 18,
  "views": 0,
  "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a> posted <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\">here</a> that there is a docker image change between stage-1 and stage-2 and I now wonder that explains my CV difference. His kernel is in R and mine is in python.</p>\n\n<p>I just ran both my kernels to see if they can complete within 2hrs which they did with time to spare. </p>\n\n<p>One model is keras based so result is not strictly reproducible but my pytorch kernel is reproducible. When I ran the pytorch kernel today, there is an improvement of 0.23+ in CV even though the training data is still the same. I ran that kernel about 5 times during the competition and every single time the results for each epoch is exactly the same.</p>\n\n<p><strong>NOTE:</strong> I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab.</p>\n\n<p>Did you observe the same thing? If so please share below.</p>",
  "messages": [
    {
      "id": "469021",
      "postDate": "02/10/2019 10:39:45",
      "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a> posted <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\">here</a> that there is a docker image change between stage-1 and stage-2 and I now wonder that explains my CV difference. His kernel is in R and mine is in python.</p>\n\n<p>I just ran both my kernels to see if they can complete within 2hrs which they did with time to spare. </p>\n\n<p>One model is keras based so result is not strictly reproducible but my pytorch kernel is reproducible. When I ran the pytorch kernel today, there is an improvement of 0.23+ in CV even though the training data is still the same. I ran that kernel about 5 times during the competition and every single time the results for each epoch is exactly the same.</p>\n\n<p><strong>NOTE:</strong> I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab.</p>\n\n<p>Did you observe the same thing? If so please share below.</p>",
      "rawMarkdown": "springmanndaniel posted [here][1] that there is a docker image change between stage-1 and stage-2 and I now wonder that explains my CV difference. His kernel is in R and mine is in python.\n\nI just ran both my kernels to see if they can complete within 2hrs which they did with time to spare. \n\nOne model is keras based so result is not strictly reproducible but my pytorch kernel is reproducible. When I ran the pytorch kernel today, there is an improvement of 0.23+ in CV even though the training data is still the same. I ran that kernel about 5 times during the competition and every single time the results for each epoch is exactly the same.\n\n**NOTE:** I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab.\n\nDid you observe the same thing? If so please share below.\n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003",
      "votes": null
    },
    {
      "id": "469041",
      "postDate": "02/10/2019 11:15:51",
      "content": "<p>The score may vary few points if you build vocabulary both with train and test but +0.23 ? You sure? I don't see any chance it possible except leaked validation process</p>",
      "rawMarkdown": "The score may vary few points if you build vocabulary both with train and test but +0.23 ? You sure? I don't see any chance it possible except leaked validation process",
      "votes": null
    },
    {
      "id": "469053",
      "postDate": "02/10/2019 11:49:06",
      "content": "<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a>, I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab. </p>",
      "rawMarkdown": "yaroshevskiy, I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab.",
      "votes": null
    },
    {
      "id": "469088",
      "postDate": "02/10/2019 13:26:38",
      "content": "<p>I ran my submitted kernels on the new test data(which is 4 days ago) and I am having same cv score. They are written in pytorch.</p>",
      "rawMarkdown": "I ran my submitted kernels on the new test data(which is 4 days ago) and I am having same cv score. They are written in pytorch.",
      "votes": null
    },
    {
      "id": "469103",
      "postDate": "02/10/2019 13:51:52",
      "content": "<p>@auchith0132, is it possible for you to run it again and let me know? I ran mine today and saw <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003</a> whilst my kernels were still running. He has screen shots in his post showing the docker difference.</p>",
      "rawMarkdown": "auchith0132, is it possible for you to run it again and let me know? I ran mine today and saw https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003 whilst my kernels were still running. He has screen shots in his post showing the docker difference.",
      "votes": null
    },
    {
      "id": "469104",
      "postDate": "02/10/2019 13:57:20",
      "content": "<p>i did run my pytorch cv but although my tokenizer was taking into account both training and test data, still local CV moved from 0.6915 to 0.69. Not big diff like yourself. </p>",
      "rawMarkdown": "i did run my pytorch cv but although my tokenizer was taking into account both training and test data, still local CV moved from 0.6915 to 0.69. Not big diff like yourself.",
      "votes": null
    },
    {
      "id": "469128",
      "postDate": "02/10/2019 14:55:49",
      "content": "<p><a href=\"/mlwhiz\">@mlwhiz</a>, that is the kind of CV difference I got from my Keras kernel but pytorch is really strange. </p>",
      "rawMarkdown": "mlwhiz, that is the kind of CV difference I got from my Keras kernel but pytorch is really strange.",
      "votes": null
    },
    {
      "id": "469334",
      "postDate": "02/11/2019 02:24:15",
      "content": "<p>Hi YaGana,</p>\n\n<p>Not sure if I understand you correctly, but +0.23 means something like</p>\n\n<p>0.7 + 0.23 = 0.93??</p>",
      "rawMarkdown": "Hi YaGana,\n\nNot sure if I understand you correctly, but +0.23 means something like\n\n0.7 + 0.23 = 0.93??",
      "votes": null
    },
    {
      "id": "469380",
      "postDate": "02/11/2019 04:59:26",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I meant my cv validation loss went down from 52.4144  to 52.1763</p>",
      "rawMarkdown": "ratthachat, I meant my cv validation loss went down from 52.4144  to 52.1763",
      "votes": null
    },
    {
      "id": "469413",
      "postDate": "02/11/2019 06:34:55",
      "content": "<p>I ran the models and my cv score remained same.</p>",
      "rawMarkdown": "I ran the models and my cv score remained same.",
      "votes": null
    },
    {
      "id": "469422",
      "postDate": "02/11/2019 07:15:16",
      "content": "<p>Ok. Thanks. I rerun my kernel and the stage-2 result is reproducible, just like the stage-1 was. Since the training data did not change I am puzzled as to why I see CV change between this week and 6 days ago (i.e. when I last ran my pytorch kernel during stage-1).</p>",
      "rawMarkdown": "Ok. Thanks. I rerun my kernel and the stage-2 result is reproducible, just like the stage-1 was. Since the training data did not change I am puzzled as to why I see CV change between this week and 6 days ago (i.e. when I last ran my pytorch kernel during stage-1).",
      "votes": null
    },
    {
      "id": "469438",
      "postDate": "02/11/2019 08:16:29",
      "content": "<p>Got it!</p>",
      "rawMarkdown": "Got it!",
      "votes": null
    },
    {
      "id": "469479",
      "postDate": "02/11/2019 09:36:51",
      "content": "<p>Hi YaGana,\ni did some tests over the weekend ... check this:\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\"> click me</a></p>",
      "rawMarkdown": "Hi YaGana,\ni did some tests over the weekend ... check this:\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\"> click me</a>",
      "votes": null
    },
    {
      "id": "469482",
      "postDate": "02/11/2019 09:54:56",
      "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a>, I just checked your update. It is strange seeing the huge time difference. The docker image change may be the reason in your case. It is hard to know how many packages are affected by that though. I did run my kernel twice with the stage-2 test data and the difference in time is only 40 seconds.</p>",
      "rawMarkdown": "springmanndaniel, I just checked your update. It is strange seeing the huge time difference. The docker image change may be the reason in your case. It is hard to know how many packages are affected by that though. I did run my kernel twice with the stage-2 test data and the difference in time is only 40 seconds.",
      "votes": null
    },
    {
      "id": "469492",
      "postDate": "02/11/2019 10:17:51",
      "content": "<p>Should have used python! </p>",
      "rawMarkdown": "Should have used python!",
      "votes": null
    },
    {
      "id": "469497",
      "postDate": "02/11/2019 10:31:06",
      "content": "<p>I usually use both R and python in order to create a diverse set of models but have not in this competition. You may still be okay. Lets hope for the best :-)</p>",
      "rawMarkdown": "I usually use both R and python in order to create a diverse set of models but have not in this competition. You may still be okay. Lets hope for the best :-)",
      "votes": null
    },
    {
      "id": "471296",
      "postDate": "02/14/2019 09:11:46",
      "content": "<p>The answer to my post is here <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80511\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80511</a>.</p>\n\n<p>I just read the competition finalized post by <a href=\"/inversion\">@inversion</a> at the link above. He confirmed in response to someone's question that they rerun our kernels with the stage-2 test which is 376k (total) = 56k (public) + 320k (private). I then checked my account again and saw my rerun kernel is shown right next to my selected kernel both with the same name but different scores. I had to sort them by scores to see them displayed next to each other. Yours may be different if you have a kernel with score that lies between your selection and the rerun.</p>",
      "rawMarkdown": "The answer to my post is here https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80511.\n\nI just read the competition finalized post by @inversion at the link above. He confirmed in response to someone's question that they rerun our kernels with the stage-2 test which is 376k (total) = 56k (public) + 320k (private). I then checked my account again and saw my rerun kernel is shown right next to my selected kernel both with the same name but different scores. I had to sort them by scores to see them displayed next to each other. Yours may be different if you have a kernel with score that lies between your selection and the rerun.",
      "votes": null
    },
    {
      "id": "471367",
      "postDate": "02/14/2019 11:05:54",
      "content": "<p>Seems like it all worked out! Congrats on placing in silver : )</p>",
      "rawMarkdown": "Seems like it all worked out! Congrats on placing in silver : )",
      "votes": null
    },
    {
      "id": "471385",
      "postDate": "02/14/2019 11:33:20",
      "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a> thanks and congrats to you too. I have been in silver but jumped up about 36 places. I have posted the original public LB pdf the day of stage-1 closing for people to check if they are curious. You can find the pdf file here <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80493\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80493</a></p>",
      "rawMarkdown": "springmanndaniel thanks and congrats to you too. I have been in silver but jumped up about 36 places. I have posted the original public LB pdf the day of stage-1 closing for people to check if they are curious. You can find the pdf file here https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80493",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 469041,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "02/10/2019 11:15:51",
      "content": "<p>The score may vary few points if you build vocabulary both with train and test but +0.23 ? You sure? I don't see any chance it possible except leaked validation process</p>",
      "votes": null,
      "replies": [
        {
          "id": 469053,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/10/2019 11:49:06",
          "content": "<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a>, I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 469088,
      "author_name": "suchith0312",
      "author_url": "",
      "post_date": "02/10/2019 13:26:38",
      "content": "<p>I ran my submitted kernels on the new test data(which is 4 days ago) and I am having same cv score. They are written in pytorch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 469103,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/10/2019 13:51:52",
          "content": "<p>@auchith0132, is it possible for you to run it again and let me know? I ran mine today and saw <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003</a> whilst my kernels were still running. He has screen shots in his post showing the docker difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 469413,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "02/11/2019 06:34:55",
          "content": "<p>I ran the models and my cv score remained same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 469422,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/11/2019 07:15:16",
          "content": "<p>Ok. Thanks. I rerun my kernel and the stage-2 result is reproducible, just like the stage-1 was. Since the training data did not change I am puzzled as to why I see CV change between this week and 6 days ago (i.e. when I last ran my pytorch kernel during stage-1).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 469104,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "02/10/2019 13:57:20",
      "content": "<p>i did run my pytorch cv but although my tokenizer was taking into account both training and test data, still local CV moved from 0.6915 to 0.69. Not big diff like yourself. </p>",
      "votes": null,
      "replies": [
        {
          "id": 469128,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/10/2019 14:55:49",
          "content": "<p><a href=\"/mlwhiz\">@mlwhiz</a>, that is the kind of CV difference I got from my Keras kernel but pytorch is really strange. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 469334,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "02/11/2019 02:24:15",
      "content": "<p>Hi YaGana,</p>\n\n<p>Not sure if I understand you correctly, but +0.23 means something like</p>\n\n<p>0.7 + 0.23 = 0.93??</p>",
      "votes": null,
      "replies": [
        {
          "id": 469380,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/11/2019 04:59:26",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I meant my cv validation loss went down from 52.4144  to 52.1763</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 469438,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "02/11/2019 08:16:29",
          "content": "<p>Got it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 469479,
      "author_name": "springmanndaniel",
      "author_url": "",
      "post_date": "02/11/2019 09:36:51",
      "content": "<p>Hi YaGana,\ni did some tests over the weekend ... check this:\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\"> click me</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 469482,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/11/2019 09:54:56",
          "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a>, I just checked your update. It is strange seeing the huge time difference. The docker image change may be the reason in your case. It is hard to know how many packages are affected by that though. I did run my kernel twice with the stage-2 test data and the difference in time is only 40 seconds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 469492,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "02/11/2019 10:17:51",
          "content": "<p>Should have used python! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 469497,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/11/2019 10:31:06",
          "content": "<p>I usually use both R and python in order to create a diverse set of models but have not in this competition. You may still be okay. Lets hope for the best :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 471367,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "02/14/2019 11:05:54",
          "content": "<p>Seems like it all worked out! Congrats on placing in silver : )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 471385,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/14/2019 11:33:20",
          "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a> thanks and congrats to you too. I have been in silver but jumped up about 36 places. I have posted the original public LB pdf the day of stage-1 closing for people to check if they are curious. You can find the pdf file here <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80493\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80493</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471296,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "02/14/2019 09:11:46",
      "content": "<p>The answer to my post is here <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80511\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80511</a>.</p>\n\n<p>I just read the competition finalized post by <a href=\"/inversion\">@inversion</a> at the link above. He confirmed in response to someone's question that they rerun our kernels with the stage-2 test which is 376k (total) = 56k (public) + 320k (private). I then checked my account again and saw my rerun kernel is shown right next to my selected kernel both with the same name but different scores. I had to sort them by scores to see them displayed next to each other. Yours may be different if you have a kernel with score that lies between your selection and the rerun.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "469021": "springmanndaniel posted [here][1] that there is a docker image change between stage-1 and stage-2 and I now wonder that explains my CV difference. His kernel is in R and mine is in python.\n\nI just ran both my kernels to see if they can complete within 2hrs which they did with time to spare. \n\nOne model is keras based so result is not strictly reproducible but my pytorch kernel is reproducible. When I ran the pytorch kernel today, there is an improvement of 0.23+ in CV even though the training data is still the same. I ran that kernel about 5 times during the competition and every single time the results for each epoch is exactly the same.\n\n**NOTE:** I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab.\n\nDid you observe the same thing? If so please share below.\n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003",
    "469041": "The score may vary few points if you build vocabulary both with train and test but +0.23 ? You sure? I don't see any chance it possible except leaked validation process",
    "469053": "yaroshevskiy, I made sure my vocabulary is built on only the training data. In fact I tested that during stage-1 and you can save roughly about 30 seconds by excluding stage-1 test data in your vocab.",
    "469088": "I ran my submitted kernels on the new test data(which is 4 days ago) and I am having same cv score. They are written in pytorch.",
    "469103": "auchith0132, is it possible for you to run it again and let me know? I ran mine today and saw https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003 whilst my kernels were still running. He has screen shots in his post showing the docker difference.",
    "469104": "i did run my pytorch cv but although my tokenizer was taking into account both training and test data, still local CV moved from 0.6915 to 0.69. Not big diff like yourself.",
    "469128": "mlwhiz, that is the kind of CV difference I got from my Keras kernel but pytorch is really strange.",
    "469334": "Hi YaGana,\n\nNot sure if I understand you correctly, but +0.23 means something like\n\n0.7 + 0.23 = 0.93??",
    "469380": "ratthachat, I meant my cv validation loss went down from 52.4144  to 52.1763",
    "469413": "I ran the models and my cv score remained same.",
    "469422": "Ok. Thanks. I rerun my kernel and the stage-2 result is reproducible, just like the stage-1 was. Since the training data did not change I am puzzled as to why I see CV change between this week and 6 days ago (i.e. when I last ran my pytorch kernel during stage-1).",
    "469438": "Got it!",
    "469479": "Hi YaGana,\ni did some tests over the weekend ... check this:\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80003\"> click me</a>",
    "469482": "springmanndaniel, I just checked your update. It is strange seeing the huge time difference. The docker image change may be the reason in your case. It is hard to know how many packages are affected by that though. I did run my kernel twice with the stage-2 test data and the difference in time is only 40 seconds.",
    "469492": "Should have used python!",
    "469497": "I usually use both R and python in order to create a diverse set of models but have not in this competition. You may still be okay. Lets hope for the best :-)",
    "471296": "The answer to my post is here https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80511.\n\nI just read the competition finalized post by @inversion at the link above. He confirmed in response to someone's question that they rerun our kernels with the stage-2 test which is 376k (total) = 56k (public) + 320k (private). I then checked my account again and saw my rerun kernel is shown right next to my selected kernel both with the same name but different scores. I had to sort them by scores to see them displayed next to each other. Yours may be different if you have a kernel with score that lies between your selection and the rerun.",
    "471367": "Seems like it all worked out! Congrats on placing in silver : )",
    "471385": "springmanndaniel thanks and congrats to you too. I have been in silver but jumped up about 36 places. I have posted the original public LB pdf the day of stage-1 closing for people to check if they are curious. You can find the pdf file here https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80493"
  },
  "source": "meta"
}