{
  "id": 79459,
  "title": "3 Critical factors that will affect stage-2 score",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79459",
  "author_name": "",
  "post_date": "2019-02-04T10:27:31.539741900Z",
  "votes": 30,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi my dear fellows, \nWe are in the last phase of the competition, and I am happy to have worked with all of you for this 3 months of marathon! </p>\n\n<p>As the end is coming soon, me and my teammate have been trying to do our best in stage-2. We found out that these 3 factors are the most critical to us. In this last period of time, we hope that this sharing will benefit some of you.</p>\n\n<p>1) <strong>Variation of LB and local test.</strong> This issue is well known to all teams here. We think that the best way is to combine Benjamin's excellent kernel on local test with shake up simulation.\n<a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a>\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821</a></p>\n\n<p>2) <strong>Running time exceeding 2 hours</strong>. At first, we were confidence that we can deal with this by saving buffer around 600s. However, we were wrong! In our experiments, we found out that the worst running time can exceed the best running time as much as 800 seconds!!  We have tried our best but we cannot solve this issue yet. If we reduce our time further, our performance will drop as well! (from the picture, we have a 1/8 chance of runtime failure ... But during the peak time, this number may get worse)\n<img src=\"https://i.ibb.co/TkxCRXw/runtest.jpg\" alt=\"See the picture\">\nBy the way, don't forget to test your stage-2 time with test data around 56k*7 = 390k </p>\n\n<p>3) <strong>Too much sensitivity</strong> As discussed in \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75753\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75753</a>\nJust by changing only slightly to any variables, our performance change can be quite large. In fact, the change magnitude is around 1-2 sigma (0.003-0.005), which is statistically OK unless 0.001 is so crucial in this competition.</p>\n\n<p>This sensitivity problem is not just for LB, <strong>we found out that even local test and shakeup simulation are also highly sensitive too</strong>. Nevertheless, if your algorithm use exactly the same information in both stage-1 and stage-2, you may be fine!  </p>\n\n<p>However, there can be the case of blindly using different information on two stages too. For example, in our case, we just found out that we use the vocabulary from both the training and test sets. And the set of vocabs is determined by Keras tokenizer. In stage-2, as  test data will change, Keras tokenizer will likely to change the set of best vocab set too. And because of this high sensitivity in scoring, we are likely doom now.</p>\n\n<p>Good luck everyone!</p>",
  "messages": [
    {
      "id": "465920",
      "postDate": "02/04/2019 10:27:31",
      "content": "<p>Hi my dear fellows, \nWe are in the last phase of the competition, and I am happy to have worked with all of you for this 3 months of marathon! </p>\n\n<p>As the end is coming soon, me and my teammate have been trying to do our best in stage-2. We found out that these 3 factors are the most critical to us. In this last period of time, we hope that this sharing will benefit some of you.</p>\n\n<p>1) <strong>Variation of LB and local test.</strong> This issue is well known to all teams here. We think that the best way is to combine Benjamin's excellent kernel on local test with shake up simulation.\n<a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a>\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821</a></p>\n\n<p>2) <strong>Running time exceeding 2 hours</strong>. At first, we were confidence that we can deal with this by saving buffer around 600s. However, we were wrong! In our experiments, we found out that the worst running time can exceed the best running time as much as 800 seconds!!  We have tried our best but we cannot solve this issue yet. If we reduce our time further, our performance will drop as well! (from the picture, we have a 1/8 chance of runtime failure ... But during the peak time, this number may get worse)\n<img src=\"https://i.ibb.co/TkxCRXw/runtest.jpg\" alt=\"See the picture\">\nBy the way, don't forget to test your stage-2 time with test data around 56k*7 = 390k </p>\n\n<p>3) <strong>Too much sensitivity</strong> As discussed in \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75753\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75753</a>\nJust by changing only slightly to any variables, our performance change can be quite large. In fact, the change magnitude is around 1-2 sigma (0.003-0.005), which is statistically OK unless 0.001 is so crucial in this competition.</p>\n\n<p>This sensitivity problem is not just for LB, <strong>we found out that even local test and shakeup simulation are also highly sensitive too</strong>. Nevertheless, if your algorithm use exactly the same information in both stage-1 and stage-2, you may be fine!  </p>\n\n<p>However, there can be the case of blindly using different information on two stages too. For example, in our case, we just found out that we use the vocabulary from both the training and test sets. And the set of vocabs is determined by Keras tokenizer. In stage-2, as  test data will change, Keras tokenizer will likely to change the set of best vocab set too. And because of this high sensitivity in scoring, we are likely doom now.</p>\n\n<p>Good luck everyone!</p>",
      "rawMarkdown": "Hi my dear fellows, \nWe are in the last phase of the competition, and I am happy to have worked with all of you for this 3 months of marathon! \n\nAs the end is coming soon, me and my teammate have been trying to do our best in stage-2. We found out that these 3 factors are the most critical to us. In this last period of time, we hope that this sharing will benefit some of you.\n\n1) **Variation of LB and local test.** This issue is well known to all teams here. We think that the best way is to combine Benjamin's excellent kernel on local test with shake up simulation.\nhttps://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\n\n2) **Running time exceeding 2 hours**. At first, we were confidence that we can deal with this by saving buffer around 600s. However, we were wrong! In our experiments, we found out that the worst running time can exceed the best running time as much as 800 seconds!!  We have tried our best but we cannot solve this issue yet. If we reduce our time further, our performance will drop as well! (from the picture, we have a 1/8 chance of runtime failure ... But during the peak time, this number may get worse)\n![See the picture][1]\n[1]: https://i.ibb.co/TkxCRXw/runtest.jpg\n\nBy the way, don't forget to test your stage-2 time with test data around 56k*7 = 390k \n\n3) **Too much sensitivity** As discussed in \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75753\nJust by changing only slightly to any variables, our performance change can be quite large. In fact, the change magnitude is around 1-2 sigma (0.003-0.005), which is statistically OK unless 0.001 is so crucial in this competition.\n\nThis sensitivity problem is not just for LB, **we found out that even local test and shakeup simulation are also highly sensitive too**. Nevertheless, if your algorithm use exactly the same information in both stage-1 and stage-2, you may be fine!  \n\nHowever, there can be the case of blindly using different information on two stages too. For example, in our case, we just found out that we use the vocabulary from both the training and test sets. And the set of vocabs is determined by Keras tokenizer. In stage-2, as  test data will change, Keras tokenizer will likely to change the set of best vocab set too. And because of this high sensitivity in scoring, we are likely doom now.\n\nGood luck everyone!",
      "votes": null
    },
    {
      "id": "466376",
      "postDate": "02/05/2019 08:45:55",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> thanks for the tips/ reminder. </p>\n\n<p>&gt; 2) Running time exceeding 2 hours. </p>\n\n<p>On the point above, I am hoping this is the reason Kaggle is going to run our scripts in batches and over the next 3 weeks. Hopefully the time variations will not be as bad as we see in your screenshot.</p>\n\n<p>On the point about the keras tokenizer, as long as you have set the num_words to some value, the effect may be negligible.</p>",
      "rawMarkdown": "ratthachat thanks for the tips/ reminder. \n\n&gt; 2) Running time exceeding 2 hours. \n\nOn the point above, I am hoping this is the reason Kaggle is going to run our scripts in batches and over the next 3 weeks. Hopefully the time variations will not be as bad as we see in your screenshot.\n\nOn the point about the keras tokenizer, as long as you have set the num_words to some value, the effect may be negligible.",
      "votes": null
    },
    {
      "id": "466383",
      "postDate": "02/05/2019 09:09:58",
      "content": "<p>Hi YaGana!  Haven’t seen you in this forum for a while (but seeing you in other competitions instead ;)</p>\n\n<p>About 2 hours limitation/variation, at the end of the day, it comes down to ‘<strong>risk tolerant</strong>’ and ‘<strong>risk management</strong>’ scheme for each one of us. If we are optimistic, i.e.  the running conditions should be better on the private test or kaggle will relax some testing conditions later, we may perhap content with around 500s buffer.</p>\n\n<p>In our case, we choose the middle path, we sacrifice our best version (best local test with has 40-50% chance of running time exceeding), and choose the mid version (moderate local test with 20-30% prob of time exceeding, risky to sensitivity), and the safest version (relatively low local test, with 0% prob of time exceeding, no sensitivity) — If at the end, kaggle choose to relax some testing conditions, our 2nd choice may be not a wise one. :)</p>\n\n<p>About the last point, we found out that even the number of vocabs is fixed, the effect is indeed severe to us (~0.003 performance drop :( ) —- As side note, 0.003 is actually statistically negligible because in our shakeup simulations, we got 0.003 as standard deviation on simulated LB.</p>",
      "rawMarkdown": "Hi YaGana!  Haven’t seen you in this forum for a while (but seeing you in other competitions instead ;)\n\nAbout 2 hours limitation/variation, at the end of the day, it comes down to ‘**risk tolerant**’ and ‘**risk management**’ scheme for each one of us. If we are optimistic, i.e.  the running conditions should be better on the private test or kaggle will relax some testing conditions later, we may perhap content with around 500s buffer.\n\nIn our case, we choose the middle path, we sacrifice our best version (best local test with has 40-50% chance of running time exceeding), and choose the mid version (moderate local test with 20-30% prob of time exceeding, risky to sensitivity), and the safest version (relatively low local test, with 0% prob of time exceeding, no sensitivity) — If at the end, kaggle choose to relax some testing conditions, our 2nd choice may be not a wise one. :)\n\nAbout the last point, we found out that even the number of vocabs is fixed, the effect is indeed severe to us (~0.003 performance drop :( ) —- As side note, 0.003 is actually statistically negligible because in our shakeup simulations, we got 0.003 as standard deviation on simulated LB.",
      "votes": null
    },
    {
      "id": "466392",
      "postDate": "02/05/2019 09:37:03",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I am not argueing your buffer recommendation, you are definitely right on that. I myself allowed about 750 seconds buffer. I have been around just really busy hence not been posting much :-)</p>\n\n<p>I  do not use vocab the same way as in the public kernels in the models I have selected for stage-2 evaluation. However, after reading your post, I rerun my fork of one of the public kernels with only the vocab changed to train data and got a slightly lower score. The difference in time is about 27 seconds which can be significant with the large stage-2 test data. So people should definitely take that into consideration.</p>",
      "rawMarkdown": "ratthachat, I am not argueing your buffer recommendation, you are definitely right on that. I myself allowed about 750 seconds buffer. I have been around just really busy hence not been posting much :-)\n\nI  do not use vocab the same way as in the public kernels in the models I have selected for stage-2 evaluation. However, after reading your post, I rerun my fork of one of the public kernels with only the vocab changed to train data and got a slightly lower score. The difference in time is about 27 seconds which can be significant with the large stage-2 test data. So people should definitely take that into consideration.",
      "votes": null
    },
    {
      "id": "466393",
      "postDate": "02/05/2019 09:42:06",
      "content": "<p>Thanks YaGana for sharing your experiment!  </p>\n\n<p>Just want to note that I didn’t intend to make any arguments too (I am totally think that choice of risk tolerance is arbitrary —- for somebody 500s buffer may be perfectly acceptable)</p>\n\n<p>Sorry again if the tone of the above post is like an argument; in fact, I just wanted to have a discussion with you ;)</p>",
      "rawMarkdown": "Thanks YaGana for sharing your experiment!  \n\nJust want to note that I didn’t intend to make any arguments too (I am totally think that choice of risk tolerance is arbitrary —- for somebody 500s buffer may be perfectly acceptable)\n\nSorry again if the tone of the above post is like an argument; in fact, I just wanted to have a discussion with you ;)",
      "votes": null
    },
    {
      "id": "466394",
      "postDate": "02/05/2019 09:52:51",
      "content": "<p>No problem, your tone was fine, I was just clarifying my point. No arguement as in I cannot say it any better than you did :-)</p>\n\n<p>Lets hope we both jump up on the private LB and not down.</p>",
      "rawMarkdown": "No problem, your tone was fine, I was just clarifying my point. No arguement as in I cannot say it any better than you did :-)\n\nLets hope we both jump up on the private LB and not down.",
      "votes": null
    },
    {
      "id": "466513",
      "postDate": "02/05/2019 14:45:37",
      "content": "<p>Thanks, tips are very well explained.</p>",
      "rawMarkdown": "Thanks, tips are very well explained.",
      "votes": null
    },
    {
      "id": "471194",
      "postDate": "02/14/2019 06:14:20",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, congratulations on the 61 points jump. We did not fall. Phew!!! :-)</p>",
      "rawMarkdown": "ratthachat, congratulations on the 61 points jump. We did not fall. Phew!!! :-)",
      "votes": null
    },
    {
      "id": "471473",
      "postDate": "02/14/2019 14:11:51",
      "content": "<p>Thanks YaGana, <a href=\"/sheriytm\">@sheriytm</a>!!</p>\n\n<p>Congratulate to you too for your yet again solid performance!\nNice to talk to you here! And see you again soon in VSB ;) </p>",
      "rawMarkdown": "Thanks YaGana, @sheriytm!!\n\nCongratulate to you too for your yet again solid performance!\nNice to talk to you here! And see you again soon in VSB ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 466376,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "02/05/2019 08:45:55",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> thanks for the tips/ reminder. </p>\n\n<p>&gt; 2) Running time exceeding 2 hours. </p>\n\n<p>On the point above, I am hoping this is the reason Kaggle is going to run our scripts in batches and over the next 3 weeks. Hopefully the time variations will not be as bad as we see in your screenshot.</p>\n\n<p>On the point about the keras tokenizer, as long as you have set the num_words to some value, the effect may be negligible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 466383,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "02/05/2019 09:09:58",
          "content": "<p>Hi YaGana!  Haven’t seen you in this forum for a while (but seeing you in other competitions instead ;)</p>\n\n<p>About 2 hours limitation/variation, at the end of the day, it comes down to ‘<strong>risk tolerant</strong>’ and ‘<strong>risk management</strong>’ scheme for each one of us. If we are optimistic, i.e.  the running conditions should be better on the private test or kaggle will relax some testing conditions later, we may perhap content with around 500s buffer.</p>\n\n<p>In our case, we choose the middle path, we sacrifice our best version (best local test with has 40-50% chance of running time exceeding), and choose the mid version (moderate local test with 20-30% prob of time exceeding, risky to sensitivity), and the safest version (relatively low local test, with 0% prob of time exceeding, no sensitivity) — If at the end, kaggle choose to relax some testing conditions, our 2nd choice may be not a wise one. :)</p>\n\n<p>About the last point, we found out that even the number of vocabs is fixed, the effect is indeed severe to us (~0.003 performance drop :( ) —- As side note, 0.003 is actually statistically negligible because in our shakeup simulations, we got 0.003 as standard deviation on simulated LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466392,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/05/2019 09:37:03",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I am not argueing your buffer recommendation, you are definitely right on that. I myself allowed about 750 seconds buffer. I have been around just really busy hence not been posting much :-)</p>\n\n<p>I  do not use vocab the same way as in the public kernels in the models I have selected for stage-2 evaluation. However, after reading your post, I rerun my fork of one of the public kernels with only the vocab changed to train data and got a slightly lower score. The difference in time is about 27 seconds which can be significant with the large stage-2 test data. So people should definitely take that into consideration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466393,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "02/05/2019 09:42:06",
          "content": "<p>Thanks YaGana for sharing your experiment!  </p>\n\n<p>Just want to note that I didn’t intend to make any arguments too (I am totally think that choice of risk tolerance is arbitrary —- for somebody 500s buffer may be perfectly acceptable)</p>\n\n<p>Sorry again if the tone of the above post is like an argument; in fact, I just wanted to have a discussion with you ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466394,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/05/2019 09:52:51",
          "content": "<p>No problem, your tone was fine, I was just clarifying my point. No arguement as in I cannot say it any better than you did :-)</p>\n\n<p>Lets hope we both jump up on the private LB and not down.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 471194,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/14/2019 06:14:20",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, congratulations on the 61 points jump. We did not fall. Phew!!! :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 471473,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "02/14/2019 14:11:51",
          "content": "<p>Thanks YaGana, <a href=\"/sheriytm\">@sheriytm</a>!!</p>\n\n<p>Congratulate to you too for your yet again solid performance!\nNice to talk to you here! And see you again soon in VSB ;) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 466513,
      "author_name": "mvs1793",
      "author_url": "",
      "post_date": "02/05/2019 14:45:37",
      "content": "<p>Thanks, tips are very well explained.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "465920": "Hi my dear fellows, \nWe are in the last phase of the competition, and I am happy to have worked with all of you for this 3 months of marathon! \n\nAs the end is coming soon, me and my teammate have been trying to do our best in stage-2. We found out that these 3 factors are the most critical to us. In this last period of time, we hope that this sharing will benefit some of you.\n\n1) **Variation of LB and local test.** This issue is well known to all teams here. We think that the best way is to combine Benjamin's excellent kernel on local test with shake up simulation.\nhttps://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\n\n2) **Running time exceeding 2 hours**. At first, we were confidence that we can deal with this by saving buffer around 600s. However, we were wrong! In our experiments, we found out that the worst running time can exceed the best running time as much as 800 seconds!!  We have tried our best but we cannot solve this issue yet. If we reduce our time further, our performance will drop as well! (from the picture, we have a 1/8 chance of runtime failure ... But during the peak time, this number may get worse)\n![See the picture][1]\n[1]: https://i.ibb.co/TkxCRXw/runtest.jpg\n\nBy the way, don't forget to test your stage-2 time with test data around 56k*7 = 390k \n\n3) **Too much sensitivity** As discussed in \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75753\nJust by changing only slightly to any variables, our performance change can be quite large. In fact, the change magnitude is around 1-2 sigma (0.003-0.005), which is statistically OK unless 0.001 is so crucial in this competition.\n\nThis sensitivity problem is not just for LB, **we found out that even local test and shakeup simulation are also highly sensitive too**. Nevertheless, if your algorithm use exactly the same information in both stage-1 and stage-2, you may be fine!  \n\nHowever, there can be the case of blindly using different information on two stages too. For example, in our case, we just found out that we use the vocabulary from both the training and test sets. And the set of vocabs is determined by Keras tokenizer. In stage-2, as  test data will change, Keras tokenizer will likely to change the set of best vocab set too. And because of this high sensitivity in scoring, we are likely doom now.\n\nGood luck everyone!",
    "466376": "ratthachat thanks for the tips/ reminder. \n\n&gt; 2) Running time exceeding 2 hours. \n\nOn the point above, I am hoping this is the reason Kaggle is going to run our scripts in batches and over the next 3 weeks. Hopefully the time variations will not be as bad as we see in your screenshot.\n\nOn the point about the keras tokenizer, as long as you have set the num_words to some value, the effect may be negligible.",
    "466383": "Hi YaGana!  Haven’t seen you in this forum for a while (but seeing you in other competitions instead ;)\n\nAbout 2 hours limitation/variation, at the end of the day, it comes down to ‘**risk tolerant**’ and ‘**risk management**’ scheme for each one of us. If we are optimistic, i.e.  the running conditions should be better on the private test or kaggle will relax some testing conditions later, we may perhap content with around 500s buffer.\n\nIn our case, we choose the middle path, we sacrifice our best version (best local test with has 40-50% chance of running time exceeding), and choose the mid version (moderate local test with 20-30% prob of time exceeding, risky to sensitivity), and the safest version (relatively low local test, with 0% prob of time exceeding, no sensitivity) — If at the end, kaggle choose to relax some testing conditions, our 2nd choice may be not a wise one. :)\n\nAbout the last point, we found out that even the number of vocabs is fixed, the effect is indeed severe to us (~0.003 performance drop :( ) —- As side note, 0.003 is actually statistically negligible because in our shakeup simulations, we got 0.003 as standard deviation on simulated LB.",
    "466392": "ratthachat, I am not argueing your buffer recommendation, you are definitely right on that. I myself allowed about 750 seconds buffer. I have been around just really busy hence not been posting much :-)\n\nI  do not use vocab the same way as in the public kernels in the models I have selected for stage-2 evaluation. However, after reading your post, I rerun my fork of one of the public kernels with only the vocab changed to train data and got a slightly lower score. The difference in time is about 27 seconds which can be significant with the large stage-2 test data. So people should definitely take that into consideration.",
    "466393": "Thanks YaGana for sharing your experiment!  \n\nJust want to note that I didn’t intend to make any arguments too (I am totally think that choice of risk tolerance is arbitrary —- for somebody 500s buffer may be perfectly acceptable)\n\nSorry again if the tone of the above post is like an argument; in fact, I just wanted to have a discussion with you ;)",
    "466394": "No problem, your tone was fine, I was just clarifying my point. No arguement as in I cannot say it any better than you did :-)\n\nLets hope we both jump up on the private LB and not down.",
    "466513": "Thanks, tips are very well explained.",
    "471194": "ratthachat, congratulations on the 61 points jump. We did not fall. Phew!!! :-)",
    "471473": "Thanks YaGana, @sheriytm!!\n\nCongratulate to you too for your yet again solid performance!\nNice to talk to you here! And see you again soon in VSB ;)"
  },
  "source": "meta"
}