{
  "id": 76271,
  "title": "[Updated with CV results] Simulating shakeup from a simulated stage 2 test set",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76271",
  "author_name": "",
  "post_date": "2018-12-31T11:49:24.127351Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Update in comment</h1>\n\n<p>Hi,</p>\n\n<p>I tried to simulate the stage 2 test set shakeup with the ideas from my previous discussion post <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821</a></p>\n\n<h3>Code:</h3>\n\n<p>All my code can be found in this kernel: <a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-meta-features-shakeup-prediction\">https://www.kaggle.com/bkkaggle/skip-rnn-meta-features-shakeup-prediction</a></p>\n\n<h3>Results</h3>\n\n<p>Val F1: 0.6775</p>\n\n<p>Simulated stage 1 test set: 0.6680</p>\n\n<p>Simulated stage 2 test set: 0.6726</p>\n\n<p>I haven't implemented a cross validation version of this kernel yet, but the stage 1 test score is a lot lower than the val F1 score which is what a lot of people on the discussion forum have been saying.</p>\n\n<p>The stage 2 test score is a slightly lower than  the val F1 score, so I believe that following the local CV F1 score is the best way to make sure that there is a minimal amount of shakeup at the end of the competition and that the public LB doesn't give anyone a reasonable idea of their stage 2 positions</p>\n\nUpdate:\n\n<p>There was a small bug in my previous code so the new shakeup results for a single model are:</p>\n\n<p>Val F1: 0.6775</p>\n\n<p>Simulated stage 1 test set: 0.6624</p>\n\n<p>Simulated stage 2 test set: 0.6662</p>",
  "messages": [
    {
      "id": "448179",
      "postDate": "12/31/2018 11:49:24",
      "content": "<h1>Update in comment</h1>\n\n<p>Hi,</p>\n\n<p>I tried to simulate the stage 2 test set shakeup with the ideas from my previous discussion post <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821</a></p>\n\n<h3>Code:</h3>\n\n<p>All my code can be found in this kernel: <a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-meta-features-shakeup-prediction\">https://www.kaggle.com/bkkaggle/skip-rnn-meta-features-shakeup-prediction</a></p>\n\n<h3>Results</h3>\n\n<p>Val F1: 0.6775</p>\n\n<p>Simulated stage 1 test set: 0.6680</p>\n\n<p>Simulated stage 2 test set: 0.6726</p>\n\n<p>I haven't implemented a cross validation version of this kernel yet, but the stage 1 test score is a lot lower than the val F1 score which is what a lot of people on the discussion forum have been saying.</p>\n\n<p>The stage 2 test score is a slightly lower than  the val F1 score, so I believe that following the local CV F1 score is the best way to make sure that there is a minimal amount of shakeup at the end of the competition and that the public LB doesn't give anyone a reasonable idea of their stage 2 positions</p>\n\nUpdate:\n\n<p>There was a small bug in my previous code so the new shakeup results for a single model are:</p>\n\n<p>Val F1: 0.6775</p>\n\n<p>Simulated stage 1 test set: 0.6624</p>\n\n<p>Simulated stage 2 test set: 0.6662</p>",
      "rawMarkdown": "# Update in comment\n\nHi,\n\nI tried to simulate the stage 2 test set shakeup with the ideas from my previous discussion post https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\n\n### Code:\n\nAll my code can be found in this kernel: https://www.kaggle.com/bkkaggle/skip-rnn-meta-features-shakeup-prediction\n\n### Results\n\nVal F1: 0.6775\n\nSimulated stage 1 test set: 0.6680\n\nSimulated stage 2 test set: 0.6726\n\nI haven't implemented a cross validation version of this kernel yet, but the stage 1 test score is a lot lower than the val F1 score which is what a lot of people on the discussion forum have been saying.\n\nThe stage 2 test score is a slightly lower than  the val F1 score, so I believe that following the local CV F1 score is the best way to make sure that there is a minimal amount of shakeup at the end of the competition and that the public LB doesn't give anyone a reasonable idea of their stage 2 positions\n\n#### Update:\nThere was a small bug in my previous code so the new shakeup results for a single model are:\n\nVal F1: 0.6775\n\nSimulated stage 1 test set: 0.6624\n\nSimulated stage 2 test set: 0.6662",
      "votes": null
    },
    {
      "id": "449705",
      "postDate": "01/03/2019 16:20:29",
      "content": "<h1>UPDATE:</h1>\n\n<h2>Cross Validation results</h2>\n\n<p>I ran the same shakeup prediction code on one of my CV kernels. To predict the range of the possible shakeup, I created 4 kernels that run the entire cross validation pipeline 5 times each with different random seeds for a total of 20 cv runs with different seeds and another kernel to aggregate their results.</p>\n\n<p>The five kernels are:</p>\n\n<p><a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds</a> <br>\n<a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds-5-10\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds-5-10</a> <br>\n<a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-10-15\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-10-15</a> <br>\n<a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-15-20\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-15-20</a></p>\n\n<p>The aggregate kernel is:</p>\n\n<p><a href=\"https://www.kaggle.com/bkkaggle/combined-shakeup-prediction\">https://www.kaggle.com/bkkaggle/combined-shakeup-prediction</a></p>\n\n<h2>Results</h2>\n\n<p>Mean and Standard Deviation  </p>\n\n<p>Stage 1 Val F1:</p>\n\n<p>0.676987, 0.001124</p>\n\n<p>Stage 1 Test F1:</p>\n\n<p>0.678823, 0.002044</p>\n\n<p>Stage 2 Test F1:</p>\n\n<p>0.686936, 0.002318</p>\n\n<p>Val - Stage 1 Shakeup:</p>\n\n<p>+0.001836, 0.002075</p>\n\n<p>Val - Stage 2 Shakeup:</p>\n\n<p>+0.009949, 0.002339</p>\n\n<p>Stage 1 - Stage 2 Shakeup:</p>\n\n<p>+0.008112, 0.001662</p>\n\n<p>Overall, my results on the stage 2 test set are higher than on the stage 1 test and the val sets and the shakeup between the val and stage 2 test sets is ~0.01 with a std of 0.002.  It should be possible to use it to estimate shakeup on some of the high scoring public kernels. I also made a separate kernel with only the shakeup prediction code as a template if anyone wants to check their possible stage 2 shakeup, <a href=\"https://www.kaggle.com/bkkaggle/shakeup-prediction-template\">https://www.kaggle.com/bkkaggle/shakeup-prediction-template</a></p>",
      "rawMarkdown": "#UPDATE:\n\n## Cross Validation results\n\nI ran the same shakeup prediction code on one of my CV kernels. To predict the range of the possible shakeup, I created 4 kernels that run the entire cross validation pipeline 5 times each with different random seeds for a total of 20 cv runs with different seeds and another kernel to aggregate their results.\n\nThe five kernels are:\n\nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds  \nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds-5-10  \nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-10-15  \nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-15-20\n\nThe aggregate kernel is:\n\nhttps://www.kaggle.com/bkkaggle/combined-shakeup-prediction\n\n## Results\n\nMean and Standard Deviation  \n\nStage 1 Val F1:\n\n0.676987, 0.001124\n\nStage 1 Test F1:\n\n0.678823, 0.002044\n\nStage 2 Test F1:\n\n0.686936, 0.002318\n\nVal - Stage 1 Shakeup:\n\n+0.001836, 0.002075\n\nVal - Stage 2 Shakeup:\n\n+0.009949, 0.002339\n\nStage 1 - Stage 2 Shakeup:\n\n+0.008112, 0.001662\n\nOverall, my results on the stage 2 test set are higher than on the stage 1 test and the val sets and the shakeup between the val and stage 2 test sets is ~0.01 with a std of 0.002.  It should be possible to use it to estimate shakeup on some of the high scoring public kernels. I also made a separate kernel with only the shakeup prediction code as a template if anyone wants to check their possible stage 2 shakeup, https://www.kaggle.com/bkkaggle/shakeup-prediction-template",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 449705,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "01/03/2019 16:20:29",
      "content": "<h1>UPDATE:</h1>\n\n<h2>Cross Validation results</h2>\n\n<p>I ran the same shakeup prediction code on one of my CV kernels. To predict the range of the possible shakeup, I created 4 kernels that run the entire cross validation pipeline 5 times each with different random seeds for a total of 20 cv runs with different seeds and another kernel to aggregate their results.</p>\n\n<p>The five kernels are:</p>\n\n<p><a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds</a> <br>\n<a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds-5-10\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds-5-10</a> <br>\n<a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-10-15\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-10-15</a> <br>\n<a href=\"https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-15-20\">https://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-15-20</a></p>\n\n<p>The aggregate kernel is:</p>\n\n<p><a href=\"https://www.kaggle.com/bkkaggle/combined-shakeup-prediction\">https://www.kaggle.com/bkkaggle/combined-shakeup-prediction</a></p>\n\n<h2>Results</h2>\n\n<p>Mean and Standard Deviation  </p>\n\n<p>Stage 1 Val F1:</p>\n\n<p>0.676987, 0.001124</p>\n\n<p>Stage 1 Test F1:</p>\n\n<p>0.678823, 0.002044</p>\n\n<p>Stage 2 Test F1:</p>\n\n<p>0.686936, 0.002318</p>\n\n<p>Val - Stage 1 Shakeup:</p>\n\n<p>+0.001836, 0.002075</p>\n\n<p>Val - Stage 2 Shakeup:</p>\n\n<p>+0.009949, 0.002339</p>\n\n<p>Stage 1 - Stage 2 Shakeup:</p>\n\n<p>+0.008112, 0.001662</p>\n\n<p>Overall, my results on the stage 2 test set are higher than on the stage 1 test and the val sets and the shakeup between the val and stage 2 test sets is ~0.01 with a std of 0.002.  It should be possible to use it to estimate shakeup on some of the high scoring public kernels. I also made a separate kernel with only the shakeup prediction code as a template if anyone wants to check their possible stage 2 shakeup, <a href=\"https://www.kaggle.com/bkkaggle/shakeup-prediction-template\">https://www.kaggle.com/bkkaggle/shakeup-prediction-template</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "448179": "# Update in comment\n\nHi,\n\nI tried to simulate the stage 2 test set shakeup with the ideas from my previous discussion post https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75821\n\n### Code:\n\nAll my code can be found in this kernel: https://www.kaggle.com/bkkaggle/skip-rnn-meta-features-shakeup-prediction\n\n### Results\n\nVal F1: 0.6775\n\nSimulated stage 1 test set: 0.6680\n\nSimulated stage 2 test set: 0.6726\n\nI haven't implemented a cross validation version of this kernel yet, but the stage 1 test score is a lot lower than the val F1 score which is what a lot of people on the discussion forum have been saying.\n\nThe stage 2 test score is a slightly lower than  the val F1 score, so I believe that following the local CV F1 score is the best way to make sure that there is a minimal amount of shakeup at the end of the competition and that the public LB doesn't give anyone a reasonable idea of their stage 2 positions\n\n#### Update:\nThere was a small bug in my previous code so the new shakeup results for a single model are:\n\nVal F1: 0.6775\n\nSimulated stage 1 test set: 0.6624\n\nSimulated stage 2 test set: 0.6662",
    "449705": "#UPDATE:\n\n## Cross Validation results\n\nI ran the same shakeup prediction code on one of my CV kernels. To predict the range of the possible shakeup, I created 4 kernels that run the entire cross validation pipeline 5 times each with different random seeds for a total of 20 cv runs with different seeds and another kernel to aggregate their results.\n\nThe five kernels are:\n\nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds  \nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-seeds-5-10  \nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-10-15  \nhttps://www.kaggle.com/bkkaggle/skip-rnn-cv-meta-shakeup-15-20\n\nThe aggregate kernel is:\n\nhttps://www.kaggle.com/bkkaggle/combined-shakeup-prediction\n\n## Results\n\nMean and Standard Deviation  \n\nStage 1 Val F1:\n\n0.676987, 0.001124\n\nStage 1 Test F1:\n\n0.678823, 0.002044\n\nStage 2 Test F1:\n\n0.686936, 0.002318\n\nVal - Stage 1 Shakeup:\n\n+0.001836, 0.002075\n\nVal - Stage 2 Shakeup:\n\n+0.009949, 0.002339\n\nStage 1 - Stage 2 Shakeup:\n\n+0.008112, 0.001662\n\nOverall, my results on the stage 2 test set are higher than on the stage 1 test and the val sets and the shakeup between the val and stage 2 test sets is ~0.01 with a std of 0.002.  It should be possible to use it to estimate shakeup on some of the high scoring public kernels. I also made a separate kernel with only the shakeup prediction code as a template if anyone wants to check their possible stage 2 shakeup, https://www.kaggle.com/bkkaggle/shakeup-prediction-template"
  },
  "source": "meta"
}