{
  "id": 78952,
  "title": "simulation experiment about shake up",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78952",
  "author_name": "",
  "post_date": "2019-01-29T11:37:18.675297100Z",
  "votes": 15,
  "comment_count": 11,
  "views": 0,
  "content": "<p>try this code, pred_train_y is the oof pred and train_y Is lable. </p>\n\n<p>It sample a same size dataset as public.</p>\n\n<pre><code>shake_up = []\nsub_pd = pd.DataFrame({\n    'pred':pred_train_y,\n    'label':train_y\n})\nfor i in range(50):\n    sample_t = sub_pd.sample(n=56370, random_state=i)\n    sr = metrics.f1_score(sample_t['label'], (sample_t['pred'] &amp;gt; best_thresh).astype(int))\n    print(sr)\n    shake_up.append(sr)\n\nprint(max(shake_up))\nprint(min(shake_up))\n</code></pre>\n\n<p>I do this and a get :\n<strong>it varies from 0.6889 to 0.7106</strong></p>",
  "messages": [
    {
      "id": "463089",
      "postDate": "01/29/2019 11:37:18",
      "content": "<p>try this code, pred_train_y is the oof pred and train_y Is lable. </p>\n\n<p>It sample a same size dataset as public.</p>\n\n<pre><code>shake_up = []\nsub_pd = pd.DataFrame({\n    'pred':pred_train_y,\n    'label':train_y\n})\nfor i in range(50):\n    sample_t = sub_pd.sample(n=56370, random_state=i)\n    sr = metrics.f1_score(sample_t['label'], (sample_t['pred'] &amp;gt; best_thresh).astype(int))\n    print(sr)\n    shake_up.append(sr)\n\nprint(max(shake_up))\nprint(min(shake_up))\n</code></pre>\n\n<p>I do this and a get :\n<strong>it varies from 0.6889 to 0.7106</strong></p>",
      "rawMarkdown": "try this code, pred\\_train\\_y is the oof pred and train\\_y Is lable. \n\nIt sample a same size dataset as public.\n\n    shake_up = []\n    sub_pd = pd.DataFrame({\n        'pred':pred_train_y,\n        'label':train_y\n    })\n    for i in range(50):\n        sample_t = sub_pd.sample(n=56370, random_state=i)\n        sr = metrics.f1_score(sample_t['label'], (sample_t['pred'] &gt; best_thresh).astype(int))\n        print(sr)\n        shake_up.append(sr)\n        \n    print(max(shake_up))\n    print(min(shake_up))\n\nI do this and a get :\n**it varies from 0.6889 to 0.7106**",
      "votes": null
    },
    {
      "id": "463094",
      "postDate": "01/29/2019 11:46:15",
      "content": "<p>Wow, the gap is soooooooooooooo large, scary. All you need is the faith in cv. LOL</p>",
      "rawMarkdown": "Wow, the gap is soooooooooooooo large, scary. All you need is the faith in cv. LOL",
      "votes": null
    },
    {
      "id": "463096",
      "postDate": "01/29/2019 11:50:42",
      "content": "<p>you'd better try to use  </p>\n\n<p>sub_pd.sample(n=56370*7, random_state=i)</p>\n\n<p>for it is same as the private lb  size</p>",
      "rawMarkdown": "you'd better try to use  \n\nsub\\_pd.sample(n=56370*7, random_state=i)\n\nfor it is same as the private lb  size",
      "votes": null
    },
    {
      "id": "463140",
      "postDate": "01/29/2019 13:55:10",
      "content": "<p>You're making me super nervous.... God bless us everyone.\nAnd thanks for the code, I'm mentally prepared for the worst result now.</p>",
      "rawMarkdown": "You're making me super nervous.... God bless us everyone.\nAnd thanks for the code, I'm mentally prepared for the worst result now.",
      "votes": null
    },
    {
      "id": "463148",
      "postDate": "01/29/2019 14:25:03",
      "content": "<p>try your local cv</p>",
      "rawMarkdown": "try your local cv",
      "votes": null
    },
    {
      "id": "463191",
      "postDate": "01/29/2019 15:36:07",
      "content": "<p>I'm going to submit one kernel of best LB, another kernel with heavy regularization, which gives me much higher CV score but slightly lower LB score.\nI just hope the variance is not that high in the private test data.</p>",
      "rawMarkdown": "I'm going to submit one kernel of best LB, another kernel with heavy regularization, which gives me much higher CV score but slightly lower LB score.\nI just hope the variance is not that high in the private test data.",
      "votes": null
    },
    {
      "id": "463240",
      "postDate": "01/29/2019 16:57:07",
      "content": "<p>Fortunately, I am able to replicate Public LB score for 0.699 and have consistent CV as well. Trust your CV. Huge shakeup expectation from #1 as well</p>",
      "rawMarkdown": "Fortunately, I am able to replicate Public LB score for 0.699 and have consistent CV as well. Trust your CV. Huge shakeup expectation from #1 as well",
      "votes": null
    },
    {
      "id": "463246",
      "postDate": "01/29/2019 17:07:06",
      "content": "<p>Well.. treating only min/max extremes is not a good practice. Take a look at the distribution.\nI did the same experiment, but with a simpler model. Got OOF and holdout predictions, together with optimal thresholds. Then subsampled 300 times a set of length 56k, got these f1 distributions for OOF and holdout:</p>\n\n<p><img src=\"https://habrastorage.org/webt/j_/lh/qu/j_lhquhuedllexg2akgi7yk_owm.png\" alt=\"enter image description here\"></p>\n\n<p>So it's 0.668 +/- 0.004 for OOF and 0.684 +/- 0.004 for holdout if you treat medians and 25%, 75% quartiles.\nNot that awful! :)</p>",
      "rawMarkdown": "Well.. treating only min/max extremes is not a good practice. Take a look at the distribution.\nI did the same experiment, but with a simpler model. Got OOF and holdout predictions, together with optimal thresholds. Then subsampled 300 times a set of length 56k, got these f1 distributions for OOF and holdout:\n\n![enter image description here][1]\n\n\n  [1]: https://habrastorage.org/webt/j_/lh/qu/j_lhquhuedllexg2akgi7yk_owm.png\n\nSo it's 0.668 +/- 0.004 for OOF and 0.684 +/- 0.004 for holdout if you treat medians and 25%, 75% quartiles.\nNot that awful! :)",
      "votes": null
    },
    {
      "id": "463475",
      "postDate": "01/30/2019 03:40:27",
      "content": "<p>waht is the relation between  local cv and two distributions</p>",
      "rawMarkdown": "waht is the relation between  local cv and two distributions",
      "votes": null
    },
    {
      "id": "463610",
      "postDate": "01/30/2019 09:36:27",
      "content": "<p>First of all, i made a holdout set. Then got OOF predicted probs for the training part and a single vector of predicted probs for the holdout set. Then found best thresholds for OOF and holdout separately. Finally, made 300 stratified splits both for OOF and holdout (with test_size=56k) and made these plots.</p>",
      "rawMarkdown": "First of all, i made a holdout set. Then got OOF predicted probs for the training part and a single vector of predicted probs for the holdout set. Then found best thresholds for OOF and holdout separately. Finally, made 300 stratified splits both for OOF and holdout (with test_size=56k) and made these plots.",
      "votes": null
    },
    {
      "id": "463612",
      "postDate": "01/30/2019 09:38:49",
      "content": "<p>A more correct way would be to perform repeared stratified KFold with fold size = 56k. Each time training a new model and searching for the best threshold. However, this would be somewhat expensive to perform if you want to gather some reasonable statistics.</p>",
      "rawMarkdown": "A more correct way would be to perform repeared stratified KFold with fold size = 56k. Each time training a new model and searching for the best threshold. However, this would be somewhat expensive to perform if you want to gather some reasonable statistics.",
      "votes": null
    },
    {
      "id": "463998",
      "postDate": "01/31/2019 02:58:52",
      "content": "<p>My best LB model is somewhere around .68 cv :(   Looks so hopeless</p>",
      "rawMarkdown": "My best LB model is somewhere around .68 cv :(   Looks so hopeless",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 463094,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "01/29/2019 11:46:15",
      "content": "<p>Wow, the gap is soooooooooooooo large, scary. All you need is the faith in cv. LOL</p>",
      "votes": null,
      "replies": [
        {
          "id": 463096,
          "author_name": "baomengjiao",
          "author_url": "",
          "post_date": "01/29/2019 11:50:42",
          "content": "<p>you'd better try to use  </p>\n\n<p>sub_pd.sample(n=56370*7, random_state=i)</p>\n\n<p>for it is same as the private lb  size</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 463140,
      "author_name": "jerrykuo7727",
      "author_url": "",
      "post_date": "01/29/2019 13:55:10",
      "content": "<p>You're making me super nervous.... God bless us everyone.\nAnd thanks for the code, I'm mentally prepared for the worst result now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 463148,
          "author_name": "baomengjiao",
          "author_url": "",
          "post_date": "01/29/2019 14:25:03",
          "content": "<p>try your local cv</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463191,
          "author_name": "jerrykuo7727",
          "author_url": "",
          "post_date": "01/29/2019 15:36:07",
          "content": "<p>I'm going to submit one kernel of best LB, another kernel with heavy regularization, which gives me much higher CV score but slightly lower LB score.\nI just hope the variance is not that high in the private test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 463240,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "01/29/2019 16:57:07",
      "content": "<p>Fortunately, I am able to replicate Public LB score for 0.699 and have consistent CV as well. Trust your CV. Huge shakeup expectation from #1 as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463246,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "01/29/2019 17:07:06",
      "content": "<p>Well.. treating only min/max extremes is not a good practice. Take a look at the distribution.\nI did the same experiment, but with a simpler model. Got OOF and holdout predictions, together with optimal thresholds. Then subsampled 300 times a set of length 56k, got these f1 distributions for OOF and holdout:</p>\n\n<p><img src=\"https://habrastorage.org/webt/j_/lh/qu/j_lhquhuedllexg2akgi7yk_owm.png\" alt=\"enter image description here\"></p>\n\n<p>So it's 0.668 +/- 0.004 for OOF and 0.684 +/- 0.004 for holdout if you treat medians and 25%, 75% quartiles.\nNot that awful! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 463475,
          "author_name": "baomengjiao",
          "author_url": "",
          "post_date": "01/30/2019 03:40:27",
          "content": "<p>waht is the relation between  local cv and two distributions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463610,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/30/2019 09:36:27",
          "content": "<p>First of all, i made a holdout set. Then got OOF predicted probs for the training part and a single vector of predicted probs for the holdout set. Then found best thresholds for OOF and holdout separately. Finally, made 300 stratified splits both for OOF and holdout (with test_size=56k) and made these plots.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463612,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "01/30/2019 09:38:49",
          "content": "<p>A more correct way would be to perform repeared stratified KFold with fold size = 56k. Each time training a new model and searching for the best threshold. However, this would be somewhat expensive to perform if you want to gather some reasonable statistics.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 463998,
      "author_name": "icemekaveli",
      "author_url": "",
      "post_date": "01/31/2019 02:58:52",
      "content": "<p>My best LB model is somewhere around .68 cv :(   Looks so hopeless</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "463089": "try this code, pred\\_train\\_y is the oof pred and train\\_y Is lable. \n\nIt sample a same size dataset as public.\n\n    shake_up = []\n    sub_pd = pd.DataFrame({\n        'pred':pred_train_y,\n        'label':train_y\n    })\n    for i in range(50):\n        sample_t = sub_pd.sample(n=56370, random_state=i)\n        sr = metrics.f1_score(sample_t['label'], (sample_t['pred'] &gt; best_thresh).astype(int))\n        print(sr)\n        shake_up.append(sr)\n        \n    print(max(shake_up))\n    print(min(shake_up))\n\nI do this and a get :\n**it varies from 0.6889 to 0.7106**",
    "463094": "Wow, the gap is soooooooooooooo large, scary. All you need is the faith in cv. LOL",
    "463096": "you'd better try to use  \n\nsub\\_pd.sample(n=56370*7, random_state=i)\n\nfor it is same as the private lb  size",
    "463140": "You're making me super nervous.... God bless us everyone.\nAnd thanks for the code, I'm mentally prepared for the worst result now.",
    "463148": "try your local cv",
    "463191": "I'm going to submit one kernel of best LB, another kernel with heavy regularization, which gives me much higher CV score but slightly lower LB score.\nI just hope the variance is not that high in the private test data.",
    "463240": "Fortunately, I am able to replicate Public LB score for 0.699 and have consistent CV as well. Trust your CV. Huge shakeup expectation from #1 as well",
    "463246": "Well.. treating only min/max extremes is not a good practice. Take a look at the distribution.\nI did the same experiment, but with a simpler model. Got OOF and holdout predictions, together with optimal thresholds. Then subsampled 300 times a set of length 56k, got these f1 distributions for OOF and holdout:\n\n![enter image description here][1]\n\n\n  [1]: https://habrastorage.org/webt/j_/lh/qu/j_lhquhuedllexg2akgi7yk_owm.png\n\nSo it's 0.668 +/- 0.004 for OOF and 0.684 +/- 0.004 for holdout if you treat medians and 25%, 75% quartiles.\nNot that awful! :)",
    "463475": "waht is the relation between  local cv and two distributions",
    "463610": "First of all, i made a holdout set. Then got OOF predicted probs for the training part and a single vector of predicted probs for the holdout set. Then found best thresholds for OOF and holdout separately. Finally, made 300 stratified splits both for OOF and holdout (with test_size=56k) and made these plots.",
    "463612": "A more correct way would be to perform repeared stratified KFold with fold size = 56k. Each time training a new model and searching for the best threshold. However, this would be somewhat expensive to perform if you want to gather some reasonable statistics.",
    "463998": "My best LB model is somewhere around .68 cv :(   Looks so hopeless"
  },
  "source": "meta"
}