{
  "id": 75182,
  "title": "what's your  cv  f1 score and lb  f1 score?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75182",
  "author_name": "AIFIRST",
  "post_date": "2018-12-19T07:58:54.718000",
  "votes": 26,
  "comment_count": 88,
  "views": 0,
  "content": "<p>my  cv  f1  is  0.678  and  lb f1  is 0.699</p>",
  "messages": [
    {
      "id": 441889,
      "postDate": "2018-12-19T07:58:54.720Z",
      "content": "<p>my  cv  f1  is  0.678  and  lb f1  is 0.699</p>",
      "rawMarkdown": "my  cv  f1  is  0.678  and  lb f1  is 0.699",
      "votes": 26
    },
    {
      "id": 447611,
      "postDate": "2018-12-30T06:59:40.227Z",
      "content": "<p>I am sure the leaderboard will shuffle like crazy at stage 2. What do you think?</p>",
      "rawMarkdown": "I am sure the leaderboard will shuffle like crazy at stage 2. What do you think?",
      "votes": 9,
      "replies": [
        {
          "id": 448289,
          "postDate": "2018-12-31T17:53:33.320Z",
          "content": "<p>Is there a stage 2 for this competition? Looking at the timeline section it says February 5, 2019 - Final submission deadline. Also, public LB is calculated on full test data. </p>",
          "rawMarkdown": "Is there a stage 2 for this competition? Looking at the timeline section it says February 5, 2019 - Final submission deadline. Also, public LB is calculated on full test data. "
        },
        {
          "id": 448294,
          "postDate": "2018-12-31T18:09:18.287Z",
          "content": "<p>Rajesh I according to the rules we won't be able to see our results on any other leaderboard apart from this. It is not specified anywhere when stage 2 data will be loaded. I think after the final submissions they will run our kernel on the whole test dataset and that is what they are calling stage 2. </p>",
          "rawMarkdown": "Rajesh I according to the rules we won't be able to see our results on any other leaderboard apart from this. It is not specified anywhere when stage 2 data will be loaded. I think after the final submissions they will run our kernel on the whole test dataset and that is what they are calling stage 2. ",
          "votes": 1
        },
        {
          "id": 448297,
          "postDate": "2018-12-31T18:15:11.910Z",
          "content": "<p>When you click on private LB it says as below:\nThe private leaderboard is calculated over the same rows as the public leaderboard in this competition.</p>\n\n<p>Guessing there shouldn't be stage 2 in this competition.</p>",
          "rawMarkdown": "When you click on private LB it says as below:\nThe private leaderboard is calculated over the same rows as the public leaderboard in this competition.\n\nGuessing there shouldn't be stage 2 in this competition."
        },
        {
          "id": 448298,
          "postDate": "2018-12-31T18:17:30.547Z",
          "content": "<p>From data page:\nTest data: This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.</p>\n\n<p>I am guessing that after final submission they will just swap the test data and run kernels. </p>",
          "rawMarkdown": "From data page:\nTest data: This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\n\nI am guessing that after final submission they will just swap the test data and run kernels. ",
          "votes": 1
        },
        {
          "id": 448299,
          "postDate": "2018-12-31T18:20:37.233Z",
          "content": "<p>Oops didnt notice that. Thanks for bringing this up :)</p>",
          "rawMarkdown": "Oops didnt notice that. Thanks for bringing this up :)",
          "votes": 1
        },
        {
          "id": 448300,
          "postDate": "2018-12-31T18:22:30.237Z",
          "content": "<p>No problem :)</p>",
          "rawMarkdown": "No problem :)"
        },
        {
          "id": 450852,
          "postDate": "2019-01-05T23:35:08.920Z",
          "content": "<p>And if with the new test file, with more rows, your kernel takes longer than 2 hours to run, will be that OK? (in a GPU kernel)</p>",
          "rawMarkdown": "And if with the new test file, with more rows, your kernel takes longer than 2 hours to run, will be that OK? (in a GPU kernel)"
        }
      ]
    },
    {
      "id": 445034,
      "postDate": "2018-12-25T12:08:31.737Z",
      "content": "<p>local cv 0.678 public lb 0.699 ,while i reach local cv 0.68+ ,get lower public lb instead</p>",
      "rawMarkdown": "local cv 0.678 public lb 0.699 ,while i reach local cv 0.68+ ,get lower public lb instead",
      "votes": 3
    },
    {
      "id": 444750,
      "postDate": "2018-12-24T18:13:32.407Z",
      "content": "<p>It took me some time to understand PyTorch, but now I am seeing a correlation between my local validation loss to local f1 to public LB. My avg. loss is about 0.0657, my local f1 is about 0.689 and the public is 0.695. PyTorch gives me a consistent result so now I feel confident that the changes I'm making aren't due to random luck or lucky seed...  </p>",
      "rawMarkdown": "It took me some time to understand PyTorch, but now I am seeing a correlation between my local validation loss to local f1 to public LB. My avg. loss is about 0.0657, my local f1 is about 0.689 and the public is 0.695. PyTorch gives me a consistent result so now I feel confident that the changes I'm making aren't due to random luck or lucky seed...  ",
      "votes": 3,
      "replies": [
        {
          "id": 445216,
          "postDate": "2018-12-26T01:18:19.590Z",
          "content": "<p>Did you modify any of the default initializer settings?</p>",
          "rawMarkdown": "Did you modify any of the default initializer settings?"
        },
        {
          "id": 445678,
          "postDate": "2018-12-26T23:32:59.663Z",
          "content": "<p>No. Not yet. I've been just trying to level out the local f1 to the public lb. I don't want to put much trust in a local f1 of 0.675 and public lb of 0.700. There's a thread about bestfitting simulating the slide in an old competition. Perhaps something there to be learned... I also can't help but to think about the Mercedes competition.</p>",
          "rawMarkdown": "No. Not yet. I've been just trying to level out the local f1 to the public lb. I don't want to put much trust in a local f1 of 0.675 and public lb of 0.700. There's a thread about bestfitting simulating the slide in an old competition. Perhaps something there to be learned... I also can't help but to think about the Mercedes competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 443456,
      "postDate": "2018-12-21T17:20:11.557Z",
      "content": "<p>My current best is a cv of 0.687 and a public lb of 0.69. Interestingly, my previous best had a cv of 0.678 and a public lb of 0.693.  So far, any of my kernels that have higher cv scores than my previous best of 0.678 keep getting lower and lower public lb scores. I know that evaluating models on the local cv should get you a better score on the private LB, but compared to most of the people in this discussion post, my public lb score is lower than what it should be for my current cv F1 score</p>",
      "rawMarkdown": "My current best is a cv of 0.687 and a public lb of 0.69. Interestingly, my previous best had a cv of 0.678 and a public lb of 0.693.  So far, any of my kernels that have higher cv scores than my previous best of 0.678 keep getting lower and lower public lb scores. I know that evaluating models on the local cv should get you a better score on the private LB, but compared to most of the people in this discussion post, my public lb score is lower than what it should be for my current cv F1 score",
      "votes": 3
    },
    {
      "id": 442841,
      "postDate": "2018-12-20T15:42:13.507Z",
      "content": "<p>local CV 0.6749 and LB 0.699</p>\n\n<p>I noticed that every time when my local cv was bigger than 0.68 my LB score always not so good. </p>",
      "rawMarkdown": "local CV 0.6749 and LB 0.699\n\nI noticed that every time when my local cv was bigger than 0.68 my LB score always not so good. ",
      "votes": 4
    },
    {
      "id": 442581,
      "postDate": "2018-12-20T06:45:29.973Z",
      "content": "<p>My local f1 score(0.10 valid) is 0.7062, but my lb is 0.697\nCry, cry, cry... <br>\nWhy are there such big differences between online and offline? </p>",
      "rawMarkdown": "My local f1 score(0.10 valid) is 0.7062, but my lb is 0.697\nCry, cry, cry...   \nWhy are there such big differences between online and offline? ",
      "votes": 4,
      "replies": [
        {
          "id": 442634,
          "postDate": "2018-12-20T08:49:44.063Z",
          "content": "<p>your  local f1 score is really high ,   I think maybe  a higher  cv f1  score  can  get   a better score  on 2nd's test data. </p>",
          "rawMarkdown": "your  local f1 score is really high ,   I think maybe  a higher  cv f1  score  can  get   a better score  on 2nd's test data. "
        },
        {
          "id": 442635,
          "postDate": "2018-12-20T08:51:44.573Z",
          "content": "<p>my cv score （5flod)     cv  f1:0.678    -&gt;  lb:0.699 \n                                         cv  f1:0.688    -&gt; lb:0.693</p>",
          "rawMarkdown": "my cv score （5flod)     cv  f1:0.678    -&gt;  lb:0.699 \n                                         cv  f1:0.688    -&gt; lb:0.693\n\n",
          "votes": 1
        },
        {
          "id": 442753,
          "postDate": "2018-12-20T12:48:41.413Z",
          "content": "<p>cv with single model always catch lower f1 than public lb, but I didn't understand why so many people set up a validation set, I think they have misunderstandings about cross-validation.</p>",
          "rawMarkdown": "cv with single model always catch lower f1 than public lb, but I didn't understand why so many people set up a validation set, I think they have misunderstandings about cross-validation."
        },
        {
          "id": 444175,
          "postDate": "2018-12-23T12:26:31.720Z",
          "content": "<p>yes, someone  just split  train data into train and valid  , valid not be used to be trained .that's not a cv </p>",
          "rawMarkdown": "yes, someone  just split  train data into train and valid  , valid not be used to be trained .that's not a cv ",
          "votes": 2
        },
        {
          "id": 444991,
          "postDate": "2018-12-25T09:34:30.867Z",
          "content": "<p><a href=\"/bestpredict\">@bestpredict</a>\nif your cv score rise and LB descent, you can go to see your confusion matrix</p>",
          "rawMarkdown": "@bestpredict\nif your cv score rise and LB descent, you can go to see your confusion matrix"
        },
        {
          "id": 447832,
          "postDate": "2018-12-30T17:22:03.277Z",
          "content": "<p>thank you for advice , I'll try it @Bai</p>",
          "rawMarkdown": "thank you for advice , I'll try it @Bai"
        }
      ]
    },
    {
      "id": 448061,
      "postDate": "2018-12-31T06:52:13.770Z",
      "content": "<p>Can someone help me make sense of this?\nI used BiLSTM-attention-Kfold-CLR-Extra Features-capsule kernel by Spiros, which gave LB 0.696 straight out of the box.\nI then added a small tweak to modify the training data which gave CV score of 0.718 but a poor 0.567 on LB.\nDoes this mean that LB test data characteristics are different from the test data characteristics given to us?</p>",
      "rawMarkdown": "Can someone help me make sense of this?\nI used BiLSTM-attention-Kfold-CLR-Extra Features-capsule kernel by Spiros, which gave LB 0.696 straight out of the box.\nI then added a small tweak to modify the training data which gave CV score of 0.718 but a poor 0.567 on LB.\nDoes this mean that LB test data characteristics are different from the test data characteristics given to us?",
      "votes": 1,
      "replies": [
        {
          "id": 448066,
          "postDate": "2018-12-31T06:57:26.390Z",
          "content": "<p>What did you do to your training data? If you want to add some noise, I thought you should firstly split the valid data out and don't modify the valid data.</p>",
          "rawMarkdown": "What did you do to your training data? If you want to add some noise, I thought you should firstly split the valid data out and don't modify the valid data.",
          "votes": 1
        },
        {
          "id": 448068,
          "postDate": "2018-12-31T07:03:34.040Z",
          "content": "<p>That's a good point. I just made a blanket change to modify the target label for questions below a certain length from 0 to 1 since most of these short questions seemed to be insincere but were marked as sincere.  So i thought its better for them to be marked as insincere with few errors than sincere with many errors.</p>",
          "rawMarkdown": "That's a good point. I just made a blanket change to modify the target label for questions below a certain length from 0 to 1 since most of these short questions seemed to be insincere but were marked as sincere.  So i thought its better for them to be marked as insincere with few errors than sincere with many errors."
        },
        {
          "id": 448096,
          "postDate": "2018-12-31T08:35:02.203Z",
          "content": "<p>That's an interesting idea, but I thought that may destroy the distribution of train data.</p>",
          "rawMarkdown": "That's an interesting idea, but I thought that may destroy the distribution of train data."
        },
        {
          "id": 448100,
          "postDate": "2018-12-31T08:45:54.600Z",
          "content": "<p>Also once you do that you actually create a rule that if question length is less than x then always predict 1. The Neural network learns that rule and will obviously do better on train data. </p>\n\n<p>Since the test data doesn't have such a rule, it doesn't work</p>",
          "rawMarkdown": "Also once you do that you actually create a rule that if question length is less than x then always predict 1. The Neural network learns that rule and will obviously do better on train data. \n\nSince the test data doesn't have such a rule, it doesn't work"
        },
        {
          "id": 448103,
          "postDate": "2018-12-31T08:51:45.217Z",
          "content": "<p>Well, I think the data is already quite noisy. There is a lot of mis-classification.</p>",
          "rawMarkdown": "Well, I think the data is already quite noisy. There is a lot of mis-classification."
        },
        {
          "id": 448106,
          "postDate": "2018-12-31T09:01:36.547Z",
          "content": "<p>I hope the test data has similar misclassifications. </p>",
          "rawMarkdown": "I hope the test data has similar misclassifications. "
        },
        {
          "id": 448113,
          "postDate": "2018-12-31T09:15:23.063Z",
          "content": "<p>Also checked the distribution of test data vs train data. Seems pretty similarly distributed. See:</p>\n\n<p><a href=\"https://www.kaggle.com/mlwhiz/adversarial-validation-and-lb-shakeup\">https://www.kaggle.com/mlwhiz/adversarial-validation-and-lb-shakeup</a></p>",
          "rawMarkdown": "Also checked the distribution of test data vs train data. Seems pretty similarly distributed. See:\n\nhttps://www.kaggle.com/mlwhiz/adversarial-validation-and-lb-shakeup"
        }
      ]
    },
    {
      "id": 446653,
      "postDate": "2018-12-28T12:43:39.030Z",
      "content": "<p>local cv about 0.68, and lb 0.704, keras is amazing</p>",
      "rawMarkdown": "local cv about 0.68, and lb 0.704, keras is amazing",
      "votes": 1,
      "replies": [
        {
          "id": 447793,
          "postDate": "2018-12-30T15:24:08.593Z",
          "content": "<p>How do you calculate your local CV in this case. Is it:</p>\n\n<ol>\n<li>Average of F1 score in each fold</li>\n<li>You predict the fold data in the trainset while doing CV and then getting the F1 score for that.</li>\n<li>Separate CV set?</li>\n</ol>\n\n<p>My CV by the second approach is &gt;0.683 yet the LB is still bad. Don't know what is happening here.</p>",
          "rawMarkdown": "How do you calculate your local CV in this case. Is it:\n\n1. Average of F1 score in each fold\n2. You predict the fold data in the trainset while doing CV and then getting the F1 score for that.\n3. Separate CV set?\n\nMy CV by the second approach is &gt;0.683 yet the LB is still bad. Don't know what is happening here."
        },
        {
          "id": 447796,
          "postDate": "2018-12-30T15:27:07.927Z",
          "content": "<p>Concat out of fold predictions and evaluate against y train.</p>",
          "rawMarkdown": "Concat out of fold predictions and evaluate against y train.",
          "votes": 1
        },
        {
          "id": 447798,
          "postDate": "2018-12-30T15:29:44.253Z",
          "content": "<p>That is what I am doing. My best model scores around 0.684 Local 5 fold CV with that. Scores like 0.69 on LB. a stack of 5 different models is scoring 0.697 Local CV and 0.692 LB....</p>",
          "rawMarkdown": "That is what I am doing. My best model scores around 0.684 Local 5 fold CV with that. Scores like 0.69 on LB. a stack of 5 different models is scoring 0.697 Local CV and 0.692 LB....\n"
        },
        {
          "id": 448315,
          "postDate": "2018-12-31T19:12:44.277Z",
          "content": "<p>hi, I calculate local CV on separate CV set.i have the same problem, my best local CV is 0.72, but its LB is 0.68, maybe overfitting?</p>",
          "rawMarkdown": "hi, I calculate local CV on separate CV set.i have the same problem, my best local CV is 0.72, but its LB is 0.68, maybe overfitting?",
          "votes": 1
        }
      ]
    },
    {
      "id": 444676,
      "postDate": "2018-12-24T14:35:34.843Z",
      "content": "<p>cv 0.69 lb 0.68. Strange</p>",
      "rawMarkdown": "cv 0.69 lb 0.68. Strange",
      "votes": 1
    },
    {
      "id": 443002,
      "postDate": "2018-12-20T21:33:43.293Z",
      "content": "<p>cv = 0.6834\nlb = 0.687\n<strong>__ update\ncv = 0.6861\nlb = 0.691\n_<em></em></strong><em> update\ncv = 0.69068\nlb = 0.697 \n_</em>__ update\ncv = 0.696892\nlb = 0.702</p>",
      "rawMarkdown": "cv = 0.6834\nlb = 0.687\n____ update\ncv = 0.6861\nlb = 0.691\n____ update\ncv = 0.69068\nlb = 0.697 \n____ update\ncv = 0.696892\nlb = 0.702",
      "votes": 1,
      "replies": [
        {
          "id": 444173,
          "postDate": "2018-12-23T12:22:56.993Z",
          "content": "<p>whant's you cv fold?</p>",
          "rawMarkdown": "whant's you cv fold?",
          "votes": 1
        },
        {
          "id": 444218,
          "postDate": "2018-12-23T15:09:22.167Z",
          "content": "<p>4 folds:\nFold: 1 Val F1 Score: 0.689334254780005 best thresh: 0.29\nFold: 2 Val F1 Score: 0.687243834863646 best thresh: 0.32\nFold: 3 Val F1 Score: 0.683858643744031 best thresh: 0.3\nFold: 4 Val F1 Score: 0.684088484904698 best thresh: 0.34</p>",
          "rawMarkdown": "4 folds:\nFold: 1 Val F1 Score: 0.689334254780005 best thresh: 0.29\nFold: 2 Val F1 Score: 0.687243834863646 best thresh: 0.32\nFold: 3 Val F1 Score: 0.683858643744031 best thresh: 0.3\nFold: 4 Val F1 Score: 0.684088484904698 best thresh: 0.34",
          "votes": 1
        },
        {
          "id": 447251,
          "postDate": "2018-12-29T13:28:55.943Z",
          "content": "<p>You are finding the threshold for each fold and then did majority voting for test data.</p>\n\n<p>Am I right ?</p>",
          "rawMarkdown": "You are finding the threshold for each fold and then did majority voting for test data.\n\nAm I right ?",
          "votes": 1
        },
        {
          "id": 451102,
          "postDate": "2019-01-06T12:30:27.397Z",
          "content": "<p>I tried a lot of stuff... i did some quick simulation on how to find the best threshold (<strong>majority voting</strong>, <strong>mean over cv-thresholds</strong>, <strong>geometric mean over cv-thresholds</strong>, <strong>static thresh of 0.33</strong>, <strong>calculate thresh via cv predictions</strong>, ...). Sadly it turns out that it is pretty unstable ... </p>",
          "rawMarkdown": "I tried a lot of stuff... i did some quick simulation on how to find the best threshold (**majority voting**, **mean over cv-thresholds**, **geometric mean over cv-thresholds**, **static thresh of 0.33**, **calculate thresh via cv predictions**, ...). Sadly it turns out that it is pretty unstable ... "
        }
      ]
    },
    {
      "id": 446060,
      "postDate": "2018-12-27T12:41:32.600Z",
      "content": "<p>CV : 0.6806\nLB: 0.698</p>\n\n<p>Although I saw some negative correlation also between CV and LB, my highest score corresponds to my highest CV yet. Will update if I see anything different. I am using Spiros approach to calculate CV. Get out of fold predictions for training set and find F1 using that with Stratified k-folds</p>\n\n<p>Also attaching chart for LB vs CV</p>",
      "rawMarkdown": "CV : 0.6806\nLB: 0.698\n\nAlthough I saw some negative correlation also between CV and LB, my highest score corresponds to my highest CV yet. Will update if I see anything different. I am using Spiros approach to calculate CV. Get out of fold predictions for training set and find F1 using that with Stratified k-folds\n\nAlso attaching chart for LB vs CV",
      "votes": 2
    },
    {
      "id": 442212,
      "postDate": "2018-12-19T16:10:46.697Z",
      "content": "<p>My current gut feeling: trust local cv score rather than public lb score.</p>",
      "rawMarkdown": "My current gut feeling: trust local cv score rather than public lb score.",
      "votes": 2,
      "replies": [
        {
          "id": 442265,
          "postDate": "2018-12-19T17:54:03.487Z",
          "content": "<p>Should you trust the val F1 score more than the local val loss? My public lb score seems more correlated with increases in local val F1 than with decreases in the local val loss</p>",
          "rawMarkdown": "Should you trust the val F1 score more than the local val loss? My public lb score seems more correlated with increases in local val F1 than with decreases in the local val loss"
        },
        {
          "id": 442320,
          "postDate": "2018-12-19T19:43:19.630Z",
          "content": "<p>Hard to say, I would say F1, but I am also checking auroc.</p>",
          "rawMarkdown": "Hard to say, I would say F1, but I am also checking auroc.",
          "votes": 1
        }
      ]
    },
    {
      "id": 447181,
      "postDate": "2018-12-29T10:10:00.597Z",
      "content": "<p>Newbie here but what is the difference between your CV and LB scores?</p>",
      "rawMarkdown": "Newbie here but what is the difference between your CV and LB scores?",
      "replies": [
        {
          "id": 447822,
          "postDate": "2018-12-30T16:44:16.920Z",
          "content": "<p>CV(Cross validation) score is the score which you get after doing a k-fold validation. </p>\n\n<p>What I mean is that if you run a 5-fold validation scheme. You train on 1,2,3,4 folds from the training data and predict the 5th fold from the training data. Once you are through with the cross validation loop, you have OOF(Out of fold) predictions for the whole training data. It is using these preds and the actual labels for training data that you get the local CV score. </p>\n\n<p>LB score is just the leaderboard score that is calculated using the predictions you provide on the test set. </p>\n\n<p>For example in the below kernel, I calculate it in this line: <a href=\"https://www.kaggle.com/mlwhiz/initializing-pytorch-layers-weight-with-kaiming\">https://www.kaggle.com/mlwhiz/initializing-pytorch-layers-weight-with-kaiming</a></p>\n\n<p>search_result = threshold_search(y_train, train_preds)</p>\n\n<p>Let me know if that clarifies your question.</p>",
          "rawMarkdown": "CV(Cross validation) score is the score which you get after doing a k-fold validation. \n\nWhat I mean is that if you run a 5-fold validation scheme. You train on 1,2,3,4 folds from the training data and predict the 5th fold from the training data. Once you are through with the cross validation loop, you have OOF(Out of fold) predictions for the whole training data. It is using these preds and the actual labels for training data that you get the local CV score. \n\nLB score is just the leaderboard score that is calculated using the predictions you provide on the test set. \n\nFor example in the below kernel, I calculate it in this line: https://www.kaggle.com/mlwhiz/initializing-pytorch-layers-weight-with-kaiming\n\nsearch_result = threshold_search(y_train, train_preds)\n\nLet me know if that clarifies your question.",
          "votes": 2
        }
      ]
    },
    {
      "id": 457467,
      "postDate": "2019-01-17T13:48:26.893Z",
      "content": "<p>Wow, you are all so lucky with 0.7 LB and 0.68 CV. For me, it's:\nCV: 0.680\nLB: 0.690</p>",
      "rawMarkdown": "Wow, you are all so lucky with 0.7 LB and 0.68 CV. For me, it's:\nCV: 0.680\nLB: 0.690"
    },
    {
      "id": 447448,
      "postDate": "2018-12-29T20:37:52.217Z",
      "content": "<p>Single model, \nValidation score: 0.699\nLB: 0.690</p>",
      "rawMarkdown": "Single model, \nValidation score: 0.699\nLB: 0.690"
    },
    {
      "id": 446501,
      "postDate": "2018-12-28T07:36:55.017Z",
      "content": "<p>Single model, 5 folders\ncv:0.687\nlb:0.691</p>",
      "rawMarkdown": "Single model, 5 folders\ncv:0.687\nlb:0.691",
      "replies": [
        {
          "id": 446518,
          "postDate": "2018-12-28T08:27:53.490Z",
          "content": "<p>best cv:0.690     (5 folders)   lb:0.694</p>",
          "rawMarkdown": "best cv:0.690     (5 folders)   lb:0.694"
        },
        {
          "id": 446564,
          "postDate": "2018-12-28T09:40:00.823Z",
          "content": "<p>Nice job !</p>",
          "rawMarkdown": "Nice job !"
        }
      ]
    },
    {
      "id": 442842,
      "postDate": "2018-12-20T15:42:42.897Z",
      "content": "<p>my cv f1 is 0.677 and lb f1 is 0.698. I improve my local f1, but get lower LB score = =</p>",
      "rawMarkdown": "my cv f1 is 0.677 and lb f1 is 0.698. I improve my local f1, but get lower LB score = =",
      "replies": [
        {
          "id": 442892,
          "postDate": "2018-12-20T16:47:50.453Z",
          "content": "<p>I think I notice a negative correlation in these levels. My best is cv f1 0.6781 and LB 0.700 . In another version local cv f1 0.6809 and LB 0.697.</p>",
          "rawMarkdown": "I think I notice a negative correlation in these levels. My best is cv f1 0.6781 and LB 0.700 . In another version local cv f1 0.6809 and LB 0.697.",
          "votes": 1
        },
        {
          "id": 443639,
          "postDate": "2018-12-22T02:08:59.307Z",
          "content": "<p><em><strong>EDIT</strong></em>: Ok, just found out <a href=\"/spirosrap\">@spirosrap</a> excellent kernel : <a href=\"https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule\">https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule</a> . That is quite clear up my question below. </p>\n\n<p><em><strong>EDIT2</strong></em> So, I think the negative correlation between local CV and LB is <em>misleading</em>. In local CV, you based your score on a single classifier (trained on K-1 folds), but in LB you based your prediction on K classifiers. Therefore, they are incomparable. If you separate one hold out set (e.g. 10% of the dataset) before doing KFolds, in my experience , local CV on that hold out set will positively correlate with LB,</p>\n\n<p>---- Original Question ----\n<a href=\"/spirosrap\">@spirosrap</a> I'm curious when you says local CV F1, how did you measure that?\nI saw in other disscusions you mentioned that you use K-Fold with no holdout set.\nTherefore, when you said you best local CV is 0.6781 did you mean this:</p>\n\n<p>First <strong>at training time</strong>,  divide ALL data into (Strastified) K Folds <strong>without \"hold out\" set</strong>\nfor each fold J in range(K),\n    Train on K-1 Folds (all folds except J), measure F1 on the remaining fold J.</p>\n\n<p>Local CV 0.6781 (mentioned above) = average of this K results.</p>\n\n<p><strong>At submission</strong>, average the prediction of these K classifiers</p>\n\n<p>Am I correct, or am I completely wrong?\nI emphasize <strong>\"without hold out\"</strong> since if we pre-allocate a hold out set to measure F1 and to select threshold, it could be a totally different situation.</p>\n\n<p><a href=\"/bestpredict\">@bestpredict</a> <a href=\"/mldevl\">@mldevl</a> May I ask you also the same question?</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "***EDIT***: Ok, just found out @spirosrap excellent kernel : https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule . That is quite clear up my question below. \n\n***EDIT2*** So, I think the negative correlation between local CV and LB is *misleading*. In local CV, you based your score on a single classifier (trained on K-1 folds), but in LB you based your prediction on K classifiers. Therefore, they are incomparable. If you separate one hold out set (e.g. 10% of the dataset) before doing KFolds, in my experience , local CV on that hold out set will positively correlate with LB,\n\n\n---- Original Question ----\n@spirosrap I'm curious when you says local CV F1, how did you measure that?\nI saw in other disscusions you mentioned that you use K-Fold with no holdout set.\nTherefore, when you said you best local CV is 0.6781 did you mean this:\n\nFirst **at training time**,  divide ALL data into (Strastified) K Folds **without \"hold out\" set**\nfor each fold J in range(K),\n    Train on K-1 Folds (all folds except J), measure F1 on the remaining fold J.\n\n\nLocal CV 0.6781 (mentioned above) = average of this K results.\n\n**At submission**, average the prediction of these K classifiers\n\nAm I correct, or am I completely wrong?\nI emphasize **\"without hold out\"** since if we pre-allocate a hold out set to measure F1 and to select threshold, it could be a totally different situation.\n\n@bestpredict @mldevl May I ask you also the same question?\n\nThanks!"
        },
        {
          "id": 443663,
          "postDate": "2018-12-22T04:27:04.217Z",
          "content": "<p>yes,you are right. k-1 folds as training set, remaining fold as validation set. After k fold, you can get the full prediction of the training set. Then you can get a best threshold based on it.</p>",
          "rawMarkdown": "yes,you are right. k-1 folds as training set, remaining fold as validation set. After k fold, you can get the full prediction of the training set. Then you can get a best threshold based on it.",
          "votes": 2
        },
        {
          "id": 443676,
          "postDate": "2018-12-22T05:50:23.943Z",
          "content": "<p><a href=\"/mldevl\">@mldevl</a> thanks!  However, I am not sure that I understand correctly, so I already edited my question to be more precise. Could you please take a look again ?</p>\n\n<p><em><strong>EDIT</strong></em>: Ok, just found out <a href=\"/spirosrap\">@spirosrap</a> excellent kernel : <a href=\"https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule\">https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule</a> . That is quite clear up my question.</p>",
          "rawMarkdown": "@mldevl thanks!  However, I am not sure that I understand correctly, so I already edited my question to be more precise. Could you please take a look again ?\n\n***EDIT***: Ok, just found out @spirosrap excellent kernel : https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule . That is quite clear up my question."
        },
        {
          "id": 443747,
          "postDate": "2018-12-22T10:26:27.650Z",
          "content": "<p>I tried to keep a 10% holdout set but the LB score decreases significantly. So, I'm not sure if this is the right track. That small difference in the training data has a significant effect in the LB score. I think it was around 0.005 in one case.</p>",
          "rawMarkdown": "I tried to keep a 10% holdout set but the LB score decreases significantly. So, I'm not sure if this is the right track. That small difference in the training data has a significant effect in the LB score. I think it was around 0.005 in one case.",
          "votes": 2
        },
        {
          "id": 443790,
          "postDate": "2018-12-22T12:48:54.053Z",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> I think the dilemma of “good performance with full data” vs. “more accurate measurement with 10% holdout” can be solved by doing two experiments with the same random seed. </p>\n\n<p>In the first experiment, we use a 10% holdout just to estimate ensemble performance in general.\nIn the second experiment, we can use full data with the same random seed just to get the good LB performance.</p>",
          "rawMarkdown": "@spirosrap I think the dilemma of “good performance with full data” vs. “more accurate measurement with 10% holdout” can be solved by doing two experiments with the same random seed. \n\nIn the first experiment, we use a 10% holdout just to estimate ensemble performance in general.\nIn the second experiment, we can use full data with the same random seed just to get the good LB performance.",
          "votes": 2
        },
        {
          "id": 443885,
          "postDate": "2018-12-22T16:55:28.607Z",
          "content": "<p>Thank you, I'll try that.</p>",
          "rawMarkdown": "Thank you, I'll try that.",
          "votes": 1
        },
        {
          "id": 445680,
          "postDate": "2018-12-26T23:36:54.970Z",
          "content": "<p>Interesting... So this is where the hold out 10% before k-folds idea stems from? It makes sense that the the f1 folds would be misleading in comparison to the final f1 check, but wouldn't the averaging of validation predictions be okay anyway to do the final check?</p>",
          "rawMarkdown": "Interesting... So this is where the hold out 10% before k-folds idea stems from? It makes sense that the the f1 folds would be misleading in comparison to the final f1 check, but wouldn't the averaging of validation predictions be okay anyway to do the final check?",
          "votes": 1
        },
        {
          "id": 447235,
          "postDate": "2018-12-29T12:30:30.813Z",
          "content": "<p>@spiro I have been using your awesome kernel and using the F1 score that we get by using train_preds as the CV F1 score. So my question is how much have you been able to improve that CV F1 score in your current implementation.</p>\n\n<p>I have not been able to improve it more than .6808 which fetched .696 while when it was at .6807, it was my current LB best at .698. </p>",
          "rawMarkdown": "@spiro I have been using your awesome kernel and using the F1 score that we get by using train_preds as the CV F1 score. So my question is how much have you been able to improve that CV F1 score in your current implementation.\n\nI have not been able to improve it more than .6808 which fetched .696 while when it was at .6807, it was my current LB best at .698. \n\n"
        },
        {
          "id": 447458,
          "postDate": "2018-12-29T21:17:29.863Z",
          "content": "<p><a href=\"/learnmower\">@learnmower</a> Regarding to your last question. When you do the averaging of validation prediction, you do an “average of each single classifier”. But when you predict the test data, you use a “combination (ensemble) of all K classifiers”. I.e. they are two different models actually.</p>\n\n<p>If each single classifier is good, will you be sure that your ensemble is also great?</p>\n\n<p>Unfortuantely, not. Imagine that all of your K classifers are almost the same [i.e. they are all good but predict the same sigmoid probability to almost test data], here the ‘average prediction’ will not improve the results of the single classifier. Why? because averaging the same prob produce nothing new i.e. [0.7+0.7+0.7+0.7 +0.7]/5 = 0.7. ; so we get nothing if our classifiers are not diversified. </p>\n\n<p>Beside good individual classifier, diversification of base classifiers is another key to ensemble performance.</p>",
          "rawMarkdown": "@learnmower Regarding to your last question. When you do the averaging of validation prediction, you do an “average of each single classifier”. But when you predict the test data, you use a “combination (ensemble) of all K classifiers”. I.e. they are two different models actually.\n\nIf each single classifier is good, will you be sure that your ensemble is also great?\n\nUnfortuantely, not. Imagine that all of your K classifers are almost the same [i.e. they are all good but predict the same sigmoid probability to almost test data], here the ‘average prediction’ will not improve the results of the single classifier. Why? because averaging the same prob produce nothing new i.e. [0.7+0.7+0.7+0.7 +0.7]/5 = 0.7. ; so we get nothing if our classifiers are not diversified. \n\nBeside good individual classifier, diversification of base classifiers is another key to ensemble performance."
        },
        {
          "id": 447475,
          "postDate": "2018-12-29T22:27:22.503Z",
          "content": "<p>How are they two different models? When you run stratified k-fold, you are partitioning your data so that you run these folds into the same model K times. </p>\n\n<p><a href=\"https://scikit-learn.org/0.16/modules/generated/sklearn.cross_validation.StratifiedKFold.html\">https://scikit-learn.org/0.16/modules/generated/sklearn.cross_validation.StratifiedKFold.html</a></p>\n\n<p>And my question pertains to people who are replying to the single model thread who are holding out 10% of the training data before running their single model into a k-fold CV. This is something that I have not seen until this competition.</p>",
          "rawMarkdown": "How are they two different models? When you run stratified k-fold, you are partitioning your data so that you run these folds into the same model K times. \n\nhttps://scikit-learn.org/0.16/modules/generated/sklearn.cross_validation.StratifiedKFold.html\n\nAnd my question pertains to people who are replying to the single model thread who are holding out 10% of the training data before running their single model into a k-fold CV. This is something that I have not seen until this competition."
        },
        {
          "id": 454052,
          "postDate": "2019-01-11T04:40:57.020Z",
          "content": "<p><a href=\"/learnmower\">@learnmower</a> Notice that I use the word ‘base classifier’ not ‘model’. Now Benjamin <a href=\"/bminixhofer\">@bminixhofer</a> has written a superb kernel of the exact same principle I have discussed here.</p>\n\n<p><a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a></p>",
          "rawMarkdown": "@learnmower Notice that I use the word ‘base classifier’ not ‘model’. Now Benjamin @bminixhofer has written a superb kernel of the exact same principle I have discussed here.\n\nhttps://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed"
        }
      ]
    },
    {
      "id": 442350,
      "postDate": "2018-12-19T21:09:17.990Z",
      "content": "<p>Anyone else having .665 on CV, and .57 on submit? I must have a bug in my code.. :-/</p>",
      "rawMarkdown": "Anyone else having .665 on CV, and .57 on submit? I must have a bug in my code.. :-/",
      "replies": [
        {
          "id": 442414,
          "postDate": "2018-12-19T23:45:35.497Z",
          "content": "<p>If you're using CuDNNLSTMs on Keras a big issue is the nondeterminism of both it and Keras</p>",
          "rawMarkdown": "If you're using CuDNNLSTMs on Keras a big issue is the nondeterminism of both it and Keras"
        },
        {
          "id": 447499,
          "postDate": "2018-12-29T23:47:43.017Z",
          "content": "<p>Not exactly but 0.69 cv and 0.68 lb. donno how others are other way round.</p>",
          "rawMarkdown": "Not exactly but 0.69 cv and 0.68 lb. donno how others are other way round."
        }
      ]
    },
    {
      "id": 442039,
      "postDate": "2018-12-19T11:57:34.843Z",
      "content": "<p>I find it hard to find a logical correlation between the two. Lower local score may give higher LB score.</p>",
      "rawMarkdown": "I find it hard to find a logical correlation between the two. Lower local score may give higher LB score.",
      "replies": [
        {
          "id": 442069,
          "postDate": "2018-12-19T12:52:14.700Z",
          "content": "<p>I think the reason may be that the test set data is not enough. and maybe  higher cv  f1 score  will get better score in  2nd's  test data</p>",
          "rawMarkdown": "I think the reason may be that the test set data is not enough. and maybe  higher cv  f1 score  will get better score in  2nd's  test data",
          "votes": 3
        }
      ]
    },
    {
      "id": 442008,
      "postDate": "2018-12-19T10:54:46.463Z",
      "content": "<p>During last week, I have raised my cv f1 score (4fold)  from ~0.680 to ~0.695, however, lb f1 score(0.04valid) remains unchanged at 0.690...o(╥﹏╥)o </p>",
      "rawMarkdown": "During last week, I have raised my cv f1 score (4fold)  from ~0.680 to ~0.695, however, lb f1 score(0.04valid) remains unchanged at 0.690...o(╥﹏╥)o ",
      "replies": [
        {
          "id": 442068,
          "postDate": "2018-12-19T12:52:03.870Z",
          "content": "<p>I think the reason may be that the test set data is not enough. and maybe  higher cv  f1 score  will get better score in  2nd's  test data</p>",
          "rawMarkdown": "I think the reason may be that the test set data is not enough. and maybe  higher cv  f1 score  will get better score in  2nd's  test data",
          "votes": 1
        },
        {
          "id": 442102,
          "postDate": "2018-12-19T13:44:54.387Z",
          "content": "<p>Sorry, there I have a question. You are doing Cross Validation, why are you set a valid dataset of 0.04? Maybe you are doing a wrong CV. I think you are doing a blending operation, but you have run the same model 4 times.</p>",
          "rawMarkdown": "Sorry, there I have a question. You are doing Cross Validation, why are you set a valid dataset of 0.04? Maybe you are doing a wrong CV. I think you are doing a blending operation, but you have run the same model 4 times."
        },
        {
          "id": 442118,
          "postDate": "2018-12-19T14:00:19.130Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 442119,
          "postDate": "2018-12-19T14:00:55.193Z",
          "content": "<p>would you like to share your local f1 score ? your lb score is highly.</p>",
          "rawMarkdown": "would you like to share your local f1 score ? your lb score is highly.\n\n"
        },
        {
          "id": 442134,
          "postDate": "2018-12-19T14:15:32.777Z",
          "content": "<p>I only do 4 fold cv in my PC for tuning the model. In kernel, I split 0.04 of the training data out as a valid set only to calculate f1 score for obtaining a reasonable threshold, that is, I use 0.96 of the training data as train set to get one classifier, and only use this one single classifier to get the prediction result...</p>",
          "rawMarkdown": "I only do 4 fold cv in my PC for tuning the model. In kernel, I split 0.04 of the training data out as a valid set only to calculate f1 score for obtaining a reasonable threshold, that is, I use 0.96 of the training data as train set to get one classifier, and only use this one single classifier to get the prediction result..."
        },
        {
          "id": 442155,
          "postDate": "2018-12-19T14:47:00.770Z",
          "content": "<p>I think 4% Hold-out set is too small, and it will fluctuate a lot (should use at least &gt;10% hold out)</p>",
          "rawMarkdown": " I think 4% Hold-out set is too small, and it will fluctuate a lot (should use at least &gt;10% hold out)\n",
          "votes": 1
        },
        {
          "id": 442218,
          "postDate": "2018-12-19T16:21:36.027Z",
          "content": "<p>Have you tried Stratified K Fold? @gkd</p>",
          "rawMarkdown": "Have you tried Stratified K Fold? @gkd"
        },
        {
          "id": 442240,
          "postDate": "2018-12-19T16:54:28.417Z",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> Yes, indeed I only used Stratified K Fold :) </p>",
          "rawMarkdown": "@spirosrap Yes, indeed I only used Stratified K Fold :) \n"
        },
        {
          "id": 442248,
          "postDate": "2018-12-19T17:12:49.223Z",
          "content": "<p>@gkd Ok I though you were talking about the validation set. I noticed significant improvements when I used the whole training set without keeping a final test set for calculating the threshold. </p>",
          "rawMarkdown": "@gkd Ok I though you were talking about the validation set. I noticed significant improvements when I used the whole training set without keeping a final test set for calculating the threshold. "
        },
        {
          "id": 442982,
          "postDate": "2018-12-20T20:30:35.127Z",
          "content": "<p>@Spiros How do you decide the threshold then? Set it to 0.5?</p>",
          "rawMarkdown": "@Spiros How do you decide the threshold then? Set it to 0.5?"
        },
        {
          "id": 444041,
          "postDate": "2018-12-23T02:01:18.503Z",
          "content": "<p>Well，if we want to ensemble the result with one single model，I think keep difference between dataset is important. So we usually perform k-fold to ensemble the result. If you split only 0.04 of the training data out，the models between 4 folds are amost the same because the train data in different folds are amost the same.</p>",
          "rawMarkdown": "Well，if we want to ensemble the result with one single model，I think keep difference between dataset is important. So we usually perform k-fold to ensemble the result. If you split only 0.04 of the training data out，the models between 4 folds are amost the same because the train data in different folds are amost the same."
        },
        {
          "id": 444195,
          "postDate": "2018-12-23T13:43:02.460Z",
          "content": "<p>@JihangZhang No I don't set it to 0.5. Depending on the final predictions, I search for the threshold with the optimal f1 score. Most of the kernels here do something similar.</p>",
          "rawMarkdown": "@JihangZhang No I don't set it to 0.5. Depending on the final predictions, I search for the threshold with the optimal f1 score. Most of the kernels here do something similar."
        },
        {
          "id": 444360,
          "postDate": "2018-12-23T21:23:59.247Z",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> Sorry how do you do this when you have the entire train set? Are you still splitting some proportion of data later as validation for the threshold?</p>",
          "rawMarkdown": "@spirosrap Sorry how do you do this when you have the entire train set? Are you still splitting some proportion of data later as validation for the threshold?"
        },
        {
          "id": 444379,
          "postDate": "2018-12-23T22:56:35.370Z",
          "content": "<p>@Learnmower In most of my experiments I have used the entire train set to calculate the threshold. I have made some experiments that I have a 0.1 hold out set to for the threshold but these didn't get me a good LB score (far lower).  It's a gamble as I see it.</p>",
          "rawMarkdown": "@Learnmower In most of my experiments I have used the entire train set to calculate the threshold. I have made some experiments that I have a 0.1 hold out set to for the threshold but these didn't get me a good LB score (far lower).  It's a gamble as I see it.",
          "votes": 1
        },
        {
          "id": 444399,
          "postDate": "2018-12-24T00:35:18.110Z",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> Ah... Hm, do you see a better correlation when using the public LB as feedback? It might be worth the gamble at this point.</p>",
          "rawMarkdown": "@spirosrap Ah... Hm, do you see a better correlation when using the public LB as feedback? It might be worth the gamble at this point."
        },
        {
          "id": 444414,
          "postDate": "2018-12-24T02:07:25.610Z",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> are you using a model architecture similar to the kernels?</p>",
          "rawMarkdown": "@spirosrap are you using a model architecture similar to the kernels?"
        },
        {
          "id": 444582,
          "postDate": "2018-12-24T10:26:44.200Z",
          "content": "<p>@StevenNguyen Yes, Bidirectional LSTM,GRU and attention.</p>",
          "rawMarkdown": "@StevenNguyen Yes, Bidirectional LSTM,GRU and attention."
        }
      ]
    },
    {
      "id": 442464,
      "postDate": "2018-12-20T01:53:43.490Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 447611,
      "author_name": "Rahul Agarwal",
      "author_url": "",
      "post_date": "2018-12-30T06:59:40.227000",
      "content": "<p>I am sure the leaderboard will shuffle like crazy at stage 2. What do you think?</p>",
      "votes": 9,
      "replies": [
        {
          "id": 448289,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2018-12-31T17:53:33.320000",
          "content": "<p>Is there a stage 2 for this competition? Looking at the timeline section it says February 5, 2019 - Final submission deadline. Also, public LB is calculated on full test data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448294,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-31T18:09:18.287000",
          "content": "<p>Rajesh I according to the rules we won't be able to see our results on any other leaderboard apart from this. It is not specified anywhere when stage 2 data will be loaded. I think after the final submissions they will run our kernel on the whole test dataset and that is what they are calling stage 2. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 448297,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2018-12-31T18:15:11.910000",
          "content": "<p>When you click on private LB it says as below:\nThe private leaderboard is calculated over the same rows as the public leaderboard in this competition.</p>\n\n<p>Guessing there shouldn't be stage 2 in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448298,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-31T18:17:30.547000",
          "content": "<p>From data page:\nTest data: This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.</p>\n\n<p>I am guessing that after final submission they will just swap the test data and run kernels. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 448299,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2018-12-31T18:20:37.233000",
          "content": "<p>Oops didnt notice that. Thanks for bringing this up :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 448300,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-31T18:22:30.237000",
          "content": "<p>No problem :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450852,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-01-05T23:35:08.920000",
          "content": "<p>And if with the new test file, with more rows, your kernel takes longer than 2 hours to run, will be that OK? (in a GPU kernel)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 445034,
      "author_name": " Aku dadi hamster",
      "author_url": "",
      "post_date": "2018-12-25T12:08:31.737000",
      "content": "<p>local cv 0.678 public lb 0.699 ,while i reach local cv 0.68+ ,get lower public lb instead</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 444750,
      "author_name": "Thomas Yokota",
      "author_url": "",
      "post_date": "2018-12-24T18:13:32.407000",
      "content": "<p>It took me some time to understand PyTorch, but now I am seeing a correlation between my local validation loss to local f1 to public LB. My avg. loss is about 0.0657, my local f1 is about 0.689 and the public is 0.695. PyTorch gives me a consistent result so now I feel confident that the changes I'm making aren't due to random luck or lucky seed...  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 445216,
          "author_name": "Alan Khoa Nguyen",
          "author_url": "",
          "post_date": "2018-12-26T01:18:19.590000",
          "content": "<p>Did you modify any of the default initializer settings?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445678,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-26T23:32:59.663000",
          "content": "<p>No. Not yet. I've been just trying to level out the local f1 to the public lb. I don't want to put much trust in a local f1 of 0.675 and public lb of 0.700. There's a thread about bestfitting simulating the slide in an old competition. Perhaps something there to be learned... I also can't help but to think about the Mercedes competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 443456,
      "author_name": "bilal2vec",
      "author_url": "",
      "post_date": "2018-12-21T17:20:11.557000",
      "content": "<p>My current best is a cv of 0.687 and a public lb of 0.69. Interestingly, my previous best had a cv of 0.678 and a public lb of 0.693.  So far, any of my kernels that have higher cv scores than my previous best of 0.678 keep getting lower and lower public lb scores. I know that evaluating models on the local cv should get you a better score on the private LB, but compared to most of the people in this discussion post, my public lb score is lower than what it should be for my current cv F1 score</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 442841,
      "author_name": "heng",
      "author_url": "",
      "post_date": "2018-12-20T15:42:13.507000",
      "content": "<p>local CV 0.6749 and LB 0.699</p>\n\n<p>I noticed that every time when my local cv was bigger than 0.68 my LB score always not so good. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 442581,
      "author_name": "xy",
      "author_url": "",
      "post_date": "2018-12-20T06:45:29.973000",
      "content": "<p>My local f1 score(0.10 valid) is 0.7062, but my lb is 0.697\nCry, cry, cry... <br>\nWhy are there such big differences between online and offline? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 442634,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-20T08:49:44.063000",
          "content": "<p>your  local f1 score is really high ,   I think maybe  a higher  cv f1  score  can  get   a better score  on 2nd's test data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442635,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-20T08:51:44.573000",
          "content": "<p>my cv score （5flod)     cv  f1:0.678    -&gt;  lb:0.699 \n                                         cv  f1:0.688    -&gt; lb:0.693</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 442753,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2018-12-20T12:48:41.413000",
          "content": "<p>cv with single model always catch lower f1 than public lb, but I didn't understand why so many people set up a validation set, I think they have misunderstandings about cross-validation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444175,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-23T12:26:31.720000",
          "content": "<p>yes, someone  just split  train data into train and valid  , valid not be used to be trained .that's not a cv </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 444991,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2018-12-25T09:34:30.867000",
          "content": "<p><a href=\"/bestpredict\">@bestpredict</a>\nif your cv score rise and LB descent, you can go to see your confusion matrix</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447832,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-30T17:22:03.277000",
          "content": "<p>thank you for advice , I'll try it @Bai</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 448061,
      "author_name": "Nikhil Utane",
      "author_url": "",
      "post_date": "2018-12-31T06:52:13.770000",
      "content": "<p>Can someone help me make sense of this?\nI used BiLSTM-attention-Kfold-CLR-Extra Features-capsule kernel by Spiros, which gave LB 0.696 straight out of the box.\nI then added a small tweak to modify the training data which gave CV score of 0.718 but a poor 0.567 on LB.\nDoes this mean that LB test data characteristics are different from the test data characteristics given to us?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 448066,
          "author_name": "SEU_Zesen_Chen",
          "author_url": "",
          "post_date": "2018-12-31T06:57:26.390000",
          "content": "<p>What did you do to your training data? If you want to add some noise, I thought you should firstly split the valid data out and don't modify the valid data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 448068,
          "author_name": "Nikhil Utane",
          "author_url": "",
          "post_date": "2018-12-31T07:03:34.040000",
          "content": "<p>That's a good point. I just made a blanket change to modify the target label for questions below a certain length from 0 to 1 since most of these short questions seemed to be insincere but were marked as sincere.  So i thought its better for them to be marked as insincere with few errors than sincere with many errors.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448096,
          "author_name": "SEU_Zesen_Chen",
          "author_url": "",
          "post_date": "2018-12-31T08:35:02.203000",
          "content": "<p>That's an interesting idea, but I thought that may destroy the distribution of train data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448100,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-31T08:45:54.600000",
          "content": "<p>Also once you do that you actually create a rule that if question length is less than x then always predict 1. The Neural network learns that rule and will obviously do better on train data. </p>\n\n<p>Since the test data doesn't have such a rule, it doesn't work</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448103,
          "author_name": "Nikhil Utane",
          "author_url": "",
          "post_date": "2018-12-31T08:51:45.217000",
          "content": "<p>Well, I think the data is already quite noisy. There is a lot of mis-classification.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448106,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-31T09:01:36.547000",
          "content": "<p>I hope the test data has similar misclassifications. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448113,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-31T09:15:23.063000",
          "content": "<p>Also checked the distribution of test data vs train data. Seems pretty similarly distributed. See:</p>\n\n<p><a href=\"https://www.kaggle.com/mlwhiz/adversarial-validation-and-lb-shakeup\">https://www.kaggle.com/mlwhiz/adversarial-validation-and-lb-shakeup</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 446653,
      "author_name": "GuanQun Wu",
      "author_url": "",
      "post_date": "2018-12-28T12:43:39.030000",
      "content": "<p>local cv about 0.68, and lb 0.704, keras is amazing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 447793,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-30T15:24:08.593000",
          "content": "<p>How do you calculate your local CV in this case. Is it:</p>\n\n<ol>\n<li>Average of F1 score in each fold</li>\n<li>You predict the fold data in the trainset while doing CV and then getting the F1 score for that.</li>\n<li>Separate CV set?</li>\n</ol>\n\n<p>My CV by the second approach is &gt;0.683 yet the LB is still bad. Don't know what is happening here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447796,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-30T15:27:07.927000",
          "content": "<p>Concat out of fold predictions and evaluate against y train.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447798,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-30T15:29:44.253000",
          "content": "<p>That is what I am doing. My best model scores around 0.684 Local 5 fold CV with that. Scores like 0.69 on LB. a stack of 5 different models is scoring 0.697 Local CV and 0.692 LB....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448315,
          "author_name": "GuanQun Wu",
          "author_url": "",
          "post_date": "2018-12-31T19:12:44.277000",
          "content": "<p>hi, I calculate local CV on separate CV set.i have the same problem, my best local CV is 0.72, but its LB is 0.68, maybe overfitting?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 444676,
      "author_name": "Kepler456b",
      "author_url": "",
      "post_date": "2018-12-24T14:35:34.843000",
      "content": "<p>cv 0.69 lb 0.68. Strange</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 443002,
      "author_name": "danzel",
      "author_url": "",
      "post_date": "2018-12-20T21:33:43.293000",
      "content": "<p>cv = 0.6834\nlb = 0.687\n<strong>__ update\ncv = 0.6861\nlb = 0.691\n_<em></em></strong><em> update\ncv = 0.69068\nlb = 0.697 \n_</em>__ update\ncv = 0.696892\nlb = 0.702</p>",
      "votes": 1,
      "replies": [
        {
          "id": 444173,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-23T12:22:56.993000",
          "content": "<p>whant's you cv fold?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 444218,
          "author_name": "danzel",
          "author_url": "",
          "post_date": "2018-12-23T15:09:22.167000",
          "content": "<p>4 folds:\nFold: 1 Val F1 Score: 0.689334254780005 best thresh: 0.29\nFold: 2 Val F1 Score: 0.687243834863646 best thresh: 0.32\nFold: 3 Val F1 Score: 0.683858643744031 best thresh: 0.3\nFold: 4 Val F1 Score: 0.684088484904698 best thresh: 0.34</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447251,
          "author_name": "redshellspy",
          "author_url": "",
          "post_date": "2018-12-29T13:28:55.943000",
          "content": "<p>You are finding the threshold for each fold and then did majority voting for test data.</p>\n\n<p>Am I right ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451102,
          "author_name": "danzel",
          "author_url": "",
          "post_date": "2019-01-06T12:30:27.397000",
          "content": "<p>I tried a lot of stuff... i did some quick simulation on how to find the best threshold (<strong>majority voting</strong>, <strong>mean over cv-thresholds</strong>, <strong>geometric mean over cv-thresholds</strong>, <strong>static thresh of 0.33</strong>, <strong>calculate thresh via cv predictions</strong>, ...). Sadly it turns out that it is pretty unstable ... </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 446060,
      "author_name": "Rahul Agarwal",
      "author_url": "",
      "post_date": "2018-12-27T12:41:32.600000",
      "content": "<p>CV : 0.6806\nLB: 0.698</p>\n\n<p>Although I saw some negative correlation also between CV and LB, my highest score corresponds to my highest CV yet. Will update if I see anything different. I am using Spiros approach to calculate CV. Get out of fold predictions for training set and find F1 using that with Stratified k-folds</p>\n\n<p>Also attaching chart for LB vs CV</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 442212,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2018-12-19T16:10:46.697000",
      "content": "<p>My current gut feeling: trust local cv score rather than public lb score.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 442265,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-12-19T17:54:03.487000",
          "content": "<p>Should you trust the val F1 score more than the local val loss? My public lb score seems more correlated with increases in local val F1 than with decreases in the local val loss</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442320,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-19T19:43:19.630000",
          "content": "<p>Hard to say, I would say F1, but I am also checking auroc.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 447181,
      "author_name": "Spencer Kraisler",
      "author_url": "",
      "post_date": "2018-12-29T10:10:00.597000",
      "content": "<p>Newbie here but what is the difference between your CV and LB scores?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 447822,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-30T16:44:16.920000",
          "content": "<p>CV(Cross validation) score is the score which you get after doing a k-fold validation. </p>\n\n<p>What I mean is that if you run a 5-fold validation scheme. You train on 1,2,3,4 folds from the training data and predict the 5th fold from the training data. Once you are through with the cross validation loop, you have OOF(Out of fold) predictions for the whole training data. It is using these preds and the actual labels for training data that you get the local CV score. </p>\n\n<p>LB score is just the leaderboard score that is calculated using the predictions you provide on the test set. </p>\n\n<p>For example in the below kernel, I calculate it in this line: <a href=\"https://www.kaggle.com/mlwhiz/initializing-pytorch-layers-weight-with-kaiming\">https://www.kaggle.com/mlwhiz/initializing-pytorch-layers-weight-with-kaiming</a></p>\n\n<p>search_result = threshold_search(y_train, train_preds)</p>\n\n<p>Let me know if that clarifies your question.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 457467,
      "author_name": "Master",
      "author_url": "",
      "post_date": "2019-01-17T13:48:26.893000",
      "content": "<p>Wow, you are all so lucky with 0.7 LB and 0.68 CV. For me, it's:\nCV: 0.680\nLB: 0.690</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 447448,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-29T20:37:52.217000",
      "content": "<p>Single model, \nValidation score: 0.699\nLB: 0.690</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446501,
      "author_name": "mr007rin",
      "author_url": "",
      "post_date": "2018-12-28T07:36:55.017000",
      "content": "<p>Single model, 5 folders\ncv:0.687\nlb:0.691</p>",
      "votes": 0,
      "replies": [
        {
          "id": 446518,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-28T08:27:53.490000",
          "content": "<p>best cv:0.690     (5 folders)   lb:0.694</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446564,
          "author_name": "mr007rin",
          "author_url": "",
          "post_date": "2018-12-28T09:40:00.823000",
          "content": "<p>Nice job !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442842,
      "author_name": "Chevalier",
      "author_url": "",
      "post_date": "2018-12-20T15:42:42.897000",
      "content": "<p>my cv f1 is 0.677 and lb f1 is 0.698. I improve my local f1, but get lower LB score = =</p>",
      "votes": 0,
      "replies": [
        {
          "id": 442892,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-20T16:47:50.453000",
          "content": "<p>I think I notice a negative correlation in these levels. My best is cv f1 0.6781 and LB 0.700 . In another version local cv f1 0.6809 and LB 0.697.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 443639,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-22T02:08:59.307000",
          "content": "<p><em><strong>EDIT</strong></em>: Ok, just found out <a href=\"/spirosrap\">@spirosrap</a> excellent kernel : <a href=\"https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule\">https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule</a> . That is quite clear up my question below. </p>\n\n<p><em><strong>EDIT2</strong></em> So, I think the negative correlation between local CV and LB is <em>misleading</em>. In local CV, you based your score on a single classifier (trained on K-1 folds), but in LB you based your prediction on K classifiers. Therefore, they are incomparable. If you separate one hold out set (e.g. 10% of the dataset) before doing KFolds, in my experience , local CV on that hold out set will positively correlate with LB,</p>\n\n<p>---- Original Question ----\n<a href=\"/spirosrap\">@spirosrap</a> I'm curious when you says local CV F1, how did you measure that?\nI saw in other disscusions you mentioned that you use K-Fold with no holdout set.\nTherefore, when you said you best local CV is 0.6781 did you mean this:</p>\n\n<p>First <strong>at training time</strong>,  divide ALL data into (Strastified) K Folds <strong>without \"hold out\" set</strong>\nfor each fold J in range(K),\n    Train on K-1 Folds (all folds except J), measure F1 on the remaining fold J.</p>\n\n<p>Local CV 0.6781 (mentioned above) = average of this K results.</p>\n\n<p><strong>At submission</strong>, average the prediction of these K classifiers</p>\n\n<p>Am I correct, or am I completely wrong?\nI emphasize <strong>\"without hold out\"</strong> since if we pre-allocate a hold out set to measure F1 and to select threshold, it could be a totally different situation.</p>\n\n<p><a href=\"/bestpredict\">@bestpredict</a> <a href=\"/mldevl\">@mldevl</a> May I ask you also the same question?</p>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443663,
          "author_name": "Chevalier",
          "author_url": "",
          "post_date": "2018-12-22T04:27:04.217000",
          "content": "<p>yes,you are right. k-1 folds as training set, remaining fold as validation set. After k fold, you can get the full prediction of the training set. Then you can get a best threshold based on it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 443676,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-22T05:50:23.943000",
          "content": "<p><a href=\"/mldevl\">@mldevl</a> thanks!  However, I am not sure that I understand correctly, so I already edited my question to be more precise. Could you please take a look again ?</p>\n\n<p><em><strong>EDIT</strong></em>: Ok, just found out <a href=\"/spirosrap\">@spirosrap</a> excellent kernel : <a href=\"https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule\">https://www.kaggle.com/spirosrap/bilstm-attention-kfold-clr-extra-features-capsule</a> . That is quite clear up my question.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443747,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-22T10:26:27.650000",
          "content": "<p>I tried to keep a 10% holdout set but the LB score decreases significantly. So, I'm not sure if this is the right track. That small difference in the training data has a significant effect in the LB score. I think it was around 0.005 in one case.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 443790,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-22T12:48:54.053000",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> I think the dilemma of “good performance with full data” vs. “more accurate measurement with 10% holdout” can be solved by doing two experiments with the same random seed. </p>\n\n<p>In the first experiment, we use a 10% holdout just to estimate ensemble performance in general.\nIn the second experiment, we can use full data with the same random seed just to get the good LB performance.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 443885,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-22T16:55:28.607000",
          "content": "<p>Thank you, I'll try that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445680,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-26T23:36:54.970000",
          "content": "<p>Interesting... So this is where the hold out 10% before k-folds idea stems from? It makes sense that the the f1 folds would be misleading in comparison to the final f1 check, but wouldn't the averaging of validation predictions be okay anyway to do the final check?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447235,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2018-12-29T12:30:30.813000",
          "content": "<p>@spiro I have been using your awesome kernel and using the F1 score that we get by using train_preds as the CV F1 score. So my question is how much have you been able to improve that CV F1 score in your current implementation.</p>\n\n<p>I have not been able to improve it more than .6808 which fetched .696 while when it was at .6807, it was my current LB best at .698. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447458,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-29T21:17:29.863000",
          "content": "<p><a href=\"/learnmower\">@learnmower</a> Regarding to your last question. When you do the averaging of validation prediction, you do an “average of each single classifier”. But when you predict the test data, you use a “combination (ensemble) of all K classifiers”. I.e. they are two different models actually.</p>\n\n<p>If each single classifier is good, will you be sure that your ensemble is also great?</p>\n\n<p>Unfortuantely, not. Imagine that all of your K classifers are almost the same [i.e. they are all good but predict the same sigmoid probability to almost test data], here the ‘average prediction’ will not improve the results of the single classifier. Why? because averaging the same prob produce nothing new i.e. [0.7+0.7+0.7+0.7 +0.7]/5 = 0.7. ; so we get nothing if our classifiers are not diversified. </p>\n\n<p>Beside good individual classifier, diversification of base classifiers is another key to ensemble performance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447475,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-29T22:27:22.503000",
          "content": "<p>How are they two different models? When you run stratified k-fold, you are partitioning your data so that you run these folds into the same model K times. </p>\n\n<p><a href=\"https://scikit-learn.org/0.16/modules/generated/sklearn.cross_validation.StratifiedKFold.html\">https://scikit-learn.org/0.16/modules/generated/sklearn.cross_validation.StratifiedKFold.html</a></p>\n\n<p>And my question pertains to people who are replying to the single model thread who are holding out 10% of the training data before running their single model into a k-fold CV. This is something that I have not seen until this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 454052,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-01-11T04:40:57.020000",
          "content": "<p><a href=\"/learnmower\">@learnmower</a> Notice that I use the word ‘base classifier’ not ‘model’. Now Benjamin <a href=\"/bminixhofer\">@bminixhofer</a> has written a superb kernel of the exact same principle I have discussed here.</p>\n\n<p><a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442350,
      "author_name": "Pollux",
      "author_url": "",
      "post_date": "2018-12-19T21:09:17.990000",
      "content": "<p>Anyone else having .665 on CV, and .57 on submit? I must have a bug in my code.. :-/</p>",
      "votes": 0,
      "replies": [
        {
          "id": 442414,
          "author_name": "MitchelFung",
          "author_url": "",
          "post_date": "2018-12-19T23:45:35.497000",
          "content": "<p>If you're using CuDNNLSTMs on Keras a big issue is the nondeterminism of both it and Keras</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447499,
          "author_name": "Kepler456b",
          "author_url": "",
          "post_date": "2018-12-29T23:47:43.017000",
          "content": "<p>Not exactly but 0.69 cv and 0.68 lb. donno how others are other way round.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442039,
      "author_name": "Spiros",
      "author_url": "",
      "post_date": "2018-12-19T11:57:34.843000",
      "content": "<p>I find it hard to find a logical correlation between the two. Lower local score may give higher LB score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 442069,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-19T12:52:14.700000",
          "content": "<p>I think the reason may be that the test set data is not enough. and maybe  higher cv  f1 score  will get better score in  2nd's  test data</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 442008,
      "author_name": "HZD",
      "author_url": "",
      "post_date": "2018-12-19T10:54:46.463000",
      "content": "<p>During last week, I have raised my cv f1 score (4fold)  from ~0.680 to ~0.695, however, lb f1 score(0.04valid) remains unchanged at 0.690...o(╥﹏╥)o </p>",
      "votes": 0,
      "replies": [
        {
          "id": 442068,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-19T12:52:03.870000",
          "content": "<p>I think the reason may be that the test set data is not enough. and maybe  higher cv  f1 score  will get better score in  2nd's  test data</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 442102,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2018-12-19T13:44:54.387000",
          "content": "<p>Sorry, there I have a question. You are doing Cross Validation, why are you set a valid dataset of 0.04? Maybe you are doing a wrong CV. I think you are doing a blending operation, but you have run the same model 4 times.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442118,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-19T14:00:19.130000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442119,
          "author_name": "AIFIRST",
          "author_url": "",
          "post_date": "2018-12-19T14:00:55.193000",
          "content": "<p>would you like to share your local f1 score ? your lb score is highly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442134,
          "author_name": "HZD",
          "author_url": "",
          "post_date": "2018-12-19T14:15:32.777000",
          "content": "<p>I only do 4 fold cv in my PC for tuning the model. In kernel, I split 0.04 of the training data out as a valid set only to calculate f1 score for obtaining a reasonable threshold, that is, I use 0.96 of the training data as train set to get one classifier, and only use this one single classifier to get the prediction result...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442155,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-19T14:47:00.770000",
          "content": "<p>I think 4% Hold-out set is too small, and it will fluctuate a lot (should use at least &gt;10% hold out)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 442218,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-19T16:21:36.027000",
          "content": "<p>Have you tried Stratified K Fold? @gkd</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442240,
          "author_name": "HZD",
          "author_url": "",
          "post_date": "2018-12-19T16:54:28.417000",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> Yes, indeed I only used Stratified K Fold :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442248,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-19T17:12:49.223000",
          "content": "<p>@gkd Ok I though you were talking about the validation set. I noticed significant improvements when I used the whole training set without keeping a final test set for calculating the threshold. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442982,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2018-12-20T20:30:35.127000",
          "content": "<p>@Spiros How do you decide the threshold then? Set it to 0.5?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444041,
          "author_name": "SEU_Zesen_Chen",
          "author_url": "",
          "post_date": "2018-12-23T02:01:18.503000",
          "content": "<p>Well，if we want to ensemble the result with one single model，I think keep difference between dataset is important. So we usually perform k-fold to ensemble the result. If you split only 0.04 of the training data out，the models between 4 folds are amost the same because the train data in different folds are amost the same.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444195,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-23T13:43:02.460000",
          "content": "<p>@JihangZhang No I don't set it to 0.5. Depending on the final predictions, I search for the threshold with the optimal f1 score. Most of the kernels here do something similar.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444360,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-23T21:23:59.247000",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> Sorry how do you do this when you have the entire train set? Are you still splitting some proportion of data later as validation for the threshold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444379,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-23T22:56:35.370000",
          "content": "<p>@Learnmower In most of my experiments I have used the entire train set to calculate the threshold. I have made some experiments that I have a 0.1 hold out set to for the threshold but these didn't get me a good LB score (far lower).  It's a gamble as I see it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 444399,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-24T00:35:18.110000",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> Ah... Hm, do you see a better correlation when using the public LB as feedback? It might be worth the gamble at this point.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444414,
          "author_name": "Alan Khoa Nguyen",
          "author_url": "",
          "post_date": "2018-12-24T02:07:25.610000",
          "content": "<p><a href=\"/spirosrap\">@spirosrap</a> are you using a model architecture similar to the kernels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 444582,
          "author_name": "Spiros",
          "author_url": "",
          "post_date": "2018-12-24T10:26:44.200000",
          "content": "<p>@StevenNguyen Yes, Bidirectional LSTM,GRU and attention.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442464,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-20T01:53:43.490000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441889": "my  cv  f1  is  0.678  and  lb f1  is 0.699",
    "447611": "I am sure the leaderboard will shuffle like crazy at stage 2. What do you think?",
    "445034": "local cv 0.678 public lb 0.699 ,while i reach local cv 0.68+ ,get lower public lb instead",
    "444750": "It took me some time to understand PyTorch, but now I am seeing a correlation between my local validation loss to local f1 to public LB. My avg. loss is about 0.0657, my local f1 is about 0.689 and the public is 0.695. PyTorch gives me a consistent result so now I feel confident that the changes I'm making aren't due to random luck or lucky seed...  ",
    "443456": "My current best is a cv of 0.687 and a public lb of 0.69. Interestingly, my previous best had a cv of 0.678 and a public lb of 0.693.  So far, any of my kernels that have higher cv scores than my previous best of 0.678 keep getting lower and lower public lb scores. I know that evaluating models on the local cv should get you a better score on the private LB, but compared to most of the people in this discussion post, my public lb score is lower than what it should be for my current cv F1 score",
    "442841": "local CV 0.6749 and LB 0.699\n\nI noticed that every time when my local cv was bigger than 0.68 my LB score always not so good. ",
    "442581": "My local f1 score(0.10 valid) is 0.7062, but my lb is 0.697\nCry, cry, cry...   \nWhy are there such big differences between online and offline? ",
    "448061": "Can someone help me make sense of this?\nI used BiLSTM-attention-Kfold-CLR-Extra Features-capsule kernel by Spiros, which gave LB 0.696 straight out of the box.\nI then added a small tweak to modify the training data which gave CV score of 0.718 but a poor 0.567 on LB.\nDoes this mean that LB test data characteristics are different from the test data characteristics given to us?",
    "446653": "local cv about 0.68, and lb 0.704, keras is amazing",
    "444676": "cv 0.69 lb 0.68. Strange",
    "443002": "cv = 0.6834\nlb = 0.687\n____ update\ncv = 0.6861\nlb = 0.691\n____ update\ncv = 0.69068\nlb = 0.697 \n____ update\ncv = 0.696892\nlb = 0.702",
    "446060": "CV : 0.6806\nLB: 0.698\n\nAlthough I saw some negative correlation also between CV and LB, my highest score corresponds to my highest CV yet. Will update if I see anything different. I am using Spiros approach to calculate CV. Get out of fold predictions for training set and find F1 using that with Stratified k-folds\n\nAlso attaching chart for LB vs CV",
    "442212": "My current gut feeling: trust local cv score rather than public lb score.",
    "447181": "Newbie here but what is the difference between your CV and LB scores?",
    "457467": "Wow, you are all so lucky with 0.7 LB and 0.68 CV. For me, it's:\nCV: 0.680\nLB: 0.690",
    "447448": "Single model, \nValidation score: 0.699\nLB: 0.690",
    "446501": "Single model, 5 folders\ncv:0.687\nlb:0.691",
    "442842": "my cv f1 is 0.677 and lb f1 is 0.698. I improve my local f1, but get lower LB score = =",
    "442350": "Anyone else having .665 on CV, and .57 on submit? I must have a bug in my code.. :-/",
    "442039": "I find it hard to find a logical correlation between the two. Lower local score may give higher LB score.",
    "442008": "During last week, I have raised my cv f1 score (4fold)  from ~0.680 to ~0.695, however, lb f1 score(0.04valid) remains unchanged at 0.690...o(╥﹏╥)o ",
    "442464": ""
  }
}