{
  "id": 75753,
  "title": "Is your best model robust or sensitive?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75753",
  "author_name": "Neuron Engineer",
  "post_date": "2018-12-26T04:45:35.157000",
  "votes": 15,
  "comment_count": 29,
  "views": 0,
  "content": "<p>I believe one critical factor of this competition is the <strong>robustness / sensitivity of our models</strong>, i.e. how much does the performance change when we adjust some variables just a little.</p>\n\n<p>For example, my best model got LB 0.699, but when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690 (shocked; if you have never tried varying some constants before, I urge you to try and see if you have the same sensitivity problem; it can be fun too!)</p>\n\n<p>For other of my models, by changing different random seeds, the performance usually fluctuated around +/- 0.005 . </p>\n\n<p>If you can make your model quite robust, you have a much better chance to be durable during the private leaderboard shuffling.</p>",
  "messages": [
    {
      "id": 445267,
      "postDate": "2018-12-26T04:45:35.157Z",
      "content": "<p>I believe one critical factor of this competition is the <strong>robustness / sensitivity of our models</strong>, i.e. how much does the performance change when we adjust some variables just a little.</p>\n\n<p>For example, my best model got LB 0.699, but when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690 (shocked; if you have never tried varying some constants before, I urge you to try and see if you have the same sensitivity problem; it can be fun too!)</p>\n\n<p>For other of my models, by changing different random seeds, the performance usually fluctuated around +/- 0.005 . </p>\n\n<p>If you can make your model quite robust, you have a much better chance to be durable during the private leaderboard shuffling.</p>",
      "rawMarkdown": "I believe one critical factor of this competition is the **robustness / sensitivity of our models**, i.e. how much does the performance change when we adjust some variables just a little.\n\nFor example, my best model got LB 0.699, but when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690 (shocked; if you have never tried varying some constants before, I urge you to try and see if you have the same sensitivity problem; it can be fun too!)\n\nFor other of my models, by changing different random seeds, the performance usually fluctuated around +/- 0.005 . \n\nIf you can make your model quite robust, you have a much better chance to be durable during the private leaderboard shuffling.",
      "votes": 14
    },
    {
      "id": 445996,
      "postDate": "2018-12-27T10:29:32.497Z",
      "content": "<p>I've been experimenting with random seeds, preprocessing, etc., and I've noted that the biggest effect came from everything you do <strong>before</strong> loading the embedding.</p>\n\n<p>So my current theory is that the most important aspect is the initialization of the embedding matrix, i.e. the word vectors for your OOV words. For some seeds, the results are excellent. Change the seed for the matrix a bit, and the LB result will fall quickly. Once you have a good seed for the embedding matrix, any change to the preprocessing, vocab size, etc. will hurt the final performance, because that will change the shape of the embedding matrix and thus the distribution of OOV word vectors.</p>\n\n<p>What do you guys think? Maybe someone has experienced similar behavior.</p>",
      "rawMarkdown": "I've been experimenting with random seeds, preprocessing, etc., and I've noted that the biggest effect came from everything you do **before** loading the embedding.\n\nSo my current theory is that the most important aspect is the initialization of the embedding matrix, i.e. the word vectors for your OOV words. For some seeds, the results are excellent. Change the seed for the matrix a bit, and the LB result will fall quickly. Once you have a good seed for the embedding matrix, any change to the preprocessing, vocab size, etc. will hurt the final performance, because that will change the shape of the embedding matrix and thus the distribution of OOV word vectors.\n\nWhat do you guys think? Maybe someone has experienced similar behavior.",
      "votes": 7,
      "replies": [
        {
          "id": 446014,
          "postDate": "2018-12-27T10:47:50.473Z",
          "content": "<p>Thanks Max for sharing the relation between seed and generalization. In fact, I have been doing something similar on seeding (this is the second time since your MTL kernel :)  . Before seeing your opinions, my theory was that a proper seed, by chance, select a good split which make a training set generalizes to the test set better. I will think more about your point on embedding matrix. Thanks for sharing!</p>",
          "rawMarkdown": "Thanks Max for sharing the relation between seed and generalization. In fact, I have been doing something similar on seeding (this is the second time since your MTL kernel :)  . Before seeing your opinions, my theory was that a proper seed, by chance, select a good split which make a training set generalizes to the test set better. I will think more about your point on embedding matrix. Thanks for sharing!",
          "votes": 1
        },
        {
          "id": 446118,
          "postDate": "2018-12-27T14:33:29.297Z",
          "content": "<p><a href=\"/mschumacher\">@mschumacher</a> Max, I'm thinking about it. Could you please clarify what you mean by <strong>\"the initialization of the embedding matrix, i.e. the word vectors for your OOV words\"</strong> . </p>\n\n<p>From what we usually do in practice, the embedding matrix should not directly depend on a seed. Since the matrix is pretrained, and is indexed by words which are in turn coming from a pre-specified training set.  So if we use all words to be tokenized, the set of vocabularies should not depend on a seed as well. (and then also the OOV). It may depend on a seed if we do a train/valid split and use only a splitted training set to specify the vocabulary. Perhaps I am missing something?</p>\n\n<p><strong>EDIT:</strong>\nDo you mean this line (my teammate just poke me on this):\nembedding_matrix = np.random.normal(emb_mean, emb_std, (max_words, embed_size))</p>",
          "rawMarkdown": "@mschumacher Max, I'm thinking about it. Could you please clarify what you mean by **\"the initialization of the embedding matrix, i.e. the word vectors for your OOV words\"** . \n\nFrom what we usually do in practice, the embedding matrix should not directly depend on a seed. Since the matrix is pretrained, and is indexed by words which are in turn coming from a pre-specified training set.  So if we use all words to be tokenized, the set of vocabularies should not depend on a seed as well. (and then also the OOV). It may depend on a seed if we do a train/valid split and use only a splitted training set to specify the vocabulary. Perhaps I am missing something?\n\n**EDIT:**\nDo you mean this line (my teammate just poke me on this):\nembedding_matrix = np.random.normal(emb_mean, emb_std, (max_words, embed_size))",
          "votes": 2
        },
        {
          "id": 446190,
          "postDate": "2018-12-27T17:16:06.667Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> Yes, what your teammate said. The embedding matrix is initialized with a random matrix. Only the <em>known</em> word vectors are overwritten. The words that don't have an embedding vector are not overwritten and thus dependent on the initial init.</p>\n\n<p>Because of that, every preprocessing step will have an effect on your word embeddings. E.g. fixing typos will remove rows from the matrix (and maybe change order), changing the distribution of your final embeddings completely.</p>\n\n<p>The first time I started thinking about this: It seemed I had found a great seed (family secret 🙂) that got me great results on the LB. Looking at the embeddings, typos, etc., especially in hard examples, made me think of several preprocessing steps that should definitely improve my accuracy. But after adding these steps, my LB score was consistently low!</p>\n\n<p>So my theory is that the initial preprocessing and seed I had resulted in some lucky word vectors for OOV vocab. Adding more preprocessing changed the word index completely, resulting in worse word vectors for some important OOVs...</p>\n\n<p>That's the only explanation I could think of, anyway, though I sometimes doubt it's statistically plausible. Experiments I've conducted since then have not disproven this theory.</p>",
          "rawMarkdown": "@ratthachat Yes, what your teammate said. The embedding matrix is initialized with a random matrix. Only the *known* word vectors are overwritten. The words that don't have an embedding vector are not overwritten and thus dependent on the initial init.\n\nBecause of that, every preprocessing step will have an effect on your word embeddings. E.g. fixing typos will remove rows from the matrix (and maybe change order), changing the distribution of your final embeddings completely.\n\nThe first time I started thinking about this: It seemed I had found a great seed (family secret 🙂) that got me great results on the LB. Looking at the embeddings, typos, etc., especially in hard examples, made me think of several preprocessing steps that should definitely improve my accuracy. But after adding these steps, my LB score was consistently low!\n\nSo my theory is that the initial preprocessing and seed I had resulted in some lucky word vectors for OOV vocab. Adding more preprocessing changed the word index completely, resulting in worse word vectors for some important OOVs...\n\nThat's the only explanation I could think of, anyway, though I sometimes doubt it's statistically plausible. Experiments I've conducted since then have not disproven this theory.",
          "votes": 4
        },
        {
          "id": 446192,
          "postDate": "2018-12-27T17:22:08.533Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 446360,
          "postDate": "2018-12-28T00:17:45.510Z",
          "content": "<p>Thanks for sharing again Max. The only problem I doubt is that (mentioned by my teammate again) we have checked that our vocab set already cover around 99.x% of texts already, so OOV shouldn't affect that much? (checked by using Dieter's or TheoVeil great kernel on preprocessing.)</p>\n\n<p>Another question: why do you not think that a good split (train/val) on KFolds should be more pluasible explanation?</p>\n\n<p>BTW, off-topic, did you have any luck on MTL? (I did not)</p>",
          "rawMarkdown": "Thanks for sharing again Max. The only problem I doubt is that (mentioned by my teammate again) we have checked that our vocab set already cover around 99.x% of texts already, so OOV shouldn't affect that much? (checked by using Dieter's or TheoVeil great kernel on preprocessing.)\n\nAnother question: why do you not think that a good split (train/val) on KFolds should be more pluasible explanation?\n\nBTW, off-topic, did you have any luck on MTL? (I did not)",
          "votes": 1
        },
        {
          "id": 447800,
          "postDate": "2018-12-30T15:31:50.003Z",
          "content": "<p>I guess that OOV affect your results when you setting the num_words too large. As you say:</p>\n\n<blockquote>\n  <p>when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690</p>\n</blockquote>\n\n<p>We should find a good num_words to build our embedding matrix so that the embedding vector of OOV words is generalization. </p>",
          "rawMarkdown": "I guess that OOV affect your results when you setting the num_words too large. As you say:\n&gt; when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690\n\nWe should find a good num_words to build our embedding matrix so that the embedding vector of OOV words is generalization. ",
          "votes": 1
        },
        {
          "id": 447970,
          "postDate": "2018-12-31T00:45:23.450Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> thanks for sharing. Hmm I'm generally very confused about this matter. For me the situation is: I have pretty bad preprocessing, but I still get a very good score. When I improve the preprocessing in any way, the score gets much worse. Same when I change the seed before creating the embedding matrix. So it may be different in your case, but in my case everything seems to point to the explanation with the lucky distribution of OOV vectors... But I'm also not happy with this explanation.</p>\n\n<p>Oh, and I haven't seriously tried MTL in one of my top kernels yet. I'm not super confident it will work, but I also have a few more ideas I didn't mention in my public kernel :)</p>",
          "rawMarkdown": "@ratthachat thanks for sharing. Hmm I'm generally very confused about this matter. For me the situation is: I have pretty bad preprocessing, but I still get a very good score. When I improve the preprocessing in any way, the score gets much worse. Same when I change the seed before creating the embedding matrix. So it may be different in your case, but in my case everything seems to point to the explanation with the lucky distribution of OOV vectors... But I'm also not happy with this explanation.\n\nOh, and I haven't seriously tried MTL in one of my top kernels yet. I'm not super confident it will work, but I also have a few more ideas I didn't mention in my public kernel :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 445461,
      "postDate": "2018-12-26T14:03:48.660Z",
      "content": "<p>Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .</p>",
      "rawMarkdown": "Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .",
      "votes": 5,
      "replies": [
        {
          "id": 446004,
          "postDate": "2018-12-27T10:36:23.630Z",
          "content": "<p>Get downvote just by clarifying the purpose of the topic :)</p>",
          "rawMarkdown": "Get downvote just by clarifying the purpose of the topic :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 446284,
      "postDate": "2018-12-27T19:56:12.633Z",
      "content": "<p>I have just made an experiment by sampling 56k rows (test size) from my validation set and the variation of f1 score was very high. So it seems that trusting CV makes more sense in this competition.</p>",
      "rawMarkdown": "I have just made an experiment by sampling 56k rows (test size) from my validation set and the variation of f1 score was very high. So it seems that trusting CV makes more sense in this competition.",
      "votes": 3,
      "replies": [
        {
          "id": 446358,
          "postDate": "2018-12-28T00:14:32.637Z",
          "content": "<p>Hi Ahmet, good point, how high was that? And that has a small correlation to LB?</p>\n\n<p>PS. since you are here may i ask an off-topic issue? Just found out your great kernel on speeding up RNN using Keras mask layer. Do you think it can also be used to speed up what we have here?</p>",
          "rawMarkdown": "Hi Ahmet, good point, how high was that? And that has a small correlation to LB?\n\nPS. since you are here may i ask an off-topic issue? Just found out your great kernel on speeding up RNN using Keras mask layer. Do you think it can also be used to speed up what we have here?"
        },
        {
          "id": 446430,
          "postDate": "2018-12-28T03:59:29.053Z",
          "content": "<p>I hope this topic is not about \"correlation\" between CV and LB. Quoting your quote <code>Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .</code></p>",
          "rawMarkdown": "I hope this topic is not about \"correlation\" between CV and LB. Quoting your quote ```Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .```",
          "votes": 1
        },
        {
          "id": 446456,
          "postDate": "2018-12-28T05:14:31.377Z",
          "content": "<p>I am sorry if you feel that i offended you sarath. I didn't mean it. (and the one who always upvote your reply was me). I just clarified the purpose of topic, but I am open to discuss everything.</p>",
          "rawMarkdown": "I am sorry if you feel that i offended you sarath. I didn't mean it. (and the one who always upvote your reply was me). I just clarified the purpose of topic, but I am open to discuss everything.",
          "votes": 1
        },
        {
          "id": 446457,
          "postDate": "2018-12-28T05:21:24.283Z",
          "content": "<p>I was about to tell you the same thing back. I upvoted you back when I commented above. I brought up the correlation thing, because I thought someone will shed new lights to that too. I just started being active in Kaggle, so not that familiar with few stuffs.  Sorry from my side :-) . </p>",
          "rawMarkdown": "I was about to tell you the same thing back. I upvoted you back when I commented above. I brought up the correlation thing, because I thought someone will shed new lights to that too. I just started being active in Kaggle, so not that familiar with few stuffs.  Sorry from my side :-) . ",
          "votes": 1
        },
        {
          "id": 446515,
          "postDate": "2018-12-28T08:23:34.957Z",
          "content": "<p>Alright! Let us continue to do our best to solve the quora mystery, my kaggle fellow :)</p>",
          "rawMarkdown": "Alright! Let us continue to do our best to solve the quora mystery, my kaggle fellow :)",
          "votes": 1
        },
        {
          "id": 446531,
          "postDate": "2018-12-28T08:45:03.460Z",
          "content": "<p>Same threshold same model scores 0.641 on 250k validation set. When I sample random 55k rows from this validation set I can get scores between 0.620-0.660. I didn't make proper experiment but the distribution looked like Gaussian to me. And no LB submission involved in this experiment.</p>\n\n<p><a href=\"/ratthachat\">@ratthachat</a> Which kernel do you refer to exactly?</p>",
          "rawMarkdown": "Same threshold same model scores 0.641 on 250k validation set. When I sample random 55k rows from this validation set I can get scores between 0.620-0.660. I didn't make proper experiment but the distribution looked like Gaussian to me. And no LB submission involved in this experiment.\n\n@ratthachat Which kernel do you refer to exactly?",
          "votes": 1
        },
        {
          "id": 448228,
          "postDate": "2018-12-31T14:40:06.393Z",
          "content": "<p>Hi Ahmet, let us move, I already made a post into your kernel :\n<a href=\"https://www.kaggle.com/divrikwicky/fast-basic-lstm-with-proper-k-fold-sentimentembed/notebook\">https://www.kaggle.com/divrikwicky/fast-basic-lstm-with-proper-k-fold-sentimentembed/notebook</a></p>",
          "rawMarkdown": "Hi Ahmet, let us move, I already made a post into your kernel :\nhttps://www.kaggle.com/divrikwicky/fast-basic-lstm-with-proper-k-fold-sentimentembed/notebook"
        },
        {
          "id": 448244,
          "postDate": "2018-12-31T15:08:40.780Z",
          "content": "<p>Sorry, I didn't get any notification on it. I have noticed that CuDNNLSTM doesn't support Masking. I will look into it and let you know.</p>",
          "rawMarkdown": "Sorry, I didn't get any notification on it. I have noticed that CuDNNLSTM doesn't support Masking. I will look into it and let you know.",
          "votes": 1
        }
      ]
    },
    {
      "id": 445301,
      "postDate": "2018-12-26T07:10:05.120Z",
      "content": "<p>My current approach is to maximize my local CV, then try my best to make it correlated with LB score. If at the end the correlation still doesn't happen then I'll take a leap of faith and trust my CV.</p>",
      "rawMarkdown": "My current approach is to maximize my local CV, then try my best to make it correlated with LB score. If at the end the correlation still doesn't happen then I'll take a leap of faith and trust my CV.",
      "votes": 3,
      "replies": [
        {
          "id": 445378,
          "postDate": "2018-12-26T10:15:37.440Z",
          "content": "<p>Hi Khoi, so what is your current best local CV?</p>",
          "rawMarkdown": "Hi Khoi, so what is your current best local CV?"
        },
        {
          "id": 445408,
          "postDate": "2018-12-26T11:53:15.630Z",
          "content": "<p>For 0.85/0.15 split its ~0.708.</p>\n\n<p>For 5-fold CV, concat everything it's 0.686.</p>",
          "rawMarkdown": "For 0.85/0.15 split its ~0.708.\n\nFor 5-fold CV, concat everything it's 0.686."
        },
        {
          "id": 445413,
          "postDate": "2018-12-26T12:02:40.863Z",
          "content": "<p>Your 15% split F1 looks impressive indeed!</p>",
          "rawMarkdown": "Your 15% split F1 looks impressive indeed!"
        },
        {
          "id": 445416,
          "postDate": "2018-12-26T12:07:59.813Z",
          "content": "<p>I think that kernel was rather badly overfitted to validation data since I used linear regression to blend 7 models together, it also almost break the time limit so it probably won't survive phase 2. I only trust CV now.</p>",
          "rawMarkdown": "I think that kernel was rather badly overfitted to validation data since I used linear regression to blend 7 models together, it also almost break the time limit so it probably won't survive phase 2. I only trust CV now.",
          "votes": 1
        },
        {
          "id": 445422,
          "postDate": "2018-12-26T12:14:55.613Z",
          "content": "<p>From what you said :  by blending 7 models, did you use the 15% validation set as a training set to the blending ? If that is the case, yes, I agree, it has a high chance to overfit the validation set. In general, a validation set is not a set to be fitted. It sole purpose is to measure the generalization of the model.</p>",
          "rawMarkdown": "From what you said :  by blending 7 models, did you use the 15% validation set as a training set to the blending ? If that is the case, yes, I agree, it has a high chance to overfit the validation set. In general, a validation set is not a set to be fitted. It sole purpose is to measure the generalization of the model."
        },
        {
          "id": 446191,
          "postDate": "2018-12-27T17:19:27.277Z",
          "content": "<p>you can try setting another small percent of data as a test set and test your final model on that to get a sense of generalization.</p>",
          "rawMarkdown": "you can try setting another small percent of data as a test set and test your final model on that to get a sense of generalization.",
          "votes": 1
        },
        {
          "id": 447180,
          "postDate": "2018-12-29T10:09:04.793Z",
          "content": "<p>Could you explain what it means to use linear regression to blend different models?</p>",
          "rawMarkdown": "Could you explain what it means to use linear regression to blend different models?"
        },
        {
          "id": 447222,
          "postDate": "2018-12-29T11:46:56.280Z",
          "content": "<p>Let's say you have five models, and then you learn with a linear model how to best combine them (weight them) for prediction based on a holdout dataset.</p>",
          "rawMarkdown": "Let's say you have five models, and then you learn with a linear model how to best combine them (weight them) for prediction based on a holdout dataset.",
          "votes": 1
        }
      ]
    },
    {
      "id": 445375,
      "postDate": "2018-12-26T10:06:56.447Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 445996,
      "author_name": "Max Schumacher",
      "author_url": "",
      "post_date": "2018-12-27T10:29:32.497000",
      "content": "<p>I've been experimenting with random seeds, preprocessing, etc., and I've noted that the biggest effect came from everything you do <strong>before</strong> loading the embedding.</p>\n\n<p>So my current theory is that the most important aspect is the initialization of the embedding matrix, i.e. the word vectors for your OOV words. For some seeds, the results are excellent. Change the seed for the matrix a bit, and the LB result will fall quickly. Once you have a good seed for the embedding matrix, any change to the preprocessing, vocab size, etc. will hurt the final performance, because that will change the shape of the embedding matrix and thus the distribution of OOV word vectors.</p>\n\n<p>What do you guys think? Maybe someone has experienced similar behavior.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 446014,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-27T10:47:50.473000",
          "content": "<p>Thanks Max for sharing the relation between seed and generalization. In fact, I have been doing something similar on seeding (this is the second time since your MTL kernel :)  . Before seeing your opinions, my theory was that a proper seed, by chance, select a good split which make a training set generalizes to the test set better. I will think more about your point on embedding matrix. Thanks for sharing!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446118,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-27T14:33:29.297000",
          "content": "<p><a href=\"/mschumacher\">@mschumacher</a> Max, I'm thinking about it. Could you please clarify what you mean by <strong>\"the initialization of the embedding matrix, i.e. the word vectors for your OOV words\"</strong> . </p>\n\n<p>From what we usually do in practice, the embedding matrix should not directly depend on a seed. Since the matrix is pretrained, and is indexed by words which are in turn coming from a pre-specified training set.  So if we use all words to be tokenized, the set of vocabularies should not depend on a seed as well. (and then also the OOV). It may depend on a seed if we do a train/valid split and use only a splitted training set to specify the vocabulary. Perhaps I am missing something?</p>\n\n<p><strong>EDIT:</strong>\nDo you mean this line (my teammate just poke me on this):\nembedding_matrix = np.random.normal(emb_mean, emb_std, (max_words, embed_size))</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 446190,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2018-12-27T17:16:06.667000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> Yes, what your teammate said. The embedding matrix is initialized with a random matrix. Only the <em>known</em> word vectors are overwritten. The words that don't have an embedding vector are not overwritten and thus dependent on the initial init.</p>\n\n<p>Because of that, every preprocessing step will have an effect on your word embeddings. E.g. fixing typos will remove rows from the matrix (and maybe change order), changing the distribution of your final embeddings completely.</p>\n\n<p>The first time I started thinking about this: It seemed I had found a great seed (family secret 🙂) that got me great results on the LB. Looking at the embeddings, typos, etc., especially in hard examples, made me think of several preprocessing steps that should definitely improve my accuracy. But after adding these steps, my LB score was consistently low!</p>\n\n<p>So my theory is that the initial preprocessing and seed I had resulted in some lucky word vectors for OOV vocab. Adding more preprocessing changed the word index completely, resulting in worse word vectors for some important OOVs...</p>\n\n<p>That's the only explanation I could think of, anyway, though I sometimes doubt it's statistically plausible. Experiments I've conducted since then have not disproven this theory.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 446192,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-27T17:22:08.533000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446360,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-28T00:17:45.510000",
          "content": "<p>Thanks for sharing again Max. The only problem I doubt is that (mentioned by my teammate again) we have checked that our vocab set already cover around 99.x% of texts already, so OOV shouldn't affect that much? (checked by using Dieter's or TheoVeil great kernel on preprocessing.)</p>\n\n<p>Another question: why do you not think that a good split (train/val) on KFolds should be more pluasible explanation?</p>\n\n<p>BTW, off-topic, did you have any luck on MTL? (I did not)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447800,
          "author_name": "Salon_sai",
          "author_url": "",
          "post_date": "2018-12-30T15:31:50.003000",
          "content": "<p>I guess that OOV affect your results when you setting the num_words too large. As you say:</p>\n\n<blockquote>\n  <p>when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690</p>\n</blockquote>\n\n<p>We should find a good num_words to build our embedding matrix so that the embedding vector of OOV words is generalization. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447970,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2018-12-31T00:45:23.450000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> thanks for sharing. Hmm I'm generally very confused about this matter. For me the situation is: I have pretty bad preprocessing, but I still get a very good score. When I improve the preprocessing in any way, the score gets much worse. Same when I change the seed before creating the embedding matrix. So it may be different in your case, but in my case everything seems to point to the explanation with the lucky distribution of OOV vectors... But I'm also not happy with this explanation.</p>\n\n<p>Oh, and I haven't seriously tried MTL in one of my top kernels yet. I'm not super confident it will work, but I also have a few more ideas I didn't mention in my public kernel :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 445461,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-26T14:03:48.660000",
      "content": "<p>Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .</p>",
      "votes": 5,
      "replies": [
        {
          "id": 446004,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-27T10:36:23.630000",
          "content": "<p>Get downvote just by clarifying the purpose of the topic :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 446284,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2018-12-27T19:56:12.633000",
      "content": "<p>I have just made an experiment by sampling 56k rows (test size) from my validation set and the variation of f1 score was very high. So it seems that trusting CV makes more sense in this competition.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 446358,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-28T00:14:32.637000",
          "content": "<p>Hi Ahmet, good point, how high was that? And that has a small correlation to LB?</p>\n\n<p>PS. since you are here may i ask an off-topic issue? Just found out your great kernel on speeding up RNN using Keras mask layer. Do you think it can also be used to speed up what we have here?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446430,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2018-12-28T03:59:29.053000",
          "content": "<p>I hope this topic is not about \"correlation\" between CV and LB. Quoting your quote <code>Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446456,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-28T05:14:31.377000",
          "content": "<p>I am sorry if you feel that i offended you sarath. I didn't mean it. (and the one who always upvote your reply was me). I just clarified the purpose of topic, but I am open to discuss everything.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446457,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2018-12-28T05:21:24.283000",
          "content": "<p>I was about to tell you the same thing back. I upvoted you back when I commented above. I brought up the correlation thing, because I thought someone will shed new lights to that too. I just started being active in Kaggle, so not that familiar with few stuffs.  Sorry from my side :-) . </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446515,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-28T08:23:34.957000",
          "content": "<p>Alright! Let us continue to do our best to solve the quora mystery, my kaggle fellow :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446531,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-12-28T08:45:03.460000",
          "content": "<p>Same threshold same model scores 0.641 on 250k validation set. When I sample random 55k rows from this validation set I can get scores between 0.620-0.660. I didn't make proper experiment but the distribution looked like Gaussian to me. And no LB submission involved in this experiment.</p>\n\n<p><a href=\"/ratthachat\">@ratthachat</a> Which kernel do you refer to exactly?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 448228,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-31T14:40:06.393000",
          "content": "<p>Hi Ahmet, let us move, I already made a post into your kernel :\n<a href=\"https://www.kaggle.com/divrikwicky/fast-basic-lstm-with-proper-k-fold-sentimentembed/notebook\">https://www.kaggle.com/divrikwicky/fast-basic-lstm-with-proper-k-fold-sentimentembed/notebook</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448244,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-12-31T15:08:40.780000",
          "content": "<p>Sorry, I didn't get any notification on it. I have noticed that CuDNNLSTM doesn't support Masking. I will look into it and let you know.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 445301,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2018-12-26T07:10:05.120000",
      "content": "<p>My current approach is to maximize my local CV, then try my best to make it correlated with LB score. If at the end the correlation still doesn't happen then I'll take a leap of faith and trust my CV.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 445378,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-26T10:15:37.440000",
          "content": "<p>Hi Khoi, so what is your current best local CV?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445408,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2018-12-26T11:53:15.630000",
          "content": "<p>For 0.85/0.15 split its ~0.708.</p>\n\n<p>For 5-fold CV, concat everything it's 0.686.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445413,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-26T12:02:40.863000",
          "content": "<p>Your 15% split F1 looks impressive indeed!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445416,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2018-12-26T12:07:59.813000",
          "content": "<p>I think that kernel was rather badly overfitted to validation data since I used linear regression to blend 7 models together, it also almost break the time limit so it probably won't survive phase 2. I only trust CV now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445422,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-26T12:14:55.613000",
          "content": "<p>From what you said :  by blending 7 models, did you use the 15% validation set as a training set to the blending ? If that is the case, yes, I agree, it has a high chance to overfit the validation set. In general, a validation set is not a set to be fitted. It sole purpose is to measure the generalization of the model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446191,
          "author_name": "Ankit Sati",
          "author_url": "",
          "post_date": "2018-12-27T17:19:27.277000",
          "content": "<p>you can try setting another small percent of data as a test set and test your final model on that to get a sense of generalization.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447180,
          "author_name": "Spencer Kraisler",
          "author_url": "",
          "post_date": "2018-12-29T10:09:04.793000",
          "content": "<p>Could you explain what it means to use linear regression to blend different models?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447222,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-29T11:46:56.280000",
          "content": "<p>Let's say you have five models, and then you learn with a linear model how to best combine them (weight them) for prediction based on a holdout dataset.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 445375,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-26T10:06:56.447000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "445267": "I believe one critical factor of this competition is the **robustness / sensitivity of our models**, i.e. how much does the performance change when we adjust some variables just a little.\n\nFor example, my best model got LB 0.699, but when I changed the number of vocabularies ( max_features variable) a little bit, (from 110000 to 100000), its LB dropped to 0.690 (shocked; if you have never tried varying some constants before, I urge you to try and see if you have the same sensitivity problem; it can be fun too!)\n\nFor other of my models, by changing different random seeds, the performance usually fluctuated around +/- 0.005 . \n\nIf you can make your model quite robust, you have a much better chance to be durable during the private leaderboard shuffling.",
    "445996": "I've been experimenting with random seeds, preprocessing, etc., and I've noted that the biggest effect came from everything you do **before** loading the embedding.\n\nSo my current theory is that the most important aspect is the initialization of the embedding matrix, i.e. the word vectors for your OOV words. For some seeds, the results are excellent. Change the seed for the matrix a bit, and the LB result will fall quickly. Once you have a good seed for the embedding matrix, any change to the preprocessing, vocab size, etc. will hurt the final performance, because that will change the shape of the embedding matrix and thus the distribution of OOV word vectors.\n\nWhat do you guys think? Maybe someone has experienced similar behavior.",
    "445461": "Note that this topic does not about correlation between CV and LB. But about sensitivity of “small changes” of your program which will affect both LB and CV .",
    "446284": "I have just made an experiment by sampling 56k rows (test size) from my validation set and the variation of f1 score was very high. So it seems that trusting CV makes more sense in this competition.",
    "445301": "My current approach is to maximize my local CV, then try my best to make it correlated with LB score. If at the end the correlation still doesn't happen then I'll take a leap of faith and trust my CV.",
    "445375": ""
  }
}