{
  "id": 71946,
  "title": "What  is your single best model doing ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71946",
  "author_name": "AcademycalBastard",
  "post_date": "2018-11-18T16:26:38.867000",
  "votes": 50,
  "comment_count": 124,
  "views": 0,
  "content": "<p>It will be great to share insights  about  how is your best single model  performming.</p>",
  "messages": [
    {
      "id": 423589,
      "postDate": "2018-11-18T16:26:38.867Z",
      "content": "<p>It will be great to share insights  about  how is your best single model  performming.</p>",
      "rawMarkdown": "It will be great to share insights  about  how is your best single model  performming.",
      "votes": 49
    },
    {
      "id": 437792,
      "postDate": "2018-12-12T13:52:43.977Z",
      "content": "<p>Migrated to Pytorch. Single model - 0.700 on LB 4 Folds 5 epochs</p>",
      "rawMarkdown": "Migrated to Pytorch. Single model - 0.700 on LB 4 Folds 5 epochs",
      "votes": 11,
      "replies": [
        {
          "id": 437814,
          "postDate": "2018-12-12T14:45:16.840Z",
          "content": "<p>So your Pytorch results are better than your Keras ones?</p>",
          "rawMarkdown": "So your Pytorch results are better than your Keras ones?"
        },
        {
          "id": 437818,
          "postDate": "2018-12-12T15:00:55.950Z",
          "content": "<p>Yes, and the training time is reduced to. And, honestly I am loving the documentation. Its lit and clean</p>",
          "rawMarkdown": "Yes, and the training time is reduced to. And, honestly I am loving the documentation. Its lit and clean",
          "votes": 3
        },
        {
          "id": 437822,
          "postDate": "2018-12-12T15:10:21.080Z",
          "content": "<p>Pytorch is the way to go..</p>\n\n<p>Keras is sometimes limited and Tensorflow is a nightmare. </p>\n\n<p>I decided to switch to Pytorch since TGS competition, but I'm still too lazy for that. </p>",
          "rawMarkdown": "Pytorch is the way to go..\n\nKeras is sometimes limited and Tensorflow is a nightmare. \n\nI decided to switch to Pytorch since TGS competition, but I'm still too lazy for that. ",
          "votes": 2
        },
        {
          "id": 437866,
          "postDate": "2018-12-12T16:56:32.443Z",
          "content": "<p>Any tutorial recommendation to get into it? Also not sure how much changed with the new 1.0 Pytorch release, which I guess won't be in the docker image anyways though.</p>",
          "rawMarkdown": "Any tutorial recommendation to get into it? Also not sure how much changed with the new 1.0 Pytorch release, which I guess won't be in the docker image anyways though."
        },
        {
          "id": 437934,
          "postDate": "2018-12-12T19:32:28.980Z",
          "content": "<p>@shahebaz, does your pytorch model give consistent result? To me that is the most important thing I am trying to figure out here. Models with inconsistent results is pretty much useless imho.</p>",
          "rawMarkdown": "@shahebaz, does your pytorch model give consistent result? To me that is the most important thing I am trying to figure out here. Models with inconsistent results is pretty much useless imho."
        },
        {
          "id": 438282,
          "postDate": "2018-12-13T12:05:14.330Z",
          "content": "<p>hey Shahebaz/serigne can you mark down your pytorch building blocks?\ni am using double lstm + attention +2 dense layers, relu  and dropout(0.1),batch size is 512, getting 0.68 single model 3 fold(unable to run more then that in 2 hours). i am not reaching 0.69+.\npreprossing : special char spacing and lower casing all sentences.\nwhat am i doing wrong?</p>",
          "rawMarkdown": "hey Shahebaz/serigne can you mark down your pytorch building blocks?\ni am using double lstm + attention +2 dense layers, relu  and dropout(0.1),batch size is 512, getting 0.68 single model 3 fold(unable to run more then that in 2 hours). i am not reaching 0.69+.\npreprossing : special char spacing and lower casing all sentences.\nwhat am i doing wrong?\n"
        },
        {
          "id": 438427,
          "postDate": "2018-12-13T16:37:33.090Z",
          "content": "<p>I'm not using Pytorch, at least not for now.  But it's my aim to switch (at least for next images competitions I enter ) </p>\n\n<p>And my building blocks is different from those typical <code>LSTM/GRU+ Attention/Pooled layers  + Dense layers</code>  I see on kernels. But you can get very good score with these architectures, try to tune more your network and check for overfitting. </p>",
          "rawMarkdown": "I'm not using Pytorch, at least not for now.  But it's my aim to switch (at least for next images competitions I enter ) \n\nAnd my building blocks is different from those typical `LSTM/GRU+ Attention/Pooled layers  + Dense layers`  I see on kernels. But you can get very good score with these architectures, try to tune more your network and check for overfitting. ",
          "votes": 2
        },
        {
          "id": 438584,
          "postDate": "2018-12-13T23:04:30.110Z",
          "content": "<p>@Shahebaz - when you migrated your model from keras to pytorch did you get similar validation results? I just migrated my model to pytorch but the results are far worse than my keras model. As far as I can tell just about everything is identical.</p>",
          "rawMarkdown": "@Shahebaz - when you migrated your model from keras to pytorch did you get similar validation results? I just migrated my model to pytorch but the results are far worse than my keras model. As far as I can tell just about everything is identical."
        },
        {
          "id": 438720,
          "postDate": "2018-12-14T04:16:03.387Z",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a> My architecture for 0.700 in Pytorch is differnet that of keras. But, I observed including cyclic CLR made the results worse in Pytorch. I am yet to implement the same architecture that gave me best score with keras. (I hate my exams) </p>",
          "rawMarkdown": "@rdizzl3 My architecture for 0.700 in Pytorch is differnet that of keras. But, I observed including cyclic CLR made the results worse in Pytorch. I am yet to implement the same architecture that gave me best score with keras. (I hate my exams) "
        },
        {
          "id": 438721,
          "postDate": "2018-12-14T04:27:35.757Z",
          "content": "<p>@Shahebaz - thanks for the info! This is my first go with pytorch so I wasn't quite sure what to expect. I will tinker around some more with it :)</p>",
          "rawMarkdown": "@Shahebaz - thanks for the info! This is my first go with pytorch so I wasn't quite sure what to expect. I will tinker around some more with it :)"
        },
        {
          "id": 459514,
          "postDate": "2019-01-21T22:26:58.513Z",
          "content": "<p>Udemy has a good an free course on Pytorch: <a href=\"https://www.udacity.com/course/deep-learning-pytorch--ud188\">https://www.udacity.com/course/deep-learning-pytorch--ud188</a></p>",
          "rawMarkdown": "Udemy has a good an free course on Pytorch: https://www.udacity.com/course/deep-learning-pytorch--ud188",
          "votes": 1
        }
      ]
    },
    {
      "id": 426445,
      "postDate": "2018-11-23T09:11:55.917Z",
      "content": "<p>I have 0.505 with naïve Bayes :-)</p>",
      "rawMarkdown": "I have 0.505 with naïve Bayes :-)",
      "votes": 10
    },
    {
      "id": 437086,
      "postDate": "2018-12-11T11:06:14.427Z",
      "content": "<p>Edit :   LB 0.701,  I still use single model , 4 folds and some processing</p>",
      "rawMarkdown": "Edit :   LB 0.701,  I still use single model , 4 folds and some processing",
      "votes": 8,
      "replies": [
        {
          "id": 437095,
          "postDate": "2018-12-11T11:21:12.927Z",
          "content": "<p>What is your execution time?\nI got 0.701 with minor changes to one of the public kernels but it exceeds execution time limit (8220.5s). Not sure if I did something wrong but didn't get any error when submitting.</p>",
          "rawMarkdown": "What is your execution time?\nI got 0.701 with minor changes to one of the public kernels but it exceeds execution time limit (8220.5s). Not sure if I did something wrong but didn't get any error when submitting."
        },
        {
          "id": 437098,
          "postDate": "2018-12-11T11:23:31.557Z",
          "rawMarkdown": "",
          "votes": -1
        },
        {
          "id": 437267,
          "postDate": "2018-12-11T16:39:27.330Z",
          "content": "<p>It run 7177s  (a bit less than 2hrs :) )\nAnyway , optimisation will be necessary , 2nd stage will have bigger test data </p>",
          "rawMarkdown": "It run 7177s  (a bit less than 2hrs :) )\nAnyway , optimisation will be necessary , 2nd stage will have bigger test data ",
          "votes": 1
        },
        {
          "id": 437273,
          "postDate": "2018-12-11T16:47:26.213Z",
          "content": "<p>Nice, Serigne. Is your local f1 score pretty close to what you're seeing on the public LB?</p>",
          "rawMarkdown": "Nice, Serigne. Is your local f1 score pretty close to what you're seeing on the public LB?",
          "votes": 1
        },
        {
          "id": 437288,
          "postDate": "2018-12-11T17:22:01.570Z",
          "content": "<p>Congrats on breaking 0.7!</p>",
          "rawMarkdown": "Congrats on breaking 0.7!",
          "votes": 1
        },
        {
          "id": 437297,
          "postDate": "2018-12-11T17:29:04.537Z",
          "content": "<blockquote>\n  <p>Following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.</p>\n</blockquote>\n\n<p><a href=\"/serigne\">@serigne</a> I saw that there was bigger test in last kernel competition. But, I came across no mention of bigger test for this competition. Can you point at the source which says bigger or how big? Thanks!</p>",
          "rawMarkdown": "&gt; Following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.\n\n@serigne I saw that there was bigger test in last kernel competition. But, I came across no mention of bigger test for this competition. Can you point at the source which says bigger or how big? Thanks!\n\n"
        },
        {
          "id": 437378,
          "postDate": "2018-12-11T19:41:34.520Z",
          "content": "<p>@Shahebaz -</p>\n\n<blockquote>\n  <p>What will be available in the 2nd stage of the competition? In the\n  second stage of the competition, we will re-run your selected Kernels.\n  The following files will be swapped with new data:</p>\n  \n  <p>test.csv - This will be swapped with the complete public and private\n  test dataset. This file will have ~56k rows in stage 1 and ~376k rows\n  in stage 2. The public leaderboard data remains the same for both\n  versions. The file name will be the same (both test.csv) to ensure\n  that your code will run. sample_submission.csv - similar to test.csv,\n  this will be changed from ~56k in stage 1 to ~376k rows in stage 2 .\n  The file name will remain the same.</p>\n</blockquote>\n\n<p>From the data page so 2nd stage will be roughly ~7 times larger. </p>",
          "rawMarkdown": "@Shahebaz -\n\n&gt; What will be available in the 2nd stage of the competition? In the\n&gt; second stage of the competition, we will re-run your selected Kernels.\n&gt; The following files will be swapped with new data:\n&gt; \n&gt; test.csv - This will be swapped with the complete public and private\n&gt; test dataset. This file will have ~56k rows in stage 1 and ~376k rows\n&gt; in stage 2. The public leaderboard data remains the same for both\n&gt; versions. The file name will be the same (both test.csv) to ensure\n&gt; that your code will run. sample_submission.csv - similar to test.csv,\n&gt; this will be changed from ~56k in stage 1 to ~376k rows in stage 2 .\n&gt; The file name will remain the same.\n\nFrom the data page so 2nd stage will be roughly ~7 times larger. ",
          "votes": 3
        },
        {
          "id": 437535,
          "postDate": "2018-12-12T05:12:28.550Z",
          "content": "<p>Oh, I missed it. Thanks a bunch. </p>",
          "rawMarkdown": "Oh, I missed it. Thanks a bunch. "
        },
        {
          "id": 437544,
          "postDate": "2018-12-12T05:35:40.773Z",
          "content": "<p>@Learnmower  I didn't compute f1 in callback in my latest kernels, but my CV  loss 0.0955  ( it scores now LB 0.702) </p>\n\n<p>Thank @Shujian, No doubt you'll do it very soon :)</p>",
          "rawMarkdown": "@Learnmower  I didn't compute f1 in callback in my latest kernels, but my CV  loss 0.0955  ( it scores now LB 0.702) \n\nThank @Shujian, No doubt you'll do it very soon :)\n",
          "votes": 4
        },
        {
          "id": 437612,
          "postDate": "2018-12-12T07:34:28.620Z",
          "content": "<p>@Serigne Thanks! Your loss looks very good. </p>",
          "rawMarkdown": "@Serigne Thanks! Your loss looks very good. ",
          "votes": 1
        },
        {
          "id": 440816,
          "postDate": "2018-12-18T01:52:29.090Z",
          "content": "<p>I got my 4-fold cv down to 0.0955 loss and 0.693 f1-score locally but it still doesn't score well on the LB. It only scores 0.690. I am going to run it multiple times to see how the score changes over each run.</p>",
          "rawMarkdown": "I got my 4-fold cv down to 0.0955 loss and 0.693 f1-score locally but it still doesn't score well on the LB. It only scores 0.690. I am going to run it multiple times to see how the score changes over each run.",
          "votes": 1
        }
      ]
    },
    {
      "id": 428628,
      "postDate": "2018-11-27T16:05:11.790Z",
      "content": "<p>Adding to the CV vs. LB train:</p>\n\n<p>5-fold avg of 1 model, LB .687</p>\n\n<p>5 fold CV (with best threshold):</p>\n\n<p>Fold 0: .697\nFold 1: .687\nFold 2: .684\nFold 3: .688\nFold 4: .688</p>\n\n<p>I guess I should be happy that the LB score is a plausible CV fold score? I'd say I'm relatively confident that the leaderboard isn't super reliable though.</p>\n\n<p>Edit: also, same model trained on all data in one run scores .685</p>",
      "rawMarkdown": "Adding to the CV vs. LB train:\n\n5-fold avg of 1 model, LB .687\n\n5 fold CV (with best threshold):\n\nFold 0: .697\nFold 1: .687\nFold 2: .684\nFold 3: .688\nFold 4: .688\n\nI guess I should be happy that the LB score is a plausible CV fold score? I'd say I'm relatively confident that the leaderboard isn't super reliable though.\n\nEdit: also, same model trained on all data in one run scores .685",
      "votes": 8,
      "replies": [
        {
          "id": 428712,
          "postDate": "2018-11-27T18:43:31.717Z",
          "content": "<p>Do these 5 folds have a separate F1 threshold each, or do you lump all the validation data into a single set and estimate a single F1 threshold which you then apply to the 5 folds to get a fold-wise F1?</p>\n\n<p>If you are indeed using a separate threshold for each fold , doesn't the variance in the threshold scare you a bit (In some of my runs, the thresholds range from 0.19 to 0.39 across folds). </p>",
          "rawMarkdown": "Do these 5 folds have a separate F1 threshold each, or do you lump all the validation data into a single set and estimate a single F1 threshold which you then apply to the 5 folds to get a fold-wise F1?\n\nIf you are indeed using a separate threshold for each fold , doesn't the variance in the threshold scare you a bit (In some of my runs, the thresholds range from 0.19 to 0.39 across folds). ",
          "votes": 2
        },
        {
          "id": 428740,
          "postDate": "2018-11-27T19:49:30.557Z",
          "content": "<p>This was looking at separate thresholds for each fold, but I also don't see as much variance in thresholds across folds as you've seen - mine range between around .31 and .37. Checking for a global threshold after across the training data (using out of sample preds) also seems to give me something roughly near the average of thresholds (.35). I've also been ensembling using your voting strategy with different thresholds instead of just averaging together the probabilities, though they've seemed to work about the same for me.</p>\n\n<p>But yes, the variance in thresholds does scare me, and I'm really wondering what the proper way to ensemble is. In an ideal world I would want to just stack with something like logreg to get a single set of reasonably calibrated probabilistic outputs, but stacking is probably too slow and so far it's unclear to me that blending based on a single out-of-sample set works better than averaging different runs of the same model (like in the k-fold strategy). It's tricky! </p>",
          "rawMarkdown": "This was looking at separate thresholds for each fold, but I also don't see as much variance in thresholds across folds as you've seen - mine range between around .31 and .37. Checking for a global threshold after across the training data (using out of sample preds) also seems to give me something roughly near the average of thresholds (.35). I've also been ensembling using your voting strategy with different thresholds instead of just averaging together the probabilities, though they've seemed to work about the same for me.\n\nBut yes, the variance in thresholds does scare me, and I'm really wondering what the proper way to ensemble is. In an ideal world I would want to just stack with something like logreg to get a single set of reasonably calibrated probabilistic outputs, but stacking is probably too slow and so far it's unclear to me that blending based on a single out-of-sample set works better than averaging different runs of the same model (like in the k-fold strategy). It's tricky! ",
          "votes": 3
        }
      ]
    },
    {
      "id": 436288,
      "postDate": "2018-12-10T04:12:02.293Z",
      "content": "<p>0.699 single model with 5 folds. Slightly modified my public kernels according to some posts above.</p>",
      "rawMarkdown": "0.699 single model with 5 folds. Slightly modified my public kernels according to some posts above.",
      "votes": 6
    },
    {
      "id": 424011,
      "postDate": "2018-11-19T12:32:02.523Z",
      "content": "<p>LB : 0.694 ( single model with k-fold)</p>",
      "rawMarkdown": "LB : 0.694 ( single model with k-fold)",
      "votes": 6,
      "replies": [
        {
          "id": 424016,
          "postDate": "2018-11-19T12:47:23.780Z",
          "content": "<p>Same for me: 0.693 single model k-fold. It varies though, with the exact same model I got 0.685 too (I have not investigated yet)</p>",
          "rawMarkdown": "Same for me: 0.693 single model k-fold. It varies though, with the exact same model I got 0.685 too (I have not investigated yet)",
          "votes": 3
        },
        {
          "id": 424041,
          "postDate": "2018-11-19T13:20:29.483Z",
          "content": "<p>Wow , without any change ? </p>\n\n<p>My score varies between 0.691 and 0.694 when I do some changes either on the network or on the data  ( I didn't try to submit the same exact model twice  to check such variation ) </p>",
          "rawMarkdown": "Wow , without any change ? \n\nMy score varies between 0.691 and 0.694 when I do some changes either on the network or on the data  ( I didn't try to submit the same exact model twice  to check such variation ) ",
          "votes": 2
        },
        {
          "id": 424076,
          "postDate": "2018-11-19T14:30:38.543Z",
          "content": "<p>Congratz, pretty good for a single model. How is the runtime of that model approximately?</p>",
          "rawMarkdown": "Congratz, pretty good for a single model. How is the runtime of that model approximately?"
        },
        {
          "id": 424081,
          "postDate": "2018-11-19T14:38:24.253Z",
          "content": "<p>Yes, without any change. My best was an earlier version, I modified it, ran a few other trials, but there was no improvement. So I reverted back to the best version and I accidentally ran it again. I did not notice until I compared the two version with the diff-tool. I only changed some comments (the results: 0.693 / 0.685). I will look at this problem closer, it is weird.</p>",
          "rawMarkdown": "Yes, without any change. My best was an earlier version, I modified it, ran a few other trials, but there was no improvement. So I reverted back to the best version and I accidentally ran it again. I did not notice until I compared the two version with the diff-tool. I only changed some comments (the results: 0.693 / 0.685). I will look at this problem closer, it is weird."
        },
        {
          "id": 424082,
          "postDate": "2018-11-19T14:39:52.207Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, runtime for me was ~6000-6500s</p>",
          "rawMarkdown": "@philippsinger, runtime for me was ~6000-6500s"
        },
        {
          "id": 424202,
          "postDate": "2018-11-19T18:18:50.357Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Are you using Keras or  pure Tensorflow or Pytorch ? </p>\n\n<p>I think Keras may have some very random behavior compared to the two others .  I tried to fix all seeds , even for kernel_initializers. </p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a>,  for my best it's about 5800s</p>",
          "rawMarkdown": "@pestipeti Are you using Keras or  pure Tensorflow or Pytorch ? \n\nI think Keras may have some very random behavior compared to the two others .  I tried to fix all seeds , even for kernel_initializers. \n\n@philippsinger,  for my best it's about 5800s",
          "votes": 1
        },
        {
          "id": 424205,
          "postDate": "2018-11-19T18:23:53.920Z",
          "content": "<p><a href=\"/serigne\">@serigne</a>, I use keras, but I will switch to pytorch. I remember earlier posts (TGS competition) about the same weird behavior, they blamed keras too..</p>",
          "rawMarkdown": "@serigne, I use keras, but I will switch to pytorch. I remember earlier posts (TGS competition) about the same weird behavior, they blamed keras too.."
        },
        {
          "id": 424366,
          "postDate": "2018-11-20T01:35:06.577Z",
          "content": "<p>Are you doing some kind of text preprocessing with single models? I understand if you don't want to answer it :)</p>",
          "rawMarkdown": "Are you doing some kind of text preprocessing with single models? I understand if you don't want to answer it :)"
        },
        {
          "id": 424399,
          "postDate": "2018-11-20T03:56:51.867Z",
          "content": "<p>My best single model scores 0.676 on the LB, and it is a single run not k-Folds. </p>\n\n<p>@Serigne, @Peter, what is the score for your single model for one run? I mean before you did k-Folds?</p>\n\n<p>I have also noticed inconsistencies in scores with another model of mine. That one is an ensemble model so I thought may be I have a bug in my code but have not had time to debug it yet. I am using keras.</p>",
          "rawMarkdown": "My best single model scores 0.676 on the LB, and it is a single run not k-Folds. \n\n@Serigne, @Peter, what is the score for your single model for one run? I mean before you did k-Folds?\n\nI have also noticed inconsistencies in scores with another model of mine. That one is an ensemble model so I thought may be I have a bug in my code but have not had time to debug it yet. I am using keras.",
          "votes": 1
        },
        {
          "id": 424424,
          "postDate": "2018-11-20T05:23:40.750Z",
          "content": "<p><a href=\"/serigne\">@serigne</a>, \nI have been pretty frustrated with the non-reproducibility aspect of Keras based kernels  as well. This is not very well documented either.</p>\n\n<p>I currently use \nnp.random.seed(1)\ntf.set_random_seed(1)</p>\n\n<p>Some Kaggle and Stackoverflow thread also suggest that using CuDNNLSTM or CuDNNGRU can lead to non-deterministic behavior. I have not explore this further.</p>\n\n<p><a href=\"/serigne\">@serigne</a>, Would you mind sharing the random intializations that you use?</p>",
          "rawMarkdown": "@serigne, \nI have been pretty frustrated with the non-reproducibility aspect of Keras based kernels  as well. This is not very well documented either.\n\n\nI currently use \nnp.random.seed(1)\ntf.set_random_seed(1)\n\nSome Kaggle and Stackoverflow thread also suggest that using CuDNNLSTM or CuDNNGRU can lead to non-deterministic behavior. I have not explore this further.\n\n@serigne, Would you mind sharing the random intializations that you use?"
        },
        {
          "id": 424446,
          "postDate": "2018-11-20T06:14:15.157Z",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a>, My single best was 0.679, but because of this keras/cudnn issue it could be 0.688 (or 0.670). If I can not reproduce then it is worthless.</p>",
          "rawMarkdown": "@sheriytm, My single best was 0.679, but because of this keras/cudnn issue it could be 0.688 (or 0.670). If I can not reproduce then it is worthless.",
          "votes": 1
        },
        {
          "id": 424500,
          "postDate": "2018-11-20T08:02:47.677Z",
          "content": "<p>I feel the same way. Running experiments right now before I rule those models out. </p>\n\n<p>One thing I know is that I have used both CuDNNLSTM and CuDNNGRU in the toxic comment competition but with k-fold and did not observe any inconsistencies in my results. I am hoping doing k-fold will eliminate that. My current single models are too heavy for k-folds at the moment, so I need to work on streamlining that first.</p>",
          "rawMarkdown": "I feel the same way. Running experiments right now before I rule those models out. \n\nOne thing I know is that I have used both CuDNNLSTM and CuDNNGRU in the toxic comment competition but with k-fold and did not observe any inconsistencies in my results. I am hoping doing k-fold will eliminate that. My current single models are too heavy for k-folds at the moment, so I need to work on streamlining that first.",
          "votes": 1
        },
        {
          "id": 424569,
          "postDate": "2018-11-20T10:32:39.297Z",
          "content": "<p><a href=\"/mihajlot\">@mihajlot</a> ....No cleaning and no special processing for now </p>\n\n<p>I think the architecture of the network  is important here ...Some others architectures I tried, failed to improve after 3 or 4 epoches.</p>",
          "rawMarkdown": "@mihajlot ....No cleaning and no special processing for now \n\nI think the architecture of the network  is important here ...Some others architectures I tried, failed to improve after 3 or 4 epoches.",
          "votes": 4
        },
        {
          "id": 424571,
          "postDate": "2018-11-20T10:36:05.723Z",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> I didn't try it.  Actually I built k-folds since my first kernel , before starting experimenting models :) </p>\n\n<p>But I will later to see. </p>",
          "rawMarkdown": "@sheriytm I didn't try it.  Actually I built k-folds since my first kernel , before starting experimenting models :) \n\nBut I will later to see. ",
          "votes": 1
        },
        {
          "id": 424573,
          "postDate": "2018-11-20T10:43:32.240Z",
          "content": "<p><a href=\"/kagsen\">@kagsen</a> \nI fixed those and the kernel_initializers features like this : </p>\n\n<pre><code>from keras.initializers import he_normal, he_uniform,  glorot_normal,  glorot_uniform\n\nBidirectional(CuDNNLSTM(self.lstm_num_feat, \n                        kernel_initializer=glorot_uniform(seed = 123),\n                         return_sequences=True)) \n\nDense(self.dense_num_feat, \n          kernel_initializer=he_uniform(seed=123), \n          activation='relu')\n</code></pre>",
          "rawMarkdown": "@kagsen \nI fixed those and the kernel_initializers features like this : \n\n    from keras.initializers import he_normal, he_uniform,  glorot_normal,  glorot_uniform\n\n    Bidirectional(CuDNNLSTM(self.lstm_num_feat, \n                            kernel_initializer=glorot_uniform(seed = 123),\n                             return_sequences=True)) \n\n    Dense(self.dense_num_feat, \n              kernel_initializer=he_uniform(seed=123), \n              activation='relu')\n\n",
          "votes": 10
        },
        {
          "id": 424966,
          "postDate": "2018-11-21T00:09:28.743Z",
          "content": "<p>@Serigne, @Peter - how many folds are you guys using?</p>",
          "rawMarkdown": "@Serigne, @Peter - how many folds are you guys using?"
        },
        {
          "id": 424975,
          "postDate": "2018-11-21T00:31:54.127Z",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a>, I used 5 folds for my best result (6 epochs/fold)</p>",
          "rawMarkdown": "@rdizzl3, I used 5 folds for my best result (6 epochs/fold)",
          "votes": 2
        },
        {
          "id": 425034,
          "postDate": "2018-11-21T03:22:04.963Z",
          "content": "<p>Hey Serigne. How many folds are you using? I saw your post in another thread and with the timings, I had assumed 5. I'm also seeing convergence around 3-4 epochs. Though I am wondering if I should be basing my decision on loss or accuracy. Accuracy starts to drop about 3-4, but loss still improves.</p>",
          "rawMarkdown": "Hey Serigne. How many folds are you using? I saw your post in another thread and with the timings, I had assumed 5. I'm also seeing convergence around 3-4 epochs. Though I am wondering if I should be basing my decision on loss or accuracy. Accuracy starts to drop about 3-4, but loss still improves.\n"
        },
        {
          "id": 425048,
          "postDate": "2018-11-21T03:56:58.250Z",
          "content": "<p>@Serigne, I am really curious about what is k in your k-fold. Seems if k is small, the training set is too small so the results are not good. If k is large, the model has to be really simple so it can finish in 2 hrs.</p>",
          "rawMarkdown": "@Serigne, I am really curious about what is k in your k-fold. Seems if k is small, the training set is too small so the results are not good. If k is large, the model has to be really simple so it can finish in 2 hrs.",
          "votes": 1
        },
        {
          "id": 425158,
          "postDate": "2018-11-21T08:00:16.887Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 425320,
          "postDate": "2018-11-21T12:47:44.883Z",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a> and <a href=\"/learnmower\">@learnmower</a>  I splitted the data in 4 folds but use only 3 folds for now. So there are three runs and still a part of the data not used for validation.  However I plan to reintegrate the last fold as soon as I optimize better my code . </p>\n\n<p><a href=\"/shujian\">@shujian</a> I suppose k small or large depends also on the size of the test set ( and we don't have many choices here , as you say, due to 2hrs limitation )   .  Anyway I did not split the data randomly :) </p>",
          "rawMarkdown": "@rdizzl3 and @learnmower  I splitted the data in 4 folds but use only 3 folds for now. So there are three runs and still a part of the data not used for validation.  However I plan to reintegrate the last fold as soon as I optimize better my code . \n\n@shujian I suppose k small or large depends also on the size of the test set ( and we don't have many choices here , as you say, due to 2hrs limitation )   .  Anyway I did not split the data randomly :) ",
          "votes": 3
        },
        {
          "id": 425897,
          "postDate": "2018-11-22T09:17:41.313Z",
          "content": "<p>@Peter, @Serigne - if you don't mind saying, do you know your average K-Fold scores and overall score on OOF predictions?</p>",
          "rawMarkdown": "@Peter, @Serigne - if you don't mind saying, do you know your average K-Fold scores and overall score on OOF predictions?"
        },
        {
          "id": 425916,
          "postDate": "2018-11-22T09:55:29.153Z",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a>, For some reason kaggle cut my logs after ~300 lines, so I don't have all of it, but here are my local f1 scores for my best model:</p>\n\n<ul>\n<li>Fold 0 - 0.6771</li>\n<li>Fold 1 - 0.6740</li>\n<li>Fold 2 - 0.6812</li>\n<li>Fold 3 - 0.6751</li>\n</ul>\n\n<p>My public score for this (average of 5 folds) was 0.693, but read <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040\">this</a> too.</p>",
          "rawMarkdown": "@rdizzl3, For some reason kaggle cut my logs after ~300 lines, so I don't have all of it, but here are my local f1 scores for my best model:\n\n- Fold 0 - 0.6771\n- Fold 1 - 0.6740\n- Fold 2 - 0.6812\n- Fold 3 - 0.6751\n\nMy public score for this (average of 5 folds) was 0.693, but read [this](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040) too.",
          "votes": 2
        },
        {
          "id": 425952,
          "postDate": "2018-11-22T10:56:10.937Z",
          "content": "<p>@RDizzl3 my f1 val_scores : </p>\n\n<p>Fold0 : 0.6668\nFold1 :  0.6761\nFold2 : (skipped for now and will be reintegrated later)\nFold3 : 0.6846</p>",
          "rawMarkdown": "@RDizzl3 my f1 val_scores : \n\nFold0 : 0.6668\nFold1 :  0.6761\nFold2 : (skipped for now and will be reintegrated later)\nFold3 : 0.6846\n",
          "votes": 2
        },
        {
          "id": 425956,
          "postDate": "2018-11-22T11:00:44.553Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a>  you can append them on a list and print them at the next cell</p>\n\n<p>I use to run my kernels in interactive session too,  just after committing, so I can check out all the output later. </p>",
          "rawMarkdown": "@pestipeti  you can append them on a list and print them at the next cell\n\nI use to run my kernels in interactive session too,  just after committing, so I can check out all the output later. ",
          "votes": 1
        },
        {
          "id": 426206,
          "postDate": "2018-11-22T21:40:53.123Z",
          "content": "<p>Thank you guys for the replies! Check out my 5-folds:</p>\n\n<p>fold-0: 0.6779\nfold-1: 0.6821\nfold-2: 0.6860\nfold-3: 0.6883\nfold-4: 0.6817</p>\n\n<p>only scores 0.683 on the LB - average of 5-folds.</p>",
          "rawMarkdown": "Thank you guys for the replies! Check out my 5-folds:\n\nfold-0: 0.6779\nfold-1: 0.6821\nfold-2: 0.6860\nfold-3: 0.6883\nfold-4: 0.6817\n\nonly scores 0.683 on the LB - average of 5-folds.",
          "votes": 1
        },
        {
          "id": 426284,
          "postDate": "2018-11-23T02:23:24.630Z",
          "content": "<p><a href=\"/serigne\">@serigne</a>, hi, I am getting closer to your model :) </p>",
          "rawMarkdown": "@serigne, hi, I am getting closer to your model :) ",
          "votes": 2
        },
        {
          "id": 426443,
          "postDate": "2018-11-23T09:10:16.157Z",
          "content": "<p>@Peter how many epochs?</p>",
          "rawMarkdown": "@Peter how many epochs?"
        },
        {
          "id": 426737,
          "postDate": "2018-11-23T18:26:02.617Z",
          "content": "<p>@Shujian I have no  doubt you will do it very soon :)</p>",
          "rawMarkdown": "@Shujian I have no  doubt you will do it very soon :)",
          "votes": 1
        },
        {
          "id": 426739,
          "postDate": "2018-11-23T18:29:57.150Z",
          "content": "<p>Thanks ;)</p>",
          "rawMarkdown": "Thanks ;)"
        },
        {
          "id": 427108,
          "postDate": "2018-11-24T15:53:27.407Z",
          "content": "<p><a href=\"/rupeshs\">@rupeshs</a>, 6 epochs/fold</p>",
          "rawMarkdown": "@rupeshs, 6 epochs/fold"
        },
        {
          "id": 427340,
          "postDate": "2018-11-25T08:50:45.697Z",
          "content": "<p>A few people have mentioned the randomness that happens in keras CuDNNs - I've was curious enough to do some digging and found <a href=\"https://github.com/tensorflow/tensorflow/issues/2732#issuecomment-224661591\">this comment</a> in the tensorflow gh</p>\n\n<blockquote>\n  <p>On GPU, small amount of non-deterministic results is expected. TensorFlow uses the Eigen library, which uses Cuda atomics to implement reduction operations, such as tf.reduce_sum etc. Those operations are non-determnistical. Each operation can introduce a small difference. If your model is not stable, it could accumulate into large errors, after many steps.</p>\n</blockquote>\n\n<p>Wonder if this is just \"a GPU thing\" in general?</p>",
          "rawMarkdown": "A few people have mentioned the randomness that happens in keras CuDNNs - I've was curious enough to do some digging and found [this comment](https://github.com/tensorflow/tensorflow/issues/2732#issuecomment-224661591) in the tensorflow gh\n\n&gt; On GPU, small amount of non-deterministic results is expected. TensorFlow uses the Eigen library, which uses Cuda atomics to implement reduction operations, such as tf.reduce_sum etc. Those operations are non-determnistical. Each operation can introduce a small difference. If your model is not stable, it could accumulate into large errors, after many steps.\n\nWonder if this is just \"a GPU thing\" in general?",
          "votes": 2
        },
        {
          "id": 429707,
          "postDate": "2018-11-29T08:14:41.880Z",
          "content": "<p>My single model is at LB 0.696, CV 0.683  with 5 epochs and 4 folds</p>",
          "rawMarkdown": "My single model is at LB 0.696, CV 0.683  with 5 epochs and 4 folds"
        }
      ]
    },
    {
      "id": 452606,
      "postDate": "2019-01-09T00:06:44.217Z",
      "content": "<p>single model with 5 Stratified KFold, CV log_loss 0.09908, f1_score 0.67972, LB 0.684</p>",
      "rawMarkdown": "single model with 5 Stratified KFold, CV log_loss 0.09908, f1_score 0.67972, LB 0.684",
      "votes": 4
    },
    {
      "id": 426817,
      "postDate": "2018-11-23T23:29:09.667Z",
      "content": "<p>Currently have a LB of 0.688 with one model, no k fold and only GloVe embeddings</p>",
      "rawMarkdown": "Currently have a LB of 0.688 with one model, no k fold and only GloVe embeddings",
      "votes": 3
    },
    {
      "id": 436389,
      "postDate": "2018-12-10T08:17:50.007Z",
      "content": "<p>Adding on my kernel the cleaning from <a href=\"https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb\">this kernel</a>  improved my score from 0.696 to 0.699</p>\n\n<p>I didn't clean anything before. </p>",
      "rawMarkdown": "Adding on my kernel the cleaning from [this kernel][1]  improved my score from 0.696 to 0.699\n\nI didn't clean anything before. \n\n  [1]: https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb",
      "votes": 1,
      "replies": [
        {
          "id": 436396,
          "postDate": "2018-12-10T08:23:51.033Z",
          "content": "<p>I noted the same. Btw, are you sure improvement is cause of cleaning and not randomness. I use different processing than that of kernel and my average score of 10-15 runtimes is slower while clearing punctuations. Thanks for reporting though!</p>",
          "rawMarkdown": "I noted the same. Btw, are you sure improvement is cause of cleaning and not randomness. I use different processing than that of kernel and my average score of 10-15 runtimes is slower while clearing punctuations. Thanks for reporting though!"
        },
        {
          "id": 436399,
          "postDate": "2018-12-10T08:29:30.980Z",
          "content": "<p>I don't know...likely randomness, given I didn't notice huge improvement on CV :)</p>",
          "rawMarkdown": "I don't know...likely randomness, given I didn't notice huge improvement on CV :)\n"
        },
        {
          "id": 436401,
          "postDate": "2018-12-10T08:33:54.237Z",
          "content": "<p>I am -1 for any new experiments until I migrate to Pytorch. I haven't used it before. Never say never :) </p>",
          "rawMarkdown": "I am -1 for any new experiments until I migrate to Pytorch. I haven't used it before. Never say never :) "
        },
        {
          "id": 436425,
          "postDate": "2018-12-10T09:23:31.390Z",
          "content": "<p>@Serigne, same here. Adding the cleaning plus a couple of tweaks, my score improved from 0.693 to 0.698 but I am still concerned about the randomness/ inconsistencies.</p>",
          "rawMarkdown": "@Serigne, same here. Adding the cleaning plus a couple of tweaks, my score improved from 0.693 to 0.698 but I am still concerned about the randomness/ inconsistencies.",
          "votes": 1
        },
        {
          "id": 437153,
          "postDate": "2018-12-11T13:10:12.433Z",
          "content": "<p>I wonder how that cleaning makes sense for most of the characters in it as keras tokenizer is removing them anyways by default.</p>",
          "rawMarkdown": "I wonder how that cleaning makes sense for most of the characters in it as keras tokenizer is removing them anyways by default."
        },
        {
          "id": 437561,
          "postDate": "2018-12-12T06:01:15.457Z",
          "content": "<p>I also got scores in the ranges of 0.692 - 0.697 with cleaning. Even the same kernel when run again is giving different results. It's frustrating in the sense that we don't really know what improved our results and what not. @Serigne, can you run your highest scoring kernel again and check if the score remains same? If you don't mind wasting a submission.</p>",
          "rawMarkdown": "I also got scores in the ranges of 0.692 - 0.697 with cleaning. Even the same kernel when run again is giving different results. It's frustrating in the sense that we don't really know what improved our results and what not. @Serigne, can you run your highest scoring kernel again and check if the score remains same? If you don't mind wasting a submission.",
          "votes": 1
        }
      ]
    },
    {
      "id": 433084,
      "postDate": "2018-12-04T17:07:53.580Z",
      "content": "<p>I've got 0.691 on 4-folds with RMSprop optimizer and mean-combined embedding. I've tried perhaps 10 different parameter configurations but I can't get past this score on LB.</p>",
      "rawMarkdown": "I've got 0.691 on 4-folds with RMSprop optimizer and mean-combined embedding. I've tried perhaps 10 different parameter configurations but I can't get past this score on LB.",
      "votes": 1,
      "replies": [
        {
          "id": 433088,
          "postDate": "2018-12-04T17:18:00.680Z",
          "content": "<p>Hi Radek, did RMSprop optimizer boost LB? Thanks.</p>",
          "rawMarkdown": "Hi Radek, did RMSprop optimizer boost LB? Thanks."
        },
        {
          "id": 433113,
          "postDate": "2018-12-04T18:03:03.787Z",
          "content": "<p>It's hard to tell due to the LB score variance, but I think it helped in this case. I've used your network from <a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr</a> but went with one less epoch and second LSTM layer instead of GRU. I've tried tweaking learning rate, number of neurons in layers, dropout and so on, but nothing seems to help getting better LB score although it gets better locally on CV.</p>",
          "rawMarkdown": "It's hard to tell due to the LB score variance, but I think it helped in this case. I've used your network from https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr but went with one less epoch and second LSTM layer instead of GRU. I've tried tweaking learning rate, number of neurons in layers, dropout and so on, but nothing seems to help getting better LB score although it gets better locally on CV.",
          "votes": 1
        },
        {
          "id": 433119,
          "postDate": "2018-12-04T18:10:57.147Z",
          "content": "<p>You can take a look at it here: <a href=\"https://www.kaggle.com/rasvob/let-s-try-clr-v3\">https://www.kaggle.com/rasvob/let-s-try-clr-v3</a>\nLB with 0.691 is in Version 2 and 7, have fun :) </p>",
          "rawMarkdown": "You can take a look at it here: https://www.kaggle.com/rasvob/let-s-try-clr-v3\nLB with 0.691 is in Version 2 and 7, have fun :) ",
          "votes": 1
        },
        {
          "id": 433120,
          "postDate": "2018-12-04T18:10:59.323Z",
          "content": "<p>Thanks for your feedback!</p>",
          "rawMarkdown": "Thanks for your feedback!"
        }
      ]
    },
    {
      "id": 433803,
      "postDate": "2018-12-05T13:49:43.647Z",
      "content": "<p>0.697 - 4 Kfold , cyclic LR, mean combined embedding. But can't go above this even if I were able to improve CV. Which is very frustrating as I can't get any idea about what is what.</p>",
      "rawMarkdown": "0.697 - 4 Kfold , cyclic LR, mean combined embedding. But can't go above this even if I were able to improve CV. Which is very frustrating as I can't get any idea about what is what.",
      "votes": 2
    },
    {
      "id": 432970,
      "postDate": "2018-12-04T14:29:03.687Z",
      "content": "<p><code>79 mins</code>  -  <code>4fold</code>  -  <code>CV 0.685</code>  -  <code>LB 0.691</code></p>",
      "rawMarkdown": "`79 mins`  -  `4fold`  -  `CV 0.685`  -  `LB 0.691`",
      "votes": 2,
      "replies": [
        {
          "id": 433672,
          "postDate": "2018-12-05T10:26:03.960Z",
          "content": "<p>thank you</p>",
          "rawMarkdown": "thank you"
        },
        {
          "id": 434998,
          "postDate": "2018-12-07T09:33:08.087Z",
          "content": "<p>small old brother , please raise me up </p>",
          "rawMarkdown": "small old brother , please raise me up ",
          "votes": -1
        },
        {
          "id": 441123,
          "postDate": "2018-12-18T09:33:02.767Z",
          "content": "<p><code>83 min</code> - <code>5fold</code> - <code>CV 0.683</code> - <code>LB 0.697</code> - <code>pytorch</code></p>",
          "rawMarkdown": "`83 min` - `5fold` - `CV 0.683` - `LB 0.697` - `pytorch`",
          "votes": 4
        },
        {
          "id": 441702,
          "postDate": "2018-12-18T23:57:07.230Z",
          "content": "<p>Did you verify how many minutes can pytorch save you ?</p>",
          "rawMarkdown": "Did you verify how many minutes can pytorch save you ?"
        },
        {
          "id": 441744,
          "postDate": "2018-12-19T02:15:21.667Z",
          "content": "<p>@Neuron</p>\n\n<p>No, in the process of migrating to pytorch, data loading and model structure will have some necessary adjustments.</p>",
          "rawMarkdown": "@Neuron\n\nNo, in the process of migrating to pytorch, data loading and model structure will have some necessary adjustments."
        },
        {
          "id": 443849,
          "postDate": "2018-12-22T15:14:21.850Z",
          "content": "<p><code>100 min</code> - <code>5fold</code> - <code>CV 0.685</code> - <code>LB 0.698</code> - <code>pytorch</code> </p>",
          "rawMarkdown": "`100 min` - `5fold` - `CV 0.685` - `LB 0.698` - `pytorch ` ",
          "votes": 2
        },
        {
          "id": 447802,
          "postDate": "2018-12-30T15:34:11.177Z",
          "content": "<p><code>110 min</code> - <code>5fold</code> - <code>CV 0.683</code> - <code>LB 0.699</code> - <code>pytorch</code></p>",
          "rawMarkdown": "`110 min` - `5fold` - `CV 0.683` - `LB 0.699` - `pytorch `"
        },
        {
          "id": 448327,
          "postDate": "2018-12-31T19:37:01.823Z",
          "content": "<p>regular structure? (2 bi lstm +attention?)</p>",
          "rawMarkdown": "regular structure? (2 bi lstm +attention?)"
        }
      ]
    },
    {
      "id": 423760,
      "postDate": "2018-11-19T02:37:36.537Z",
      "content": "<p>My Best Single Model Public Leaderboard Score: 0.675</p>",
      "rawMarkdown": "My Best Single Model Public Leaderboard Score: 0.675",
      "votes": 2
    },
    {
      "id": 423869,
      "postDate": "2018-11-19T06:57:52.167Z",
      "content": "<p>0.685</p>",
      "rawMarkdown": "0.685"
    },
    {
      "id": 446459,
      "postDate": "2018-12-28T05:22:53.320Z",
      "content": "<p>change parameters improve LB score to 0.702, it may be overfit on testset,so,I'm worried about the second stage</p>",
      "rawMarkdown": "change parameters improve LB score to 0.702, it may be overfit on testset,so,I'm worried about the second stage",
      "votes": 1
    },
    {
      "id": 426442,
      "postDate": "2018-11-23T09:07:00.243Z",
      "content": "<p>LB: 66.3 (Single model ,no kfold)</p>",
      "rawMarkdown": "LB: 66.3 (Single model ,no kfold)",
      "votes": -1
    },
    {
      "id": 460246,
      "postDate": "2019-01-23T09:30:24.160Z",
      "content": "<p>Change parameters of public kernel\nLB: 0.701 CV 0.6823</p>",
      "rawMarkdown": "Change parameters of public kernel\nLB: 0.701 CV 0.6823"
    },
    {
      "id": 439764,
      "postDate": "2018-12-16T09:58:22.037Z",
      "content": "<p>Hi, is anybody able to go beyond both “LB 0.69” and “Local CV 0.69” , using only ‘single classifier’ ? (single model with no KFold, no ensemble) . In my case, even though my ensemble get LB0.697, my single only get LB0.675 (with 0.693 Local CV)</p>",
      "rawMarkdown": "Hi, is anybody able to go beyond both “LB 0.69” and “Local CV 0.69” , using only ‘single classifier’ ? (single model with no KFold, no ensemble) . In my case, even though my ensemble get LB0.697, my single only get LB0.675 (with 0.693 Local CV)",
      "replies": [
        {
          "id": 440268,
          "postDate": "2018-12-17T10:01:04.153Z",
          "content": "<p>I think your model might be overfitting...</p>",
          "rawMarkdown": "I think your model might be overfitting...",
          "votes": 1
        },
        {
          "id": 440281,
          "postDate": "2018-12-17T10:27:25.433Z",
          "content": "<p>You are right. However, I just would like to know if anybody can do &gt;= 0.69 on both LB and Local CV. Can you ?</p>",
          "rawMarkdown": "You are right. However, I just would like to know if anybody can do &gt;= 0.69 on both LB and Local CV. Can you ?"
        },
        {
          "id": 440285,
          "postDate": "2018-12-17T10:46:43.780Z",
          "content": "<p>What do you mean with Local CV no KFold?</p>",
          "rawMarkdown": "What do you mean with Local CV no KFold?",
          "votes": 1
        },
        {
          "id": 440297,
          "postDate": "2018-12-17T11:05:03.713Z",
          "content": "<p>Hi Philipp, </p>\n\n<p>Since there might be confusion about prediction using 'a single model with KFold' and 'a single classifier'.</p>\n\n<p>When we do KFold (although with a single model), in the end we have K classifiers (one for each validation fold), and usually we predict the final result using this ensemble of K classifiers. </p>\n\n<p>What I want to know is the performance of a single classifier (no average with other classifiers). Previously, I see one of the kaggler (top 10 in the current LB) claimed that he could get his very high performance by not using ensemble at all (but with some trick), but now his comment is deleted. </p>",
          "rawMarkdown": "Hi Philipp, \n\nSince there might be confusion about prediction using 'a single model with KFold' and 'a single classifier'.\n\nWhen we do KFold (although with a single model), in the end we have K classifiers (one for each validation fold), and usually we predict the final result using this ensemble of K classifiers. \n\nWhat I want to know is the performance of a single classifier (no average with other classifiers). Previously, I see one of the kaggler (top 10 in the current LB) claimed that he could get his very high performance by not using ensemble at all (but with some trick), but now his comment is deleted. ",
          "votes": 1
        },
        {
          "id": 440306,
          "postDate": "2018-12-17T11:25:57.840Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I think I may have also read the comment you mentioned in the forum. I think by single model, he meant KFold but not a kernel like SKR's that had 3-different types of models averaging the result.</p>",
          "rawMarkdown": "@ratthachat, I think I may have also read the comment you mentioned in the forum. I think by single model, he meant KFold but not a kernel like SKR's that had 3-different types of models averaging the result.",
          "votes": 1
        },
        {
          "id": 440320,
          "postDate": "2018-12-17T11:57:25.627Z",
          "content": "<p>That's also how I understood it.</p>",
          "rawMarkdown": "That's also how I understood it."
        },
        {
          "id": 440382,
          "postDate": "2018-12-17T13:28:29.993Z",
          "content": "<p>Thanks YaGana <a href=\"/sheriytm\">@sheriytm</a> ; But I remember vividly that he wrote ‘Single Model, No KFold, with some tricks’. Perhaps my memory is wrong :D</p>",
          "rawMarkdown": "Thanks YaGana @sheriytm ; But I remember vividly that he wrote ‘Single Model, No KFold, with some tricks’. Perhaps my memory is wrong :D"
        },
        {
          "id": 440636,
          "postDate": "2018-12-17T20:08:45.577Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, may be you remembered correctly and I missed that comment. I usually start with a single run model, get it to a good score before I go to KFold which is why I also searched the discussion section thoroughly for such results when I started. But I found most people were either doing ensembles or KFold you can see an example from my exchanges with Peter above.</p>",
          "rawMarkdown": "@ratthachat, may be you remembered correctly and I missed that comment. I usually start with a single run model, get it to a good score before I go to KFold which is why I also searched the discussion section thoroughly for such results when I started. But I found most people were either doing ensembles or KFold you can see an example from my exchanges with Peter above.",
          "votes": 1
        },
        {
          "id": 440915,
          "postDate": "2018-12-18T04:57:05.850Z",
          "content": "<p>Thanks YaGana.</p>",
          "rawMarkdown": "Thanks YaGana.",
          "votes": 1
        },
        {
          "id": 441663,
          "postDate": "2018-12-18T22:53:25.947Z",
          "content": "<p>@Neuron Engineer - I have finally managed to get my local CV at 0.692 F1 score and on the LB it is 0.691. The only thing is this is with K-fold but the validation is extremely close now and translating on the LB (so far). In the coming days I think I will check performance without K-fold and i'll report any findings here.</p>",
          "rawMarkdown": "@Neuron Engineer - I have finally managed to get my local CV at 0.692 F1 score and on the LB it is 0.691. The only thing is this is with K-fold but the validation is extremely close now and translating on the LB (so far). In the coming days I think I will check performance without K-fold and i'll report any findings here.",
          "votes": 1
        },
        {
          "id": 441700,
          "postDate": "2018-12-18T23:54:35.787Z",
          "content": "<p>Thanks <a href=\"/rdizzl3\">@rdizzl3</a>! Looking forward to hear that!</p>",
          "rawMarkdown": "Thanks @rdizzl3! Looking forward to hear that!"
        },
        {
          "id": 441979,
          "postDate": "2018-12-19T10:10:11.730Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 434126,
      "postDate": "2018-12-05T23:31:15.243Z",
      "content": "<p>Hm... some people have now reported having large differences between their local and public LB score... especially where these lower scores result in very high public LB scores. I am going to have to side with Eddy in that I'm probably much more dubious about standings on the public LB at this point.</p>",
      "rawMarkdown": "Hm... some people have now reported having large differences between their local and public LB score... especially where these lower scores result in very high public LB scores. I am going to have to side with Eddy in that I'm probably much more dubious about standings on the public LB at this point.",
      "replies": [
        {
          "id": 435702,
          "postDate": "2018-12-08T15:36:59.880Z",
          "content": "<p>It is very weird, I have some ensembles that score much higher on local CV and then suck badly on public LB, and then there are single models which score much worse on local CV, but get a boost on public LB. Not really sure which direction to go from here...</p>",
          "rawMarkdown": "It is very weird, I have some ensembles that score much higher on local CV and then suck badly on public LB, and then there are single models which score much worse on local CV, but get a boost on public LB. Not really sure which direction to go from here...",
          "votes": 1
        },
        {
          "id": 437275,
          "postDate": "2018-12-11T16:50:15.853Z",
          "content": "<p>Yeah. I think there's a lot of noise... It feels like this competition will be a toss-up.</p>",
          "rawMarkdown": "Yeah. I think there's a lot of noise... It feels like this competition will be a toss-up.",
          "votes": 1
        }
      ]
    },
    {
      "id": 429622,
      "postDate": "2018-11-29T04:38:37.280Z",
      "content": "<p><strong>Update:</strong> .690 LB for a super simple RNN with data that I briefly preprocessed beforehand.  Still looking into data augmentation, folds, post-processing, and hyper parameters to improve my solo model score even more.  My validation score for this was .6782.</p>",
      "rawMarkdown": "**Update:** .690 LB for a super simple RNN with data that I briefly preprocessed beforehand.  Still looking into data augmentation, folds, post-processing, and hyper parameters to improve my solo model score even more.  My validation score for this was .6782."
    },
    {
      "id": 427339,
      "postDate": "2018-11-25T08:44:31.513Z",
      "content": "<p>with a GRU I managed to get 0.677 ... I've only seen this once due to CuDNN randomness</p>\n\n<p>I'm really impressed some people are getting over 0.68!!!</p>",
      "rawMarkdown": "with a GRU I managed to get 0.677 ... I've only seen this once due to CuDNN randomness\n\nI'm really impressed some people are getting over 0.68!!!"
    },
    {
      "id": 425512,
      "postDate": "2018-11-21T17:58:24.650Z",
      "content": "<p>0.669 no kfold</p>",
      "rawMarkdown": "0.669 no kfold"
    },
    {
      "id": 425052,
      "postDate": "2018-11-21T04:04:21.527Z",
      "content": "<p>LB 0.679 in 1310 sec (GPU): <a href=\"https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7593124\">https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7593124</a></p>",
      "rawMarkdown": "LB 0.679 in 1310 sec (GPU): https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7593124\n",
      "replies": [
        {
          "id": 425718,
          "postDate": "2018-11-22T02:24:33.100Z",
          "content": "<p><strong>Update:</strong></p>\n\n<p>0.685 with 5 folds: <a href=\"https://www.kaggle.com/shujian/single-rnn-with-5-folds\">https://www.kaggle.com/shujian/single-rnn-with-5-folds</a></p>",
          "rawMarkdown": "**Update:**\n\n0.685 with 5 folds: https://www.kaggle.com/shujian/single-rnn-with-5-folds"
        },
        {
          "id": 426208,
          "postDate": "2018-11-22T21:42:58.660Z",
          "content": "<p>interesting! you’re local score is much lower than public...</p>\n\n<p>my f1 scores for a fold are as follows: \n1: 0.68603\n2: 0.68225\n3: 0.68375\n4: 0.68435\n5: 0.68521</p>\n\n<p>I trained up to 6 epochs.</p>\n\n<p>I tested the blend used by @kagSen and it scores a 0.687. I will try your blending method.</p>",
          "rawMarkdown": "interesting! you’re local score is much lower than public...\n\nmy f1 scores for a fold are as follows: \n1: 0.68603\n2: 0.68225\n3: 0.68375\n4: 0.68435\n5: 0.68521\n\nI trained up to 6 epochs.\n\nI tested the blend used by @kagSen and it scores a 0.687. I will try your blending method.\n",
          "votes": 2
        },
        {
          "id": 426209,
          "postDate": "2018-11-22T21:44:07.803Z",
          "content": "<p>Cool!</p>",
          "rawMarkdown": "Cool!"
        },
        {
          "id": 426234,
          "postDate": "2018-11-23T00:02:30.547Z",
          "content": "<p><strong>Update:</strong></p>\n\n<p>LB 0.691 single model with 5 folds and 6 epochs each: <a href=\"https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-8\">https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-8</a></p>\n\n<p>LB 0.692 single model with 4 folds and 8 epochs each: <a href=\"https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-9\">https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-9</a></p>",
          "rawMarkdown": "**Update:**\n\nLB 0.691 single model with 5 folds and 6 epochs each: https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-8\n\n\nLB 0.692 single model with 4 folds and 8 epochs each: https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-9",
          "votes": 3
        },
        {
          "id": 426309,
          "postDate": "2018-11-23T03:28:36.837Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 445650,
      "postDate": "2018-12-26T21:34:47Z",
      "content": "<p>LB: 0.701, Keras, single model, 5-fold CV, 5 epochs.\nF1  score from CV: 0.677!</p>\n\n<p>Edit: LB: 0.702 same model, different optim. parameters ...</p>",
      "rawMarkdown": "LB: 0.701, Keras, single model, 5-fold CV, 5 epochs.\nF1  score from CV: 0.677!\n\nEdit: LB: 0.702 same model, different optim. parameters ...",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 446091,
          "postDate": "2018-12-27T13:33:07.310Z",
          "content": "<p>Nice! Are you using an attention mechanism or capsule like I have seen others try?</p>",
          "rawMarkdown": "Nice! Are you using an attention mechanism or capsule like I have seen others try?"
        },
        {
          "id": 446320,
          "postDate": "2018-12-27T21:45:48.587Z",
          "content": "<p>No, only what comes with Keras: mostly fully-connected,  convolutional, and LSTM layers.\nI try to keep it simple to allow reasonable Nb CV folds </p>",
          "rawMarkdown": "No, only what comes with Keras: mostly fully-connected,  convolutional, and LSTM layers.\nI try to keep it simple to allow reasonable Nb CV folds ",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 448440,
          "postDate": "2019-01-01T07:33:51.297Z",
          "content": "<p>What is the execution time @Annabelle ? </p>",
          "rawMarkdown": "What is the execution time @Annabelle ? "
        },
        {
          "id": 448477,
          "postDate": "2019-01-01T09:37:48Z",
          "content": "<p>Too much to add another model: approx 6500 s total kernel execution time </p>",
          "rawMarkdown": "Too much to add another model: approx 6500 s total kernel execution time ",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 429623,
      "postDate": "2018-11-29T04:44:05.307Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 426610,
      "postDate": "2018-11-23T14:37:30.877Z",
      "content": "<p>F1 score = 0.622 , word2vec embedding, single XGB model</p>",
      "rawMarkdown": "F1 score = 0.622 , word2vec embedding, single XGB model",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 437792,
      "author_name": "Shahebaz Mohammad",
      "author_url": "",
      "post_date": "2018-12-12T13:52:43.977000",
      "content": "<p>Migrated to Pytorch. Single model - 0.700 on LB 4 Folds 5 epochs</p>",
      "votes": 11,
      "replies": [
        {
          "id": 437814,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-12T14:45:16.840000",
          "content": "<p>So your Pytorch results are better than your Keras ones?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437818,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-12T15:00:55.950000",
          "content": "<p>Yes, and the training time is reduced to. And, honestly I am loving the documentation. Its lit and clean</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 437822,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-12T15:10:21.080000",
          "content": "<p>Pytorch is the way to go..</p>\n\n<p>Keras is sometimes limited and Tensorflow is a nightmare. </p>\n\n<p>I decided to switch to Pytorch since TGS competition, but I'm still too lazy for that. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 437866,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-12T16:56:32.443000",
          "content": "<p>Any tutorial recommendation to get into it? Also not sure how much changed with the new 1.0 Pytorch release, which I guess won't be in the docker image anyways though.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437934,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-12T19:32:28.980000",
          "content": "<p>@shahebaz, does your pytorch model give consistent result? To me that is the most important thing I am trying to figure out here. Models with inconsistent results is pretty much useless imho.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438282,
          "author_name": "Alkimay",
          "author_url": "",
          "post_date": "2018-12-13T12:05:14.330000",
          "content": "<p>hey Shahebaz/serigne can you mark down your pytorch building blocks?\ni am using double lstm + attention +2 dense layers, relu  and dropout(0.1),batch size is 512, getting 0.68 single model 3 fold(unable to run more then that in 2 hours). i am not reaching 0.69+.\npreprossing : special char spacing and lower casing all sentences.\nwhat am i doing wrong?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438427,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-13T16:37:33.090000",
          "content": "<p>I'm not using Pytorch, at least not for now.  But it's my aim to switch (at least for next images competitions I enter ) </p>\n\n<p>And my building blocks is different from those typical <code>LSTM/GRU+ Attention/Pooled layers  + Dense layers</code>  I see on kernels. But you can get very good score with these architectures, try to tune more your network and check for overfitting. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 438584,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-12-13T23:04:30.110000",
          "content": "<p>@Shahebaz - when you migrated your model from keras to pytorch did you get similar validation results? I just migrated my model to pytorch but the results are far worse than my keras model. As far as I can tell just about everything is identical.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438720,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-14T04:16:03.387000",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a> My architecture for 0.700 in Pytorch is differnet that of keras. But, I observed including cyclic CLR made the results worse in Pytorch. I am yet to implement the same architecture that gave me best score with keras. (I hate my exams) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438721,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-12-14T04:27:35.757000",
          "content": "<p>@Shahebaz - thanks for the info! This is my first go with pytorch so I wasn't quite sure what to expect. I will tinker around some more with it :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 459514,
          "author_name": "JM100",
          "author_url": "",
          "post_date": "2019-01-21T22:26:58.513000",
          "content": "<p>Udemy has a good an free course on Pytorch: <a href=\"https://www.udacity.com/course/deep-learning-pytorch--ud188\">https://www.udacity.com/course/deep-learning-pytorch--ud188</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 426445,
      "author_name": "Andrzej Kuro",
      "author_url": "",
      "post_date": "2018-11-23T09:11:55.917000",
      "content": "<p>I have 0.505 with naïve Bayes :-)</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 437086,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-12-11T11:06:14.427000",
      "content": "<p>Edit :   LB 0.701,  I still use single model , 4 folds and some processing</p>",
      "votes": 8,
      "replies": [
        {
          "id": 437095,
          "author_name": "Mihajlo T.",
          "author_url": "",
          "post_date": "2018-12-11T11:21:12.927000",
          "content": "<p>What is your execution time?\nI got 0.701 with minor changes to one of the public kernels but it exceeds execution time limit (8220.5s). Not sure if I did something wrong but didn't get any error when submitting.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437098,
          "author_name": "Link",
          "author_url": "",
          "post_date": "2018-12-11T11:23:31.557000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 437267,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-11T16:39:27.330000",
          "content": "<p>It run 7177s  (a bit less than 2hrs :) )\nAnyway , optimisation will be necessary , 2nd stage will have bigger test data </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437273,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-11T16:47:26.213000",
          "content": "<p>Nice, Serigne. Is your local f1 score pretty close to what you're seeing on the public LB?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437288,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-12-11T17:22:01.570000",
          "content": "<p>Congrats on breaking 0.7!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437297,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-11T17:29:04.537000",
          "content": "<blockquote>\n  <p>Following the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.</p>\n</blockquote>\n\n<p><a href=\"/serigne\">@serigne</a> I saw that there was bigger test in last kernel competition. But, I came across no mention of bigger test for this competition. Can you point at the source which says bigger or how big? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437378,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-12-11T19:41:34.520000",
          "content": "<p>@Shahebaz -</p>\n\n<blockquote>\n  <p>What will be available in the 2nd stage of the competition? In the\n  second stage of the competition, we will re-run your selected Kernels.\n  The following files will be swapped with new data:</p>\n  \n  <p>test.csv - This will be swapped with the complete public and private\n  test dataset. This file will have ~56k rows in stage 1 and ~376k rows\n  in stage 2. The public leaderboard data remains the same for both\n  versions. The file name will be the same (both test.csv) to ensure\n  that your code will run. sample_submission.csv - similar to test.csv,\n  this will be changed from ~56k in stage 1 to ~376k rows in stage 2 .\n  The file name will remain the same.</p>\n</blockquote>\n\n<p>From the data page so 2nd stage will be roughly ~7 times larger. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 437535,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-12T05:12:28.550000",
          "content": "<p>Oh, I missed it. Thanks a bunch. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437544,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-12T05:35:40.773000",
          "content": "<p>@Learnmower  I didn't compute f1 in callback in my latest kernels, but my CV  loss 0.0955  ( it scores now LB 0.702) </p>\n\n<p>Thank @Shujian, No doubt you'll do it very soon :)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 437612,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-12T07:34:28.620000",
          "content": "<p>@Serigne Thanks! Your loss looks very good. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440816,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-12-18T01:52:29.090000",
          "content": "<p>I got my 4-fold cv down to 0.0955 loss and 0.693 f1-score locally but it still doesn't score well on the LB. It only scores 0.690. I am going to run it multiple times to see how the score changes over each run.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 428628,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-11-27T16:05:11.790000",
      "content": "<p>Adding to the CV vs. LB train:</p>\n\n<p>5-fold avg of 1 model, LB .687</p>\n\n<p>5 fold CV (with best threshold):</p>\n\n<p>Fold 0: .697\nFold 1: .687\nFold 2: .684\nFold 3: .688\nFold 4: .688</p>\n\n<p>I guess I should be happy that the LB score is a plausible CV fold score? I'd say I'm relatively confident that the leaderboard isn't super reliable though.</p>\n\n<p>Edit: also, same model trained on all data in one run scores .685</p>",
      "votes": 8,
      "replies": [
        {
          "id": 428712,
          "author_name": "kagSen",
          "author_url": "",
          "post_date": "2018-11-27T18:43:31.717000",
          "content": "<p>Do these 5 folds have a separate F1 threshold each, or do you lump all the validation data into a single set and estimate a single F1 threshold which you then apply to the 5 folds to get a fold-wise F1?</p>\n\n<p>If you are indeed using a separate threshold for each fold , doesn't the variance in the threshold scare you a bit (In some of my runs, the thresholds range from 0.19 to 0.39 across folds). </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 428740,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-11-27T19:49:30.557000",
          "content": "<p>This was looking at separate thresholds for each fold, but I also don't see as much variance in thresholds across folds as you've seen - mine range between around .31 and .37. Checking for a global threshold after across the training data (using out of sample preds) also seems to give me something roughly near the average of thresholds (.35). I've also been ensembling using your voting strategy with different thresholds instead of just averaging together the probabilities, though they've seemed to work about the same for me.</p>\n\n<p>But yes, the variance in thresholds does scare me, and I'm really wondering what the proper way to ensemble is. In an ideal world I would want to just stack with something like logreg to get a single set of reasonably calibrated probabilistic outputs, but stacking is probably too slow and so far it's unclear to me that blending based on a single out-of-sample set works better than averaging different runs of the same model (like in the k-fold strategy). It's tricky! </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 436288,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-12-10T04:12:02.293000",
      "content": "<p>0.699 single model with 5 folds. Slightly modified my public kernels according to some posts above.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 424011,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-11-19T12:32:02.523000",
      "content": "<p>LB : 0.694 ( single model with k-fold)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 424016,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-19T12:47:23.780000",
          "content": "<p>Same for me: 0.693 single model k-fold. It varies though, with the exact same model I got 0.685 too (I have not investigated yet)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 424041,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-19T13:20:29.483000",
          "content": "<p>Wow , without any change ? </p>\n\n<p>My score varies between 0.691 and 0.694 when I do some changes either on the network or on the data  ( I didn't try to submit the same exact model twice  to check such variation ) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 424076,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-11-19T14:30:38.543000",
          "content": "<p>Congratz, pretty good for a single model. How is the runtime of that model approximately?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424081,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-19T14:38:24.253000",
          "content": "<p>Yes, without any change. My best was an earlier version, I modified it, ran a few other trials, but there was no improvement. So I reverted back to the best version and I accidentally ran it again. I did not notice until I compared the two version with the diff-tool. I only changed some comments (the results: 0.693 / 0.685). I will look at this problem closer, it is weird.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424082,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-19T14:39:52.207000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, runtime for me was ~6000-6500s</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424202,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-19T18:18:50.357000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Are you using Keras or  pure Tensorflow or Pytorch ? </p>\n\n<p>I think Keras may have some very random behavior compared to the two others .  I tried to fix all seeds , even for kernel_initializers. </p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a>,  for my best it's about 5800s</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424205,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-19T18:23:53.920000",
          "content": "<p><a href=\"/serigne\">@serigne</a>, I use keras, but I will switch to pytorch. I remember earlier posts (TGS competition) about the same weird behavior, they blamed keras too..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424366,
          "author_name": "Mihajlo T.",
          "author_url": "",
          "post_date": "2018-11-20T01:35:06.577000",
          "content": "<p>Are you doing some kind of text preprocessing with single models? I understand if you don't want to answer it :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424399,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-20T03:56:51.867000",
          "content": "<p>My best single model scores 0.676 on the LB, and it is a single run not k-Folds. </p>\n\n<p>@Serigne, @Peter, what is the score for your single model for one run? I mean before you did k-Folds?</p>\n\n<p>I have also noticed inconsistencies in scores with another model of mine. That one is an ensemble model so I thought may be I have a bug in my code but have not had time to debug it yet. I am using keras.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424424,
          "author_name": "kagSen",
          "author_url": "",
          "post_date": "2018-11-20T05:23:40.750000",
          "content": "<p><a href=\"/serigne\">@serigne</a>, \nI have been pretty frustrated with the non-reproducibility aspect of Keras based kernels  as well. This is not very well documented either.</p>\n\n<p>I currently use \nnp.random.seed(1)\ntf.set_random_seed(1)</p>\n\n<p>Some Kaggle and Stackoverflow thread also suggest that using CuDNNLSTM or CuDNNGRU can lead to non-deterministic behavior. I have not explore this further.</p>\n\n<p><a href=\"/serigne\">@serigne</a>, Would you mind sharing the random intializations that you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424446,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-20T06:14:15.157000",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a>, My single best was 0.679, but because of this keras/cudnn issue it could be 0.688 (or 0.670). If I can not reproduce then it is worthless.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424500,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-11-20T08:02:47.677000",
          "content": "<p>I feel the same way. Running experiments right now before I rule those models out. </p>\n\n<p>One thing I know is that I have used both CuDNNLSTM and CuDNNGRU in the toxic comment competition but with k-fold and did not observe any inconsistencies in my results. I am hoping doing k-fold will eliminate that. My current single models are too heavy for k-folds at the moment, so I need to work on streamlining that first.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424569,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-20T10:32:39.297000",
          "content": "<p><a href=\"/mihajlot\">@mihajlot</a> ....No cleaning and no special processing for now </p>\n\n<p>I think the architecture of the network  is important here ...Some others architectures I tried, failed to improve after 3 or 4 epoches.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 424571,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-20T10:36:05.723000",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> I didn't try it.  Actually I built k-folds since my first kernel , before starting experimenting models :) </p>\n\n<p>But I will later to see. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 424573,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-20T10:43:32.240000",
          "content": "<p><a href=\"/kagsen\">@kagsen</a> \nI fixed those and the kernel_initializers features like this : </p>\n\n<pre><code>from keras.initializers import he_normal, he_uniform,  glorot_normal,  glorot_uniform\n\nBidirectional(CuDNNLSTM(self.lstm_num_feat, \n                        kernel_initializer=glorot_uniform(seed = 123),\n                         return_sequences=True)) \n\nDense(self.dense_num_feat, \n          kernel_initializer=he_uniform(seed=123), \n          activation='relu')\n</code></pre>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 424966,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-21T00:09:28.743000",
          "content": "<p>@Serigne, @Peter - how many folds are you guys using?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424975,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-21T00:31:54.127000",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a>, I used 5 folds for my best result (6 epochs/fold)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425034,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-11-21T03:22:04.963000",
          "content": "<p>Hey Serigne. How many folds are you using? I saw your post in another thread and with the timings, I had assumed 5. I'm also seeing convergence around 3-4 epochs. Though I am wondering if I should be basing my decision on loss or accuracy. Accuracy starts to drop about 3-4, but loss still improves.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425048,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-21T03:56:58.250000",
          "content": "<p>@Serigne, I am really curious about what is k in your k-fold. Seems if k is small, the training set is too small so the results are not good. If k is large, the model has to be really simple so it can finish in 2 hrs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425158,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-21T08:00:16.887000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425320,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-21T12:47:44.883000",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a> and <a href=\"/learnmower\">@learnmower</a>  I splitted the data in 4 folds but use only 3 folds for now. So there are three runs and still a part of the data not used for validation.  However I plan to reintegrate the last fold as soon as I optimize better my code . </p>\n\n<p><a href=\"/shujian\">@shujian</a> I suppose k small or large depends also on the size of the test set ( and we don't have many choices here , as you say, due to 2hrs limitation )   .  Anyway I did not split the data randomly :) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 425897,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-22T09:17:41.313000",
          "content": "<p>@Peter, @Serigne - if you don't mind saying, do you know your average K-Fold scores and overall score on OOF predictions?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425916,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-22T09:55:29.153000",
          "content": "<p><a href=\"/rdizzl3\">@rdizzl3</a>, For some reason kaggle cut my logs after ~300 lines, so I don't have all of it, but here are my local f1 scores for my best model:</p>\n\n<ul>\n<li>Fold 0 - 0.6771</li>\n<li>Fold 1 - 0.6740</li>\n<li>Fold 2 - 0.6812</li>\n<li>Fold 3 - 0.6751</li>\n</ul>\n\n<p>My public score for this (average of 5 folds) was 0.693, but read <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72040\">this</a> too.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425952,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-22T10:56:10.937000",
          "content": "<p>@RDizzl3 my f1 val_scores : </p>\n\n<p>Fold0 : 0.6668\nFold1 :  0.6761\nFold2 : (skipped for now and will be reintegrated later)\nFold3 : 0.6846</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425956,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-22T11:00:44.553000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a>  you can append them on a list and print them at the next cell</p>\n\n<p>I use to run my kernels in interactive session too,  just after committing, so I can check out all the output later. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426206,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-11-22T21:40:53.123000",
          "content": "<p>Thank you guys for the replies! Check out my 5-folds:</p>\n\n<p>fold-0: 0.6779\nfold-1: 0.6821\nfold-2: 0.6860\nfold-3: 0.6883\nfold-4: 0.6817</p>\n\n<p>only scores 0.683 on the LB - average of 5-folds.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426284,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-23T02:23:24.630000",
          "content": "<p><a href=\"/serigne\">@serigne</a>, hi, I am getting closer to your model :) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 426443,
          "author_name": "Rupesh Sreeraman",
          "author_url": "",
          "post_date": "2018-11-23T09:10:16.157000",
          "content": "<p>@Peter how many epochs?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426737,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-11-23T18:26:02.617000",
          "content": "<p>@Shujian I have no  doubt you will do it very soon :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426739,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-23T18:29:57.150000",
          "content": "<p>Thanks ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427108,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-11-24T15:53:27.407000",
          "content": "<p><a href=\"/rupeshs\">@rupeshs</a>, 6 epochs/fold</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427340,
          "author_name": "Hamish",
          "author_url": "",
          "post_date": "2018-11-25T08:50:45.697000",
          "content": "<p>A few people have mentioned the randomness that happens in keras CuDNNs - I've was curious enough to do some digging and found <a href=\"https://github.com/tensorflow/tensorflow/issues/2732#issuecomment-224661591\">this comment</a> in the tensorflow gh</p>\n\n<blockquote>\n  <p>On GPU, small amount of non-deterministic results is expected. TensorFlow uses the Eigen library, which uses Cuda atomics to implement reduction operations, such as tf.reduce_sum etc. Those operations are non-determnistical. Each operation can introduce a small difference. If your model is not stable, it could accumulate into large errors, after many steps.</p>\n</blockquote>\n\n<p>Wonder if this is just \"a GPU thing\" in general?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 429707,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-11-29T08:14:41.880000",
          "content": "<p>My single model is at LB 0.696, CV 0.683  with 5 epochs and 4 folds</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 452606,
      "author_name": "Qing Liu",
      "author_url": "",
      "post_date": "2019-01-09T00:06:44.217000",
      "content": "<p>single model with 5 Stratified KFold, CV log_loss 0.09908, f1_score 0.67972, LB 0.684</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 426817,
      "author_name": "Tim Johnson",
      "author_url": "",
      "post_date": "2018-11-23T23:29:09.667000",
      "content": "<p>Currently have a LB of 0.688 with one model, no k fold and only GloVe embeddings</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 436389,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-12-10T08:17:50.007000",
      "content": "<p>Adding on my kernel the cleaning from <a href=\"https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb\">this kernel</a>  improved my score from 0.696 to 0.699</p>\n\n<p>I didn't clean anything before. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 436396,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-10T08:23:51.033000",
          "content": "<p>I noted the same. Btw, are you sure improvement is cause of cleaning and not randomness. I use different processing than that of kernel and my average score of 10-15 runtimes is slower while clearing punctuations. Thanks for reporting though!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436399,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-10T08:29:30.980000",
          "content": "<p>I don't know...likely randomness, given I didn't notice huge improvement on CV :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436401,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-10T08:33:54.237000",
          "content": "<p>I am -1 for any new experiments until I migrate to Pytorch. I haven't used it before. Never say never :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436425,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-10T09:23:31.390000",
          "content": "<p>@Serigne, same here. Adding the cleaning plus a couple of tweaks, my score improved from 0.693 to 0.698 but I am still concerned about the randomness/ inconsistencies.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437153,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-11T13:10:12.433000",
          "content": "<p>I wonder how that cleaning makes sense for most of the characters in it as keras tokenizer is removing them anyways by default.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437561,
          "author_name": "Ankit Sati",
          "author_url": "",
          "post_date": "2018-12-12T06:01:15.457000",
          "content": "<p>I also got scores in the ranges of 0.692 - 0.697 with cleaning. Even the same kernel when run again is giving different results. It's frustrating in the sense that we don't really know what improved our results and what not. @Serigne, can you run your highest scoring kernel again and check if the score remains same? If you don't mind wasting a submission.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 433084,
      "author_name": "Radek Svoboda",
      "author_url": "",
      "post_date": "2018-12-04T17:07:53.580000",
      "content": "<p>I've got 0.691 on 4-folds with RMSprop optimizer and mean-combined embedding. I've tried perhaps 10 different parameter configurations but I can't get past this score on LB.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 433088,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-12-04T17:18:00.680000",
          "content": "<p>Hi Radek, did RMSprop optimizer boost LB? Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433113,
          "author_name": "Radek Svoboda",
          "author_url": "",
          "post_date": "2018-12-04T18:03:03.787000",
          "content": "<p>It's hard to tell due to the LB score variance, but I think it helped in this case. I've used your network from <a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr</a> but went with one less epoch and second LSTM layer instead of GRU. I've tried tweaking learning rate, number of neurons in layers, dropout and so on, but nothing seems to help getting better LB score although it gets better locally on CV.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433119,
          "author_name": "Radek Svoboda",
          "author_url": "",
          "post_date": "2018-12-04T18:10:57.147000",
          "content": "<p>You can take a look at it here: <a href=\"https://www.kaggle.com/rasvob/let-s-try-clr-v3\">https://www.kaggle.com/rasvob/let-s-try-clr-v3</a>\nLB with 0.691 is in Version 2 and 7, have fun :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433120,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-12-04T18:10:59.323000",
          "content": "<p>Thanks for your feedback!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433803,
      "author_name": "Ankit Sati",
      "author_url": "",
      "post_date": "2018-12-05T13:49:43.647000",
      "content": "<p>0.697 - 4 Kfold , cyclic LR, mean combined embedding. But can't go above this even if I were able to improve CV. Which is very frustrating as I can't get any idea about what is what.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 432970,
      "author_name": "Hikkiiiiiiiii",
      "author_url": "",
      "post_date": "2018-12-04T14:29:03.687000",
      "content": "<p><code>79 mins</code>  -  <code>4fold</code>  -  <code>CV 0.685</code>  -  <code>LB 0.691</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 433672,
          "author_name": "jetou Xu",
          "author_url": "",
          "post_date": "2018-12-05T10:26:03.960000",
          "content": "<p>thank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434998,
          "author_name": "AlexYung",
          "author_url": "",
          "post_date": "2018-12-07T09:33:08.087000",
          "content": "<p>small old brother , please raise me up </p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 441123,
          "author_name": "Hikkiiiiiiiii",
          "author_url": "",
          "post_date": "2018-12-18T09:33:02.767000",
          "content": "<p><code>83 min</code> - <code>5fold</code> - <code>CV 0.683</code> - <code>LB 0.697</code> - <code>pytorch</code></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 441702,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-18T23:57:07.230000",
          "content": "<p>Did you verify how many minutes can pytorch save you ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441744,
          "author_name": "Hikkiiiiiiiii",
          "author_url": "",
          "post_date": "2018-12-19T02:15:21.667000",
          "content": "<p>@Neuron</p>\n\n<p>No, in the process of migrating to pytorch, data loading and model structure will have some necessary adjustments.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443849,
          "author_name": "Hikkiiiiiiiii",
          "author_url": "",
          "post_date": "2018-12-22T15:14:21.850000",
          "content": "<p><code>100 min</code> - <code>5fold</code> - <code>CV 0.685</code> - <code>LB 0.698</code> - <code>pytorch</code> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 447802,
          "author_name": "Hikkiiiiiiiii",
          "author_url": "",
          "post_date": "2018-12-30T15:34:11.177000",
          "content": "<p><code>110 min</code> - <code>5fold</code> - <code>CV 0.683</code> - <code>LB 0.699</code> - <code>pytorch</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448327,
          "author_name": "Alkimay",
          "author_url": "",
          "post_date": "2018-12-31T19:37:01.823000",
          "content": "<p>regular structure? (2 bi lstm +attention?)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423760,
      "author_name": "Shreyans Dhankhar",
      "author_url": "",
      "post_date": "2018-11-19T02:37:36.537000",
      "content": "<p>My Best Single Model Public Leaderboard Score: 0.675</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423869,
      "author_name": "Si Feng",
      "author_url": "",
      "post_date": "2018-11-19T06:57:52.167000",
      "content": "<p>0.685</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446459,
      "author_name": "GuanQun Wu",
      "author_url": "",
      "post_date": "2018-12-28T05:22:53.320000",
      "content": "<p>change parameters improve LB score to 0.702, it may be overfit on testset,so,I'm worried about the second stage</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426442,
      "author_name": "Rupesh Sreeraman",
      "author_url": "",
      "post_date": "2018-11-23T09:07:00.243000",
      "content": "<p>LB: 66.3 (Single model ,no kfold)</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 460246,
      "author_name": "tamanyan",
      "author_url": "",
      "post_date": "2019-01-23T09:30:24.160000",
      "content": "<p>Change parameters of public kernel\nLB: 0.701 CV 0.6823</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 439764,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-12-16T09:58:22.037000",
      "content": "<p>Hi, is anybody able to go beyond both “LB 0.69” and “Local CV 0.69” , using only ‘single classifier’ ? (single model with no KFold, no ensemble) . In my case, even though my ensemble get LB0.697, my single only get LB0.675 (with 0.693 Local CV)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 440268,
          "author_name": "Link",
          "author_url": "",
          "post_date": "2018-12-17T10:01:04.153000",
          "content": "<p>I think your model might be overfitting...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440281,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-17T10:27:25.433000",
          "content": "<p>You are right. However, I just would like to know if anybody can do &gt;= 0.69 on both LB and Local CV. Can you ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440285,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-17T10:46:43.780000",
          "content": "<p>What do you mean with Local CV no KFold?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440297,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-17T11:05:03.713000",
          "content": "<p>Hi Philipp, </p>\n\n<p>Since there might be confusion about prediction using 'a single model with KFold' and 'a single classifier'.</p>\n\n<p>When we do KFold (although with a single model), in the end we have K classifiers (one for each validation fold), and usually we predict the final result using this ensemble of K classifiers. </p>\n\n<p>What I want to know is the performance of a single classifier (no average with other classifiers). Previously, I see one of the kaggler (top 10 in the current LB) claimed that he could get his very high performance by not using ensemble at all (but with some trick), but now his comment is deleted. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440306,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-17T11:25:57.840000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I think I may have also read the comment you mentioned in the forum. I think by single model, he meant KFold but not a kernel like SKR's that had 3-different types of models averaging the result.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440320,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-17T11:57:25.627000",
          "content": "<p>That's also how I understood it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440382,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-17T13:28:29.993000",
          "content": "<p>Thanks YaGana <a href=\"/sheriytm\">@sheriytm</a> ; But I remember vividly that he wrote ‘Single Model, No KFold, with some tricks’. Perhaps my memory is wrong :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440636,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-12-17T20:08:45.577000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, may be you remembered correctly and I missed that comment. I usually start with a single run model, get it to a good score before I go to KFold which is why I also searched the discussion section thoroughly for such results when I started. But I found most people were either doing ensembles or KFold you can see an example from my exchanges with Peter above.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 440915,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-18T04:57:05.850000",
          "content": "<p>Thanks YaGana.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441663,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-12-18T22:53:25.947000",
          "content": "<p>@Neuron Engineer - I have finally managed to get my local CV at 0.692 F1 score and on the LB it is 0.691. The only thing is this is with K-fold but the validation is extremely close now and translating on the LB (so far). In the coming days I think I will check performance without K-fold and i'll report any findings here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441700,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-12-18T23:54:35.787000",
          "content": "<p>Thanks <a href=\"/rdizzl3\">@rdizzl3</a>! Looking forward to hear that!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441979,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-19T10:10:11.730000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434126,
      "author_name": "Thomas Yokota",
      "author_url": "",
      "post_date": "2018-12-05T23:31:15.243000",
      "content": "<p>Hm... some people have now reported having large differences between their local and public LB score... especially where these lower scores result in very high public LB scores. I am going to have to side with Eddy in that I'm probably much more dubious about standings on the public LB at this point.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 435702,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-08T15:36:59.880000",
          "content": "<p>It is very weird, I have some ensembles that score much higher on local CV and then suck badly on public LB, and then there are single models which score much worse on local CV, but get a boost on public LB. Not really sure which direction to go from here...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437275,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-11T16:50:15.853000",
          "content": "<p>Yeah. I think there's a lot of noise... It feels like this competition will be a toss-up.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 429622,
      "author_name": "Matthew Anderson",
      "author_url": "",
      "post_date": "2018-11-29T04:38:37.280000",
      "content": "<p><strong>Update:</strong> .690 LB for a super simple RNN with data that I briefly preprocessed beforehand.  Still looking into data augmentation, folds, post-processing, and hyper parameters to improve my solo model score even more.  My validation score for this was .6782.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 427339,
      "author_name": "Hamish",
      "author_url": "",
      "post_date": "2018-11-25T08:44:31.513000",
      "content": "<p>with a GRU I managed to get 0.677 ... I've only seen this once due to CuDNN randomness</p>\n\n<p>I'm really impressed some people are getting over 0.68!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425512,
      "author_name": "Ankit Sati",
      "author_url": "",
      "post_date": "2018-11-21T17:58:24.650000",
      "content": "<p>0.669 no kfold</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425052,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-11-21T04:04:21.527000",
      "content": "<p>LB 0.679 in 1310 sec (GPU): <a href=\"https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7593124\">https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7593124</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 425718,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-22T02:24:33.100000",
          "content": "<p><strong>Update:</strong></p>\n\n<p>0.685 with 5 folds: <a href=\"https://www.kaggle.com/shujian/single-rnn-with-5-folds\">https://www.kaggle.com/shujian/single-rnn-with-5-folds</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426208,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-11-22T21:42:58.660000",
          "content": "<p>interesting! you’re local score is much lower than public...</p>\n\n<p>my f1 scores for a fold are as follows: \n1: 0.68603\n2: 0.68225\n3: 0.68375\n4: 0.68435\n5: 0.68521</p>\n\n<p>I trained up to 6 epochs.</p>\n\n<p>I tested the blend used by @kagSen and it scores a 0.687. I will try your blending method.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 426209,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-22T21:44:07.803000",
          "content": "<p>Cool!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426234,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-23T00:02:30.547000",
          "content": "<p><strong>Update:</strong></p>\n\n<p>LB 0.691 single model with 5 folds and 6 epochs each: <a href=\"https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-8\">https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-8</a></p>\n\n<p>LB 0.692 single model with 4 folds and 8 epochs each: <a href=\"https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-9\">https://www.kaggle.com/shujian/single-rnn-with-5-folds-v1-9</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 426309,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-23T03:28:36.837000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 445650,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-26T21:34:47",
      "content": "<p>LB: 0.701, Keras, single model, 5-fold CV, 5 epochs.\nF1  score from CV: 0.677!</p>\n\n<p>Edit: LB: 0.702 same model, different optim. parameters ...</p>",
      "votes": 3,
      "replies": [
        {
          "id": 446091,
          "author_name": "Bjenk Ellefsen",
          "author_url": "",
          "post_date": "2018-12-27T13:33:07.310000",
          "content": "<p>Nice! Are you using an attention mechanism or capsule like I have seen others try?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446320,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-27T21:45:48.587000",
          "content": "<p>No, only what comes with Keras: mostly fully-connected,  convolutional, and LSTM layers.\nI try to keep it simple to allow reasonable Nb CV folds </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 448440,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2019-01-01T07:33:51.297000",
          "content": "<p>What is the execution time @Annabelle ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448477,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-01T09:37:48",
          "content": "<p>Too much to add another model: approx 6500 s total kernel execution time </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 429623,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-29T04:44:05.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426610,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-23T14:37:30.877000",
      "content": "<p>F1 score = 0.622 , word2vec embedding, single XGB model</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "423589": "It will be great to share insights  about  how is your best single model  performming.",
    "437792": "Migrated to Pytorch. Single model - 0.700 on LB 4 Folds 5 epochs",
    "426445": "I have 0.505 with naïve Bayes :-)",
    "437086": "Edit :   LB 0.701,  I still use single model , 4 folds and some processing",
    "428628": "Adding to the CV vs. LB train:\n\n5-fold avg of 1 model, LB .687\n\n5 fold CV (with best threshold):\n\nFold 0: .697\nFold 1: .687\nFold 2: .684\nFold 3: .688\nFold 4: .688\n\nI guess I should be happy that the LB score is a plausible CV fold score? I'd say I'm relatively confident that the leaderboard isn't super reliable though.\n\nEdit: also, same model trained on all data in one run scores .685",
    "436288": "0.699 single model with 5 folds. Slightly modified my public kernels according to some posts above.",
    "424011": "LB : 0.694 ( single model with k-fold)",
    "452606": "single model with 5 Stratified KFold, CV log_loss 0.09908, f1_score 0.67972, LB 0.684",
    "426817": "Currently have a LB of 0.688 with one model, no k fold and only GloVe embeddings",
    "436389": "Adding on my kernel the cleaning from [this kernel][1]  improved my score from 0.696 to 0.699\n\nI didn't clean anything before. \n\n  [1]: https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb",
    "433084": "I've got 0.691 on 4-folds with RMSprop optimizer and mean-combined embedding. I've tried perhaps 10 different parameter configurations but I can't get past this score on LB.",
    "433803": "0.697 - 4 Kfold , cyclic LR, mean combined embedding. But can't go above this even if I were able to improve CV. Which is very frustrating as I can't get any idea about what is what.",
    "432970": "`79 mins`  -  `4fold`  -  `CV 0.685`  -  `LB 0.691`",
    "423760": "My Best Single Model Public Leaderboard Score: 0.675",
    "423869": "0.685",
    "446459": "change parameters improve LB score to 0.702, it may be overfit on testset,so,I'm worried about the second stage",
    "426442": "LB: 66.3 (Single model ,no kfold)",
    "460246": "Change parameters of public kernel\nLB: 0.701 CV 0.6823",
    "439764": "Hi, is anybody able to go beyond both “LB 0.69” and “Local CV 0.69” , using only ‘single classifier’ ? (single model with no KFold, no ensemble) . In my case, even though my ensemble get LB0.697, my single only get LB0.675 (with 0.693 Local CV)",
    "434126": "Hm... some people have now reported having large differences between their local and public LB score... especially where these lower scores result in very high public LB scores. I am going to have to side with Eddy in that I'm probably much more dubious about standings on the public LB at this point.",
    "429622": "**Update:** .690 LB for a super simple RNN with data that I briefly preprocessed beforehand.  Still looking into data augmentation, folds, post-processing, and hyper parameters to improve my solo model score even more.  My validation score for this was .6782.",
    "427339": "with a GRU I managed to get 0.677 ... I've only seen this once due to CuDNN randomness\n\nI'm really impressed some people are getting over 0.68!!!",
    "425512": "0.669 no kfold",
    "425052": "LB 0.679 in 1310 sec (GPU): https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7593124\n",
    "445650": "LB: 0.701, Keras, single model, 5-fold CV, 5 epochs.\nF1  score from CV: 0.677!\n\nEdit: LB: 0.702 same model, different optim. parameters ...",
    "429623": "",
    "426610": "F1 score = 0.622 , word2vec embedding, single XGB model"
  }
}