{
  "id": 75507,
  "title": "Can a single model exceed 0.702+",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75507",
  "author_name": "A.Barqawi",
  "post_date": "2018-12-22T14:10:24.462000",
  "votes": 29,
  "comment_count": 33,
  "views": 0,
  "content": "<p>My best single model reached 0.700 using Pytoch with LSTM+GRU+Attention which reached top 35 models up to this moment.<br>\nI used @Shahebaz advice to move into Pytoch which helped me to get better score comparing to Tensorflow/Keras models.<br>\nCan I continue with a single model to reach 0.702+ or should I start shifting the gears to ensembling?</p>",
  "messages": [
    {
      "id": 443825,
      "postDate": "2018-12-22T14:10:24.463Z",
      "content": "<p>My best single model reached 0.700 using Pytoch with LSTM+GRU+Attention which reached top 35 models up to this moment.<br>\nI used @Shahebaz advice to move into Pytoch which helped me to get better score comparing to Tensorflow/Keras models.<br>\nCan I continue with a single model to reach 0.702+ or should I start shifting the gears to ensembling?</p>",
      "rawMarkdown": "My best single model reached 0.700 using Pytoch with LSTM+GRU+Attention which reached top 35 models up to this moment.<br>\nI used @Shahebaz advice to move into Pytoch which helped me to get better score comparing to Tensorflow/Keras models.<br>\nCan I continue with a single model to reach 0.702+ or should I start shifting the gears to ensembling?",
      "votes": 28
    },
    {
      "id": 443903,
      "postDate": "2018-12-22T17:31:04.237Z",
      "content": "<p>Interesting that Pytorch gives you better results, what do you think leads to that? </p>",
      "rawMarkdown": "Interesting that Pytorch gives you better results, what do you think leads to that? ",
      "votes": 6,
      "replies": [
        {
          "id": 443940,
          "postDate": "2018-12-22T19:02:48.450Z",
          "content": "<p>Glad you found it useful. Thanks <a href=\"/jaguar00\">@jaguar00</a> (A.Barqawi)</p>\n\n<blockquote>\n  <p>what do you think leads to that? </p>\n</blockquote>\n\n<p>In my case. Pytorch trains faster and address bit of non-determism that give us a better understanding what works or not. But after a 0.700 I am experiencing lot of inconsistency with Pytorch too. </p>",
          "rawMarkdown": "Glad you found it useful. Thanks @jaguar00 (A.Barqawi)\n\n&gt;  what do you think leads to that? \n\nIn my case. Pytorch trains faster and address bit of non-determism that give us a better understanding what works or not. But after a 0.700 I am experiencing lot of inconsistency with Pytorch too. ",
          "votes": 3
        },
        {
          "id": 443943,
          "postDate": "2018-12-22T19:03:33.700Z",
          "content": "<p>The same structure I tried it with Tensorflow/Keras and Pytorch however, with Pytorch the results were higher, one reason maybe its a bit faster so more cross validation folds are allowed, another thing the results were less distracted (get almost same result range with each run) which helped me to know which things really work and enhance on top of it.</p>",
          "rawMarkdown": "The same structure I tried it with Tensorflow/Keras and Pytorch however, with Pytorch the results were higher, one reason maybe its a bit faster so more cross validation folds are allowed, another thing the results were less distracted (get almost same result range with each run) which helped me to know which things really work and enhance on top of it.",
          "votes": 3
        },
        {
          "id": 443944,
          "postDate": "2018-12-22T19:05:54.360Z",
          "content": "<p>Here's a speed comparision :) </p>\n\n<p><img src=\"https://wrosinski.github.io/assets/images/framework_plots/epochs_times_comparison.png\" alt=\"\"></p>\n\n<p>Source - <a href=\"https://wrosinski.github.io/deep-learning-frameworks/\">https://wrosinski.github.io/deep-learning-frameworks/</a></p>",
          "rawMarkdown": "Here's a speed comparision :) \n\n\n![](https://wrosinski.github.io/assets/images/framework_plots/epochs_times_comparison.png)\n\n\nSource - https://wrosinski.github.io/deep-learning-frameworks/",
          "votes": 14
        },
        {
          "id": 443945,
          "postDate": "2018-12-22T19:06:09.083Z",
          "content": "<p>HH I wrote the reason at the same time then relies you post the same thing Shahebaz. Thanks for the advice to move into Pytorch helped me a lot to enhance the model</p>",
          "rawMarkdown": "HH I wrote the reason at the same time then relies you post the same thing Shahebaz. Thanks for the advice to move into Pytorch helped me a lot to enhance the model",
          "votes": 1
        },
        {
          "id": 443948,
          "postDate": "2018-12-22T19:09:35.850Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 443950,
          "postDate": "2018-12-22T19:20:08.267Z",
          "content": "<p>Interesting, probably I am doing something wrong with Pytorch because it is training slower :) I will stick to Keras for now most likely.</p>",
          "rawMarkdown": "Interesting, probably I am doing something wrong with Pytorch because it is training slower :) I will stick to Keras for now most likely.",
          "votes": 3
        },
        {
          "id": 444012,
          "postDate": "2018-12-22T23:15:47.223Z",
          "content": "<p>In my case, it's faster than Keras but I couldn't replicate Keras score. I could be missing something also :)</p>",
          "rawMarkdown": "In my case, it's faster than Keras but I couldn't replicate Keras score. I could be missing something also :)",
          "votes": 3
        },
        {
          "id": 444067,
          "postDate": "2018-12-23T04:06:10.150Z",
          "content": "<blockquote>\n  <p>another thing the results were less distracted (get almost same result range with each run) which helped me to know which things really work and enhance on top of it.</p>\n</blockquote>\n\n<p>Totally agree with this.</p>",
          "rawMarkdown": "&gt; another thing the results were less distracted (get almost same result range with each run) which helped me to know which things really work and enhance on top of it.\n\nTotally agree with this.",
          "votes": 3
        },
        {
          "id": 444068,
          "postDate": "2018-12-23T04:07:59.210Z",
          "content": "<p><a href=\"/mihajlot\">@mihajlot</a> probably because the default settings of Pytorch are totally different with Keras, such as the initializers.</p>",
          "rawMarkdown": "@mihajlot probably because the default settings of Pytorch are totally different with Keras, such as the initializers.",
          "votes": 2
        },
        {
          "id": 446539,
          "postDate": "2018-12-28T08:58:30.007Z",
          "content": "<p>Do u try to modify the initializers as same as Keras ? I try to make Pytorch initializers as same as Keras default settings, but I cannot get the same result.</p>",
          "rawMarkdown": "Do u try to modify the initializers as same as Keras ? I try to make Pytorch initializers as same as Keras default settings, but I cannot get the same result."
        },
        {
          "id": 448058,
          "postDate": "2018-12-31T06:45:18.127Z",
          "content": "<p>I've tried different ways of initializing (such as xavier normal and orthogonal), they were not able to lift the LB score.</p>",
          "rawMarkdown": "I've tried different ways of initializing (such as xavier normal and orthogonal), they were not able to lift the LB score."
        }
      ]
    },
    {
      "id": 447636,
      "postDate": "2018-12-30T07:55:30.973Z",
      "content": "<p>The same structure to you, capsule layer helps me improve to 0.701 score but the threshold is too low, I doubt it is not robust.</p>",
      "rawMarkdown": "The same structure to you, capsule layer helps me improve to 0.701 score but the threshold is too low, I doubt it is not robust.",
      "votes": 2,
      "replies": [
        {
          "id": 450246,
          "postDate": "2019-01-04T14:49:37.373Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 450264,
          "postDate": "2019-01-04T15:35:08.217Z",
          "content": "<p>Capsule layer cost too much time.</p>",
          "rawMarkdown": "Capsule layer cost too much time."
        },
        {
          "id": 450465,
          "postDate": "2019-01-05T02:31:40.417Z",
          "content": "<p>Yes, so I don't have too much choice on the parameters setting. </p>",
          "rawMarkdown": "Yes, so I don't have too much choice on the parameters setting. "
        }
      ]
    },
    {
      "id": 444518,
      "postDate": "2018-12-24T07:06:25.213Z",
      "content": "<p>For me Tensorflow produce better results than Pytorch and keras. But, I am struggling with CV and LB variations. My CV score for 4 fold CV, is 0.701, but in LB, it is 0.688. My best 0.691 is using simple LSTM only.</p>",
      "rawMarkdown": "For me Tensorflow produce better results than Pytorch and keras. But, I am struggling with CV and LB variations. My CV score for 4 fold CV, is 0.701, but in LB, it is 0.688. My best 0.691 is using simple LSTM only.",
      "votes": 1
    },
    {
      "id": 448667,
      "postDate": "2019-01-01T20:02:44.480Z",
      "content": "<p>my best is 0.697 with blending lstm+attention capture cnn and gru+attention  how could you achieve over 0,7 with only one model ,did u clean the data will well?    by the way  what the CV and LB are ?</p>",
      "rawMarkdown": "my best is 0.697 with blending lstm+attention capture cnn and gru+attention  how could you achieve over 0,7 with only one model ,did u clean the data will well?    by the way  what the CV and LB are ?",
      "replies": [
        {
          "id": 448775,
          "postDate": "2019-01-02T04:24:30.977Z",
          "content": "<p>Cleaning data enhance the score and make the model more robust, my local CV was 0.68 for public LB 0.700 </p>",
          "rawMarkdown": "Cleaning data enhance the score and make the model more robust, my local CV was 0.68 for public LB 0.700 ",
          "votes": 4
        }
      ]
    },
    {
      "id": 446076,
      "postDate": "2018-12-27T13:12:33.613Z",
      "content": "<p>All designed RNN, GRU, CuDNNGRU etc. layers (at least I know) could not be used effectively to be used to form deep networks (deep enough to solve QIQC problem which has too many features) .  They do not have effective batch normalization, dropout, activation layer interfaces. Attention layers etc. can only extend the learning mechanism with miner improvements for these much features like in QIQC classification training sets... Without redesign the AM layers or using completely different topology, over-fitting is inevitable.  Best wishes...   </p>",
      "rawMarkdown": "All designed RNN, GRU, CuDNNGRU etc. layers (at least I know) could not be used effectively to be used to form deep networks (deep enough to solve QIQC problem which has too many features) .  They do not have effective batch normalization, dropout, activation layer interfaces. Attention layers etc. can only extend the learning mechanism with miner improvements for these much features like in QIQC classification training sets... Without redesign the AM layers or using completely different topology, over-fitting is inevitable.  Best wishes...   "
    },
    {
      "id": 445039,
      "postDate": "2018-12-25T12:28:54.150Z",
      "content": "<p>So what worked for you A.Barqawi?\nMy single model still reaches only  0.697</p>",
      "rawMarkdown": "So what worked for you A.Barqawi?\nMy single model still reaches only  0.697"
    },
    {
      "id": 445029,
      "postDate": "2018-12-25T12:01:05.570Z",
      "content": "<p>Any trick you apply to get 0.700 using LSTM+GRU+Attention? I reach 0.699 with LTSM+GRU architecture with hidden size modify</p>",
      "rawMarkdown": "Any trick you apply to get 0.700 using LSTM+GRU+Attention? I reach 0.699 with LTSM+GRU architecture with hidden size modify",
      "replies": [
        {
          "id": 445102,
          "postDate": "2018-12-25T16:27:36.543Z",
          "content": "<p>LSTM+GRU+Attention should be enough to reach the score, use cross validation, tune the params, check what works and enhance on top of it. Pytorch worked for my case to reach 0.700. cc: <a href=\"/mlwhiz\">@mlwhiz</a></p>",
          "rawMarkdown": "LSTM+GRU+Attention should be enough to reach the score, use cross validation, tune the params, check what works and enhance on top of it. Pytorch worked for my case to reach 0.700. cc: @mlwhiz",
          "votes": 5
        },
        {
          "id": 445232,
          "postDate": "2018-12-26T02:44:50.617Z",
          "content": "<p>thanks! your cross validation means split fixed percent data to validate , not only k-fold validation?</p>",
          "rawMarkdown": "thanks! your cross validation means split fixed percent data to validate , not only k-fold validation?"
        },
        {
          "id": 445242,
          "postDate": "2018-12-26T03:38:06.623Z",
          "content": "<p>This what is use for cross validation using sklearn: StratifiedKFold</p>",
          "rawMarkdown": "This what is use for cross validation using sklearn: StratifiedKFold"
        }
      ]
    },
    {
      "id": 447364,
      "postDate": "2018-12-29T18:04:32.023Z",
      "content": "<p>Yes. I use single Keras model, trained 5 times within a CV , and get LB f1 = 0.703. Using single model makes it easier to understand how NN architecture impacts performance.</p>\n\n<p>Edit: \nRe-run same model gave lower f1 = 0.701 on test set  :-{\nThe former 0.703 submission probably just had a 'lucky' value of the decision threshold.</p>",
      "rawMarkdown": "Yes. I use single Keras model, trained 5 times within a CV , and get LB f1 = 0.703. Using single model makes it easier to understand how NN architecture impacts performance.\n\nEdit: \nRe-run same model gave lower f1 = 0.701 on test set  :-{\nThe former 0.703 submission probably just had a 'lucky' value of the decision threshold.",
      "votes": 6,
      "isDeleted": true,
      "replies": [
        {
          "id": 447507,
          "postDate": "2018-12-30T01:18:28.210Z",
          "content": "<p>Hi, What is the total run time ( approx ) for the single model?</p>",
          "rawMarkdown": "Hi, What is the total run time ( approx ) for the single model?"
        },
        {
          "id": 447644,
          "postDate": "2018-12-30T08:14:39.910Z",
          "content": "<p>Not so short :-) It was approx. 6500 s total kernel run time.</p>",
          "rawMarkdown": "Not so short :-) It was approx. 6500 s total kernel run time.",
          "isDeleted": true
        },
        {
          "id": 447672,
          "postDate": "2018-12-30T09:56:45.520Z",
          "content": "<p>Thanks for the reply :-) </p>",
          "rawMarkdown": "Thanks for the reply :-) "
        },
        {
          "id": 450500,
          "postDate": "2019-01-05T05:22:50.693Z",
          "content": "<p>Now since you breakthrough 0.707, anymore update ? ;) <a href=\"/sezaugg\">@sezaugg</a></p>\n\n<p>EDIT: Thanks so much for your sharing below.</p>",
          "rawMarkdown": "Now since you breakthrough 0.707, anymore update ? ;) @sezaugg\n\nEDIT: Thanks so much for your sharing below.\n",
          "votes": 1
        },
        {
          "id": 450559,
          "postDate": "2019-01-05T08:59:35.233Z",
          "content": "<p>First, the LB 0.707 submission was a bit of luck, I could not reproduce.\nHowever I was able to consistently reproduce submissions with f1 = 0.704, 0.705 but not higher.\nI still use Keras, single architecture (=single model) trained multiple times. </p>",
          "rawMarkdown": "First, the LB 0.707 submission was a bit of luck, I could not reproduce.\nHowever I was able to consistently reproduce submissions with f1 = 0.704, 0.705 but not higher.\nI still use Keras, single architecture (=single model) trained multiple times. ",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 450565,
          "postDate": "2019-01-05T09:13:46.843Z",
          "content": "<p>Is local CV correlating well? Single model superb :-) . </p>",
          "rawMarkdown": "Is local CV correlating well? Single model superb :-) . "
        },
        {
          "id": 450571,
          "postDate": "2019-01-05T09:40:00.630Z",
          "content": "<p>I have the feeling with a single model trained multiple times on train/test split, it is extremely crucial to get a lucky split in train/test. For me it is very weird, my CV results are much more stable on LB than my train/test splits and I cannot get as high with them.</p>",
          "rawMarkdown": "I have the feeling with a single model trained multiple times on train/test split, it is extremely crucial to get a lucky split in train/test. For me it is very weird, my CV results are much more stable on LB than my train/test splits and I cannot get as high with them."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 443903,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2018-12-22T17:31:04.237000",
      "content": "<p>Interesting that Pytorch gives you better results, what do you think leads to that? </p>",
      "votes": 6,
      "replies": [
        {
          "id": 443940,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-22T19:02:48.450000",
          "content": "<p>Glad you found it useful. Thanks <a href=\"/jaguar00\">@jaguar00</a> (A.Barqawi)</p>\n\n<blockquote>\n  <p>what do you think leads to that? </p>\n</blockquote>\n\n<p>In my case. Pytorch trains faster and address bit of non-determism that give us a better understanding what works or not. But after a 0.700 I am experiencing lot of inconsistency with Pytorch too. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 443943,
          "author_name": "A.Barqawi",
          "author_url": "",
          "post_date": "2018-12-22T19:03:33.700000",
          "content": "<p>The same structure I tried it with Tensorflow/Keras and Pytorch however, with Pytorch the results were higher, one reason maybe its a bit faster so more cross validation folds are allowed, another thing the results were less distracted (get almost same result range with each run) which helped me to know which things really work and enhance on top of it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 443944,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-22T19:05:54.360000",
          "content": "<p>Here's a speed comparision :) </p>\n\n<p><img src=\"https://wrosinski.github.io/assets/images/framework_plots/epochs_times_comparison.png\" alt=\"\"></p>\n\n<p>Source - <a href=\"https://wrosinski.github.io/deep-learning-frameworks/\">https://wrosinski.github.io/deep-learning-frameworks/</a></p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 443945,
          "author_name": "A.Barqawi",
          "author_url": "",
          "post_date": "2018-12-22T19:06:09.083000",
          "content": "<p>HH I wrote the reason at the same time then relies you post the same thing Shahebaz. Thanks for the advice to move into Pytorch helped me a lot to enhance the model</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 443948,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-22T19:09:35.850000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443950,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-22T19:20:08.267000",
          "content": "<p>Interesting, probably I am doing something wrong with Pytorch because it is training slower :) I will stick to Keras for now most likely.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 444012,
          "author_name": "Mihajlo T.",
          "author_url": "",
          "post_date": "2018-12-22T23:15:47.223000",
          "content": "<p>In my case, it's faster than Keras but I couldn't replicate Keras score. I could be missing something also :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 444067,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-12-23T04:06:10.150000",
          "content": "<blockquote>\n  <p>another thing the results were less distracted (get almost same result range with each run) which helped me to know which things really work and enhance on top of it.</p>\n</blockquote>\n\n<p>Totally agree with this.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 444068,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-12-23T04:07:59.210000",
          "content": "<p><a href=\"/mihajlot\">@mihajlot</a> probably because the default settings of Pytorch are totally different with Keras, such as the initializers.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 446539,
          "author_name": "Salon_sai",
          "author_url": "",
          "post_date": "2018-12-28T08:58:30.007000",
          "content": "<p>Do u try to modify the initializers as same as Keras ? I try to make Pytorch initializers as same as Keras default settings, but I cannot get the same result.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 448058,
          "author_name": "Laevatein",
          "author_url": "",
          "post_date": "2018-12-31T06:45:18.127000",
          "content": "<p>I've tried different ways of initializing (such as xavier normal and orthogonal), they were not able to lift the LB score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 447636,
      "author_name": "SEU_Zesen_Chen",
      "author_url": "",
      "post_date": "2018-12-30T07:55:30.973000",
      "content": "<p>The same structure to you, capsule layer helps me improve to 0.701 score but the threshold is too low, I doubt it is not robust.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 450246,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-04T14:49:37.373000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450264,
          "author_name": "Salon_sai",
          "author_url": "",
          "post_date": "2019-01-04T15:35:08.217000",
          "content": "<p>Capsule layer cost too much time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450465,
          "author_name": "SEU_Zesen_Chen",
          "author_url": "",
          "post_date": "2019-01-05T02:31:40.417000",
          "content": "<p>Yes, so I don't have too much choice on the parameters setting. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 444518,
      "author_name": "aintnosunshine",
      "author_url": "",
      "post_date": "2018-12-24T07:06:25.213000",
      "content": "<p>For me Tensorflow produce better results than Pytorch and keras. But, I am struggling with CV and LB variations. My CV score for 4 fold CV, is 0.701, but in LB, it is 0.688. My best 0.691 is using simple LSTM only.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 448667,
      "author_name": "asvgdsf",
      "author_url": "",
      "post_date": "2019-01-01T20:02:44.480000",
      "content": "<p>my best is 0.697 with blending lstm+attention capture cnn and gru+attention  how could you achieve over 0,7 with only one model ,did u clean the data will well?    by the way  what the CV and LB are ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 448775,
          "author_name": "A.Barqawi",
          "author_url": "",
          "post_date": "2019-01-02T04:24:30.977000",
          "content": "<p>Cleaning data enhance the score and make the model more robust, my local CV was 0.68 for public LB 0.700 </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 446076,
      "author_name": "Güner ALPAYDIN",
      "author_url": "",
      "post_date": "2018-12-27T13:12:33.613000",
      "content": "<p>All designed RNN, GRU, CuDNNGRU etc. layers (at least I know) could not be used effectively to be used to form deep networks (deep enough to solve QIQC problem which has too many features) .  They do not have effective batch normalization, dropout, activation layer interfaces. Attention layers etc. can only extend the learning mechanism with miner improvements for these much features like in QIQC classification training sets... Without redesign the AM layers or using completely different topology, over-fitting is inevitable.  Best wishes...   </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445039,
      "author_name": "Rahul Agarwal",
      "author_url": "",
      "post_date": "2018-12-25T12:28:54.150000",
      "content": "<p>So what worked for you A.Barqawi?\nMy single model still reaches only  0.697</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445029,
      "author_name": " Aku dadi hamster",
      "author_url": "",
      "post_date": "2018-12-25T12:01:05.570000",
      "content": "<p>Any trick you apply to get 0.700 using LSTM+GRU+Attention? I reach 0.699 with LTSM+GRU architecture with hidden size modify</p>",
      "votes": 0,
      "replies": [
        {
          "id": 445102,
          "author_name": "A.Barqawi",
          "author_url": "",
          "post_date": "2018-12-25T16:27:36.543000",
          "content": "<p>LSTM+GRU+Attention should be enough to reach the score, use cross validation, tune the params, check what works and enhance on top of it. Pytorch worked for my case to reach 0.700. cc: <a href=\"/mlwhiz\">@mlwhiz</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 445232,
          "author_name": " Aku dadi hamster",
          "author_url": "",
          "post_date": "2018-12-26T02:44:50.617000",
          "content": "<p>thanks! your cross validation means split fixed percent data to validate , not only k-fold validation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445242,
          "author_name": "A.Barqawi",
          "author_url": "",
          "post_date": "2018-12-26T03:38:06.623000",
          "content": "<p>This what is use for cross validation using sklearn: StratifiedKFold</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 447364,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-29T18:04:32.023000",
      "content": "<p>Yes. I use single Keras model, trained 5 times within a CV , and get LB f1 = 0.703. Using single model makes it easier to understand how NN architecture impacts performance.</p>\n\n<p>Edit: \nRe-run same model gave lower f1 = 0.701 on test set  :-{\nThe former 0.703 submission probably just had a 'lucky' value of the decision threshold.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 447507,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2018-12-30T01:18:28.210000",
          "content": "<p>Hi, What is the total run time ( approx ) for the single model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447644,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-30T08:14:39.910000",
          "content": "<p>Not so short :-) It was approx. 6500 s total kernel run time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447672,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2018-12-30T09:56:45.520000",
          "content": "<p>Thanks for the reply :-) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450500,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-01-05T05:22:50.693000",
          "content": "<p>Now since you breakthrough 0.707, anymore update ? ;) <a href=\"/sezaugg\">@sezaugg</a></p>\n\n<p>EDIT: Thanks so much for your sharing below.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 450559,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-05T08:59:35.233000",
          "content": "<p>First, the LB 0.707 submission was a bit of luck, I could not reproduce.\nHowever I was able to consistently reproduce submissions with f1 = 0.704, 0.705 but not higher.\nI still use Keras, single architecture (=single model) trained multiple times. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 450565,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-01-05T09:13:46.843000",
          "content": "<p>Is local CV correlating well? Single model superb :-) . </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450571,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-01-05T09:40:00.630000",
          "content": "<p>I have the feeling with a single model trained multiple times on train/test split, it is extremely crucial to get a lucky split in train/test. For me it is very weird, my CV results are much more stable on LB than my train/test splits and I cannot get as high with them.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "443825": "My best single model reached 0.700 using Pytoch with LSTM+GRU+Attention which reached top 35 models up to this moment.<br>\nI used @Shahebaz advice to move into Pytoch which helped me to get better score comparing to Tensorflow/Keras models.<br>\nCan I continue with a single model to reach 0.702+ or should I start shifting the gears to ensembling?",
    "443903": "Interesting that Pytorch gives you better results, what do you think leads to that? ",
    "447636": "The same structure to you, capsule layer helps me improve to 0.701 score but the threshold is too low, I doubt it is not robust.",
    "444518": "For me Tensorflow produce better results than Pytorch and keras. But, I am struggling with CV and LB variations. My CV score for 4 fold CV, is 0.701, but in LB, it is 0.688. My best 0.691 is using simple LSTM only.",
    "448667": "my best is 0.697 with blending lstm+attention capture cnn and gru+attention  how could you achieve over 0,7 with only one model ,did u clean the data will well?    by the way  what the CV and LB are ?",
    "446076": "All designed RNN, GRU, CuDNNGRU etc. layers (at least I know) could not be used effectively to be used to form deep networks (deep enough to solve QIQC problem which has too many features) .  They do not have effective batch normalization, dropout, activation layer interfaces. Attention layers etc. can only extend the learning mechanism with miner improvements for these much features like in QIQC classification training sets... Without redesign the AM layers or using completely different topology, over-fitting is inevitable.  Best wishes...   ",
    "445039": "So what worked for you A.Barqawi?\nMy single model still reaches only  0.697",
    "445029": "Any trick you apply to get 0.700 using LSTM+GRU+Attention? I reach 0.699 with LTSM+GRU architecture with hidden size modify",
    "447364": "Yes. I use single Keras model, trained 5 times within a CV , and get LB f1 = 0.703. Using single model makes it easier to understand how NN architecture impacts performance.\n\nEdit: \nRe-run same model gave lower f1 = 0.701 on test set  :-{\nThe former 0.703 submission probably just had a 'lucky' value of the decision threshold."
  }
}