{
  "id": 80696,
  "title": "18th place solution from 300-th at Public LB",
  "url": "/competitions/quora-insincere-questions-classification/writeups/toguro-18th-place-solution-from-300-th-at-public-l",
  "author_name": "",
  "post_date": "2019-02-15T14:41:29.322032200Z",
  "votes": 19,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thanks Quora for great competition and thanks all who participated and contributed here.\nWe were surprised at this result since we were in bronze in Public LB.\nOur (me and <a href=\"https://www.kaggle.com/hattan0523\">@hattan0523</a> ) kernel is here.\n<a href=\"https://www.kaggle.com/kentaronakanishi/18th-place-solution\">https://www.kaggle.com/kentaronakanishi/18th-place-solution</a>\nThe local CV of our model is around 0.696-0.698.</p>\n\n<p>Our points are below:</p>\n\n<ul>\n<li>simple 2 layers RNN model with units=96 by 5 fold</li>\n<li>use semantic bernoulli dropout for embedding layer</li>\n<li>remove any filter when using keras tokenizer</li>\n<li>batch size control to get larger batch size in later epochs</li>\n<li>cut data length at max length in a batch</li>\n<li>learn embedding weights only at last epoch</li>\n<li>lots of lucks by using my wedding anniversary as seed</li>\n</ul>\n\n<p>We don't use in our case:</p>\n\n<ul>\n<li>capsules: contribute but take too much time</li>\n<li>attention: only effect val_loss and no change in CV score</li>\n<li>CNN based model: It worked but increasing rnn width is better for us</li>\n<li>GBDT models (XGboost, LightBGM): They don’t achieve our desired CV score, and a lot of time are needed to run. Thus, we don’t use an ensemble or stacking method. </li>\n<li>word typos: a kernel (<a href=\"https://www.kaggle.com/sunnymarkliu/more-text-cleaning-to-increase-word-coverage\">https://www.kaggle.com/sunnymarkliu/more-text-cleaning-to-increase-word-coverage</a>) shows many typos of the data set at In[18] of kernel. We rechecked and modified them. This is good for public LB, however not good for private LB.</li>\n</ul>\n\n<p>All questions and suggestions are welcome.</p>",
  "messages": [
    {
      "id": "472224",
      "postDate": "02/15/2019 14:41:29",
      "content": "<p>Thanks Quora for great competition and thanks all who participated and contributed here.\nWe were surprised at this result since we were in bronze in Public LB.\nOur (me and <a href=\"https://www.kaggle.com/hattan0523\">@hattan0523</a> ) kernel is here.\n<a href=\"https://www.kaggle.com/kentaronakanishi/18th-place-solution\">https://www.kaggle.com/kentaronakanishi/18th-place-solution</a>\nThe local CV of our model is around 0.696-0.698.</p>\n\n<p>Our points are below:</p>\n\n<ul>\n<li>simple 2 layers RNN model with units=96 by 5 fold</li>\n<li>use semantic bernoulli dropout for embedding layer</li>\n<li>remove any filter when using keras tokenizer</li>\n<li>batch size control to get larger batch size in later epochs</li>\n<li>cut data length at max length in a batch</li>\n<li>learn embedding weights only at last epoch</li>\n<li>lots of lucks by using my wedding anniversary as seed</li>\n</ul>\n\n<p>We don't use in our case:</p>\n\n<ul>\n<li>capsules: contribute but take too much time</li>\n<li>attention: only effect val_loss and no change in CV score</li>\n<li>CNN based model: It worked but increasing rnn width is better for us</li>\n<li>GBDT models (XGboost, LightBGM): They don’t achieve our desired CV score, and a lot of time are needed to run. Thus, we don’t use an ensemble or stacking method. </li>\n<li>word typos: a kernel (<a href=\"https://www.kaggle.com/sunnymarkliu/more-text-cleaning-to-increase-word-coverage\">https://www.kaggle.com/sunnymarkliu/more-text-cleaning-to-increase-word-coverage</a>) shows many typos of the data set at In[18] of kernel. We rechecked and modified them. This is good for public LB, however not good for private LB.</li>\n</ul>\n\n<p>All questions and suggestions are welcome.</p>",
      "rawMarkdown": "Thanks Quora for great competition and thanks all who participated and contributed here.\nWe were surprised at this result since we were in bronze in Public LB.\nOur (me and [@hattan0523][1] ) kernel is here.\nhttps://www.kaggle.com/kentaronakanishi/18th-place-solution\nThe local CV of our model is around 0.696-0.698.\n\nOur points are below:\n\n- simple 2 layers RNN model with units=96 by 5 fold\n- use semantic bernoulli dropout for embedding layer\n- remove any filter when using keras tokenizer\n- batch size control to get larger batch size in later epochs\n- cut data length at max length in a batch\n- learn embedding weights only at last epoch\n- lots of lucks by using my wedding anniversary as seed\n\nWe don't use in our case:\n\n- capsules: contribute but take too much time\n- attention: only effect val_loss and no change in CV score\n- CNN based model: It worked but increasing rnn width is better for us\n- GBDT models (XGboost, LightBGM): They don’t achieve our desired CV score, and a lot of time are needed to run. Thus, we don’t use an ensemble or stacking method. \n- word typos: a kernel (https://www.kaggle.com/sunnymarkliu/more-text-cleaning-to-increase-word-coverage) shows many typos of the data set at In[18] of kernel. We rechecked and modified them. This is good for public LB, however not good for private LB.\n\nAll questions and suggestions are welcome.\n\n\n  [1]: https://www.kaggle.com/hattan0523",
      "votes": null
    },
    {
      "id": "472308",
      "postDate": "02/15/2019 17:01:11",
      "content": "<p>Congratulations @cfiken. Keep up the good work going.</p>",
      "rawMarkdown": "Congratulations @cfiken. Keep up the good work going.",
      "votes": null
    },
    {
      "id": "472323",
      "postDate": "02/15/2019 17:36:22",
      "content": "<p>That's great !!!</p>",
      "rawMarkdown": "That's great !!!",
      "votes": null
    },
    {
      "id": "472626",
      "postDate": "02/16/2019 10:29:15",
      "content": "<p>Congratulations! I know how to get magic random seed, I need to get married first haha. :)</p>",
      "rawMarkdown": "Congratulations! I know how to get magic random seed, I need to get married first haha. :)",
      "votes": null
    },
    {
      "id": "472678",
      "postDate": "02/16/2019 13:08:15",
      "content": "<p>If the typo correction was good for public LB, how did you make the decision not to use it? Did it also pull the CV down?</p>",
      "rawMarkdown": "If the typo correction was good for public LB, how did you make the decision not to use it? Did it also pull the CV down?",
      "votes": null
    },
    {
      "id": "472727",
      "postDate": "02/16/2019 14:54:29",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "472728",
      "postDate": "02/16/2019 14:54:37",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "472730",
      "postDate": "02/16/2019 14:56:03",
      "content": "<p>haha thanks a lot and have a nice partner ;)</p>",
      "rawMarkdown": "haha thanks a lot and have a nice partner ;)",
      "votes": null
    },
    {
      "id": "472736",
      "postDate": "02/16/2019 15:01:15",
      "content": "<p>We just select two submissions, one use hard typo correcting and the another doesn't.\nAt first this is just for a reason of time limiting (fixing typo takes much time) but correcting typo results in getting lower score.</p>",
      "rawMarkdown": "We just select two submissions, one use hard typo correcting and the another doesn't.\nAt first this is just for a reason of time limiting (fixing typo takes much time) but correcting typo results in getting lower score.",
      "votes": null
    },
    {
      "id": "472950",
      "postDate": "02/17/2019 00:46:09",
      "content": "<p>This is great result, thanks for sharing !! </p>",
      "rawMarkdown": "This is great result, thanks for sharing !!",
      "votes": null
    },
    {
      "id": "473421",
      "postDate": "02/18/2019 01:14:51",
      "content": "<p>Thanks!!</p>",
      "rawMarkdown": "Thanks!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 472308,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/15/2019 17:01:11",
      "content": "<p>Congratulations @cfiken. Keep up the good work going.</p>",
      "votes": null,
      "replies": [
        {
          "id": 472727,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "02/16/2019 14:54:29",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472323,
      "author_name": "chanchalkumarmaji",
      "author_url": "",
      "post_date": "02/15/2019 17:36:22",
      "content": "<p>That's great !!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 472728,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "02/16/2019 14:54:37",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472626,
      "author_name": "sunnymarkliu",
      "author_url": "",
      "post_date": "02/16/2019 10:29:15",
      "content": "<p>Congratulations! I know how to get magic random seed, I need to get married first haha. :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 472730,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "02/16/2019 14:56:03",
          "content": "<p>haha thanks a lot and have a nice partner ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472678,
      "author_name": "mlisovyi",
      "author_url": "",
      "post_date": "02/16/2019 13:08:15",
      "content": "<p>If the typo correction was good for public LB, how did you make the decision not to use it? Did it also pull the CV down?</p>",
      "votes": null,
      "replies": [
        {
          "id": 472736,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "02/16/2019 15:01:15",
          "content": "<p>We just select two submissions, one use hard typo correcting and the another doesn't.\nAt first this is just for a reason of time limiting (fixing typo takes much time) but correcting typo results in getting lower score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472950,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "02/17/2019 00:46:09",
      "content": "<p>This is great result, thanks for sharing !! </p>",
      "votes": null,
      "replies": [
        {
          "id": 473421,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "02/18/2019 01:14:51",
          "content": "<p>Thanks!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "472224": "Thanks Quora for great competition and thanks all who participated and contributed here.\nWe were surprised at this result since we were in bronze in Public LB.\nOur (me and [@hattan0523][1] ) kernel is here.\nhttps://www.kaggle.com/kentaronakanishi/18th-place-solution\nThe local CV of our model is around 0.696-0.698.\n\nOur points are below:\n\n- simple 2 layers RNN model with units=96 by 5 fold\n- use semantic bernoulli dropout for embedding layer\n- remove any filter when using keras tokenizer\n- batch size control to get larger batch size in later epochs\n- cut data length at max length in a batch\n- learn embedding weights only at last epoch\n- lots of lucks by using my wedding anniversary as seed\n\nWe don't use in our case:\n\n- capsules: contribute but take too much time\n- attention: only effect val_loss and no change in CV score\n- CNN based model: It worked but increasing rnn width is better for us\n- GBDT models (XGboost, LightBGM): They don’t achieve our desired CV score, and a lot of time are needed to run. Thus, we don’t use an ensemble or stacking method. \n- word typos: a kernel (https://www.kaggle.com/sunnymarkliu/more-text-cleaning-to-increase-word-coverage) shows many typos of the data set at In[18] of kernel. We rechecked and modified them. This is good for public LB, however not good for private LB.\n\nAll questions and suggestions are welcome.\n\n\n  [1]: https://www.kaggle.com/hattan0523",
    "472308": "Congratulations @cfiken. Keep up the good work going.",
    "472323": "That's great !!!",
    "472626": "Congratulations! I know how to get magic random seed, I need to get married first haha. :)",
    "472678": "If the typo correction was good for public LB, how did you make the decision not to use it? Did it also pull the CV down?",
    "472727": "Thanks!",
    "472728": "Thanks!",
    "472730": "haha thanks a lot and have a nice partner ;)",
    "472736": "We just select two submissions, one use hard typo correcting and the another doesn't.\nAt first this is just for a reason of time limiting (fixing typo takes much time) but correcting typo results in getting lower score.",
    "472950": "This is great result, thanks for sharing !!",
    "473421": "Thanks!!"
  },
  "source": "meta"
}