{
  "id": 80509,
  "title": "117th solution. Achieve 0.701 PB and 0.703 LB in 4000 seconds.",
  "url": "/competitions/quora-insincere-questions-classification/writeups/together-to-the-better-117th-solution-achieve-0-70",
  "author_name": "",
  "post_date": "2019-02-14T06:43:11.617Z",
  "votes": 12,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Here is my solution. Simple one directional GRU, with some effective preprocessing for OOV words. No spell correction or anything. I have used 5 fold cross validation with exponential moving average ( which is not necessary ). Code is bit clumsy and keras lovers, don't find it happy as it is in pure tensorflow :-) . If any doubts, please do comment. Hope it will be useful.</p>\n\n<p><a href=\"https://www.kaggle.com/s4sarath/cudnngru-best-final\">https://www.kaggle.com/s4sarath/cudnngru-best-final</a></p>\n\n<p>1.) I have <code>preprocessed by word basis</code>. Normally, I will split by space and check if a word is present in the vocab. If present, I will add it to my train embedding else I will preprocess only that word and cache it to a dictionary. For eg: <code>goood</code> might not be in glove or fasttext, then I will preprocess (a lot of sub and small functions are there) will give me <code>good</code> which is in glove. So, I will add it to mapping_dict = {<code>goood</code> : <code>good</code>}. Then, next time <code>goood</code> is appearing, I dont have to preprocess it, I will first look at <code>mapping_dict</code> and so on. My final vocab size was ```192000</p>\n\n<p>2.) I have used <code>GRU</code> one side, with <code>256</code> dimensional embeddings. <code>Cross validation</code> of <code>5</code> fold, where each folds have different epochs, I used <code>[5,5,5,4,4]</code>. Used <code>Adam Optimizer</code> for minimzing the objective.</p>\n\n<p>3.) I have used <code>meta embeddings</code> <code>Glove+Paragram</code> as per <code>@shijuan's</code> Kernel</p>\n\n<p>4.) Used tensorflow,<code>CudnnGRU</code>, it is very fast and max len is <code>150 for a sentence.</code></p>\n\n<p>Pretty much that's it. </p>",
  "messages": [
    {
      "id": "471169",
      "postDate": "02/14/2019 05:29:56",
      "content": "<p>Here is my solution. Simple one directional GRU, with some effective preprocessing for OOV words. No spell correction or anything. I have used 5 fold cross validation with exponential moving average ( which is not necessary ). Code is bit clumsy and keras lovers, don't find it happy as it is in pure tensorflow :-) . If any doubts, please do comment. Hope it will be useful.</p>\n\n<p><a href=\"https://www.kaggle.com/s4sarath/cudnngru-best-final\">https://www.kaggle.com/s4sarath/cudnngru-best-final</a></p>\n\n<p>1.) I have <code>preprocessed by word basis</code>. Normally, I will split by space and check if a word is present in the vocab. If present, I will add it to my train embedding else I will preprocess only that word and cache it to a dictionary. For eg: <code>goood</code> might not be in glove or fasttext, then I will preprocess (a lot of sub and small functions are there) will give me <code>good</code> which is in glove. So, I will add it to mapping_dict = {<code>goood</code> : <code>good</code>}. Then, next time <code>goood</code> is appearing, I dont have to preprocess it, I will first look at <code>mapping_dict</code> and so on. My final vocab size was ```192000</p>\n\n<p>2.) I have used <code>GRU</code> one side, with <code>256</code> dimensional embeddings. <code>Cross validation</code> of <code>5</code> fold, where each folds have different epochs, I used <code>[5,5,5,4,4]</code>. Used <code>Adam Optimizer</code> for minimzing the objective.</p>\n\n<p>3.) I have used <code>meta embeddings</code> <code>Glove+Paragram</code> as per <code>@shijuan's</code> Kernel</p>\n\n<p>4.) Used tensorflow,<code>CudnnGRU</code>, it is very fast and max len is <code>150 for a sentence.</code></p>\n\n<p>Pretty much that's it. </p>",
      "rawMarkdown": "Here is my solution. Simple one directional GRU, with some effective preprocessing for OOV words. No spell correction or anything. I have used 5 fold cross validation with exponential moving average ( which is not necessary ). Code is bit clumsy and keras lovers, don't find it happy as it is in pure tensorflow :-) . If any doubts, please do comment. Hope it will be useful.\n\nhttps://www.kaggle.com/s4sarath/cudnngru-best-final\n\n1.) I have ```preprocessed by word basis```. Normally, I will split by space and check if a word is present in the vocab. If present, I will add it to my train embedding else I will preprocess only that word and cache it to a dictionary. For eg: ```goood``` might not be in glove or fasttext, then I will preprocess (a lot of sub and small functions are there) will give me ```good``` which is in glove. So, I will add it to mapping_dict = {```goood``` : ```good```}. Then, next time ```goood``` is appearing, I dont have to preprocess it, I will first look at ```mapping_dict``` and so on. My final vocab size was ```192000\n\n2.) I have used ```GRU``` one side, with ```256``` dimensional embeddings. ```Cross validation``` of ```5``` fold, where each folds have different epochs, I used ```[5,5,5,4,4]```. Used ```Adam Optimizer``` for minimzing the objective.\n\n3.) I have used ```meta embeddings``` ```Glove+Paragram``` as per ```@shijuan's``` Kernel\n\n4.) Used tensorflow,``` CudnnGRU```, it is very fast and max len is ```150 for a sentence. ```\n\nPretty much that's it.",
      "votes": null
    },
    {
      "id": "471183",
      "postDate": "02/14/2019 06:01:20",
      "content": "<p>Why it should be downvoted?</p>",
      "rawMarkdown": "Why it should be downvoted?",
      "votes": null
    },
    {
      "id": "471193",
      "postDate": "02/14/2019 06:14:18",
      "content": "<p>??</p>",
      "rawMarkdown": "??",
      "votes": null
    },
    {
      "id": "471200",
      "postDate": "02/14/2019 06:20:33",
      "content": "<p>I think people downvoted from frustration after shakeup. Do not strain about this. You did a great job. Thank you for sharing.</p>\n\n<p>But formally, of course, you could write more about how you came to your result:)</p>",
      "rawMarkdown": "I think people downvoted from frustration after shakeup. Do not strain about this. You did a great job. Thank you for sharing.\n\nBut formally, of course, you could write more about how you came to your result:)",
      "votes": null
    },
    {
      "id": "471477",
      "postDate": "02/14/2019 14:15:00",
      "content": "<p>Hey Sarath, I just know the result!\nCongratulation for your great performance here!</p>",
      "rawMarkdown": "Hey Sarath, I just know the result!\nCongratulation for your great performance here!",
      "votes": null
    },
    {
      "id": "471856",
      "postDate": "02/15/2019 01:44:33",
      "content": "<p>Thanks @Neuron Engineer. Glad that you also did well. I had a great time learning new things.</p>",
      "rawMarkdown": "Thanks @Neuron Engineer. Glad that you also did well. I had a great time learning new things.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471183,
      "author_name": "s4sarath",
      "author_url": "",
      "post_date": "02/14/2019 06:01:20",
      "content": "<p>Why it should be downvoted?</p>",
      "votes": null,
      "replies": [
        {
          "id": 471200,
          "author_name": "rooshroosh",
          "author_url": "",
          "post_date": "02/14/2019 06:20:33",
          "content": "<p>I think people downvoted from frustration after shakeup. Do not strain about this. You did a great job. Thank you for sharing.</p>\n\n<p>But formally, of course, you could write more about how you came to your result:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471193,
      "author_name": "s4sarath",
      "author_url": "",
      "post_date": "02/14/2019 06:14:18",
      "content": "<p>??</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471477,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "02/14/2019 14:15:00",
      "content": "<p>Hey Sarath, I just know the result!\nCongratulation for your great performance here!</p>",
      "votes": null,
      "replies": [
        {
          "id": 471856,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "02/15/2019 01:44:33",
          "content": "<p>Thanks @Neuron Engineer. Glad that you also did well. I had a great time learning new things.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "471169": "Here is my solution. Simple one directional GRU, with some effective preprocessing for OOV words. No spell correction or anything. I have used 5 fold cross validation with exponential moving average ( which is not necessary ). Code is bit clumsy and keras lovers, don't find it happy as it is in pure tensorflow :-) . If any doubts, please do comment. Hope it will be useful.\n\nhttps://www.kaggle.com/s4sarath/cudnngru-best-final\n\n1.) I have ```preprocessed by word basis```. Normally, I will split by space and check if a word is present in the vocab. If present, I will add it to my train embedding else I will preprocess only that word and cache it to a dictionary. For eg: ```goood``` might not be in glove or fasttext, then I will preprocess (a lot of sub and small functions are there) will give me ```good``` which is in glove. So, I will add it to mapping_dict = {```goood``` : ```good```}. Then, next time ```goood``` is appearing, I dont have to preprocess it, I will first look at ```mapping_dict``` and so on. My final vocab size was ```192000\n\n2.) I have used ```GRU``` one side, with ```256``` dimensional embeddings. ```Cross validation``` of ```5``` fold, where each folds have different epochs, I used ```[5,5,5,4,4]```. Used ```Adam Optimizer``` for minimzing the objective.\n\n3.) I have used ```meta embeddings``` ```Glove+Paragram``` as per ```@shijuan's``` Kernel\n\n4.) Used tensorflow,``` CudnnGRU```, it is very fast and max len is ```150 for a sentence. ```\n\nPretty much that's it.",
    "471183": "Why it should be downvoted?",
    "471193": "??",
    "471200": "I think people downvoted from frustration after shakeup. Do not strain about this. You did a great job. Thank you for sharing.\n\nBut formally, of course, you could write more about how you came to your result:)",
    "471477": "Hey Sarath, I just know the result!\nCongratulation for your great performance here!",
    "471856": "Thanks @Neuron Engineer. Glad that you also did well. I had a great time learning new things."
  },
  "source": "meta"
}