{
  "id": 80544,
  "title": "From 400-ish public to 26 private",
  "url": "/competitions/quora-insincere-questions-classification/writeups/theo-viel-from-400-ish-public-to-26-private",
  "author_name": "",
  "post_date": "2019-02-14T10:18:18.110760400Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello, and congratz to everybody who made it 'till the end! Special thanks to people who shared stuff during the challenge, I learnt a lot.</p>\n\n<p>I'll give you a brief overview of my model that made it top the 26th:</p>\n\n<ul>\n<li><p>Preprocessing : Some special characters cleaning, number processing, contractions &amp; mispells replacement and latex tags cleaning. No lowering though.</p></li>\n<li><p>Embeddings : Concatenation of glove, fasttext and paragram.</p></li>\n<li><p>Some features : Toxic words ratio, Total length, word vs unique words, ratio of capital letters.</p></li>\n</ul>\n\n<p><strong>Model:</strong> </p>\n\n<ul>\n<li>I used PyTorch</li>\n<li>Single model , 5 folds, 4 epochs :\n<ul><li>Embedding layer + some noise</li>\n<li>LSTM, 64 Units (unidirectional)</li>\n<li>GRU, 32 Units (unidirectional)</li>\n<li>Attention, maxpool &amp;  average pool on the outputs of both rnns</li>\n<li>Concatenating them with features</li>\n<li>32 units dense + reLu + Batchnorm + Dropout</li>\n<li>And the final layer</li></ul></li>\n</ul>\n\n<p>CV : 0.688, Public LB : 0.700 </p>\n\n<p>This model was not my best one on the LB, but it had a good CV and an average LB which made me trust it more than the others.</p>\n\n<p>Thanks for reading, feel free to ask me any question! I'll probably make my code public, but it needs some cleaning first. </p>",
  "messages": [
    {
      "id": "471340",
      "postDate": "02/14/2019 10:18:18",
      "content": "<p>Hello, and congratz to everybody who made it 'till the end! Special thanks to people who shared stuff during the challenge, I learnt a lot.</p>\n\n<p>I'll give you a brief overview of my model that made it top the 26th:</p>\n\n<ul>\n<li><p>Preprocessing : Some special characters cleaning, number processing, contractions &amp; mispells replacement and latex tags cleaning. No lowering though.</p></li>\n<li><p>Embeddings : Concatenation of glove, fasttext and paragram.</p></li>\n<li><p>Some features : Toxic words ratio, Total length, word vs unique words, ratio of capital letters.</p></li>\n</ul>\n\n<p><strong>Model:</strong> </p>\n\n<ul>\n<li>I used PyTorch</li>\n<li>Single model , 5 folds, 4 epochs :\n<ul><li>Embedding layer + some noise</li>\n<li>LSTM, 64 Units (unidirectional)</li>\n<li>GRU, 32 Units (unidirectional)</li>\n<li>Attention, maxpool &amp;  average pool on the outputs of both rnns</li>\n<li>Concatenating them with features</li>\n<li>32 units dense + reLu + Batchnorm + Dropout</li>\n<li>And the final layer</li></ul></li>\n</ul>\n\n<p>CV : 0.688, Public LB : 0.700 </p>\n\n<p>This model was not my best one on the LB, but it had a good CV and an average LB which made me trust it more than the others.</p>\n\n<p>Thanks for reading, feel free to ask me any question! I'll probably make my code public, but it needs some cleaning first. </p>",
      "rawMarkdown": "Hello, and congratz to everybody who made it 'till the end! Special thanks to people who shared stuff during the challenge, I learnt a lot.\n\nI'll give you a brief overview of my model that made it top the 26th:\n\n- Preprocessing : Some special characters cleaning, number processing, contractions &amp; mispells replacement and latex tags cleaning. No lowering though.\n\n- Embeddings : Concatenation of glove, fasttext and paragram.\n\n- Some features : Toxic words ratio, Total length, word vs unique words, ratio of capital letters.\n\n**Model:** \n\n- I used PyTorch\n- Single model , 5 folds, 4 epochs :\n - Embedding layer + some noise\n - LSTM, 64 Units (unidirectional)\n - GRU, 32 Units (unidirectional)\n - Attention, maxpool &amp;  average pool on the outputs of both rnns\n - Concatenating them with features\n - 32 units dense + reLu + Batchnorm + Dropout\n - And the final layer\n\nCV : 0.688, Public LB : 0.700 \n\nThis model was not my best one on the LB, but it had a good CV and an average LB which made me trust it more than the others.\n\nThanks for reading, feel free to ask me any question! I'll probably make my code public, but it needs some cleaning first.",
      "votes": null
    },
    {
      "id": "471344",
      "postDate": "02/14/2019 10:29:11",
      "content": "<p>Congrats @theo, Keep up the good work going. Is unidirectional giving better results than bidirectional. My structure is the same as yours, I have 2 gru's and have done some post processing on final results to only improve <strong>recall</strong> but I think the distribution of private test data is same as public test data. Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats @theo, Keep up the good work going. Is unidirectional giving better results than bidirectional. My structure is the same as yours, I have 2 gru's and have done some post processing on final results to only improve **recall** but I think the distribution of private test data is same as public test data. Thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "471345",
      "postDate": "02/14/2019 10:29:19",
      "content": "<p>Wow congratulations!</p>\n\n<p>Did you expect to jump so many places? Sounds like you had a feeling you'd jump a bit given your CV was good?</p>",
      "rawMarkdown": "Wow congratulations!\n\nDid you expect to jump so many places? Sounds like you had a feeling you'd jump a bit given your CV was good?",
      "votes": null
    },
    {
      "id": "471380",
      "postDate": "02/14/2019 11:19:43",
      "content": "<p>Performances are roughly the same, but Attention works way better on an unidirectional layer.</p>",
      "rawMarkdown": "Performances are roughly the same, but Attention works way better on an unidirectional layer.",
      "votes": null
    },
    {
      "id": "471381",
      "postDate": "02/14/2019 11:21:13",
      "content": "<p>Thanks !\nI expected/hoped to jump over people that forked public kernels and did not put much effort in the competition, but surely did not expect to end up 26th !</p>",
      "rawMarkdown": "Thanks !\nI expected/hoped to jump over people that forked public kernels and did not put much effort in the competition, but surely did not expect to end up 26th !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471344,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/14/2019 10:29:11",
      "content": "<p>Congrats @theo, Keep up the good work going. Is unidirectional giving better results than bidirectional. My structure is the same as yours, I have 2 gru's and have done some post processing on final results to only improve <strong>recall</strong> but I think the distribution of private test data is same as public test data. Thanks for sharing your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 471380,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/14/2019 11:19:43",
          "content": "<p>Performances are roughly the same, but Attention works way better on an unidirectional layer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471345,
      "author_name": "hamishdickson",
      "author_url": "",
      "post_date": "02/14/2019 10:29:19",
      "content": "<p>Wow congratulations!</p>\n\n<p>Did you expect to jump so many places? Sounds like you had a feeling you'd jump a bit given your CV was good?</p>",
      "votes": null,
      "replies": [
        {
          "id": 471381,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/14/2019 11:21:13",
          "content": "<p>Thanks !\nI expected/hoped to jump over people that forked public kernels and did not put much effort in the competition, but surely did not expect to end up 26th !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "471340": "Hello, and congratz to everybody who made it 'till the end! Special thanks to people who shared stuff during the challenge, I learnt a lot.\n\nI'll give you a brief overview of my model that made it top the 26th:\n\n- Preprocessing : Some special characters cleaning, number processing, contractions &amp; mispells replacement and latex tags cleaning. No lowering though.\n\n- Embeddings : Concatenation of glove, fasttext and paragram.\n\n- Some features : Toxic words ratio, Total length, word vs unique words, ratio of capital letters.\n\n**Model:** \n\n- I used PyTorch\n- Single model , 5 folds, 4 epochs :\n - Embedding layer + some noise\n - LSTM, 64 Units (unidirectional)\n - GRU, 32 Units (unidirectional)\n - Attention, maxpool &amp;  average pool on the outputs of both rnns\n - Concatenating them with features\n - 32 units dense + reLu + Batchnorm + Dropout\n - And the final layer\n\nCV : 0.688, Public LB : 0.700 \n\nThis model was not my best one on the LB, but it had a good CV and an average LB which made me trust it more than the others.\n\nThanks for reading, feel free to ask me any question! I'll probably make my code public, but it needs some cleaning first.",
    "471344": "Congrats @theo, Keep up the good work going. Is unidirectional giving better results than bidirectional. My structure is the same as yours, I have 2 gru's and have done some post processing on final results to only improve **recall** but I think the distribution of private test data is same as public test data. Thanks for sharing your solution.",
    "471345": "Wow congratulations!\n\nDid you expect to jump so many places? Sounds like you had a feeling you'd jump a bit given your CV was good?",
    "471380": "Performances are roughly the same, but Attention works way better on an unidirectional layer.",
    "471381": "Thanks !\nI expected/hoped to jump over people that forked public kernels and did not put much effort in the competition, but surely did not expect to end up 26th !"
  },
  "source": "meta"
}