{
  "id": 80665,
  "title": "1st place solution (public LB)",
  "url": "/competitions/quora-insincere-questions-classification/writeups/ods-ai-toulouse-goose-1st-place-solution-public-lb",
  "author_name": "",
  "post_date": "2019-02-15T11:07:11.150Z",
  "votes": 97,
  "comment_count": 28,
  "views": 0,
  "content": "<h2>Motivation</h2>\n\n<p>The purpose of this competition wasn't really clear to me. So many SOTA NLP models were introduced in 2018 and none of them were allowed during the competition. It sounds like a joke to approach a complicated NLP task when even ULMFiT (kudos to <a href=\"/jhoward\">@jhoward</a>) is prohibited. So I decided to have some fun together with <a href=\"/evgeny000\">@evgeny000</a> and <a href=\"/mathurinache\">@mathurinache</a>.</p>\n\n<p><a href=\"https://i.imgur.com/IH4fZyx.gif?1\"><img src=\"https://i.imgur.com/IH4fZyx.gif?1\"></a></p>\n\n<h2>Getting true labels</h2>\n\n<p>Many of you were wondering about 0.782 score. The solution is quite straightforward: we just <strong>scraped the answers</strong>. The process was the following:</p>\n\n<ol>\n<li>It's easy to notice that Quora links are very much similar to the actual question asked. By applying several heuristics you can back reverse the link from the question. For example, <a href=\"https://www.quora.com/Why-did-you-quit-your-job-at-Amazon\">https://www.quora.com/Why-did-you-quit-your-job-at-Amazon</a></li>\n<li>Insincere questions can be then detected by the tag <em>QuestionRestrictedInsincerePrompt</em> in the HTML code which can be obtained by python requests library.  However, there is even more elegant solution to that.</li>\n<li>At some point, we've noticed that if you add <em>/log</em> to any Quora question link than you get the full history of page changes including topics, comment, users, etc. Try it <a href=\"https://www.quora.com/unanswered/Are-Muslims-really-ashamed-of-their-religion/log\">here</a>. As you may see there is a log line \"Question marked as possibly insincere by Quora Content Review\". That is exactly what we were looking for.  </li>\n</ol>\n\n<p><a href=\"https://imgflip.com/i/2tqg7u\"><img src=\"https://i.imgflip.com/2tqg7u.jpg\"></a></p>\n\n<h2>Second stage labels</h2>\n\n<p>Of course, that's not the end of the story. Scraping 1 stage answers would be useless without getting 2 stage labels. How can we guess them? First, we tried using data from the <a href=\"https://www.kaggle.com/c/quora-question-pairs\">previous Quora competition</a>. We collected around 2000 insincere questions from there. The second option was using Related Questions but none of them were insincere. </p>\n\n<p>Our last resort was <strong>toxic users</strong>, i.e. non-anonymous users with many insincere questions. For example, <a href=\"https://www.quora.com/profile/Jim-Lunde/questions\">this guy</a>. We collected a list of such users based on train/test datasets and then scraped all of their questions. Some of these questions were not in 1 stage datasets which seemed to be very promising - they should have been in the 2 stage data. Unfortunately, as you can see by our private score we were wrong (for the best, probably).</p>\n\n<p><a href=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"><img src=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"></a></p>\n\n<h2>Small technical challenges</h2>\n\n<ul>\n<li>Notebooks and scripts on Kaggle are limited by the size of 1Mb. How could we squeeze several thousand questions into the code? We converted an array of questions into txt file, then compressed it with tar.gz, then converted the archive to base64 code and inserted it into the script. 10K questions were stored using only 500Kb of space. </li>\n<li>We used <a href=\"https://www.ip-adress.com/proxy-list\">list of 50 proxies</a> and 100 threads to parallel our scraping process. At some point, my own IP address was completely banned from requests to Quora. The access was restored after a couple of days, though.</li>\n<li>Part of code that shows user questions was written in JS so some advanced knowledge of Selenium framework was required.</li>\n</ul>\n\n<p><a href=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"><img src=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"></a></p>\n\n<h2>Takeaways</h2>\n\n<ul>\n<li>There is always a room for a creative approach to a Kaggle competition</li>\n<li>You never fail if you learn something new</li>\n<li>Whatever you do - make it fun</li>\n</ul>",
  "messages": [
    {
      "id": "472042",
      "postDate": "02/15/2019 09:06:48",
      "content": "<h2>Motivation</h2>\n\n<p>The purpose of this competition wasn't really clear to me. So many SOTA NLP models were introduced in 2018 and none of them were allowed during the competition. It sounds like a joke to approach a complicated NLP task when even ULMFiT (kudos to <a href=\"/jhoward\">@jhoward</a>) is prohibited. So I decided to have some fun together with <a href=\"/evgeny000\">@evgeny000</a> and <a href=\"/mathurinache\">@mathurinache</a>.</p>\n\n<p><a href=\"https://i.imgur.com/IH4fZyx.gif?1\"><img src=\"https://i.imgur.com/IH4fZyx.gif?1\"></a></p>\n\n<h2>Getting true labels</h2>\n\n<p>Many of you were wondering about 0.782 score. The solution is quite straightforward: we just <strong>scraped the answers</strong>. The process was the following:</p>\n\n<ol>\n<li>It's easy to notice that Quora links are very much similar to the actual question asked. By applying several heuristics you can back reverse the link from the question. For example, <a href=\"https://www.quora.com/Why-did-you-quit-your-job-at-Amazon\">https://www.quora.com/Why-did-you-quit-your-job-at-Amazon</a></li>\n<li>Insincere questions can be then detected by the tag <em>QuestionRestrictedInsincerePrompt</em> in the HTML code which can be obtained by python requests library.  However, there is even more elegant solution to that.</li>\n<li>At some point, we've noticed that if you add <em>/log</em> to any Quora question link than you get the full history of page changes including topics, comment, users, etc. Try it <a href=\"https://www.quora.com/unanswered/Are-Muslims-really-ashamed-of-their-religion/log\">here</a>. As you may see there is a log line \"Question marked as possibly insincere by Quora Content Review\". That is exactly what we were looking for.  </li>\n</ol>\n\n<p><a href=\"https://imgflip.com/i/2tqg7u\"><img src=\"https://i.imgflip.com/2tqg7u.jpg\"></a></p>\n\n<h2>Second stage labels</h2>\n\n<p>Of course, that's not the end of the story. Scraping 1 stage answers would be useless without getting 2 stage labels. How can we guess them? First, we tried using data from the <a href=\"https://www.kaggle.com/c/quora-question-pairs\">previous Quora competition</a>. We collected around 2000 insincere questions from there. The second option was using Related Questions but none of them were insincere. </p>\n\n<p>Our last resort was <strong>toxic users</strong>, i.e. non-anonymous users with many insincere questions. For example, <a href=\"https://www.quora.com/profile/Jim-Lunde/questions\">this guy</a>. We collected a list of such users based on train/test datasets and then scraped all of their questions. Some of these questions were not in 1 stage datasets which seemed to be very promising - they should have been in the 2 stage data. Unfortunately, as you can see by our private score we were wrong (for the best, probably).</p>\n\n<p><a href=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"><img src=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"></a></p>\n\n<h2>Small technical challenges</h2>\n\n<ul>\n<li>Notebooks and scripts on Kaggle are limited by the size of 1Mb. How could we squeeze several thousand questions into the code? We converted an array of questions into txt file, then compressed it with tar.gz, then converted the archive to base64 code and inserted it into the script. 10K questions were stored using only 500Kb of space. </li>\n<li>We used <a href=\"https://www.ip-adress.com/proxy-list\">list of 50 proxies</a> and 100 threads to parallel our scraping process. At some point, my own IP address was completely banned from requests to Quora. The access was restored after a couple of days, though.</li>\n<li>Part of code that shows user questions was written in JS so some advanced knowledge of Selenium framework was required.</li>\n</ul>\n\n<p><a href=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"><img src=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"></a></p>\n\n<h2>Takeaways</h2>\n\n<ul>\n<li>There is always a room for a creative approach to a Kaggle competition</li>\n<li>You never fail if you learn something new</li>\n<li>Whatever you do - make it fun</li>\n</ul>",
      "rawMarkdown": "## Motivation ##\nThe purpose of this competition wasn't really clear to me. So many SOTA NLP models were introduced in 2018 and none of them were allowed during the competition. It sounds like a joke to approach a complicated NLP task when even ULMFiT (kudos to @jhoward) is prohibited. So I decided to have some fun together with @evgeny000 and @mathurinache.\n\n<a href=\"https://i.imgur.com/IH4fZyx.gif?1\"><img src=\"https://i.imgur.com/IH4fZyx.gif?1\"></a>\n\n## Getting true labels ##\nMany of you were wondering about 0.782 score. The solution is quite straightforward: we just **scraped the answers**. The process was the following:\n\n 1. It's easy to notice that Quora links are very much similar to the actual question asked. By applying several heuristics you can back reverse the link from the question. For example, https://www.quora.com/Why-did-you-quit-your-job-at-Amazon\n 2. Insincere questions can be then detected by the tag *QuestionRestrictedInsincerePrompt* in the HTML code which can be obtained by python requests library.  However, there is even more elegant solution to that.\n 3. At some point, we've noticed that if you add */log* to any Quora question link than you get the full history of page changes including topics, comment, users, etc. Try it [here][1]. As you may see there is a log line \"Question marked as possibly insincere by Quora Content Review\". That is exactly what we were looking for.  \n\n<a href=\"https://imgflip.com/i/2tqg7u\"><img src=\"https://i.imgflip.com/2tqg7u.jpg\"></a>\n\n## Second stage labels ##\nOf course, that's not the end of the story. Scraping 1 stage answers would be useless without getting 2 stage labels. How can we guess them? First, we tried using data from the [previous Quora competition][2]. We collected around 2000 insincere questions from there. The second option was using Related Questions but none of them were insincere. \n\nOur last resort was **toxic users**, i.e. non-anonymous users with many insincere questions. For example, [this guy][3]. We collected a list of such users based on train/test datasets and then scraped all of their questions. Some of these questions were not in 1 stage datasets which seemed to be very promising - they should have been in the 2 stage data. Unfortunately, as you can see by our private score we were wrong (for the best, probably).\n\n<a href=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"><img src=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"></a>\n\n\n## Small technical challenges ##\n- Notebooks and scripts on Kaggle are limited by the size of 1Mb. How could we squeeze several thousand questions into the code? We converted an array of questions into txt file, then compressed it with tar.gz, then converted the archive to base64 code and inserted it into the script. 10K questions were stored using only 500Kb of space. \n- We used [list of 50 proxies][4] and 100 threads to parallel our scraping process. At some point, my own IP address was completely banned from requests to Quora. The access was restored after a couple of days, though.\n- Part of code that shows user questions was written in JS so some advanced knowledge of Selenium framework was required.\n\n<a href=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"><img src=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"></a>\n\n\n## Takeaways ##\n\n- There is always a room for a creative approach to a Kaggle competition\n- You never fail if you learn something new\n- Whatever you do - make it fun\n\n\n  [1]: https://www.quora.com/unanswered/Are-Muslims-really-ashamed-of-their-religion/log\n  [2]: https://www.kaggle.com/c/quora-question-pairs\n  [3]: https://www.quora.com/profile/Jim-Lunde/questions\n  [4]: https://www.ip-adress.com/proxy-list",
      "votes": null
    },
    {
      "id": "472050",
      "postDate": "02/15/2019 09:17:49",
      "content": "<p>Awesome, By seeing ULMFiT in this topic. I was literally shocked. I tried it for days and left as it is of no use (just to check the prominence of that). Anyways great hack by scraping. One of your team members should be JS freak. Not fruitful in long run.</p>",
      "rawMarkdown": "Awesome, By seeing ULMFiT in this topic. I was literally shocked. I tried it for days and left as it is of no use (just to check the prominence of that). Anyways great hack by scraping. One of your team members should be JS freak. Not fruitful in long run.",
      "votes": null
    },
    {
      "id": "472053",
      "postDate": "02/15/2019 09:19:33",
      "content": "<p>So, this outstanding score was not a result of a good model? </p>\n\n<p><img src=\"https://habrastorage.org/webt/tm/pi/3y/tmpi3yxnd_6mvff0z2fujqaqtjw.png\" alt=\"\"></p>",
      "rawMarkdown": "So, this outstanding score was not a result of a good model? \n\n![](https://habrastorage.org/webt/tm/pi/3y/tmpi3yxnd_6mvff0z2fujqaqtjw.png)",
      "votes": null
    },
    {
      "id": "472059",
      "postDate": "02/15/2019 09:23:18",
      "content": "<p>I have been playing around with ULMFIT and BERT as well for the past few days and could not get  a significantly better score than the ones on Private LB. So calling the whole task a joke is a bit of a stretch. I will continue to tinker with those SOTA models though.</p>",
      "rawMarkdown": "I have been playing around with ULMFIT and BERT as well for the past few days and could not get  a significantly better score than the ones on Private LB. So calling the whole task a joke is a bit of a stretch. I will continue to tinker with those SOTA models though.",
      "votes": null
    },
    {
      "id": "472061",
      "postDate": "02/15/2019 09:26:13",
      "content": "<p>who would've thought</p>",
      "rawMarkdown": "who would've thought",
      "votes": null
    },
    {
      "id": "472062",
      "postDate": "02/15/2019 09:26:32",
      "content": "<p><code>we just scraped the answers</code></p>\n\n<p>my favorite part </p>",
      "rawMarkdown": "`we just scraped the answers`\n\nmy favorite part",
      "votes": null
    },
    {
      "id": "472064",
      "postDate": "02/15/2019 09:28:51",
      "content": "<p>it seems like annotation was a bit different from Quora Content Review (which we all were trying to improve)</p>",
      "rawMarkdown": "it seems like annotation was a bit different from Quora Content Review (which we all were trying to improve)",
      "votes": null
    },
    {
      "id": "472068",
      "postDate": "02/15/2019 09:33:17",
      "content": "<p>It was very smart</p>",
      "rawMarkdown": "It was very smart",
      "votes": null
    },
    {
      "id": "472088",
      "postDate": "02/15/2019 10:17:24",
      "content": "<p>This will further increase the number of jokes about Russian hackers that I hear from my Dutch colleagues :)</p>",
      "rawMarkdown": "This will further increase the number of jokes about Russian hackers that I hear from my Dutch colleagues :)",
      "votes": null
    },
    {
      "id": "472113",
      "postDate": "02/15/2019 11:29:58",
      "content": "<p><a href=\"/ppleskov\">@ppleskov</a> хаха офигенно</p>",
      "rawMarkdown": "ppleskov хаха офигенно",
      "votes": null
    },
    {
      "id": "472123",
      "postDate": "02/15/2019 11:49:54",
      "content": "<p>parallel GPU used to win competitions, now it's parallel scraping ;)</p>",
      "rawMarkdown": "parallel GPU used to win competitions, now it's parallel scraping ;)",
      "votes": null
    },
    {
      "id": "472132",
      "postDate": "02/15/2019 11:59:43",
      "content": "<p>cores before hoes</p>",
      "rawMarkdown": "cores before hoes",
      "votes": null
    },
    {
      "id": "472181",
      "postDate": "02/15/2019 13:04:58",
      "content": "<p>It seems you have a lot of fun during this competition :)</p>",
      "rawMarkdown": "It seems you have a lot of fun during this competition :)",
      "votes": null
    },
    {
      "id": "472488",
      "postDate": "02/16/2019 02:38:55",
      "content": "<p>compete for not only learning, but also for fun ...</p>",
      "rawMarkdown": "compete for not only learning, but also for fun ...",
      "votes": null
    },
    {
      "id": "472629",
      "postDate": "02/16/2019 10:38:30",
      "content": "<p>Happy Kaggling! That is the point. :)</p>",
      "rawMarkdown": "Happy Kaggling! That is the point. :)",
      "votes": null
    },
    {
      "id": "472732",
      "postDate": "02/16/2019 14:57:28",
      "content": "<p>What a fun and creative way to go into this competition! </p>",
      "rawMarkdown": "What a fun and creative way to go into this competition!",
      "votes": null
    },
    {
      "id": "473861",
      "postDate": "02/18/2019 15:54:34",
      "content": "<p>I am not taking anything away from your greatness <a href=\"/ppleskov\">@ppleskov</a>. you are one of those whom I admire the most here on kaggle. Some times posed restrictions may generate that need which can result in new discoveries perhaps new architectures or alternative solutions. \nThis competition you have chosen Plan B. And even from this we learnt about some real world scenarios.\nEven if you had stick to Plan A  I am sure you could have produced a great solution for us</p>",
      "rawMarkdown": "I am not taking anything away from your greatness @ppleskov. you are one of those whom I admire the most here on kaggle. Some times posed restrictions may generate that need which can result in new discoveries perhaps new architectures or alternative solutions. \nThis competition you have chosen Plan B. And even from this we learnt about some real world scenarios.\nEven if you had stick to Plan A  I am sure you could have produced a great solution for us",
      "votes": null
    },
    {
      "id": "474195",
      "postDate": "02/19/2019 04:09:57",
      "content": "<p>congrats and thanks for sharing your approach:)</p>",
      "rawMarkdown": "congrats and thanks for sharing your approach:)",
      "votes": null
    },
    {
      "id": "474220",
      "postDate": "02/19/2019 04:58:47",
      "content": "<p><img src=\"https://pics.me.me/actually-im-not-even-mad-thats-amazing-13517597.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://pics.me.me/actually-im-not-even-mad-thats-amazing-13517597.png)",
      "votes": null
    },
    {
      "id": "474246",
      "postDate": "02/19/2019 06:00:39",
      "content": "<p>This is a creative approach and the squeeze trick is interesting too. Thanks for sharing!</p>",
      "rawMarkdown": "This is a creative approach and the squeeze trick is interesting too. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "474251",
      "postDate": "02/19/2019 06:07:48",
      "content": "<p>Great hack !! Thanks for sharing !</p>",
      "rawMarkdown": "Great hack !! Thanks for sharing !",
      "votes": null
    },
    {
      "id": "475176",
      "postDate": "02/20/2019 11:45:28",
      "content": "<p>The toxic user toxic user you point out (<a href=\"https://www.quora.com/profile/Jim-Lunde/questions\">this guy</a>)  is priceless!</p>",
      "rawMarkdown": "The toxic user toxic user you point out (<a href=\"https://www.quora.com/profile/Jim-Lunde/questions\">this guy</a>)  is priceless!",
      "votes": null
    },
    {
      "id": "475617",
      "postDate": "02/21/2019 01:46:38",
      "content": "<p>Nice, thanks for sharing!</p>",
      "rawMarkdown": "Nice, thanks for sharing!",
      "votes": null
    },
    {
      "id": "475919",
      "postDate": "02/21/2019 11:26:59",
      "content": "<p>cool cool cool!</p>",
      "rawMarkdown": "cool cool cool!",
      "votes": null
    },
    {
      "id": "476870",
      "postDate": "02/23/2019 11:04:29",
      "content": "<p>So, just trolling the competition? You must have lots of free time to do all that just for this :)</p>\n\n<p>I think the competition was fine, no need to always play with the latest tools made by some big corporations etc.</p>",
      "rawMarkdown": "So, just trolling the competition? You must have lots of free time to do all that just for this :)\n\nI think the competition was fine, no need to always play with the latest tools made by some big corporations etc.",
      "votes": null
    },
    {
      "id": "476976",
      "postDate": "02/23/2019 15:36:07",
      "content": "<p>Very interesting, thanks for sharing your experience.</p>",
      "rawMarkdown": "Very interesting, thanks for sharing your experience.",
      "votes": null
    },
    {
      "id": "717777",
      "postDate": "01/13/2020 15:53:10",
      "content": "<p>Wow, this topic did NOT age well, did it?</p>",
      "rawMarkdown": "Wow, this topic did NOT age well, did it?",
      "votes": null
    },
    {
      "id": "727612",
      "postDate": "01/23/2020 22:09:05",
      "content": "<p>The amount of \"awesome idea!\" comments about cheating in the competition is pretty baffling as well. I understand it's fun to see a unique solution, but the community (at least most commenters on the thread) are actively agreeing with the behavior. There's a larger issue here.</p>",
      "rawMarkdown": "The amount of \"awesome idea!\" comments about cheating in the competition is pretty baffling as well. I understand it's fun to see a unique solution, but the community (at least most commenters on the thread) are actively agreeing with the behavior. There's a larger issue here.",
      "votes": null
    },
    {
      "id": "790994",
      "postDate": "03/30/2020 02:46:01",
      "content": "<p>I was here for a good model.👀 </p>",
      "rawMarkdown": "I was here for a good model.👀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 472050,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/15/2019 09:17:49",
      "content": "<p>Awesome, By seeing ULMFiT in this topic. I was literally shocked. I tried it for days and left as it is of no use (just to check the prominence of that). Anyways great hack by scraping. One of your team members should be JS freak. Not fruitful in long run.</p>",
      "votes": null,
      "replies": [
        {
          "id": 472059,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/15/2019 09:23:18",
          "content": "<p>I have been playing around with ULMFIT and BERT as well for the past few days and could not get  a significantly better score than the ones on Private LB. So calling the whole task a joke is a bit of a stretch. I will continue to tinker with those SOTA models though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472053,
      "author_name": "",
      "author_url": "",
      "post_date": "02/15/2019 09:19:33",
      "content": "<p>So, this outstanding score was not a result of a good model? </p>\n\n<p><img src=\"https://habrastorage.org/webt/tm/pi/3y/tmpi3yxnd_6mvff0z2fujqaqtjw.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 472061,
          "author_name": "",
          "author_url": "",
          "post_date": "02/15/2019 09:26:13",
          "content": "<p>who would've thought</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472062,
      "author_name": "hireme",
      "author_url": "",
      "post_date": "02/15/2019 09:26:32",
      "content": "<p><code>we just scraped the answers</code></p>\n\n<p>my favorite part </p>",
      "votes": null,
      "replies": [
        {
          "id": 472064,
          "author_name": "",
          "author_url": "",
          "post_date": "02/15/2019 09:28:51",
          "content": "<p>it seems like annotation was a bit different from Quora Content Review (which we all were trying to improve)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472068,
          "author_name": "hireme",
          "author_url": "",
          "post_date": "02/15/2019 09:33:17",
          "content": "<p>It was very smart</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472088,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "02/15/2019 10:17:24",
      "content": "<p>This will further increase the number of jokes about Russian hackers that I hear from my Dutch colleagues :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472113,
      "author_name": "kcostya",
      "author_url": "",
      "post_date": "02/15/2019 11:29:58",
      "content": "<p><a href=\"/ppleskov\">@ppleskov</a> хаха офигенно</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472123,
      "author_name": "ksevta",
      "author_url": "",
      "post_date": "02/15/2019 11:49:54",
      "content": "<p>parallel GPU used to win competitions, now it's parallel scraping ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 472132,
          "author_name": "",
          "author_url": "",
          "post_date": "02/15/2019 11:59:43",
          "content": "<p>cores before hoes</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472181,
      "author_name": "demonplus",
      "author_url": "",
      "post_date": "02/15/2019 13:04:58",
      "content": "<p>It seems you have a lot of fun during this competition :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472488,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "02/16/2019 02:38:55",
      "content": "<p>compete for not only learning, but also for fun ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472629,
      "author_name": "sunnymarkliu",
      "author_url": "",
      "post_date": "02/16/2019 10:38:30",
      "content": "<p>Happy Kaggling! That is the point. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472732,
      "author_name": "pbenavides",
      "author_url": "",
      "post_date": "02/16/2019 14:57:28",
      "content": "<p>What a fun and creative way to go into this competition! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 473861,
      "author_name": "reachkishore",
      "author_url": "",
      "post_date": "02/18/2019 15:54:34",
      "content": "<p>I am not taking anything away from your greatness <a href=\"/ppleskov\">@ppleskov</a>. you are one of those whom I admire the most here on kaggle. Some times posed restrictions may generate that need which can result in new discoveries perhaps new architectures or alternative solutions. \nThis competition you have chosen Plan B. And even from this we learnt about some real world scenarios.\nEven if you had stick to Plan A  I am sure you could have produced a great solution for us</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474195,
      "author_name": "econdata",
      "author_url": "",
      "post_date": "02/19/2019 04:09:57",
      "content": "<p>congrats and thanks for sharing your approach:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474220,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "02/19/2019 04:58:47",
      "content": "<p><img src=\"https://pics.me.me/actually-im-not-even-mad-thats-amazing-13517597.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474246,
      "author_name": "leekltw1",
      "author_url": "",
      "post_date": "02/19/2019 06:00:39",
      "content": "<p>This is a creative approach and the squeeze trick is interesting too. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474251,
      "author_name": "priteshshrivastava",
      "author_url": "",
      "post_date": "02/19/2019 06:07:48",
      "content": "<p>Great hack !! Thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 475176,
      "author_name": "mhonkaggle",
      "author_url": "",
      "post_date": "02/20/2019 11:45:28",
      "content": "<p>The toxic user toxic user you point out (<a href=\"https://www.quora.com/profile/Jim-Lunde/questions\">this guy</a>)  is priceless!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 475617,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "02/21/2019 01:46:38",
      "content": "<p>Nice, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 475919,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "02/21/2019 11:26:59",
      "content": "<p>cool cool cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476870,
      "author_name": "donkeys",
      "author_url": "",
      "post_date": "02/23/2019 11:04:29",
      "content": "<p>So, just trolling the competition? You must have lots of free time to do all that just for this :)</p>\n\n<p>I think the competition was fine, no need to always play with the latest tools made by some big corporations etc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476976,
      "author_name": "carlosalvarado",
      "author_url": "",
      "post_date": "02/23/2019 15:36:07",
      "content": "<p>Very interesting, thanks for sharing your experience.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 717777,
      "author_name": "lucabasa",
      "author_url": "",
      "post_date": "01/13/2020 15:53:10",
      "content": "<p>Wow, this topic did NOT age well, did it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 727612,
          "author_name": "maltonji",
          "author_url": "",
          "post_date": "01/23/2020 22:09:05",
          "content": "<p>The amount of \"awesome idea!\" comments about cheating in the competition is pretty baffling as well. I understand it's fun to see a unique solution, but the community (at least most commenters on the thread) are actively agreeing with the behavior. There's a larger issue here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 790994,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "03/30/2020 02:46:01",
      "content": "<p>I was here for a good model.👀 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "472042": "## Motivation ##\nThe purpose of this competition wasn't really clear to me. So many SOTA NLP models were introduced in 2018 and none of them were allowed during the competition. It sounds like a joke to approach a complicated NLP task when even ULMFiT (kudos to @jhoward) is prohibited. So I decided to have some fun together with @evgeny000 and @mathurinache.\n\n<a href=\"https://i.imgur.com/IH4fZyx.gif?1\"><img src=\"https://i.imgur.com/IH4fZyx.gif?1\"></a>\n\n## Getting true labels ##\nMany of you were wondering about 0.782 score. The solution is quite straightforward: we just **scraped the answers**. The process was the following:\n\n 1. It's easy to notice that Quora links are very much similar to the actual question asked. By applying several heuristics you can back reverse the link from the question. For example, https://www.quora.com/Why-did-you-quit-your-job-at-Amazon\n 2. Insincere questions can be then detected by the tag *QuestionRestrictedInsincerePrompt* in the HTML code which can be obtained by python requests library.  However, there is even more elegant solution to that.\n 3. At some point, we've noticed that if you add */log* to any Quora question link than you get the full history of page changes including topics, comment, users, etc. Try it [here][1]. As you may see there is a log line \"Question marked as possibly insincere by Quora Content Review\". That is exactly what we were looking for.  \n\n<a href=\"https://imgflip.com/i/2tqg7u\"><img src=\"https://i.imgflip.com/2tqg7u.jpg\"></a>\n\n## Second stage labels ##\nOf course, that's not the end of the story. Scraping 1 stage answers would be useless without getting 2 stage labels. How can we guess them? First, we tried using data from the [previous Quora competition][2]. We collected around 2000 insincere questions from there. The second option was using Related Questions but none of them were insincere. \n\nOur last resort was **toxic users**, i.e. non-anonymous users with many insincere questions. For example, [this guy][3]. We collected a list of such users based on train/test datasets and then scraped all of their questions. Some of these questions were not in 1 stage datasets which seemed to be very promising - they should have been in the 2 stage data. Unfortunately, as you can see by our private score we were wrong (for the best, probably).\n\n<a href=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"><img src=\"http://www.quickmeme.com/img/98/98e8d07eb0ee43919ffe3526f0037200e832438fe2d9a89f71bdb209f321df01.jpg\"></a>\n\n\n## Small technical challenges ##\n- Notebooks and scripts on Kaggle are limited by the size of 1Mb. How could we squeeze several thousand questions into the code? We converted an array of questions into txt file, then compressed it with tar.gz, then converted the archive to base64 code and inserted it into the script. 10K questions were stored using only 500Kb of space. \n- We used [list of 50 proxies][4] and 100 threads to parallel our scraping process. At some point, my own IP address was completely banned from requests to Quora. The access was restored after a couple of days, though.\n- Part of code that shows user questions was written in JS so some advanced knowledge of Selenium framework was required.\n\n<a href=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"><img src=\"http://www.quickmeme.com/img/8c/8c7e8d075b2177dd011ad4ba4657b5164b80eff9af8ec16105c8856b23ffe240.jpg\"></a>\n\n\n## Takeaways ##\n\n- There is always a room for a creative approach to a Kaggle competition\n- You never fail if you learn something new\n- Whatever you do - make it fun\n\n\n  [1]: https://www.quora.com/unanswered/Are-Muslims-really-ashamed-of-their-religion/log\n  [2]: https://www.kaggle.com/c/quora-question-pairs\n  [3]: https://www.quora.com/profile/Jim-Lunde/questions\n  [4]: https://www.ip-adress.com/proxy-list",
    "472050": "Awesome, By seeing ULMFiT in this topic. I was literally shocked. I tried it for days and left as it is of no use (just to check the prominence of that). Anyways great hack by scraping. One of your team members should be JS freak. Not fruitful in long run.",
    "472053": "So, this outstanding score was not a result of a good model? \n\n![](https://habrastorage.org/webt/tm/pi/3y/tmpi3yxnd_6mvff0z2fujqaqtjw.png)",
    "472059": "I have been playing around with ULMFIT and BERT as well for the past few days and could not get  a significantly better score than the ones on Private LB. So calling the whole task a joke is a bit of a stretch. I will continue to tinker with those SOTA models though.",
    "472061": "who would've thought",
    "472062": "`we just scraped the answers`\n\nmy favorite part",
    "472064": "it seems like annotation was a bit different from Quora Content Review (which we all were trying to improve)",
    "472068": "It was very smart",
    "472088": "This will further increase the number of jokes about Russian hackers that I hear from my Dutch colleagues :)",
    "472113": "ppleskov хаха офигенно",
    "472123": "parallel GPU used to win competitions, now it's parallel scraping ;)",
    "472132": "cores before hoes",
    "472181": "It seems you have a lot of fun during this competition :)",
    "472488": "compete for not only learning, but also for fun ...",
    "472629": "Happy Kaggling! That is the point. :)",
    "472732": "What a fun and creative way to go into this competition!",
    "473861": "I am not taking anything away from your greatness @ppleskov. you are one of those whom I admire the most here on kaggle. Some times posed restrictions may generate that need which can result in new discoveries perhaps new architectures or alternative solutions. \nThis competition you have chosen Plan B. And even from this we learnt about some real world scenarios.\nEven if you had stick to Plan A  I am sure you could have produced a great solution for us",
    "474195": "congrats and thanks for sharing your approach:)",
    "474220": "![](https://pics.me.me/actually-im-not-even-mad-thats-amazing-13517597.png)",
    "474246": "This is a creative approach and the squeeze trick is interesting too. Thanks for sharing!",
    "474251": "Great hack !! Thanks for sharing !",
    "475176": "The toxic user toxic user you point out (<a href=\"https://www.quora.com/profile/Jim-Lunde/questions\">this guy</a>)  is priceless!",
    "475617": "Nice, thanks for sharing!",
    "475919": "cool cool cool!",
    "476870": "So, just trolling the competition? You must have lots of free time to do all that just for this :)\n\nI think the competition was fine, no need to always play with the latest tools made by some big corporations etc.",
    "476976": "Very interesting, thanks for sharing your experience.",
    "717777": "Wow, this topic did NOT age well, did it?",
    "727612": "The amount of \"awesome idea!\" comments about cheating in the competition is pretty baffling as well. I understand it's fun to see a unique solution, but the community (at least most commenters on the thread) are actively agreeing with the behavior. There's a larger issue here.",
    "790994": "I was here for a good model.👀"
  },
  "source": "meta"
}