{
  "id": 75903,
  "title": "Synthetic test data? ",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75903",
  "author_name": "",
  "post_date": "2018-12-27T11:45:59.928240300Z",
  "votes": 15,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hey. </p>\n\n<p>I have a hypothesis that all questions from the test dataset are generated. If you look for the questions from the training dataset, you might easily find them on Quora platform (except toxic questions, which are deleted). See the example below. </p>\n\n<p>Train questions example (Q - question from the dataset, q - first result from Quora search) </p>\n\n<pre><code>Q: What are the nomenclature of the microorganism?\nq: What are the nomenclatures of a microorganism?\n\nQ: Should the U.S. keep troops in Syria?\nq: Should the U.S. keep troops in Syria?\n\nQ: What role does Serena play in the TV show \"Bewitched\"?\nq: What role does Serena play in the TV show \"Bewitched\"?\n</code></pre>\n\n<p>Question from train part and the first question from Quora search result are exactly the same. And now let's have a look at the test questions: </p>\n\n<pre><code>Q: What is the working of thermistors?\nq: How do PTC thermistors work?\n\nQ: Where can I find car insurance on a budget?\nq: Where can I find tips on finding cheap car insurance?\n\nQ: How can I check how old is my iPhone?\nq: Can I find out how old my iPhone is?\n</code></pre>\n\n<p>As you could see, questions from the test part are similar but do not match exactly. They might be generated or obfuscated with Quora's pretrained language model in order to overcome any cheating attempts.  </p>\n\n<p>What do you think? </p>",
  "messages": [
    {
      "id": "446040",
      "postDate": "12/27/2018 11:45:59",
      "content": "<p>Hey. </p>\n\n<p>I have a hypothesis that all questions from the test dataset are generated. If you look for the questions from the training dataset, you might easily find them on Quora platform (except toxic questions, which are deleted). See the example below. </p>\n\n<p>Train questions example (Q - question from the dataset, q - first result from Quora search) </p>\n\n<pre><code>Q: What are the nomenclature of the microorganism?\nq: What are the nomenclatures of a microorganism?\n\nQ: Should the U.S. keep troops in Syria?\nq: Should the U.S. keep troops in Syria?\n\nQ: What role does Serena play in the TV show \"Bewitched\"?\nq: What role does Serena play in the TV show \"Bewitched\"?\n</code></pre>\n\n<p>Question from train part and the first question from Quora search result are exactly the same. And now let's have a look at the test questions: </p>\n\n<pre><code>Q: What is the working of thermistors?\nq: How do PTC thermistors work?\n\nQ: Where can I find car insurance on a budget?\nq: Where can I find tips on finding cheap car insurance?\n\nQ: How can I check how old is my iPhone?\nq: Can I find out how old my iPhone is?\n</code></pre>\n\n<p>As you could see, questions from the test part are similar but do not match exactly. They might be generated or obfuscated with Quora's pretrained language model in order to overcome any cheating attempts.  </p>\n\n<p>What do you think? </p>",
      "rawMarkdown": "Hey. \n\nI have a hypothesis that all questions from the test dataset are generated. If you look for the questions from the training dataset, you might easily find them on Quora platform (except toxic questions, which are deleted). See the example below. \n\nTrain questions example (Q - question from the dataset, q - first result from Quora search) \n\n    Q: What are the nomenclature of the microorganism?\n    q: What are the nomenclatures of a microorganism?\n    \n    Q: Should the U.S. keep troops in Syria?\n    q: Should the U.S. keep troops in Syria?\n    \n    Q: What role does Serena play in the TV show \"Bewitched\"?\n    q: What role does Serena play in the TV show \"Bewitched\"?\n\nQuestion from train part and the first question from Quora search result are exactly the same. And now let's have a look at the test questions: \n\n    Q: What is the working of thermistors?\n    q: How do PTC thermistors work?\n    \n    Q: Where can I find car insurance on a budget?\n    q: Where can I find tips on finding cheap car insurance?\n    \n    Q: How can I check how old is my iPhone?\n    q: Can I find out how old my iPhone is?\n\nAs you could see, questions from the test part are similar but do not match exactly. They might be generated or obfuscated with Quora's pretrained language model in order to overcome any cheating attempts.  \n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "446081",
      "postDate": "12/27/2018 13:18:12",
      "content": "<p>Any Idea how we could have cheated if the questions were not generated synthetically? \nI just think by using Kernels, time limits and 2nd stage test data they need not have thought of generating the test questions.</p>",
      "rawMarkdown": "Any Idea how we could have cheated if the questions were not generated synthetically? \nI just think by using Kernels, time limits and 2nd stage test data they need not have thought of generating the test questions.",
      "votes": null
    },
    {
      "id": "446092",
      "postDate": "12/27/2018 13:33:43",
      "content": "<ol>\n<li>Parse questions from Quora;</li>\n<li>If question is found in the database -&gt; prediction = 0 (found == levenshtein distance &lt;= 1);\nThis simple approach gives ~0.85 F1 on the train dataset (~500 samples). </li>\n</ol>\n\n<p>However, this is against the rules (usage of additional data) and participants might be disqualified. </p>",
      "rawMarkdown": "1. Parse questions from Quora;\n2. If question is found in the database -&gt; prediction = 0 (found == levenshtein distance &lt;= 1);\nThis simple approach gives ~0.85 F1 on the train dataset (~500 samples). \n\nHowever, this is against the rules (usage of additional data) and participants might be disqualified.",
      "votes": null
    },
    {
      "id": "446096",
      "postDate": "12/27/2018 13:42:54",
      "content": "<p>Exactly. They took care of it using kernel constraints.</p>",
      "rawMarkdown": "Exactly. They took care of it using kernel constraints.",
      "votes": null
    },
    {
      "id": "446451",
      "postDate": "12/28/2018 04:51:27",
      "content": "<p>very nice! pay tribute to the organizers.</p>",
      "rawMarkdown": "very nice! pay tribute to the organizers.",
      "votes": null
    },
    {
      "id": "446513",
      "postDate": "12/28/2018 08:14:38",
      "content": "<p>It seems most questions from the test sets &amp; train sets  are generated . I tried synthetic some of train data by shuffling, but  it helpless.</p>",
      "rawMarkdown": "It seems most questions from the test sets &amp; train sets  are generated . I tried synthetic some of train data by shuffling, but  it helpless.",
      "votes": null
    },
    {
      "id": "446537",
      "postDate": "12/28/2018 08:57:21",
      "content": "<p>Very good finding! The test data may be a sample from duplicate questions which are not available online. See: <a href=\"https://www.kaggle.com/c/quora-question-pairs\">https://www.kaggle.com/c/quora-question-pairs</a> </p>\n\n<p>Also the examples you shared seemed like malformed sentences that people form. Maybe Quora corrected them or if they are duplicated, Quora selected the one with better grammar.</p>",
      "rawMarkdown": "Very good finding! The test data may be a sample from duplicate questions which are not available online. See: https://www.kaggle.com/c/quora-question-pairs \n\nAlso the examples you shared seemed like malformed sentences that people form. Maybe Quora corrected them or if they are duplicated, Quora selected the one with better grammar.",
      "votes": null
    },
    {
      "id": "446541",
      "postDate": "12/28/2018 09:02:43",
      "content": "<blockquote>\n  <p>I tried synthetic some of train data by shuffling</p>\n</blockquote>\n\n<p>Could you give more details on this? </p>",
      "rawMarkdown": "&gt;I tried synthetic some of train data by shuffling\n\nCould you give more details on this?",
      "votes": null
    },
    {
      "id": "447915",
      "postDate": "12/30/2018 20:37:38",
      "content": "<p>On the other hand, the test set has tremendous amount of mispellings. Any idea how to clean them within time limits?</p>",
      "rawMarkdown": "On the other hand, the test set has tremendous amount of mispellings. Any idea how to clean them within time limits?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 446081,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "12/27/2018 13:18:12",
      "content": "<p>Any Idea how we could have cheated if the questions were not generated synthetically? \nI just think by using Kernels, time limits and 2nd stage test data they need not have thought of generating the test questions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 446092,
          "author_name": "",
          "author_url": "",
          "post_date": "12/27/2018 13:33:43",
          "content": "<ol>\n<li>Parse questions from Quora;</li>\n<li>If question is found in the database -&gt; prediction = 0 (found == levenshtein distance &lt;= 1);\nThis simple approach gives ~0.85 F1 on the train dataset (~500 samples). </li>\n</ol>\n\n<p>However, this is against the rules (usage of additional data) and participants might be disqualified. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446096,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/27/2018 13:42:54",
          "content": "<p>Exactly. They took care of it using kernel constraints.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 446451,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/28/2018 04:51:27",
      "content": "<p>very nice! pay tribute to the organizers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 446513,
      "author_name": "mr007rin",
      "author_url": "",
      "post_date": "12/28/2018 08:14:38",
      "content": "<p>It seems most questions from the test sets &amp; train sets  are generated . I tried synthetic some of train data by shuffling, but  it helpless.</p>",
      "votes": null,
      "replies": [
        {
          "id": 446541,
          "author_name": "",
          "author_url": "",
          "post_date": "12/28/2018 09:02:43",
          "content": "<blockquote>\n  <p>I tried synthetic some of train data by shuffling</p>\n</blockquote>\n\n<p>Could you give more details on this? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 446537,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "12/28/2018 08:57:21",
      "content": "<p>Very good finding! The test data may be a sample from duplicate questions which are not available online. See: <a href=\"https://www.kaggle.com/c/quora-question-pairs\">https://www.kaggle.com/c/quora-question-pairs</a> </p>\n\n<p>Also the examples you shared seemed like malformed sentences that people form. Maybe Quora corrected them or if they are duplicated, Quora selected the one with better grammar.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 447915,
      "author_name": "isikkuntay",
      "author_url": "",
      "post_date": "12/30/2018 20:37:38",
      "content": "<p>On the other hand, the test set has tremendous amount of mispellings. Any idea how to clean them within time limits?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "446040": "Hey. \n\nI have a hypothesis that all questions from the test dataset are generated. If you look for the questions from the training dataset, you might easily find them on Quora platform (except toxic questions, which are deleted). See the example below. \n\nTrain questions example (Q - question from the dataset, q - first result from Quora search) \n\n    Q: What are the nomenclature of the microorganism?\n    q: What are the nomenclatures of a microorganism?\n    \n    Q: Should the U.S. keep troops in Syria?\n    q: Should the U.S. keep troops in Syria?\n    \n    Q: What role does Serena play in the TV show \"Bewitched\"?\n    q: What role does Serena play in the TV show \"Bewitched\"?\n\nQuestion from train part and the first question from Quora search result are exactly the same. And now let's have a look at the test questions: \n\n    Q: What is the working of thermistors?\n    q: How do PTC thermistors work?\n    \n    Q: Where can I find car insurance on a budget?\n    q: Where can I find tips on finding cheap car insurance?\n    \n    Q: How can I check how old is my iPhone?\n    q: Can I find out how old my iPhone is?\n\nAs you could see, questions from the test part are similar but do not match exactly. They might be generated or obfuscated with Quora's pretrained language model in order to overcome any cheating attempts.  \n\nWhat do you think?",
    "446081": "Any Idea how we could have cheated if the questions were not generated synthetically? \nI just think by using Kernels, time limits and 2nd stage test data they need not have thought of generating the test questions.",
    "446092": "1. Parse questions from Quora;\n2. If question is found in the database -&gt; prediction = 0 (found == levenshtein distance &lt;= 1);\nThis simple approach gives ~0.85 F1 on the train dataset (~500 samples). \n\nHowever, this is against the rules (usage of additional data) and participants might be disqualified.",
    "446096": "Exactly. They took care of it using kernel constraints.",
    "446451": "very nice! pay tribute to the organizers.",
    "446513": "It seems most questions from the test sets &amp; train sets  are generated . I tried synthetic some of train data by shuffling, but  it helpless.",
    "446537": "Very good finding! The test data may be a sample from duplicate questions which are not available online. See: https://www.kaggle.com/c/quora-question-pairs \n\nAlso the examples you shared seemed like malformed sentences that people form. Maybe Quora corrected them or if they are duplicated, Quora selected the one with better grammar.",
    "446541": "&gt;I tried synthetic some of train data by shuffling\n\nCould you give more details on this?",
    "447915": "On the other hand, the test set has tremendous amount of mispellings. Any idea how to clean them within time limits?"
  },
  "source": "meta"
}