{
  "id": 74074,
  "title": "What about guessing \"sincere\" all the time?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74074",
  "author_name": "Junlin Shang",
  "post_date": "2018-12-08T06:23:07.254000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi guys, I'm new at this so this may be a dumb question:</p>\n\n<p>I noticed only 6% of the training samples are labeled \"insincere\". So as a baseline, I made a predictions full of 0s. After I submitted it, I got 0 f1 score.</p>\n\n<p>I know this is a stupid predictions, but I didn't expect such a low score. There must be at least <strong>some sincere questions</strong> inside the test set, right? What did I miss?</p>",
  "messages": [
    {
      "id": 435514,
      "postDate": "2018-12-08T06:23:07.253Z",
      "content": "<p>Hi guys, I'm new at this so this may be a dumb question:</p>\n\n<p>I noticed only 6% of the training samples are labeled \"insincere\". So as a baseline, I made a predictions full of 0s. After I submitted it, I got 0 f1 score.</p>\n\n<p>I know this is a stupid predictions, but I didn't expect such a low score. There must be at least <strong>some sincere questions</strong> inside the test set, right? What did I miss?</p>",
      "rawMarkdown": "Hi guys, I'm new at this so this may be a dumb question:\n\nI noticed only 6% of the training samples are labeled \"insincere\". So as a baseline, I made a predictions full of 0s. After I submitted it, I got 0 f1 score.\n\nI know this is a stupid predictions, but I didn't expect such a low score. There must be at least **some sincere questions** inside the test set, right? What did I miss?",
      "votes": 1
    },
    {
      "id": 435558,
      "postDate": "2018-12-08T08:44:00.790Z",
      "content": "<p>The metric chosen for this competition was actually chosen to avoid the trick you want to use. F1 score depends on which label is chosen as the \"positive\" one, which is the most relevant one (here the \"1\" label). If you have zero true positives, you have 0 precision and recall.</p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\">https://en.wikipedia.org/wiki/Precision_and_recall</a></p>\n\n<p>Probably you are choosing in scikit-learn a bad labelling.</p>",
      "rawMarkdown": "The metric chosen for this competition was actually chosen to avoid the trick you want to use. F1 score depends on which label is chosen as the \"positive\" one, which is the most relevant one (here the \"1\" label). If you have zero true positives, you have 0 precision and recall.\n\nhttps://en.wikipedia.org/wiki/Precision_and_recall\n\nProbably you are choosing in scikit-learn a bad labelling.",
      "replies": [
        {
          "id": 435909,
          "postDate": "2018-12-09T03:40:55.243Z",
          "content": "<p>Thank you for your reply! Now I see it, you are absolutely right! </p>",
          "rawMarkdown": "Thank you for your reply! Now I see it, you are absolutely right! "
        }
      ]
    },
    {
      "id": 435517,
      "postDate": "2018-12-08T06:27:22.373Z",
      "content": "<p>Before I submitted the all 0s prediction, I split 15% out of the training data as validation set. And the all 0s prediction got me a 0.9 f1 score <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">by <code>f1_score</code> method from sklearn</a>. So I must made a stupid mistake somewhere...</p>",
      "rawMarkdown": "Before I submitted the all 0s prediction, I split 15% out of the training data as validation set. And the all 0s prediction got me a 0.9 f1 score [by `f1_score ` method from sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html). So I must made a stupid mistake somewhere..."
    },
    {
      "id": 435519,
      "postDate": "2018-12-08T06:35:14.610Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 435558,
      "author_name": "Daniel Cañueto",
      "author_url": "",
      "post_date": "2018-12-08T08:44:00.790000",
      "content": "<p>The metric chosen for this competition was actually chosen to avoid the trick you want to use. F1 score depends on which label is chosen as the \"positive\" one, which is the most relevant one (here the \"1\" label). If you have zero true positives, you have 0 precision and recall.</p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\">https://en.wikipedia.org/wiki/Precision_and_recall</a></p>\n\n<p>Probably you are choosing in scikit-learn a bad labelling.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 435909,
          "author_name": "Junlin Shang",
          "author_url": "",
          "post_date": "2018-12-09T03:40:55.243000",
          "content": "<p>Thank you for your reply! Now I see it, you are absolutely right! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435517,
      "author_name": "Junlin Shang",
      "author_url": "",
      "post_date": "2018-12-08T06:27:22.373000",
      "content": "<p>Before I submitted the all 0s prediction, I split 15% out of the training data as validation set. And the all 0s prediction got me a 0.9 f1 score <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">by <code>f1_score</code> method from sklearn</a>. So I must made a stupid mistake somewhere...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435519,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T06:35:14.610000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "435514": "Hi guys, I'm new at this so this may be a dumb question:\n\nI noticed only 6% of the training samples are labeled \"insincere\". So as a baseline, I made a predictions full of 0s. After I submitted it, I got 0 f1 score.\n\nI know this is a stupid predictions, but I didn't expect such a low score. There must be at least **some sincere questions** inside the test set, right? What did I miss?",
    "435558": "The metric chosen for this competition was actually chosen to avoid the trick you want to use. F1 score depends on which label is chosen as the \"positive\" one, which is the most relevant one (here the \"1\" label). If you have zero true positives, you have 0 precision and recall.\n\nhttps://en.wikipedia.org/wiki/Precision_and_recall\n\nProbably you are choosing in scikit-learn a bad labelling.",
    "435517": "Before I submitted the all 0s prediction, I split 15% out of the training data as validation set. And the all 0s prediction got me a 0.9 f1 score [by `f1_score ` method from sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html). So I must made a stupid mistake somewhere...",
    "435519": ""
  }
}