{
  "id": 403011,
  "title": "Is it wrong to predict all data with true?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/403011",
  "author_name": "",
  "post_date": "2023-04-20T16:44:45.286017700Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am a beginner in machine learning and not very good at English. Sorry if I sound crazy. This text is machine translated.</p>\n<p>Looking at the correct answer labels in the training data for this competition, the percentage of true(1) is clearly high.<br>\nIn the case of such data, if we assume that all the predicted results are true, the F1 would be higher.<br>\nFor example, if there are 10 data, 8 of which are true and the remaining 2 are false, and we predict that all the data are true, the RECALL would be 1 and the PRECISION would be 0.8.<br>\nAs a result, I think the F1 score, which is considered the evaluation value in this competition, would also be higher. Is this thinking incorrect?</p>",
  "messages": [
    {
      "id": "2228613",
      "postDate": "04/20/2023 16:44:45",
      "content": "<p>I am a beginner in machine learning and not very good at English. Sorry if I sound crazy. This text is machine translated.</p>\n<p>Looking at the correct answer labels in the training data for this competition, the percentage of true(1) is clearly high.<br>\nIn the case of such data, if we assume that all the predicted results are true, the F1 would be higher.<br>\nFor example, if there are 10 data, 8 of which are true and the remaining 2 are false, and we predict that all the data are true, the RECALL would be 1 and the PRECISION would be 0.8.<br>\nAs a result, I think the F1 score, which is considered the evaluation value in this competition, would also be higher. Is this thinking incorrect?</p>",
      "rawMarkdown": "I am a beginner in machine learning and not very good at English. Sorry if I sound crazy. This text is machine translated.\n\nLooking at the correct answer labels in the training data for this competition, the percentage of true(1) is clearly high.\nIn the case of such data, if we assume that all the predicted results are true, the F1 would be higher.\nFor example, if there are 10 data, 8 of which are true and the remaining 2 are false, and we predict that all the data are true, the RECALL would be 1 and the PRECISION would be 0.8.\nAs a result, I think the F1 score, which is considered the evaluation value in this competition, would also be higher. Is this thinking incorrect?",
      "votes": null
    },
    {
      "id": "2228807",
      "postDate": "04/20/2023 20:39:53",
      "content": "<p>The metric is F1 but with macro averaging, therefore the class imbalancedness is not taken into account. I struggled with this issue for a long time, but this thread cleared it up. <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/400819\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/400819</a></p>",
      "rawMarkdown": "The metric is F1 but with macro averaging, therefore the class imbalancedness is not taken into account. I struggled with this issue for a long time, but this thread cleared it up. https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/400819",
      "votes": null
    },
    {
      "id": "2229524",
      "postDate": "04/21/2023 12:53:55",
      "content": "<p>Thank you very much for taking the time to answer questions that have already been answered elsewhere! This is exactly what I wanted to know!</p>",
      "rawMarkdown": "Thank you very much for taking the time to answer questions that have already been answered elsewhere! This is exactly what I wanted to know!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228807,
      "author_name": "shinomoriaoshi",
      "author_url": "",
      "post_date": "04/20/2023 20:39:53",
      "content": "<p>The metric is F1 but with macro averaging, therefore the class imbalancedness is not taken into account. I struggled with this issue for a long time, but this thread cleared it up. <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/400819\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/400819</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2229524,
          "author_name": "shiki980",
          "author_url": "",
          "post_date": "04/21/2023 12:53:55",
          "content": "<p>Thank you very much for taking the time to answer questions that have already been answered elsewhere! This is exactly what I wanted to know!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2228613": "I am a beginner in machine learning and not very good at English. Sorry if I sound crazy. This text is machine translated.\n\nLooking at the correct answer labels in the training data for this competition, the percentage of true(1) is clearly high.\nIn the case of such data, if we assume that all the predicted results are true, the F1 would be higher.\nFor example, if there are 10 data, 8 of which are true and the remaining 2 are false, and we predict that all the data are true, the RECALL would be 1 and the PRECISION would be 0.8.\nAs a result, I think the F1 score, which is considered the evaluation value in this competition, would also be higher. Is this thinking incorrect?",
    "2228807": "The metric is F1 but with macro averaging, therefore the class imbalancedness is not taken into account. I struggled with this issue for a long time, but this thread cleared it up. https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/400819",
    "2229524": "Thank you very much for taking the time to answer questions that have already been answered elsewhere! This is exactly what I wanted to know!"
  },
  "source": "meta"
}