{
  "id": 410599,
  "title": "Use predictions for other questions as a features🤔",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/410599",
  "author_name": "",
  "post_date": "2023-05-15T23:15:58.558050Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi there!✋🙂<br>\nI think that it is a possibly a good idea to use predicted labels for seventeen other questions to to predict the last one 🤔 . For each question of course 🙂. <br>\nWe can use a labels file in such case to train a meta model and after we score test data with the main model we can use this meta model to correct and improve the final result. <br>\nDid somebody tried this? Could you share your insights?👀<br>\nThank you!😉🤔    </p>",
  "messages": [
    {
      "id": "2260828",
      "postDate": "05/15/2023 23:15:58",
      "content": "<p>Hi there!✋🙂<br>\nI think that it is a possibly a good idea to use predicted labels for seventeen other questions to to predict the last one 🤔 . For each question of course 🙂. <br>\nWe can use a labels file in such case to train a meta model and after we score test data with the main model we can use this meta model to correct and improve the final result. <br>\nDid somebody tried this? Could you share your insights?👀<br>\nThank you!😉🤔    </p>",
      "rawMarkdown": "Hi there!✋🙂\nI think that it is a possibly a good idea to use predicted labels for seventeen other questions to to predict the last one 🤔 . For each question of course 🙂. \nWe can use a labels file in such case to train a meta model and after we score test data with the main model we can use this meta model to correct and improve the final result. \nDid somebody tried this? Could you share your insights?👀\nThank you!😉🤔",
      "votes": null
    },
    {
      "id": "2260961",
      "postDate": "05/16/2023 02:48:11",
      "content": "<p>FYI, <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/392131\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/392131</a></p>",
      "rawMarkdown": "FYI, https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/392131",
      "votes": null
    },
    {
      "id": "2262297",
      "postDate": "05/16/2023 20:09:27",
      "content": "<p><a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> thank you! 🙏 I also have an idea to compute individual threshold for each question to convert probabilities to 0's and 1's. Is there a valid idea in your opinion? May be discussion tread for this topic also already exists?🤔</p>",
      "rawMarkdown": "shinomoriaoshi thank you! 🙏 I also have an idea to compute individual threshold for each question to convert probabilities to 0's and 1's. Is there a valid idea in your opinion? May be discussion tread for this topic also already exists?🤔",
      "votes": null
    },
    {
      "id": "2262471",
      "postDate": "05/17/2023 01:02:41",
      "content": "<p>I am not sure the topic has been discussed somewhere in the forum, but nice idea. In my opinion, we have less data to validate with this idea, making the validation a bit overfitting and less reliable. There are some Kagglers reporting that scores of individual questions improve but the overall score decreases.</p>",
      "rawMarkdown": "I am not sure the topic has been discussed somewhere in the forum, but nice idea. In my opinion, we have less data to validate with this idea, making the validation a bit overfitting and less reliable. There are some Kagglers reporting that scores of individual questions improve but the overall score decreases.",
      "votes": null
    },
    {
      "id": "2264582",
      "postDate": "05/18/2023 14:51:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ivanisaev\" target=\"_blank\">@ivanisaev</a>,</p>\n<p>When optimization is done by finding the best threshold <strong>for each question</strong> (<em>i.e.,</em> optimize locally), it's not guaranteed that the global optimal is found. And the key is that the final evaluation is done by deriving Macro-F1 on <strong>flattened</strong> predicting results and the groundtruth. That is, our goal is to balance out the <strong>overall</strong> precision and recall, not question-wise. Some forums have discussed this topic, which are shown as follows:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/391062\" target=\"_blank\">What happen with F1-score for all question?</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217\" target=\"_blank\">Different Thresholds for Different Question Models</a></li>\n</ol>\n<p>Hope this helps, thanks!</p>",
      "rawMarkdown": "Hi @ivanisaev,\n\nWhen optimization is done by finding the best threshold **for each question** (*i.e.,* optimize locally), it's not guaranteed that the global optimal is found. And the key is that the final evaluation is done by deriving Macro-F1 on **flattened** predicting results and the groundtruth. That is, our goal is to balance out the **overall** precision and recall, not question-wise. Some forums have discussed this topic, which are shown as follows:\n\n1. [What happen with F1-score for all question?](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/391062)\n2. [Different Thresholds for Different Question Models](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217)\n\nHope this helps, thanks!",
      "votes": null
    },
    {
      "id": "2264852",
      "postDate": "05/18/2023 19:22:08",
      "content": "<p><a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> thank you for useful thoughts and links! It helps!👍</p>",
      "rawMarkdown": "abaojiang thank you for useful thoughts and links! It helps!👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2260961,
      "author_name": "shinomoriaoshi",
      "author_url": "",
      "post_date": "05/16/2023 02:48:11",
      "content": "<p>FYI, <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/392131\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/392131</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2262297,
          "author_name": "ivanisaev",
          "author_url": "",
          "post_date": "05/16/2023 20:09:27",
          "content": "<p><a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> thank you! 🙏 I also have an idea to compute individual threshold for each question to convert probabilities to 0's and 1's. Is there a valid idea in your opinion? May be discussion tread for this topic also already exists?🤔</p>",
          "votes": null,
          "replies": [
            {
              "id": 2262471,
              "author_name": "shinomoriaoshi",
              "author_url": "",
              "post_date": "05/17/2023 01:02:41",
              "content": "<p>I am not sure the topic has been discussed somewhere in the forum, but nice idea. In my opinion, we have less data to validate with this idea, making the validation a bit overfitting and less reliable. There are some Kagglers reporting that scores of individual questions improve but the overall score decreases.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2264582,
              "author_name": "abaojiang",
              "author_url": "",
              "post_date": "05/18/2023 14:51:26",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ivanisaev\" target=\"_blank\">@ivanisaev</a>,</p>\n<p>When optimization is done by finding the best threshold <strong>for each question</strong> (<em>i.e.,</em> optimize locally), it's not guaranteed that the global optimal is found. And the key is that the final evaluation is done by deriving Macro-F1 on <strong>flattened</strong> predicting results and the groundtruth. That is, our goal is to balance out the <strong>overall</strong> precision and recall, not question-wise. Some forums have discussed this topic, which are shown as follows:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/391062\" target=\"_blank\">What happen with F1-score for all question?</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217\" target=\"_blank\">Different Thresholds for Different Question Models</a></li>\n</ol>\n<p>Hope this helps, thanks!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2264852,
                  "author_name": "ivanisaev",
                  "author_url": "",
                  "post_date": "05/18/2023 19:22:08",
                  "content": "<p><a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> thank you for useful thoughts and links! It helps!👍</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2260828": "Hi there!✋🙂\nI think that it is a possibly a good idea to use predicted labels for seventeen other questions to to predict the last one 🤔 . For each question of course 🙂. \nWe can use a labels file in such case to train a meta model and after we score test data with the main model we can use this meta model to correct and improve the final result. \nDid somebody tried this? Could you share your insights?👀\nThank you!😉🤔",
    "2260961": "FYI, https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/392131",
    "2262297": "shinomoriaoshi thank you! 🙏 I also have an idea to compute individual threshold for each question to convert probabilities to 0's and 1's. Is there a valid idea in your opinion? May be discussion tread for this topic also already exists?🤔",
    "2262471": "I am not sure the topic has been discussed somewhere in the forum, but nice idea. In my opinion, we have less data to validate with this idea, making the validation a bit overfitting and less reliable. There are some Kagglers reporting that scores of individual questions improve but the overall score decreases.",
    "2264582": "Hi @ivanisaev,\n\nWhen optimization is done by finding the best threshold **for each question** (*i.e.,* optimize locally), it's not guaranteed that the global optimal is found. And the key is that the final evaluation is done by deriving Macro-F1 on **flattened** predicting results and the groundtruth. That is, our goal is to balance out the **overall** precision and recall, not question-wise. Some forums have discussed this topic, which are shown as follows:\n\n1. [What happen with F1-score for all question?](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/391062)\n2. [Different Thresholds for Different Question Models](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217)\n\nHope this helps, thanks!",
    "2264852": "abaojiang thank you for useful thoughts and links! It helps!👍"
  },
  "source": "meta"
}