{
  "id": 404803,
  "title": "What is the effectiveness of using group kfold in this competition?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/404803",
  "author_name": "",
  "post_date": "2023-04-24T23:58:46.969813300Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Please give advice on the use of groupkfold, which is seen in multiple code notebooks. In this learning data, I plan to create a prediction model for the assigned quests for each of the three level_groups (04, 512, 1322). However, in each of the three training datasets, there is only one row of data per session_id. If we treat session_id as a group, why is it meaningful to use group kfold? I would appreciate any advice if there are any misunderstandings in my understanding.</p>",
  "messages": [
    {
      "id": "2234143",
      "postDate": "04/24/2023 23:58:46",
      "content": "<p>Please give advice on the use of groupkfold, which is seen in multiple code notebooks. In this learning data, I plan to create a prediction model for the assigned quests for each of the three level_groups (04, 512, 1322). However, in each of the three training datasets, there is only one row of data per session_id. If we treat session_id as a group, why is it meaningful to use group kfold? I would appreciate any advice if there are any misunderstandings in my understanding.</p>",
      "rawMarkdown": "Please give advice on the use of groupkfold, which is seen in multiple code notebooks. In this learning data, I plan to create a prediction model for the assigned quests for each of the three level_groups (04, 512, 1322). However, in each of the three training datasets, there is only one row of data per session_id. If we treat session_id as a group, why is it meaningful to use group kfold? I would appreciate any advice if there are any misunderstandings in my understanding.",
      "votes": null
    },
    {
      "id": "2234179",
      "postDate": "04/25/2023 01:36:54",
      "content": "<p>You are correct, we do not need Group KFold if we create a new train dataframe that only has 1 row of features per <code>session_id x level_group</code> combination. If you feature engineer inside your KFold, then you need to use Group KFold, but most notebooks including my own perform all feature engineering outside KFold and generate a new train dataframe that only has 1 row of features per <code>session_id x level_group</code> combination.</p>\n<p>If general, any time the group column (of the train dataframe) only has 1 row per group, then Group KFold is not needed.</p>",
      "rawMarkdown": "You are correct, we do not need Group KFold if we create a new train dataframe that only has 1 row of features per `session_id x level_group` combination. If you feature engineer inside your KFold, then you need to use Group KFold, but most notebooks including my own perform all feature engineering outside KFold and generate a new train dataframe that only has 1 row of features per `session_id x level_group` combination.\n\nIf general, any time the group column (of the train dataframe) only has 1 row per group, then Group KFold is not needed.",
      "votes": null
    },
    {
      "id": "2234212",
      "postDate": "04/25/2023 02:30:11",
      "content": "<p>Thank you for the easy-to-understand advice!</p>\n<p>I understand now. If there are no cases where the same user is included in the dataset with different IDs, then I think there would be no leak even with the normal k-fold validation.</p>\n<p>I posted this because my learning model's F1 score for validation was producing an unusually high score (such as 0.9 or 0.8), and I wondered if there was some kind of leak happening. (The score is calculated by averaging the F1 scores of the k-fold CV for quest1-18 assigned to each level group, and then calculating the average for each quest. For example, if the k-fold averages for q1-3 are 0.6, 0.7, and 0.8, the overall model score would be 0.7.)</p>\n<p>Perhaps the excessively high score is due to class label imbalance. I am creating a decision tree model, but I wonder if it makes sense to reduce the imbalance.</p>\n<p>I apologize for continuing to ask for advice, but if you have time, please let me know.</p>",
      "rawMarkdown": "Thank you for the easy-to-understand advice!\n\nI understand now. If there are no cases where the same user is included in the dataset with different IDs, then I think there would be no leak even with the normal k-fold validation.\n\nI posted this because my learning model's F1 score for validation was producing an unusually high score (such as 0.9 or 0.8), and I wondered if there was some kind of leak happening. (The score is calculated by averaging the F1 scores of the k-fold CV for quest1-18 assigned to each level group, and then calculating the average for each quest. For example, if the k-fold averages for q1-3 are 0.6, 0.7, and 0.8, the overall model score would be 0.7.)\n\nPerhaps the excessively high score is due to class label imbalance. I am creating a decision tree model, but I wonder if it makes sense to reduce the imbalance.\n\nI apologize for continuing to ask for advice, but if you have time, please let me know.",
      "votes": null
    },
    {
      "id": "2234223",
      "postDate": "04/25/2023 02:59:53",
      "content": "<p><a href=\"https://www.kaggle.com/kazumasatokaggle\" target=\"_blank\">@kazumasatokaggle</a> This competition metric is <strong>macro F1</strong>. Perhaps your high F1 score is because you are not using <code>sklearn.metrics.f1_score(true, pred, average='macro')</code> where we set <code>average='macro'</code></p>",
      "rawMarkdown": "kazumasatokaggle This competition metric is **macro F1**. Perhaps your high F1 score is because you are not using `sklearn.metrics.f1_score(true, pred, average='macro')` where we set `average='macro'`",
      "votes": null
    },
    {
      "id": "2234238",
      "postDate": "04/25/2023 03:43:38",
      "content": "<p>Thank you for your response.<br>\nIt was a beginner's mistake.<br>\nBased on the advice you gave me, I will continue to make improvements.</p>",
      "rawMarkdown": "Thank you for your response.\nIt was a beginner's mistake.\nBased on the advice you gave me, I will continue to make improvements.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2234179,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/25/2023 01:36:54",
      "content": "<p>You are correct, we do not need Group KFold if we create a new train dataframe that only has 1 row of features per <code>session_id x level_group</code> combination. If you feature engineer inside your KFold, then you need to use Group KFold, but most notebooks including my own perform all feature engineering outside KFold and generate a new train dataframe that only has 1 row of features per <code>session_id x level_group</code> combination.</p>\n<p>If general, any time the group column (of the train dataframe) only has 1 row per group, then Group KFold is not needed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2234212,
          "author_name": "kazumasatokaggle",
          "author_url": "",
          "post_date": "04/25/2023 02:30:11",
          "content": "<p>Thank you for the easy-to-understand advice!</p>\n<p>I understand now. If there are no cases where the same user is included in the dataset with different IDs, then I think there would be no leak even with the normal k-fold validation.</p>\n<p>I posted this because my learning model's F1 score for validation was producing an unusually high score (such as 0.9 or 0.8), and I wondered if there was some kind of leak happening. (The score is calculated by averaging the F1 scores of the k-fold CV for quest1-18 assigned to each level group, and then calculating the average for each quest. For example, if the k-fold averages for q1-3 are 0.6, 0.7, and 0.8, the overall model score would be 0.7.)</p>\n<p>Perhaps the excessively high score is due to class label imbalance. I am creating a decision tree model, but I wonder if it makes sense to reduce the imbalance.</p>\n<p>I apologize for continuing to ask for advice, but if you have time, please let me know.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2234223,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "04/25/2023 02:59:53",
              "content": "<p><a href=\"https://www.kaggle.com/kazumasatokaggle\" target=\"_blank\">@kazumasatokaggle</a> This competition metric is <strong>macro F1</strong>. Perhaps your high F1 score is because you are not using <code>sklearn.metrics.f1_score(true, pred, average='macro')</code> where we set <code>average='macro'</code></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2234238,
                  "author_name": "kazumasatokaggle",
                  "author_url": "",
                  "post_date": "04/25/2023 03:43:38",
                  "content": "<p>Thank you for your response.<br>\nIt was a beginner's mistake.<br>\nBased on the advice you gave me, I will continue to make improvements.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2234143": "Please give advice on the use of groupkfold, which is seen in multiple code notebooks. In this learning data, I plan to create a prediction model for the assigned quests for each of the three level_groups (04, 512, 1322). However, in each of the three training datasets, there is only one row of data per session_id. If we treat session_id as a group, why is it meaningful to use group kfold? I would appreciate any advice if there are any misunderstandings in my understanding.",
    "2234179": "You are correct, we do not need Group KFold if we create a new train dataframe that only has 1 row of features per `session_id x level_group` combination. If you feature engineer inside your KFold, then you need to use Group KFold, but most notebooks including my own perform all feature engineering outside KFold and generate a new train dataframe that only has 1 row of features per `session_id x level_group` combination.\n\nIf general, any time the group column (of the train dataframe) only has 1 row per group, then Group KFold is not needed.",
    "2234212": "Thank you for the easy-to-understand advice!\n\nI understand now. If there are no cases where the same user is included in the dataset with different IDs, then I think there would be no leak even with the normal k-fold validation.\n\nI posted this because my learning model's F1 score for validation was producing an unusually high score (such as 0.9 or 0.8), and I wondered if there was some kind of leak happening. (The score is calculated by averaging the F1 scores of the k-fold CV for quest1-18 assigned to each level group, and then calculating the average for each quest. For example, if the k-fold averages for q1-3 are 0.6, 0.7, and 0.8, the overall model score would be 0.7.)\n\nPerhaps the excessively high score is due to class label imbalance. I am creating a decision tree model, but I wonder if it makes sense to reduce the imbalance.\n\nI apologize for continuing to ask for advice, but if you have time, please let me know.",
    "2234223": "kazumasatokaggle This competition metric is **macro F1**. Perhaps your high F1 score is because you are not using `sklearn.metrics.f1_score(true, pred, average='macro')` where we set `average='macro'`",
    "2234238": "Thank you for your response.\nIt was a beginner's mistake.\nBased on the advice you gave me, I will continue to make improvements."
  },
  "source": "meta"
}