{
  "id": 412279,
  "title": "The number of features used in tree-based models",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/412279",
  "author_name": "",
  "post_date": "2023-05-23T04:24:11.591081900Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<pre><code>group 0-4: about 400\ngroup 5-12: about 800\ngroup 13-22: about 1000\n</code></pre>\n<p>Does generating more features help you get better result?</p>",
  "messages": [
    {
      "id": "2270181",
      "postDate": "05/23/2023 04:24:11",
      "content": "<pre><code>group 0-4: about 400\ngroup 5-12: about 800\ngroup 13-22: about 1000\n</code></pre>\n<p>Does generating more features help you get better result?</p>",
      "rawMarkdown": "```\ngroup 0-4: about 400\ngroup 5-12: about 800\ngroup 13-22: about 1000\n```\nDoes generating more features help you get better result?",
      "votes": null
    },
    {
      "id": "2270326",
      "postDate": "05/23/2023 05:55:41",
      "content": "<p>According to my experience, yes. Generating more features and then eliminating them to a marginally smaller size (very close to 400/800/1000) help me get a better score. </p>\n<p>But your results are still better than mine and most of us, so maybe your strategy is better after all.  </p>",
      "rawMarkdown": "According to my experience, yes. Generating more features and then eliminating them to a marginally smaller size (very close to 400/800/1000) help me get a better score. \n\nBut your results are still better than mine and most of us, so maybe your strategy is better after all.",
      "votes": null
    },
    {
      "id": "2270340",
      "postDate": "05/23/2023 06:09:26",
      "content": "<p>I have not yet found the way to eliminate redundant features. The CV-LB makes me confused and 704 is a lucky result with CV only 0.70008.</p>",
      "rawMarkdown": "I have not yet found the way to eliminate redundant features. The CV-LB makes me confused and 704 is a lucky result with CV only 0.70008.",
      "votes": null
    },
    {
      "id": "2270353",
      "postDate": "05/23/2023 06:22:03",
      "content": "<p>I feel the data is kind of noisy. Once you pass the 0.700 line and start to eliminate redundant features, CV-LB correlation becomes very unstable. High LB scores are not necessarily tied to high CV scores, could only mean some lucky elimination split.</p>\n<p>Looks like one of the overfitting competitions for me, maybe expect a huge shake-up for private LB. </p>",
      "rawMarkdown": "I feel the data is kind of noisy. Once you pass the 0.700 line and start to eliminate redundant features, CV-LB correlation becomes very unstable. High LB scores are not necessarily tied to high CV scores, could only mean some lucky elimination split.\n\nLooks like one of the overfitting competitions for me, maybe expect a huge shake-up for private LB.",
      "votes": null
    },
    {
      "id": "2270391",
      "postDate": "05/23/2023 06:45:06",
      "content": "<p>For me it's 300, 600, 900. Can't seem to get pass CV0.700 to see the world you guys are seeing. 👀</p>",
      "rawMarkdown": "For me it's 300, 600, 900. Can't seem to get pass CV0.700 to see the world you guys are seeing. 👀",
      "votes": null
    },
    {
      "id": "2270806",
      "postDate": "05/23/2023 12:49:20",
      "content": "<p>After I test my pipeline in local, I guess the lb has some leak issue, maybe public test set includes some data in train data set.</p>",
      "rawMarkdown": "After I test my pipeline in local, I guess the lb has some leak issue, maybe public test set includes some data in train data set.",
      "votes": null
    },
    {
      "id": "2270816",
      "postDate": "05/23/2023 12:54:21",
      "content": "<p>Thank you for your sharing.<br>\nbut how did you know the leakage?</p>",
      "rawMarkdown": "Thank you for your sharing.\nbut how did you know the leakage?",
      "votes": null
    },
    {
      "id": "2270840",
      "postDate": "05/23/2023 13:09:29",
      "content": "<p>a overfitting model get high lb score</p>",
      "rawMarkdown": "a overfitting model get high lb score",
      "votes": null
    },
    {
      "id": "2271084",
      "postDate": "05/23/2023 16:01:34",
      "content": "<p>Not related, but would you mind sharing if you're using one or multiple model types?</p>",
      "rawMarkdown": "Not related, but would you mind sharing if you're using one or multiple model types?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2270326,
      "author_name": "ataraxian",
      "author_url": "",
      "post_date": "05/23/2023 05:55:41",
      "content": "<p>According to my experience, yes. Generating more features and then eliminating them to a marginally smaller size (very close to 400/800/1000) help me get a better score. </p>\n<p>But your results are still better than mine and most of us, so maybe your strategy is better after all.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2270340,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "05/23/2023 06:09:26",
          "content": "<p>I have not yet found the way to eliminate redundant features. The CV-LB makes me confused and 704 is a lucky result with CV only 0.70008.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2270353,
              "author_name": "ataraxian",
              "author_url": "",
              "post_date": "05/23/2023 06:22:03",
              "content": "<p>I feel the data is kind of noisy. Once you pass the 0.700 line and start to eliminate redundant features, CV-LB correlation becomes very unstable. High LB scores are not necessarily tied to high CV scores, could only mean some lucky elimination split.</p>\n<p>Looks like one of the overfitting competitions for me, maybe expect a huge shake-up for private LB. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2270806,
                  "author_name": "jimmyliao86204",
                  "author_url": "",
                  "post_date": "05/23/2023 12:49:20",
                  "content": "<p>After I test my pipeline in local, I guess the lb has some leak issue, maybe public test set includes some data in train data set.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2270816,
                      "author_name": "taruto1215",
                      "author_url": "",
                      "post_date": "05/23/2023 12:54:21",
                      "content": "<p>Thank you for your sharing.<br>\nbut how did you know the leakage?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2270840,
                          "author_name": "jimmyliao86204",
                          "author_url": "",
                          "post_date": "05/23/2023 13:09:29",
                          "content": "<p>a overfitting model get high lb score</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2270391,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "05/23/2023 06:45:06",
      "content": "<p>For me it's 300, 600, 900. Can't seem to get pass CV0.700 to see the world you guys are seeing. 👀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2271084,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "05/23/2023 16:01:34",
      "content": "<p>Not related, but would you mind sharing if you're using one or multiple model types?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2270181": "```\ngroup 0-4: about 400\ngroup 5-12: about 800\ngroup 13-22: about 1000\n```\nDoes generating more features help you get better result?",
    "2270326": "According to my experience, yes. Generating more features and then eliminating them to a marginally smaller size (very close to 400/800/1000) help me get a better score. \n\nBut your results are still better than mine and most of us, so maybe your strategy is better after all.",
    "2270340": "I have not yet found the way to eliminate redundant features. The CV-LB makes me confused and 704 is a lucky result with CV only 0.70008.",
    "2270353": "I feel the data is kind of noisy. Once you pass the 0.700 line and start to eliminate redundant features, CV-LB correlation becomes very unstable. High LB scores are not necessarily tied to high CV scores, could only mean some lucky elimination split.\n\nLooks like one of the overfitting competitions for me, maybe expect a huge shake-up for private LB.",
    "2270391": "For me it's 300, 600, 900. Can't seem to get pass CV0.700 to see the world you guys are seeing. 👀",
    "2270806": "After I test my pipeline in local, I guess the lb has some leak issue, maybe public test set includes some data in train data set.",
    "2270816": "Thank you for your sharing.\nbut how did you know the leakage?",
    "2270840": "a overfitting model get high lb score",
    "2271084": "Not related, but would you mind sharing if you're using one or multiple model types?"
  },
  "source": "meta"
}