{
  "id": 92612,
  "title": "Fewer features have better performance in high TTF",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92612",
  "author_name": "",
  "post_date": "2019-05-18T14:19:51.632418300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<ol>\n<li>CV:1.9798,LB:1.424</li>\n<li>CV:1.9254,LB:1.477\nThe first CV-LB has 2000 features, and the second CV-LB has  200 features.\nI can not understand this strange condition. \nInteresting is 200 features model  hav low mae in high ttf , but have high mae in low ttf compare with 2000 features. \nSo, my conclusion is fewer features have more performance in high ttf.\nCan you share the relationship between the numbers of features and CV-LB?</li>\n</ol>",
  "messages": [
    {
      "id": "533139",
      "postDate": "05/18/2019 14:19:51",
      "content": "<ol>\n<li>CV:1.9798,LB:1.424</li>\n<li>CV:1.9254,LB:1.477\nThe first CV-LB has 2000 features, and the second CV-LB has  200 features.\nI can not understand this strange condition. \nInteresting is 200 features model  hav low mae in high ttf , but have high mae in low ttf compare with 2000 features. \nSo, my conclusion is fewer features have more performance in high ttf.\nCan you share the relationship between the numbers of features and CV-LB?</li>\n</ol>",
      "rawMarkdown": "1. CV:1.9798,LB:1.424\n2. CV:1.9254,LB:1.477\nThe first CV-LB has 2000 features, and the second CV-LB has  200 features.\nI can not understand this strange condition. \nInteresting is 200 features model  hav low mae in high ttf , but have high mae in low ttf compare with 2000 features. \nSo, my conclusion is fewer features have more performance in high ttf.\nCan you share the relationship between the numbers of features and CV-LB?",
      "votes": null
    },
    {
      "id": "533179",
      "postDate": "05/18/2019 15:46:04",
      "content": "<p><a href=\"/haonanyao\">@haonanyao</a>, thanks for your sharing. May I ask what ur CV scheme is?</p>",
      "rawMarkdown": "haonanyao, thanks for your sharing. May I ask what ur CV scheme is?",
      "votes": null
    },
    {
      "id": "533199",
      "postDate": "05/18/2019 16:51:37",
      "content": "<p>I don't think you should think about the direct correlation between the number of features to the performance in high / low ttf. The question is more about trusting the LB or your CV. Public LB is only 341 segments, while your CV has potentially up to 4000+ segments. And the distribution high / low ttf might be different. So it's possible that your CV and LB are not correlated depending on which CV you use.</p>",
      "rawMarkdown": "I don't think you should think about the direct correlation between the number of features to the performance in high / low ttf. The question is more about trusting the LB or your CV. Public LB is only 341 segments, while your CV has potentially up to 4000+ segments. And the distribution high / low ttf might be different. So it's possible that your CV and LB are not correlated depending on which CV you use.",
      "votes": null
    },
    {
      "id": "533369",
      "postDate": "05/19/2019 05:02:36",
      "content": "<p>Thanks for your reply.\nI think there is only one variable(numbers of features) in this discussing CV-LB, and the mean of train set is larger than the public. This is why I put forward my conclusion.\nSo, my immature model can have both good performances in high ttf and low ttf.\nIt is my fault, the problem is how to balance it when we don't know the private set.</p>",
      "rawMarkdown": "Thanks for your reply.\nI think there is only one variable(numbers of features) in this discussing CV-LB, and the mean of train set is larger than the public. This is why I put forward my conclusion.\nSo, my immature model can have both good performances in high ttf and low ttf.\nIt is my fault, the problem is how to balance it when we don't know the private set.",
      "votes": null
    },
    {
      "id": "533383",
      "postDate": "05/19/2019 05:41:40",
      "content": "<p>By the way, looking for team up or  friends, working together for the next kaggle competition</p>",
      "rawMarkdown": "By the way, looking for team up or  friends, working together for the next kaggle competition",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 533179,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "05/18/2019 15:46:04",
      "content": "<p><a href=\"/haonanyao\">@haonanyao</a>, thanks for your sharing. May I ask what ur CV scheme is?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 533199,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "05/18/2019 16:51:37",
      "content": "<p>I don't think you should think about the direct correlation between the number of features to the performance in high / low ttf. The question is more about trusting the LB or your CV. Public LB is only 341 segments, while your CV has potentially up to 4000+ segments. And the distribution high / low ttf might be different. So it's possible that your CV and LB are not correlated depending on which CV you use.</p>",
      "votes": null,
      "replies": [
        {
          "id": 533369,
          "author_name": "haonanyao",
          "author_url": "",
          "post_date": "05/19/2019 05:02:36",
          "content": "<p>Thanks for your reply.\nI think there is only one variable(numbers of features) in this discussing CV-LB, and the mean of train set is larger than the public. This is why I put forward my conclusion.\nSo, my immature model can have both good performances in high ttf and low ttf.\nIt is my fault, the problem is how to balance it when we don't know the private set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 533383,
      "author_name": "haonanyao",
      "author_url": "",
      "post_date": "05/19/2019 05:41:40",
      "content": "<p>By the way, looking for team up or  friends, working together for the next kaggle competition</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "533139": "1. CV:1.9798,LB:1.424\n2. CV:1.9254,LB:1.477\nThe first CV-LB has 2000 features, and the second CV-LB has  200 features.\nI can not understand this strange condition. \nInteresting is 200 features model  hav low mae in high ttf , but have high mae in low ttf compare with 2000 features. \nSo, my conclusion is fewer features have more performance in high ttf.\nCan you share the relationship between the numbers of features and CV-LB?",
    "533179": "haonanyao, thanks for your sharing. May I ask what ur CV scheme is?",
    "533199": "I don't think you should think about the direct correlation between the number of features to the performance in high / low ttf. The question is more about trusting the LB or your CV. Public LB is only 341 segments, while your CV has potentially up to 4000+ segments. And the distribution high / low ttf might be different. So it's possible that your CV and LB are not correlated depending on which CV you use.",
    "533369": "Thanks for your reply.\nI think there is only one variable(numbers of features) in this discussing CV-LB, and the mean of train set is larger than the public. This is why I put forward my conclusion.\nSo, my immature model can have both good performances in high ttf and low ttf.\nIt is my fault, the problem is how to balance it when we don't know the private set.",
    "533383": "By the way, looking for team up or  friends, working together for the next kaggle competition"
  },
  "source": "meta"
}