{
  "id": 92197,
  "title": "Please Help me out . . .",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92197",
  "author_name": "Monster",
  "post_date": "2019-05-14T10:03:53.954000",
  "votes": 4,
  "comment_count": 20,
  "views": 0,
  "content": "<p>when i using 40 features my CV  = 2.067 and PL = 1.471\nwhen i am excluding  many features  and using only 4 features from those \nmy  CV  = 2.0102 and PL = 1.527\ni cant find which one to trust plzz give some suggestion . . . .\n<strong>THANKS IN ADVANCE</strong></p>",
  "messages": [
    {
      "id": 531107,
      "postDate": "2019-05-14T10:12:42.650Z",
      "content": "<p>Hint: look at the features which are different between train and test set :)</p>",
      "rawMarkdown": "Hint: look at the features which are different between train and test set :)",
      "votes": 12,
      "replies": [
        {
          "id": 531354,
          "postDate": "2019-05-14T17:54:26.040Z",
          "content": "<p>At what point do you define them as different? Can you just do something like calculating a distance measure based on their train/test histograms (and looking at them), or are there more sophisticated techniques?</p>",
          "rawMarkdown": "At what point do you define them as different? Can you just do something like calculating a distance measure based on their train/test histograms (and looking at them), or are there more sophisticated techniques?"
        },
        {
          "id": 531376,
          "postDate": "2019-05-14T18:55:54.490Z",
          "content": "<p>I am afraid there is no clear cut. You could look at the KS statistics. This is an implementation in python:\n<a href=\"https://docs.scipy.org/doc/scipy-0.15.1/reference/generated/scipy.stats.ks_2samp.html\">https://docs.scipy.org/doc/scipy-0.15.1/reference/generated/scipy.stats.ks_2samp.html</a></p>",
          "rawMarkdown": "I am afraid there is no clear cut. You could look at the KS statistics. This is an implementation in python:\nhttps://docs.scipy.org/doc/scipy-0.15.1/reference/generated/scipy.stats.ks_2samp.html",
          "votes": 3
        },
        {
          "id": 531378,
          "postDate": "2019-05-14T19:01:05.450Z",
          "content": "<p>thanks sir <a href=\"/stecasasso\">@stecasasso</a> </p>",
          "rawMarkdown": "thanks sir @stecasasso "
        },
        {
          "id": 531403,
          "postDate": "2019-05-14T20:30:21.357Z",
          "content": "<p>Thanks, that seems to be what i was looking for ;)</p>",
          "rawMarkdown": "Thanks, that seems to be what i was looking for ;)"
        },
        {
          "id": 531676,
          "postDate": "2019-05-15T10:25:03.657Z",
          "content": "<p>You can also build a classifier that distinguishes train vs test and drop important features. Here is a good explanation <a href=\"https://towardsdatascience.com/how-dis-similar-are-my-train-and-test-data-56af3923de9b\">https://towardsdatascience.com/how-dis-similar-are-my-train-and-test-data-56af3923de9b</a></p>",
          "rawMarkdown": "You can also build a classifier that distinguishes train vs test and drop important features. Here is a good explanation https://towardsdatascience.com/how-dis-similar-are-my-train-and-test-data-56af3923de9b",
          "votes": 4
        },
        {
          "id": 532247,
          "postDate": "2019-05-16T14:04:36.447Z",
          "content": "<p>I tried it but gave me worse results. Anyway, it's a very nice idea </p>",
          "rawMarkdown": "I tried it but gave me worse results. Anyway, it's a very nice idea "
        }
      ]
    },
    {
      "id": 531138,
      "postDate": "2019-05-14T11:24:07.670Z",
      "content": "<p>Feature selection can leak if not done properly. If that is the case, your CV is over optimistic without real improvement in the model.</p>",
      "rawMarkdown": "Feature selection can leak if not done properly. If that is the case, your CV is over optimistic without real improvement in the model.",
      "votes": 3
    },
    {
      "id": 531137,
      "postDate": "2019-05-14T11:20:54.713Z",
      "content": "<p>You can use a combined CV weighted by the number of instances:</p>\n\n<p><code>CCV = (CV*Ntrain+LBscore*Npublic)/(Ntrain+Npublic)</code></p>",
      "rawMarkdown": "You can use a combined CV weighted by the number of instances:\n\n`CCV = (CV*Ntrain+LBscore*Npublic)/(Ntrain+Npublic)`",
      "votes": 3,
      "replies": [
        {
          "id": 531372,
          "postDate": "2019-05-14T18:46:08.117Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks"
        },
        {
          "id": 531552,
          "postDate": "2019-05-15T05:54:50.710Z",
          "content": "<p>Hi <a href=\"/sionek\">@sionek</a> , can I know what is Ntrain and Npublic in the above equation. I believe Ntrain means no. of models trained and Npublic means no.of people in LB. Correct me if I am worng. Thaks.</p>",
          "rawMarkdown": "Hi @sionek , can I know what is Ntrain and Npublic in the above equation. I believe Ntrain means no. of models trained and Npublic means no.of people in LB. Correct me if I am worng. Thaks."
        },
        {
          "id": 531575,
          "postDate": "2019-05-15T06:35:11.627Z",
          "content": "<p>Not exactly ;) I am a chemist, so I would never propose a formula containing a sum of values of different units/sense.  Npublic is a number of segments/files in the public part of test directory (or a number of \"public\" 13% instances in the submission file), and Ntrain is a number of segments after the cut of the train dataset on pieces containing the same number of rows (150000) as each test file (or a number of instances in the training file containing your features), i.e. <code>Npublic = 341 , Ntrain = 4194</code></p>",
          "rawMarkdown": "Not exactly ;) I am a chemist, so I would never propose a formula containing a sum of values of different units/sense.  Npublic is a number of segments/files in the public part of test directory (or a number of \"public\" 13% instances in the submission file), and Ntrain is a number of segments after the cut of the train dataset on pieces containing the same number of rows (150000) as each test file (or a number of instances in the training file containing your features), i.e. `Npublic = 341 , Ntrain = 4194`",
          "votes": 3
        },
        {
          "id": 531603,
          "postDate": "2019-05-15T07:35:53.613Z",
          "content": "<p>Yeah, now the equation makes sense to me. Thanks for clarification. You are calculating the weighted score of train and public set based on number of instances. But can it really help to judge the model?</p>",
          "rawMarkdown": "Yeah, now the equation makes sense to me. Thanks for clarification. You are calculating the weighted score of train and public set based on number of instances. But can it really help to judge the model?"
        },
        {
          "id": 531612,
          "postDate": "2019-05-15T07:45:20.857Z",
          "content": "<p>To paraphrase one of the Polish presidents and CPMP: \"In this competition, public LB score is useful and even useless\" ;)</p>",
          "rawMarkdown": "To paraphrase one of the Polish presidents and CPMP: \"In this competition, public LB score is useful and even useless\" ;)",
          "votes": 3
        },
        {
          "id": 531670,
          "postDate": "2019-05-15T10:11:20.963Z",
          "content": "<p><a href=\"/sionek\">@sionek</a> </p>\n\n<p>Not sure I am happy to be in the same bag as that Polish president ;)</p>\n\n<p>But you're right, I said useless, then said useful as one fold, which is precisely what you describe I think.</p>",
          "rawMarkdown": "@sionek \n\nNot sure I am happy to be in the same bag as that Polish president ;)\n\nBut you're right, I said useless, then said useful as one fold, which is precisely what you describe I think.",
          "votes": 2
        }
      ]
    },
    {
      "id": 531229,
      "postDate": "2019-05-14T14:11:31.560Z",
      "content": "<p>You have two submissions so why not use both - traditionally quite a few people use their best Cross-Validation submission and their best LB submission to increase their chances. </p>",
      "rawMarkdown": "You have two submissions so why not use both - traditionally quite a few people use their best Cross-Validation submission and their best LB submission to increase their chances. ",
      "votes": 2,
      "replies": [
        {
          "id": 531375,
          "postDate": "2019-05-14T18:47:13.650Z",
          "content": "<p>thanks sir</p>",
          "rawMarkdown": "thanks sir",
          "votes": 1
        },
        {
          "id": 531787,
          "postDate": "2019-05-15T14:40:57.080Z",
          "content": "<p>Agreed, a good submission strategy can make a difference</p>",
          "rawMarkdown": "Agreed, a good submission strategy can make a difference"
        }
      ]
    },
    {
      "id": 531099,
      "postDate": "2019-05-14T10:03:53.953Z",
      "content": "<p>when i using 40 features my CV  = 2.067 and PL = 1.471\nwhen i am excluding  many features  and using only 4 features from those \nmy  CV  = 2.0102 and PL = 1.527\ni cant find which one to trust plzz give some suggestion . . . .\n<strong>THANKS IN ADVANCE</strong></p>",
      "rawMarkdown": "when i using 40 features my CV  = 2.067 and PL = 1.471\nwhen i am excluding  many features  and using only 4 features from those \nmy  CV  = 2.0102 and PL = 1.527\ni cant find which one to trust plzz give some suggestion . . . .\n**THANKS IN ADVANCE**",
      "votes": 2
    },
    {
      "id": 532902,
      "postDate": "2019-05-18T02:28:28.223Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 533401,
          "postDate": "2019-05-19T06:21:11.533Z",
          "content": "<p>just by looking CV score  , feature importance and feature correlation with other features</p>",
          "rawMarkdown": "just by looking CV score  , feature importance and feature correlation with other features"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 531107,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2019-05-14T10:12:42.650000",
      "content": "<p>Hint: look at the features which are different between train and test set :)</p>",
      "votes": 12,
      "replies": [
        {
          "id": 531354,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-14T17:54:26.040000",
          "content": "<p>At what point do you define them as different? Can you just do something like calculating a distance measure based on their train/test histograms (and looking at them), or are there more sophisticated techniques?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531376,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2019-05-14T18:55:54.490000",
          "content": "<p>I am afraid there is no clear cut. You could look at the KS statistics. This is an implementation in python:\n<a href=\"https://docs.scipy.org/doc/scipy-0.15.1/reference/generated/scipy.stats.ks_2samp.html\">https://docs.scipy.org/doc/scipy-0.15.1/reference/generated/scipy.stats.ks_2samp.html</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 531378,
          "author_name": "Monster",
          "author_url": "",
          "post_date": "2019-05-14T19:01:05.450000",
          "content": "<p>thanks sir <a href=\"/stecasasso\">@stecasasso</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531403,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-14T20:30:21.357000",
          "content": "<p>Thanks, that seems to be what i was looking for ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531676,
          "author_name": "ivan",
          "author_url": "",
          "post_date": "2019-05-15T10:25:03.657000",
          "content": "<p>You can also build a classifier that distinguishes train vs test and drop important features. Here is a good explanation <a href=\"https://towardsdatascience.com/how-dis-similar-are-my-train-and-test-data-56af3923de9b\">https://towardsdatascience.com/how-dis-similar-are-my-train-and-test-data-56af3923de9b</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 532247,
          "author_name": "Matias Thayer",
          "author_url": "",
          "post_date": "2019-05-16T14:04:36.447000",
          "content": "<p>I tried it but gave me worse results. Anyway, it's a very nice idea </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 531138,
      "author_name": "Amjad",
      "author_url": "",
      "post_date": "2019-05-14T11:24:07.670000",
      "content": "<p>Feature selection can leak if not done properly. If that is the case, your CV is over optimistic without real improvement in the model.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 531137,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2019-05-14T11:20:54.713000",
      "content": "<p>You can use a combined CV weighted by the number of instances:</p>\n\n<p><code>CCV = (CV*Ntrain+LBscore*Npublic)/(Ntrain+Npublic)</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 531372,
          "author_name": "Monster",
          "author_url": "",
          "post_date": "2019-05-14T18:46:08.117000",
          "content": "<p>thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531552,
          "author_name": "Guna Sekhar",
          "author_url": "",
          "post_date": "2019-05-15T05:54:50.710000",
          "content": "<p>Hi <a href=\"/sionek\">@sionek</a> , can I know what is Ntrain and Npublic in the above equation. I believe Ntrain means no. of models trained and Npublic means no.of people in LB. Correct me if I am worng. Thaks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531575,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2019-05-15T06:35:11.627000",
          "content": "<p>Not exactly ;) I am a chemist, so I would never propose a formula containing a sum of values of different units/sense.  Npublic is a number of segments/files in the public part of test directory (or a number of \"public\" 13% instances in the submission file), and Ntrain is a number of segments after the cut of the train dataset on pieces containing the same number of rows (150000) as each test file (or a number of instances in the training file containing your features), i.e. <code>Npublic = 341 , Ntrain = 4194</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 531603,
          "author_name": "Guna Sekhar",
          "author_url": "",
          "post_date": "2019-05-15T07:35:53.613000",
          "content": "<p>Yeah, now the equation makes sense to me. Thanks for clarification. You are calculating the weighted score of train and public set based on number of instances. But can it really help to judge the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531612,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2019-05-15T07:45:20.857000",
          "content": "<p>To paraphrase one of the Polish presidents and CPMP: \"In this competition, public LB score is useful and even useless\" ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 531670,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-15T10:11:20.963000",
          "content": "<p><a href=\"/sionek\">@sionek</a> </p>\n\n<p>Not sure I am happy to be in the same bag as that Polish president ;)</p>\n\n<p>But you're right, I said useless, then said useful as one fold, which is precisely what you describe I think.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 531229,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-05-14T14:11:31.560000",
      "content": "<p>You have two submissions so why not use both - traditionally quite a few people use their best Cross-Validation submission and their best LB submission to increase their chances. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 531375,
          "author_name": "Monster",
          "author_url": "",
          "post_date": "2019-05-14T18:47:13.650000",
          "content": "<p>thanks sir</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531787,
          "author_name": "Eric R",
          "author_url": "",
          "post_date": "2019-05-15T14:40:57.080000",
          "content": "<p>Agreed, a good submission strategy can make a difference</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 532902,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-18T02:28:28.223000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 533401,
          "author_name": "Monster",
          "author_url": "",
          "post_date": "2019-05-19T06:21:11.533000",
          "content": "<p>just by looking CV score  , feature importance and feature correlation with other features</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "531107": "Hint: look at the features which are different between train and test set :)",
    "531138": "Feature selection can leak if not done properly. If that is the case, your CV is over optimistic without real improvement in the model.",
    "531137": "You can use a combined CV weighted by the number of instances:\n\n`CCV = (CV*Ntrain+LBscore*Npublic)/(Ntrain+Npublic)`",
    "531229": "You have two submissions so why not use both - traditionally quite a few people use their best Cross-Validation submission and their best LB submission to increase their chances. ",
    "531099": "when i using 40 features my CV  = 2.067 and PL = 1.471\nwhen i am excluding  many features  and using only 4 features from those \nmy  CV  = 2.0102 and PL = 1.527\ni cant find which one to trust plzz give some suggestion . . . .\n**THANKS IN ADVANCE**",
    "532902": ""
  }
}