{
  "id": 53303,
  "title": "problem with submission format big difference cv/LB",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53303",
  "author_name": "",
  "post_date": "2018-03-29T08:10:04.028358Z",
  "votes": null,
  "comment_count": 16,
  "views": 0,
  "content": "<p>i face some problem when i submit, my socre LB is very low 0.9011 when my cv is 0.97. I think it's problem with the submission format. If some one could help me, i will be grateful :p</p>",
  "messages": [
    {
      "id": "305653",
      "postDate": "03/29/2018 08:10:04",
      "content": "<p>i face some problem when i submit, my socre LB is very low 0.9011 when my cv is 0.97. I think it's problem with the submission format. If some one could help me, i will be grateful :p</p>",
      "rawMarkdown": "i face some problem when i submit, my socre LB is very low 0.9011 when my cv is 0.97. I think it's problem with the submission format. If some one could help me, i will be grateful :p",
      "votes": null
    },
    {
      "id": "305657",
      "postDate": "03/29/2018 08:14:44",
      "content": "<p>Are you sure it's not an overfitting problem? Unless you are rounding to integers, I don't how the discrepancy could get that big due to format of the submission...</p>",
      "rawMarkdown": "Are you sure it's not an overfitting problem? Unless you are rounding to integers, I don't how the discrepancy could get that big due to format of the submission...",
      "votes": null
    },
    {
      "id": "305663",
      "postDate": "03/29/2018 08:22:35",
      "content": "<p>i don't think it is an overfitting problem this my train/validation results with : </p>\n\n<p>[175]   train-auc:0.974679  valid-auc:0.973527</p>\n\n<p>[180]   train-auc:0.974804  valid-auc:0.973635</p>\n\n<p>[185]   train-auc:0.974972  valid-auc:0.973767</p>\n\n<p>[190]   train-auc:0.97512   valid-auc:0.973881</p>\n\n<p>[195]   train-auc:0.975252  valid-auc:0.973981</p>\n\n<p>[199]   train-auc:0.975353  valid-auc:0.974065</p>\n\n<p>but when i submit i get 0.9011, i don't know but some has said that click_id must be in order 0,1,2,3,4 in the submission file. I did like he said and the result still the same :(</p>",
      "rawMarkdown": "i don't think it is an overfitting problem this my train/validation results with : \n\n[175]\ttrain-auc:0.974679\tvalid-auc:0.973527\n\n[180]\ttrain-auc:0.974804\tvalid-auc:0.973635\n\n[185]\ttrain-auc:0.974972\tvalid-auc:0.973767\n\n[190]\ttrain-auc:0.97512\tvalid-auc:0.973881\n\n[195]\ttrain-auc:0.975252\tvalid-auc:0.973981\n\n[199]\ttrain-auc:0.975353\tvalid-auc:0.974065\n\n\nbut when i submit i get 0.9011, i don't know but some has said that click_id must be in order 0,1,2,3,4 in the submission file. I did like he said and the result still the same :(",
      "votes": null
    },
    {
      "id": "306195",
      "postDate": "03/30/2018 02:48:07",
      "content": "<p>Maybe it's the problem of data distribution. I had the same problem at first. But it got better after adjusting the distribution.</p>",
      "rawMarkdown": "Maybe it's the problem of data distribution. I had the same problem at first. But it got better after adjusting the distribution.",
      "votes": null
    },
    {
      "id": "306199",
      "postDate": "03/30/2018 02:52:38",
      "content": "<p>If it's not an overfitting problem,  I guess there are something wrong in your feature engineering on testing data.</p>",
      "rawMarkdown": "If it's not an overfitting problem,  I guess there are something wrong in your feature engineering on testing data.",
      "votes": null
    },
    {
      "id": "306267",
      "postDate": "03/30/2018 06:30:32",
      "content": "<p>how do you create your submission file?  It looks like you don't get the right ids for your predictions.</p>",
      "rawMarkdown": "how do you create your submission file?  It looks like you don't get the right ids for your predictions.",
      "votes": null
    },
    {
      "id": "306349",
      "postDate": "03/30/2018 09:15:34",
      "content": "<p>No, it's not the problem, i have checked the ids before submission. I have used the validation strategy train on 07,08-11 data and validate on 09-11 between 4h-16h and i have the same problem a validation socre 0.974 and LB of 0.94 a huge difference </p>",
      "rawMarkdown": "No, it's not the problem, i have checked the ids before submission. I have used the validation strategy train on 07,08-11 data and validate on 09-11 between 4h-16h and i have the same problem a validation socre 0.974 and LB of 0.94 a huge difference",
      "votes": null
    },
    {
      "id": "306352",
      "postDate": "03/30/2018 09:17:52",
      "content": "<p>may be what i do to select train data is like : </p>\n\n<p>train_1 = train.loc[train.is_attributed == 1,]</p>\n\n<p>train_0 = train.loc[train.is_attributed == 0,]</p>\n\n<p>train_samp = train_0.sample(frac=0.1, random_state=42)</p>\n\n<p>train = pd.concat([train_samp,train_1])</p>\n\n<p>an simple undersampling startegy </p>",
      "rawMarkdown": "may be what i do to select train data is like : \n\ntrain_1 = train.loc[train.is_attributed == 1,]\n\ntrain_0 = train.loc[train.is_attributed == 0,]\n\ntrain_samp = train_0.sample(frac=0.1, random_state=42)\n\ntrain = pd.concat([train_samp,train_1])\n\nan simple undersampling startegy",
      "votes": null
    },
    {
      "id": "306356",
      "postDate": "03/30/2018 09:23:59",
      "content": "<p>i checked i did the same thing for both train and test, it's really confusing. I don't have the answer until now</p>",
      "rawMarkdown": "i checked i did the same thing for both train and test, it's really confusing. I don't have the answer until now",
      "votes": null
    },
    {
      "id": "306786",
      "postDate": "03/31/2018 02:30:17",
      "content": "<p>No comments on your updersampling strategy, but just be aware of the difference in hours between train &amp; test, bro.</p>",
      "rawMarkdown": "No comments on your updersampling strategy, but just be aware of the difference in hours between train &amp; test, bro.",
      "votes": null
    },
    {
      "id": "306877",
      "postDate": "03/31/2018 08:12:53",
      "content": "<p>You're ,not answering my question, hence I cannot help.  Good luck.</p>",
      "rawMarkdown": "You're ,not answering my question, hence I cannot help.  Good luck.",
      "votes": null
    },
    {
      "id": "306880",
      "postDate": "03/31/2018 08:22:03",
      "content": "<p>yeah i have answered your question that i checked the ids before the submission  so get the right ids. I found out what's the problem and it's a little bit strange. The problem is i created some features and these features do well in validation(0.975) but in the LB i get a huge difference 0.9011. It's not clear for me yet</p>",
      "rawMarkdown": "yeah i have answered your question that i checked the ids before the submission  so get the right ids. I found out what's the problem and it's a little bit strange. The problem is i created some features and these features do well in validation(0.975) but in the LB i get a huge difference 0.9011. It's not clear for me yet",
      "votes": null
    },
    {
      "id": "306882",
      "postDate": "03/31/2018 08:28:46",
      "content": "<p>I believe the comment you mentioned is not about ordering within submission, it is about the fact that ids are not sorted within test.csv, so you should read ids from test.csv, and not assume they are in order.</p>",
      "rawMarkdown": "I believe the comment you mentioned is not about ordering within submission, it is about the fact that ids are not sorted within test.csv, so you should read ids from test.csv, and not assume they are in order.",
      "votes": null
    },
    {
      "id": "306884",
      "postDate": "03/31/2018 08:37:53",
      "content": "<p>sorry i don't understand which comment you mean. </p>",
      "rawMarkdown": "sorry i don't understand which comment you mean.",
      "votes": null
    },
    {
      "id": "306885",
      "postDate": "03/31/2018 08:38:27",
      "content": "<p>sorry i don't understand which comment you mean.</p>",
      "rawMarkdown": "sorry i don't understand which comment you mean.",
      "votes": null
    },
    {
      "id": "306888",
      "postDate": "03/31/2018 08:42:10",
      "content": "<p>i am referring to this \"some has said that click_id must be in order\", original comment is <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/52354/why-different-auc-local-train-test-sets-and-leaderboard-set?forumMessageId=298395#post298395\">here</a> </p>",
      "rawMarkdown": "i am referring to this \"some has said that click_id must be in order\", original comment is [here][1] \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/52354/why-different-auc-local-train-test-sets-and-leaderboard-set?forumMessageId=298395#post298395",
      "votes": null
    },
    {
      "id": "306890",
      "postDate": "03/31/2018 08:43:12",
      "content": "<p>thanks bro</p>",
      "rawMarkdown": "thanks bro",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 305657,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "03/29/2018 08:14:44",
      "content": "<p>Are you sure it's not an overfitting problem? Unless you are rounding to integers, I don't how the discrepancy could get that big due to format of the submission...</p>",
      "votes": null,
      "replies": [
        {
          "id": 305663,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/29/2018 08:22:35",
          "content": "<p>i don't think it is an overfitting problem this my train/validation results with : </p>\n\n<p>[175]   train-auc:0.974679  valid-auc:0.973527</p>\n\n<p>[180]   train-auc:0.974804  valid-auc:0.973635</p>\n\n<p>[185]   train-auc:0.974972  valid-auc:0.973767</p>\n\n<p>[190]   train-auc:0.97512   valid-auc:0.973881</p>\n\n<p>[195]   train-auc:0.975252  valid-auc:0.973981</p>\n\n<p>[199]   train-auc:0.975353  valid-auc:0.974065</p>\n\n<p>but when i submit i get 0.9011, i don't know but some has said that click_id must be in order 0,1,2,3,4 in the submission file. I did like he said and the result still the same :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306882,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "03/31/2018 08:28:46",
          "content": "<p>I believe the comment you mentioned is not about ordering within submission, it is about the fact that ids are not sorted within test.csv, so you should read ids from test.csv, and not assume they are in order.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306884,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/31/2018 08:37:53",
          "content": "<p>sorry i don't understand which comment you mean. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306885,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/31/2018 08:38:27",
          "content": "<p>sorry i don't understand which comment you mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306888,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "03/31/2018 08:42:10",
          "content": "<p>i am referring to this \"some has said that click_id must be in order\", original comment is <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/52354/why-different-auc-local-train-test-sets-and-leaderboard-set?forumMessageId=298395#post298395\">here</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306890,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/31/2018 08:43:12",
          "content": "<p>thanks bro</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306195,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "03/30/2018 02:48:07",
      "content": "<p>Maybe it's the problem of data distribution. I had the same problem at first. But it got better after adjusting the distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 306352,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/30/2018 09:17:52",
          "content": "<p>may be what i do to select train data is like : </p>\n\n<p>train_1 = train.loc[train.is_attributed == 1,]</p>\n\n<p>train_0 = train.loc[train.is_attributed == 0,]</p>\n\n<p>train_samp = train_0.sample(frac=0.1, random_state=42)</p>\n\n<p>train = pd.concat([train_samp,train_1])</p>\n\n<p>an simple undersampling startegy </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306786,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "03/31/2018 02:30:17",
          "content": "<p>No comments on your updersampling strategy, but just be aware of the difference in hours between train &amp; test, bro.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306199,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "03/30/2018 02:52:38",
      "content": "<p>If it's not an overfitting problem,  I guess there are something wrong in your feature engineering on testing data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 306356,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/30/2018 09:23:59",
          "content": "<p>i checked i did the same thing for both train and test, it's really confusing. I don't have the answer until now</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306267,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/30/2018 06:30:32",
      "content": "<p>how do you create your submission file?  It looks like you don't get the right ids for your predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 306349,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/30/2018 09:15:34",
          "content": "<p>No, it's not the problem, i have checked the ids before submission. I have used the validation strategy train on 07,08-11 data and validate on 09-11 between 4h-16h and i have the same problem a validation socre 0.974 and LB of 0.94 a huge difference </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306877,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/31/2018 08:12:53",
          "content": "<p>You're ,not answering my question, hence I cannot help.  Good luck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306880,
          "author_name": "mohamed1",
          "author_url": "",
          "post_date": "03/31/2018 08:22:03",
          "content": "<p>yeah i have answered your question that i checked the ids before the submission  so get the right ids. I found out what's the problem and it's a little bit strange. The problem is i created some features and these features do well in validation(0.975) but in the LB i get a huge difference 0.9011. It's not clear for me yet</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "305653": "i face some problem when i submit, my socre LB is very low 0.9011 when my cv is 0.97. I think it's problem with the submission format. If some one could help me, i will be grateful :p",
    "305657": "Are you sure it's not an overfitting problem? Unless you are rounding to integers, I don't how the discrepancy could get that big due to format of the submission...",
    "305663": "i don't think it is an overfitting problem this my train/validation results with : \n\n[175]\ttrain-auc:0.974679\tvalid-auc:0.973527\n\n[180]\ttrain-auc:0.974804\tvalid-auc:0.973635\n\n[185]\ttrain-auc:0.974972\tvalid-auc:0.973767\n\n[190]\ttrain-auc:0.97512\tvalid-auc:0.973881\n\n[195]\ttrain-auc:0.975252\tvalid-auc:0.973981\n\n[199]\ttrain-auc:0.975353\tvalid-auc:0.974065\n\n\nbut when i submit i get 0.9011, i don't know but some has said that click_id must be in order 0,1,2,3,4 in the submission file. I did like he said and the result still the same :(",
    "306195": "Maybe it's the problem of data distribution. I had the same problem at first. But it got better after adjusting the distribution.",
    "306199": "If it's not an overfitting problem,  I guess there are something wrong in your feature engineering on testing data.",
    "306267": "how do you create your submission file?  It looks like you don't get the right ids for your predictions.",
    "306349": "No, it's not the problem, i have checked the ids before submission. I have used the validation strategy train on 07,08-11 data and validate on 09-11 between 4h-16h and i have the same problem a validation socre 0.974 and LB of 0.94 a huge difference",
    "306352": "may be what i do to select train data is like : \n\ntrain_1 = train.loc[train.is_attributed == 1,]\n\ntrain_0 = train.loc[train.is_attributed == 0,]\n\ntrain_samp = train_0.sample(frac=0.1, random_state=42)\n\ntrain = pd.concat([train_samp,train_1])\n\nan simple undersampling startegy",
    "306356": "i checked i did the same thing for both train and test, it's really confusing. I don't have the answer until now",
    "306786": "No comments on your updersampling strategy, but just be aware of the difference in hours between train &amp; test, bro.",
    "306877": "You're ,not answering my question, hence I cannot help.  Good luck.",
    "306880": "yeah i have answered your question that i checked the ids before the submission  so get the right ids. I found out what's the problem and it's a little bit strange. The problem is i created some features and these features do well in validation(0.975) but in the LB i get a huge difference 0.9011. It's not clear for me yet",
    "306882": "I believe the comment you mentioned is not about ordering within submission, it is about the fact that ids are not sorted within test.csv, so you should read ids from test.csv, and not assume they are in order.",
    "306884": "sorry i don't understand which comment you mean.",
    "306885": "sorry i don't understand which comment you mean.",
    "306888": "i am referring to this \"some has said that click_id must be in order\", original comment is [here][1] \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/52354/why-different-auc-local-train-test-sets-and-leaderboard-set?forumMessageId=298395#post298395",
    "306890": "thanks bro"
  },
  "source": "meta"
}