{
  "id": 45921,
  "title": "How to validate if one didn't use old members.csv?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45921",
  "author_name": "",
  "post_date": "2017-12-18T00:42:55.589186Z",
  "votes": 7,
  "comment_count": 17,
  "views": 0,
  "content": "<p>As mentioned <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363#250087\">here</a>, the old members.csv is not allowed to use.</p>\n\n<p>Based on the current top Private LB scores, I assume that some teams still selected submissions that used the expiration dates in the old members.csv (Our best score with the expiration date was 0.07817 on Private LB and 0.07899 on Public LB).</p>\n\n<p>My question is how the organizer are going to validate if one did not use such data sets that are not permitted.</p>\n\n<p>Please advise.</p>\n\n<p>Thanks in advance.</p>",
  "messages": [
    {
      "id": "259216",
      "postDate": "12/18/2017 00:42:55",
      "content": "<p>As mentioned <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363#250087\">here</a>, the old members.csv is not allowed to use.</p>\n\n<p>Based on the current top Private LB scores, I assume that some teams still selected submissions that used the expiration dates in the old members.csv (Our best score with the expiration date was 0.07817 on Private LB and 0.07899 on Public LB).</p>\n\n<p>My question is how the organizer are going to validate if one did not use such data sets that are not permitted.</p>\n\n<p>Please advise.</p>\n\n<p>Thanks in advance.</p>",
      "rawMarkdown": "As mentioned [here][1], the old members.csv is not allowed to use.\n\nBased on the current top Private LB scores, I assume that some teams still selected submissions that used the expiration dates in the old members.csv (Our best score with the expiration date was 0.07817 on Private LB and 0.07899 on Public LB).\n\nMy question is how the organizer are going to validate if one did not use such data sets that are not permitted.\n\nPlease advise.\n\nThanks in advance.\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363#250087",
      "votes": null
    },
    {
      "id": "259228",
      "postDate": "12/18/2017 01:06:15",
      "content": "<p>One idea is to have the top teams post screenshots of how they now score with their final best model WITH the members1.csv added back into their model.  As you probably know, when a Kaggle competition ends users can still keep making submissions to view what their LB scores would have been.</p>\n\n<p>Presumably if data from members1 was not used in their model, then they should see a huge jump in LB score by adding it back in.  I believe .015-.025 can be expected.  So a current .09 score should achieve at LEAST a .075 score with members1 csv data.  If the teams can show proof that their model can achieve a &gt;.015 LB gain by adding member1 back and submitting it, then posting the screenshot, then I would say they deserve their spot.  Selfishly, I just want to see how low the scoreboard can go :)  I'm betting I can score below .055 if I use member1 data and spend time re-tuning everything.</p>\n\n<p>Just a thought, upvote if you like the idea.</p>",
      "rawMarkdown": "One idea is to have the top teams post screenshots of how they now score with their final best model WITH the members1.csv added back into their model.  As you probably know, when a Kaggle competition ends users can still keep making submissions to view what their LB scores would have been.\n\nPresumably if data from members1 was not used in their model, then they should see a huge jump in LB score by adding it back in.  I believe .015-.025 can be expected.  So a current .09 score should achieve at LEAST a .075 score with members1 csv data.  If the teams can show proof that their model can achieve a &gt;.015 LB gain by adding member1 back and submitting it, then posting the screenshot, then I would say they deserve their spot.  Selfishly, I just want to see how low the scoreboard can go :)  I'm betting I can score below .055 if I use member1 data and spend time re-tuning everything.\n\nJust a thought, upvote if you like the idea.",
      "votes": null
    },
    {
      "id": "259230",
      "postDate": "12/18/2017 01:14:48",
      "content": "<p>Thanks @Bryan for the idea. I like the idea. It might be little inconvenient for teams who haven't used it to add it and rerun the pipeline though. </p>\n\n<p>By the way, is your Private LB score with members? If so, what's your score without old members? If not, it's really impressive and hats off. Either way, great job and congrats! :)</p>",
      "rawMarkdown": "Thanks @Bryan for the idea. I like the idea. It might be little inconvenient for teams who haven't used it to add it and rerun the pipeline though. \n\nBy the way, is your Private LB score with members? If so, what's your score without old members? If not, it's really impressive and hats off. Either way, great job and congrats! :)",
      "votes": null
    },
    {
      "id": "259232",
      "postDate": "12/18/2017 01:17:19",
      "content": "<p>It includes member3 data, not member1.  I'll spend some time tomorrow re-modeling with members1 and will post a screenshot of what my model can achieve with it.  \"Proof is in the pudding\"</p>",
      "rawMarkdown": "It includes member3 data, not member1.  I'll spend some time tomorrow re-modeling with members1 and will post a screenshot of what my model can achieve with it.  \"Proof is in the pudding\"",
      "votes": null
    },
    {
      "id": "259233",
      "postDate": "12/18/2017 01:22:56",
      "content": "<p>Awesome! Congrats for the 1st place! Looking forward to your solution. :)</p>",
      "rawMarkdown": "Awesome! Congrats for the 1st place! Looking forward to your solution. :)",
      "votes": null
    },
    {
      "id": "259237",
      "postDate": "12/18/2017 01:27:12",
      "content": "<p>Also you can't download members_v1 anymore... for teams who joined this competition after data refresh, they can't get that file.</p>",
      "rawMarkdown": "Also you can't download members_v1 anymore... for teams who joined this competition after data refresh, they can't get that file.",
      "votes": null
    },
    {
      "id": "259242",
      "postDate": "12/18/2017 01:38:25",
      "content": "<p>I was wondering what is the secret to developing the 0.07 score, and here is the answer! The old members.csv...!\nWe all notice that there is a huge score drop between public leaderboard and private leaderboard for some rankers.\nDid it contain the expiration date of customers?</p>\n\n<p>I also assume that Kaggle scoring system cannot find who exploit the features from that data or not. <br>\nThe most fairness solution is to make available the old members1.csv and give time to re-submit it. (although it might be troublesome to most competitors)   </p>",
      "rawMarkdown": "I was wondering what is the secret to developing the 0.07 score, and here is the answer! The old members.csv...!\nWe all notice that there is a huge score drop between public leaderboard and private leaderboard for some rankers.\nDid it contain the expiration date of customers?\n\nI also assume that Kaggle scoring system cannot find who exploit the features from that data or not.  \nThe most fairness solution is to make available the old members1.csv and give time to re-submit it. (although it might be troublesome to most competitors)",
      "votes": null
    },
    {
      "id": "259244",
      "postDate": "12/18/2017 01:41:49",
      "content": "<p>Congrats Bryan!! I'm looking forward to your solutions - especially for features which made the outstanding score among all.</p>",
      "rawMarkdown": "Congrats Bryan!! I'm looking forward to your solutions - especially for features which made the outstanding score among all.",
      "votes": null
    },
    {
      "id": "259248",
      "postDate": "12/18/2017 02:14:23",
      "content": "<p>I think the right thing organizers have to to is to examine the candidate winning solutions to see if they use some not allowed data and re run them to check if they can reach their scores.</p>",
      "rawMarkdown": "I think the right thing organizers have to to is to examine the candidate winning solutions to see if they use some not allowed data and re run them to check if they can reach their scores.",
      "votes": null
    },
    {
      "id": "259252",
      "postDate": "12/18/2017 02:28:51",
      "content": "<p>By the way, to be clear, my solution did use members3 but also included data from the members file from the other kkbox competition, and adding features from it netted a significant gain for a subset of the data  (LB improvement of .003-.004 if I recall).  That was allowed per Kaggle admins.  Sorry to anyone that missed the chance to net a gain from that.</p>",
      "rawMarkdown": "By the way, to be clear, my solution did use members3 but also included data from the members file from the other kkbox competition, and adding features from it netted a significant gain for a subset of the data  (LB improvement of .003-.004 if I recall).  That was allowed per Kaggle admins.  Sorry to anyone that missed the chance to net a gain from that.",
      "votes": null
    },
    {
      "id": "259255",
      "postDate": "12/18/2017 02:39:22",
      "content": "<p>Congrats Bryan! I also use use data from the other competition but I was not able to get such an improvement!</p>",
      "rawMarkdown": "Congrats Bryan! I also use use data from the other competition but I was not able to get such an improvement!",
      "votes": null
    },
    {
      "id": "259258",
      "postDate": "12/18/2017 02:53:01",
      "content": "<p>From that file I used these features:  a time scaled membership days remaining, a user change flag (flag users that had member data [bd, reg method, reg init time] that changed between the time of that members file and the members3 file), and a city change flag (flag users that had a change in city between the files).  I then built a separate base model for just the subset of data that had coverage in that file (~5% of the data in my training set I believe).  Then it was ensembled in with the other models.  </p>\n\n<p>Did you not gain any improvement from using it?</p>",
      "rawMarkdown": "From that file I used these features:  a time scaled membership days remaining, a user change flag (flag users that had member data [bd, reg method, reg init time] that changed between the time of that members file and the members3 file), and a city change flag (flag users that had a change in city between the files).  I then built a separate base model for just the subset of data that had coverage in that file (~5% of the data in my training set I believe).  Then it was ensembled in with the other models.  \n\nDid you not gain any improvement from using it?",
      "votes": null
    },
    {
      "id": "259486",
      "postDate": "12/18/2017 13:57:49",
      "content": "<p>I got some improvment, but not that much!</p>",
      "rawMarkdown": "I got some improvment, but not that much!",
      "votes": null
    },
    {
      "id": "259717",
      "postDate": "12/18/2017 22:37:48",
      "content": "<p>Bryan, congrats on your win!  </p>\n\n<p>Do you mind telling me whether you used the scala file to generate the labels?  </p>\n\n<p>Cheers, and congrats again!</p>",
      "rawMarkdown": "Bryan, congrats on your win!  \n\nDo you mind telling me whether you used the scala file to generate the labels?  \n\nCheers, and congrats again!",
      "votes": null
    },
    {
      "id": "260151",
      "postDate": "12/19/2017 17:39:18",
      "content": "<p>Thank you, STA.  I did use the scala script to generate new training data.  I'll go over it in more detail when I overview my solution, but I (and others as well) found that the training1 and training2 sets made available here were generated using a different scala script with different criteria then what was used for the test set and was later made available for use.  I don't think it was possible to achieve a high score (say, top 30 spot) using the original training files.  And moving to the new training sets definitely made my CV scores more in sync with the LB.</p>",
      "rawMarkdown": "Thank you, STA.  I did use the scala script to generate new training data.  I'll go over it in more detail when I overview my solution, but I (and others as well) found that the training1 and training2 sets made available here were generated using a different scala script with different criteria then what was used for the test set and was later made available for use.  I don't think it was possible to achieve a high score (say, top 30 spot) using the original training files.  And moving to the new training sets definitely made my CV scores more in sync with the LB.",
      "votes": null
    },
    {
      "id": "260186",
      "postDate": "12/19/2017 20:34:06",
      "content": "<p>Bryan, thanks again for the comments.  I could never get my personal CV scores to align with the LB scores.  My CV scores were always about .5 points \"higher\" than my LB scores.  It now makes more sense.  My model was trying to explain customers who were incorrectly labeled as churners in the training set.  </p>\n\n<p>Again, thanks for your comments.  They answer a lot of questions.  I appreciate your sharing...</p>",
      "rawMarkdown": "Bryan, thanks again for the comments.  I could never get my personal CV scores to align with the LB scores.  My CV scores were always about .5 points \"higher\" than my LB scores.  It now makes more sense.  My model was trying to explain customers who were incorrectly labeled as churners in the training set.  \n\nAgain, thanks for your comments.  They answer a lot of questions.  I appreciate your sharing...",
      "votes": null
    },
    {
      "id": "263390",
      "postDate": "12/30/2017 01:10:37",
      "content": "<p>Sorry I didn't get to this sooner. With work and holidays hitting hard, I've been extremely crunched for time.   But as promised, I spent a little time today adding an expiration_dt feature from the members1 (prohibited) file to my main base model, and below are my submission results.  It achieved a .0579, which is a .022 gain in accuracy over my winning model.  </p>\n\n<p>Keep in mind what I submitted is just a quick base XGB model, not tuned, and not stacked or blended with any other learners.  If I spent time tuning it properly and scaling it, and then built out an ensemble as I did with my winning submission, then it would probably score &lt;.055 easily and maybe even &lt;.05.  There could also be other features derived from the members1 expiration_dt that add signal (I only used one raw date feature).  </p>\n\n<p>Just wanted to  show what is possible with my model and set of features  if the members1 file is used, to prove that it was not used in my winning submission.  I feel I owed that to the Kagglers who felt like they potentially lost unfairly.  A thorough overview of the model will be presented in the WSDM paper.</p>\n\n<p><img src=\"https://bryangregory.com/Kaggle/kaggle_kkbox.png\" alt=\"Submission Score\" title=\"\"></p>\n\n<p>Link:  <a href=\"https://bryangregory.com/Kaggle/kaggle_kkbox.png\">https://bryangregory.com/Kaggle/kaggle_kkbox.png</a></p>",
      "rawMarkdown": "Sorry I didn't get to this sooner. With work and holidays hitting hard, I've been extremely crunched for time.   But as promised, I spent a little time today adding an expiration_dt feature from the members1 (prohibited) file to my main base model, and below are my submission results.  It achieved a .0579, which is a .022 gain in accuracy over my winning model.  \n\nKeep in mind what I submitted is just a quick base XGB model, not tuned, and not stacked or blended with any other learners.  If I spent time tuning it properly and scaling it, and then built out an ensemble as I did with my winning submission, then it would probably score &lt;.055 easily and maybe even &lt;.05.  There could also be other features derived from the members1 expiration_dt that add signal (I only used one raw date feature).  \n\nJust wanted to  show what is possible with my model and set of features  if the members1 file is used, to prove that it was not used in my winning submission.  I feel I owed that to the Kagglers who felt like they potentially lost unfairly.  A thorough overview of the model will be presented in the WSDM paper.\n\n![Submission Score][1]\n\n\n  [1]: https://bryangregory.com/Kaggle/kaggle_kkbox.png\n\nLink:  https://bryangregory.com/Kaggle/kaggle_kkbox.png",
      "votes": null
    },
    {
      "id": "263824",
      "postDate": "01/01/2018 06:18:24",
      "content": "<p>Really cool! Congrats!</p>",
      "rawMarkdown": "Really cool! Congrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 259228,
      "author_name": "bryangregory",
      "author_url": "",
      "post_date": "12/18/2017 01:06:15",
      "content": "<p>One idea is to have the top teams post screenshots of how they now score with their final best model WITH the members1.csv added back into their model.  As you probably know, when a Kaggle competition ends users can still keep making submissions to view what their LB scores would have been.</p>\n\n<p>Presumably if data from members1 was not used in their model, then they should see a huge jump in LB score by adding it back in.  I believe .015-.025 can be expected.  So a current .09 score should achieve at LEAST a .075 score with members1 csv data.  If the teams can show proof that their model can achieve a &gt;.015 LB gain by adding member1 back and submitting it, then posting the screenshot, then I would say they deserve their spot.  Selfishly, I just want to see how low the scoreboard can go :)  I'm betting I can score below .055 if I use member1 data and spend time re-tuning everything.</p>\n\n<p>Just a thought, upvote if you like the idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 259230,
          "author_name": "jeongyoonlee",
          "author_url": "",
          "post_date": "12/18/2017 01:14:48",
          "content": "<p>Thanks @Bryan for the idea. I like the idea. It might be little inconvenient for teams who haven't used it to add it and rerun the pipeline though. </p>\n\n<p>By the way, is your Private LB score with members? If so, what's your score without old members? If not, it's really impressive and hats off. Either way, great job and congrats! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259232,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/18/2017 01:17:19",
          "content": "<p>It includes member3 data, not member1.  I'll spend some time tomorrow re-modeling with members1 and will post a screenshot of what my model can achieve with it.  \"Proof is in the pudding\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259233,
          "author_name": "jeongyoonlee",
          "author_url": "",
          "post_date": "12/18/2017 01:22:56",
          "content": "<p>Awesome! Congrats for the 1st place! Looking forward to your solution. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259237,
          "author_name": "soundwaveli00",
          "author_url": "",
          "post_date": "12/18/2017 01:27:12",
          "content": "<p>Also you can't download members_v1 anymore... for teams who joined this competition after data refresh, they can't get that file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259244,
          "author_name": "sundong",
          "author_url": "",
          "post_date": "12/18/2017 01:41:49",
          "content": "<p>Congrats Bryan!! I'm looking forward to your solutions - especially for features which made the outstanding score among all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263390,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/30/2017 01:10:37",
          "content": "<p>Sorry I didn't get to this sooner. With work and holidays hitting hard, I've been extremely crunched for time.   But as promised, I spent a little time today adding an expiration_dt feature from the members1 (prohibited) file to my main base model, and below are my submission results.  It achieved a .0579, which is a .022 gain in accuracy over my winning model.  </p>\n\n<p>Keep in mind what I submitted is just a quick base XGB model, not tuned, and not stacked or blended with any other learners.  If I spent time tuning it properly and scaling it, and then built out an ensemble as I did with my winning submission, then it would probably score &lt;.055 easily and maybe even &lt;.05.  There could also be other features derived from the members1 expiration_dt that add signal (I only used one raw date feature).  </p>\n\n<p>Just wanted to  show what is possible with my model and set of features  if the members1 file is used, to prove that it was not used in my winning submission.  I feel I owed that to the Kagglers who felt like they potentially lost unfairly.  A thorough overview of the model will be presented in the WSDM paper.</p>\n\n<p><img src=\"https://bryangregory.com/Kaggle/kaggle_kkbox.png\" alt=\"Submission Score\" title=\"\"></p>\n\n<p>Link:  <a href=\"https://bryangregory.com/Kaggle/kaggle_kkbox.png\">https://bryangregory.com/Kaggle/kaggle_kkbox.png</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263824,
          "author_name": "soundwaveli00",
          "author_url": "",
          "post_date": "01/01/2018 06:18:24",
          "content": "<p>Really cool! Congrats!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 259242,
      "author_name": "sundong",
      "author_url": "",
      "post_date": "12/18/2017 01:38:25",
      "content": "<p>I was wondering what is the secret to developing the 0.07 score, and here is the answer! The old members.csv...!\nWe all notice that there is a huge score drop between public leaderboard and private leaderboard for some rankers.\nDid it contain the expiration date of customers?</p>\n\n<p>I also assume that Kaggle scoring system cannot find who exploit the features from that data or not. <br>\nThe most fairness solution is to make available the old members1.csv and give time to re-submit it. (although it might be troublesome to most competitors)   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 259248,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "12/18/2017 02:14:23",
      "content": "<p>I think the right thing organizers have to to is to examine the candidate winning solutions to see if they use some not allowed data and re run them to check if they can reach their scores.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 259252,
      "author_name": "bryangregory",
      "author_url": "",
      "post_date": "12/18/2017 02:28:51",
      "content": "<p>By the way, to be clear, my solution did use members3 but also included data from the members file from the other kkbox competition, and adding features from it netted a significant gain for a subset of the data  (LB improvement of .003-.004 if I recall).  That was allowed per Kaggle admins.  Sorry to anyone that missed the chance to net a gain from that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 259255,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "12/18/2017 02:39:22",
          "content": "<p>Congrats Bryan! I also use use data from the other competition but I was not able to get such an improvement!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259258,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/18/2017 02:53:01",
          "content": "<p>From that file I used these features:  a time scaled membership days remaining, a user change flag (flag users that had member data [bd, reg method, reg init time] that changed between the time of that members file and the members3 file), and a city change flag (flag users that had a change in city between the files).  I then built a separate base model for just the subset of data that had coverage in that file (~5% of the data in my training set I believe).  Then it was ensembled in with the other models.  </p>\n\n<p>Did you not gain any improvement from using it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259486,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "12/18/2017 13:57:49",
          "content": "<p>I got some improvment, but not that much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259717,
          "author_name": "",
          "author_url": "",
          "post_date": "12/18/2017 22:37:48",
          "content": "<p>Bryan, congrats on your win!  </p>\n\n<p>Do you mind telling me whether you used the scala file to generate the labels?  </p>\n\n<p>Cheers, and congrats again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260151,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/19/2017 17:39:18",
          "content": "<p>Thank you, STA.  I did use the scala script to generate new training data.  I'll go over it in more detail when I overview my solution, but I (and others as well) found that the training1 and training2 sets made available here were generated using a different scala script with different criteria then what was used for the test set and was later made available for use.  I don't think it was possible to achieve a high score (say, top 30 spot) using the original training files.  And moving to the new training sets definitely made my CV scores more in sync with the LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260186,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2017 20:34:06",
          "content": "<p>Bryan, thanks again for the comments.  I could never get my personal CV scores to align with the LB scores.  My CV scores were always about .5 points \"higher\" than my LB scores.  It now makes more sense.  My model was trying to explain customers who were incorrectly labeled as churners in the training set.  </p>\n\n<p>Again, thanks for your comments.  They answer a lot of questions.  I appreciate your sharing...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "259216": "As mentioned [here][1], the old members.csv is not allowed to use.\n\nBased on the current top Private LB scores, I assume that some teams still selected submissions that used the expiration dates in the old members.csv (Our best score with the expiration date was 0.07817 on Private LB and 0.07899 on Public LB).\n\nMy question is how the organizer are going to validate if one did not use such data sets that are not permitted.\n\nPlease advise.\n\nThanks in advance.\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/44363#250087",
    "259228": "One idea is to have the top teams post screenshots of how they now score with their final best model WITH the members1.csv added back into their model.  As you probably know, when a Kaggle competition ends users can still keep making submissions to view what their LB scores would have been.\n\nPresumably if data from members1 was not used in their model, then they should see a huge jump in LB score by adding it back in.  I believe .015-.025 can be expected.  So a current .09 score should achieve at LEAST a .075 score with members1 csv data.  If the teams can show proof that their model can achieve a &gt;.015 LB gain by adding member1 back and submitting it, then posting the screenshot, then I would say they deserve their spot.  Selfishly, I just want to see how low the scoreboard can go :)  I'm betting I can score below .055 if I use member1 data and spend time re-tuning everything.\n\nJust a thought, upvote if you like the idea.",
    "259230": "Thanks @Bryan for the idea. I like the idea. It might be little inconvenient for teams who haven't used it to add it and rerun the pipeline though. \n\nBy the way, is your Private LB score with members? If so, what's your score without old members? If not, it's really impressive and hats off. Either way, great job and congrats! :)",
    "259232": "It includes member3 data, not member1.  I'll spend some time tomorrow re-modeling with members1 and will post a screenshot of what my model can achieve with it.  \"Proof is in the pudding\"",
    "259233": "Awesome! Congrats for the 1st place! Looking forward to your solution. :)",
    "259237": "Also you can't download members_v1 anymore... for teams who joined this competition after data refresh, they can't get that file.",
    "259242": "I was wondering what is the secret to developing the 0.07 score, and here is the answer! The old members.csv...!\nWe all notice that there is a huge score drop between public leaderboard and private leaderboard for some rankers.\nDid it contain the expiration date of customers?\n\nI also assume that Kaggle scoring system cannot find who exploit the features from that data or not.  \nThe most fairness solution is to make available the old members1.csv and give time to re-submit it. (although it might be troublesome to most competitors)",
    "259244": "Congrats Bryan!! I'm looking forward to your solutions - especially for features which made the outstanding score among all.",
    "259248": "I think the right thing organizers have to to is to examine the candidate winning solutions to see if they use some not allowed data and re run them to check if they can reach their scores.",
    "259252": "By the way, to be clear, my solution did use members3 but also included data from the members file from the other kkbox competition, and adding features from it netted a significant gain for a subset of the data  (LB improvement of .003-.004 if I recall).  That was allowed per Kaggle admins.  Sorry to anyone that missed the chance to net a gain from that.",
    "259255": "Congrats Bryan! I also use use data from the other competition but I was not able to get such an improvement!",
    "259258": "From that file I used these features:  a time scaled membership days remaining, a user change flag (flag users that had member data [bd, reg method, reg init time] that changed between the time of that members file and the members3 file), and a city change flag (flag users that had a change in city between the files).  I then built a separate base model for just the subset of data that had coverage in that file (~5% of the data in my training set I believe).  Then it was ensembled in with the other models.  \n\nDid you not gain any improvement from using it?",
    "259486": "I got some improvment, but not that much!",
    "259717": "Bryan, congrats on your win!  \n\nDo you mind telling me whether you used the scala file to generate the labels?  \n\nCheers, and congrats again!",
    "260151": "Thank you, STA.  I did use the scala script to generate new training data.  I'll go over it in more detail when I overview my solution, but I (and others as well) found that the training1 and training2 sets made available here were generated using a different scala script with different criteria then what was used for the test set and was later made available for use.  I don't think it was possible to achieve a high score (say, top 30 spot) using the original training files.  And moving to the new training sets definitely made my CV scores more in sync with the LB.",
    "260186": "Bryan, thanks again for the comments.  I could never get my personal CV scores to align with the LB scores.  My CV scores were always about .5 points \"higher\" than my LB scores.  It now makes more sense.  My model was trying to explain customers who were incorrectly labeled as churners in the training set.  \n\nAgain, thanks for your comments.  They answer a lot of questions.  I appreciate your sharing...",
    "263390": "Sorry I didn't get to this sooner. With work and holidays hitting hard, I've been extremely crunched for time.   But as promised, I spent a little time today adding an expiration_dt feature from the members1 (prohibited) file to my main base model, and below are my submission results.  It achieved a .0579, which is a .022 gain in accuracy over my winning model.  \n\nKeep in mind what I submitted is just a quick base XGB model, not tuned, and not stacked or blended with any other learners.  If I spent time tuning it properly and scaling it, and then built out an ensemble as I did with my winning submission, then it would probably score &lt;.055 easily and maybe even &lt;.05.  There could also be other features derived from the members1 expiration_dt that add signal (I only used one raw date feature).  \n\nJust wanted to  show what is possible with my model and set of features  if the members1 file is used, to prove that it was not used in my winning submission.  I feel I owed that to the Kagglers who felt like they potentially lost unfairly.  A thorough overview of the model will be presented in the WSDM paper.\n\n![Submission Score][1]\n\n\n  [1]: https://bryangregory.com/Kaggle/kaggle_kkbox.png\n\nLink:  https://bryangregory.com/Kaggle/kaggle_kkbox.png",
    "263824": "Really cool! Congrats!"
  },
  "source": "meta"
}