{
  "id": 40142,
  "title": "Please open training data generation scripts to make this competition fair",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/40142",
  "author_name": "",
  "post_date": "2017-09-28T09:47:08.569254600Z",
  "votes": 33,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear organizers, thank you very much for organizing an interesting competition. I have a suggestion to keep this competition fair. Please open your script which was used to create the training and the test dataset. I saw many questions so far are around how to re-engineer the rules to label the data and to ignore training example. I think it is crucial to give all people access to the scripts or at least define the set of rules  in a formal, complete and concrete way. The current rules are very ambiguous (the rule to label churned users and prediction out of the scope).   Otherwise, I doubt that this competition is all about re-engineering those rules. Thank you! </p>",
  "messages": [
    {
      "id": "225119",
      "postDate": "09/28/2017 09:47:08",
      "content": "<p>Dear organizers, thank you very much for organizing an interesting competition. I have a suggestion to keep this competition fair. Please open your script which was used to create the training and the test dataset. I saw many questions so far are around how to re-engineer the rules to label the data and to ignore training example. I think it is crucial to give all people access to the scripts or at least define the set of rules  in a formal, complete and concrete way. The current rules are very ambiguous (the rule to label churned users and prediction out of the scope).   Otherwise, I doubt that this competition is all about re-engineering those rules. Thank you! </p>",
      "rawMarkdown": "Dear organizers, thank you very much for organizing an interesting competition. I have a suggestion to keep this competition fair. Please open your script which was used to create the training and the test dataset. I saw many questions so far are around how to re-engineer the rules to label the data and to ignore training example. I think it is crucial to give all people access to the scripts or at least define the set of rules  in a formal, complete and concrete way. The current rules are very ambiguous (the rule to label churned users and prediction out of the scope).   Otherwise, I doubt that this competition is all about re-engineering those rules. Thank you!",
      "votes": null
    },
    {
      "id": "233001",
      "postDate": "10/19/2017 01:07:47",
      "content": "<p>I totally agree on this. After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns. </p>",
      "rawMarkdown": "I totally agree on this. After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns.",
      "votes": null
    },
    {
      "id": "233079",
      "postDate": "10/19/2017 06:27:38",
      "content": "<p>Based on this <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/39913\">post</a>,\nsnapshot will affect whether a user falls into the scope. Can organizers also provide the snapshots which are used to generate training and testing data set? </p>",
      "rawMarkdown": "Based on this [post](https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/39913),\nsnapshot will affect whether a user falls into the scope. Can organizers also provide the snapshots which are used to generate training and testing data set?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 233001,
      "author_name": "shunjiangxu",
      "author_url": "",
      "post_date": "10/19/2017 01:07:47",
      "content": "<p>I totally agree on this. After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 233079,
      "author_name": "soundwaveli00",
      "author_url": "",
      "post_date": "10/19/2017 06:27:38",
      "content": "<p>Based on this <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/39913\">post</a>,\nsnapshot will affect whether a user falls into the scope. Can organizers also provide the snapshots which are used to generate training and testing data set? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "225119": "Dear organizers, thank you very much for organizing an interesting competition. I have a suggestion to keep this competition fair. Please open your script which was used to create the training and the test dataset. I saw many questions so far are around how to re-engineer the rules to label the data and to ignore training example. I think it is crucial to give all people access to the scripts or at least define the set of rules  in a formal, complete and concrete way. The current rules are very ambiguous (the rule to label churned users and prediction out of the scope).   Otherwise, I doubt that this competition is all about re-engineering those rules. Thank you!",
    "233001": "I totally agree on this. After spending so much time processing the transaction data for the users in the train set, I found out the majority (~90%) of them will be 'out of scope' based on the definition given in the competition. I don't know what's really happening in the test set. Also someone was asking the Log-loss question and how come the loss is not log(2) when submitting all 0.5 as the prediction. Then another question is the 'expiration date' feature and if it's futuristic. In the XGboost model, you can see it's 'the' most important factor for the prediction. It certainly looks like futuristic and should be removed from the data if so. When you have a model with so much difference in the 'validation' and final 'test' result, it points to high probabilities of issues with the data itself, such as they are very different distribution. These questions point to some serious issues of this competition. I hope more people upvote to catch the organizers' attention. If they are serious about it and try to get some good results please do respond to these serious concerns.",
    "233079": "Based on this [post](https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/39913),\nsnapshot will affect whether a user falls into the scope. Can organizers also provide the snapshots which are used to generate training and testing data set?"
  },
  "source": "meta"
}