{
  "id": 44363,
  "title": "Urgent question to admins: What data can we use?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44363",
  "author_name": "",
  "post_date": "2017-11-27T18:42:20.125706700Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>After a lot of confusion regarding this competition data and since external data is allowed, I think a clear statement about the data that the competitors may use should be given. Could you please confirm if we can or we cannot use this data:</p>\n\n<ul>\n<li>version 1, version 2 and version 3 (if available) of all competition files</li>\n<li>data extracted with  provided scala labeller</li>\n<li>data from the other WSDM competition</li>\n</ul>",
  "messages": [
    {
      "id": "249116",
      "postDate": "11/27/2017 18:42:20",
      "content": "<p>After a lot of confusion regarding this competition data and since external data is allowed, I think a clear statement about the data that the competitors may use should be given. Could you please confirm if we can or we cannot use this data:</p>\n\n<ul>\n<li>version 1, version 2 and version 3 (if available) of all competition files</li>\n<li>data extracted with  provided scala labeller</li>\n<li>data from the other WSDM competition</li>\n</ul>",
      "rawMarkdown": "After a lot of confusion regarding this competition data and since external data is allowed, I think a clear statement about the data that the competitors may use should be given. Could you please confirm if we can or we cannot use this data:\n\n- version 1, version 2 and version 3 (if available) of all competition files\n- data extracted with  provided scala labeller\n- data from the other WSDM competition",
      "votes": null
    },
    {
      "id": "249209",
      "postDate": "11/28/2017 00:36:44",
      "content": "<p>Thanks for clarifying that, Aloisio.  As a late entry competitor that joined after the new files were released, I can say I don't even have access to a v2 members file.  I also haven't brought in any WSDM comp data files or data extracted via the scala labeller.  So your question and the subsequent answer could greatly change the data I'm currently building my models off of. And I'm sure there are others like me.</p>\n\n<p>Admins, please clarify as quickly as possible, thanks!</p>",
      "rawMarkdown": "Thanks for clarifying that, Aloisio.  As a late entry competitor that joined after the new files were released, I can say I don't even have access to a v2 members file.  I also haven't brought in any WSDM comp data files or data extracted via the scala labeller.  So your question and the subsequent answer could greatly change the data I'm currently building my models off of. And I'm sure there are others like me.\n\nAdmins, please clarify as quickly as possible, thanks!",
      "votes": null
    },
    {
      "id": "249212",
      "postDate": "11/28/2017 00:48:56",
      "content": "<p>Did you got the original members.csv (v1) file? I think the member_v2.csv file was broken, so they replaced with the v3 file. See Kyle question and Wendy answer in <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169</a></p>",
      "rawMarkdown": "Did you got the original members.csv (v1) file? I think the member_v2.csv file was broken, so they replaced with the v3 file. See Kyle question and Wendy answer in [https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169][1]\n\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169",
      "votes": null
    },
    {
      "id": "249328",
      "postDate": "11/28/2017 07:57:46",
      "content": "<p>Hi Bryan even I do not have access to latest data, can you please tell me if I use (train_version1 + member_version1+user_logs_version1) and test set is (submission_v2.csv) </p>",
      "rawMarkdown": "Hi Bryan even I do not have access to latest data, can you please tell me if I use (train_version1 + member_version1+user_logs_version1) and test set is (submission_v2.csv)",
      "votes": null
    },
    {
      "id": "249965",
      "postDate": "11/29/2017 15:07:39",
      "content": "<p>You can use whatever mix of files (it's all valid data) that you find works best for training, but the test set is definitely submission_v2.csv.  v2 versions of the other files are just updates with March data (so you will probably want to include them with your training).</p>",
      "rawMarkdown": "You can use whatever mix of files (it's all valid data) that you find works best for training, but the test set is definitely submission_v2.csv.  v2 versions of the other files are just updates with March data (so you will probably want to include them with your training).",
      "votes": null
    },
    {
      "id": "249972",
      "postDate": "11/29/2017 15:22:29",
      "content": "<p>I think the organizers will not answer my question. I was wondering how would they prevent people from using the data from the first member.csv file, since its already spread, if they decide that we are not allowed to use that file. </p>",
      "rawMarkdown": "I think the organizers will not answer my question. I was wondering how would they prevent people from using the data from the first member.csv file, since its already spread, if they decide that we are not allowed to use that file.",
      "votes": null
    },
    {
      "id": "250016",
      "postDate": "11/29/2017 16:38:53",
      "content": "<p>thanks a lot Bryan ...</p>",
      "rawMarkdown": "thanks a lot Bryan ...",
      "votes": null
    },
    {
      "id": "250087",
      "postDate": "11/29/2017 18:47:47",
      "content": "<p>All - As data extracted with provided scala labeller and data from the other WSDM competition are both currently available, they are considered \"public\" and can be used. </p>\n\n<p>As only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "All - As data extracted with provided scala labeller and data from the other WSDM competition are both currently available, they are considered \"public\" and can be used. \n\nAs only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.\n\nThanks!",
      "votes": null
    },
    {
      "id": "250173",
      "postDate": "11/29/2017 22:19:20",
      "content": "<p>Thank you very much Addison Howard! I think it would be a good idea to state that at an official thread. You also could state what will the admins do to ensure that, in the end,  the 3 top teams follow the rules, since some no more allowed files went public in the past.</p>",
      "rawMarkdown": "Thank you very much Addison Howard! I think it would be a good idea to state that at an official thread. You also could state what will the admins do to ensure that, in the end,  the 3 top teams follow the rules, since some no more allowed files went public in the past.",
      "votes": null
    },
    {
      "id": "250184",
      "postDate": "11/29/2017 22:51:07",
      "content": "<p>Thanks Aloisio! We do not reveal all of the anti-cheating methods we employ to protect that process securely. And the official rules state that external, public data is allowed :)</p>",
      "rawMarkdown": "Thanks Aloisio! We do not reveal all of the anti-cheating methods we employ to protect that process securely. And the official rules state that external, public data is allowed :)",
      "votes": null
    },
    {
      "id": "250251",
      "postDate": "11/30/2017 00:13:21",
      "content": "<p>Ok! Thanks.</p>",
      "rawMarkdown": "Ok! Thanks.",
      "votes": null
    },
    {
      "id": "250371",
      "postDate": "11/30/2017 01:44:48",
      "content": "<p>Thanks for clarifying, Addison!</p>",
      "rawMarkdown": "Thanks for clarifying, Addison!",
      "votes": null
    },
    {
      "id": "255075",
      "postDate": "12/08/2017 07:39:50",
      "content": "<p>This was also a question in my mind, i am jumping in. I loaded all available files v1 v2... Basically i used group by and aggregate functions on all(transactions, logs...). However, dimensions doesn't seem to match. So you are saying sample_submissions_v2 is the only test set? Does anyone have this problem of different dimensions you get with group by&amp;aggregate functions?</p>",
      "rawMarkdown": "This was also a question in my mind, i am jumping in. I loaded all available files v1 v2... Basically i used group by and aggregate functions on all(transactions, logs...). However, dimensions doesn't seem to match. So you are saying sample_submissions_v2 is the only test set? Does anyone have this problem of different dimensions you get with group by&amp;aggregate functions?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 249209,
      "author_name": "bryangregory",
      "author_url": "",
      "post_date": "11/28/2017 00:36:44",
      "content": "<p>Thanks for clarifying that, Aloisio.  As a late entry competitor that joined after the new files were released, I can say I don't even have access to a v2 members file.  I also haven't brought in any WSDM comp data files or data extracted via the scala labeller.  So your question and the subsequent answer could greatly change the data I'm currently building my models off of. And I'm sure there are others like me.</p>\n\n<p>Admins, please clarify as quickly as possible, thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 249212,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "11/28/2017 00:48:56",
          "content": "<p>Did you got the original members.csv (v1) file? I think the member_v2.csv file was broken, so they replaced with the v3 file. See Kyle question and Wendy answer in <a href=\"https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169\">https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249328,
      "author_name": "mainakdatageek",
      "author_url": "",
      "post_date": "11/28/2017 07:57:46",
      "content": "<p>Hi Bryan even I do not have access to latest data, can you please tell me if I use (train_version1 + member_version1+user_logs_version1) and test set is (submission_v2.csv) </p>",
      "votes": null,
      "replies": [
        {
          "id": 249965,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "11/29/2017 15:07:39",
          "content": "<p>You can use whatever mix of files (it's all valid data) that you find works best for training, but the test set is definitely submission_v2.csv.  v2 versions of the other files are just updates with March data (so you will probably want to include them with your training).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249972,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "11/29/2017 15:22:29",
      "content": "<p>I think the organizers will not answer my question. I was wondering how would they prevent people from using the data from the first member.csv file, since its already spread, if they decide that we are not allowed to use that file. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250016,
      "author_name": "mainakdatageek",
      "author_url": "",
      "post_date": "11/29/2017 16:38:53",
      "content": "<p>thanks a lot Bryan ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250087,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "11/29/2017 18:47:47",
      "content": "<p>All - As data extracted with provided scala labeller and data from the other WSDM competition are both currently available, they are considered \"public\" and can be used. </p>\n\n<p>As only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 250173,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "11/29/2017 22:19:20",
          "content": "<p>Thank you very much Addison Howard! I think it would be a good idea to state that at an official thread. You also could state what will the admins do to ensure that, in the end,  the 3 top teams follow the rules, since some no more allowed files went public in the past.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250184,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "11/29/2017 22:51:07",
          "content": "<p>Thanks Aloisio! We do not reveal all of the anti-cheating methods we employ to protect that process securely. And the official rules state that external, public data is allowed :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250251,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "11/30/2017 00:13:21",
          "content": "<p>Ok! Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250371,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "11/30/2017 01:44:48",
          "content": "<p>Thanks for clarifying, Addison!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255075,
          "author_name": "berkkarahan",
          "author_url": "",
          "post_date": "12/08/2017 07:39:50",
          "content": "<p>This was also a question in my mind, i am jumping in. I loaded all available files v1 v2... Basically i used group by and aggregate functions on all(transactions, logs...). However, dimensions doesn't seem to match. So you are saying sample_submissions_v2 is the only test set? Does anyone have this problem of different dimensions you get with group by&amp;aggregate functions?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "249116": "After a lot of confusion regarding this competition data and since external data is allowed, I think a clear statement about the data that the competitors may use should be given. Could you please confirm if we can or we cannot use this data:\n\n- version 1, version 2 and version 3 (if available) of all competition files\n- data extracted with  provided scala labeller\n- data from the other WSDM competition",
    "249209": "Thanks for clarifying that, Aloisio.  As a late entry competitor that joined after the new files were released, I can say I don't even have access to a v2 members file.  I also haven't brought in any WSDM comp data files or data extracted via the scala labeller.  So your question and the subsequent answer could greatly change the data I'm currently building my models off of. And I'm sure there are others like me.\n\nAdmins, please clarify as quickly as possible, thanks!",
    "249212": "Did you got the original members.csv (v1) file? I think the member_v2.csv file was broken, so they replaced with the v3 file. See Kyle question and Wendy answer in [https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169][1]\n\n\n  [1]: https://www.kaggle.com/c/kkbox-churn-prediction-challenge/discussion/43169",
    "249328": "Hi Bryan even I do not have access to latest data, can you please tell me if I use (train_version1 + member_version1+user_logs_version1) and test set is (submission_v2.csv)",
    "249965": "You can use whatever mix of files (it's all valid data) that you find works best for training, but the test set is definitely submission_v2.csv.  v2 versions of the other files are just updates with March data (so you will probably want to include them with your training).",
    "249972": "I think the organizers will not answer my question. I was wondering how would they prevent people from using the data from the first member.csv file, since its already spread, if they decide that we are not allowed to use that file.",
    "250016": "thanks a lot Bryan ...",
    "250087": "All - As data extracted with provided scala labeller and data from the other WSDM competition are both currently available, they are considered \"public\" and can be used. \n\nAs only the latest versions of the competition files are public, those are the only versions of the competition data that can be used.\n\nThanks!",
    "250173": "Thank you very much Addison Howard! I think it would be a good idea to state that at an official thread. You also could state what will the admins do to ensure that, in the end,  the 3 top teams follow the rules, since some no more allowed files went public in the past.",
    "250184": "Thanks Aloisio! We do not reveal all of the anti-cheating methods we employ to protect that process securely. And the official rules state that external, public data is allowed :)",
    "250251": "Ok! Thanks.",
    "250371": "Thanks for clarifying, Addison!",
    "255075": "This was also a question in my mind, i am jumping in. I loaded all available files v1 v2... Basically i used group by and aggregate functions on all(transactions, logs...). However, dimensions doesn't seem to match. So you are saying sample_submissions_v2 is the only test set? Does anyone have this problem of different dimensions you get with group by&amp;aggregate functions?"
  },
  "source": "meta"
}