{
  "id": 45991,
  "title": "Accuracy of the labels in training_v2.csv",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45991",
  "author_name": "",
  "post_date": "2017-12-19T04:27:19.834244200Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Congrats to the winners and competitors.......</p>\n\n<p>InfiniteWing has posted training labels (is_churn) using the scala file.  Thanks InfiniteWing.  Here is what I found, using his 201703 file.</p>\n\n<p>There are 87,330 churners in the training_v2.csv file.   That is the file that I used to train my model.  Compared with the scala program, which presumably was used to label the submission file, </p>\n\n<p>24,714 (28.3%) are dropped, 28,373 (32.5%) are labeled as non-churners, and 34,343 (39.2%) are labeled as churners.</p>\n\n<p>So, the scala program and training_v2.csv agree on only 40% of the labeled churners. </p>\n\n<p>If would be great if someone can confirm this or tell me where my logic is wrong.  I understand that the scala program was public, but you do expect that the provided labels in the training set would be correct, right?</p>",
  "messages": [
    {
      "id": "259854",
      "postDate": "12/19/2017 04:27:19",
      "content": "<p>Congrats to the winners and competitors.......</p>\n\n<p>InfiniteWing has posted training labels (is_churn) using the scala file.  Thanks InfiniteWing.  Here is what I found, using his 201703 file.</p>\n\n<p>There are 87,330 churners in the training_v2.csv file.   That is the file that I used to train my model.  Compared with the scala program, which presumably was used to label the submission file, </p>\n\n<p>24,714 (28.3%) are dropped, 28,373 (32.5%) are labeled as non-churners, and 34,343 (39.2%) are labeled as churners.</p>\n\n<p>So, the scala program and training_v2.csv agree on only 40% of the labeled churners. </p>\n\n<p>If would be great if someone can confirm this or tell me where my logic is wrong.  I understand that the scala program was public, but you do expect that the provided labels in the training set would be correct, right?</p>",
      "rawMarkdown": "Congrats to the winners and competitors.......\n\nInfiniteWing has posted training labels (is_churn) using the scala file.  Thanks InfiniteWing.  Here is what I found, using his 201703 file.\n\nThere are 87,330 churners in the training_v2.csv file.   That is the file that I used to train my model.  Compared with the scala program, which presumably was used to label the submission file, \n\n24,714 (28.3%) are dropped, 28,373 (32.5%) are labeled as non-churners, and 34,343 (39.2%) are labeled as churners.\n\nSo, the scala program and training_v2.csv agree on only 40% of the labeled churners. \n\nIf would be great if someone can confirm this or tell me where my logic is wrong.  I understand that the scala program was public, but you do expect that the provided labels in the training set would be correct, right?",
      "votes": null
    },
    {
      "id": "260152",
      "postDate": "12/19/2017 17:48:11",
      "content": "<p>My stats differ slightly from what you posted, so InfiniteWing and I must have done something slightly different.  If I recall, I found a minor discrepancy in the scala script that was creating further noise (related to date cutoffs like \"MSNO=amZm5p3h8Nw6CxciKVOjcckI5inwCEkWQFcwWfXxvTQ=\" which should not be a churn but the scala script is labelling as a churn) and I attempted to correct them.  I'll post in more detail later in my overview.  </p>\n\n<p>But your premise is correct that the training sets generated using the scala script differed greatly from the training sets posted for the contest.  Here are some quick stats I show for my scala-generated training sets:  </p>\n\n<p>Scala training set for January (training 1):  </p>\n\n<ul>\n<li>Total churn candidates (total records):  879,478</li>\n<li>Total churned:  34,771</li>\n<li>Churn avg:  .039536</li>\n</ul>\n\n<p>Scala training set for Februray (training 2):  </p>\n\n<ul>\n<li>Total churn candidates (total records):  862,158</li>\n<li>Total churned:  42,681</li>\n<li>Churn avg:  .049505</li>\n</ul>\n\n<p>And presumed churn average for March (test):  ~.036</p>",
      "rawMarkdown": "My stats differ slightly from what you posted, so InfiniteWing and I must have done something slightly different.  If I recall, I found a minor discrepancy in the scala script that was creating further noise (related to date cutoffs like \"MSNO=amZm5p3h8Nw6CxciKVOjcckI5inwCEkWQFcwWfXxvTQ=\" which should not be a churn but the scala script is labelling as a churn) and I attempted to correct them.  I'll post in more detail later in my overview.  \n\nBut your premise is correct that the training sets generated using the scala script differed greatly from the training sets posted for the contest.  Here are some quick stats I show for my scala-generated training sets:  \n\nScala training set for January (training 1):  \n\n - Total churn candidates (total records):  879,478\n - Total churned:  34,771\n - Churn avg:  .039536\n\nScala training set for Februray (training 2):  \n\n - Total churn candidates (total records):  862,158\n - Total churned:  42,681\n - Churn avg:  .049505\n\nAnd presumed churn average for March (test):  ~.036",
      "votes": null
    },
    {
      "id": "260174",
      "postDate": "12/19/2017 19:36:32",
      "content": "<p>For 201703, I don't think we have enough transaction to generate label?</p>",
      "rawMarkdown": "For 201703, I don't think we have enough transaction to generate label?",
      "votes": null
    },
    {
      "id": "260178",
      "postDate": "12/19/2017 20:03:44",
      "content": "<p>Bryan, thanks for your comments.  The administrators suggested that the scala program was a nice-to-have, meaning that it could provide additional information.   I didn't realize it was a prerequisite to generating the training set itself.  </p>\n\n<p>I started to get the sense that the training set was wrong when I did a deep-dive on the user logs.  Many of the identified churners were streaming music past their membership expiration date.  I rationalized that it was some special incentive deal to entice them to stay.  It turns out they were not churners at all.  The user logs were the least important type of features in my model.  Of course, that is surprising because you would expect it to be rich with information because it represents user behavior.  Now I know that it was not as useful because 60% of the churners were invalid.  </p>\n\n<p>In any case, the scala file was fair game.  Like you commented elsewhere, it was difficult to get it to run on a windows-based machine, so I did not pursue further.  My bad....</p>\n\n<p>Congrats on your win again!</p>",
      "rawMarkdown": "Bryan, thanks for your comments.  The administrators suggested that the scala program was a nice-to-have, meaning that it could provide additional information.   I didn't realize it was a prerequisite to generating the training set itself.  \n\nI started to get the sense that the training set was wrong when I did a deep-dive on the user logs.  Many of the identified churners were streaming music past their membership expiration date.  I rationalized that it was some special incentive deal to entice them to stay.  It turns out they were not churners at all.  The user logs were the least important type of features in my model.  Of course, that is surprising because you would expect it to be rich with information because it represents user behavior.  Now I know that it was not as useful because 60% of the churners were invalid.  \n\nIn any case, the scala file was fair game.  Like you commented elsewhere, it was difficult to get it to run on a windows-based machine, so I did not pursue further.  My bad....\n\nCongrats on your win again!",
      "votes": null
    },
    {
      "id": "260182",
      "postDate": "12/19/2017 20:19:29",
      "content": "<p>Yep, exactly.  And just to add to that, the way the contest was structured (by calendar month) made it difficult as well to generate signal from UL and TRX data (particularly UL data).  Of course in a real world  churn model we would be interested in what a user has done recently relative to the day of the month their membership is expiring, structuring the problem by calendar months.  In other words, Mar UL activity data has more meaning for a user expiring in early Apr, than it does for a user expiring in late Apr.  </p>",
      "rawMarkdown": "Yep, exactly.  And just to add to that, the way the contest was structured (by calendar month) made it difficult as well to generate signal from UL and TRX data (particularly UL data).  Of course in a real world  churn model we would be interested in what a user has done recently relative to the day of the month their membership is expiring, structuring the problem by calendar months.  In other words, Mar UL activity data has more meaning for a user expiring in early Apr, than it does for a user expiring in late Apr.",
      "votes": null
    },
    {
      "id": "260183",
      "postDate": "12/19/2017 20:25:27",
      "content": "<p>It is presumed based on LB feedback.  See Saigon Apps's post.  My submission feedback closely mirrored what he posted. I.e.-My submissions with means near .035-.037 scored best.  </p>\n\n<p>I controlled the mean using a scale_pos_weight &lt; 1 on the XGBoost portion of my model.  The final submission used scale_pos_weight=.8.  If anyone used XGBoost and wants a quick increase in accuracy (LB improvement), then lower the scale_pos_weight until your submission is hitting a mean in the above range.  You will likely see a very large increase in accuracy.</p>",
      "rawMarkdown": "It is presumed based on LB feedback.  See Saigon Apps's post.  My submission feedback closely mirrored what he posted. I.e.-My submissions with means near .035-.037 scored best.  \n\nI controlled the mean using a scale_pos_weight &lt; 1 on the XGBoost portion of my model.  The final submission used scale_pos_weight=.8.  If anyone used XGBoost and wants a quick increase in accuracy (LB improvement), then lower the scale_pos_weight until your submission is hitting a mean in the above range.  You will likely see a very large increase in accuracy.",
      "votes": null
    },
    {
      "id": "260253",
      "postDate": "12/19/2017 23:35:15",
      "content": "<p>Yep, I just did the simple computation of adjusting my predictions down by 60% to a mean of 3.5 churn rate.  My LB score improves from .1245 to .1140, a jump of about 40 positions.  That is remarkable and pretty funny, given how much time I was spending during the last week to eek out changes of .001.  :)   </p>\n\n<p>I understand why making the adjustment helps me, since my models were trained on an invalid training set with a 9% churn rate.  But I am not sure why they would help you if you trained on the proper data set, unless 1) the submission file respondents are meaningfully different than the ones we trained on and / or 2) the scala program, as people understand it, was not used to generate the submission labels.  Both are probably true. </p>\n\n<p>Still, I am surprised that simply adjusting probabilities downwards like that would work, given how nonlinear the loss function is.  If the loss function were closer to linear I can more easily understand how reducing your predicted probabilities to match the target population would work.</p>\n\n<p>Cheers again Bryan. </p>",
      "rawMarkdown": "Yep, I just did the simple computation of adjusting my predictions down by 60% to a mean of 3.5 churn rate.  My LB score improves from .1245 to .1140, a jump of about 40 positions.  That is remarkable and pretty funny, given how much time I was spending during the last week to eek out changes of .001.  :)   \n\nI understand why making the adjustment helps me, since my models were trained on an invalid training set with a 9% churn rate.  But I am not sure why they would help you if you trained on the proper data set, unless 1) the submission file respondents are meaningfully different than the ones we trained on and / or 2) the scala program, as people understand it, was not used to generate the submission labels.  Both are probably true. \n\nStill, I am surprised that simply adjusting probabilities downwards like that would work, given how nonlinear the loss function is.  If the loss function were closer to linear I can more easily understand how reducing your predicted probabilities to match the target population would work.\n\nCheers again Bryan.",
      "votes": null
    },
    {
      "id": "260379",
      "postDate": "12/20/2017 05:00:59",
      "content": "<p>Hi @Hang, thanks for your reply, I did not consider that. Yes, there might have some different if one have renew their membership in April. But if someone still have transactions in march(201703) after his expiration date in march, the scala label code can still dig them out. (I think I should re-run my model later, currently I am using the churn label from 201701 to 201703)</p>\n\n<p>Some part of my scala code is like this:</p>\n\n<pre><code>val targetDIR = \"user_label_201703\"\nval historyCutoff = \"20170228\"\nval targetExpireStart = \"20170301\"\nval targetExpireEnd = \"20170331\"\n</code></pre>\n\n<p>P.S.\nI think train_v2.csv is the churn result for member whos membership expire in 201703, am I right?</p>",
      "rawMarkdown": "Hi @Hang, thanks for your reply, I did not consider that. Yes, there might have some different if one have renew their membership in April. But if someone still have transactions in march(201703) after his expiration date in march, the scala label code can still dig them out. (I think I should re-run my model later, currently I am using the churn label from 201701 to 201703)\n\n\nSome part of my scala code is like this:\n\n    val targetDIR = \"user_label_201703\"\n    val historyCutoff = \"20170228\"\n    val targetExpireStart = \"20170301\"\n    val targetExpireEnd = \"20170331\"\n\nP.S.\nI think train_v2.csv is the churn result for member whos membership expire in 201703, am I right?",
      "votes": null
    },
    {
      "id": "260382",
      "postDate": "12/20/2017 05:14:32",
      "content": "<p>In addition, I think you can estimate the churn rate by calculating the log loss with all 0 submission on public LB. Correct me if I am wrong.</p>",
      "rawMarkdown": "In addition, I think you can estimate the churn rate by calculating the log loss with all 0 submission on public LB. Correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "260424",
      "postDate": "12/20/2017 07:01:22",
      "content": "<p>Great point!  </p>\n\n<p>I think you can compute the average provided you use a constant probability other than .5.  I didn't think you could use zero but I see that it is replaced with 10**-15.</p>\n\n<p>This is my first Kaggle competition.  It is interesting to learn all of the techniques used.    </p>",
      "rawMarkdown": "Great point!  \n\nI think you can compute the average provided you use a constant probability other than .5.  I didn't think you could use zero but I see that it is replaced with 10**-15.\n\nThis is my first Kaggle competition.  It is interesting to learn all of the techniques used.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 260152,
      "author_name": "bryangregory",
      "author_url": "",
      "post_date": "12/19/2017 17:48:11",
      "content": "<p>My stats differ slightly from what you posted, so InfiniteWing and I must have done something slightly different.  If I recall, I found a minor discrepancy in the scala script that was creating further noise (related to date cutoffs like \"MSNO=amZm5p3h8Nw6CxciKVOjcckI5inwCEkWQFcwWfXxvTQ=\" which should not be a churn but the scala script is labelling as a churn) and I attempted to correct them.  I'll post in more detail later in my overview.  </p>\n\n<p>But your premise is correct that the training sets generated using the scala script differed greatly from the training sets posted for the contest.  Here are some quick stats I show for my scala-generated training sets:  </p>\n\n<p>Scala training set for January (training 1):  </p>\n\n<ul>\n<li>Total churn candidates (total records):  879,478</li>\n<li>Total churned:  34,771</li>\n<li>Churn avg:  .039536</li>\n</ul>\n\n<p>Scala training set for Februray (training 2):  </p>\n\n<ul>\n<li>Total churn candidates (total records):  862,158</li>\n<li>Total churned:  42,681</li>\n<li>Churn avg:  .049505</li>\n</ul>\n\n<p>And presumed churn average for March (test):  ~.036</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 260174,
      "author_name": "soundwaveli00",
      "author_url": "",
      "post_date": "12/19/2017 19:36:32",
      "content": "<p>For 201703, I don't think we have enough transaction to generate label?</p>",
      "votes": null,
      "replies": [
        {
          "id": 260183,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/19/2017 20:25:27",
          "content": "<p>It is presumed based on LB feedback.  See Saigon Apps's post.  My submission feedback closely mirrored what he posted. I.e.-My submissions with means near .035-.037 scored best.  </p>\n\n<p>I controlled the mean using a scale_pos_weight &lt; 1 on the XGBoost portion of my model.  The final submission used scale_pos_weight=.8.  If anyone used XGBoost and wants a quick increase in accuracy (LB improvement), then lower the scale_pos_weight until your submission is hitting a mean in the above range.  You will likely see a very large increase in accuracy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260253,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2017 23:35:15",
          "content": "<p>Yep, I just did the simple computation of adjusting my predictions down by 60% to a mean of 3.5 churn rate.  My LB score improves from .1245 to .1140, a jump of about 40 positions.  That is remarkable and pretty funny, given how much time I was spending during the last week to eek out changes of .001.  :)   </p>\n\n<p>I understand why making the adjustment helps me, since my models were trained on an invalid training set with a 9% churn rate.  But I am not sure why they would help you if you trained on the proper data set, unless 1) the submission file respondents are meaningfully different than the ones we trained on and / or 2) the scala program, as people understand it, was not used to generate the submission labels.  Both are probably true. </p>\n\n<p>Still, I am surprised that simply adjusting probabilities downwards like that would work, given how nonlinear the loss function is.  If the loss function were closer to linear I can more easily understand how reducing your predicted probabilities to match the target population would work.</p>\n\n<p>Cheers again Bryan. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260379,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/20/2017 05:00:59",
          "content": "<p>Hi @Hang, thanks for your reply, I did not consider that. Yes, there might have some different if one have renew their membership in April. But if someone still have transactions in march(201703) after his expiration date in march, the scala label code can still dig them out. (I think I should re-run my model later, currently I am using the churn label from 201701 to 201703)</p>\n\n<p>Some part of my scala code is like this:</p>\n\n<pre><code>val targetDIR = \"user_label_201703\"\nval historyCutoff = \"20170228\"\nval targetExpireStart = \"20170301\"\nval targetExpireEnd = \"20170331\"\n</code></pre>\n\n<p>P.S.\nI think train_v2.csv is the churn result for member whos membership expire in 201703, am I right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260382,
          "author_name": "infinitewing",
          "author_url": "",
          "post_date": "12/20/2017 05:14:32",
          "content": "<p>In addition, I think you can estimate the churn rate by calculating the log loss with all 0 submission on public LB. Correct me if I am wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260424,
          "author_name": "",
          "author_url": "",
          "post_date": "12/20/2017 07:01:22",
          "content": "<p>Great point!  </p>\n\n<p>I think you can compute the average provided you use a constant probability other than .5.  I didn't think you could use zero but I see that it is replaced with 10**-15.</p>\n\n<p>This is my first Kaggle competition.  It is interesting to learn all of the techniques used.    </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 260178,
      "author_name": "",
      "author_url": "",
      "post_date": "12/19/2017 20:03:44",
      "content": "<p>Bryan, thanks for your comments.  The administrators suggested that the scala program was a nice-to-have, meaning that it could provide additional information.   I didn't realize it was a prerequisite to generating the training set itself.  </p>\n\n<p>I started to get the sense that the training set was wrong when I did a deep-dive on the user logs.  Many of the identified churners were streaming music past their membership expiration date.  I rationalized that it was some special incentive deal to entice them to stay.  It turns out they were not churners at all.  The user logs were the least important type of features in my model.  Of course, that is surprising because you would expect it to be rich with information because it represents user behavior.  Now I know that it was not as useful because 60% of the churners were invalid.  </p>\n\n<p>In any case, the scala file was fair game.  Like you commented elsewhere, it was difficult to get it to run on a windows-based machine, so I did not pursue further.  My bad....</p>\n\n<p>Congrats on your win again!</p>",
      "votes": null,
      "replies": [
        {
          "id": 260182,
          "author_name": "bryangregory",
          "author_url": "",
          "post_date": "12/19/2017 20:19:29",
          "content": "<p>Yep, exactly.  And just to add to that, the way the contest was structured (by calendar month) made it difficult as well to generate signal from UL and TRX data (particularly UL data).  Of course in a real world  churn model we would be interested in what a user has done recently relative to the day of the month their membership is expiring, structuring the problem by calendar months.  In other words, Mar UL activity data has more meaning for a user expiring in early Apr, than it does for a user expiring in late Apr.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "259854": "Congrats to the winners and competitors.......\n\nInfiniteWing has posted training labels (is_churn) using the scala file.  Thanks InfiniteWing.  Here is what I found, using his 201703 file.\n\nThere are 87,330 churners in the training_v2.csv file.   That is the file that I used to train my model.  Compared with the scala program, which presumably was used to label the submission file, \n\n24,714 (28.3%) are dropped, 28,373 (32.5%) are labeled as non-churners, and 34,343 (39.2%) are labeled as churners.\n\nSo, the scala program and training_v2.csv agree on only 40% of the labeled churners. \n\nIf would be great if someone can confirm this or tell me where my logic is wrong.  I understand that the scala program was public, but you do expect that the provided labels in the training set would be correct, right?",
    "260152": "My stats differ slightly from what you posted, so InfiniteWing and I must have done something slightly different.  If I recall, I found a minor discrepancy in the scala script that was creating further noise (related to date cutoffs like \"MSNO=amZm5p3h8Nw6CxciKVOjcckI5inwCEkWQFcwWfXxvTQ=\" which should not be a churn but the scala script is labelling as a churn) and I attempted to correct them.  I'll post in more detail later in my overview.  \n\nBut your premise is correct that the training sets generated using the scala script differed greatly from the training sets posted for the contest.  Here are some quick stats I show for my scala-generated training sets:  \n\nScala training set for January (training 1):  \n\n - Total churn candidates (total records):  879,478\n - Total churned:  34,771\n - Churn avg:  .039536\n\nScala training set for Februray (training 2):  \n\n - Total churn candidates (total records):  862,158\n - Total churned:  42,681\n - Churn avg:  .049505\n\nAnd presumed churn average for March (test):  ~.036",
    "260174": "For 201703, I don't think we have enough transaction to generate label?",
    "260178": "Bryan, thanks for your comments.  The administrators suggested that the scala program was a nice-to-have, meaning that it could provide additional information.   I didn't realize it was a prerequisite to generating the training set itself.  \n\nI started to get the sense that the training set was wrong when I did a deep-dive on the user logs.  Many of the identified churners were streaming music past their membership expiration date.  I rationalized that it was some special incentive deal to entice them to stay.  It turns out they were not churners at all.  The user logs were the least important type of features in my model.  Of course, that is surprising because you would expect it to be rich with information because it represents user behavior.  Now I know that it was not as useful because 60% of the churners were invalid.  \n\nIn any case, the scala file was fair game.  Like you commented elsewhere, it was difficult to get it to run on a windows-based machine, so I did not pursue further.  My bad....\n\nCongrats on your win again!",
    "260182": "Yep, exactly.  And just to add to that, the way the contest was structured (by calendar month) made it difficult as well to generate signal from UL and TRX data (particularly UL data).  Of course in a real world  churn model we would be interested in what a user has done recently relative to the day of the month their membership is expiring, structuring the problem by calendar months.  In other words, Mar UL activity data has more meaning for a user expiring in early Apr, than it does for a user expiring in late Apr.",
    "260183": "It is presumed based on LB feedback.  See Saigon Apps's post.  My submission feedback closely mirrored what he posted. I.e.-My submissions with means near .035-.037 scored best.  \n\nI controlled the mean using a scale_pos_weight &lt; 1 on the XGBoost portion of my model.  The final submission used scale_pos_weight=.8.  If anyone used XGBoost and wants a quick increase in accuracy (LB improvement), then lower the scale_pos_weight until your submission is hitting a mean in the above range.  You will likely see a very large increase in accuracy.",
    "260253": "Yep, I just did the simple computation of adjusting my predictions down by 60% to a mean of 3.5 churn rate.  My LB score improves from .1245 to .1140, a jump of about 40 positions.  That is remarkable and pretty funny, given how much time I was spending during the last week to eek out changes of .001.  :)   \n\nI understand why making the adjustment helps me, since my models were trained on an invalid training set with a 9% churn rate.  But I am not sure why they would help you if you trained on the proper data set, unless 1) the submission file respondents are meaningfully different than the ones we trained on and / or 2) the scala program, as people understand it, was not used to generate the submission labels.  Both are probably true. \n\nStill, I am surprised that simply adjusting probabilities downwards like that would work, given how nonlinear the loss function is.  If the loss function were closer to linear I can more easily understand how reducing your predicted probabilities to match the target population would work.\n\nCheers again Bryan.",
    "260379": "Hi @Hang, thanks for your reply, I did not consider that. Yes, there might have some different if one have renew their membership in April. But if someone still have transactions in march(201703) after his expiration date in march, the scala label code can still dig them out. (I think I should re-run my model later, currently I am using the churn label from 201701 to 201703)\n\n\nSome part of my scala code is like this:\n\n    val targetDIR = \"user_label_201703\"\n    val historyCutoff = \"20170228\"\n    val targetExpireStart = \"20170301\"\n    val targetExpireEnd = \"20170331\"\n\nP.S.\nI think train_v2.csv is the churn result for member whos membership expire in 201703, am I right?",
    "260382": "In addition, I think you can estimate the churn rate by calculating the log loss with all 0 submission on public LB. Correct me if I am wrong.",
    "260424": "Great point!  \n\nI think you can compute the average provided you use a constant probability other than .5.  I didn't think you could use zero but I see that it is replaced with 10**-15.\n\nThis is my first Kaggle competition.  It is interesting to learn all of the techniques used."
  },
  "source": "meta"
}